VLDB 2026 Research / reviewers in the wild / expert
Sambaran Bandyopadhyay
dblp:83/10310
· DBLP profile ↗
24ranked-venue papers
14as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 13 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deep Submodular Optimization and LLM for Multimodal Content Extraction and Automatic Poster Generation from Long DocumentabstractA poster from a long input document can be considered as a one-page easy-to-read multimodal (text and images) summary presented on a nice template with good design elements. Automatic transformation of a long document into a poster is a very less studied but challenging task. It involves content summarization of the input document followed by template generation and harmonization. In this work, we propose a novel deep submodular function which can be trained on ground truth summaries to extract multimodal content from the document and explicitly ensures good coverage, diversity and alignment of text and images. Then, we use an LLM based paraphraser and propose to generate a template with various design aspects conditioned on the input content. We show the merits of our approach through extensive automated and human evaluations. Vijay Jaisankar, Sambaran Bandyopadhyay, Kalp Vyas, Varre Chaitanya, Shwetha Somasundaram |
AAAI | 2 |
| 2025 | Language Models of Code Are Few-Shot Planners and Reasoners for Multi-Document Summarization with AttributionabstractDocument summarization has greatly benefited from advances in large language models (LLMs). In real-world situations, summaries often need to be generated from multiple documents with diverse sources and authors, lacking a clear information flow. Naively concatenating these documents and generating a summary can lead to poorly structured narratives and redundancy. Additionally, attributing each part of the generated summary to a specific source is crucial for reliability. In this study, we address multi-document summarization with attribution using our proposed solution ***MiDAS-PRo***, consisting of three stages: (i) Planning the hierarchical organization of source documents, (ii) Reasoning by generating relevant entities/topics, and (iii) Summary Generation. We treat the first two sub-problems as a code completion task for LLMs. By incorporating well-selected in-context learning examples through a graph attention network, LLMs effectively generate plans and reason topics for a document collection. Experiments on summarizing scientific articles from public datasets show that our approach outperforms state-of-the-art baselines in both automated and human evaluations. Abhilash Nandy, Sambaran Bandyopadhyay |
AAAI | 2 |
| 2025 | Infogen: Generating Complex Statistical Infographics from DocumentsabstractAkash Ghosh, Aparna Garimella, Pritika Ramu, Sambaran Bandyopadhyay, Sriparna Saha. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Akash Ghosh, Aparna Garimella, Pritika Ramu, Sambaran Bandyopadhyay, Sriparna Saha 0001 |
ACL (1) | 4 |
| 2024 | Presentations by the Humans and For the Humans: Harnessing LLMs for Generating Persona-Aware Slides from DocumentsabstractIshani Mondal, Shwetha S, Anandhavelu Natarajan, Aparna Garimella, Sambaran Bandyopadhyay, Jordan Boyd-Graber. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ishani Mondal, Shwetha S, Anandhavelu Natarajan, Aparna Garimella, Sambaran Bandyopadhyay, Jordan L. Boyd-Graber |
EACL (1) | 5 |
| 2024 | Is This a Bad Table? A Closer Look at the Evaluation of Table Generation from TextabstractUnderstanding whether a generated table is of good quality is important to be able to use it in creating or editing documents using automatic methods.In this work, we underline that existing measures for table quality evaluation fail to capture the overall semantics of the tables, and sometimes unfairly penalize good tables and reward bad ones.We propose TABEVAL, a novel table evaluation strategy that captures table semantics by first breaking down a table into a list of natural language atomic statements and then compares them with ground truth statements using entailment-based measures.To validate our approach, we curate a dataset comprising of text descriptions for 1,250 diverse Wikipedia tables, covering a range of topics and structures, in contrast to the limited scope of existing datasets.We compare TABEVAL with existing metrics using unsupervised and supervised textto-table generation methods, demonstrating its stronger correlation with human judgments of table quality across four datasets. Pritika Ramu, Aparna Garimella, Sambaran Bandyopadhyay |
EMNLP | 3 |
| 2024 | Enhancing Presentation Slide Generation by LLMs with a Multi-Staged End-to-End ApproachabstractGenerating presentation slides from a long document with multimodal elements such as text and images is an important task.This is time consuming and needs domain expertise if done manually.Existing approaches for generating a rich presentation from a document are often semi-automatic or only put a flat summary into the slides ignoring the importance of a good narrative.In this paper, we address this research gap by proposing a multi-staged end-to-end model which uses a combination of LLM and VLM.We have experimentally shown that compared to applying LLMs directly with state-ofthe-art prompting, our proposed multi-staged solution is better in terms of automated metrics and human evaluation. Sambaran Bandyopadhyay, Himanshu Maheshwari, Anandhavelu Natarajan, Apoorv Saxena |
INLG | 1 |
| 2022 | Monolith to Microservices: Representing Application Software through Heterogeneous Graph Neural NetworkabstractMonolithic software encapsulates all functional capabilities into a single deployable unit. But managing it becomes harder as the demand for new functionalities grow. Microservice architecture is seen as an alternative as it advocates building an application through a set of loosely coupled small services wherein each service owns a single functional responsibility. But the challenges associated with the separation of functional modules, slows down the migration of a monolithic code into microservices. In this work, we propose a representation learning based solution to tackle this problem. We use a heterogeneous graph to jointly represent software artifacts (like programs and resources) and the different relationships they share (function calls, inheritance, etc.), and perform a constraint-based clustering through a novel heterogeneous graph neural network. Experimental studies show that our approach is effective on monoliths of different types. Alex Mathai, Sambaran Bandyopadhyay, Utkarsh Desai, Srikanth Tamilselvam |
IJCAI | 2 |
| 2022 | Dynamic Structure Learning through Graph Neural Network for Forecasting Soil Moisture in Precision AgricultureabstractSoil moisture is an important component of precision agriculture as it directly impacts the growth and quality of vegetation. Forecasting soil moisture is essential to schedule the irrigation and optimize the use of water. Physics based soil moisture models need rich features and heavy computation which is not scalable. In recent literature, conventional machine learning models have been applied for this problem. These models are fast and simple, but they often fail to capture the spatio-temporal correlation that soil moisture exhibits over a region. In this work, we propose a novel graph neural network based solution that learns temporal graph structures and forecast soil moisture in an end-to-end framework. Our solution is able to handle the problem of missing ground truth soil moisture which is common in practice. We show the merit of our algorithm on real-world soil moisture data. Anoushka Vyas, Sambaran Bandyopadhyay |
IJCAI | 2 |
| 2021 | Graph Neural Network to Dilute Outliers for Refactoring Monolith ApplicationabstractMicroservices are becoming the defacto design choice for software architecture. It involves partitioning the software components into finer modules such that the development can happen independently. It also provides natural benefits when deployed on the cloud since resources can be allocated dynamically to necessary components based on demand. Therefore, enterprises as part of their journey to cloud, are increasingly looking to refactor their monolith application into one or more candidate microservices; wherein each service contains a group of software entities (e.g., classes) that are responsible for a common functionality. Graphs are a natural choice to represent a software system. Each software entity can be represented as nodes and its dependencies with other entities as links. Therefore, this problem of refactoring can be viewed as a graph based clustering task. In this work, we propose a novel method to adapt the recent advancements in graph neural networks in the context of code to better understand the software and apply them in the clustering task. In that process, we also identify the outliers in the graph which can be directly mapped to top refactor candidates in the software. Our solution is able to improve state-of-the-art performance compared to works from both software engineering and existing graph representation based techniques. Utkarsh Desai, Sambaran Bandyopadhyay, Srikanth Tamilselvam |
AAAI | 2 |
| 2021 | Data Quality for Machine Learning TasksabstractThe quality of training data has a huge impact on the efficiency, accuracy and complexity of machine learning tasks. Data remains susceptible to errors or irregularities that may be introduced during collection, aggregation or annotation stage. This necessitates profiling and assessment of data to understand its suitability for machine learning tasks and failure to do so can result in inaccurate analytics and unreliable decisions. While researchers and practitioners have focused on improving the quality of models, there are limited efforts towards improving the data quality. Nitin Gupta 0005, Shashank Mujumdar, Hima Patel, Satoshi Masuda, Naveen Panwar, Sambaran Bandyopadhyay, Sameep Mehta, Shanmukha C. Guttula, Shazia Afzal, Ruhi Sharma Mittal, Vitobha Munigala |
KDD | 6 |
| 2021 | A Deep Hybrid Pooling Architecture for Graph Classification with Hierarchical Attention
Sambaran Bandyopadhyay, Manasvi Aggarwal, M. Narasimha Murty |
PAKDD (1) | 1 |
| 2021 | Unsupervised constrained community detection via self-expressive graph neural networkabstractGraph neural networks (GNNs) are able to achieve promising performance on multiple graph downstream tasks such as node classification and link prediction. Comparatively lesser work has been done to design GNNs which can operate directly for community detection on graphs. Traditionally, GNNs are trained on a semi-supervised or self-supervised loss function and then clustering algorithms are applied to detect communities. However, such decoupled approaches are inherently sub-optimal. Designing an unsupervised loss function to train a GNN and extract communities in an integrated manner is a fundamental challenge. To tackle this problem, we combine the principle of self-expressiveness with the framework of self-supervised graph neural network for unsupervised community detection for the first time in literature. Our solution is trained in an end-to-end fashion and achieves state-of-the-art community detection performance on multiple publicly available datasets. Sambaran Bandyopadhyay, Vishal Peter |
UAI | 1 |
| 2020 | Self-supervised Hierarchical Graph Neural Network for Graph RepresentationabstractGraph neural networks (GNNs) gain significant interest in the domain of network representation learning. To obtain a graph level vector representation from individual node embeddings, hierarchical pooling algorithms are proposed in the recent literature which adhere the hierarchical structure of an input graph. A major limitation for most of the existing supervised GNNs is their dependency on large number of graph labels (often 80%-90%) to train the parameters of the neural architecture. But obtaining labels of a large number of graphs is expensive for real world applications. So in this work, we propose an unsupervised hierarchical neural network, referred as GraPHmax, for obtaining graph level representation. We propose the concept of periphery representation and show its effectiveness to obtain discriminative features of an input graph. Further, inspired by the concepts from self-supervised learning, we propose to maximize periphery and hierarchical information in the context of hierarchical GNN. Thorough experimentation on both synthetic and real-world graph datasets shows that GraPHmax is not only able to outperform unsupervised graph embedding techniques, it often achieves state-of-the-art performance even with respect to a set of popular supervised GNN algorithms. Sambaran Bandyopadhyay, Manasvi Aggarwal, M. Narasimha Murty |
IEEE BigData | 1 |
| 2020 | Hypergraph Attention Isomorphism Network by Learning Line Graph ExpansionabstractGraph neural networks (GNNs) are able to achieve state-of-the-art performance for node representation and classification in a network. But, most of the existing GNNs can be applied to simple graphs, where an edge connects only a pair of nodes. Studies have shown that hypergraphs are effective to model real-world relationships which are of higher order in nature. Recently, graph neural networks are proposed for hypergraphs, but they implicitly use clique or star expansions to convert the hypergraph to a simple graph, or use computationally expensive hypergraph Laplacian.In this work, we propose a novel hypergraph neural network for semi-supervised hypernode classification, which operates directly on the hypergraphs with varying hyperedge sizes. Within each layer, it indirectly works on the line graph of the given hypergraph, without actually forming the line graph explicitly. Moreover, it also employs a self-attention mechanism to learn the weights of those edge relationships. Experimentally, HAIN is able to improve the state-of-the-art hypernode classification performance on all the datasets we use. We make the source code available to ease the reproducibility of the results. Sambaran Bandyopadhyay, Kishalay Das, M. Narasimha Murty |
IEEE BigData | 1 |
| 2020 | A Multilayered Informative Random Walk for Attributed Social Network EmbeddingabstractNetwork representation learning (also known as Graph embedding) is a technique to map the nodes of a network to a lower dimensional vector space.Random walk based representation techniques are found to be efficient as they can easily preserve different orders of proximities between the nodes in the embedding space.Most of the social networks now-a-days have some content (or attributes) associated with each node.These attributes can provide complementary information along with the link structure of the network.But in a real life network, the information carried by the link structure and that by the attributes vary significantly over the nodes.Most of the existing unsupervised attributed network embedding algorithms do not distinguish between the link structure and the attributes of a node depending on their informativeness.In this work, we propose an unsupervised node embedding technique that exploits both the structure and attributes by intelligently prioritizing one of them, in the random walk, for each node separately.We convert the network into a multi-layered graph and propose a novel random walk based on the informativeness of a node in different layers.This unified approach is simple and computationally fast, yet able to use the content as a complement to structure and viceversa.Experimental evaluations on four real world publicly available datasets show the merit of our approach (up to 168.75% improvement) compared to the state-of-the-art algorithms in the domain.We make the source code available to download. Sambaran Bandyopadhyay, Anirban Biswas, Harsh Kara, M. N. Murty |
ECAI | 1 |
| 2020 | Integrating Network Embedding and Community Outlier Detection via Multiclass Graph DescriptionabstractNetwork (or graph) embedding is the task to map the nodes of a graph to a lower dimensional vector space, such that it preserves the graph properties and facilitates the downstream network mining tasks.Real world networks often come with (community) outlier nodes, which behave differently from the regular nodes of the community.These outlier nodes can affect the embedding of the regular nodes, if not handled carefully.In this paper, we propose a novel unsupervised graph embedding approach (called DMGD) which integrates outlier and community detection with node embedding.We extend the idea of deep support vector data description to the framework of graph embedding when there are multiple communities present in the given network, and an outlier is characterized relative to its community.We also show the theoretical bounds on the number of outliers detected by DMGD.Our formulation boils down to an interesting minimax game between the outliers, community assignments and the node embedding function.We also propose an efficient algorithm to solve this optimization framework.Experimental results on both synthetic and real world networks show the merit of our approach compared to state-of-the-arts. Sambaran Bandyopadhyay, Saley Vishal Vivek, M. Narasimha Murty |
ECAI | 1 |
| 2020 | Outlier Resistant Unsupervised Deep Architectures for Attributed Network EmbeddingabstractAttributed network embedding is the task to learn a lower dimensional vector representation of the nodes of an attributed network, which can be used further for downstream network mining tasks. Nodes in a network exhibit community structure and most of the network embedding algorithms work well when the nodes, along with their attributes, adhere to the community structure of the network. But real life networks come with community outlier nodes, which deviate significantly in terms of their link structure or attribute similarities from the other nodes of the community they belong to. These outlier nodes, if not processed carefully, can even affect the embeddings of the other nodes in the network. Thus, a node embedding framework for dealing with both the link structure and attributes in the presence of outliers in an unsupervised setting is practically important. In this work, we propose a deep unsupervised autoencoders based solution which minimizes the effect of outlier nodes while generating the network embedding. We use both stochastic gradient descent and closed form updates for faster optimization of the network parameters. We further explore the role of adversarial learning for this task, and propose a second unsupervised deep model which learns by discriminating the structure and the attribute based embeddings of the network and minimizes the effect of outliers in a coupled way. Our experiments show the merit of these deep models to detect outliers and also the superiority of the generated network embeddings for different downstream mining tasks. To the best of our knowledge, these are the first unsupervised non linear approaches that reduce the effect of the outlier nodes while generating Network Embedding. Sambaran Bandyopadhyay, Lokesh N, Saley Vishal Vivek, M. Narasimha Murty |
WSDM | 1 |
| 2019 | Outlier Aware Network Embedding for Attributed NetworksabstractAttributed network embedding has received much interest from the research community as most of the networks come with some content in each node, which is also known as node attributes. Existing attributed network approaches work well when the network is consistent in structure and attributes, and nodes behave as expected. But real world networks often have anomalous nodes. Typically these outliers, being relatively unexplainable, affect the embeddings of other nodes in the network. Thus all the downstream network mining tasks fail miserably in the presence of such outliers. Hence an integrated approach to detect anomalies and reduce their overall effect on the network embedding is required.Towards this end, we propose an unsupervised outlier aware network embedding algorithm (ONE) for attributed networks, which minimizes the effect of the outlier nodes, and hence generates robust network embeddings. We align and jointly optimize the loss functions coming from structure and attributes of the network. To the best of our knowledge, this is the first generic network embedding approach which incorporates the effect of outliers for an attributed network without any supervision. We experimented on publicly available real networks and manually planted different types of outliers to check the performance of the proposed algorithm. Results demonstrate the superiority of our approach to detect the network outliers compared to the state-of-the-art approaches. We also consider different downstream machine learning applications on networks to show the efficiency of ONE as a generic network embedding technique. The source code is made available at https://github.com/sambaranban/ONE. Sambaran Bandyopadhyay, Lokesh Nagalapatti, M. Narasimha Murty |
AAAI | 1 |
| 2018 | DivGroup: A Diversified Approach to Divide Collection of Patterns into Uniform GroupsabstractSimilarity based grouping of patterns has been explored profusely under the well celebrated clustering paradigm in pattern recognition and machine learning. In clustering, objects in the same cluster are similar to each other and objects belonging to different clusters are dissimilar in a corresponding sense. However, it is not rare to come across situations where instead of a similarity based grouping, forming groups of diverse objects is needed. Resource allocation across different parts of an organization, performing cross-validation splits of dataset with class imbalance, heterogeneous or mixed ability partitioning of students, etc. are the applications of grouping which require each group to contain diverse set of patterns. Moreover, these applications also demand different groups to be similar to each other in some sense. In this work, we propose a generic framework for partitioning a collection of patterns into a set of groups such that the above two criteria are fulfilled. To the best of our knowledge, this is the first work to propose such a framework irrespective of any particular application. Towards this end, it turns out that finding an optimal solution to the problem that we developed is NP Hard. So we Propose an approximate solution for the same. We conduct experiments on both synthetic and real world datasets to evaluate the performance of the proposed algorithm. We show the merit of the algorithm by comparing the results with some related state-of-the-art baseline methods. Sambaran Bandyopadhyay, Sharad Nandanwar, Rishabh Deshmukh, M. Narasimha Murty |
ICPR | 1 |
| 2018 | A Generic Axiomatic Characterization for Measuring Influence in Social NetworksabstractMeasuring influence, through centrality measures, has been a center-piece of research in the analysis of complex social networks, such as finding coherent communities (clusters) and locating trend setters (prototypes) in viral marketing. Even though there exists a few axiomatic frameworks associated with some specific forms of influence measures in the literature, these formal frameworks are not generic in nature in terms of characterizing the space of influence measures for complex social networks. To address this research gap, we propose a generic axiomatic framework, in this paper, to capture most of the key intrinsic properties of any influence measure in networks. We further analyze certain popular centrality measures using this framework. Interestingly, our analysis reveals that none of the centrality measures considered satisfies all the desirable axioms. We finally conclude this paper by stating an appealing conjecture on a potential impossibility theorem associated with the proposed axiomatic framework. Sambaran Bandyopadhyay, Ramasuri Narayanam, M. Narasimha Murty |
ICPR | 1 |
| 2016 | An Axiomatic Framework for Ex-Ante Dynamic Pricing Mechanisms in Smart GridabstractIn electricity markets, the choice of the right pricing regime is crucial for the utilities because the price they charge to their consumers, in anticipation of their demand in real-time, is a key determinant of their profits and ultimately their survival in competitive energy markets. Among the existing pricing regimes, in this paper, we consider ex-ante dynamic pricing schemes as (i) they help to address the peak demand problem (a crucial problem in smart grids), and (ii) they are transparent and fair to consumers as the cost of electricity can be calculated before the actual consumption. In particular, we propose an axiomatic framework that establishes the conceptual underpinnings of the class of ex-ante dynamic pricing schemes. We first propose five key axioms that reflect the criteria that are vital for energy utilities and their relationship with consumers. We then prove an impossibility theorem to show that there is no pricing regime that satisfies all the five axioms simultaneously. We also study multiple cost functions arising from various pricing regimes to examine the subset of axioms that they satisfy. We believe that our proposed framework in this paper is first of its kind to evaluate the class of ex-ante dynamic pricing schemes in a manner that can be operationalised by energy utilities. Sambaran Bandyopadhyay, Ramasuri Narayanam, Sarvapali D. Ramchurn, Vijay Arya, Iskandarbin Petra |
AAAI | 1 |
| 2016 | Axioms to characterize efficient incremental clusteringabstractAlthough clustering is one of the central tasks in machine learning for the last few decades, analysis of clustering irrespective of any particular algorithm was not undertaken for a long time. In the recent literature, axiomatic frameworks have been proposed for clustering and its quality. But none of the proposed frameworks has concentrated on the computational aspects of clustering, which is essential in current big data analytics. In this paper, we propose an axiomatic framework for clustering which considers both the quality and the computational complexity of clustering algorithms. The axioms proposed by us necessarily associate the problem of clustering with the important concept of incremental learning and divide and conquer learning. We also propose an order independent incremental clustering algorithm which satisfies all of these axioms in some constrained manner. Sambaran Bandyopadhyay, M. Narasimha Murty |
ICPR | 1 |
| 2015 | Aggregate Demand-Based Real-Time Pricing Mechanism for the Smart Grid: A Game-Theoretic Analysis
Sambaran Bandyopadhyay, Ramasuri Narayanam, Ramachandra Kota, Mohamad Iskandar Petra, Zainul Charbiwala |
IJCAI | 1 |
| 2015 | Voltage Correlations in Smart Meter DataabstractThe connectivity model of a power distribution network can easily become outdated due to system changes occurring in the field. Maintaining and sustaining an accurate connectivity model is a key challenge for distribution utilities worldwide. This work shows that voltage time series measurements collected from customer smart meters exhibit correlations that are consistent with the hierarchical structure of the distribution network. These correlations may be leveraged to cluster customers based on common ancestry and help verify and correct an existing connectivity model. Additionally, customers may be clustered in combination with voltage data from circuit metering points, spatial data from the geographical information system, and any existing but partially accurate connectivity model to infer customer to transformer and phase connectivity relationships with high accuracy. Rajendu Mitra, Ramachandra Kota, Sambaran Bandyopadhyay, Vijay Arya, Brian Sullivan, Richard Mueller, Heather Storey, Gerard Labut |
KDD | 3 |