EDBT 2026 Demo / reviewers in the wild / expert
Avirup Saha
dblp:228/2505
· DBLP profile ↗
13ranked-venue papers
5as first author
7since 2021 · last 2024
0000-0002-5014-3582ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Deep learning architectures and training · 52% Learning paradigms · 26% Language models and text generation · 22% | |
| Databases, data mining, and information retrieval
3 papers |
Data mining · 71% Recommender systems · 21% Data models and query languages · 8% | |
| Computer networks
1 paper |
Network performance modeling · 38% Network measurement and analytics · 38% Content delivery and video streaming · 23% | |
| Theoretical computer science
2 papers |
Algorithmic game theory and mechanism design · 54% Graph algorithms and graph theory · 46% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Services computing and microservices · 100% |
Topics — the 18 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
foundation model |
0.8 | 1 | 2024 | AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability Data · AAAI 2024 |
Natural language and speech › Language models and text generation › agentic language model › tool-augmented language models
function calling |
0.8 | 1 | 2024 | Sequential API Function Calling Using GraphQL Schema · EMNLP 2024 |
Machine learning › Deep learning architectures and training › foundation model
time series foundation model |
0.8 | 1 | 2024 | AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability Data · AAAI 2024 |
Data mining › time series analysis › time series forecasting
multivariate time series forecasting |
0.8 | 1 | 2024 | AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability Data · AAAI 2024 |
Data mining › time series analysis
time series forecasting |
0.8 | 1 | 2024 | AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability Data · AAAI 2024 |
Cloud and datacenter computing
cloud federation |
0.5 | 1 | 2021 | Quality and Profit Assured Trusted Cloud Federation Formation: Game Theory Based Approach · IEEE Trans. Serv. Comput. 2021 |
Algorithmic game theory and mechanism design
coalitional game |
0.5 | 1 | 2021 | Quality and Profit Assured Trusted Cloud Federation Formation: Game Theory Based Approach · IEEE Trans. Serv. Comput. 2021 |
Machine learning › Learning paradigms › semi-supervised learning
graph-based semi-supervised learning |
0.4 | 1 | 2020 | Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block Model · IJCAI 2020 |
Machine learning › Learning paradigms › semi-supervised learning › graph-based semi-supervised learning
label propagation |
0.4 | 1 | 2020 | Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block Model · IJCAI 2020 |
Data mining
multi-view learning |
0.4 | 1 | 2020 | MVL: Multi-View Learning for News Recommendation · SIGIR 2020 |
Recommender systems
news recommendation |
0.4 | 1 | 2020 | MVL: Multi-View Learning for News Recommendation · SIGIR 2020 |
Graph algorithms and graph theory › graph clustering › community detection
stochastic block model |
0.4 | 1 | 2020 | Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block Model · IJCAI 2020 |
Network performance modeling
traffic modeling |
0.4 | 1 | 2019 | Learning Network Traffic Dynamics Using Temporal Point Process · INFOCOM 2019 |
Network measurement and analytics
traffic prediction |
0.4 | 1 | 2019 | Learning Network Traffic Dynamics Using Temporal Point Process · INFOCOM 2019 |
Cloud and datacenter computing
quality of service |
0.1 | 1 | 2021 | Quality and Profit Assured Trusted Cloud Federation Formation: Game Theory Based Approach · IEEE Trans. Serv. Comput. 2021 |
Cloud and datacenter computing
resource management |
0.1 | 1 | 2021 | Quality and Profit Assured Trusted Cloud Federation Formation: Game Theory Based Approach · IEEE Trans. Serv. Comput. 2021 |
Content delivery and video streaming › caching
cache miss reduction |
0.1 | 1 | 2019 | Learning Network Traffic Dynamics Using Temporal Point Process · INFOCOM 2019 |
Content delivery and video streaming
caching |
0.1 | 1 | 2019 | Learning Network Traffic Dynamics Using Temporal Point Process · INFOCOM 2019 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.3autoencoder · 1.5TSMixer · 1.5simulation · 1.0nash stability analysis · 1.0hedonic coalitional game · 1.0news encoder · 0.4graph neural network · 0.4attention mechanism · 0.4temporal point process · 0.4deep learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability DataabstractThe efficiency of business processes relies on business key performance indicators (Biz-KPIs), that can be negatively impacted by IT failures. Business and IT Observability (BizITObs) data fuses both Biz-KPIs and IT event channels together as multivariate time series data. Forecasting Biz-KPIs in advance can enhance efficiency and revenue through proactive corrective measures. However, BizITObs data generally exhibit both useful and noisy inter-channel interactions between Biz-KPIs and IT events that need to be effectively decoupled. This leads to suboptimal forecasting performance when existing multivariate forecasting models are employed. To address this, we introduce AutoMixer, a time-series Foundation Model (FM) approach, grounded on the novel technique of channel-compressed pretrain and finetune workflows. AutoMixer leverages an AutoEncoder for channel-compressed pretraining and integrates it with the advanced TSMixer model for multivariate time series forecasting. This fusion greatly enhances the potency of TSMixer for accurate forecasts and also generalizes well across several downstream tasks. Through detailed experiments and dashboard analytics, we show AutoMixer's capability to consistently improve the Biz-KPI's forecasting accuracy (by 11-15%) which directly translates to actionable business insights. Santosh Palaskar, Vijay Ekambaram, Arindam Jati, Neelamadhav Gantayat, Avirup Saha, Seema Nagar, Nam H. Nguyen, Pankaj Dayama 0001, Renuka Sindhgatta, Prateeti Mohapatra, Jayant Kalagnanam, Nandyala Hemachandra, Narayan Rangaraj |
AAAI | 5 |
| 2024 | Sequential API Function Calling Using GraphQL SchemaabstractAvirup Saha, Lakshmi Mandal, Balaji Ganesan, Sambit Ghosh, Renuka Sindhgatta, Carlos Eberhardt, Dan Debrunner, Sameep Mehta. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Avirup Saha, Lakshmi Mandal, Balaji Ganesan, Sambit Ghosh, Renuka Sindhgatta, Carlos Eberhardt, Dan Debrunner, Sameep Mehta |
EMNLP | 1 |
| 2022 | Proactive Fault-Tolerance Technique to Enhance Reliability of Cloud Service in Cloud Federation EnvironmentabstractCloud federation is a new computing paradigm that has paved the way for cloud service providers (CSPs) to offer their unused resources (virtual machine) to other CSPs when their resource demands are low. Federation also allows CSPs to outsource their resource requests to other CSPs when their computing resources’ demands are high. Thus, in cloud federation environment reliability and availability of services offered by service providers increase as the CSPs are able to share their resources among themselves. Moreover, to maintain the reliability and availability of cloud services offered through federation, it is important that the computational environment of member CSPs within the federation is fault tolerant. Therefore, there is a need for fault tolerant system to guarantee cloud service reliability and availability in cloud federation environment. In this article, we propose a proactive fault tolerance system that preempts faults within the federation on the basis of CPU temperature. The fault tolerance system within the federation is modeled as a multi-objective optimization problem of maximizing profit and minimizing migration cost while redistributing resources (virtual machine) from faulty CSPs to non-faulty CSPs within the federation. To address this issue, we have also proposed an algorithm called Preference Based Fault Management (PBFM) to manage the federation in the event of faults. We perform extensive experiments to evaluate the effectiveness of our proposed mechanism and compare it with two other mechanisms MCAFM (Migration Cost Assured Fault Management) and PAFM (Profit Assured Fault Management). Results show that our proposed mechanism PBFM yields an optimized solution to the general problem of profit and migration cost trade-off in presence of faulty CSPs. Benay Kumar Ray, Avirup Saha, Sunirmal Khatua, Sarbani Roy |
IEEE Trans. Cloud Comput. | 2 |
| 2021 | tWT-WT: A Dataset to Assert the Role of Target Entities for Detecting Stance of TweetsabstractThe stance detection task aims at detecting the stance of a tweet or a text for a target.These targets can be named entities or free-form sentences (claims).Though the task involves reasoning of the tweet with respect to a target, we find that it is possible to achieve high accuracy on several publicly available Twitter stance detection datasets without looking at the target sentence.Specifically, a simple tweet classification model achieved human-level performance on the WT-WT dataset and more than two-third accuracy on various other datasets.We investigate the existence of biases in such datasets to find the potential spurious correlations of sentiment-stance relations and lexical choice associated with the stance category.Furthermore, we propose a new large dataset free of such biases and demonstrate its aptness on the existing stance detection systems.Our empirical findings show much scope for research on the stance detection task and proposes several considerations for creating future stance detection datasets.1 Ayush Kaushal, Avirup Saha, Niloy Ganguly |
NAACL-HLT | 2 |
| 2021 | Graph-based semi-supervised learning through the lens of safetyabstractGraph-based semi-supervised learning (G-SSL) algorithms have witnessed rapid development and widespread usage across a variety of applications in recent years. However, the theoretical characterisation of the efficacy of such algorithms has remained an under-explored area. We introduce a novel algorithm for G-SSL, CSX, whose objective function extends those of Label Propagation and Expander, two popular G-SSL algorithms. We provide data-dependent generalisation error bounds for all three aforementioned algorithms when they are applied to graphs drawn from a partially labelled extension of a versatile latent space graph generative model. The bounds we obtain enable us to characterise the predictive performance as measured by accuracy in terms of homophily and label quantity. Building on this we develop a key notion of GLM-safety which enables us to compare G-SSL algorithms on the basis of the range of graphs on which they obtain a guaranteed accuracy. We show that the proposed algorithm CSX has a better GLM-safety profile than Label Propagation and Expander while achieving comparable or better accuracy on synthetic as well as real-world benchmark networks. Shreyas Sheshadri, Avirup Saha, Priyank Patel, Samik Datta, Niloy Ganguly |
UAI | 2 |
| 2021 | Demarcating Endogenous and Exogenous Opinion Dynamics: An Experimental Design ApproachabstractThe networked opinion diffusion in online social networks is often governed by the two genres of opinions— endogenous opinions that are driven by the influence of social contacts among users, and exogenous opinions which are formed by external effects like news and feeds. Accurate demarcation of endogenous and exogenous messages offers an important cue to opinion modeling, thereby enhancing its predictive performance. In this article, we design a suite of unsupervised classification methods based on experimental design approaches, in which, we aim to select the subsets of events which minimize different measures of mean estimation error. In more detail, we first show that these subset selection tasks are NP-Hard. Then we show that the associated objective functions are weakly submodular, which allows us to cast efficient approximation algorithms with guarantees. Finally, we validate the efficacy of our proposal on various real-world datasets crawled from Twitter as well as diverse synthetic datasets. Our experiments range from validating prediction performance on unsanitized and sanitized events to checking the effect of selecting optimal subsets of various sizes. Through various experiments, we have found that our method offers a significant improvement in accuracy in terms of opinion forecasting, against several competitors. Paramita Koley, Avirup Saha, Sourangshu Bhattacharya, Niloy Ganguly, Abir De |
ACM Trans. Knowl. Discov. Data | 2 |
| 2021 | Quality and Profit Assured Trusted Cloud Federation Formation: Game Theory Based ApproachabstractWith more awareness and growth in the cloud market, demands for computational resources have increased in order to provide services to the cloud users. Sometimes it is difficult for an individual cloud service provider (CSP) to meet the level of promised quality of service (QoS) and to fulfill all types of resource requests dynamically. Cloud federation has become a consolidated paradigm in which group of cooperative CSPs share their unused resources with peers to gain some economic benefit. Hence, the cloud federation overcomes the limitation of each CSP for maintaining QoS during sudden spikes in resource demand. However, the presence of untrusted CSPs degrades the QoS of the services delivered through federation. Trusted CSPs are highly reputed in the federation as they can extend their resources and services to maintain the level of committed QoS by the member CSPs of the federation. Therefore, to guarantee delivery of committed QoS, it will be necessary to form a federation with trusted CSPs only. In this paper, we present a broker based cloud federation architecture. The cloud federation formation is modeled as a hedonic coalitional game. The main objective of this work is to find the most suitable and stable federation of trusted CSPs that will maximize the satisfaction level of each individual CSP on the basis of QoS and profit. The proposed coalitional game inspired cloud federation formation (CGCFF) algorithm has been extensively compared with selected existing techniques. Simulation results show that the set of federation formed by CGCFF is Nash-stable and performs better than these techniques in terms of satisfaction, quality and profit. Benay Kumar Ray, Avirup Saha, Sunirmal Khatua, Sarbani Roy |
IEEE Trans. Serv. Comput. | 2 |
| 2020 | A GAN-based Framework for Modeling Hashtag Popularity Dynamics Using Assistive InformationabstractTemporal point process (TPP) models have hitherto been moderately good at nowcasting hashtag popularity, but have been very poor at forecasting due to insufficient modeling of Twitter microdynamics. Recent studies have shown that the highly fluctuating nature of hashtag popularity dynamics is due to the influence of two external factors: (i) hashtag-tweet reinforcement and (ii) inter-hashtag competition. In this paper, we propose a marked TPP based on Generative Adversarial Networks (GANs) which can seamlessly incorporate the assistive information necessary to capture the above effects and successfully forecast distant popularity trends. To achieve this, we employ a unique linear semi-autoregressive model for mark generation and couple the time and mark generative aspects. On seven diverse datasets crawled from Twitter covering several real-world events, our model yields remarkably stable performance in predicting hashtag popularity in diverse situations and offers a substantial improvement over the existing state of the art generative models. Avirup Saha, Niloy Ganguly |
CIKM | 1 |
| 2020 | Understanding the Success of Graph-based Semi-Supervised Learning using Partially Labelled Stochastic Block ModelabstractWith the proliferation of learning scenarios with an abundance of instances, but limited amount of high-quality labels, semi-supervised learning algorithms came to prominence. Graph-based semi-supervised learning (G-SSL) algorithms, of which Label Propagation (LP) is a prominent example, are particularly well-suited for these problems. The premise of LP is the existence of homophily in the graph, but beyond that nothing is known about the efficacy of LP. In particular, there is no characterisation that connects the structural constraints, volume and quality of the labels to the accuracy of LP. In this work, we draw upon the notion of recovery from the literature on community detection, and provide guarantees on accuracy for partially-labelled graphs generated from the Partially-Labelled Stochastic Block Model (PLSBM). Extensive experiments performed on synthetic data verify the theoretical findings. Avirup Saha, Shreyas Sheshadri, Samik Datta, Niloy Ganguly, Disha Makhija, Priyank Patel |
IJCAI | 1 |
| 2020 | MVL: Multi-View Learning for News RecommendationabstractIn this paper, we propose a Multi-View Learning (MVL) framework for news recommendation which uses both the content view and the user-news interaction graph view. In the content view, we use a news encoder to learn news representations from different information like titles, bodies and categories. We obtain representation of user from his/her browsed news conditioned on the candidate news article to be recommended. In the graph-view, we propose to use a graph neural network to capture the user-news, user-user and news-news relatedness in the user-news bipartite graphs by modeling the interactions between different users and news. In addition, we propose to incorporate attention mechanism into the graph neural network to model the importance of these interactions for more informative representation learning of user and news. Experiments on a real world dataset validate the effectiveness of MVL. T. Y. S. S. Santosh, Avirup Saha, Niloy Ganguly |
SIGIR | 2 |
| 2019 | Learning Network Traffic Dynamics Using Temporal Point ProcessabstractAccurate modeling of network traffic has a wide variety of applications. In this paper, we propose Network Transmission Point Process (NTPP), a probabilistic deep machinery that models the traffic characteristics of hosts on a network and effectively forecasts the network traffic patterns, such as load spikes. Existing stochastic models relied on the network traffic being self-similar in nature, thus failing to account for traffic anomalies. These anomalies, such as short-term traffic bursts, are very prevalent in certain modern-day traffic conditions, e.g. datacenter traffic, thus refuting the assumption of self-similarity. Our model is robust to such anomalies since it effectively leverages the self-exciting nature of the bursty network traffic using a temporal point process model.On seven diverse datasets collected from the fields of cyberdefense exercises (CDX), website access logs, datacenter traffic, and P2P traffic, NTPP offers a substantial performance boost in predicting network traffic characteristics against several baselines, ranging from forecasting the network traffic volume to detecting traffic spikes. We also demonstrate an application of our model to a caching scenario, showing that it can be used to effectively lower the cache miss rate. Avirup Saha, Niloy Ganguly, Sandip Chakraborty 0001, Abir De |
INFOCOM | 1 |
| 2019 | Toward maximization of profit and quality of cloud federation: solution to cloud federation formation problem
Benay Kumar Ray, Avirup Saha, Sunirmal Khatua, Sarbani Roy |
J. Supercomput. | 2 |
| 2018 | CRPP: Competing Recurrent Point Process for Modeling Visibility Dynamics in Information DiffusionabstractAccurate modeling of how the visibility of a piece of information varies across time has a wide variety of applications. For example, in an e-commerce site like Amazon, it can help to identify which product is preferred over others; in Twitter, it can predict which hashtag may go viral against others. Visibility of a piece of information, therefore, indicates the ability of a piece of information to attract the attention of the users, against the rest. Therefore, apart from the individual information diffusion processes, the information visibility dynamics also involves a competition process, where each information diffusion process competes against each other to draw the attention of users. Despite models of individual information diffusion processes abounding in literature, modeling the competition process is left unaddressed. In this paper, we propose Competing Recurrent Point Process (CRPP), a probabilistic deep machinery that unifies the nonlinear generative dynamics of a collection of diffusion processes, and inter-process competition - the two ingredients of visibility dynamics. To design this model, we rely on a recurrent neural network (RNN) guided generative framework, where the recurrent unit captures the joint temporal dynamics of a group of processes. This is aided by a discriminative model which captures the underlying competition process by discriminating among the various processes using several ranking functions. On ten diverse datasets crawled from Amazon and Twitter, CRPP offers a substantial performance boost in predicting item visibility against several baselines, thereby achieving significant accuracy in predicting both the collective diffusion mechanism and the underlying competition processes. Avirup Saha, Bidisha Samanta, Niloy Ganguly, Abir De |
CIKM | 1 |