VLDB 2026 Research / reviewers in the wild / expert
Koustuv Dasgupta
dblp:06/111
· DBLP profile ↗
43ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 18 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 9Artificial intelligence and machine learning · 8 · 4 since 2021Computer networks · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 5Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 62% Question answering and dialogue systems · 38% | |
| Human-computer interaction and pervasive computing
5 papers |
Ubiquitous computing and smart environments · 75% Wearable and physiological sensing · 15% Collaborative and social computing · 10% | |
| Databases, data mining, and information retrieval
3 papers |
Recommender systems · 82% Web and social media mining · 18% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Distributed systems · 68% Storage systems · 20% Cloud and datacenter computing · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% | |
| Software engineering, system software, and programming languages
2 papers |
Services computing and microservices · 100% |
Topics — the 24 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
hallucination detection |
0.9 | 1 | 2025 | PHANTOM: A Benchmark for Hallucination Detection in Financial Long-Context QA · NeurIPS 2025 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
long-context question answering |
0.9 | 1 | 2025 | PHANTOM: A Benchmark for Hallucination Detection in Financial Long-Context QA · NeurIPS 2025 |
Natural language and speech › Language models and text generation
text summarization |
0.6 | 1 | 2022 | ECTSum: A New Benchmark Dataset For Bullet Point Summarization of Long Earnings Call Transcripts · EMNLP 2022 |
Ubiquitous computing and smart environments
mobile crowdsourcing |
0.5 | 2 | 2016 | TASKer: behavioral insights via campus-based experimental mobile crowd-sourcing · UbiComp 2016 Campus-Scale Mobile Crowd-Tasking: Deployment & Behavioral Insights · CSCW 2016 |
Computational finance and economics › financial data analysis
financial document analysis |
0.3 | 1 | 2025 | PHANTOM: A Benchmark for Hallucination Detection in Financial Long-Context QA · NeurIPS 2025 |
Recommender systems › personalized ranking
top-n recommendation |
0.2 | 1 | 2016 | CAPReS: Context Aware Persona Based Recommendation for Shoppers · AAAI 2016 |
Ubiquitous computing and smart environments › context recognition
activity recognition |
0.2 | 1 | 2016 | IRIS: Tapping wearable sensing to capture in-store retail insights on shoppers · PerCom 2016 |
Ubiquitous computing and smart environments › context recognition › activity recognition
complex activity recognition |
0.2 | 1 | 2016 | IRIS: Tapping wearable sensing to capture in-store retail insights on shoppers · PerCom 2016 |
Wearable and physiological sensing
smartwatch sensing |
0.2 | 1 | 2016 | IRIS: Tapping wearable sensing to capture in-store retail insights on shoppers · PerCom 2016 |
Services computing and microservices › service composition
web service composition |
0.1 | 2 | 2005 | A service creation environment based on end to end composition of Web services · WWW 2005 Building Applications Using End to End Composition of Web Services · AAAI 2005 |
Ubiquitous computing and smart environments
context-aware computing |
0.1 | 1 | 2009 | Programmable Presence Virtualization for Next-Generation Context-Based Applications · PerCom 2009 |
Distributed systems
publish/subscribe systems |
0.1 | 1 | 2009 | Programmable Presence Virtualization for Next-Generation Context-Based Applications · PerCom 2009 |
Web and social media mining
social network analysis |
0.1 | 1 | 2008 | Analyzing the Structure and Evolution of Massive Telecom Graphs · IEEE Trans. Knowl. Data Eng. 2008 |
Collaborative and social computing › social media
social network sites |
0.1 | 1 | 2008 | R-U-in?: doing what you like, with people whom you like · WWW 2008 |
Services computing and microservices › service composition
automated service composition |
0.1 | 1 | 2005 | A service creation environment based on end to end composition of Web services · WWW 2005 |
Storage systems
data migration |
0.1 | 1 | 2005 | QoSMig: Adaptive Rate-Controlled Migration of Bulk Data in Storage Systems · ICDE 2005 |
Distributed systems › replication › replica management
replica placement |
0.0 | 1 | 2001 | Optimal Placement of Replicas in Trees with Read, Write, and Storage Costs · IEEE Trans. Parallel Distributed Syst. 2001 |
Distributed systems
replication |
0.0 | 1 | 2001 | Optimal Placement of Replicas in Trees with Read, Write, and Storage Costs · IEEE Trans. Parallel Distributed Syst. 2001 |
Mathematical optimization
optimization |
0.0 | 1 | 2001 | Optimal Placement of Replicas in Trees with Read, Write, and Storage Costs · IEEE Trans. Parallel Distributed Syst. 2001 |
Software-defined and programmable networks
programmable network services |
0.0 | 1 | 2009 | Programmable Presence Virtualization for Next-Generation Context-Based Applications · PerCom 2009 |
Cloud and datacenter computing › quality of service
differentiated service |
0.0 | 1 | 2005 | QoSMig: Adaptive Rate-Controlled Migration of Bulk Data in Storage Systems · ICDE 2005 |
Cloud and datacenter computing
quality of service |
0.0 | 1 | 2005 | QoSMig: Adaptive Rate-Controlled Migration of Bulk Data in Storage Systems · ICDE 2005 |
Distributed systems
service-oriented architecture |
0.0 | 1 | 2005 | Building Applications Using End to End Composition of Web Services · AAAI 2005 |
Automated reasoning and model checking
planning |
0.0 | 1 | 2005 | A service creation environment based on end to end composition of Web services · WWW 2005 |
Methods — techniques the papers use, named apart from their topics
large language model fine-tuning · 1.7benchmark construction · 1.7dynamic programming · 0.3XSLT transformation · 0.3XML processing · 0.3segmentation · 0.2hidden markov model · 0.2heuristic · 0.2field deployment · 0.2differential pricing · 0.2deployment study · 0.2cheating analytics · 0.2change point detection · 0.2behavioral analysis · 0.2web 2.0 · 0.2IMS converged network · 0.2ontology · 0.1WSDL · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptation of Embedding Models to Financial Filings Via LLM Distillation
Eliot Brenner, Dominic Seyler, Manjunath Hegde, Andrei Simion, Koustuv Dasgupta, Bing Xiang |
IEEE Big Data | 5 |
| 2025 | PHANTOM: A Benchmark for Hallucination Detection in Financial Long-Context QAabstractWhile Large Language Models (LLMs) show great promise, their tendencies to hallucinate pose significant risks in high-stakes domains like finance, especially when used for regulatory reporting and decision-making. Existing hallucination detection benchmarks fail to capture the complexities of financial benchmarks, which require high numerical precision, nuanced understanding of the language of finance, and ability to handle long-context documents. To address this, we introduce PHANTOM, a novel benchmark dataset for evaluating hallucination detection in long-context financial QA. Our approach first generates a seed dataset of high-quality "query-answer-document (chunk)" triplets, with either hallucinated or correct answers - that are validated by human annotators and subsequently expanded to capture various context lengths and information placements. We demonstrate how PHANTOM allows fair comparison of hallucination detection models and provides insights into LLM performance, offering a valuable resource for improving hallucination detection in financial applications. Further, our benchmarking results highlight the severe challenges out-of-the-box models face in detecting real-world hallucinations on long context data, and establish some promising directions towards alleviating these challenges, by fine-tuning open-source LLMs using PHANTOM. Lanlan Ji, Dominic Seyler, Gunkirat Kaur, Manjunath Hegde, Koustuv Dasgupta, Bing Xiang |
NeurIPS | 5 |
| 2024 | Parameter-Efficient Instruction Tuning of Large Language Models For Extreme Financial Numeral LabellingabstractSubhendu Khatuya, Rajdeep Mukherjee, Akash Ghosh, Manjunath Hegde, Koustuv Dasgupta, Niloy Ganguly, Saptarshi Ghosh, Pawan Goyal. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Subhendu Khatuya, Rajdeep Mukherjee, Akash Ghosh, Manjunath Hegde, Koustuv Dasgupta, Niloy Ganguly, Saptarshi Ghosh 0001, Pawan Goyal 0002 |
NAACL-HLT | 5 |
| 2022 | ECTSum: A New Benchmark Dataset For Bullet Point Summarization of Long Earnings Call TranscriptsabstractRajdeep Mukherjee, Abhinav Bohra, Akash Banerjee, Soumya Sharma, Manjunath Hegde, Afreen Shaikh, Shivani Shrivastava, Koustuv Dasgupta, Niloy Ganguly, Saptarshi Ghosh, Pawan Goyal. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Rajdeep Mukherjee, Abhinav Bohra, Akash Banerjee, Soumya Sharma, Manjunath Hegde, Afreen Shaikh, Shivani Shrivastava, Koustuv Dasgupta, Niloy Ganguly, Saptarshi Ghosh 0001, Pawan Goyal 0002 |
EMNLP | 8 |
| 2016 | CAPReS: Context Aware Persona Based Recommendation for ShoppersabstractNowadays, brick-and-mortar stores are finding it extremely difficult to retain their customers due to the ever increasing competition from the online stores. One of the key reasons for this is the lack of personalized shopping experience offered by the brick-and-mortar stores. This work considers the problem of persona based shopping recommendation for such stores to maximize the value for money of the shoppers. For this problem, it proposes a non-polynomial time-complexity optimal dynamic program and a polynomial time-complexity non-optimal heuristic, for making top-k recommendations by taking into account shopper persona and her time and budget constraints. In our empirical evaluations with a mix of real-world data and simulated data, the performance of the heuristic in terms of the persona based recommendations (quantified by similarity scores and items recommended) closely matched (differed by only 8% each with) that of the dynamic program and at the same time heuristic ran at least twice faster compared to the dynamic program. Joydeep Banerjee, Gurulingesh Raravi, Manoj Gupta 0002, Sindhu Kiranmai Ernala, Shruti Kunde, Koustuv Dasgupta |
AAAI | 6 |
| 2016 | Campus-Scale Mobile Crowd-Tasking: Deployment & Behavioral InsightsabstractMobile crowd-tasking markets are growing at an unprecedented rate with increasing number of smartphone users. Such platforms differ from their online counterparts in that they demand physical mobility and can benefit from smartphone processors and sensors for verification purposes. Despite the importance of such mobile crowd-tasking markets, little is known about the labor supply dynamics and mobility patterns of the users. Thivya Kandappu, Archan Misra, Shih-Fen Cheng, Nikita Jaiman, Randy Tandriansyah, Cen Chen 0001, Hoong Chuin Lau, Deepthi Chander, Koustuv Dasgupta |
CSCW | 9 |
| 2016 | TASKer: behavioral insights via campus-based experimental mobile crowd-sourcingabstractWhile mobile crowd-sourcing has become a game-changer for many urban operations, such as last mile logistics and municipal monitoring, we believe that the design of such crowd-sourcing strategies must better accommodate the real-world behavioral preferences and characteristics of users. To provide a real-world testbed to study the impact of novel mobile crowd-sourcing strategies, we have designed, developed and experimented with a real-world mobile crowd-tasking platform on the SMU campus, called TA&Sslash;Ker. We enhanced the TA$Ker platform to support several new features (e.g., task bundling, differential pricing and cheating analytics) and experimentally investigated these features via a two-month deployment of TA$Ker, involving 900 real users on the SMU campus who performed over 30,000 tasks. Our studies (i) show the benefits of bundling tasks as a combined package, (ii) reveal the effectiveness of differential pricing strategies and (iii) illustrate key aspects of cheating (false reporting) behavior observed among workers. Thivya Kandappu, Nikita Jaiman, Randy Tandriansyah, Archan Misra, Shih-Fen Cheng, Cen Chen 0001, Hoong Chuin Lau, Deepthi Chander, Koustuv Dasgupta |
UbiComp | 9 |
| 2016 | IRIS: Tapping wearable sensing to capture in-store retail insights on shoppersabstractWe investigate the possibility of using a combination of a smartphone and a smartwatch, carried by a shopper, to get insights into the shopper's behavior inside a retail store. The proposed IRIS framework uses standard locomotive and gestural micro-activities as building blocks to define novel composite features that help classify different facets of a shopper's interaction/experience with individual items, as well as attributes of the overall shopping episode or the store. Besides defining such novel features, IRIS builds a novel segmentation algorithm, which partitions the duration of an entire shopping episode into atomic item-level interactions, by using a combination of feature-based landmarking, change point detection and variable-order HMM-based sequence prediction. Experiments with 50 real-life grocery shopping episodes, collected from 25 shoppers, we show that IRIS can demarcate item-level interactions with an accuracy of approx. 91%, and subsequently characterize item-and-episode level shopper behavior with accuracies of over 90%. Meera Radhakrishnan, Sharanya Eswaran, Archan Misra, Deepthi Chander, Koustuv Dasgupta |
PerCom | 5 |
| 2015 | PISCES: Participatory Incentive Strategies for Effective Community Engagement in Smart CitiesabstractA key challenge in participatory sensing systems has been the design of incentive mechanisms that motivate individuals to contribute data to consuming applications. Emerging trends in urban development and smart city planning indicate the use of citizen reports to gather insights and identify areas for transformation. Consumers of these reports (e.g. city agencies) typically associate non-uniform utility (or values) to different reports based on the spatio-temporal context of the reports. For example, a report indicating traffic congestion near an airport, in early morning hours, would tend to have much higher utility than a similar report from a sparse residential area. In such cases, the design of an incentive mechanism must motivate participants, via appropriate rewards (or payments), to provide higher utility reports when compared to less valued ones. The main challenge in designing such an incentive scheme is two-fold: (i) lack of prior knowledge of participants in terms of their availability (i.e. who are in the vicinity) and reporting behaviour (i.e. what are the rewards expected); and (ii) minimizing payments to the reporters while ensuring that the desired number of reports are collected. In this paper, we propose STOC-PISCES, an algorithm that guarantees a stochastic optimal solution in the generalized setting of an unknown set of participants, with non-deterministic availabilities and stochastically rational reporting behaviour. The superior performance of STOC-PISCES in experimental settings, based on real-world data, endorses its adoption as an incentive strategy in participatory sensing applications like smart city management. Arpita Biswas, Deepthi Chander, Koustuv Dasgupta, Koyel Mukherjee 0001, Mridula Singh, Tridib Mukherjee |
HCOMP | 3 |
| 2015 | Opportunities for Process Improvement: A Cross-Clientele Analysis of Event Data Using Process Mining
R. P. Jagadeesh Chandra Bose, Avantika Gupta, Deepthi Chander, Ajith Ramanath, Koustuv Dasgupta |
ICSOC | 5 |
| 2015 | Personalized Messaging Engine: The Next Step in Employee Engagement
Abhishek Tripathi, Aditya Hegde 0001, Koustuv Dasgupta |
ICSOC | 5 |
| 2015 | Fair Resource Allocation for Heterogeneous TasksabstractWe consider the problem of fair resource allocation for tasks where a resource can be assigned to at most one task, without any fractional allocation. The system is heterogeneous: capacity and cost may vary across resources, and different tasks may have different resource demand. Due to heterogeneity of resources, the cost of allocating a task in isolation, without any other competing task, may differ significantly from its allocation cost when the task is allocated along with other tasks. In this context, we consider the problem of allocating resource to tasks, while ensuring that the cost is distributed fairly across the tasks, namely, the ratio of allocation cost of a task to its isolation cost is minimized over all tasks. We show that this fair resource allocation problem is strongly NP-Hard even when the resources are of unit size by a reduction from 3-partition. Our central results are an LP rounding based algorithm with an approximation ratio of 2+ O(ϵ) for the problem when resources are of unit size, and a near-optimal greedy algorithm for a more restricted version. The above fair allocation problem arises for resource allocation in various context, such as, allocating computing resources for reservations requests from tenants in a data centre, allocating resources to computing tasks in grid computing, or allocating personnel for tasks in service delivery organizations. Koyel Mukherjee 0001, Partha Dutta, Gurulingesh Raravi, Thangaraj Rajasubramaniam, Koustuv Dasgupta, Atul Singh |
IPDPS | 5 |
| 2014 | CrowdUtility: A Recommendation System for Crowdsourcing PlatformsabstractCrowd workers exhibit varying work patterns, expertise, and quality leading to wide variability in the performance of crowdsourcing platforms. The onus of choosing a suitable platform to post tasks is mostly with the requester, often leading to poor guarantees and unmet requirements due to the dynamism in performance of crowd platforms. Towards this end, we demonstrate CrowdUtility, a statistical modelling based tool for evaluating multiple crowdsourcing platforms and recommending a platform that best suits the requirements of the requester. CrowdUtility uses an online Multi-Armed Bandit framework, to schedule tasks while optimizing platform performance. We demonstrate an end-to end system starting from requirements specification, to platform recommendation, to real-time monitoring. Deepthi Chander, Sakyajit Bhattacharya, L. Elisa Celis, Koustuv Dasgupta, Saraschandra Karanam, Vaibhav Rajan, Avantika Gupta |
HCOMP | 4 |
| 2014 | TRACCS: A Framework for Trajectory-Aware Coordinated Urban Crowd-SourcingabstractWe investigate the problem of large-scale mobile crowd-tasking, where a large pool of citizen crowd-workers are used to perform a variety of location-specific urban logistics tasks. Current approaches to such mobile crowd-tasking are very decentralized: a crowd-tasking platform usually provides each worker a set of available tasks close to the worker's current location; each worker then independently chooses which tasks she wants to accept and perform. In contrast, we propose TRACCS, a more coordinated task assignment approach, where the crowd-tasking platform assigns a sequence of tasks to each worker, taking into account their expected location trajectory over a wider time horizon, as opposed to just instantaneous location. We formulate such task assignment as an optimization problem, that seeks to maximize the total payoff from all assigned tasks, subject to a maximum bound on the detour (from the expected path) that a worker will experience to complete her assigned tasks. We develop credible computationally-efficient heuristics to address this optimization problem (whose exact solution requires solving a complex integer linear program), and show, via simulations with realistic topologies and commuting patterns, that a specific heuristic (called Greedy-ILS) increases the fraction of assigned tasks by more than 20%, and reduces the average detour overhead by more than 60%, compared to the current decentralized approach. Cen Chen 0001, Shih-Fen Cheng, Aldy Gunawan, Archan Misra, Koustuv Dasgupta, Deepthi Chander |
HCOMP | 5 |
| 2014 | Adaptive Performance Optimization over Crowd Labor ChannelsabstractWe describe a system which monitors the performance of labor channels within a crowdsourcing platform in an online manner. This allows us to automatically determine if and when to switch between labor channels in order to improve overall performance of crowd tasks. Saraschandra Karanam, Deepthi Chander, L. Elisa Celis, Koustuv Dasgupta, Vaibhav Rajan |
HCOMP | 4 |
| 2014 | Post It or Not: Viewership Based Posting of Crowdsourced TasksabstractWe propose an online scheduling algorithm for posting crowdsourcing tasks which maximizes a novel metric called task viewership. This metric is computed using stochastic model based on coverage process and it measures the likelihood that a task is viewed by multiple crowd workers, which is correlated to the likelihood that it will be selected and completed. Pallavi Manohar, Deepthi Chander, L. Elisa Celis, Koustuv Dasgupta, Sakyajit Bhattacharya |
HCOMP | 4 |
| 2014 | CityZen: A Cost-Effective City Management System with Incentive-Driven Resident EngagementabstractCities typically face a wide gamut of management and maintenance problems. Existing automated sensor based solutions are prohibitively expensive to deploy. Furthermore, these solutions need to be complemented by incorporating human judgment for accurate, timely and cost-effective city-related event identification. To this end, this work proposes City Zen, which is a novel platform for event reporting and analytics to engage residents towards city management through incentives. Key contributions include: (a) the City Zen platform with an app for enabling authenticated residents to report events in a city for end-to-end integrated smart city management, (b) Differentiated incentive management based on types and priorities of events, quality and timeliness of event reports as well as resident intent, and (c) a social dashboard for searching events and subscribing for event alerts with additional ability to provide feedback on the event reports. Ongoing pilots and our performance study indicate that the platform indeed performs city management cost-effectively depending on the incentive mechanisms used for engaging residents. Tridib Mukherjee, Deepthi Chander, Anirban Mondal, Koustuv Dasgupta, Ashwin Venkat |
MDM (1) | 4 |
| 2014 | CloudRank: A statistical modelling framework for characterizing user behaviour towards targeted cloud managementabstractA rank clustering system, CloudRank, is proposed that takes into account cloud user preference data to characterize cloud user behaviour and also identify (an initially unknown set of) groups of users with similar behaviour in an unsupervised manner. The user groups are determined based on fitting mixture models on the cloud user preference observations. A preference can be anything that a system designer would like to include to characterize high-level user requirements such as demands on performance, cost, security, availability, etc. CloudRank can be useful for: (i) cloud providers to target their service offerings according to the user groups (i.e. customer segments) through appropriate customization of services pertaining to the user groups typical requirements; (ii) recommendation systems or a marketplace (that enables inter-operability among different providers) to determine which offerings best suit certain user groups; and (iii) prediction of any new users behaviour based on their preference information. Results on realistic feedbacks from internal cloud service providers show an average of 80% accuracy of the proposed unsupervised technique. When compared with a supervised technique, i.e. when the number of user groups are known beforehand, the error is within 15%, thus making it a promising technique for realistic deployments, particularly when there is no prior knowledge regarding the clusters. Sakyajit Bhattacharya, Tridib Mukherjee, Koustuv Dasgupta |
NOMS | 3 |
| 2009 | User interests in social media sites: an exploration with micro-blogsabstractRecent technological advances in mobile-based access to social networking platforms and facilities to update information in real{time (e.g. in Facebook) have allowed an individual's online presence to be as ephemeral and dynamic in nature, as her very thoughts and interests. In this context, micro-blogging has been widely adopted by users as an effective means to capture and disseminate their thoughts and actions to a larger audience on a daily basis. Interestingly, daily chatters of a user obtained from her micro-blogs offer a unique information source to analyze and interpret her context in real-time - i.e. interests, intentions,and activities. In this paper, we gather data from the public timeline of Twitter spanning across ten worldwide cities over a period of four weeks. We use this dataset to (a) explore how users express interests in real-time through micro-blogs, and (b) understand how text mining techniques can be applied to interpret real-time context of a user based on her tweets. Initial findings reported herein suggest that social media sites like Twitter constitute a promising source for extracting user context that can be exploited by novel social networking applications. Nilanjan Banerjee, Dipanjan Chakraborty 0001, Koustuv Dasgupta, Sumit Mittal, Anupam Joshi, Seema Nagar, Angshu Rai, Sameer Madan |
CIKM | 3 |
| 2009 | R-U-In? - Exploiting Rich Presence and Converged Communications for Next-Generation Activity-Oriented Social NetworkingabstractWith the growing popularity of social networking, traditional Internet Service Providers (ISPs) and telecom operators have both started exploring new opportunities to boost their revenue streams. The efforts have facilitated consumers to stay connected to their favorite social networks,be it from an ISP portal or a mobile device. The use of Web 2.0 technologies and converged communication tools has further led to a rise in both user-generated content as well as contextual information (i.e. rich presence) about users - including their current location, availability, interests and moods. In this evolving landscape, social networking players need to innovate for value-centric usage models that increase customer stickiness,along with business models to monetize the social media. To this end, we present R-U-In? - an activity-oriented social networking system for users to collaborate and participate in activities of mutual interest. Activities can be initiated and scheduled on-demand and be as ephemeral as the user interests themselves. R-U-In? leverages contextual modeling and reasoning techniques to enable ldquosocial searchrdquo based on real-time user interests and finds potential matches for the proposed activity. Further, it exploits next-generation presence and communication technologies to manage the entire activity lifecycle in real-time. Initial survey results, based on a prototype implementation of R-U-In?, attest to the promise of realtime activity-oriented social networking - both in terms of an effective collaboration tool for value-oriented social networking users and an enhanced end-user experience. Nilanjan Banerjee, Dipanjan Chakraborty 0001, Koustuv Dasgupta, Sumit Mittal, Seema Nagar, Saguna Saguna |
Mobile Data Management | 3 |
| 2009 | Programmable Presence Virtualization for Next-Generation Context-Based ApplicationsabstractPresence, broadly defined as an event publish-notification infrastructure for converged applications, has emerged as a key mechanism for collecting and disseminating context attributes for next-generation services in both enterprise and provider domains. Current presence-based solutions and products lack in the ability to a) support flexible user-defined queries over dynamic presence data and b) derive composite presence from multiple provider domains. Accordingly, current uses of context are limited to individual domains/organizations and do not provide a programmable mechanism for rapid creation of context-aware services. This paper describes a presence virtualization architecture, where a Virtualized Presence Server receives customizable queries from multiple presence clients, retrieves the necessary data from the base presence servers, applies the required virtualization logic and notifies the presence clients. To support both query expressiveness and computational efficiency, virtualization queries are structured to separately identify both the XSLT-based transformation primitives and the presence sources over which the transformation occurs. For improved scalability, the proposed architecture offloads the XSLT-related processing to a high-performance XML processing engine. We describe our current implementation and present performance results that attest to the promise of this virtualization approach. Arup Acharya, Nilanjan Banerjee, Dipanjan Chakraborty 0001, Koustuv Dasgupta, Archan Misra, Shachi Sharma, Xiping Wang, Charles Wright |
PerCom | 4 |
| 2008 | Social ties and their relevance to churn in mobile telecom networksabstractSocial Network Analysis has emerged as a key paradigm in modern sociology, technology, and information sciences. The paradigm stems from the view that the attributes of an individual in a network are less important than their ties (relationships) with other individuals in the network. Exploring the nature and strength of these ties can help understand the structure and dynamics of social networks and explain real-world phenomena, ranging from organizational efficiency to the spread of information and disease. Koustuv Dasgupta, Balaji Viswanathan, Dipanjan Chakraborty 0001, Sougata Mukherjea, Amit Anil Nanavati, Anupam Joshi |
EDBT | 1 |
| 2008 | Data-WISE: Efficient management of data-intensive workflows in scheduled grid environmentsabstractThe execution of data-intensive workflow applications in scientific and enterprise grids has gained popularity in recent times. Such applications process large and dynamic data sets, and often present scope for optimized data handling that can be exploited for performance. Traditionally, core grid middleware technologies of scheduling and orchestration, have treated data management as a background activity - decoupled from job management and handled at the storage and/or network protocol level. We believe that an important requirement for building data-aware grid technologies lies in managing data flows at the application level, in conjunction with their computation counterparts. To this end, we present Data-WISE, an end-to-end framework for management of data-intensive workflows as first class citizens, that addresses aspects of data flow orchestration, co-scheduling and runtime management. The optimizations are focused on exploiting application structure for use of data parallelism, replication, and runtime adaptations. We implement data-WISE on a real testbed and demonstrate significant improvements in terms of application response time, resource utilization, and adaptability to varying resource conditions. The proposed framework acts as an important step towards making distributed execution of data-intensive workflows a reality. Gargi Dasgupta, Koustuv Dasgupta, Balaji Viswanathan |
NOMS | 2 |
| 2008 | R-U-in?: doing what you like, with people whom you likeabstractThis paper presents R-U-In? - a social networking application that leverages Web 2.0 and IMS-based Converged Networks technologies to create a rich next-generation service. R-U-In? allows a user to search (in real-time) and solicit participation of like-minded partners for an activity of mutual interest (e.g. a rock concert, a soccer game, or a movie). It is an example of a situational mashup application that exploits content and capabilities of a Telecom operator, blended with Web 2.0 technologies, to provide an enhanced, value-added service experience. Nilanjan Banerjee, Dipanjan Chakraborty 0001, Koustuv Dasgupta, Sumit Mittal, Seema Nagar |
WWW | 3 |
| 2008 | Analyzing the Structure and Evolution of Massive Telecom GraphsabstractWith the ever-growing competition in telecommunications markets, operators have to increasingly rely on business intelligence to offer the right incentives to their customers. Existing approaches for telecom business intelligence have almost solely focused on the individual behavior of customers. In this paper, we use the call detail records of a mobile operator to construct call graphs, that is, graphs induced by people calling each other. We determine the structural properties of these graphs and also introduce the Treasure-Hunt model to describe the shape of mobile call graphs. Moreover, we determine how the structure of these call graphs evolve over time. Finally, since short messaging service (SMS) is becoming a preferred mode of communication among many sections of the society, we study the properties of the SMS graph. Our analysis indicates several interesting similarities and differences between the SMS graph and the corresponding call graph. We believe that our analysis techniques can allow telecom operators to better understand the social behavior of their customers and potentially provide major insights for designing effective incentives. Amit Anil Nanavati, Dipanjan Chakraborty 0001, Koustuv Dasgupta, Sougata Mukherjea, Gautam Das 0005, Siva Gurumurthy, Anupam Joshi |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2008 | Compass: optimizing the migration cost vs. application performance tradeoffabstractWe investigate methodologies for placement and migration of logical data stores in virtualized storage systems leading to optimum system configuration in a dynamic workload scenario. The aim is to optimize the tradeoff between the performance or operational cost improvement resulting from changes in store placement, and the cost imposed by the involved data migration step. We propose a unified economic utility based framework in which the tradeoff can be formulated as a utility maximization problem where the utility of a configuration is defined as the difference between the benefit of a configuration and the cost of moving to the configuration. We present a storage management middleware framework and architecture Compass that allows systems designers to plug-in different placement as well as migration techniques for estimation of utilities associated with different configurations. The biggest obstacle in optimizing the placement benefit and migration cost tradeoff is the exponential number of possible configurations that one may have to evaluate. We present algorithms that explore the configuration space efficiently and compute a candidate set of configurations that optimize this cost-benefit tradeoff. Our algorithms have many desirable properties including local optimality. Comprehensive experimental studies demonstrate the efficacy of the proposed framework and exploration algorithms, as our algorithms outperform migration cost-oblivious placement strategies by up to 40% on real OLTP traces for many settings. Akshat Verma, Upendra Sharma, Rohit Jain, Koustuv Dasgupta |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2007 | An Integrated Development Environment for Web Service CompositionabstractWeb services provide an instantiation of the loosely coupled service–oriented architecture and facilitate the process of enterprise application integration by encapsulating information, software, and other resources. However, to exploit the true potential of web services, it is critical to develop technologies and tools for composing new services from existing ones. While numerous composition approaches have been developed in the past, very little has been done towards tooling. What is clearly lacking is an Integrated Development Environment (IDE) to ease the process of composition, thereby reducing development time and integration efforts. In this paper, we build on our previous work on service composition, and present an IDE for end–to–end composition of web services. We elaborate on the design of the IDE, describe its integration with existing technologies, and discuss its usability based on the findings of a user survey. Girish Chafle, Gautam Das 0005, Koustuv Dasgupta, Arun Kumar 0002, Sumit Mittal, Sougata Mukherjea, Biplav Srivastava |
ICWS | 3 |
| 2007 | Compass: Cost of Migration-aware Placement in Storage SystemsabstractWe investigate methodologies for placement and migration of logical data stores in virtualized storage systems leading to optimum system configuration in a dynamic workload scenario. The aim is to optimize the tradeoff between the performance or operational cost improvement resulting from changes in store placement, and the cost imposed by the involved data migration step. We propose a unified economic utility based framework in which the tradeoff can be formulated as a utility maximization problem where the utility of a configuration is defined as the difference between the benefit of a configuration and the cost of moving to the configuration. We present a storage management middleware framework and architecture Compass that allows systems designers to plug-in different placement as well as migration techniques for estimation of utilities associated with different configurations. The biggest obstacle in optimizing the placement benefit and migration cost tradeoff is the exponential number of possible configurations that one may have to evaluate. We present algorithms that explore the configuration space efficiently and compute a candidate set of configurations that optimize this cost-benefit tradeoff. Our algorithms have many desirable properties including local optimality. Comprehensive experimental studies demonstrate the efficacy of the proposed framework and exploration algorithms, as our algorithms outperform migration cost-oblivious placement strategies by up to 40% on real OLTP traces for many settings. Akshat Verma, Upendra Sharma, Rohit Jain, Koustuv Dasgupta |
Integrated Network Management | 4 |
| 2006 | On the structural properties of massive telecom call graphs: findings and implicationsabstractWith ever growing competition in telecommunications markets, operators have to increasingly rely on business intelligence to offer the right incentives to their customers. Toward this end, existing approaches have almost solely focussed on the individual behaviour of customers. Call graphs, that is, graphs induced by people calling each other, can allow telecom operators to better understand the interaction behaviour of their customers, and potentially provide major insights for designing effective incentives.In this paper, we use the Call Detail Records of a mobile operator from four geographically disparate regions to construct call graphs, and analyse their structural properties. Our findings provide business insights and help devise strategies for Mobile Telecom operators. Another goal of this paper is to identify the shape of such graphs. In order to do so, we extend the well-known reachability analysis approach with some of our own techniques to reveal the shape of such massive graphs. Based on our analysis, we introduce the Treasure-Hunt model to describe the shape of mobile call graphs. The proposed techniques are general enough for analysing any large graph. Finally, how well the proposed model captures the shape of other mobile call graphs needs to be the subject of future studies. Amit Anil Nanavati, Siva Gurumurthy, Gautam Das 0005, Dipanjan Chakraborty 0001, Koustuv Dasgupta, Sougata Mukherjea, Anupam Joshi |
CIKM | 5 |
| 2006 | DECO: Data Replication and Execution CO-scheduling for Utility Grids
Vikas Agarwal, Gargi Dasgupta, Koustuv Dasgupta, Amit Purohit, Balaji Viswanathan |
ICSOC | 3 |
| 2006 | Adaptation inWeb Service Composition and ExecutionabstractWeb services simplify enterprise application integration by facilitating reuse of existing components for creating new services. In a dynamic environment, it is imperative to design a Web Service Composition and Execution (WSCE) system that adapts to failure of component services or changes in their QoS offerings. In this paper, we motivate a staged approach for adaptive WSCE (A-WSCE) that cleanly separates the functional and non-functional requirements of a new service, and enables different environmental changes to be absorbed at different stages of composition and execution. We use Synthy, a prototype service creation environment, to implement our solution and demonstrate its effectiveness. Girish Chafle, Koustuv Dasgupta, Arun Kumar 0002, Sumit Mittal, Biplav Srivastava |
ICWS | 2 |
| 2006 | QoS-GRAF: A Framework for QoS based Grid Resource Allocation with Failure ProvisioningabstractIn this paper, it describes the QoS-GRAF, a framework for providing revenue maximization in a utility computing grid where jobs have multiple resource dependencies and differentiated QoS pricing. To solve the revenue maximization problem the linear relaxation based algorithms, MRPA and MLBA, that achieve performance within 1-5% of the optimal solution and significantly outperform alternative approaches are used. Both show better revenue earnings across small, medium and large jobs, with efficient resource utilization. As a part ongoing work, the backup algorithms for multiple failures is developed. Scheduling algorithms are incorporated to produce maximum profitable schedule considering job deadlines Gargi Dasgupta, Koustuv Dasgupta, Amit Purohit, Balaji Viswanathan |
IWQoS | 2 |
| 2006 | Efficient Querying and Resource Management Using Distributed Presence Information in Converged NetworksabstractNext-generation converged networks shall deliver many innovative services over the standardized SIPbased IMS signaling infrastructure. Several such services exploit the joint presence information of a consumer, i.e. SIP entity requesting a service, and a vendor, i.e. SIP resource providing a service. Presence information is a collection of contextual attributes (e.g. location, availability, reputation), some of which change dynamically. Moreover, this collective presence information is distributed across multiple presence servers. While performing query matching based on joint presence information, a server usually routes each query to a locally available resource. However, skews in the spatio-temporal distribution of queries and resources may require queries to be routed to alternate servers with available resources. We propose a novel Resource-Aware Query Routing scheme, called RAQR, where each server proactively establishes gradients to suitable servers via a diffusion-based algorithm. Gradients are set up whenever a server anticipates scarcity of resources and withdrawn when the resource crunch is mitigated. We compare RAQR with alternative resource matching schemes and show that it adapts to spatio-temporal variations in resource availability, thereby leading to effective query matching with minimal control overhead. Dipanjan Chakraborty 0001, Koustuv Dasgupta, Archan Misra |
MDM | 2 |
| 2005 | Building Applications Using End to End Composition of Web Services
Vikas Agarwal, Girish Chafle, Koustuv Dasgupta, Neeran M. Karnik, Arun Kumar 0002, Ashish Kundu, Anupam Mediratta, Sumit Mittal, Biplav Srivastava |
AAAI | 3 |
| 2005 | QoSMig: Adaptive Rate-Controlled Migration of Bulk Data in Storage SystemsabstractLogical reorganization of data and requirements of differentiated QoS in information systems necessitate bulk data migration by the underlying storage layer. Such data migration needs to ensure that regular client I/Os are not impacted significantly while migration is in progress. We formalize the data migration problem in a unified admission control framework that captures both the performance requirements of client I/Os and the constraints associated with migration. We propose an adaptive rate-control based data migration methodology, QoSMig, that achieves the optimal client performance in a differentiated QoS setting, while ensuring that the specified migration constraints are met QoSMig uses both long term averages and short term forecasts of client traffic to compute a migration schedule. We present an architecture based on Service Level Enforcement Discipline for Storage (SLEDS) that supports QoSMig. Our trace-driven experimental study demonstrates that QoSMig provides significantly better I/O performance as compared to existing migration methodologies. Koustuv Dasgupta, Sugata Ghosal, Rohit Jain, Upendra Sharma, Akshat Verma |
ICDE | 1 |
| 2005 | A service creation environment based on end to end composition of Web servicesabstractThe demand for quickly delivering new applications is increasingly becoming a business imperative today. Application development is often done in an ad hoc manner, without standard frameworks or libraries, thus resulting in poor reuse of software assets. Web services have received much interest in industry due to their potential in facilitating seamless business-to-business or enterprise application integration. A web services composition tool can help automate the process, from creating business process functionality, to developing executable workflows, to deploying them on an execution environment. However, we find that the main approaches taken thus far to standardize and compose web services are piecemeal and insufficient. The business world has adopted a (distributed) programming approach in which web service instances are described using WSDL, composed into flows with a language like BPEL and invoked with the SOAP protocol. Academia has propounded the AI approach of formally representing web service capabilities in ontologies, and reasoning about their composition using goal-oriented inferencing techniques from planning. We present the first integrated work in composing web services end to end from specification to deployment by synergistically combining the strengths of the above approaches. We describe a prototype service creation environment along with a use-case scenario, and demonstrate how it can significantly speed up the time-to-market for new services. Vikas Agarwal, Koustuv Dasgupta, Neeran M. Karnik, Arun Kumar 0002, Ashish Kundu, Sumit Mittal, Biplav Srivastava |
WWW | 2 |
| 2005 | Synthy: A system for end to end composition of web services
Vikas Agarwal, Girish Chafle, Koustuv Dasgupta, Neeran M. Karnik, Arun Kumar 0002, Sumit Mittal, Biplav Srivastava |
J. Web Semant. | 3 |
| 2003 | Topology-Aware Placement and Role Assignment for Energy-Efficient Information Gathering in Sensor NetworksabstractConsider a network of energy-constrained wireless nodes, capable of sensing and communicating, to be deployed over an area to be monitored. There is a set of points or regions of interest in that are, each of which must be sensed (covered) by at least one node. The nodes are allowed to perform in-network data aggregation. As a node may or may not cover one or more points/regions of interest, we allow nodes to assume two roles sensor (nodes that sense their vicinity and generate data packets) and relay (nodes that only aggregate and transmit data packets). We consider the problem of placing nodes in the monitoring area and assigning roles to them such that the system lifetime is maximized, while ensuring that each point/region of interest is covered by at least one sensor node. This is the maximum lifetime sensor deployment problem with coverage constraints. The paper presents a novel algorithm to solve this problem and provides experimental results to demonstrate the effectiveness of the proposed algorithm. Koustuv Dasgupta, Meghna Kukreja, Konstantinos Kalpakis |
ISCC | 1 |
| 2003 | An efficient clustering-based heuristic for data gathering and aggregation in sensor networksabstractThe rapid advances in processor, memory, and radio technology have enabled the development of distributed networks of small, inexpensive nodes that are capable of sensing, computation, and wireless communication. Sensor networks of the future are envisioned to revolutionize the paradigm of collecting and processing information in diverse environments. However, the severe energy constraints and limited computing resources of the sensors, present major challenges for such a vision to become a reality. We consider a network of energy-constrained sensors that are deployed over a region. Each sensor periodically produces information as it monitors its vicinity. The basic operation in such a network is the systematic gathering and transmission of sensed data gathering and transmission of sensed data to a base station for further processing. During data gathering, sensors have the ability to perform in-network aggregation (fusion) of data packets enroute to the base station. The lifetime of such a sensor system is the time during which we can gather information from all the sensors to the base station. A key challenge in data gathering is to maximize the system lifetime, given the energy constraints of the sensors. Given the location of sensors and the base station and the available energy at each sensor, we are interested in finding an efficient manner in which data should be collected from all the sensors and transmitted to the base station, such that the system lifetime is maximized. This is the maximum lifetime data-gathering problem. In this paper, we describe a heuristic to solve the data-gathering problem with aggregation in sensor networks. Our experimental results demonstrate that the proposed algorithm significantly outperform previous methods, in terms of system lifetime. Koustuv Dasgupta, Konstantinos Kalpakis, Parag Namjoshi |
WCNC | 1 |
| 2003 | Efficient algorithms for maximum lifetime data gathering and aggregation in wireless sensor networks
Konstantinos Kalpakis, Koustuv Dasgupta, Parag Namjoshi |
Comput. Networks | 2 |
| 2002 | Understanding service demand for adaptive allocation of distributed resourcesabstractInternet services experience frequent changes in demand, resource characteristics and service requirements. A service can meet its requirements with minimal cost only if the allocation of its distributed resources can be adaptively controlled. The design of adaptive control schemes requires a solid understanding of the dynamic characteristics of the distributed service, demand and resources. This study examines the dynamic properties of distributed demand and differs from prior demand characterization work by focusing on dynamic variations across clients, time and region which are crucial in adaptively allocating distributed resources. Our analysis of the demand for the 1998 World Cup web site finds that the dynamic behavior of a small subset of clients is representative of the entire demand, the churn in the active set of clients from day to day is relatively small and that regional demand shows significant, and predictable, variations in the hourly scales. These results will provide guidance for the design of adaptive policies. Jose Renato Santos, Koustuv Dasgupta, G. John Janakiraman, Yoshio Turner |
GLOBECOM | 2 |
| 2001 | Steiner-Optimal Data Replication in Tree Networks with Storage CostsabstractWe consider the problem of placing copies of objects at multiple locations in a distributed system, whose interconnection network is a tree, in order to minimize the cost of servicing read and write requests to the objects. We assume that the tree nodes have limited storage and the number of copies permitted may be limited. The set of nodes that have a copy of the object, called replica nodes, constitute the replica set of the object. Read requests of a node are serviced from the closest replica node. Write requests of a node are propagated to all the replicas of the object using a minimum cost Steiner tree that includes the writer and all replica nodes. The total cost associated with a replica set equals the cost of servicing all the read and write requests, plus the storage cost at all the replica nodes. We are interested in finding a replica set with minimum total cost, i.e. a Steiner-optimal replica set. Given a tree with n nodes, we provide an O(n/sup 6/p/sup 2/)-time algorithm for finding a Steiner-optimal replica set of size p, taking into consideration the read, write, and storage costs. Our algorithm can also find a Steiner-optimal replica set for a tree with n nodes in time O(n/sup 8/). We also demonstrate that the policy used to propagate write requests to all the replica nodes in the network affects the cost and configuration of the optimal replica set for the object. Konstantinos Kalpakis, Koustuv Dasgupta, Ouri Wolfson |
IDEAS | 2 |
| 2001 | Optimal Placement of Replicas in Trees with Read, Write, and Storage CostsabstractWe consider the problem of placing copies of objects in a tree network in order to minimize the cost of servicing read and write requests to objects when the tree nodes have limited storage and the number of copies permitted is limited. The set of nodes that have a copy of the object is the residence set of the object. A node wishing to read the object will read the object from the closest node in the residence set. A node wishing to update the object will update the copy of the object at all the nodes in the residence set. Updates are propagated over a certain minimum spanning tree. The cost associated with a residence set equals the cost of servicing all the read and write requests and the storage costs for those copies. We describe an O(n/sup 3/p/sup 2/)-time algorithm for finding an optimal residence set of size p for an object in a tree with n nodes, taking into consideration the read, write, and storage costs. Furthermore, we describe a O(n/sup 3/p/sup 2//spl Lambda//sub max//sup 2/)-time algorithm for finding a minimum cost normal p-residence set for an object in a tree, this time also taking into account the load imposed by the nodes of the tree on the nodes in a residence set and their capacity constraints, where /spl Lambda//sub max/ is an upper bound on the capacity of each node of the tree. Konstantinos Kalpakis, Koustuv Dasgupta, Ouri Wolfson |
IEEE Trans. Parallel Distributed Syst. | 2 |