Jaideep Srivastava

dblp:s/JaideepSrivastava · DBLP profile ↗
← Back
158ranked-venue papers
12as first author
13since 2021 · last 2026
0000-0001-9385-7545ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 96 · 8 first-author · 4 since 2021Artificial intelligence and machine learning · 66 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 19 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 since 2021Systems, architecture and hardware · 9Software engineering, systems software and programming languages · 8 · 2 first-authorComputer networks · 6Theory of computation · 4 · 1 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2026 BiMind: A Dual-Head Reasoning Model with Attention-Geometry Adapter for Incorrect Information Detection
abstract
Incorrect information poses significant challenges by disrupting content veracity and integrity, yet most detection approaches struggle to jointly balance textual content verification with external knowledge modification under collapsed attention geometries.To address this issue, we propose a dual-head reasoning framework, BiMind, which disentangles content-internal reasoning from knowledgeaugmented reasoning.In BiMind, we introduce three core innovations: (i) an attention geometry adapter that reshapes attention logits via token-conditioned offsets and mitigates attention collapse; (ii) a self-retrieval knowledge mechanism, which constructs an in-domain semantic memory through kNN retrieval and injects retrieved neighbors via feature-wise linear modulation; (iii) the uncertainty-aware fusion strategies, including entropy-gated fusion and a trainable agreement head, stabilized by a symmetric Kullback-Leibler agreement regularizer.To quantify the knowledge contributions, we define a novel metric, Value-of-eXperience (VoX), to measure instance-wise logit gains from knowledge-augmented reasoning.Experiment results on public datasets demonstrate that our BiMind model outperforms advanced detection approaches and provides interpretable diagnostics on when and why knowledge matters.Our BiMind model and the tested datasets are available at https://github.com/cvzh/BiMind.
Emily K. Vraga, Jisu Huh, Jaideep Srivastava
ACL (1)4
2026 Data Skeleton Learning: Scalable active clustering with sparse graph structures
Xun Fu, Bin Chen 0034, Yan-Li Lee 0001, Tian Zou, Xin Wang 0064, Zhen Liu 0006, Jaideep Srivastava
Pattern Recognit.9
2025 The Psychology of Falsehood: A Human-Centric Survey of Misinformation Detection
abstract
Arghodeep Nandi, Megha Sundriyal, Euna Mehnaz Khan, Jikai Sun, Emily K. Vraga, Jaideep Srivastava, Tanmoy Chakraborty. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Arghodeep Nandi, Megha Sundriyal, Euna Mehnaz Khan, Jikai Sun, Emily K. Vraga, Jaideep Srivastava, Tanmoy Chakraborty 0002
EMNLP6
2024 Skepticism and Sentiment Analysis to Understand Consumer Perception of Virtual Influencers in Social Media Advertising
abstract
Social media advertising has become one of the most popular forms of advertising in recent years. Advancements in video and image generation have led to the creation of avatars, which have changed the landscape of online marketing. The avatars used for social media advertising are known as virtual influencers. In this paper, we use Instagram comments to observe consumer perceptions of virtual influencers (VIs) and human influencers (HIs) in social media advertising. By investigating how consumers engage with and respond to VIs, the study aims to provide insights into their effectiveness compared to HIs. In this study, we consider two timelines, pre- and post-pandemic, for different types of posts by VIs and HIs. We focus mainly on two user-affective states: sentiment and skepticism. The results show that VIs are perceived with more negativity and skepticism compared to HIs, irrespective of the type of post in both time periods. However, negative and skeptical responses towards VIs decreased post-pandemic. This empirical study provides insights into the evolving field of influencer advertising and highlights changes in consumer perception pre- and post-pandemic towards VIs and HIs.
Smitha Muthya Sudheendra, Maral Abdollahi, Dongyeop Kang, Jisu Huh, Jaideep Srivastava
IEEE Big Data5
2024 SkOTaPA: A Dataset for Skepticism Detection in Online Text after Persuasion Attempt
abstract
Individuals often encounter persuasion attempts, during which a persuasion agent aims to persuade a target to change the target’s emotions, beliefs, and behaviors. These persuasion attempts can be observed in various social settings, such as advertising, public health, political campaigns, and personal relationships. During these persuasion attempts, targets generally like to preserve their autonomy, so their responses often manifest in some form of resistance, like a skeptical reaction. In order to detect such skepticism in response to persuasion attempts on social media, we developed a corpus based on consumer psychology. In this paper, we consider one of the most prominent areas in which persuasion attempts unfold: social media influencer marketing. In this paper, we introduce the skepticism detection corpus, SkOTaPA, which was developed using multiple independent human annotations, and inter-coder reliability was evaluated with Krippendorff’s alpha (0.709). We performed validity tests to show skepticism cannot be detected using other potential proxy variables like sentiment and sarcasm.
Smitha Muthya Sudheendra, Maral Abdollahi, Dongyeop Kang, Jisu Huh, Jaideep Srivastava
LREC/COLING5
2024 Boosting cluster tree with reciprocal nearest neighbors scoring
Zhen Liu 0006, Bin Chen 0034, Jaideep Srivastava
Eng. Appl. Artif. Intell.4
2024 Low-Light Phase Retrieval With Implicit Generative Priors
abstract
Phase retrieval (PR) is fundamentally important in scientific imaging and is crucial for nanoscale techniques like coherent diffractive imaging (CDI). Low radiation dose imaging is essential for applications involving radiation-sensitive samples. However, most PR methods struggle in low-dose scenarios due to high shot noise. Recent advancements in optical data acquisition setups, such as in-situ CDI, have shown promise for low-dose imaging, but they rely on a time series of measurements, making them unsuitable for single-image applications. Similarly, data-driven phase retrieval techniques are not easily adaptable to data-scarce situations. Zero-shot deep learning methods based on pre-trained and implicit generative priors have been effective in various imaging tasks but have shown limited success in PR. In this work, we propose low-dose deep image prior (LoDIP), which combines in-situ CDI with the power of implicit generative priors to address single-image low-dose phase retrieval. Quantitative evaluations demonstrate LoDIP's superior performance in this task and its applicability to real experimental scenarios.
Raunak Manekar, Elisa Negrini, Minh Pham 0003, Daniel Jacobs, Jaideep Srivastava, Stanley J. Osher, Jianwei Miao
IEEE Trans. Image Process.5
2023 Scalable clustering by aggregating representatives in hierarchical groups
Zhen Liu 0006, Debarati Das 0004, Bin Chen 0034, Jaideep Srivastava
Pattern Recognit.5
2021 Facilitating CPAP Adherence with Personalized Recommendations Using Artificial Neural Networks
abstract
Sleep apnea is a common sleep disorder that, if left untreated, can have critical complications to the individual. The most common and effective treatment for sleep apnea is the Continuous Positive Airway Pressure (CPAP) therapy. But it has a long-term adherence rate as low as 60% due to discomfort and other factors. Although previous research has attempted to increase CPAP usage, there has been little to no change in its average adherence for the past two decades. This paper attempts to change this scenario using a large longitudinal dataset combined with a Recurrent Neural Network model to generate therapy use recommendations after one month of therapy. We performed a retrospective cohort analysis on 3380 patients during their first six months of therapy and compared our personalized recommendation system with the current generic recommendations made by sleep physicians. We show that recommendations generated by our artificial neural network model are easier to achieve since they are significantly closer to patients' therapy progress while being equally successful in maintaining therapy adherence.
Matheus Araújo 0001, Tara Pereira, Jaideep Srivastava, Conrad Iber
CBMS3
2021 Defining and Monitoring Patient Clusters Based on Therapy Adherence in Sleep Apnea Management
abstract
Obstructive Sleep Apnea (OSA) is a disorder in which breathing repeatedly stops and starts due to recurrent episodes of partial and complete airway obstruction during sleep. One of the common treatments for moderate and severe OSA cases is the use of Continuous Positive Airway Pressure (CPAP) devices that keep the airways open. Unfortunately, about 40% of the patients using CPAP devices abandon their therapy within six months. In this work, we propose a method to cluster and monitor patients according to their therapy usage behavior aiming for a timely and appropriate intervention. Our data corresponds to 1815 CPAP users in their first six months of therapy. In contrast to the simple rule-based methods currently employed by sleep clinics to identify non-adherent behavior, our approach uses clustering techniques to group patients based on their CPAP usage patterns. After identifying four main clusters, we investigate how patients can change between clusters over the months using Markov Chain analysis. We observed that patients who change to a healthy cluster have a higher probability of staying there in the future, reinforcing the need for early intervention. Finally, we use machine learning-based models to predict the next month's probability of adherence and nonadherence according to our pre-defined cluster definitions.
Mourya Karan Reddy Baddam, Matheus Araújo 0001, Jaideep Srivastava
CBMS3
2021 IACN: Influence-Aware and Attention-Based Co-evolutionary Network for Recommendation
Shalini Pandey, George Karypis, Jaideep Srivastava
PAKDD (2)3
2021 SCARLET: Explainable Attention Based Graph Neural Network for Fake News Spreader Prediction
Bhavtosh Rath, Xavier Morales, Jaideep Srivastava
PAKDD (1)3
2021 Learning Spatiotemporal Latent Factors of Traffic via Regularized Tensor Factorization: Imputing Missing Values and Forecasting
abstract
Intelligent transportation systems are a key component in smart cities, and the estimation and prediction of the spatiotemporal traffic state is critical to capture the dynamics of traffic congestion, i.e., its generation, propagation and mitigation, in order to increase operational efficiency and improve livability within smart cities. And while spatiotemporal data related to traffic is becoming common place due to the wide availability of cheap sensors and the rapid deployment of IoT platforms, the data still suffer some challenges related to sparsity, incompleteness, and noise which makes the traffic analytics difficult. In this article, we investigate the problem of missing data or noisy information in the context of real-time monitoring and forecasting of traffic congestion for road networks in a city. The road network is represented as a directed graph in which nodes are junctions (intersections) and edges are road segments. We assume that the city has deployed high-fidelity sensors for speed reading in a subset of edges; and the objective is to infer the speed readings for the remaining edges in the network; and to estimate the missing values in the segments for which sensors have stopped generating data due to technical problems (e.g., battery, network, etc.). We propose a tensor representation for the series of road network snapshots, and develop a regularized factorization method to estimate the missing values, while learning the latent factors of the network. The regularizer, which incorporates spatial properties of the road network, improves the quality of the results. The learned factors, with a graph-based temporal dependency, are then used in an autoregressive algorithm to predict the future state of the road network with a large horizon. Extensive numerical experiments with real traffic data from the cities of Doha (Qatar) and Aarhus (Denmark) demonstrate that the proposed approach is appropriate for imputing the missing data and predicting the traffic state. It is accurate and efficient and can easily be applied to other traffic datasets.
Abdelkader Baggag, Sofiane Abbar, Ankit Sharma 0004, Tahar Zanouda, Abdulaziz Al-Homaid, Abhiraj Mohan, Jaideep Srivastava
IEEE Trans. Knowl. Data Eng.7
2020 Detecting Fake News Spreaders in Social Networks using Inductive Representation Learning
abstract
An important aspect of preventing fake news dissemination is to proactively detect the likelihood of its spreading. Research in the domain of fake news spreader detection has not been explored much from a network analysis perspective. In this paper, we propose a graph neural network based approach to identify nodes that are likely to become spreaders of false information. Using the community health assessment model and interpersonal trust we propose an inductive representation learning framework to predict nodes of densely-connected community structures that are most likely to spread fake news, thus making the entire community vulnerable to the infection. Using topology and interaction based trust properties of nodes in real-world Twitter networks, we are able to predict false information spreaders with an accuracy of over 90%.
Bhavtosh Rath, Aadesh Salecha, Jaideep Srivastava
ASONAM3
2020 RKT: Relation-Aware Self-Attention for Knowledge Tracing
abstract
The world has transitioned into a new phase of online learning in response to the recent Covid19 pandemic. Now more than ever, it has become paramount to push the limits of online learning in every manner to keep flourishing the education system. One crucial component of online learning is Knowledge Tracing (KT). The aim of KT is to model student's knowledge level based on their answers to a sequence of exercises referred as interactions. Students acquire their skills while solving exercises and each such interaction has a distinct impact on student ability to solve a future exercise. This impact is characterized by 1) the relation between exercises involved in the interactions and 2) student forget behavior. Traditional studies on knowledge tracing do not explicitly model both the components jointly to estimate the impact of these interactions. In this paper, we propose a novel Relation-aware self-attention model for Knowledge Tracing (RKT). We introduce a relation-aware self-attention layer that incorporates the contextual information. This contextual information integrates both the exercise relation information through their textual content as well as student performance data and the forget behavior information through modeling an exponentially decaying kernel function. Extensive experiments on three real-world datasets, among which two new collections are released to the public, show that our model outperforms state-of-the-art knowledge tracing methods. Furthermore, the interpretable attention weights help visualize the relation between interactions and temporal patterns in the human learning process.
Shalini Pandey, Jaideep Srivastava
CIKM2
2019 Adversarial Unsupervised Representation Learning for Activity Time-Series
abstract
Sufficient physical activity and restful sleep play a major role in the prevention and cure of many chronic conditions. Being able to proactively screen and monitor such chronic conditions would be a big step forward for overall health. The rapid increase in the popularity of wearable devices pro-vides a significant new source, making it possible to track the user’s lifestyle real-time. In this paper, we propose a novel unsupervised representation learning technique called activ-ity2vecthat learns and “summarizes” the discrete-valued ac-tivity time-series. It learns the representations with three com-ponents: (i) the co-occurrence and magnitude of the activ-ity levels in a time-segment, (ii) neighboring context of the time-segment, and (iii) promoting subject-invariance with ad-versarial training. We evaluate our method on four disorder prediction tasks using linear classifiers. Empirical evaluation demonstrates that our proposed method scales and performs better than many strong baselines. The adversarial regime helps improve the generalizability of our representations by promoting subject invariant features. We also show that using the representations at the level of a day works the best since human activity is structured in terms of daily routines.
Karan Aggarwal, Shafiq R. Joty, Luis Fernández-Luque, Jaideep Srivastava
AAAI4
2019 On churn and social contagion
abstract
Massively Multiplayer Online Role-Playing Games (MMORPGs) are persistent virtual environments where millions of players interact in an online manner. We study the problem of player churn and social contagion using MMORPG game logs by analyzing the impact of a node's churn behavior on its immediate neighborhood or group. The two key research questions in this paper are - When an active node, ego, becomes dormant, what is the impact on the activity behavior of ego's immediate neighbor, alter, 1) based on ego's characteristics and ego's relationship with alter and 2) based on the activity behavior of alter's remaining neighbors. We use a supervised learning framework to study the impact of player churn and social contagion. Experimental results show that the classification models perform substantially better than random for both the research problems. Finally, we use a data-driven approach to propose a player typology based on degree of socialization and analyze churn behavior among these player types. Experimental results show that the loner player type is much more likely to churn than the socializer player types and as the degree of socialization decreases among socializers, the propensity to churn increases.
Zoheb Borbora, Arpita Chandra, Ponnurangam Kumaraguru, Jaideep Srivastava
ASONAM4
2019 Finding your social space: empirical study of social exploration in multiplayer online games
abstract
Social dynamics are based on human needs for trust, support, resource sharing, irrespective of whether they operate in real life or in a virtual setting. Massively multiplayer online role-playing games (MMORPGS) serve as enablers of leisurely social activity and are important tools for social interactions. Past research has shown that socially dense gaming environments like MMORPGs can be used to study important social phenomena, which may operate in real life, too. We describe the process of social exploration to entail the following components 1) finding the balance between personal and social time 2) making choice between a large number of weak ties or few strong social ties. 3) finding a social group. In general, these are the major determinants of an individual's social life. This paper looks into the phenomenon of social exploration in an activity based online social environment. We study this process through the lens of the following research questions, 1) What are the different social behavior types? 2) Is there a change in a player's social behavior over time? 3) Are certain social behaviors more stable than the others? 4) Can longitudinal research of player behavior help shed light on the social dynamics and processes in the network? We use an unsupervised machine learning approach to come up with 4 different social behavior types - Lone Wolf, Pack Wolf of Small Pack, Pack Wolf of a Large Pack and Social Butterfly. The types represent the degree of socialization of players in the game. Our research reveals that social behaviors change with time. While lone wolf and pack wolf of small pack are more stable social behaviors, pack wolf of large pack and social butterflies are more transient. We also observe that players progressively move from large groups with weak social ties to settle in small groups with stronger ties.
Arpita Chandra, Zoheb Borbora, Ponnurangam Kumaraguru, Jaideep Srivastava
ASONAM4
2019 Evaluating vulnerability to fake news in social networks: a community health assessment model
abstract
Understanding the spread of false information in social networks has gained a lot of recent attention. In this paper, we explore the role community structures play in determining how people get exposed to fake news. Inspired by approaches in epidemiology, we propose a novel Community Health Assessment model, whose goal is to understand the vulnerability of communities to fake news spread. We define the concepts of neighbor, boundary and core nodes of a community and propose appropriate metrics to quantify the vulnerability of nodes (individual-level) and communities (group-level) to spreading fake news. We evaluate our model on communities identified using three popular community detection algorithms for twelve real-world news spreading networks collected from Twitter. Experimental results show that the proposed metrics perform significantly better on the fake news spreading networks than on the true news, indicating that our community health assessment model is effective.
Bhavtosh Rath, Wei Gao 0001, Jaideep Srivastava
ASONAM3
2019 A Data-Driven Approach for Continuous Adherence Predictions in Sleep Apnea Therapy Management
abstract
The abandonment rate of patients who use CPAP devices for obstructive sleep apnea (OSA) therapy is as high as 60%. However, there is growing evidence that timely and appropriate intervention can improve long-term adherence to therapy. Current practice in sleep clinics of identifying potential patients who will abandon the treatment is not sufficiently effective in terms of accuracy and timeliness. Recent proposals in the literature have tried to identify non-adherent patients in a specific period of their therapy; however, there is no generalized approach by which clinical providers can monitor their patients continually with the goal of maximizing adherence. Towards this more generic goal, we propose CTAP-CPAP, a Continuous Treatment Adherence Prediction framework. With CTAP-CPAP, we address the problem of generalizing the prediction for any day in the treatment, where a robust framework with multiple machine learning models is implemented to assist medical practitioners keep track of the patient risk of non-adherence. Aiming the parallel progress of both machine learning and health informatics fields, we complement the study with a transparent discussion on the machine learning techniques used to build CTAP-CPAP and our view of its operationalization in a sleep clinic.
Matheus Araújo 0001, Louis Kazaglis, Conrad Iber, Jaideep Srivastava
IEEE BigData4
2019 Using Clinical Notes with Time Series Data for ICU Management
abstract
Swaraj Khadanga, Karan Aggarwal, Shafiq Joty, Jaideep Srivastava. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Swaraj Khadanga, Karan Aggarwal, Shafiq R. Joty, Jaideep Srivastava
EMNLP/IJCNLP (1)4
2018 A Structured Learning Approach with Neural Conditional Random Fields for Sleep Staging
abstract
Sleep plays a vital role in human health, both mental and physical. Sleep disorders like sleep apnea are increasing in prevalence, with the rapid increase in factors like obesity. Sleep apnea is most commonly treated with Continuous Positive Air Pressure (CPAP) therapy. Presently, however, there is no mechanism to monitor a patient's progress with CPAP. Accurate detection of sleep stages from CPAP flow signal is crucial for such a mechanism. We propose, for the first time, an automated sleep staging model based only on the flow signal. Deep neural networks have recently shown high accuracy on sleep staging by eliminating handcrafted features. However, these methods focus exclusively on extracting informative features from the input signal, without paying much attention to the dynamics of sleep stages in the output sequence. We propose an end-to-end framework that uses a combination of deep convolution and recurrent neural networks to extract high-level features from raw flow signal with a structured output layer based on a conditional random field to model the temporal transition structure of the sleep stages. We improve upon the previous methods by 10% using our model, that can be augmented to the previous sleep staging deep learning methods. We also show that our method can be used to accurately track sleep metrics like sleep efficiency calculated from sleep stages that can be deployed for monitoring the response of CPAP therapy on sleep apnea patients. Apart from the technical contributions, we expect this study to motivate new research questions in sleep science.
Karan Aggarwal, Swaraj Khadanga, Shafiq R. Joty, Louis Kazaglis, Jaideep Srivastava
IEEE BigData5
2017 Participatory Health Informatics for Data-Driven Precision Medicine
Jaideep Srivastava, Luis Fernández-Luque, Fernando Martín-Sánchez, Kenneth D. Mandl
AMIA1
2017 From Retweet to Believability: Utilizing Trust to Identify Rumor Spreaders on Twitter
abstract
Ubiquitous use of social media such as microblogging platforms brings about ample opportunities for the false information to diffuse online. It is very important not just to determine the veracity of information but also the authenticity of the users who spread the information, especially in time-critical situations like real-world emergencies, where urgent measures have to be taken for stopping the spread of fake information. In this work, we propose a novel machine learning based approach for automatic identification of the users spreading rumorous information by leveraging the concept of believability, i.e., the extent to which the propagated information is likely to be perceived as truthful, based on the trust measures of users in Twitter's retweet network. We hypothesize that the believability between two users is proportional to the trustingness of the retweeter and the trustworthiness of the tweeter, which are two complementary measures of user trust and can be inferred from retweeting behaviors using a variant of HITS algorithm. With the retweet network edge-weighted by believability scores, we use network representation learning to generate user embeddings, which are then leveraged to classify users into as rumor spreaders or not. Based on experiments on a very large real-world rumor dataset collected from Twitter, we demonstrate that our method can effectively identify rumor spreaders and outperform four strong baselines with large margin.
Bhavtosh Rath, Wei Gao 0001, Jing Ma 0004, Jaideep Srivastava
ASONAM4
2017 Visualization of Wearable Data and Biometrics for Analysis and Recommendations in Childhood Obesity
abstract
Obesity is one of the major health risk factors behind the rise of non-communicable conditions. Understanding the factors influencing obesity is very complex since there are many variables that can affect the health behaviors leading to it. Nowadays, multiple data sources can be used to study health behaviors, such as wearable sensors for physical activity and sleep, social media, mobile and health data. In this paper we describe the design of a dashboard for the visualization of actigraphy and biometric data from a childhood obesity camp in Qatar. This dashboard allows quantitative discoveries that can be used to guide patient behavior and orient qualitative research.
Michaël Aupetit 0001, Luis Fernández-Luque, Meghna Singh, Jaideep Srivastava
CBMS4
2017 The 360QS Toolkit for Sleep and Physical Activity Analysis Based on Wearables
abstract
Sleep and physical activity are human behaviours that play a major role in our health. Poor sleep or lack of physical activity have been found to increase health risks and reduce quality of life. The rapid adoption and evolution of pervasive computing systems, both in the health and wellness domain, are creating a new data-intensive context in which we can learn about the sleep and physical activity patterns of individuals. In this paper we provide an overview of the toolkit we have developed to conduct research on personal health data about sleep and physical activity. This toolkit has been used to develop predictive models of sleep quality based on wearable data, and also to create data visualizations to help healthcare professionals in making decisions.
Meghna Singh, Luis Fernández-Luque, Jaideep Srivastava
CBMS3
2017 Advanced Computation of Sparse Precision Matrices for Big Data
Abdelkader Baggag, Halima Bensmail, Jaideep Srivastava
PAKDD (2)3
2017 Weighted Simplicial Complex: A Novel Approach for Predicting Small Group Evolution
Ankit Sharma 0004, Terrence J. Moore, Ananthram Swami, Jaideep Srivastava
PAKDD (1)4
2017 Effective Urban Structure Inference from Traffic Flow Dynamics
abstract
Mobility in a city is represented as traffic flows in and out of defined urban travel or administrative zones. While the zones and the road networks connecting them are fixed in space, traffic flows between pairs of zones are dynamic through the day. Understanding these dynamics in real time is crucial for real time traffic planning in the city. In this paper, we use real time traffic flow data to generate dense functional correlation matrices between zones during different times of the day. Then, we derive optimal sparse representations of these dense functional matrices, that accurately recover not only the existing road network connectivity between zones, but also reveal new latent links between zones that do not yet exist but are suggested by traffic flow dynamics. We call this sparse representation the time-varying effective traffic connectivity of the city. A convex optimization problem is formulated and used to infer the sparse effective traffic network from time series data of traffic flow for arbitrary levels of temporal granularity. We demonstrate the results for the city of Doha, Qatar on data collected from several hundred bluetooth sensors deployed across the city to record vehicular activity through the city's traffic zones. While the static road network connectivity between zones is accurately inferred, other long range connections are also predicted that could be useful in planning future road linkages in the city. Further, the proposed model can be applied to socio-economic activity other than traffic, such as new housing, construction, or economic activity captured as functional correlations between zones, and can also be similarly used to predict new traffic linkages that are latently needed but as yet do not exist. Preliminary experiments suggest that our framework can be used by urban transportation experts and policy specialists to take a real time data-driven approach towards urban planning and real time traffic planning in the city, especially at the level of administrative zones of a city.
Somwrita Sarkar, Sanjay Chawla, Shameem Ahmad, Jaideep Srivastava, Hossam M. Hammady, Fethi Filali, Wassim Znaidi, Javier Borge-Holthoefer
IEEE Trans. Big Data4
2017 Formation and Reciprocation of Dyadic Trust
abstract
This paper reports a detailed empirical study of interpersonal trust in a multi-relational online social network. This study addresses two main aspects of interpersonal trust: formation and reciprocation. Computational models developed, using multi-relational networks, for these processes provide interesting insights about online social interactions. Our findings for trust formation (initiation) indicate a strong role of lower familiarity interactions before trust(high familiarity relationship) is formed. Similarly, trust reciprocation is not automatic, but strongly depends on enough lower familiarity interactions. This study is the first quantification of the “scaffolding role” played by lower familiarity interactions, in formation of high familiarity relationships.
Atanu Roy, Ayush Singhal, Jaideep Srivastava
ACM Trans. Internet Techn.3
2017 Research dataset discovery from research publications using web context
abstract
Scientific datasets play a crucial role in data-driven research. While there are several repositories that curate public datasets, several more datasets and their usage is “hidden” in the research publications. Hence, discovering a relevant dataset for a research topic requires in-depth investigation of several publications, tracking dataset usage and in-exhaustive literature search. To this end, a search engine to directly handle the research dataset discovery problem is extremely useful for the scientific community. In this work, we define an important paradigm of dataset search known as “ dataset discovery in application context”. Unlike dataset look-up type search where the user looks up for dataset in a repository, application context based search corresponds to search without information about the name of the dataset. Such searches arise when the user is looking a best fit dataset for his research problem. We show that in this paradigm of search, conventional methods of indexing the little text about the dataset description do not work due to lack of application text content within the description text for a dataset. To alleviate this problem we propose two models of search, namely, (1) a user profile based search and (2) a keyword based search. We show that in both these models the dataset discovery is done in the application context by leveraging information from open source web resources such as scholarly articles repositories and academic search engines. The performance of the proposed models were tested with simulated test queries (user profiles) as well as with real world user studies.
Ayush Singhal, Jaideep Srivastava
Web Intell.2
2016 Trustingness & trustworthiness: A pair of complementary trust measures in a social network
abstract
The increase in analysis of real life social networks has led to a better understanding of the ways humans socialize in a group. Since trust is an important part of any social interaction, researchers use such networks to understand the nuances of trust relationships. One of the major requirements in trust applications is identifying the trustworthy actors in these networks. This paper proposes a pair of complementary measures that can be used to measure trust scores of actors in a social network using involvement of social networks. Based on the proposed measures, an iterative matrix convergence algorithm is developed that calculates the trustingness and the trustworthiness of each actor in the network. Trustingness of an actor is defined as the propensity of an actor to trust his neighbors in the network. Trustworthiness, on the other hand, is defined as the willingness of the network to trust an individual actor. The algorithm is proposed based on the idea that a person having higher trustingness score contributes to the trustworthiness of its neighbors to a lower degree. Conversely, a higher trustworthiness score is a result of lots of neighbors linked to the actor having low trustingness scores. The algorithm runs in O(k × |E|) time where k denotes the number of iterations and |E| denotes the number of edges in the network. Moreover, the paper shows that the algorithm converges to a finite value quickly. Finally the proposed scores for trust prediction is implemented for various social networks and is shown that the proposed algorithm performs better (average 5%) than the state of the art trust scoring algorithms. The full version of the paper along with the coding implementation can be found at [3].
Atanu Roy, Chandrima Sarkar, Jaideep Srivastava, Jisu Huh
ASONAM3
2016 Querying and Tracking Influencers in Social Streams
abstract
Influence analysis is an important problem in social network analysis due to its impact on viral marketing and targeted advertisements. Most of the existing influence analysis methods determine the influencers in a static network with an influence propagation model based on pre-defined edge propagation probabilities. However, none of these models can be queried to find influencers in both context and time-sensitive fashion from a streaming social data. In this paper, we propose an approach to maintain real-time influence scores of users in a social stream using a topic and time-sensitive approach, while the network and topic is constantly evolving over time. We show that our approach is efficient in terms of online maintenance and effective in terms various types of real-time context- and time-sensitive queries. We evaluate our results on both social and collaborative network data sets.
Karthik Subbian, Charu C. Aggarwal, Jaideep Srivastava
WSDM3
2016 Social network regularized Sparse Linear Model for Top-N recommendation
Xiaodong Feng 0001, Ankit Sharma 0004, Jaideep Srivastava, Sen Wu 0001, Zhiwei Tang
Eng. Appl. Artif. Intell.3
2016 Mining Influencers Using Information Flows in Social Streams
abstract
The problem of discovering information flow trends in social networks has become increasingly relevant due to the increasing amount of content in online social networks, and its relevance as a tool for research into the content trends analysis in the network. An important part of this analysis is to determine the key patterns of flow in the underlying network. Almost all the work in this area has focused on fixed models of the network structure, and edge-based transmission between nodes. In this article, we propose a fully content-centered model of flow analysis in networks, in which the analysis is based on actual content transmissions in the underlying social stream, rather than a static model of transmission on the edges. First, we introduce the problem of influence analysis in the context of information flow in networks. We then propose a novel algorithm InFlowMine to discover the information flow patterns in the network and demonstrate the effectiveness of the discovered information flows using an influence mining application. This application illustrates the flexibility and effectiveness of our information flow model to find topic- or network-specific influencers, or their combinations. We empirically show that our information flow mining approach is effective and efficient than the existing methods on a number of different measures.
Karthik Subbian, Charu C. Aggarwal, Jaideep Srivastava
ACM Trans. Knowl. Discov. Data3
2015 Predicting Small Group Accretion in Social Networks: A topology based incremental approach
abstract
Small Group evolution has been of central importance in social sciences and also in the industry for understanding dynamics of team formation. While most of research works studying groups deal at a macro level with evolution of arbitrary size communities, in this paper we restrict ourselves to studying evolution of small group (size ≤ 20) which is governed by contrasting sociological phenomenon. Given a previous history of group collaboration between a set of actors, we address the problem of predicting likely future group collaborations. Unfortunately, predicting groups requires choosing from (n r) possibilities (where r is group size and n is total number of actors), which becomes computationally intractable as group size increases. However, our statistical analysis of a real world dataset has shown that two processes: an external actor joining an existing group (incremental accretion (IA)) or collaborating with a subset of actors of an exiting group (subgroup accretion (SA)), are largely responsible for future group formation. This helps to drastically reduce the (n r) possibilities. We therefore, model the attachment of a group for different actors outside this group. In this paper, we have built three topology based prediction models to study these phenomena. The performance of these models is evaluated using extensive experiments over DBLP dataset. Our prediction results shows that the proposed models are significantly useful for future group predictions both for IA and SA.
Ankit Sharma 0004, Rui Kuang, Jaideep Srivastava, Xiaodong Feng 0001, Kartik Singhal 0001
ASONAM3
2015 Rare Class Detection in Networks
abstract
The problem of node classification in networks is an important one in a wide variety of social networking domains. In many real applications such as product recommendations, the class of interest may be very rare. In such scenarios, it is often very difficult to learn the most relevant node classification characteristics, both because of the paucity of training data, and because of poor connectivity among rare class nodes in the network structure. Node classification methods crucially dependent upon structural homophily, and a lack of connectivity among rare class nodes can create significant challenges. However, many such social networks are content-rich, and the content-rich nature of such networks can be leveraged to compensate for the lack of structural connectivity among rare class nodes. While content-centric and semi-supervised methods have been used earlier in the context of paucity of labeled data, the rare class scenario has not been investigated in this context. In fact, we are not aware of any known classification method which is tailored towards rare class detection in networks. This paper will present a spectral approach for rare-class detection, which uses a distance-preserving transform, in order to combine the structural information in the network with the available content. We will show the advantage of this approach over traditional methods for collective classification.
Karthik Subbian, Charu C. Aggarwal, Jaideep Srivastava, Vipin Kumar 0001
SDM3
2015 Just in Time Recommendations: Modeling the Dynamics of Boredom in Activity Streams
abstract
Recommendation methods have mainly dealt with the problem of recommending new items to the user while user visitation behavior to the familiar items (items which have been consumed before) are little understood. In this paper, we analyze user activity streams and show that user's temporal consumption of familiar items is driven by boredom. Specifically, users move on to a different item when bored and return to the same item when their interest is restored. To model this behavior we include two latent psychological states of preference for items - sensitization and boredom. In the sensitization state the user is highly engaged with the item, while in the boredom state the user is disinterested. We model this behavior using a Hidden Semi-Markov Model for the gaps between user consumption activities. We show that our model performs much better than the state-of-the-art temporal recommendation models at predicting the revisit time to the item. Moreover, we attribute two main reasons for this: (1) recommending items that are not in the bored state for the user, (2) recommending items where user has restored her interests.
Komal Kapoor, Karthik Subbian, Jaideep Srivastava, Paul Schrater
WSDM3
2015 Automatic instance selection via locality constrained sparse representation for missing value estimation
Xiaodong Feng 0001, Sen Wu 0001, Jaideep Srivastava, Prasanna Kumar Desikan
Knowl. Based Syst.3
2014 A hazard based approach to user return time prediction
abstract
In the competitive environment of the internet, retaining and growing one's user base is of major concern to most web services. Furthermore, the economic model of many web services is allowing free access to most content, and generating revenue through advertising. This unique model requires securing user time on a site rather than the purchase of good which makes it crucially important to create new kinds of metrics and solutions for growth and retention efforts for web services. In this work, we address this problem by proposing a new retention metric for web services by concentrating on the rate of user return. We further apply predictive analysis to the proposed retention metric on a service, as a means for characterizing lost customers. Finally, we set up a simple yet effective framework to evaluate a multitude of factors that contribute to user return. Specifically, we define the problem of return time prediction for free web services. Our solution is based on the Cox's proportional hazard model from survival analysis. The hazard based approach offers several benefits including the ability to work with censored data, to model the dynamics in user return rates, and to easily incorporate different types of covariates in the model. We compare the performance of our hazard based model in predicting the user return time and in categorizing users into buckets based on their predicted return time, against several baseline regression and classification methods and find the hazard based approach to be superior.
Komal Kapoor, Mingxuan Sun 0001, Jaideep Srivastava
KDD3
2014 Scalable Information Flow Mining in Networks
Karthik Subbian, Chidananda Sridhar, Charu C. Aggarwal, Jaideep Srivastava
ECML/PKDD (3)4
2014 Behavioral data mining and network analysis in massive online games
abstract
The last decade has been characterized by an explosion of social media in a variety of forms. Since the data is captured in digital form it has become possible for the first time study human behavior at a massive scale. Not only is it possible to address traditional questions in the social sciences regarding collective dynamics of human behaviors but it is also possible to study new types of human behaviors which have arisen as a result of usage of new mediums like twitter, YouTube, Facebook, one games etc. Each of these mediums has its respective limitations and affordances. Out of all these mediums the most complex and data rich medium is that of Massive Online Games (MOGs). MOGs refer to massive online persistent environments (World of Warcraft, EVE Online, EverQuest etc) shared by millions of people . In general these environments are characterized by a rich array of activities and social interactions with a wide array of behaviors e.g., cooperation, trade, quest, deceit, mentoring etc. Such environments allow one to study human behavior at a level of granularity where it was not possible to do so previously. Given the challenges associated with analyzing this type of data traditional techniques in data mining and social network analysis have to be extended with insights from the social sciences. The tutorial will cover predictive and generative models in the study of MOGs. Additionally we will cover some SNA techniques which are more appropriate for MOGs given the multi-dimensionality of the data (P*/ERGM Models, IR Based Network Analysis, Hypergrah based Techniques, Coextensive Social Networks etc). We also describe the various ways in which MOGs exhibit similarities to the real world e.g., economic behaviors, clandestine behaviors, mentoring etc).
Muhammad Aurangzeb Ahmad, Jaideep Srivastava
WSDM2
2014 Topic formation and development: a core-group evolving process
Tieyun Qian, Qing Li 0001, Bing Liu 0001, Hui Xiong 0001, Jaideep Srivastava, Phillip C.-Y. Sheu
World Wide Web5
2014 Exploiting small world property for network clustering
Tieyun Qian, Qing Li 0001, Jaideep Srivastava, Zhiyong Peng 0001, Yang Yang 0009, Shuo Wang 0011
World Wide Web3
2013 Guilt by association?: network based propagation approaches for gold farmer detection
abstract
The term 'Gold Farmer' refers to a class of players in massive online games (MOGs) involved in a set of interrelated activities which are considered to be deviant activities. Consequently these gold farmers are actively banned by game administrators. The task of gold farmer detection is to identify gold farmers in a population of players but just like other clandestine actors they not labeled as such. In this paper the problem of extending the label of gold farmers to players which are not labeled as such is considered. Two main classes of techniques are described and evaluated: Network-based approaches and similarity based approaches. It is also explored how dividing the problem further by relabeling the data based on behavioral patterns can further improve the results
Muhammad Aurangzeb Ahmad, Brian Keegan, Atanu Roy, Dmitri Williams, Jaideep Srivastava, Noshir S. Contractor
ASONAM5
2013 TeamSkill and the NBA: applying lessons from virtual worlds to the real-world
abstract
In this paper, we build on our previous work by evaluating several approaches for assessing the skill of players and teams on the basis of both individual performance and group cohesion, or "team chemistry", using game data from the National Basketball Association (NBA). Previously developed for skill assessment in team-based multi-player video games (e.g., Halo 3), we find that group cohesion is a predictive feature in virtual and real-world team-based games, and that methods utilizing such features can often outperform the baseline in both contexts. Additionally, we observe a strong positive correlation between the predictive accuracy of our group cohesion-based approaches and the duration of playing time between a particular configuration of players on a team and their opponents, or "match-up" length.
Colin DeLong, Loren G. Terveen, Jaideep Srivastava
ASONAM3
2013 Socialization and trust formation: a mutual reinforcement? An exploratory analysis in an online virtual setting
abstract
Social interactions preceding and succeeding trust formation can be significant indicators of formation of trust in online social networks. In this research we analyze the social interaction trends that lead and follow formation of trust in these networks. This enables us to hypothesize novel theories responsible for explaining formation of trust in online social settings and provide key insights. We find that a certain level of socialization threshold needs to be met in order for trust to develop between two individuals. This threshold differs across persons and across networks. Once the trust relation has developed between a pair of characters connected by some social relation (also referred to as a character dyad), trust can be maintained with a lower rate of socialization.
Atanu Roy, Zoheb Borbora, Jaideep Srivastava
ASONAM3
2013 Dynamics of trust reciprocation in multi-relational networks
abstract
Understanding the dynamics of reciprocation is of great interest in sociology and computational social science. The recent growth of Massively Multi-player Online Games (MMOGs) has provided unprecedented access to large-scale data which enables us to study such complex human behavior in a more systematic manner. In this paper, we consider three different networks in the EverQuest2 game: chat, trade, and trust. The chat network has the highest level of reciprocation (33%) because there are essentially no barriers to it. The trade network has a lower rate of reciprocation (27%) because it has the obvious barrier of requiring goods or money for exchange; morever, there is no clear benefit to returning a trade link except in terms of social connections. The trust network has the lowest reciprocation (14%) because this equates to sharing certain within-game assets such as weapons, and so there is a high barrier for such connections In general, we observe that reciprocation rate is inversely related to the barrier level in these networks. We also note that reciprocation has connections across the heterogeneous networks. Our experiments indicate that players make use of the medium-barrier reciprocations to strengthen a relationship. We hypothesize that lower-barrier interactions are an important component to predicting higher-barrier ones. We verify our hypothesis using predictive models for trust reciprocations with features from trade interactions. Incorporating the number of trades (both before and after the initial trust link) boosts our ability to predict if the trust will be reciprocated up to 11% with respect to the AUC. More generally, we see strong correlations across the different networks and emphasize that network dynamics, such as reciprocation, cannot be studied in isolation on just a single type of connection.
Ayush Singhal, Karthik Subbian, Jaideep Srivastava, Tamara G. Kolda, Ali Pinar
ASONAM3
2013 Finding influencers in networks using social capital
abstract
The existing methods for finding influencers use the process of information diffusion to discover the nodes with maximum information spread. These models capture only the process of information diffusion and not the actual social value of collaborations in the network. We have proposed a method for finding influencers using the idea that people generate more value for their work by collaborating with peers of high influence. The social value generated through such collaborations denotes the notion of individual social capital. We hypothesize and show that players with high social capital are often key influencers in the network. We propose a value-allocation model to compute the social capital and allocate the fair share of this capital to each individual involved in the collaboration. We show that our allocation satisfies several axioms of fairness and falls in the same class as the Myerson's allocation function. We implement our allocation rule using an efficient algorithm SoCap and show that our algorithm outperforms the baselines in several real-life data sets. Specifically, in DBLP network, our algorithm outperforms PageRank, PMIA and Weighted Degree baselines up to 8% in terms of precision, recall and F1-measure.
Karthik Subbian, Dhruv Sharma, Jaideep Srivastava
ASONAM4
2013 Content-centric flow mining for influence analysis in social streams
abstract
The problem of discovering information flow trends and influencers in social networks has become increasingly relevant both because of the increasing amount of content available from online networks in the form of social streams, and because of its relevance as a tool for content trends analysis. An important part of this analysis is to determine the key patterns of flow and corresponding influencers in the underlying network. Almost all the work on influence analysis has focused on fixed models of the network structure, and edge-based transmission between nodes. In this paper, we propose a fully content-centered model of flow analysis in social network streams, in which the analysis is based on actual content transmissions in the network, rather than a static model of transmission on the edges. First, we introduce the problem of information flow mining in social streams, and then propose a novel algorithm InFlowMine to discover the information flow patterns in the network. We then leverage this approach to determine the key influencers in the network. Our approach is flexible, since it can also determine topic-specific influencers. We experimentally show the effectiveness and efficiency of our model.
Karthik Subbian, Charu C. Aggarwal, Jaideep Srivastava
CIKM3
2013 Measuring spontaneous devaluations in user preferences
abstract
Spontaneous devaluation in preferences is ubiquitous, where yesterday's hit is today's affliction. Despite technological advances facilitating access to a wide range of media commodities, finding engaging content is a major enterprise with few principled solutions. Systems tracking spontaneous devaluation in user preferences can allow prediction of the onset of boredom in users potentially catering to their changed needs. In this work, we study the music listening histories of Last.fm users focusing on the changes in their preferences based on their choices for different artists at different points in time. A hazard function, commonly used in statistics for survival analysis, is used to capture the rate at which a user returns to an artist as a function of exposure to the artist. The analysis provides the first evidence of spontaneous devaluation in preferences of music listeners. Better understanding of the temporal dynamics of this phenomenon can inform solutions to the similarity-diversity dilemma of recommender systems.
Komal Kapoor, Nisheeth Srivastava, Jaideep Srivastava, Paul Schrater
KDD3
2013 Community Detection with Prior Knowledge
abstract
The problem of community detection is a challenging one because of the presence of hubs and noisy links, which tend to create highly imbalanced graph clusters. Often, these resulting clusters are not very intuitive and difficult to interpret. With the growing availability of network information, there is a significant amount of prior knowledge available about the communities in social, communication and several other networks. These community labels may be noisy and very limited, though they do help in community detection. In this paper, we explore the use of such noisy labeled information for finding high quality communities. We will present an adaptive density-based clustering which allows flexible incorporation of prior knowledge in to the community detection process. We use a random walk framework to compute the node densities and the level of supervision regulates the node densities and the quality of resulting density based clusters. Our framework is general enough to produce both overlapping and non-overlapping clusters. We empirically show that even with a tiny amount of supervision, our approach can produce superior communities compared to popular baselines.
Karthik Subbian, Charu C. Aggarwal, Jaideep Srivastava, Philip S. Yu
SDM3
2013 Automating Document Annotation Using Open Source Knowledge
abstract
Annotating documents with relevant and comprehensive keywords offers invaluable assistance to the readers to quickly overview any document. The problem of document annotation is addressed in the literature under two broad classes of techniques namely, key phrase extraction and key phrase abstraction. In this paper, we propose a novel approach to generate summary phrases for research documents. Given the dynamic nature of scientific research, it has become important to incorporate new and popular scientific terminologies in document annotations. For this purpose, we have used crowd-source knowledge bases like Wikipedia and WikiCFP (a open source information source for call for papers) for automating key phrase generation. Also, we have taken into account the lack of availability of the document's content (due to protective policies) and developed a global context based key-phrase identification approach. We show that given only the title of a document, the proposed approach generates its global context information using academic search engines like Google Scholar. We evaluated the performance of the proposed approach on real-world dataset obtained from a computer science research document corpus. We quantitatively evaluated the performance of the proposed approach and compared it with two baseline approaches.
Ayush Singhal, Ravindra Kasturi, Jaideep Srivastava
Web Intelligence3
2013 Leveraging Web Intelligence for Finding Interesting Research Datasets
abstract
The problem of user's interest to item matching is at the core of recommendation systems and search engines. This problem is well studied in different contexts such as item, document, music and movie recommendations. For the purpose of recommendation these systems store the context or the meta-data information about the item of interest (e.g. user rating for books, tags, price etc). However, the general approaches for finding relevant items for recommendation cannot be directly applied in the case when the context or meta-data information about the item of interest is missing. In this paper we describe an algorithmic approach to handle this problem of missing context for items. In the proposed approach we have extended the context of user's interest and developed an unsupervised algorithm to find the items of interest for the user. Finally the items are ranked based on their relevance to the user's interest. We study this problem in the domain of dataset recommendation where the meta-data information about the datasets is missing due to lack of coherent and complete repository for the research datasets. We evaluate the performance of the proposed framework with real world dataset consisting of 20 user queries. We find that the proposed framework can recommend datasets for user queries with a recall of 90% in the top-4 recommendations. We also compared the performance of the dataset finding algorithm with the state of art supervised classification approach. We get a significant improvement of 36% using the proposed algorithm.
Ayush Singhal, Ravindra Kasturi, Vidyashankar Sivakumar, Jaideep Srivastava
Web Intelligence4
2012 Improved Feature Selection for Hematopoietic Cell Transplantation Outcome Prediction using Rank Aggregation
Chandrima Sarkar, Sarah Cooley, Jaideep Srivastava
FedCSIS3
2012 TeamSkill Evolved: Mixed Classification Schemes for Team-Based Multi-player Games
Colin DeLong, Jaideep Srivastava
PAKDD (1)2
2012 Improving bagging performance through multi-algorithm ensembles
Kuo-Wei Hsu, Jaideep Srivastava
Frontiers Comput. Sci.2
2011 Modeling Player Performance in Massively Multiplayer Online Role-Playing Games: The Effects of Diversity in Mentoring Network
abstract
This study investigates and reports preliminary findings on player performance prediction approaches which model player's past performance and social diversity in mentoring network in Ever Quest II, a popular massively multiplayer online role-playing game (MMORPG) developed by Sony Online Entertainment. Our contributions include a better understanding of performance metrics used in the game and a foundation of recommendation systems for mentors and apprentices. We examined three different game servers from the Ever Quest II game logs. In all three servers, the results from our analyses suggest that increase in social diversity in terms of characters and classes encountered moderately negatively correlates with player performance. Based on this finding, we built predictive models to predict player's future performance based on past performance and social diversity in terms of mentoring activities. Our results indicate that 1) models employing past performance and social diversity perform better and 2) prediction for mentors is generally better than that for apprentices.
Kyong Jin Shim, Kuo-Wei Hsu, Jaideep Srivastava
ASONAM3
2011 Effects of Mentoring on Player Performance in Massively Multiplayer Online Role-Playing Games (MMORPGs)
abstract
Massively Multiplayer Online Role-Playing Games (MMORPGs) have become increasingly popular and have communities comprising millions of subscribers. With their increasing popularity, researchers are realizing that video games can be a means to fully observe an entire isolated universe. In this study, we examine and report our findings on the effects of mentoring activities on player performance in Ever Quest II, a popular MMORPG developed by Sony Online Entertainment.
Kyong Jin Shim, Kuo-Wei Hsu, Jaideep Srivastava
ASONAM3
2011 Trust Amongst Rogues? A Hypergraph Approach for Comparing Clandestine Trust Networks in MMOGs
Muhammad Aurangzeb Ahmad, Brian Keegan, Dmitri Williams, Jaideep Srivastava, Noshir S. Contractor
ICWSM4
2011 Augmenting Image Processing with Social Tag Mining for Landmark Recognition
Amogh Mahapatra, Yonghong Tian 0001, Jaideep Srivastava
MMM (1)4
2011 TeamSkill: Modeling Team Chemistry in Online Multi-player Games
Colin DeLong, Nishith Pathak, Kendrick Erickson, Eric Perrino, Kyong Jin Shim, Jaideep Srivastava
PAKDD (2)6
2011 Mining temporal patterns in popularity of web items
Woong-Kee Loh, Sandeep Mane, Jaideep Srivastava
Inf. Sci.3
2010 GTPA: A Generative Model For Online Mentor-Apprentice Networks
abstract
There is a large body of work on the evolution of graphs in various domains, which shows that many real graphs evolve in a similar manner. In this paper we study a novel type of network formed by mentor-apprentice relationships in a massively multiplayer online role playing game. We observe that some of the static and dynamic laws which have been observed in many other real world networks are not observed in this network. Consequently well known graph generators like Preferential Attachment, Forest Fire, Butterfly, RTM, etc., cannot be applied to such mentoring networks. We propose a novel generative model to generate networks with the characteristics of mentoring networks.
Muhammad Aurangzeb Ahmad, David A. Huffaker, Jeffrey William Treem, Marshall Scott Poole, Jaideep Srivastava
AAAI6
2010 A Generalized Linear Threshold Model for Multiple Cascades
abstract
This paper presents a generalized version of the linear threshold model for simulating multiple cascades on a network while allowing nodes to switch between them. The proposed model is shown to be a rapidly mixing Markov chain and the corresponding steady state distribution is used to estimate highly likely states of the cascades' spread in the network. Results on a variety of real world networks demonstrate the high quality of the estimated solution.
Nishith Pathak, Arindam Banerjee 0001, Jaideep Srivastava
ICDM3
2010 Relationship between Diversity and Correlation in Multi-Classifier Systems
Kuo-Wei Hsu, Jaideep Srivastava
PAKDD (2)2
2010 Player Performance Prediction in Massively Multiplayer Online Role-Playing Games (MMORPGs)
Kyong Jin Shim, Richa Sharan, Jaideep Srivastava
PAKDD (2)3
2010 Distortion-free predictive streaming time-series matching
Woong-Kee Loh, Yang-Sae Moon, Jaideep Srivastava
Inf. Sci.3
2009 What's behind topic formation and development: a perspective of community core groups
abstract
Over the past several years, there has been a great interest in topic detection and tracking (TDT). Recently, analyzing general research trend from the huge amount of history documents also arouses considerable attention. However, existing work on TDT mainly focuses on overall trend analysis, and is unable to address questions such as "what determines the evolution of a topic?" and "when and how does a new topic get formed?".
Tieyun Qian, Qing Li 0001, Bing Liu 0001, Hui Xiong 0001, Jaideep Srivastava, Phillip C.-Y. Sheu
CIKM5
2009 Diversity in Combinations of Heterogeneous Classifiers
Kuo-Wei Hsu, Jaideep Srivastava
PAKDD2
2009 Simultaneously Finding Fundamental Articles and New Topics Using a Community Tracking Method
Tieyun Qian, Jaideep Srivastava, Zhiyong Peng 0001, Phillip C.-Y. Sheu
PAKDD2
2008 I/O Scalable Bregman Co-clustering
Kuo-Wei Hsu, Arindam Banerjee 0001, Jaideep Srivastava
PAKDD3
2008 Predicting trusts among users of online communities: an epinions case study
abstract
Trust between a pair of users is an important piece of information for users in an online community (such as electronic commerce websites and product review websites) where users may rely on trust information to make decisions. In this paper, we address the problem of predicting whether a user trusts another user. Most prior work infers unknown trust ratings from known trust ratings. The effectiveness of this approach depends on the connectivity of the known web of trust and can be quite poor when the connectivity is very sparse which is often the case in an online community. In this paper, we therefore propose a classification approach to address the trust prediction problem. We develop a taxonomy to obtain an extensive set of relevant features derived from user attributes and user interactions in an online community. As a test case, we apply the approach to data collected from Epinions, a large product review community that supports various types of interactions as well as a web of trust that can be used for training and evaluation. Empirical results show that the trust among users can be effectively predicted using pre-trained classifiers.
Haifeng Liu 0007, Ee-Peng Lim, Hady Wirawan Lauw, Minh-Tam Le, Aixin Sun, Jaideep Srivastava, Young Ae Kim
EC6
2007 Impact of social influence in e-commerce decision making
abstract
Purchasing decisions are often strongly influenced by people who the consumer knows and trusts. Moreover, many online shoppers tend to wait for the opinions of early adopters before making a purchase decision to reduce the risk of buying a new product. Web-based social communities, actively fostered by E-commerce companies, allow consumers to share their personal experiences by writing reviews, rating reviews, and chatting among trusting members. They drive the volume of traffic to retail sites and become a starting point for Web shoppers. E-commerce companies have recently started to capture data on the social interaction between consumers in their websites, with the potential objective of understanding and leveraging social influence in customers' purchase decision making to improve customer relationship management and increase sales. In this paper, we present an overview of the impact of social influence in E-commerce decision making to provide guidance to researchers and companies who have an interest in related issues. We identify how data about social influence can be captured from online customer behaviors and how social influence can be used by E-commerce websites to aid the user decision making process. We also provide a summary of technology for social network analysis and identify the research challenges of measuring and leveraging the impact of social influence on E-commerce decision making.
Young Ae Kim, Jaideep Srivastava
ICEC2
2007 PWave: A Multi-source Multi-sink Anycast Routing Framework for Wireless Sensor Networks
Zhi-Li Zhang, Jaideep Srivastava, Victor Firoiu
Networking3
2007 A probabilistic approach to modeling and estimating the QoS of web-services-based workflows
San-Yih Hwang, Haojun Wang, Jian Tang 0001, Jaideep Srivastava
Inf. Sci.4
2007 Tag-Splitting: Adaptive Collision Arbitration Protocols for RFID Tag Identification
abstract
Tag identification is an important tool in RFID systems with applications for monitoring and tracking. A RFID reader recognizes tags through communication over a shared wireless channel. When multiple tags transmit their IDs simultaneously, the tag-to-reader signals collide and this collision disturbs a reader's identification process. Therefore, tag collision arbitration for passive tags is a significant issue for fast identification. This paper presents two adaptive tag anticollision protocols: an adaptive query splitting protocol (AQS), which is an improvement on the query tree protocol, and an adaptive binary splitting protocol (ABS), which is based on the binary tree protocol and is a de facto standard for RFID anticollision protocols. To reduce collisions and identify tags efficiently, adaptive tag anticollision protocols use information obtained from the last process of tag identification. Our performance evaluation shows that AQS and ABS outperform other tree-based tag anticollision protocols
Jihoon Myung, Wonjun Lee 0001, Jaideep Srivastava, Timothy K. Shih
IEEE Trans. Parallel Distributed Syst.3
2006 Who Thinks Who Knows Who? Socio-cognitive Analysis of Email Networks
abstract
Interpersonal interaction plays an important role in organizational dynamics, and understanding these interaction networks is a key issue for any organization, since these can be tapped to facilitate various organizational processes. However, the approaches of collecting data about them using surveys/interviews are fraught with problems of scalability, logistics and reporting biases, especially since such surveys may be perceived to be intrusive. Widespread use of computer networks for organizational communication provides a unique opportunity to overcome these difficulties and automatically map the organizational networks with a high degree of detail and accuracy. This paper describes an effective and scalable approach for modeling organizational networks by tapping into an organization's email communication. The approach models communication between actors as non-stationary Bernoulli trials and Bayesian inference is used for estimating model parameters over time. This approach is useful for socio-cognitive analysis (who knows who knows who) of organizational communication networks. Using this approach, novel measures for analysis of (i) closeness between actors' perceptions about such organizational networks (agreement), (ii) divergence of an actor's perceptions about organizational network from reality (misperception) are explained. Using the Enron email data, we show that these techniques provide sociologists with a new tool to understand organizational networks.
Nishith Pathak, Sandeep Mane, Jaideep Srivastava
ICDM3
2006 Divide and conquer approach for efficient pagerank computation
abstract
PageRank is a popular ranking metric for large graphs such as theWorld Wide Web. Current research techniques for improving computational efficiency of PageRank have focussed on improving the I/O cost, convergence and parallelizing the computation process. In this paper, we propose a divide and conquer strategy for efficient computation of PageRank. The strategy is different from contemporary improvements in that itcan be combined with any existing enhancements to PageRank, giving way to an entire class of more efficient algorithms. Wepresent a novel graph-partitioning technique for dividing thegraph into subgraphs, on which computation can be performed independently. This approach has two significant benefits. Firstly, since the approach focuses on work-reduction, it can be combined with any existing enhancements to PageRank. Secondly, the proposed approach leads naturally into developing an incremental approach for computation of such ranking metrics given that these large graphs evolve over a period of time. The partitioning technique is both lossless and independent of the type (variant) ofPageRank computation algorithm used. The experimental results for a static single graph (graph at a single time instance) as well as for the incremental computation in case of evolving graphs, illustrate the utility of our novel partitioning approach. The proposed approach can also be applied for the computation of anyother metric based on first order Markov chain model.
Prasanna Kumar Desikan, Nishith Pathak, Jaideep Srivastava, Vipin Kumar 0001
ICWE3
2006 eSENSE: energy efficient stochastic sensing framework scheme for wireless sensor platforms
abstract
Energy is a precious resource in wireless sensor networks as sensor nodes are typically powered by batteries with high replacement cost. This paper presents eSENSE: an energy-efficient stochastic sensing framework for wireless sensor platforms. eSENSE is a node-level framework that utilizes knowledge of the underlying data streams as well as application data quality requirements to conserve energy on a sensor node. eSENSE employs a stochastic scheduling algorithm to dynamically control the operating modes of the sensor node components. This scheduling algorithm enables an adaptive sampling strategy that aggressively conserves power by adjusting sensing activity to the application requirements. Using experimental results obtained on Power-TOSSIM with a real-world data trace, we demonstrate that our approach reduces energy consumption by 29-36% while providing strong statistical guarantees on data quality.
Abhishek Chandra, Jaideep Srivastava
IPSN3
2005 Feature Mining for Prediction of Degree of Liver Fibrosis
Benjamin W. Mayer, Huzefa Rangwala, Rohit Gupta 0003, Jaideep Srivastava, George Karypis, Vipin Kumar 0001, Piet C. de Groen
AMIA4
2005 Spatial Clustering of Chimpanzee Locations for Neighborhood Identification
abstract
Since 1960, the chimpanzees (Pan troglodytes) of Gombe National Park, Tanzania, have been studied by behavioral ecologists, including Jane Goodall. Data have been collected for more than 40 years and are being analyzed by researchers in order to increase our understanding of the social structure of chimpanzees. In this paper, we consider the following question of interest to behavioral ecologists - "Does clustering exist among female chimpanzees in terms of their spatial locations ?" The analysis of this question will help behavioral ecologists to learn about the space use and the social interactions between female chimpanzees. The data collected for this analysis are marked spatial point patterns over the park. Current spatial clustering methods lack the ability to handle such marked point patterns directly. This paper presents a novel application of spatial point pattern analysis and data mining techniques to the ecological problem of clustering female chimpanzees. We found that Ripley's K-function provides a powerful statistical tool for evaluating clustering behavior among spatial point patterns. We then proposed two clustering approaches for marked point patterns using the K-function. Experimental results using the proposed clustering methods provide significant insight into the dynamics of female chimpanzee space use and into the overall social stucture of the species. In addition, the proposed methods can be extended to also include temporal information.
Sandeep Mane, Carson Murray, Shashi Shekhar 0001, Jaideep Srivastava, Anne Pusey
ICDM4
2005 Estimating missed actual positives using independent classifiers
abstract
Data mining is increasingly being applied in environments having very high rate of data generation like network intrusion detection [7], where routers generate about 300,000 -- 500,000 connections every minute. In such rare class data domains, the cost of missing a rare-class instance is much higher than that of other classes. However, the high cost for manual labeling of instances, the high rate at which data is collected as well as real-time response constraints do not always allow one to determine the actual classes for the collected unlabeled datasets. In our previous work [9], this problem of missed false negatives was explained in context of two different domains -- "network intrusion detection" and "business opportunity classification". In such cases, an estimate for the number of such missed high-cost, rare instances will aid in the evaluation of the performance of the modeling technique (e.g. classification) used. A capture-recapture method was used for estimating false negatives, using two or more learning methods (i.e. classifiers). This paper focuses on the dependence between the class labels assigned by such learners. We define the conditional independence for classifiers given a class label and show its relation to the conditional independence of the features sets (used by the classifiers) given a class label. The later is a computationally expensive problem and hence, a heuristic algorithm is proposed for obtaining conditionally independent (or less dependent) feature sets for the classifiers. Initial results of this algorithm on synthetic datasets are promising and further research is being pursued.
Sandeep Mane, Jaideep Srivastava, San-Yih Hwang
KDD2
2005 WICER: A Weighted Inter-Cluster Edge Ranking for Clustered Graphs
abstract
Several algorithms based on link analysis have been developed to measure the importance of nodes on a graph such as pages on the World Wide Web. PageRank and HITS are the most popular ranking algorithms to rank the nodes of any directed graph. But, both these algorithms assign equal importance to all the edges and nodes, ignoring the semantically rich information from nodes and edges. Therefore, in the case of a graph containing natural clusters, these algorithms do not differentiate between inter-cluster edges and intra-cluster edges. Based on this parameter, we propose a weighted inter-cluster edge ranking for clustered graphs that weighs edges (based on whether it is an inter-cluster or an intra-cluster edge) and nodes (based on the number of clusters it connects). We introduce a parameter '/spl alpha/' which can be adjusted depending on the bias desired in a clustered graph. Our experiments were two fold. We implemented our algorithm to a relationship set representing legal entities and documents and the results indicate the significance of the weighted edge approach. We also generated biased and random walks to quantitatively study the performance.
Divya Padmanabhan, Prasanna Kumar Desikan, Jaideep Srivastava, Kashif Riaz
Web Intelligence3
2005 QoS-Aware Admission Control and Dynamic Resource Provisioning Framework in Ubiquitous Multimedia Computing Environments
Wonjun Lee 0001, Jaideep Srivastava, Bikash Sabata
J. Supercomput.2
2004 A Probabilistic QoS Model and Computation Framework for Web Services-Based Workflows
San-Yih Hwang, Haojun Wang, Jaideep Srivastava, Raymond A. Paul
ER3
2004 Estimation of False Negatives in Classification
abstract
In many classification problems such as spam detection and network intrusion, a large number of unlabeled test instances are predicted negative by the classifier However, the high costs as well as time constraints on an expert's time prevent further analysis of the "predicted false" class instances in order to segregate the false negatives from the true negatives. A systematic method is thus required to obtain an estimate of the number of false negatives. A capture-recapture based method can be used to obtain an ML-estimate of false negatives when two or more independent classifiers are available. In the case for which independence does not hold, we can apply log-linear models to obtain an estimate of false negatives. However, as shown in this paper, lesser the dependencies among the classifiers, better is the estimate obtained for false negatives. Thus, ideally independent classifiers should be used to estimate the false negatives in an unlabeled dataset. Experimental results on the spam dataset from the UCI machine learning repository are presented.
Sandeep Mane, Jaideep Srivastava, San-Yih Hwang, Jamshid A. Vayghan
ICDM2
2004 Selecting the right objective measure for association analysis
Pang-Ning Tan, Vipin Kumar 0001, Jaideep Srivastava
Inf. Syst.3
2004 Blocking Reduction Strategies in Hierarchical Text Classification
abstract
One common approach in hierarchical text classification involves associating classifiers with nodes in the category tree and classifying text documents in a top-down manner. Classification methods using this top-down approach can scale well and cope with changes to the category trees. However, all these methods suffer from blocking which refers to documents wrongly rejected by the classifiers at higher-levels and cannot be passed to the classifiers at lower-levels. We propose a classifier-centric performance measure known as blocking factor to determine the extent of the blocking. Three methods are proposed to address the blocking problem, namely, threshold reduction, restricted voting, and extended multiplicative. Our experiments using support vector machine (SVM) classifiers on the Reuters collection have shown that they all could reduce blocking and improve the classification accuracy. Our experiments have also shown that the Restricted Voting method delivered the best performance.
Aixin Sun, Ee-Peng Lim, Wee Keong Ng, Jaideep Srivastava
IEEE Trans. Knowl. Data Eng.4
2003 A Comparative Study of Anomaly Detection Schemes in Network Intrusion Detection
abstract
Intrusion detection corresponds to a suite of techniques that are used to identify attacks against computers and network infrastructures. Anomaly detection is a key element of intrusion detection in which perturbations of normal behavior suggest the presence of intentionally or unintentionally induced attacks, faults, defects, etc. This paper focuses on a detailed comparative study of several anomaly detection schemes for identifying different network intrusions. Several existing supervised and unsupervised anomaly detection schemes and their variations are evaluated on the DARPA 1998 data set of network connections [9] as well as on real network data using existing standard evaluation techniques as well as using several specific metrics that are appropriate when detecting attacks that involve a large number of connections. Our experimental results indicate that some anomaly detection schemes appear very promising when detecting novel intrusions in both DARPA'98 data and real network data.
Aleksandar Lazarevic, Levent Ertöz, Vipin Kumar 0001, Aysel Ozgur, Jaideep Srivastava
SDM5
2003 Web Business Intelligence: Mining the Web for Actionable Knowledge
abstract
It is estimated that over seven billion static pages exist in the Web today, and backend databases can potentially produce at least three times as many dynamic pages. However, the best search engines index only approximately 20% of the static pages. So the real question is: While the Web is certainly the most amazing and comprehensive information source ever created, are you really getting all the information you need for your specific purpose? The answer to this question is mostly “yes” for the individual user, who uses the Web as an information source for casual purposes. However, for an individual who uses the Web as an essential and comprehensive source of information—for business or research—the answer is quite the opposite. Even a sophisticated Web user requires a significant amount of time and effort to find all of the information needed for a given task. In this paper the concept of Web Business Intelligence (WBI) is introduced, an emerging class of software that leverages the unprecedented content on the Web to extract actionable knowledge in an organizational setting. The contributions include an architecture for WBI, a survey of technologies relevant to the various components of the architecture, and illustration of the value of WBI by means of a detailed example from the e-finance domain. This article concludes with a discussion on the future of WBI.
Jaideep Srivastava, Robert Cooley
INFORMS J. Comput.1
2002 Selecting the right interestingness measure for association patterns
abstract
Many techniques for association rule mining and feature selection require a suitable metric to capture the dependencies among variables in a data set. For example, metrics such as support, confidence, lift, correlation, and collective strength are often 'used to determine the interestingness of association patterns. However, many such measures provide conflicting information about the interestingness of a pattern, and the best metric to use for a given application domain is rarely known. In this paper, we present an overview of various measures proposed in the statistics, machine learning and data mining literature. We describe several key properties one should examine in order to select the right measure for a given application domain. A comparative study of these properties is made using twenty one of the existing measures. We show that each measure has different properties which make them useful for some application domains, but not for others. We also present two scenarios in which most of the existing measures agree with each other, namely, support-based pruning and table standardization. Finally, we present an algorithm to select a small set of tables such that an expert can select a desirable measure by looking at just this small set of tables.
Pang-Ning Tan, Vipin Kumar 0001, Jaideep Srivastava
KDD3
2002 A Case for Analytical Customer Relationship Management
Jaideep Srivastava, Jau-Hwang Wang, Ee-Peng Lim, San-Yih Hwang
PAKDD1
2002 Web Mining
Ron Kohavi, Brij M. Masand, Myra Spiliopoulou, Jaideep Srivastava
Data Min. Knowl. Discov.4
2002 QoS-adaptive bandwidth scheduling in continuous media streaming
Wonjun Lee 0001, Jaideep Srivastava, Hojung Cha
Inf. Softw. Technol.2
2002 Efficient Aggregation Algorithms for Compressed Data Warehouses
abstract
Aggregation and cube are important operations for online analytical processing (OLAP). Many efficient algorithms to compute aggregation and cube for relational OLAP have been developed. Some work has been done on efficiently computing cube for multidimensional data warehouses that store data sets in multidimensional arrays rather than in tables. However, to our knowledge, there is nothing to date in the literature describing aggregation algorithms on compressed data warehouses for multidimensional OLAP. This paper presents a set of aggregation algorithms on compressed data warehouses for multidimensional OLAP. These algorithms operate directly on compressed data sets, which are compressed by the mapping-complete compression methods, without the need to first decompress them. The algorithms have different performance behaviors as a function of the data set parameters, sizes of outputs and main memory availability. The algorithms are described and the I/O and CPU cost functions are presented in this paper. A decision procedure to select the most efficient algorithm for a given aggregation request is also proposed. The analysis and experimental results show that the algorithms have better performance on sparse data than the previous aggregation algorithms.
Jianzhong Li 0001, Jaideep Srivastava
IEEE Trans. Knowl. Data Eng.2
2002 Error spreading: a perception-driven approach to handling error in continuous media streaming
abstract
With the growing popularity of the Internet, there is increasing interest in using it for audio and video transmission. Perceptual studies of audio and video viewing have shown that viewers find bursty losses, mostly caused by congestion, to be the most annoying disturbance, and hence these are critical issues to be addressed for continuous media streaming applications. Classical error handling techniques have mostly been geared toward ensuring that the transmission is correct, with no attention to timeliness. For isochronous traffic like audio and video, timeliness is a key criterion, and given the high degree of content redundancy, some loss of content is quite acceptable. We introduce the concept of error spreading, which is a transformation technique that permutes the input sequence of packets (from a continuous stream of data) before transmission. The packets are unscrambled at the receiving end. The transformation is designed to ensure that bursty losses in the transformed domain get spread all over the sequence in the original domain, thus improving the perceptual quality of the stream. Our error spreading idea deals with both cases where the stream has or does not have inter-frame dependencies. We next describe a continuous media transmission protocol and experimentally validate its performance based on this idea. We also show that our protocol can be used complementary to other error handling protocols.
Srivatsan Varadarajan, Hung Q. Ngo 0001, Jaideep Srivastava
IEEE/ACM Trans. Netw.3
2001 An Algebraic QoS-Based Resource Allocation Model for Competitive Multimedia Applications
Wonjun Lee 0001, Jaideep Srivastava
Multim. Tools Appl.2
2001 Normal forms and syntactic completeness proofs for functional independencies
Duminda Wijesekera, M. Ganesh 0001, Jaideep Srivastava, Anil Nerode
Theor. Comput. Sci.3
2000 A Market-based Resource Management and QoS Support Framework for Distributed Multimedia Systems
abstract
The last decade has seen an explosive growth in multimedia applications and considerable research in related technologies, with special emphasis on Quality of Service (QoS) requirements, such as timeliness, precision, and accuracy [12]. Meeting QoS guarantees in a distributed real-time multimedia systems is an end-to-end issue because the users are interested in the end results � i.e., from application to application. Allocating proper resources to the applications with respect to QoS pro visioning to maximize the total system utilization or bene t, is a fundamental probleminallmultimedia systems today. The term utility or bene t may take on the meaning of usefulness, satisfaction, or of pleasure, depending on the context.
Wonjun Lee 0001, Jaideep Srivastava
CIKM2
2000 Reserve-based disk admission control and bandwidth scheduling strategy for continuous media servers
abstract
In this paper we present a novel admission control algorithm that exploits the degradability property of continuous media applications to improve the performance of the system. The algorithm is based on setting aside a portion of the resources as reserves and managing it intelligently so that the total utility of the system is enhanced. This reserve-based admission control strategy (RAC) is a compromise between purely greedy and non-greedy strategy, and it leads to an efficient protocol that improves the performance of the system. While the protocol is simple for admission decision, it also results in better performance for the system by reserving some resources for important future applications.
Wonjun Lee 0001, Jaideep Srivastava
ICCCN2
2000 An Adaptive, Perception-Driven Error Spreading Scheme in Continuous Media Streaming
abstract
For transmission of continuous media (CM) streams such as audio and video over the Internet, a critical issue is that periodic network overloads cause bursty packet losses. Studies on human perception of audio and video streams have shown that bursty losses have the most annoying affects. Hence, addressing this issue is critical for multimedia applications such as Internet telephony, videoconferencing, distance learning, etc. Classical error handling schemes like retransmission and forward error recovery have the undesirable effects of (a) introducing timing variations, which is unacceptable for isochronous traffic, and (b) using up valuable bandwidth, potentially exacerbating the problem. This paper introduces a new concept called error spreading, which is a transformation technique that takes an input sequence of packets (from an audio or video stream) and permutes them before transmission. The packets are then un-permuted at the receiver before delivery to the application. The purpose is to spread out bursty network errors, in order to achieve better perceptual quality of the transmitted stream. Analysis has been done to determine the provable lower bound on bursty errors tolerable by this class of protocols. An algorithm to generate the optimal permutation for a given network loss rate is presented. While our previous work had focused on streams with no inter-frame dependencies, e.g. MJPEG encoded video, in this paper the technique is generalized to streams with inter-frame dependencies, e.g. MPEG encoded video.
Srivatsan Varadarajan, Hung Q. Ngo 0001, Jaideep Srivastava
ICDCS3
2000 Web mining for e-commerce (workshop session - title only)
abstract
No abstract available.
Ron Kohavi, Myra Spiliopoulou, Jaideep Srivastava
KDD3
2000 Indirect Association: Mining Higher Order Dependencies in Data
Pang-Ning Tan, Vipin Kumar 0001, Jaideep Srivastava
PKDD3
2000 QoS-based evaluation of file systems and distributed system services for continuous media provisioning
Wonjun Lee 0001, Difu Su, Jaideep Srivastava
Inf. Softw. Technol.3
2000 Efficient Aggregation Algorithms on Very Large Compressed Data Warehouses
Jianzhong Li 0001, Yingshu Li 0001, Jaideep Srivastava
J. Comput. Sci. Technol.3
2000 SMDP: Minimizing Buffer Requirements for Continuous Media Servers
Youjip Won, Jaideep Srivastava
Multim. Syst.2
1999 High Performance Data Mining
Vipin Kumar 0001, Jaideep Srivastava
HiPC2
1999 Adaptive Disk Scheduling Algorithms for Video Servers
abstract
Soft-real time applications, such as continuous media (CM) systems, have an important property, namely, they allow for graceful adaptation of the application Quality-of-Service (QoS), and therefore are able to have acceptable performance with reduced resource utilization. This can be used by the admission control process to decide if an application can be admitted, even if the resource is congested. In this paper, we present a Soft-QoS framework for Continuous Media servers, which provides a dynamic and adaptive admission control and scheduling algorithm. Using our policy, we could increase the number of simultaneously running clients that could be supported and could ensure a good response ratio and better resource utilization under heavy traffic requirements. The observations and findings from the model are validated with simulation studies.
Wonjun Lee 0001, Jaideep Srivastava, Won-Ho Lee
ICPP2
1999 Event Detection from Time Series Data
abstract
Article Event detection from time series data Share on Authors: Valery Guralnik Department of Computer Science, University of Minnesota Department of Computer Science, University of MinnesotaView Profile , Jaideep Srivastava Department of Computer Science, University of Minnesota Department of Computer Science, University of MinnesotaView Profile Authors Info & Claims KDD '99: Proceedings of the fifth ACM SIGKDD international conference on Knowledge discovery and data miningAugust 1999 Pages 33–42https://doi.org/10.1145/312129.312190Online:01 August 1999Publication History 284citation6,313DownloadsMetricsTotal Citations284Total Downloads6,313Last 12 Months264Last 6 weeks21 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Valery Guralnik, Jaideep Srivastava
KDD2
1999 Aggregation Algorithms for Very Large Compressed Data Warehouses
Jianzhong Li 0001, Doron Rotem, Jaideep Srivastava
VLDB3
1999 First 20 Precision Among World Web Search Services (Search Engines)
abstract
Five search engines, Alta Vista, Excite, HotBot, Infoseek, and Lycos, are compared for precision on the first 20 results returned for 15 queries, adding weight for ranking effectiveness. All searching was done from January 31 to March 12, 1997. In the study, steps are taken to ensure that bias has not unduly influenced the evaluation. Friedmann's randomized block design is used to perform multiple comparisons for significance. Analysis shows that Alta Vista, Excite and Infoseek are the top three services, with their relative rank changing depending on how one operationally defines the concept of relevance. Correspondence analysis shows that Lycos performed better on short, unstructured queries, whereas HotBot performed better on structured queries.
H. Vernon Leighton, Jaideep Srivastava
J. Am. Soc. Inf. Sci.2
1999 Data Preparation for Mining World Wide Web Browsing Patterns
Robert Cooley, Bamshad Mobasher, Jaideep Srivastava
Knowl. Inf. Syst.3
1999 Experimental Evaluation of Loss Perception in Continuous Media
Duminda Wijesekera, Jaideep Srivastava, Anil Nerode, Mark Foresti
Multim. Syst.2
1999 Strategic Replication of Video Files in a Distributed Environment
Youjip Won, Jaideep Srivastava
Multim. Tools Appl.2
1999 Warehouse Creation - A Potential Roadblock to Data Warehousing
abstract
Data warehousing is gaining in popularity as organizations realize the benefits of being able to perform sophisticated analyses of their data. Recent years have seen the introduction of a number of data-warehousing engines, from both established database vendors as well as new players. The engines themselves are relatively easy to use and come with a good set of end-user tools. However, there is one key stumbling block to the rapid development of data warehouses, namely that of warehouse population. Specifically, problems arise in populating a warehouse with existing data since it has various types of heterogeneity. Given the lack of good tools, this task has generally been performed by various system integrators, e.g., software consulting organizations which have developed in-house tools and processes for the task. The general conclusion is that the task has proven to be labor-intensive, error-prone, and generally frustrating, leading a number of warehousing projects to be abandoned mid-way through development. However, the picture is not as grim as it appears. The problems that are being encountered in warehouse creation are very similar to those encountered in data integration, and they have been studied for about two decades. However, not all problems relevant to warehouse creation have been solved, and a number of research issues remain. The principal goal of this paper is to identify the common issues in data integration and data-warehouse creation.
Jaideep Srivastava, Ping-Yao Chen
IEEE Trans. Knowl. Data Eng.1
1998 Pattern Directed Mining of Sequence Data
Valery Guralnik, Duminda Wijesekera, Jaideep Srivastava
KDD3
1998 CORBA Evaluation of Video Streaming wrt QoS Provisioning
abstract
Describes the design, implementation and evaluation of CORBA- and socket-based continuous media (CM) systems. TCP/IP is not suitable for distributed applications which require high network bandwidth and timing criticality. UDP/IP is one of the alternatives. However, due to the fact that UDP is a lossy protocol, many issues arise when implementing distributed CM applications. Most of the QoS (quality of service) metrics known so far assume that the communication channel is lossless. In this paper, since we use UDP for CM data transmission, we adopt a new QoS metric that is applicable to lossy streams to evaluate the performance of our CM server. To reduce QoS loss factors and drift factors, we adopt a new strategy, called the QoS-driven dropping mechanism, for the CM server. Besides the traditional C-socket (TCP-UDP/IP) based CM server mechanisms, we implemented our CM server on CORBA. It turns out that the CORBA-based implementation runs considerably slower than the UDP-version, but faster than the TCP version.
Wonjun Lee 0001, Jaideep Srivastava
SRDS2
1997 Experimental Evaluation of PFS Continuous Media File System
abstract
Article Experimental evaluation of PFS continuous media file system Share on Authors: Wonjun Lee Department of Computer Science, University of Minnesota, Minneapolis, MN Department of Computer Science, University of Minnesota, Minneapolis, MNView Profile , Difu Su Department of Computer Science, University of Minnesota, Minneapolis, MN Department of Computer Science, University of Minnesota, Minneapolis, MNView Profile , Duminda Wijesekera Department of Computer Science, University of Minnesota, Minneapolis, MN Department of Computer Science, University of Minnesota, Minneapolis, MNView Profile , Jaideep Srivastava Department of Computer Science, University of Minnesota, Minneapolis, MN Department of Computer Science, University of Minnesota, Minneapolis, MNView Profile , Deepak Kenchammana-Hosekote IBM Almaden Research Center, San Hose CA IBM Almaden Research Center, San Hose CAView Profile , Mark Foresti Rome Laboratory, Griffs Air Force Base, Rome NY Rome Laboratory, Griffs Air Force Base, Rome NYView Profile Authors Info & Claims CIKM '97: Proceedings of the sixth international conference on Information and knowledge managementJanuary 1997 Pages 246–253https://doi.org/10.1145/266714.266905Published:01 January 1997 8citation342DownloadsMetricsTotal Citations8Total Downloads342Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Wonjun Lee 0001, Difu Su, Duminda Wijesekera, Jaideep Srivastava, Deepak R. Kenchammana-Hosekote, Mark Foresti
CIKM4
1997 Distributed Service Paradigm for Remote Video Retrieval Request
abstract
The per service cost have been serious impediment to wide spread usage of on-line digital continuous media service, especially in the entertainment arena. Although handling the continuous media may be achievable due to the technology advances in past few years, its competitiveness in the market with the existing service type such as video rental is still in question. In this paper, we propose a service paradigm for continuous media delivery in a distributed infrastructure in an effort to reduce the resource requirement to support a set of service requests. The storage resource and network resource to support a set of requests should be properly quantified to a uniform metric to measure the efficiency of the service schedule. We developed a cost model which maps the given service schedule to a quantity. The proposed cost model is used to capture the amortized resource requirement of the schedule and thus to measure the efficiency of the schedule. We develop a scheduling algorithm which strategically replicates the requested continuous media files at the various intermediate storages.
Youjip Won, Jaideep Srivastava
HPDC2
1997 A multimedia programming toolkit/environment
abstract
This paper provides details and implementation experiences of a multimedia programming language and associated toolkits. The language, a data-flow paradigm for multimedia streams, consists of blocks of code that can be connected through their data ports. Continuous media flows through these ports into and out of blocks. The blocks are responsible for the processing of continuous media data. Examples of such processing include capturing, displaying, storing, retrieving and analyzing their contents. The blocks also have parameter ports that specify other pertinent parameters, such as location, and display characteristics such as geometry, etc. The connection topology of blocks is specified using a graphical editor called the Program Development Tool (PDT) and the geometric parameters are specified by using another graphical editor called the User Interface Development Tool (UIDT). Experience with modeling multimedia presentations in our environment and the enhancements provided by the two graphical editors are discussed in detail.
Raja Harinath, Wonjun Lee 0001, Shwetal S. Parikh, Difu Su, Sunil Wadhwa, Duminda Wijesekera, Jaideep Srivastava, Deepak R. Kenchammana-Hosekote
ICPADS7
1997 Web Mining: Information and Pattern Discovery on the World Wide Web
abstract
Application of data mining techniques to the World Wide Web, referred to as Web mining, has been the focus of several recent research projects and papers. However, there is no established vocabulary, leading to confusion when comparing research efforts. The term Web mining has been used in two distinct ways. The first, called Web content mining in this paper, is the process of information discovery from sources across the World Wide Web. The second, called Web usage mining, is the process of mining for user browsing and access patterns. We define Web mining and present an overview of the various research issues, techniques, and development efforts. We briefly describe WEBMINER, a system for Web usage mining, and conclude the paper by listing research issues.
Robert Cooley, Bamshad Mobasher, Jaideep Srivastava
ICTAI3
1997 Tableaux for Functional Dependencies and Independencies
Duminda Wijesekera, M. Ganesh 0001, Jaideep Srivastava, Anil Nerode
TABLEAUX3
1997 I/O Scheduling for Digital Continuous Media
Deepak R. Kenchammana-Hosekote, Jaideep Srivastava
Multim. Syst.2
1996 Mining Entity-Identification Rules for Database Integration
M. Ganesh 0001, Jaideep Srivastava, Travis Richardson
KDD2
1996 Entity Identification in Database Integration
Ee-Peng Lim, Jaideep Srivastava, Satya Prabhakar, James Richardson
Inf. Sci.2
1996 Quality of Service (QoS) Metrics for Continuous Media
Duminda Wijesekera, Jaideep Srivastava
Multim. Tools Appl.2
1996 An Evidential Reasoning Approach to Attribute Value Conflict Resolution in Database Integration
abstract
Resolving domain incompatibility among independently developed databases often involves uncertain information. DeMichiel (1989) showed that uncertain information can be generated by the mapping of conflicting attributes to a common domain, based on some domain knowledge. We show that uncertain information can also arise when the database integration process requires information not directly represented in the component databases, but can be obtained through some summary of data. We therefore propose an extended relational model based on Dempster-Shafer theory of evidence to incorporate such uncertain knowledge about the source databases. The extended relation uses evidence sets to represent uncertainty in information, which allow probabilities to be attached to subsets of possible domain values. We also develop a full set of extended relational operations over the extended relations. In particular, an extended union operation has been formalized to combine two extended relations using Dempster's rule of combination. The closure and boundedness properties of our proposed extended operations are formulated. We also illustrate the use of extended operations by some query examples.
Ee-Peng Lim, Jaideep Srivastava, Shashi Shekhar 0001
IEEE Trans. Knowl. Data Eng.2
1995 An Algebraic Transformation Framework for Multidatabase Queries
Ee-Peng Lim, Jaideep Srivastava, San-Yih Hwang
Distributed Parallel Databases2
1995 Myriad: Design and Implementation of a Federated Database Prototype
abstract
Abstract A key problem in providing ‘enterprise‐wide’ information is the integration of databases that have been independently developed. An important requirement is to accommodate heterogeneity and maintain the autonomy of component databases. Myriad is a federated database prototype developed at the University of Minnesota, to provide a testbed for investigating alternatives in architecture and algorithms for database integration, query processing and optimization, and concurrency control and recovery. The system incorporates our group's research results in these areas. This paper describes our experiences in the design and implementation of Myriad, and in the project management. Special emphasis is given to discussing design alternatives and their impact on Myriad. This paper also presents the software engineering principles and the project management techniques we used in developing Myriad and the lessons we learned. We believe these lessons would be useful for practitioners who wish to develop a similar system. Handling heterogeneity and autonomy were prime objectives throughout the prototyping effort. We are convinced that a prototype federated database is an important infrastructural requirement for the overall goal of ‘enterprise‐integration’, and believe Myriad to be a significant contribution towards this.
Ee-Peng Lim, San-Yih Hwang, Jaideep Srivastava, Dave Clements, M. Ganesh 0001
Softw. Pract. Exp.3
1994 Performance Evaluation of Grid Based Multi-Attibute Record Declustering Methods
abstract
We focus on multi-attribute declustering methods which are based on some type of grid-based partitioning of the data space. Theoretical results are derived which show that no declustering method can be strictly optimal for range queries if the number of disks is greater than 5. A detailed performance evaluation is carried out to see how various declustering schemes perform under a wide range of query and database scenarios (both relative to each other and to the optimal). Parameters that are varied include shape and size of queries, database size, number of attributes and the number of disks. The results show that information about common queries on a relation is very important and ought to be used in deciding the declustering for it, and that this is especially crucial for small queries. Also, there is no clear winner, and as such parallel database systems must support a number of declustering methods.>
Bhaskar Himatsingka, Jaideep Srivastava
ICDE2
1994 Resolving Attribute Incompatibility in Database Integration: An Evidential Reasoning Approach
abstract
Resolving domain incompatibility among independently developed databases often involves uncertain information. DeMichiel (1989) showed that uncertain information can be generated by the mapping of conflicting attributes to a common domain, based on some domain knowledge. The authors show that uncertain information can also arise when the database integration process requires information not directly represented in the component databases, but can be obtained through some summary of data. They therefore propose an extended relational model based on Dempster-Shafer theory of evidence (1976) to incorporate such uncertain knowledge about the source databases. They also develop a full set of extended relational operations over the extended relations. In particular, an extended union operation has been formalized to combine two extended relations using Dempster's rule of combination. The closure and boundedness properties of the proposed extended operations are formulated.>
Ee-Peng Lim, Jaideep Srivastava, Shashi Shekhar 0001
ICDE2
1994 The MYRIAD Federated Database Prototype
abstract
No abstract available.
San-Yih Hwang, Ee-Peng Lim, H.-R. Yang, S. Musukula, K. Mediratta, M. Ganesh 0001, Dave Clements, J. Stenoien, Jaideep Srivastava
SIGMOD Conference9
1994 Transaction Recovery in Federated Autonomous Databases
San-Yih Hwang, Jaideep Srivastava, Jianzhong Li 0001
Distributed Parallel Databases2
1993 Multiple Query Optimization with Depth-First Branch-and-Bound and Dynamic Query Ordering
abstract
Ahmet Cosar, Ee-Peng Lim, Jaideep SrivastavaDepartmentof Computer ScienceUniversity of MinnesotaMinneapolis, MN 55455AbstractIn certain database applications such as deductivedatabases, batch query processing, and recursive queryprocessing etc., a single query can be transformed into aset ofclosely related database queries. Great benefits canbe obtained by executing a group of related queries all to-gether in a single unijied multi-plan instead of executingeach query separately. In order to achieve this, MultipleQuery Optimization (MQO) identifies common task(s)(e.g. common subezpressions, joins, etc.) among a setof query plans and creates a single unified plan (multi-plan) which can be executed to obtain the required out-puts forall queries at once. In this paper, anew heuris-tic function (f=), dynamic query ordering heuristics,and Depth-First Branch-and-Bound (DFBB) are de-jined and experimentally evaluated, and compared withexisting methods which use A* and static query order-ing. Our experiments show that all three of f., DFBB,and dynamic query ordering help to improve the perfor-mance of our h4Q0 algorithm.1 IntroductionThe objective of multiple query optimization (MQO)is to exploit the benefits of sharing common tasks inthe access plans for a group of queries. In certaindatabase applications, e.g. deductive query processing,batch query processing and recursive query processing,often a group of queries are submitted together to theDBMS for execution. The traditional approach of pro-cessing queries one at a time will be inefficient espe-cially when there is a high number of queries sharingPermission to copy without fee all or part of this materisi isgmnted provided that the copies m. not mad. or distributed fordirect commercial advantage, tha ACM copyright notica and thatitla of tha publication and ita data appaar, and
Ahmet Cosar, Ee-Peng Lim, Jaideep Srivastava
CIKM3
1993 Concurrency Control in Federated Databases: A Dynamic Approach
abstract
Article Free Access Share on Concurrency control in federated databases: a dynamic approach Authors: San-Yih Hwang Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MN Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MNView Profile , Jiandong Huang Sensor and System Development Center, Honeywell, 3660 Technology Drive, Minneapolis, MN Sensor and System Development Center, Honeywell, 3660 Technology Drive, Minneapolis, MNView Profile , Jaideep Srivastava Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MN Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MNView Profile Authors Info & Claims CIKM '93: Proceedings of the second international conference on Information and knowledge managementDecember 1993Pages 694–703https://doi.org/10.1145/170088.170458Published:01 December 1993Publication History 3citation370DownloadsMetricsTotal Citations3Total Downloads370Last 12 Months8Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
San-Yih Hwang, Jiandong Huang, Jaideep Srivastava
CIKM3
1993 Query Optimization and Processing in Federated Database Systems
abstract
this paper, we have selected a minimal set of core operations that includes the set of relational operations, i.e. foe,ß,1,\\Gamma,[g, as well as three other operators which are useful in specifying the federated views. These are 2-way Outerjoin Operator (
Ee-Peng Lim, Jaideep Srivastava
CIKM2
1993 Entity Identification in Database Integration
abstract
The objective of entity identification is to determine the correspondence between object instances from more than one database. Entity identification at the instance level, assuming that schema level heterogeneity has been resolved a priori, is examined. Soundness and completeness are defined as the desired properties of any entity identification technique. To achieve soundness, a set of identity and distinctness rules are established for entities in the integrated world. The use of extended key, which is the union of keys, and possibly other attributes, from the relations to be matched, and its corresponding identify rule are proposed to determine the equivalence between tuples from relations which may not share any common key. Instance level functional dependencies (ILFD), a form of semantic constraint information about the real-world entities, are used to derive the missing extended key attribute values of a tuple.>
Ee-Peng Lim, Jaideep Srivastava, Satya Prabhakar, James Richardson
ICDE2
1993 Algorithms for Loading Parallel Grid Files
abstract
The paper describes three fast loading algorithms for grid files on a parallel shared nothing architecture. The algorithms use dynamic programming and sampling to effectively partition the data file among the processors to achieve maximum parallelism in answering range queries. Each processor then constructs in parallel its own portion of the grid file. Analytical results and simulations are given for the three algorithms.
Jianzhong Li 0001, Doron Rotem, Jaideep Srivastava
SIGMOD Conference3
1993 Database Concurrency Control in Multilevel Secure Database Management Systems
abstract
Concurrent execution of transactions in database management systems (DBMSs) may lead to contention for access to data, which in a multilevel secure DBMS (MLS/DBMS) may lead to insecurity. Security issues involved in database concurrency control for MLS/DBMSs are examined, and it is shown how a scheduler can affect security. Data conflict security, (DC-security), a property that implies a system is free of covert channels due to contention for access to data, is introduced. A definition of DC-security based on noninterference is presented. Two properties that constitute a necessary condition for DC-security are introduced along with two simpler necessary conditions. A class of schedulers called output-state-equivalent is identified for which another criterion implies DC-security. The criterion considers separately the behavior of the scheduler in response to those inputs that cause rollback and those that do not. The security properties of several existing scheduling protocols are characterized. Many are found to be insecure.>
Thomas F. Keefe, Wei-Tek Tsai, Jaideep Srivastava
IEEE Trans. Knowl. Data Eng.3
1992 A High Dimensional Array Assignment Method for Parallel Computing Systems
Jianzhong Li 0001, Jaideep Srivastava
ICPP (2)2
1992 An Analytical Performance Model of Parallel Production Systems
abstract
An analytical model for parallel production systems is developed, where rule firing is modeled as a transaction. Both resource contention and data contention are modeled in detail, and the performance of locking-based and optimistic approaches is analyzed. It is shown that significant speedup can be gained in parallel rule execution. The main contribution is the insight into parallel rule firing provided by the parametric model.>
Jau-Hwang Wang, Jaideep Srivastava
ICTAI2
1992 CMD: A Multidimensional Declustering Method for Parallel Data Systems
Jianzhong Li 0001, Jaideep Srivastava, Doron Rotem
VLDB2
1992 A formal model of trade-off between optimization and execution costs in semantic query optimization
Shashi Shekhar 0001, Jaideep Srivastava, Soumitra Dutta
Data Knowl. Eng.2
1991 Conditional transactions: a model of computation for active databases
abstract
A transaction model for active databases is introduced. Concurrent execution of rules needs to be carefully managed to ensure consistent semantics. It is shown that serializability is not always sufficient for correctness. A new criterion, conditional conflict serializability (CCS) is developed and shown to ensure the desired correctness. A graph-based scheduler for it is presented. Practical schedulers should also be recoverable, which the graph-based scheduler is not. The authors prove that conventional two-phase locking also achieves CCS, and can thus be used in practice.>
Jaideep Srivastava, Kuo-Wei Hwang, Wei-Tek Tsai
COMPSAC1
1991 Experiences with IFS: a distributed image filing system
abstract
The authors have designed and implemented a distributed image filing system. Users sitting at workstations can scan, store and retrieve images stored on optical disk at an image server. The system is built as a set of client and server processes performing specialized tasks. The system is network independent, in the sense that it runs on any of the local networks currently available on the market, and location transparent, as any mapping between processes and machines is allowed, including the extreme case of a standalone system where all processes run on the same machine. It is believed that the system is interesting from the perspective of designing a network independent distributed system, and also from a software engineering point of view, as its modular design permits simpler development, debugging and testing, as well as future extensibility.>
Gabriel Broner, Jaideep Srivastava
LCN2
1991 Enhancing Real-Time DBMS Performance with Multiversion Data and Priority Based Disk Scheduling
abstract
Real-time multiversion concurrency control algorithms are proposed, to: increase concurrency, adjust the serialization order dynamically and work without an estimate of a transaction's runtime. The authors also propose disk scheduling algorithms which consider not only the transactions which request input/output (I/O) but also those affected by I/O. They consider transactions which are directly affected when priorities of I/O requests are assigned, in addition to transactions which generates these requests. The real-time disk resident database system model and the multiversion concurrency control algorithms are described. The real-time disk scheduling algorithms and their properties are also described. Real-time multiversion concurrency control and disk scheduling algorithms are shown to decrease miss ratio significantly.>
Jaideep Srivastava
RTSS2
1990 Real-time scheduling of multiple segment tasks
abstract
The authors study the problem of on-line non-preemptive scheduling of multiple segment real-time tasks. Task segments alternate between using CPU and I/O resources. A task model is proposed which encompasses a wider class of tasks than models proposed earlier. Instead of developing new scheduling algorithms, the authors develop a class of slack distribution policies which use varying degrees of information about task structure and device utilization to budget task slack. Slack distribution policies are shown to improve the performance of all scheduling algorithms studied. Two key observations are: slack distribution is helpful beyond a certain threshold of task arrival rate, and algorithms which normally perform poorly are helped to a greater degree by slack distribution. A study of various scheduling algorithms for a constant value function reveals that all of them favor tasks with a large number of small segments to tasks with a small number of large segments. It is shown that the Moore ordering algorithm is not optimal for multiple segment tasks.>
Kamhing Ho, James H. Rice, Jaideep Srivastava
COMPSAC3
1990 Multilevel Secure Database Concurrency Control
abstract
The implications of multilevel security on database concurrency control are explored. Transactions are vital for multilevel secure database management systems (MLS/DBMSs) because they provide transparency to concurrency and to failure. Concurrent execution of transactions may lead to contention among subjects for access to data, which in MLS/DBMSs may lead to security problems. An abstraction of security models in terms of the transactions which they produce is presented. The notion of DC-Security which identifies a class of covert channels that are caused by contention for access to shared data, is introduced. This notion is useful for evaluating the security of transaction schedulers. A framework for multilevel secure schedulers which allows analysis of a schedulers' security properties at the protocol level is presented. Necessary and sufficient conditions are developed for DC-Security in this framework and proved using noninterference. A wide range of schedulers is evaluated against these conditions.>
Thomas F. Keefe, Wei-Tek Tsai, Jaideep Srivastava
ICDE3
1990 Parallelism in Database Production Systems
abstract
A study is made of the issues in parallelizing database production systems. Various kinds of parallelism possible in production systems are classified. Parallel execution of production systems leads to some subtle problems which, if not handled carefully, can cause altered execution semantics. The precise conditions that any parallel implementation of production systems must fulfil in order to be semantically consistent are identified. A mechanism that guarantees the semantic consistency of parallel execution and proves its correctness is presented. It is based on a novel locking mechanism that provides more parallelism than conventional two-phase locking. The various factors that can affect the actual speedup of a database production system are discussed.>
Jaideep Srivastava, Kuo-Wei Hwang, Jack S. Eddy Tan
ICDE1
1989 A framework for expressing and controlling imprecision in databases
abstract
The concept of precision in databases for both queries and data is introduced and formalized. Algorithms for maintaining a database at a specified degree of precision are presented and analyzed. It is shown how these algorithms can be used to choose an optimal query plan by means of a query optimizer.>
Jaideep Srivastava, Doron Rotem
COMPSAC1
1989 TBSAM: An Access Method for Efficient Processing of Statistical Queries
abstract
The domain of statistical and scientific databases is targeted here, and the class of aggregate queries which are very often encountered in this domain is considered. Such a query is aimed at retrieving some aggregate characteristics of the raw data. The tree-based statistics access method (TBSAM), which provides support for the efficient processing of aggregate queries, is presented. It is related to the B/sup +/-tree and also processes the B/sup +/-tree's efficient update properties. Complementing TBSAM is the provision of a grouped update algorithm for minimizing expensive indexed database updates.>
Jaideep Srivastava, Jack S. Eddy Tan, Vincent Y. Lum
IEEE Trans. Knowl. Data Eng.1
1988 A Tree Based Access Method (TBSAM) for Fast Processing of Aggregate Queries
abstract
A novel database access method for statistical database processing is discussed. The structure of the access method and its application for the processing of statistical queries is shown. Descriptive statistics on attributes of the data can be calculated very efficiently. In addition, the structure is suited for queries involving order statistics on an index attribute. An extra advantage provided by the structure is the ability to do various kinds of sampling. The cost of processing various queries is analyzed. Structural updates to the access method are discussed.>
Jaideep Srivastava, Vincent Y. Lum
ICDE1
1988 Efficient Algorithms for Maintenance of Large Database
abstract
Installing an update in an indexed database of size N takes O(log N) disk accesses, because of a root-leaf path traversal. If the installing of updates is deferred for a while, until a certain number n of them have been collected, substantial savings can be obtained. Two efficient algorithms for performing such grouped updates are presented and analyzed. The conditions on n are derived for which these algorithms perform better than straightforward update installation. A technique called model scaling, which is borrowed from the domain of fluid mechanics, is used to show how small simulations can be used for the accurate modeling of real-life environments. Some experimental results are presented to validate the analysis as well as justify the assumptions made.>
Jaideep Srivastava, C. V. Ramamoorthy
ICDE1
1988 Analytical Modeling of Materialized View Maintenance
abstract
Article Analytical modeling of materialized view maintenance Share on Authors: Jaideep Srivastava Computer Science Division, University of California Berkeley and Computer Science Research, Lawrence Berkeley Laboratories, University of California, Berkeley, CA Computer Science Division, University of California Berkeley and Computer Science Research, Lawrence Berkeley Laboratories, University of California, Berkeley, CAView Profile , Doron Rotem Computer Science Research, Lawrence Berkeley Laboratories, University of California, Berkeley, CA Computer Science Research, Lawrence Berkeley Laboratories, University of California, Berkeley, CAView Profile Authors Info & Claims PODS '88: Proceedings of the seventh ACM SIGACT-SIGMOD-SIGART symposium on Principles of database systemsMarch 1988 Pages 126–134https://doi.org/10.1145/308386.308425Published:01 March 1988 26citation292DownloadsMetricsTotal Citations26Total Downloads292Last 12 Months8Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jaideep Srivastava, Doron Rotem
PODS1
1988 Precision-Time Tradeoffs: A Paradigm for Processing Statistical Queries on Databases
Jaideep Srivastava, Doron Rotem
SSDBM1
1988 A Formal Model of Trade-off between Optimization and Execution Costs in Semantic Query Optimization
Shashi Shekhar 0001, Jaideep Srivastava, Soumitra Dutta
VLDB2
1986 A Distributed Clustering Algorithm for Large Computer Networks
C. V. Ramamoorthy, Jaideep Srivastava, Wei-Tek Tsai
ICDCS2