EDBT 2026 Demo / reviewers in the wild / expert
Aravindan Raghuveer
dblp:20/1664
· DBLP profile ↗
24ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0001-5006-4385ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 17 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FRACTAL: Fine-Grained Scoring from Aggregate Text LabelsabstractFine-Tuning of LLMs using RLHF / RLAIF has been shown as a critical step to improve the performance of LLMs in complex generation tasks.These methods typically use responselevel human or model feedback for alignment.Recent works indicate that finer sentence or span-level labels provide more accurate and interpretable feedback for LLM optimization.In this work, we propose FRACTAL, a suite of models to disaggregate responselevel labels into sentence-level (pseudo-)labels through Multiple Instance Learning (MIL) and Learning from Label Proportions (LLP) formulations, novel usage of prior information, and maximum likelihood calibration.We perform close to 2000 experiments across 6 datasets and 4 tasks that show that FRACTAL can reach up to 93% of the performance of the fully supervised baseline while requiring only around 10% of the gold labels.Furthermore, in a downstream eval, employing steplevel pseudo scores in RLHF for a math reasoning task leads to 5% absolute improvement in performance.Our work is the first to develop response-level feedback to sentence-level scoring techniques leveraging sentence-level prior information, along with comprehensive evaluations on multiple tasks as well as end-to-end finetuning evaluations. Yukti Makhija, Priyanka Agrawal, Rishi Saket, Aravindan Raghuveer |
ACL (1) | 4 |
| 2025 | Aggregating Data for Optimal LearningabstractMultiple Instance Regression (MIR) and Learning from Label Proportions (LLP) are useful learning frameworks, where the training data is partitioned into disjoint sets or bags, and only an aggregate label, i.e., bag-label for each bag is available to the learner. In the case of MIR, the bag-label is the label of an undisclosed instance from the bag, while in LLP, the bag-label is the mean of the bag’s labels. In this paper, we study for various loss functions in MIR and LLP, what is the optimal way to partition the dataset into bags such that the utility for downstream tasks like linear regression is maximized. We theoretically provide utility guarantees, and show that in each case, the optimal bagging strategy (approximately) reduces to finding an optimal clustering of the feature vectors and/or the labels with respect to natural objectives such as $k$-means. We also show that our bagging mechanisms can be made label-differentially private, incurring an additional utility error. We then generalize our results to the setting of Generalized Linear Models (GLMs). Finally, we experimentally validate our theoretical results. Sushant Agarwal, Yukti Makhija, Rishi Saket, Aravindan Raghuveer |
UAI | 4 |
| 2025 | Learning from Label Proportions and Covariate-shifted InstancesabstractIn many applications, especially due to lack of supervision or privacy concerns, the training data is grouped into bags of instances (feature-vectors) and for each bag we have only an aggregate label derived from the instance-labels in the bag. In learning from label proportions (LLP) the aggregate label is the average of the instance-labels in a bag, and a significant body of work has focused on training models in the LLP setting to predict instance-labels. In practice however, the training data may have fully supervised albeit covariate-shifted source data, along with the usual target data with bag-labels, and we wish to train a good instance-level predictor on the target domain. We call this the covariate-shifted hybrid LLP problem. Fully supervised covariate shifted data often has useful training signals and the goal is to leverage them for better predictive performance in the hybrid LLP setting. To achieve this, we develop methods for hybrid LLP which naturally incorporate the target bag-labels along with the source instance-labels, in the domain adaptation framework. Apart from proving theoretical guarantees bounding the target generalization error, we also conduct experiments on several publicly available datasets showing that our methods outperform LLP and domain adaptation baselines as well techniques from previous related work. Sagalpreet Singh, Navodita Sharma, Shreyas Havaldar, Rishi Saket, Aravindan Raghuveer |
UAI | 5 |
| 2024 | Fairness under Covariate Shift: Improving Fairness-Accuracy Tradeoff with Few Unlabeled Test SamplesabstractCovariate shift in the test data is a common practical phenomena that can significantly downgrade both the accuracy and the fairness performance of the model. Ensuring fairness across different sensitive groups under covariate shift is of paramount importance due to societal implications like criminal justice. We operate in the unsupervised regime where only a small set of unlabeled test samples along with a labeled training set is available. Towards improving fairness under this highly challenging yet realistic scenario, we make three contributions. First is a novel composite weighted entropy based objective for prediction accuracy which is optimized along with a representation matching loss for fairness. We experimentally verify that optimizing with our loss formulation outperforms a number of state-of-the-art baselines in the pareto sense with respect to the fairness-accuracy tradeoff on several standard datasets. Our second contribution is a new setting we term Asymmetric Covariate Shift that, to the best of our knowledge, has not been studied before. Asymmetric covariate shift occurs when distribution of covariates of one group shifts significantly compared to the other groups and this happens when a dominant group is over-represented. While this setting is extremely challenging for current baselines, We show that our proposed method significantly outperforms them. Our third contribution is theoretical, where we show that our weighted entropy term along with prediction loss on the training set approximates test loss under covariate shift. Empirically and through formal sample complexity bounds, we show that this approximation to the unseen test loss does not depend on importance sampling variance which affects many other baselines. Shreyas Havaldar, Jatin Chauhan, Karthikeyan Shanmugam 0001, Jay Nandy, Aravindan Raghuveer |
AAAI | 5 |
| 2024 | LLP-Bench: A Large Scale Tabular Benchmark for Learning from Label ProportionsabstractWith large neural models becoming increasingly accurate and powerful, they have raised privacy and transparency concerns on data usage. Therefore, data platforms, regulations and user expectations are rapidly evolving leading to enforcing privacy via aggregation. We focus on the use case of online advertising where the emergence of aggregate data is imminent and can significantly impact the multi-billion dollar industry. In aggregated datasets, labels are assigned to groups of data points rather than individual data points. This leads to a formulation of a weakly supervised task - Learning from Label Proportions where a model is trained on groups (a.k.a bags) of instances and their corresponding label proportions to predict labels for individual instances. While learning on aggregate data due to privacy concerns is becoming increasingly popular there is no large scale benchmark for measuring performance and guiding improvements on this important task. We propose LLP-Bench - a web scale benchmark with ~ 70 datasets and 45 million datapoints. To the best of our knowledge, LLP-Bench is the first large scale tabular LLP benchmark with an extensive diversity in constituent datasets, realistic in terms of the sponsored search datasets used and aggregation mechanisms followed. Through more than 3000 experiments we compare the performance of 9 SOTA methods in detail. To the best of our knowledge, this is the first study that compares diverse approaches in such depth. Anand Brahmbhatt, Mohith Pokala, Rishi Saket, Aravindan Raghuveer |
CIKM | 4 |
| 2024 | Learning from Label Proportions: Bootstrapping Supervised Learners via Belief PropagationabstractLearning from Label Proportions (LLP) is a learning problem where only aggregate level labels are available for groups of instances, called bags, during training, and the aim is to get the best performance at the instance-level on the test data. This setting arises in domains like advertising and medicine due to privacy considerations. We propose a novel algorithmic framework for this problem that iteratively performs two main steps. For the first step (Pseudo Labeling) in every iteration, we define a Gibbs distribution over binary instance labels that incorporates a) covariate information through the constraint that instances with similar covariates should have similar labels and b) the bag level aggregated label. We then use Belief Propagation (BP) to marginalize the Gibbs distribution to obtain pseudo labels. In the second step (Embedding Refinement), we use the pseudo labels to provide supervision for a learner that yields a better embedding. Further, we iterate on the two steps again by using the second step's embeddings as new covariates for the next iteration. In the final iteration, a classifier is trained using the pseudo labels. Our algorithm displays strong gains against several SOTA baselines (upto **15%**) for the LLP Binary Classification problem on various dataset types - tabular and Image. We achieve these improvements with minimal computational overhead above standard supervised learning due to Belief Propagation, for large bag sizes, even for a million samples. Shreyas Havaldar, Navodita Sharma, Shubhi Sareen, Karthikeyan Shanmugam 0001, Aravindan Raghuveer |
ICLR | 5 |
| 2024 | Generalization and Learnability in Multiple Instance RegressionabstractMultiple instance regression (MIR) was introduced by Ray and Page (2001) as an analogue of multiple instance learning (MIL) in which we are given bags of feature-vectors (instances) and for each bag there is a bag-label which matches the label of one (unknown) primary instance from that bag. The goal is to compute a hypothesis regressor consistent with the underlying instance-labels. A natural approach is to find the best primary instance assignment and regressor optimizing the mse loss on the bags though no formal generalization guarantees were known. Our work is the first to prove generalization error bounds for MIR when the bags are drawn i.i.d. at random. Essentially, with high probability any MIR regressor with low error on sampled bags also has low error on the underlying instance-label distribution. We next study the complexity of linear regression on MIR bags, shown to be NP-hard in general by Ray and Page (2001), who however left open the possibility of arbitrarily good approximations. Significantly strengthening previous work, we prove a strong inapproximability bound: even if there exists zero bag-loss MIR linear regressor on a collection of $2$-sized bags with labels in $[-1,1]$, it is NP-hard to find an MIR linear regressor with bag-loss $< C$ for some absolute constant $C > 0$. Our work also proposes a model training method for MIR based on a novel weighted assignment loss, geared towards handling overlapping bags which have not received much attention previously. We conduct empirical evaluations on synthetic and real-world datasets showing that our method outperforms the baseline MIR methods. Kushal Chauhan, Rishi Saket, Lorne Applebaum, Ashwinkumar Badanidiyuru, Chandan Giri, Aravindan Raghuveer |
UAI | 6 |
| 2023 | Bi-Phone: Modeling Inter Language Phonetic Influences in TextabstractAbhirut Gupta, Ananya B. Sai, Richard Sproat, Yuri Vasilevski, James Ren, Ambarish Jash, Sukhdeep Sodhi, Aravindan Raghuveer. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Abhirut Gupta, Ananya Sai, Richard Sproat, Yuri Vasilevski, James S. Ren, Ambarish Jash, Sukhdeep S. Sodhi, Aravindan Raghuveer |
ACL (1) | 8 |
| 2023 | Non-Uniform Adversarial Perturbations for Discrete Tabular DatasetsabstractWe study the problem of adversarial attack and robustness on tabular datasets with discrete features. The discrete features of a tabular dataset represent high-level meaningful concepts, with different sets of vocabularies, leading to requiring non-uniform robustness. Further, the notion of distance between tabular input instances is not well defined, making the problem of producing adversarial examples with minor perturbations qualitatively more challenging compared to existing methods. Towards this, our paper defines the notion of distance through the lens of feature embeddings, learnt to represent the discrete features. We then formulate the task of generating adversarial examples as abinary set selection problem under non-uniform feature importance. Next, we propose an efficient approximate gradient-descent based algorithm, calledDiscrete Non-uniform Approximation (DNA) attack, by reformulating the problem into a continuous domain to solve the original optimization problem for generating adversarial examples. We demonstrate the effectiveness of our proposed DNA attack using two large real-world discrete tabular datasets from e-commerce domains for binary classification, where the datasets are heavily biased for one-class. We also analyze challenges for existing adversarial training frameworks for such datasets under our DNA attack. Jay Nandy, Jatin Chauhan, Rishi Saket, Aravindan Raghuveer |
CIKM | 4 |
| 2023 | ReTAG: Reasoning Aware Table to Analytic Text GenerationabstractThe task of table summarization involves generating text that both succinctly and accurately represents the table or a specific set of highlighted cells within a table.While significant progress has been made in table to text generation techniques, models still mostly generate descriptive summaries, which reiterates the information contained within the table in sentences.Through analysis of popular table to text benchmarks (ToTTo (Parikh et al., 2020) and InfoTabs (Gupta et al., 2020)) we observe that in order to generate the ideal summary, multiple types of reasoning is needed coupled with access to knowledge beyond the scope of the table.To address this gap, we propose RETAG, a table and reasoning aware model that uses vector-quantization to infuse different types of analytical reasoning into the output.RETAG achieves 2.2%, 2.9% improvement on the PARENT metric in the relevant slice of ToTTo and InfoTabs for the table to text generation task over state of the art baselines.Through human evaluation, we observe that output from RETAG is upto 12% more faithful and analytical compared to a strong table-aware model.To the best of our knowledge, RETAG is the first model that can controllably use multiple reasoning methods within a structure-aware sequence to sequence model to surpass state of the art performance in multiple table to text tasks.We extend (and open source 35.6K analytical, 55.9k descriptive instances) the ToTTo, InfoTabs datasets with the reasoning categories used in each reference sentences. Deepanway Ghosal, Preksha Nema, Aravindan Raghuveer |
EMNLP | 3 |
| 2023 | PAC Learning Linear Thresholds from Label ProportionsabstractLearning from label proportions (LLP) is a generalization of supervised learning in which the training data is available as sets or bags of feature-vectors (instances) along with the average instance-label of each bag. The goal is to train a good instance classifier. While most previous works on LLP have focused on training models on such training data, computational learnability of LLP was only
recently explored by Saket (2021, 2022) who showed worst case intractability of properly learning linear threshold functions (LTFs) from label proportions. However, their work did not rule out efficient algorithms for this problem for natural distributions.
In this work we show that it is indeed possible to efficiently learn LTFs using LTFs when given access to random bags of some label proportion in which feature-vectors are, conditioned on their labels, independently sampled from a Gaussian distribution $N(µ, Σ)$. Our work shows that a certain matrix – formed using covariances of the differences of feature-vectors sampled from the bags with and without replacement – necessarily has its principal component, after a transformation, in the direction of the normal vector of the LTF. Our algorithm estimates the means and covariance matrices using subgaussian concentration bounds which we show can be applied to efficiently sample bags for approximating the normal direction. Using this in conjunction with novel generalization error bounds in the bag setting, we show that a low error hypothesis LTF can be identified. For some special cases of the $N(0, I)$ distribution we provide a simpler mean estimation based algorithm. We include an experimental evaluation of our learning algorithms along with a comparison with those of Saket (2021, 2022) and random LTFs, demonstrating the effectiveness of our techniques. Anand Brahmbhatt, Rishi Saket, Aravindan Raghuveer |
NeurIPS | 3 |
| 2022 | On Combining Bags to Better Learn from Label ProportionsabstractIn the framework of learning from label proportions (LLP) the goal is to learn a good instance-level label predictor from the observed label proportions of bags of instances. Most of the LLP algorithms either explicitly or implicitly assume the nature of bag distributions with respect to the actual labels and instances, or cleverly adapt supervised learning techniques to suit LLP. In practical applications however, the scale and nature of data could render such assumptions invalid and the many of the algorithms impractical. In this paper we address the hard problem of solving LLP with provable error bounds while being bag distribution agnostic and model agnostic. We first propose the concept of generalized bags, an extension of bags and then devise an algorithm to combine bag distributions, if possible, into good generalized bag distributions. We show that (w.h.p) any classifier optimizing the squared Euclidean label-proportion loss on such a generalized bag distribution is guaranteed to minimize the instance-level loss as well. The predictive quality of our method is experimentally evaluated and it equals or betters the previous methods on pseudo-synthetic and real-world datasets. Rishi Saket, Aravindan Raghuveer, Balaraman Ravindran |
AISTATS | 2 |
| 2022 | Domain-Agnostic Contrastive Representations for Learning from Label ProportionsabstractWe study the weak supervision learning problem of Learning from Label Proportions (LLP) where the goal is to learn an instance-level classifier using proportions of various class labels in a bag -- a collection of input instances that often can be highly correlated. While representation learning for weakly-supervised tasks is found to be effective, they often require domain knowledge. To the best of our knowledge, representation learning for tabular data (unstructured data containing both continuous and categorical features) are not studied. In this paper, we propose to learn diverse representations of instances within the same bags to effectively utilize the weak bag-level supervision. We propose a domain agnostic LLP method, called "Self Contrastive Representation Learning for LLP" (SelfCLR-LLP) that incorporates a novel self--contrastive function as an auxiliary loss to learn representations on tabular data for LLP. We show that diverse representations for instances within the same bags aid efficient usage of the weak bag-level LLP supervision. We evaluate the proposed method through extensive experiments on real-world LLP datasets from e-commerce applications to demonstrate the effectiveness of our proposed SelfCLR-LLP. In this paper, we propose to learn diverse representations of instances within the same bags to effectively utilize the weak bag-level supervision. We propose a domain agnostic LLP method, called "Self Contrastive Representation Learning for LLP" (SelfCLR-LLP) that incorporates a novel self--contrastive function as an auxiliary loss to learn representations on tabular data for LLP. We show that diverse representations for instances within the same bags aid efficient usage of the weak bag-level LLP supervision. We evaluate the proposed method through extensive experiments on real-world LLP datasets from e-commerce applications to demonstrate the effectiveness of our proposed SelfCLR-LLP. Jay Nandy, Rishi Saket, Jatin Chauhan, Balaraman Ravindran, Aravindan Raghuveer |
CIKM | 6 |
| 2022 | T-STAR: Truthful Style Transfer using AMR Graph as Intermediate RepresentationabstractUnavailability of parallel corpora for training text style transfer (TST) models is a very challenging yet common scenario.Also, TST models implicitly need to preserve the content while transforming a source sentence into the target style.To tackle these problems, an intermediate representation is often constructed that is devoid of style while still preserving the meaning of the source sentence.In this work, we study the usefulness of Abstract Meaning Representation (AMR) graph as the intermediate style agnostic representation.We posit that semantic notations like AMR are a natural choice for an intermediate representation.Hence, we propose T-STAR: a model comprising of two components, text-to-AMR encoder and a AMR-to-text decoder.We propose several modeling improvements to enhance the style agnosticity of the generated AMR.To the best of our knowledge, T-STAR is the first work that uses AMR as an intermediate representation for TST.With thorough experimental evaluation we show T-STAR significantly outperforms state of the art techniques by achieving on an average 15.2% higher content preservation with negligible loss (∼3%) in style accuracy.Through detailed human evaluation with 90, 000 ratings, we also show that T-STAR has upto 50% lesser hallucinations compared to state of the art TST models. Anubhav Jangra, Preksha Nema, Aravindan Raghuveer |
EMNLP | 3 |
| 2022 | CoCoa: An Encoder-Decoder Model for Controllable Code-switched GenerationabstractCode-switching has seen growing interest in recent years as an important multilingual NLP phenomenon.Generating code-switched text for data augmentation has been sufficiently well-explored.However, there is no prior work on generating code-switched text with fine-grained control on the degree of codeswitching and the lexical choices used to convey formality.We present COCOA, an encoder-decoder translation model that converts monolingual Hindi text to Hindi-English code-switched text with both encoder-side and decoder-side interventions to achieve finegrained controllable generation.COCOA can be invoked at test-time to synthesize codeswitched text that is simultaneously faithful to syntactic and lexical attributes relevant to code-switching.COCOA outputs were subjected to rigorous subjective and objective evaluations.Human evaluations establish that our outputs are of superior quality while being faithful to desired attributes.We show significantly improved BLEU scores when compared with human-generated code-switched references.Compared to competitive baselines, we show 10% reduction in perplexity on a language modeling task and also demonstrate clear improvements on a downstream code-switched sentiment analysis task. Sneha Mondal, Ritika, Shreya Pathak, Preethi Jyothi, Aravindan Raghuveer |
EMNLP | 5 |
| 2022 | Multi-Variate Time Series Forecasting on Variable SubsetsabstractWe formulate a new inference task in the domain of multivariate time series forecasting (MTSF), called Variable Subset Forecast (VSF), where only a small subset of the variables is available during inference. Variables are absent during inference because of long-term data loss (eg. sensor failures) or high -> low-resource domain shift between train / test. To the best of our knowledge, robustness of MTSF models in presence of such failures, has not been studied in the literature. Through extensive evaluation, we first show that the performance of state of the art methods degrade significantly in the VSF setting. We propose a non-parametric, wrapper technique that can be applied on top any existing forecast models. Through systematic experiments across 4 datasets and 5 forecast models, we show that our technique is able to recover close to 95% performance of the models even when only 15% of the original variables are present. Jatin Chauhan, Aravindan Raghuveer, Rishi Saket, Jay Nandy, Balaraman Ravindran |
KDD | 2 |
| 2022 | Walking with PACE - Personalized and Automated Coaching EngineabstractWe design and implement a personalized and automated physical activity coaching engine, PACE, which uses the Fogg’s behavioral model (FBM) to engage users in mini-conversation based coaching sessions. It is a chat-based nudge assistant that can boost (encourage) and sense (ask) the motivation, ability and propensity of users to walk and help them in achieving their step count targets, similar to a human coach. We demonstrate the feasibility, effectiveness and acceptability of PACE by directly comparing to human coaches in a Wizard-of-Oz deployment study with 33 participants over 21 days. We tracked coach-participant conversations, step counts and qualitative survey feedback. Our findings indicate that the PACE framework strongly emulated human coaching with no significant differences in the overall number of active days, step count and engagement patterns. The qualitative user feedback suggests that PACE cultivated a coach-like experience, offering barrier resolution via motivational and educational support. We use traditional human-computer interaction approaches, to interrogate the conversational data and report positive PACE-participant interaction patterns with respect to addressal, disclosure, collaborative target settings, and reflexivity. As a post-hoc analysis, we annotated the conversation logs from the human coaching arm and trained machine learning (ML) models on these data sets to predict the next boost (AUC 0.73 ± 0.02) and sense (AUC 0.83 ± 0.01) action. In future, such ML-based models could be made increasingly personalized and adaptive based on user behaviors. Madhurima Vardhan, Narayan Hegde, Srujana Merugu, Shantanu Prabhat, Deepak Nathani, Martin G. Seneviratne, Nur Muhammad, Pranay Reddy, Sriram Lakshminarasimhan, Karina Lorenzana, Eshan Motwani, Partha Talukdar, Aravindan Raghuveer |
UMAP | 14 |
| 2021 | HintedBT: Augmenting Back-Translation with Quality and Transliteration HintsabstractBack-translation (BT) of target monolingual corpora is a widely used data augmentation strategy for neural machine translation (NMT), especially for low-resource language pairs.To improve effectiveness of the available BT data, we introduce HintedBT-a family of techniques which provides hints (through tags) to the encoder and decoder.First, we propose a novel method of using both high and low quality BT data by providing hints (as source tags on the encoder) to the model about the quality of each source-target pair.We don't filter out low quality data but instead show that these hints enable the model to learn effectively from noisy data.Second, we address the problem of predicting whether a source token needs to be translated or transliterated to the target language, which is common in crossscript translation tasks (i.e., where source and target do not share the written script).For such cases, we propose training the model with additional hints (as target tags on the decoder) that provide information about the operation required on the source (translation or both translation and transliteration).We conduct experiments and detailed analyses on standard WMT benchmarks for three crossscript low/medium-resource language pairs: {Hindi,Gujarati,Tamil}→English.Our methods compare favorably with five strong and well established baselines.We show that using these hints, both separately and together, significantly improves translation quality and leads to state-of-the-art performance in all three language pairs in corresponding bilingual settings. Sahana Ramnath, Melvin Johnson, Abhirut Gupta, Aravindan Raghuveer |
EMNLP (1) | 4 |
| 2018 | Modeling Mobile User Actions for Purchase Recommendation using Deep Memory NetworksabstractRapid expansion of mobile devices has brought an unprecedented opportunity for mobile operators and content publishers to reach many users at any point in time. Understanding usage patterns of mobile applications (apps) is an integral task that precedes advertising efforts of providing relevant recommendations to users. However, this task can be very arduous due to the unstructured nature of app data, with sparseness in available information. This study proposes a novel approach to learn representations of mobile user actions using Deep Memory Networks. We validate the proposed approach on millions of app usage sessions built from large scale feeds of mobile app events and mobile purchase receipts. The empirical study demonstrates that the proposed approach performed better compared to several competitive baselines in terms of recommendation precision quality. To the best of our knowledge this is the first study analyzing app usage patterns for purchase recommendation. Djordje Gligorijevic, Jelena Gligorijevic, Aravindan Raghuveer, Mihajlo Grbovic, Zoran Obradovic |
SIGIR | 3 |
| 2007 | Towards efficient search on unstructured data: an intelligent-storage approachabstractApplications that create and consume unstructured data have grown both in scale of storage requirements and complexity of search primitives. We consider two such applications: exhaustive search and integration of structured and unstructured data. Current block-based storage systems are either incapable or inefficient to address the challenges bought forth by the above applications. We propose a storage framework to efficiently store and search unstructured and structured data while controlling storage management costs. Experimental results based on our prototype show that the proposed system can provide impressive performance and feature benefits. Aravindan Raghuveer, Meera Jindal, Mohamed F. Mokbel, Biplob K. Debnath, David Hung-Chang Du |
CIKM | 1 |
| 2007 | Enabling database-aware storage with OSD
Aravindan Raghuveer, Steven W. Schlosser, Sami Iren |
MSST | 1 |
| 2007 | A Network-Aware Approach for Video and Metadata StreamingabstractProviding quality of service (QoS) for Internet-based video streaming applications requires the server and/or client to be network-aware and adaptive. We present a dynamic rate and quality adaptation algorithm where the server varies its sending rate (without varying the quality level) to adapt to the network and client conditions and only as a last resort, does quality adaptation. We place the adaptation logic at the client since it has better knowledge about both the demand (buffer conditions, variable bit-rate requirements) and supply (network conditions). Our approach is unique because the server's sending rate is calculated based on the client's varying demand (consumption rate) and the network status. Also, we do not model the network as a black-box but instead augment endpoint observations with a feedback from the network to represent its status more precisely. To make an informed adaptation decision, the client requires sizes of all frames in the variable bit rate video. But the overhead involved in sending this metadata is significant. So we propose a lossy compression technique to reduce the amount of control information and consequently the overhead. We also present a scheduling algorithm, dynamic scheduling algorithm for reduced trace delivery (DART), to deliver the compressed control information to the client. This algorithm can be used to deliver any form of metadata (like subtitles, alerts, etc.), especially in applications like IP-TV. Simulations show that the proposed techniques can significantly improve user perceived QoS when compared to other popular adaptation methods. Aravindan Raghuveer, Ewa Kusmierek, David Hung-Chang Du |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | Techniques for efficient stream of layered video in heterogeneous client environmentsabstractUniversal multimedia access (UMA) refers to accessing multimedia content over a wide range of client terminals and network capacities. Scalable coding is a very popular technique to enable UMA for video. Overhead introduced by the scalable coding approach limits the number of layers that can be stored for each video. Therefore some clients may be served the closest available quality than the best-fit quality. This is a major drawback of scalable coding from the end-user perspective. We propose to employ transcoding to tailor content exactly to the client's best-fit quality when the required layer is not stored. Inserting a transcoder in the server-client path introduces new challenges in deciding the layering structure (number of layers, bandwidth per layer) of a video. The optimal layering structure should be decided based on factors like total I/O bandwidth penalty incurred due to layering and transcoding effort required to service the "non-layered" versions. The solution to this problem is further complicated by practical issues like diverse popularity of video objects and resource availability. Another issue that we address in this paper is reducing WAN bandwidth penalty incurred due to transport and coding overhead inherent to scalable coding. This particular problem applies to all schemes that use layered encoding to broadcast video. We map the above mentioned problems onto a 0-1 multiple choice knapsack structure and propose an algorithm to find a near-optimal solution. The uniqueness of our approach not only lies in the streaming model but also in the integrated manner in which we address a variety of issues put forth by layered coding. Aravindan Raghuveer, Nam Oh Kang, David Hung-Chang Du |
GLOBECOM | 1 |
| 2004 | Network-aware rate adaptation for video streamingabstractSince the current Internet provides a best-effort service, timely delivery of data is not guaranteed for video streaming applications. Providing quality of service (QoS) in such a setting requires server and client to be network-aware and adaptive. We present a dynamic rate and quality adaptation scheme where the server varies its sending rate (without varying the quality) to adapt to the network and client conditions, and, only as a last resort, does quality adaptation. Our approach is unique because the server's sending rate is calculated based on the client's varying demand, and the network status. We do not model the network as a black-box but instead augment endpoint observations with feedback from the network and hence represent its status more precisely. We also propose a method to reduce the amount of control information needed to make adaptation decisions Aravindan Raghuveer, Ewa Kusmierek, David Hung-Chang Du |
ICME | 1 |