EDBT 2026 Demo / reviewers in the wild / expert
Arnab Nandi 0001
dblp:28/568
· DBLP profile ↗
53ranked-venue papers
15as first author
5since 2021 · last 2025
0000-0002-4138-603XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 43 · 14 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 4Applied, interdisciplinary, general and emerging computing · 4Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating Video Segment Access via Quality-Aware Multi-Source SelectionabstractVideo data can be slow to process due to the size of video streams and the computational complexity needed to decode, transform, and encode them. These challenges are particularly significant in interactive applications, such as quickly generating compilation videos from a user search. We look at optimizing access to source video segments in multimedia systems where multiple separately encoded copies of video sources are available, such as proxy/optimized media in conventional non-linear video editors or VOD streams in content distribution networks. Rather than selecting a single source to use (e.g., "use the lowest-bitrate 720p source"), we specify a minimum visual quality (e.g., "use any frames with VMAF ≥ 85"). This quality constraint and the needed segment bounds are used to find the lowest-latency operations to decode a segment from multiple available sources with diverse bitrates, resolutions, and codecs. This uses higher-quality/slower-to-decode sources if the encoding is better aligned for the specific segment bounds, which can provide faster access than using just one lower-quality source. We provide a general solution to this Quality-Aware Multi-Source Selection problem with optimal computational complexity. We create a dataset using adaptive-bitrate streaming Video on Demand sources from YouTube's CDN. We evaluate our algorithm on simple segment decoding as well as embedded into a larger editing system---a declarative video editor. Our evaluation shows up to 23% lower latency access, depending on segment length, at identical visual quality levels. Dominik Winecki, Arnab Nandi 0001 |
MMSys | 2 |
| 2025 | OmniMesh: Addressing Findability Challenges in Distributed Nature Data Repositories
Arnab Nandi 0001, Wei-Lun Chao, Rongjun Qin, Carl Boettiger, Hilmar Lapp, Tanya Y. Berger-Wolf |
SSDBM | 1 |
| 2024 | V2V: Efficiently Synthesizing Video Results for Video QueriesabstractQuerying video data has become increasingly popular and useful. Video queries can be complex, ranging from retrieval tasks (“find me the top videos that have … ”), to analytics (“how many videos contained object X per day?”), to excerpting tasks (“highlight and zoom into scenes with object X near object Y”), or combinations thereof. Results for video queries are still typically shown as either relational data or a primitive collection of clickable thumbnails on a web page. Presenting query results in this form is an impedance mismatch with the video medium: they are cumbersome to skim through and are in a different modality and information density compared to the source data. We describe V2V, a system to efficiently synthesize video results for video queries. V2V returns a fully-edited video, allowing the user to consume results in the same manner as the source videos. A key challenge is that synthesizing video results from a collection of videos is computationally intensive, especially within interactive query response times. To address this, V2V features a grammar to express video transformations in a declarative manner and a heuristic optimizer that improves the efficiency of V2V processing in a manner similar to how databases execute relational queries. Experiments show that our V2V optimizer enables video synthesis to run 3x faster. Dominik Winecki, Arnab Nandi 0001 |
ICDE | 2 |
| 2021 | DreamStore: A Data Platform for Enabling Shared Augmented RealityabstractUnlike traditional object stores, Augmented Reality (AR) query workloads possess several unique characteristics, such as spatial and visual information. Such workloads are often keyed on a variety of attributes simultaneously, such as device orientation and position, the scene in view, and spatial anchors. The natural mode of user-interaction in these devices triggers queries implicitly based on the field in the user's view at any instant, generating data queries in excess of the device frame rate. Ensuring a smooth user experience in such a scenario requires a systemic solution exploiting the unique characteristics of the AR workloads. For exploration in such contexts, we are presented with a view-maintenance or cache-prefetching problem; how do we download the smallest subset from the server to the mixed reality device such that latency and device space constraints are met? We present a novel data platform - DreamStore, that considers AR queries as first-class queries, and view-maintenance and large-scale analytics infrastructure around this design choice. Through performance experiments on large-scale and query-intensive AR workloads on DreamStore, we show the advantages and the capabilities of our proposed platform. Meraj Ahmed Khan, Arnab Nandi 0001 |
VR | 2 |
| 2021 | Improving Information Extraction from Visually Rich Documents using Visual Span RepresentationsabstractAlong with textual content, visual features play an essential role in the semantics of visually rich documents. Information extraction (IE) tasks perform poorly on these documents if these visual cues are not taken into account. In this paper, we present Artemis - a visually aware, machine-learning-based IE method for heterogeneous visually rich documents. Artemis represents a visual span in a document by jointly encoding its visual and textual context for IE tasks. Our main contribution is two-fold. First, we develop a deep-learning model that identifies the local context boundary of a visual span with minimal human-labeling. Second, we describe a deep neural network that encodes the multimodal context of a visual span into a fixed-length vector by taking its textual and layout-specific features into account. It identifies the visual span(s) containing a named entity by leveraging this learned representation followed by an inference task. We evaluate Artemis on four heterogeneous datasets from different domains over a suite of information extraction tasks. Results show that it outperforms state-of-the-art text-based methods by up to 17 points in F1-score. Ritesh Sarkhel, Arnab Nandi 0001 |
Proc. VLDB Endow. | 2 |
| 2020 | Amplifying Domain Expertise to Combat Antimicrobial Resistance
Protiva Rahman, Arnab Nandi 0001, Erinn Hade, Emily S. Patterson, Courtney Hebert |
AMIA | 2 |
| 2020 | Interpretable Multi-headed Attention for Abstractive Summarization at Controllable LengthsabstractAbstractive summarization at controllable lengths is a challenging task in natural language processing.It is even more challenging for domains where limited training data is available or scenarios in which the length of the summary is not known beforehand.At the same time, when it comes to trusting machine-generated summaries, explaining how a summary was constructed in human-understandable terms may be critical.We propose Multi-level Summarizer (MLS), a supervised method to construct abstractive summaries of a text document at controllable lengths.The key enabler of our method is an interpretable multi-headed attention mechanism that computes attention distribution over an input document using an array of timestep independent semantic kernels.Each kernel optimizes a human-interpretable syntactic or semantic property.Exhaustive experiments on two low-resource datasets in English language show that MLS outperforms strong baselines by up to 14.70% in the METEOR score.Human evaluation of the summaries also suggests that they capture the key concepts of the document at various length-budgets. Ritesh Sarkhel, Moniba Keymanesh, Arnab Nandi 0001, Srinivasan Parthasarathy 0001 |
COLING | 3 |
| 2020 | Evaluating interactive data systems
Protiva Rahman, Lilong Jiang, Arnab Nandi 0001 |
VLDB J. | 3 |
| 2019 | ARQuery: Hallucinating Analytics over Real-World Data using Augmented Reality
Codi J. Burley, Arnab Nandi 0001 |
CIDR | 2 |
| 2019 | Deterministic Routing between Layout Abstractions for Multi-Scale Classification of Visually Rich DocumentsabstractClassifying heterogeneous visually rich documents is a challenging task. Difficulty of this task increases even more if the maximum allowed inference turnaround time is constrained by a threshold. The increased overhead in inference cost, compared to the limited gain in classification capabilities make current multi-scale approaches infeasible in such scenarios. There are two major contributions of this work. First, we propose a spatial pyramid model to extract highly discriminative multi-scale feature descriptors from a visually rich document by leveraging the inherent hierarchy of its layout. Second, we propose a deterministic routing scheme for accelerating end-to-end inference by utilizing the spatial pyramid model. A depth-wise separable multi-column convolutional network is developed to enable our method. We evaluated the proposed approach on four publicly available, benchmark datasets of visually rich documents. Results suggest that our proposed approach demonstrates robust performance compared to the state-of-the-art methods in both classification accuracy and total inference turnaround. Ritesh Sarkhel, Arnab Nandi 0001 |
IJCAI | 2 |
| 2019 | Flux capacitors for JavaScript deloreans: approximate caching for physics-based data interactionabstractInteractive visualizations have become an effective and pervasive mode of allowing users to explore the data in a visual, fluid, and immersive manner. While modern web, mobile, touch, and gesturedriven next-generation interfaces such as Leap Motion allow for highly interactive experiences, they pose unique and unprecedented workloads to the underlying data platform. Usually, these visualizations do not need precise results for most queries generated during an interaction, and the users require the intermediate results as feedback only to guide them towards their goal query. We present a middleware component - Flux Capacitor, that insulates the backend from bursty and query-intensive workloads. Flux Capacitor uses prefetching and caching strategies devised by exploiting the inherent physics-metaphor of UI widgets such as friction and inertia in range sliders, and typical characteristics of user-interaction. This enables low interaction response times while intelligently trading off accuracy Meraj Ahmed Khan, Arnab Nandi 0001 |
IUI | 2 |
| 2019 | Transformer: a database-driven approach to generating forms for constrained interactionabstractForm-based data insertion or querying is often one of the most time-consuming steps in data-driven workflows. The small screen and lack of physical keyboard in devices such as smartphones and smartwatches introduce imprecision during user input. This can lead to data quality issues such as incomplete responses and errors, increasing user input time. We present Transformer, a system that leverages the contents of the database to automatically optimize forms for constrained input settings. Our cost function models the user input effort based on the schema and data distribution. This is used by Transformer to find the user interface (UI) widget and layout with ideal input cost for each form field. We demonstrate through user studies that Transformer provides a significantly improved user experience, with up to 50% and 57% reduction in form completion time for smartphones and smartwatches respectively. Protiva Rahman, Arnab Nandi 0001 |
IUI | 2 |
| 2019 | Short and Long-term Pattern Discovery Over Large-Scale Geo-Spatiotemporal DataabstractPattern discovery in geo-spatiotemporal data (such as traffic and weather data) is about finding patterns of collocation, co-occurrence, cascading, or cause and effect between geospatial entities. Using simplistic definitions of spatiotemporal neighborhood (a common characteristic of the existing general-purpose frameworks) is not semantically representative of geo-spatiotemporal data. We therefore introduce a new geo-spatiotemporal pattern discovery framework which defines a semantically correct definition of neighborhood; and then provides two capabilities, one to explore propagation patterns and the other to explore influential patterns. Propagation patterns reveal common cascading forms of geospatial entities in a region. Influential patterns demonstrate the impact of temporally long-term geospatial entities on their neighborhood. We apply this framework on a large dataset of traffic and weather data at countrywide scale, collected for the contiguous United States over two years. Our important findings include the identification of 90 common propagation patterns of traffic and weather entities (e.g., rain --> accident --> congestion), which results in identification of four categories of states within the US; and interesting influential patterns with respect to the "location", "duration", and "type" of long-term entities (e.g., a major construction --> more traffic incidents). These patterns and the categorization of the states provide useful insights on the driving habits and infrastructure characteristics of different regions in the US, and could be of significant value for applications such as urban planning and personalized insurance. Sobhan Moosavi, Mohammad Hossein Samavatian, Arnab Nandi 0001, Srinivasan Parthasarathy 0001, Rajiv Ramnath |
KDD | 3 |
| 2019 | International Workshop on Human-In-the-Loop Data Analytics (HILDA)abstractThe Human In the Loop Data Analytics (HILDA) workshop aims to foster interdisciplinary efforts that tackle important challenges in better supporting humans in the loop in the context of data-intensive computations, such as interactive data exploration, integration, analytics, and machine learning. Over the past several years, HILDA has brought together DB researchers interested in the distinctive ways that people impact data management tasks, as well as like-minded researchers in other communities, such as Information Visualization, Data Mining/Machine Learning, and HCI. The work presented at HILDA covers a broad range of topics, from algorithmic, interface, and system design to the user's cognitive, physical, and goal-seeking perspectives when managing and exploring data as well as notions of approximation/prediction. This year, we continued to encourage submissions for initial ideas and visions, early reports of work in progress, as well as reflections on completed projects. Leilani Battle, Surajit Chaudhuri, Arnab Nandi 0001 |
SIGMOD Conference | 3 |
| 2019 | Visual Segmentation for Information Extraction from Heterogeneous Visually Rich DocumentsabstractPhysical and digital documents often contain visually rich information. With such information, there is no strict ordering or positioning in the document where the data values must appear. Along with textual cues, these documents often also rely on salient visual features to define distinct semantic boundaries and augment the information they disseminate. When performing information extraction (IE), traditional techniques fall short, as they use a text-only representation and do not consider the visual cues inherent to the layout of these documents. We propose VS2, a generalized approach for information extraction from heterogeneous visually rich documents. There are two major contributions of this work. First, we propose a robust segmentation algorithm that decomposes a visually rich document into a bag of visually isolated but semantically coherent areas, called logical blocks. Document type agnostic low-level visual and semantic features are used in this process. Our second contribution is a distantly supervised search-and-select method for identifying the named entities within these documents by utilizing the context boundaries defined by these logical blocks. Experimental results on three heterogeneous datasets suggest that the proposed approach significantly outperforms its text-only counterparts on all datasets. Comparing it against the state-of-the-art methods also reveal that VS2 performs comparably or better on all datasets. Ritesh Sarkhel, Arnab Nandi 0001 |
SIGMOD Conference | 2 |
| 2018 | Derivation of expert consensus rules for missing antimicrobial susceptibility data
Protiva Rahman, Erinn Hade, Arnab Nandi 0001, Preeti Pancholi, Mark Lusberg, Kurt Stevenson, Courtney Hebert |
AMIA | 3 |
| 2018 | Evaluating Interactive Data Systems: Workloads, Metrics, and GuidelinesabstractHighly interactive query interfaces have become a popular tool for ad-hoc data analysis and exploration, posing a new kind of workload to the underlying data infrastructure. Compared with traditional systems that are optimized for throughput or batched performance, ad-hoc and interactive data exploration systems focus more on user-centric interactivity, which raises a new class of performance challenges. Further, with the advent of new interaction devices~(e.g., touch, gesture) and different query interface paradigms~(e.g., sliders), maintaining interactive performance becomes even more challenging. Thus, when building interactive data systems, there is a clear need to articulate the design space. Lilong Jiang, Protiva Rahman, Arnab Nandi 0001 |
SIGMOD Conference | 3 |
| 2018 | ICARUS: Minimizing Human Effort in Iterative Data CompletionabstractAn important step in data preparation involves dealing with incomplete datasets. In some cases, the missing values are unreported because they are characteristics of the domain and are known by practitioners. Due to this nature of the missing values, imputation and inference methods do not work and input from domain experts is required. A common method for experts to fill missing values is through rules. However, for large datasets with thousands of missing data points, it is laborious and time consuming for a user to make sense of the data and formulate effective completion rules. Thus, users need to be shown subsets of the data that will have the most impact in completing missing fields. Further, these subsets should provide the user with enough information to make an update. Choosing subsets that maximize the probability of filling in missing data from a large dataset is computationally expensive. To address these challenges, we present Icarus, which uses a heuristic algorithm to show the user small subsets of the database in the form of a matrix. This allows the user to iteratively fill in data by applying suggested rules based on their direct edits to the matrix. The suggested rules amplify the users' input to multiple missing fields by using the database schema to infer hierarchies. Simulations show Icarus has an average improvement of 50% across three datasets over the baseline system. Further, in-person user studies demonstrate that naive users can fill in 68% of missing data within an hour, while manual rule specification spans weeks. Protiva Rahman, Courtney Hebert, Arnab Nandi 0001 |
Proc. VLDB Endow. | 3 |
| 2018 | A Session-Based Approach to Fast-But-Approximate Interactive Data Cube ExplorationabstractWith the proliferation of large datasets, sampling has become pervasive in data analysis. Sampling has numerous benefits—from reducing the computation time and cost to increasing the scope of interactive analysis. A popular task in data science, well-suited toward sampling, is the computation of fast-but-approximate aggregations over sampled data. Aggregation is a foundational block of data analysis, with data cube being its primary construct. We observe that such aggregation queries are typically issued in an ad-hoc, interactive setting. In contrast to one-off queries, a typical query session consists of a series of quick queries, interspersed with the user inspecting the results and formulating the next query. The similarity between session queries opens up opportunities for reusing computation of not just query results, but also error estimates. Error estimates need to be provided alongside sampled results for the results to be meaningful. We propose Sesame , a rewrite and caching framework that accelerates the entire interactive session of aggregation queries over sampled data. We focus on two unique and computationally expensive aspects of this use case: query speculation in the presence of sampling, and error computation, and provide novel strategies for result and error reuse. We demonstrate that our approach outperforms conventional sampled aggregation techniques by at least an order of magnitude, without modifying the underlying database. Niranjan Kamat, Arnab Nandi 0001 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2017 | Casual Querying: Facilitating Information Dissemination in Ad-hoc Environments
Arnab Nandi 0001 |
CIDR | 1 |
| 2017 | Characterizing Driving Context from Driver BehaviorabstractBecause of the increasing availability of spatiotemporal data, a variety of data-analytic applications have become possible. Characterizing driving context, where context may be thought of as a combination of location and time, is a new challenging application. An example of such a characterization is finding the correlation between driving behavior and traffic conditions. This contextual information enables analysts to validate observation-based hypotheses about the driving of an individual. In this paper, we present DriveContext, a novel framework to find the characteristics of a context, by extracting significant driving patterns (e.g., a slow-down), and then identifying the set of potential causes behind patterns (e.g., traffic congestion). Our experimental results confirm the feasibility of the framework in identifying meaningful driving patterns, with improvements in comparison with the state-of-the-art. We also demonstrate how the framework derives interesting characteristics for different contexts, through real-world examples. Sobhan Moosavi, Behrooz Omidvar-Tehrani, R. Bruce Craig, Arnab Nandi 0001, Rajiv Ramnath |
SIGSPATIAL/GIS | 4 |
| 2017 | Small DataabstractData is becoming increasingly personal. Individuals regularly interact with a wide variety of structured data, from SQLite databases on phones, to HR spreadsheets, to personal sensors, to open government data appearing in news articles. Although these workloads are important, many of the classical challenges associated with scale and Big Data do not apply. This panel brings together experts in a variety of fields to explore the new opportunities and challenges presented by "Small Data". Oliver Kennedy, D. Richard Hipp, Stratos Idreos, Amélie Marian, Arnab Nandi 0001, Carmela Troncoso, Eugene Wu 0002 |
ICDE | 5 |
| 2017 | DV8: Interactive Analysis of Aviation DataabstractThe vast volume of real-time air traffic data, being produced through new digital transmissions of the movement of aircraft throughout the US National Airspace System (NAS), is a rich resource for evaluating the performance of the system. To date, the potential for comprehensively analyzing this data has yet to be tapped, precisely due to a lack of tools that limit fully interactive data visualization. In this paper, we propose DV8, an interactive data visualization framework which provides in immediately visualized aviation-oriented insights, with a focus on evaluating the deviations among flights by route, type, airport, and aircraft performance. By providing scenarios validated by aviation experts, we illustrate different utilities of DV8 in areas such as capacity planning, flight route prediction, and fuel consumption. Behrooz Omidvar-Tehrani, Arnab Nandi 0001, Dalton Flanagan, Seth Young |
ICDE | 2 |
| 2017 | Don't Just Swipe Left, Tell Me Why: Enhancing Gesture-based Feedback with Reason BinsabstractDespite several advances in information retrieval systems and user interfaces, the specification of queries over text-based document collections remains a challenging problem. Query specification with keywords is a popular solution. However, given the widespread adoption of gesture-driven interfaces such as multitouch technologies in smartphones and tablets, the lack of a physical keyboard makes query specification with keywords inconvenient. We present BinGO, a novel gestural approach to querying text databases that allows users to refine their queries using a swipe gesture to either "like" or "dislike" candidate documents as well as express the reasons they like or dislike a document by swiping through automatically generated "reason bins". Such reasons refine a user's query with additional keywords. We present an online and efficient bin generation algorithm that presents reason bins at gesture articulation. We motivate and describe BinGo's unique interface design choices. Based on our analysis and user studies, we demonstrate that query specification by swiping through reason bins is easy and expressive. Juan Felipe Beltran, Azza Abouzeid, Arnab Nandi 0001 |
IUI | 4 |
| 2017 | On Designing a GeoViz-Aware Database System - Challenges and Opportunities
Mohamed Sarwat, Arnab Nandi 0001 |
SSTD | 2 |
| 2017 | A Unified Correlation-based Approach to Sampling Over JoinsabstractSupporting sampling in the presence of joins is an important problem in data analysis, but is inherently challenging due to the need to avoid correlation between output tuples. Current solutions provide either correlated or non-correlated samples. Sampling might not always be feasible in the non-correlated sampling-based approaches -- the sample size or intermediate data size might be exceedingly large. On the other hand, a correlated sample may not be representative of the join. This paper presents a unified strategy towards join sampling, while considering sample correlation every step of the way. We provide two key contributions. First, in the case where a correlated sample is acceptable, we provide techniques, for all join types, to sample base relations so that their join is as random as possible. Second, in the case where a correlated sample is not acceptable, we provide enhancements to the state-of-the-art algorithms to reduce their execution time and intermediate data size. Niranjan Kamat, Arnab Nandi 0001 |
SSDBM | 2 |
| 2017 | Data Tweening: Incremental Visualization of Data TransformsabstractIn the context of interactive query sessions, it is common to issue a succession of queries, transforming a dataset to the desired result. It is often difficult to comprehend a succession of transformations, especially for complex queries. Thus, to facilitate understanding of each data transformation and to provide continuous feedback, we introduce the concept of "data tweening", i.e., interpolating between resultsets, presenting to the user a series of incremental visual representations of a resultset transformation. We present tweening methods that consider not just the changes in the result, but also the changes in the query. Through user studies, we show that data tweening allows users to efficiently comprehend data transforms, and also enables them to gain a better understanding of the underlying query operations. Meraj Ahmed Khan, Larry Xu, Arnab Nandi 0001, Joseph M. Hellerstein |
Proc. VLDB Endow. | 3 |
| 2017 | DataTweener: A Demonstration of a Tweening Engine for Incremental Visualization of Data TransformsabstractWith the development and advancement of new data interaction modalities, data exploration and analysis has become a highly interactive process situating the user in a session of successive queries. With rapidly changing results, it becomes difficult for the end user to fully comprehend transformations, especially the transforms corresponding to complex queries. We introduce "data tweening" as an informative way of visualizing structural data transforms, presenting the users with a series of incremental visual representations of a resultset transformation. We present transformations as ordered sequences of basic structural transforms and visual cues. The sequences are generated using an automated framework which utilizes differences between the consecutive resultsets and queries in a query session. We evaluate the effectiveness of tweening as a visualization method through a user study. Meraj Ahmed Khan, Larry Xu, Arnab Nandi 0001, Joseph M. Hellerstein |
Proc. VLDB Endow. | 3 |
| 2016 | Hackathons as an Informal Learning PlatformabstractHackathons are fast-paced events where competitors work in teams to go from an idea to working software or hardware within a single day or a weekend and demonstrate their creation to a live audience of peers. Due to the "fun" and informal nature of such events, they make for excellent informal learning platforms that attract a diverse spectrum of students, especially those typically uninterested in traditional classroom settings. In this paper, we investigate the informal learning aspects of Ohio State's annual hackathon events over the past two years, with over 100 student participants in 2013 and over 200 student participants in 2014. Despite the competitive nature of such events, we observed a significant amount of peer-learning -- students teaching each other how to solve specific challenges and learn new skills. The events featured mentors from both the university and industry, who provided round-the-clock hands-on support, troubleshooting and advice. Due to the gamified format of the events, students were heavily motivated to learn new skills due to practical applicability and peer effects, rather than merely academic metrics. Some teams continued their hacks as long-term projects, while others formed new student groups to host lectures and practice building prototypes on a regular basis. Using a combined analysis of post-event surveys, student academic records and source-code commit log data from the event, we share insights, demographics, statistics and anecdotes from hosting these hackathons. Arnab Nandi 0001, Meris Mandernach |
SIGCSE | 1 |
| 2016 | FluxQuery: An Execution Framework for Highly Interactive Query WorkloadsabstractModern computing devices and user interfaces have necessitated highly interactive querying. Some of these interfaces issue a large number of dynamically changing and continuous queries to the backend. In others, users expect to inspect results during the query formulation process, in order to guide or help them towards specifying a full-fledged query. Thus, users end up issuing a fast-changing workload to the underlying database. In such situations, the user's query intent can be thought of as being in flux. In this paper, we show that the traditional query execution engines are not well-suited for this new class of highly interactive workloads. We propose a novel model to interpret the variability of likely queries in a workload. We implemented a cyclic scan-based approach to process queries from such workloads in an efficient and practical manner while reducing the overall system load. We evaluate and compare our methods with traditional systems and demonstrate the scalability of our approach, enabling thousands of queries to run simultaneously within interactive response times given low memory and CPU requirements. Roee Ebenstein, Niranjan Kamat, Arnab Nandi 0001 |
SIGMOD Conference | 3 |
| 2015 | Breathing Life into Database Textbooks
Arnab Nandi 0001 |
CIDR | 1 |
| 2015 | Surpassing Humans and Computers with JELLYBEAN: Crowd-Vision-Hybrid Counting AlgorithmsabstractCounting objects is a fundamental image processisng primitive, and has many scientific, health, surveillance, security, and military applications. Existing supervised computer vision techniques typically require large quantities of labeled training data, and even with that, fail to return accurate results in all but the most stylized settings. Using vanilla crowdsourcing, on the other hand, can lead to significant errors, especially on images with many objects. In this paper, we present our JellyBean suite of algorithms, that combines the best of crowds and computer vision to count objects in images, and uses judicious decomposition of images to greatly improve accuracy at low cost. Our algorithms have several desirable properties: (i) they are theoretically optimal or near-optimal, in that they ask as few questions as possible to humans (under certain intuitively reasonable assumptions that we justify in our paper experimentally); (ii) they operate under stand-alone or hybrid modes, in that they can either work independent of computer vision algorithms, or work in concert with them, depending on whether the computer vision techniques are available or useful for the given setting; (iii) they perform very well in practice, returning accurate counts on images that no individual worker or computer vision algorithm can count correctly, while not incurring a high cost. Akash Das Sarma, Arnab Nandi 0001, Aditya G. Parameswaran, Jennifer Widom |
HCOMP | 3 |
| 2015 | SnapToQuery: Providing Interactive Feedback during Exploratory Query SpecificationabstractA critical challenge in the data exploration process is discovering and issuing the "right" query, especially when the space of possible queries is large. This problem of exploratory query specification is exacerbated by the use of interactive user interfaces driven by mouse, touch, or next-generation, three-dimensional, motion capture-based devices; which, are often imprecise due to jitter and sensitivity issues. In this paper, we propose SnapToQuery , a novel technique that guides users through the query space by providing interactive feedback during the query specification process by "snapping" to the user's likely intended queries. These intended queries can be derived from prior query logs, or from the data itself, using methods described in this paper. In order to provide interactive response times over large datasets, we propose two data reduction techniques when snapping to these queries. Performance experiments demonstrate that our algorithms help maintain an interactive experience while allowing for accurate guidance. User studies over three kinds of devices (mouse, touch, and motion capture) show that SnapToQuery can help users specify queries quicker and more accurately; resulting in a query specification time speedup of 1.4× for mouse and touch-based devices and 2.2× for motion capture-based devices. Lilong Jiang, Arnab Nandi 0001 |
Proc. VLDB Endow. | 2 |
| 2014 | Distributed and interactive cube explorationabstractInteractive ad-hoc analytics over large datasets has become an increasingly popular use case. We detail the challenges encountered when building a distributed system that allows the interactive exploration of a data cube. We introduce DICE, a distributed system that uses a novel session-oriented model for data cube exploration, designed to provide the user with interactive sub-second latencies for specified accuracy levels. A novel framework is provided that combines three concepts: faceted exploration of data cubes, speculative execution of queries and query execution over subsets of data. We discuss design considerations, implementation details and optimizations of our system. Experiments demonstrate that DICE provides a sub-second interactive cube exploration experience at the billion-tuple scale that is at least 33% faster than current approaches. Niranjan Kamat, Prasanth Jayachandran, Karthik Tunga, Arnab Nandi 0001 |
ICDE | 4 |
| 2014 | SAGA: array storage as a DB with support for structural aggregationsabstractIn recent years, many Array DBMSs, including SciDB and RasDaMan have emerged to meet the needs of data management applications where the natural structures are the arrays. These systems, like their relational counterparts, involve an expensive data ingestion phase. The paradigm of using native storage as a DB and providing database-like support (e.g., the NoDB approach) has recently been shown to be an effective approach for dealing with infrequently queried data, where data ingestion costs cannot be justified, though only in context of relational data. Arnab Nandi 0001, Gagan Agrawal |
SSDBM | 2 |
| 2014 | Combining User Interaction, Speculative Query Execution and Sampling in the DICE SystemabstractThe interactive exploration of data cubes has become a popular application, especially over large datasets. In this paper, we present DICE , a combination of a novel frontend query interface and distributed aggregation backend that enables interactive cube exploration. DICE provides a convenient, practical alternative to the typical offline cube materialization strategy by allowing the user to explore facets of the data cube, trading off accuracy for interactive response-times, by sampling the data. We consider the time spent by the user perusing the results of their current query as an opportunity to execute and cache the most likely followup queries. The frontend presents a novel intuitive interface that allows for sampling-aware aggregations, and encourages interaction via our proposed faceted model. The design of our backend is tailored towards the low-latency user interaction at the frontend, and vice-versa. We discuss the synergistic design behind both the frontend user experience and the backend architecture of DICE ; and, present a demonstration that allows the user to fluidly interact with billion-tuple datasets within sub-second interactive response times. Prasanth Jayachandran, Karthik Tunga, Niranjan Kamat, Arnab Nandi 0001 |
Proc. VLDB Endow. | 4 |
| 2013 | Keyboards, R.I.P
Arnab Nandi 0001 |
CIDR | 1 |
| 2013 | Querying Without Keyboards
Arnab Nandi 0001 |
CIDR | 1 |
| 2013 | With a little help from my friendsabstractA typical person has numerous online friends that, according to studies, the person often consults for opinions and advice. However, public broadcasting a question to all friends risks social capital when repeated too often, is not tolerant to topic sensitivity, and can result in no response, as the message is lost in a myriad of status updates. Direct messaging is more personal and avoids these pitfalls, but requires manual selection of friends to contact, which can be time consuming and challenging. A user may have difficulty guessing which of their numerous online friends can provide a high quality and timely response. We demonstrate a working system that addresses these issues by returning an ordered subset of friends predicting (a) near-term availability, (b) willingness to respond and (c) topical knowledge, given a query. The combination of these three aspects are unique to our solution, and all are critical to the problem of obtaining timely and relevant responses. Our system acts as a decision aid - we give insight into why each friend was recommended and let the user decide whom to contact. Arnab Nandi 0001, Stelios Paparizos, John C. Shafer, Rakesh Agrawal 0001 |
ICDE | 1 |
| 2013 | PLASMA-HD: Probing the LAttice Structure and MAkeup of High-dimensional DataabstractRapidly making sense of, analyzing, and extracting useful information from large and complex data is a grand challenge. A user tasked with meeting this challenge is often befuddled with questions on where and how to begin to understand the relevant characteristics of such data. Real-world problem scenarios often involve scalability limitations and time constraints. In this paper we present an incremental interactive data analysis system as a step to address this challenge. This system builds on recent progress in the fields of interactive data exploration, locality sensitive hashing, knowledge caching, and graph visualization. Using visual clues based on rapid incremental estimates, a user is provided a multi-level capability to probe and interrogate the intrinsic structure of data. Throughout the interactive process, the output of previous probes can be used to construct increasingly tight coherence estimates across the parameter space, providing strong hints to the user about promising analysis steps to perform next. We present examples, interactive scenarios, and experimental results on several synthetic and real-world datasets which show the effectiveness and efficiency of our approach. The implications of this work are quite broad and can impact fields ranging from top-k algorithms to data clustering and from manifold learning to similarity search. David Fuhry, Venu Satuluri, Arnab Nandi 0001, Srinivasan Parthasarathy 0001 |
Proc. VLDB Endow. | 4 |
| 2013 | GestureQuery: A Multitouch Database Query InterfaceabstractMultitouch interfaces allow users to directly and interactively manipulate data. We propose bringing such interactive manipulation to the task of querying SQL databases. This paper describes an initial implementation of such an interface for multitouch tablet devices called GestureQuery that translates multitouch gestures into database queries. It provides database users with immediate constructive feedback on their queries, allowing rapid iteration and refinement of those queries. Based on preliminary user studies, Gesture-Query is easier to use, and lets users construct target queries quicker than console-based SQL and visual query builders while maintaining interactive performance. Lilong Jiang, Michael I. Mandel, Arnab Nandi 0001 |
Proc. VLDB Endow. | 3 |
| 2013 | Gestural Query SpecificationabstractDirect, ad-hoc interaction with databases has typically been performed over console-oriented conversational interfaces using query languages such as SQL. With the rise in popularity of gestural user interfaces and computing devices that use gestures as their exclusive modes of interaction, database query interfaces require a fundamental rethinking to work without keyboards. We present a novel query specification system that allows the user to query databases using a series of gestures. We present a novel gesture recognition system that uses both the interaction and the state of the database to classify gestural input into relational database queries. We conduct exhaustive systems performance tests and user studies to demonstrate that our system is not only performant and capable of interactive latencies, but it is also more usable, faster to use and more intuitive than existing systems. Arnab Nandi 0001, Lilong Jiang, Michael I. Mandel |
Proc. VLDB Endow. | 1 |
| 2012 | Skimmer: rapid scrolling of relational query resultsabstractA relational database often yields a large set of tuples as the result of a query. Users browse this result set to find the information they require. If the result set is large, there may be many pages of data to browse. Since results comprise tuples of alphanumeric values that have few visual markers, it is hard to browse the data quickly, even if it is sorted. Manish Singh 0002, Arnab Nandi 0001, H. V. Jagadish |
SIGMOD Conference | 2 |
| 2012 | Data Cube Materialization and Mining over MapReduceabstractComputing interesting measures for data cubes and subsequent mining of interesting cube groups over massive data sets are critical for many important analyses done in the real world. Previous studies have focused on algebraic measures such as SUM that are amenable to parallel computation and can easily benefit from the recent advancement of parallel computing infrastructure such as MapReduce. Dealing with holistic measures such as TOP-K, however, is nontrivial. In this paper, we detail real-world challenges in cube materialization and mining tasks on web-scale data sets. Specifically, we identify an important subset of holistic measures and introduce MR-Cube, a MapReduce-based framework for efficient cube computation and identification of interesting cube groups on holistic measures. We provide extensive experimental analyses over both real and synthetic data. We demonstrate that, unlike existing techniques which cannot scale to the 100 million tuple mark for our data sets, MR-Cube successfully and efficiently computes cubes with holistic measures over billion-tuple data sets. Arnab Nandi 0001, Cong Yu 0001, Philip Bohannon, Raghu Ramakrishnan 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2011 | Distributed cube materialization on holistic measuresabstractCube computation over massive datasets is critical for many important analyses done in the real world. Unlike commonly studied algebraic measures such as SUM that are amenable to parallel computation, efficient cube computation of holistic measures such as TOP-K is non-trivial and often impossible with current methods. In this paper we detail real-world challenges in cube materialization tasks on Web-scale datasets. Specifically, we identify an important subset of holistic measures and introduce MR-Cube, a MapReduce based framework for efficient cube computation on these measures. We provide extensive experimental analyses over both real and synthetic data. We demonstrate that, unlike existing techniques which cannot scale to the 100 million tuple mark for our datasets, MR-Cube successfully and efficiently computes cubes with holistic measures over billion-tuple datasets. Arnab Nandi 0001, Cong Yu 0001, Philip Bohannon, Raghu Ramakrishnan 0001 |
ICDE | 1 |
| 2011 | Guided Interaction: Rethinking the Query-Result Paradigm
Arnab Nandi 0001, H. V. Jagadish |
Proc. VLDB Endow. | 1 |
| 2009 | Qunits: queried units in database search
Arnab Nandi 0001, H. V. Jagadish |
CIDR | 1 |
| 2009 | PrivatePond: Outsourced Management of Web Corpuses
Daniel Fabbri, Arnab Nandi 0001, Kristen LeFevre, H. V. Jagadish |
WebDB | 2 |
| 2009 | HAMSTER: Using Search Clicklogs for Schema and Taxonomy MatchingabstractWe address the problem of unsupervised matching of schema information from a large number of data sources into the schema of a data warehouse. The matching process is the first step of a framework to integrate data feeds from third-party data providers into a structured-search engine's data warehouse. Our experiments show that traditional schema-based and instance-based schema matching methods fall short. We propose a new technique based on the search engine's clicklogs. Two schema elements are matched if the distribution of keyword queries that cause click-throughs on their instances are similar. We present experiments on large commercial datasets that show the new technique has much better accuracy than traditional techniques. Arnab Nandi 0001, Philip A. Bernstein |
Proc. VLDB Endow. | 1 |
| 2007 | Making database systems usableabstractDatabase researchers have striven to improve the capability of a database in terms of both performance and functionality. We assert that the usability of a database is as important as its capability. In this paper, we study why database systems today are so difficult to use. We identify a set of five pain points and propose a research agenda to address these. In particular, we introduce a presentation data model and recommend direct data manipulation with a schema later approach. We also stress the importance of provenance and of consistency across presentation models. H. V. Jagadish, Adriane Chapman, Aaron Elkiss, Magesh Jayapandian, Yunyao Li 0001, Arnab Nandi 0001, Cong Yu 0001 |
SIGMOD Conference | 6 |
| 2007 | Assisted querying using instant-response interfacesabstractWe demonstrate a novel query interface that enables users to construct a rich search query without any prior knowledge of the underlying schema or data. The interface, which is in the form of a single text input box, interacts in real-time with the users as they type, guiding them through the query construction. We discuss the issues of schema and data complexity, result size estimation, and query validity; and provide novel approaches to solving these problems. We demonstrate our query interface on two popular applications; an enterprise-wide personnel search, and a biological information database. Arnab Nandi 0001, H. V. Jagadish |
SIGMOD Conference | 1 |
| 2007 | Effective Phrase Prediction
Arnab Nandi 0001, H. V. Jagadish |
VLDB | 1 |
| 2005 | SPIN: searching personal information networksabstractNo abstract available. Soumen Chakrabarti, Jeetendra Mirchandani, Arnab Nandi 0001 |
SIGIR | 3 |