Amarnath Gupta

dblp:g/AmarnathGupta · DBLP profile ↗
← Back
68ranked-venue papers
16as first author
6since 2021 · last 2026
0000-0003-0897-120XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 41 · 9 first-author · 4 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Feasible Plan Generation with Ambiguity-Boundedness in Cross-Model Query Processing
Subhasis Dasgupta, Amarnath Gupta
DEXA (1)2
2026 Micro: a Lightweight Middleware for Optimizing Cross-Store Cross-Model Graph-Relation Joins
Xiuwen Zheng 0002, Arun Kumar 0001, Amarnath Gupta
ICDE3
2025 Scientific Knowledge Graph Construction Needs an AI-Mediated, Scientist-in-the-Loop Workflow (A Blue Sky Paper)
abstract
Scientific knowledge graph (KG) construction increasingly relies on large language models (LLMs) and heuristic pipelines to extract semantic structure from unstructured corpora. While automation enables scale, it often introduces ambiguity, overgeneralization, and inconsistencies—especially in biomedical and grey-literature domains. To address this, we propose a reflexive, AI-mediated, scientist-in-the-loop workflow that elevates human participation from labeling to semantic guidance. Unlike traditional HITL systems where humans serve as annotators or post hoc validators, our architecture enables the AI system to monitor confidence, detect failure modes, and initiate structured consultations with domain experts. Each intervention is triggered by statistical anomalies, graph discrepancies, or semantic drift, and is coupled with diagnostics, summaries, and provenance trails. We introduce the design of a modular enrichment pipeline augmented by diagnostic triggers, GGD-based corrections, and cube-based lineage tracking. This position paper articulates the motivation, architecture, and open challenges for realizing truly collaborative, transparent, and high-quality knowledge graph construction in science.
Jon Stephens, Shruti Sawant, Amarnath Gupta
eScience3
2024 Generating Cross-model Analytics Workloads Using LLMs
abstract
Data analytics applications today often require processing heterogeneous data from different data models, including relational, graph, and text data, for more holistic analytics. While query optimization for single data models, especially relational data, has been studied for decades, there is surprisingly little work on query optimization for cross-model data analytics. Cross-model query optimization can benefit from the long line of prior work in query optimization in the relational realm, wherein cost-based and/or machine learning-based (ML-based) optimizers are common. Both approaches require a large and diverse set of query workloads to measure, tune, and evaluate a query optimizer. To the best of our knowledge, there are still no large public cross-model benchmark workloads, a significant obstacle for systems researchers in this space. In this paper, we take a step toward filling this research gap by generating new query workloads spanning relational and graph data, which are ubiquitous in analytics applications. Our approach leverages large language models (LLMs) via different prompting strategies to generate queries and proposes new rule-based post-processing methods to ensure query correctness. We evaluate the pros and cons of each strategy and perform an in-depth analysis by categorizing the syntactic and semantic errors of the generated queries. So far, we have produced over 4000 correct cross-model queries, the largest set ever. Our code, prompts, data, and query workloads will all be released publicly.
Xiuwen Zheng 0002, Arun Kumar 0001, Amarnath Gupta
CIKM3
2024 On Modeling User-Aware Proactive Communication in Conversational Recommendation Systems*
abstract
Despite tremendous advances in large language models and conversational recommendation systems, these system are unable to handle many complex recommendation tasks. We present a new recommendation task called business ideation and show that the baseline performance of LLM-only recommendation is far from satisfactory. A part of the reason is neither the LLM not the recommendation has a proper user model and does not have the mechanisms to effective elicit a user model via interactions. We propose borrowing insights from decades of human communication research may bring in new ideas into the development of LLM-centric recommenders and present a research agenda to advance new generation systems.
Amarnath Gupta
e-Science1
2021 TemPredict: A Big Data Analytical Platform for Scalable Exploration and Monitoring of Personalized Multimodal Data for COVID-19
abstract
A key takeaway from the COVID-19 crisis is the need for scalable methods and systems for ingestion of big data related to the disease, such as models of the virus, health surveys, and social data, and the ability to integrate and analyze the ingested data rapidly. One specific example is the use of the Internet of Things and wearables (i.e., the Oura ring) to collect large-scale individualized data (e.g., temperature and heart rate) continuously and to create personalized baselines for detection of disease symptoms. Individualized data, when collected, has great potential to be linked with other datasets making it possible to combine individual and societal scale models for further understanding the disease. However, the volume and variability of such data require novel big data approaches to be developed as infrastructure for scalable use. This paper presents the data pipeline and big data infrastructure for the TemPredict project, which, to the best of our knowledge, is the largest public effort to gather continuous physiological data for time-series analysis. This effort unifies data ingestion with the development of a novel end-to-end cyberinfrastructure to enable the curation, cleaning, alignment, sketching, and passing of the data, in a secure manner, by the researchers making use of the ingested data for their COVID-19 detection algorithm development efforts. We present the challenges, the closed-loop data pipelines, and the secure infrastructure to support the development of time-sensitive algorithms for alerting individuals based on physiological predictors illness, enabling early intervention.
Shweta Purawat, Subhasis Dasgupta, Jining Song, Shakti Davis, Kajal T. Claypool, Sandeep Chandra, Ashley E. Mason, Varun K. Viswanath, Amit Klein 0002, Patrick Kasl, YingJing Wen, Benjamin L. Smarr, Amarnath Gupta, Ilkay Altintas
IEEE BigData13
2020 Discovering Interesting Subgraphs in Social Media Networks
abstract
Social media data are often modeled as heterogeneous graphs with multiple types of nodes and edges. We present a discovery algorithm that first chooses a “background” graph based on a user's analytical interest and then automatically discovers subgraphs that are structurally and content-wise distinctly different from the background graph. The technique combines the notion of a group-by operation on a graph and the notion of subjective interestingness, resulting in an automated discovery of interesting subgraphs. Our experiments on a socio-political database show the effectiveness of our technique.
Subhasis Dasgupta, Amarnath Gupta
ASONAM2
2020 An Algebraic Approach for High-level Text Analytics
abstract
Text analytical tasks like word embedding, phrase mining and topic modeling, are placing increasing demands as well as challenges to existing database management systems. In this paper, we provide a novel algebraic approach based on associative arrays. Our data model and algebra can bring together relational operators and text operators, which enables interesting optimization opportunities for hybrid data sources that have both relational and textual data. We demonstrate its expressive power in text analytics using several real-world tasks.
Xiuwen Zheng 0002, Amarnath Gupta
SSDBM2
2020 Towards finding the best-fit distribution for OSN data
Subhayan Bhattacharya, Sankhamita Sinha, Sarbani Roy, Amarnath Gupta
J. Supercomput.4
2019 Social network of extreme tweeters: a case study
abstract
The number of posts made by a single user account on a social media platform Twitter in any given time interval is usually low. However, there is a subset of users whose volume of posts is much higher than the median. In this paper, we investigate the content diversity and the social neighborhood of these extreme users and others. We define a metric called "interest narrowness", and identify that a subset of extreme users, termed anomalous users, post with very low topic diversity. We show that anomalous groups have the strongest within-group interactions, compared to their interaction with others, and exhibit different information sharing behaviors with other anomalous users compared to non-anomalous extreme tweeters.
Xiuwen Zheng 0002, Amarnath Gupta
ASONAM2
2019 Streaming Graph Ingestion with Resource-Aware Buffering and Graph Compression
abstract
Ingesting high-speed streaming data from social media into a graph database must overcome three problems – 1) the data can be really bursty, 2) the data must be transformed into a graph and 3) the graph database may not be able to ingest high-burst, high-velocity data. We have developed an adaptive buffering mechanism and a graph compression technique that effectively mitigate the problem.
Subhasis Dasgupta, Aditya Bagchi, Amarnath Gupta
eScience3
2019 Multi-model Investigative Exploration of Social Media Data with BOUTIQUE: A Case Study in Public Health
abstract
We present our experience with a data science problem in Public Health, where researchers use social media (Twitter) to determine whether the public shows awareness of HIV prevention measures offered by Public Health campaigns. To help the researcher, we develop a investigative exploration system called BOUTIQUE that allows a user to perform a multistep visualization and exploration of data through a dashboard interface. Unique features of BOUTIQUE includes its ability to handle heterogeneous types of data provided by a polystore, and its ability to use computation as part of the investigative exploration process. In this paper, we present the design of the BOUTIQUE middleware and walk through an investigation process for a real-life problem.
Junan Guo, Subhasis Dasgupta, Amarnath Gupta
eScience3
2019 On Constructing a Knowledge Base of Chinese Criminal Cases
abstract
We are developing a knowledge base over Chinese judicial decision documents to facilitate landscape analyses of Chinese Criminal Cases. We view judicial decision documents as a mixed-granularity semi-structured text where different levels of the text carry different semantic constructs and entailments.We use a combination of context-sensitive grammar, dependency parsing and discourse analysis to extract a formal and interpretable representation of these documents. Our knowledge base is developed by constructing associations between different elements of these documents. The interpretability is contributed in part by our formal representation of the Chinese criminal laws, also as semi-structured documents. The landscape analyses utilizes these two representations and enables a law researcher to ask legal pattern analysis queries.
Benjamin L. Liebman, Rachel E. Stern, Margaret E. Roberts, Amarnath Gupta
JURIX5
2017 Generating polystore ingestion plans - A demonstration with the AWESOME system
abstract
AWESOME is a polystore system that enables a data analyst to create a data ingestion script that specifies how it should collect, organize, run a data-derivation pipeline and reports results of the analysis. The collected data can be stored in different component stores under AWESOME for subsequent secondary analysis. This paper demonstrates the process by which AWESOME analyzes the script to construct an efficient ingestion plan, an executable database-centric dataflow specification which populates the raw and all derived data into the polystore. The demonstration will show how changing the script will alter the ingestion plan using a combination of rules and ingestion cost estimation.
Subhasis Dasgupta, Charles McKay, Amarnath Gupta
IEEE BigData3
2017 Toward Building a Legal Knowledge-Base of Chinese Judicial Documents for Large-Scale Analytics
abstract
We present an approach for constructing a legal knowledge-base that is sufficiently scalable to allow for large-scale corpus-level analyses. We do this by creating a polymorphic knowledge representation that includes hybrid ontologies, semistructured representations of sentences, and unsupervised statistical extraction of topics. We apply our approach to over one million judicial decision documents from Henan, China. Our knowledge-base allows us to make corpus-level queries that enable discovery, retrieval, and legal pattern analysis that shed new light on everyday law in China.
Amarnath Gupta, Alice Z. Wang, Haoshen Hong, Benjamin L. Liebman, Rachel E. Stern, Subhasis Dasgupta, Margaret E. Roberts
JURIX1
2016 Analytics-driven data ingestion and derivation in the AWESOME polystore
abstract
Polystores, i.e., data management systems that use multiple stores for different data models, are gaining popularity. We are developing a polystore-based system called AWESOME to support social data analytics. The AWESOME polystore can support relational, semistructured, graph and text data and houses a Spark computation engine to produce derived data during ingestion. ADIL, the data ingestion language of AWESOME allows a user to flexibly specify the placement of original and derived data into and across component stores and the computation engine. The paper also outlines a number of optimization strategies for managing data placement in AWESOME.
Subhasis Dasgupta, Kevin L. Coakley, Amarnath Gupta
IEEE BigData3
2013 Visualizing progressive discovery
abstract
Computational problems are increasingly relying on context-aware approaches for tractable solutions. Usually, these approaches statically link additional sources of information to those already present in the problem space. We have been building CueNet, a context discovery framework, which will dynamically discover the most relevant context for a given application problem. In this demonstration, we will show how the identities of people in personal photos can be discovered through contextual information. In this demonstration, we present Picatrix: an event based photo browsing web interface. Users can select a photo, and see a live visualization of how our context discovery algorithm, seeded with the initial information, discovers context from different data sources, and uses it to tag the faces in the given photo.
Arjun Satish, Ramesh Jain 0001, Amarnath Gupta
ICMR3
2013 Social life networks: a multimedia problem?
abstract
Connecting people to the resources they need is a fundamental task for any society. We present the idea of a technology that can be used by the middle tier of a society so that it uses people's mobile devices and social networks to connect the needy with providers. We conceive of a world observatory called the Social Life Network (SLN) that connects together people and things and monitors for people's needs as their life situations evolve. To enable such a system we need SLN to register and recognize situations by combining people's activities and data streaming from personal devices and environment sensors, and based on the situations make the connections when possible. But is this a multimedia problem? We show that many pattern recognition, machine learning, sensor fusion and information retrieval techniques used in multimedia-related research are deeply connected to the SLN problem. We sketch the functional architecture of such a system and show the place for these techniques.
Amarnath Gupta, Ramesh Jain 0001
ACM Multimedia1
2013 Semantic query reformulation: the NIF experience
abstract
The NIF system is a semantic search engine that uses an ontology to improve search quality. In this experience paper we present SKEYQL, our semantic keyword query language and describe a number of ontology-based query reformulation strategies that go beyond standard query expansion techniques. We also present a set of lessons learnt and strategies that did not work. We reaffirm the importance of pre-annotating data to ensure quality query results.
Amarnath Gupta, Anita E. Bandrowski, Christopher Condit, Xufei Qian, Jeffrey S. Grethe, Maryann E. Martone
SSDBM1
2012 Maturation of Neuroscience Information Framework: An Ontology Driven Information System for Neuroscience
abstract
The numbers of available neuroscience resources (databases, tools, materials and networks) on the web have, and continue to expand; particularly in light of newly implemented data sharing policies required by funding agencies and journals. However, the nature of dense, multi-faceted neuroscience data and the design of classic search engine systems makes efficient, reliable, and relevant discovery of such resources a significant challenge. This challenge is especially pertinent for online databases, whose dynamic content is largely opaque to contemporary search engines. The Neuroscience Information FrameworkThe Neuroscience Information Framework (NIF), http://neuinfo.org
Fahim T. Imam, Stephen D. Larson, Anita E. Bandrowski, Jeffrey S. Grethe, Amarnath Gupta, Maryann E. Martone
FOIS5
2012 Processing Semantic Keyword Queries for Scientific Literature
Ibrahim Burak Özyurt, Christopher Condit, Amarnath Gupta
NLDB3
2011 Corrigendum to "BioDB: An ontology-enhanced information system for heterogeneous biological information" [Data & Knowledge Engineering 69 (11) (2010) 1084-1102]
Amarnath Gupta, Christopher Condit, Xufei Qian
Data Knowl. Eng.1
2010 "Disputatio" on the Use of Ontologies in Multimedia
abstract
panel "Disputatio" on the Use of Ontologies in Multimedia Share on Authors: Simone Santini Universidad Autónoma de Madrid, Madrid, Spain Universidad Autónoma de Madrid, Madrid, SpainView Profile , Amarnath Gupta University of California San Diego, La Jolla, CA, USA University of California San Diego, La Jolla, CA, USAView Profile Authors Info & Claims MM '10: Proceedings of the 18th ACM international conference on MultimediaOctober 2010 Pages 1723–1728https://doi.org/10.1145/1873951.1874340Online:25 October 2010Publication History 2citation133DownloadsMetricsTotal Citations2Total Downloads133Last 12 Months1Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Simone Santini, Amarnath Gupta
ACM Multimedia2
2010 BiologicalNetworks 2.0 - an integrative view of genome biology data
abstract
BACKGROUND: A significant problem in the study of mechanisms of an organism's development is the elucidation of interrelated factors which are making an impact on the different levels of the organism, such as genes, biological molecules, cells, and cell systems. Numerous sources of heterogeneous data which exist for these subsystems are still not integrated sufficiently enough to give researchers a straightforward opportunity to analyze them together in the same frame of study. Systematic application of data integration methods is also hampered by a multitude of such factors as the orthogonal nature of the integrated data and naming problems. RESULTS: Here we report on a new version of BiologicalNetworks, a research environment for the integral visualization and analysis of heterogeneous biological data. BiologicalNetworks can be queried for properties of thousands of different types of biological entities (genes/proteins, promoters, COGs, pathways, binding sites, and other) and their relations (interactions, co-expression, co-citations, and other). The system includes the build-pathways infrastructure for molecular interactions/relations and module discovery in high-throughput experiments. Also implemented in BiologicalNetworks are the Integrated Genome Viewer and Comparative Genomics Browser applications, which allow for the search and analysis of gene regulatory regions and their conservation in multiple species in conjunction with molecular pathways/networks, experimental data and functional annotations. CONCLUSIONS: The new release of BiologicalNetworks together with its back-end database introduces extensive functionality for a more efficient integrated multi-level analysis of microarray, sequence, regulatory, and other data. BiologicalNetworks is freely available at http://www.biologicalnetworks.org.
Sergey Kozhenkov, Yulia Dubinina, Mayya Sedova, Amarnath Gupta, Julia V. Ponomarenko, Michael Baitaluk
BMC Bioinform.4
2010 BioDB: An ontology-enhanced information system for heterogeneous biological information
Amarnath Gupta, Christopher Condit, Xufei Qian
Data Knowl. Eng.1
2009 Ontology driven data integration for autism research
abstract
Autism spectrum disorder is an inherently complex phenomenon requiring large studies of many different types to further understanding of its causes. The National Database for Autism Research (NDAR) is being constructed to aid in this effort by providing a means for researchers to share and integrate data. An autism ontology drafted by a group at Stanford is being incorporated for use by NDAR to allow semantic data integration. The architecture upon which NDAR is built - the UCSD developed data integration environment - supports the use of this autism ontology, including annotation of data with ontological concepts and ontology enhanced queries on databases, both central and federated.
Lynn Young, Samson W. Tu, Lakshika Tennakoon, David Vismer, Vadim Astakhov, Amarnath Gupta, Jeffrey S. Grethe, Maryann E. Martone, Amar K. Das, Matthew J. McAuliffe
CBMS6
2009 MEDIALIFE: from images to a life chronicle
abstract
demonstration Share on MEDIALIFE: from images to a life chronicle Authors: Amarnath Gupta University of California San Diego, La Jolla, CA, USA University of California San Diego, La Jolla, CA, USAView Profile , Setareh Rafatirad University of California Irvine, Irvine, CA, USA University of California Irvine, Irvine, CA, USAView Profile , Mingyan Gao University of California Irvine, Irvine, CA, USA University of California Irvine, Irvine, CA, USAView Profile , Ramesh Jain University of California Irvine, Irvine, CA, USA University of California Irvine, Irvine, CA, USAView Profile Authors Info & Claims SIGMOD '09: Proceedings of the 2009 ACM SIGMOD International Conference on Management of dataJune 2009 Pages 1119–1122https://doi.org/10.1145/1559845.1559998Published:29 June 2009Publication History 4citation280DownloadsMetricsTotal Citations4Total Downloads280Last 12 Months2Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Amarnath Gupta, Setareh Rafatirad, Mingyan Gao, Ramesh Jain 0001
SIGMOD Conference1
2009 Tolkien: An Event Based Storytelling System
abstract
Since the dawn of human civilization, stories have been a popular medium of communication, both synchronously and asynchronously. Technically, a story is a time-ordered coherent sequence of events. In many applications, heterogeneous data is collected and organized so appropriate stories could be told. In this paper, we present a system that helps in generation of stories using a large database of events with associated multimodal data, called eventbase. We define storytelling as a two step process in which a storyteller can retrieve appropriate events and associated data, and then those are further filtered using preferences of the viewer. We develop this model using a measure of interestingness based on attributes of selected events and the preferences. Using an event system developed in our laboratory, we demonstrate the story telling process in Tolkien as one that generates multiple queries to select coherent interesting events to form a story.
Arjun Satish, Ramesh Jain 0001, Amarnath Gupta
Proc. VLDB Endow.3
2008 Graphitti: An Annotation Management System for Heterogeneous Objects
abstract
Annotation is the process of supplementing data with additional information that was not part of the actual observation, but reflects post-facto comments and associations made by a user who analyzes the data. While annotation management systems are emerging in the field of relational data, such systems for scientific applications, where there is a wide heterogeneity in the types of annotable data, are almost nonexistent. In this demonstration paper, we describeGraphitti, a tool that (i) allows a user to annotate a wide variety of scientific data, and (ii) allows a user to query data and their annotations in a seamless manner.
Christopher Condit, Amarnath Gupta
ICDE3
2008 MedSMan: a live multimedia stream querying system
Bin Liu 0010, Amarnath Gupta, Ramesh Jain 0001
Multim. Tools Appl.2
2006 Semantically Based Data Integration Environment for Biomedical Research
abstract
This paper presents an overview of the data integration mediation system developed as part of the Biomedical Informatics Research Network (BIRN; http://www.nbirn.net) project. BIRN is sponsored by the National Center for Research Resources (NCRR), a component of the National Institutes of Health (NIH). A core BIRN goal is the development of a multi-institution information management system to support biomedical research. Each participating institution maintains a database of their experimental or computationally derived data, and the data integration system performs semantic integration over the databases to enable researchers to perform analyses based on larger and broader datasets than would be available from any single institution's data. This demonstration paper describes architecture, implementation, and capabilities of the semantically based data integration system for BIRN
Vadim Astakhov, Amarnath Gupta, Jeffrey S. Grethe, Edward Ross, Aylin Yilmaz, Xufei Qian, Simone Santini, Maryann E. Martone, Mark H. Ellisman
CBMS2
2006 OntoQuest: Exploring Ontological Data Made Easy
Li Chen 0027, Maryann E. Martone, Amarnath Gupta, Lisa Fong, Mona Wong-Barnum
VLDB3
2006 PathSys: integrating molecular interaction graphs for systems biology
abstract
BACKGROUND: The goal of information integration in systems biology is to combine information from a number of databases and data sets, which are obtained from both high and low throughput experiments, under one data management scheme such that the cumulative information provides greater biological insight than is possible with individual information sources considered separately. RESULTS: Here we present PathSys, a graph-based system for creating a combined database of networks of interaction for generating integrated view of biological mechanisms. We used PathSys to integrate over 14 curated and publicly contributed data sources for the budding yeast (S. cerevisiae) and Gene Ontology. A number of exploratory questions were formulated as a combination of relational and graph-based queries to the integrated database. Thus, PathSys is a general-purpose, scalable, graph-data warehouse of biological information, complete with a graph manipulation and a query language, a storage mechanism and a generic data-importing mechanism through schema-mapping. CONCLUSION: Results from several test studies demonstrate the effectiveness of the approach in retrieving biologically interesting relations between genes and proteins, the networks connecting them, and of the utility of PathSys as a scalable graph-based warehouse for interaction-network integration and a hypothesis generator system. The PathSys's client software, named BiologicalNetworks, developed for navigation and analyses of molecular networks, is available as a Java Web Start application at http://brak.sdsc.edu/pub/BiologicalNetworks.
Michael Baitaluk, Xufei Qian, Shubhada Godbole, Alpan Raval, Animesh Ray, Amarnath Gupta
BMC Bioinform.6
2005 Efficient Algorithms for Pattern Matching on Directed Acyclic Graphs
abstract
Recently graph data models have become increasingly popular in many scientific fields. Efficient query processing over such data is critical. Existing works often rely on index structures that store pre-computed transitive relations to achieve efficient graph matching. In this paper, we present a family of stack-based algorithms to handle path and twig pattern queries for directed acyclic graphs (DAGs) in particular. With the worst-case space cost linearly bounded by the number of edges in the graph, our algorithms achieve a quadratic runtime complexity in the average size of the query variable bindings. This is optimal among the navigation-based graph matching algorithms.
Li Chen 0027, Amarnath Gupta, M. Erdem Kurul
ICDE2
2005 Spatiotemporal Annotation Graph (STAG): A Data Model for Composite Digital Objects
abstract
In this demonstration, we present a database over complex documents, which, in addition to a structured text content, also has update information, annotations, and embedded objects. We propose a new data model called spatiotemporal annotation graphs (STAG) for a database of composite digital objects and present a system that shows a query language to efficiently and effectively query such database. The particular application to be demonstrated is a database over annotated MS Word and PowerPoint presentations with embedded multimedia objects.
Smriti Yamini, Amarnath Gupta
ICDE2
2005 MedSMan: a streaming data management system over live multimedia
abstract
Querying live media streams is a challenging problem that is becoming an essential requirement in a growing number of applications. Research in multimedia information systems has addressed and made good progress in dealing with archived data. Meanwhile, research in stream databases has received significant attention for querying alphanumeric symbolic streams. The lack of a unifying data model capable of representing multimedia data and providing reasonable abstractions for querying live multimedia streams poses the challenge of how to make the best use of data in video and other sensor networks for various applications including video surveillance, live conferencing and Eventweb. This paper presents a system that enables direct capture of media streams from sensors and automatically generates meaningful feature streams that can be queried by a data stream processor. The system provides an effective combination of extensible digital processing techniques and general data stream management research.
Bin Liu 0010, Amarnath Gupta, Ramesh Jain 0001
ACM Multimedia2
2005 NSF Long Term Ecological Research Sites - Praxis et Theoria
Judith Bayard Cushing, Kristin Vanderbilt, James Brunt, Amarnath Gupta, Matthew B. Jones, Peter McCartney
SSDBM4
2005 Stack-based Algorithms for Pattern Matching on DAGs
Li Chen 0027, Amarnath Gupta, M. Erdem Kurul
VLDB2
2004 Modeling Functional Data Sources as Relations
Simone Santini, Amarnath Gupta
ER2
2004 Using Stream Semantics for Continuous Queries in Media Stream Processors
abstract
In the case of media and feature streams, explicit inter-stream constraints exist and can be exploited in the evaluation of continuous queries in the spirit of semantic query optimization. We express these properties using a media stream declaration language MSDL. In the demonstration, we present IMMERSI-MEET, an application built around an immersive environment. The IMMERSI-MEET system distinguishes between continuous streams, where values of different types come at a specified data rate, and discrete streams where sources push values intermittently. In MSDL, any dependence declaration must have at least one dependency specifying predicate in the body. As stream declarations are registered, the stream constraints are interpreted to construct a set of evaluation directives.
Amarnath Gupta, Bin Liu 0010, Pilho Kim, Ramesh Jain 0001
ICDE1
2004 Designing and executing scientific workflows with a programmable integrator
abstract
MOTIVATION: As in many other fields of science, computational methods in molecular biology need to intersperse information access and algorithm execution in a computational workflow. Users often find difficulties when transferring data between data sources and applications. In most cases there is no standard solution for workflow design and execution and tailored scripting mechanisms are implemented in a case by case basis. RESULTS: In this paper, we present a general purpose 'programmable integrator' that can access information from a variety of sources in a coordinated manner. Its usefulness in complex bioinformatics applications is claimed and supported by some application examples. AVAILABILITY: Tools are freely available to non-profit educations and research institutions. Usage by commercial organizations requires a license agreement. Software requirements: Java v1.3 (http://java.sun.com), Xerces XML Parser (http://xml.apache.org/xerces-j) and Kweelt implementation of XQuery (http://kweelt.sourceforge.net/).
Monica Chagoyen, M. Erdem Kurul, Pedro A. de Alarcón, José María Carazo, Amarnath Gupta
Bioinform.5
2003 An Interpolated Volume Model for Databases
Tianqiu Wang, Simone Santini, Amarnath Gupta
ER3
2003 BIRN-M: A Semantic Mediator for Solving Real-World Neuroscience Problems
abstract
No abstract available.
Amarnath Gupta, Bertram Ludäscher, Maryann E. Martone
SIGMOD Conference1
2003 A Modeling and Execution Environment for Distributed Scientific Workflows
abstract
We illustrate how a domain scientist can perform a complex scientific task by interleaving data access, querying, and manipulation, as well as analytical steps and computations in complex, problem specific ways. We show how our system is used by a geneticist for solving the problem of discovering the so-called "co-regulated" genes by interlinking data and computation from several Web sites, local computations, as well as local and remote databases. The main distinctive features of our system (compared, e.g., to the ZOO environment (Ioannidis et al., 1996)) include: (i) executable workflows run as Web services; (ii) abstract workflows employ concept names and semantic types that are higher-level (and thus more "scientist friendly") than executable workflows; and (iii) our system supports automatic translation of the latter into the former.
Ilkay Altintas, Sangeeta Bhagwanani, David Buttler, Sandeep Chandra, Zhengang Cheng, Matthew Coleman, Terence Critchlow, Amarnath Gupta, Ling Liu 0001, Bertram Ludäscher, Calton Pu, Reagan W. Moore, Arie Shoshani, Mladen A. Vouk
SSDBM8
2003 Compiling Abstract Scientific Workflows into Web Service Workflows
abstract
The authors present an approach for compiling "scientist-friendly" abstract workflow specifications into real-world executable workflows of Web service invocations, using a set of abstract-as-view definitions from a repository of abstract tasks. There have been a number of proposals and systems for scientific workflow management. However, our approach features unique aspects, in particular: the separation of abstract and concrete executable workflows; and the use of database mediation techniques to automatically translate AWFs into EWFs.
Bertram Ludäscher, Ilkay Altintas, Amarnath Gupta
SSDBM3
2003 A Practical Approach for Microscopy Imaging Data Management (MIDM) in Neuroscience
abstract
Current data management approaches can easily handle the relatively simple requirements for molecular biology research but not the more varied and sophisticated microscopy imaging data in neuroscience research. We developed a project-oriented experimental imaging data management system through integration of the object-relational Oracle DBMS (database management system) and a distributed file management system, the storage resource broker (SRB). The data model we developed on Oracle9i supports semantic and analytical queries and image content mining. The MIDM provides comprehensive descriptive, structural, spatial and administrative information on microscopy image datasets. The current MIDM is web accessible at http://ncmir.ucsd.edu/CCDB. This paper describes the MIDM architecture and data mode in MIDM.
Shenglan Zhang, Xufei Qian, Amarnath Gupta, Maryann E. Martone
SSDBM3
2003 An Interpolated Volume Data Model
Tianqiu Wang, Simone Santini, Amarnath Gupta
VLDB3
2003 Towards a formalization of disease-specific ontologies for neuroinformatics
Amarnath Gupta, Bertram Ludäscher, Jeffrey S. Grethe, Maryann E. Martone
Neural Networks1
2002 Navigating Virtual Information Sources with Know-ME
Xufei Qian, Bertram Ludäscher, Maryann E. Martone, Amarnath Gupta
EDBT4
2002 Registering Scientific Information Sources for Semantic Mediation
Amarnath Gupta, Bertram Ludäscher, Maryann E. Martone
ER1
2002 Conceptual Integration of Multiple Partial Geometric Models
Simone Santini, Amarnath Gupta
ER2
2002 GeMBASE: A Geometric Mediator for Brain Analysis with Surface Ensembles
Simone Santini, Amarnath Gupta
VLDB2
2002 Principles of schema design for multimedia databases
abstract
This paper presents the rudiments of a theory of schema design for databases containing high dimensional features of the type used for describing multimedia data. We introduce a model of multimedia database based on tables containing feature types, and the concept of schema design which is based on splitting tables depending on the functional relations between different parts of the features. We show that certain relations between substructures of a same feature structure can lead to schemas for which efficient algorithms for k-nearest neighbor and range searches can be defined.
Simone Santini, Amarnath Gupta
IEEE Trans. Multim.2
2001 Model-Based Mediation with Domain Maps
abstract
Proposes an extension to current view-based mediator systems called "model-based mediation", in which views are defined and executed at the level of conceptual models (CMs) rather than at the structural level. Structural integration and lifting of data to the conceptual level is "pushed down" from the mediator to wrappers which, in our system, export the classes, associations, constraints and query capabilities of a source. Another novel feature of our architecture is the use of domain maps - semantic nets of concepts and relationships that are used to mediate across sources from multiple worlds (i.e. whose data are related in indirect and often complex ways). As part of registering a source's CM with the mediator, the wrapper creates a "semantic index" of its data into the domain map. We show that these indexes not only semantically correlate the multiple-worlds data, and thereby support the definition of the integrated CM, but they are also useful during query processing, for example, to select relevant sources. A first prototype of the system has been implemented for a complex neuroscience mediation problem.
Bertram Ludäscher, Amarnath Gupta, Maryann E. Martone
ICDE2
2001 A Wavelet Data Model For Image Databases
abstract
This paper defines wavelet transforms as a datatype suitable for inclusion in databases and an algebra for the manipulation of this data type.
Simone Santini, Amarnath Gupta
ICME2
2001 Towards a federated neuroscientific knowledge management system using brain atlases
Gully A. P. C. Burns, Klaas E. Stephan, Bertram Ludäscher, Amarnath Gupta, Rolf Kötter
Neurocomputing4
2001 Emergent Semantics through Interaction in Image Databases
abstract
In this paper, we briefly discuss some aspects of image semantics and the role that it plays for the design of image databases. We argue that images don't have an intrinsic meaning, but that they are endowed with a meaning by placing them in the context of other images and by the user interaction. From this observation, we conclude that, in an image, database users should be allowed to manipulate not only the individual images, but also the relation between them. We present an interface model based on the manipulation of configurations of images.
Simone Santini, Amarnath Gupta, Ramesh Jain 0001
IEEE Trans. Knowl. Data Eng.2
2000 Semiorder Database for Complex Activity Recognition in Multi-Sensory Environments
abstract
A prototype semiorder database used for activity recognition in multi-sensory monitoring environments is described. Activities are spatio-temporal compositions of events, which are a type of atomic semantic units for such compositions. The focus is on the temporal composition of activities from events in the presence of bounded duration of temporal uncertainty in an event occurrence. Such temporal uncertainty forces the concurrency between event occurrences to be intransitive. Under certain assumptions, a subclass of partial orders, known as semiorders, models such intransitive concurrency appropriately. The semiorder database stores events and their semiorder temporal order of occurrences. A semiorder data model and the corresponding query language that embeds a semiorder pattern language are the main constituents of the semiorder database. We demonstrate this database and queries for activity recognition in a real time environment. The demonstration also includes a transducer subsystem for detection of events.
Shailendra K. Bhonsle, Amarnath Gupta, Simone Santini, Ramesh Jain 0001
ICDE2
2000 Database Architecture for Autonomous Transportation Agents for On-Scene Networked Incident Management (ATON)
abstract
A collection of distributed databases forms an important architectural component of the ATON project for networked incidence management of highway traffic. The database sub-architecture supports the architectural integration of many thematic areas of the ATON, and provides many high level abstractions that semantically correspond to traffic incidents. These databases are queried for the detection of local or distributed traffic incidents by many distributed control and analysis algorithms. The abstract database sub-architecture leads to many databases of different types having different functions. The semantic event/activity database is one of them. The distributed multi-sensory sub-architecture detects atomic semantic events that occur in the environment. The event/activity database stores them and their temporal structures. An activity is a temporal composition of atomic events. A specialized query language is used to flexibly define and detect activities of interest. The query language embodies high level semantic pattern matching abstractions. Preliminary results of using the event/activity database are also presented.
Mohan M. Trivedi, Shailendra K. Bhonsle, Amarnath Gupta
ICPR3
2000 Knowledge-Based Integration of Neuroscience Data Sources
abstract
The need for information integration is paramount in many biological disciplines, because of the large heterogeneity in both the types of data involved and in the diversity of approaches (physiological, anatomical, biochemical, etc.) taken by biologists to study the same or correlated phenomena. However, the very heterogeneity makes the task of information integration very difficult since two approaches studying different aspects of the same phenomena may not even share common attributes in their schema description. The paper develops a wrapper-mediator architecture which extends the conventional data- and view-oriented information mediation approach by incorporating additional knowledge modules that bridge the gap between the heterogeneous data sources. The semantic integration of the disparate local data sources employs F-logic as a data and knowledge representation and reasoning formalism. We show that the rich object oriented modeling features of F-logic together with its declarative rule language and the uniform treatment of data and metadata (schema information) make it an ideal candidate for complex integration tasks. We substantiate this claim by elaborating on our integration architecture and illustrating the approach using real world examples from the neuroscience domain. The complete integration framework is currently under development; a first prototype establishing the viability of the approach is operational.
Amarnath Gupta, Bertram Ludäscher, Maryann E. Martone
SSDBM1
2000 Model-Based Information Integration in a Neuroscience Mediator System
Bertram Ludäscher, Amarnath Gupta, Maryann E. Martone
VLDB2
2000 Content-Based Image Retrieval at the End of the Early Years
abstract
Presents a review of 200 references in content-based image retrieval. The paper starts with discussing the working conditions of content-based retrieval: patterns of use, types of pictures, the role of semantics, and the sensory gap. Subsequent sections discuss computational steps for image retrieval systems. Step one of the review is image processing for retrieval sorted by color, texture, and local geometry. Features for retrieval are discussed next, sorted by: accumulative and global features, salient points, object and shape features, signs, and structural combinations thereof. Similarity of pictures and objects in pictures is reviewed for each of the feature types, in close connection to the types and means of feedback the user of the systems is capable of giving by interaction. We briefly discuss aspects of system engineering: databases, system architecture, and evaluation. In the concluding section, we present our view on: the driving force of the field, the heritage from computer vision, the influence on computer vision, the role of similarity and of interaction, the need for databases, the problem of evaluation, and the role of the semantic gap.
Arnold W. M. Smeulders, Marcel Worring, Simone Santini, Amarnath Gupta, Ramesh Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
1999 XML-Based Information Mediation with MIX
abstract
The MIX mediator system, MIXm, is developed as part of the MIX Project at the San Diego Supercomputer Center, and the University of California, San Diego.1 MIXm uses XML as the common model for data exchange. Mediator views are expressed in XMAS (XML Matching And Structuring Language), a declarative XML query language. To facilitate user-friendly query formulation and for optimization purposes, MIXm employs XML DTDs as a structural description (in effect, a “schema”) of the exchanged data. The novel features of the system include:
Chaitanya K. Baru, Amarnath Gupta, Bertram Ludäscher, Richard Marciano, Yannis Papakonstantinou, Pavel E. Velikhov, Vincent Chu
SIGMOD Conference2
1999 An extensible information model for shared scientific data collections
Amarnath Gupta, Chaitanya K. Baru
Future Gener. Comput. Syst.1
1996 Content-based retrieval of ophthalmological images
abstract
This paper describes steps towards an information system for the storage and content-based retrieval of ocular fundus images. Based on the Virage Incorporated framework for defining similarity metrics, the authors have developed a number of primitives for the representation of ocular fundus images. A prototype Query By Pictorial Example (QBPE) system yields similarity rankings in approximate agreement with those of a human expert.
Amarnath Gupta, Saied Moezzi, Adam L. Taylor, Shankar Chatterjee, Ramesh Jain 0001, Michael H. Goldbaum, S. Burgess
ICIP (3)1
1996 Interactive Video on WWW: Beyond VCR-Like Interfaces
Arun Katkere, Jennifer Schlenzig, Amarnath Gupta, Ramesh Jain 0001
Comput. Networks3
1996 A hue preserving enhancement scheme for a class of colour images
Amarnath Gupta, Bhabatosh Chanda
Pattern Recognit. Lett.1
1991 Semantic Queries with Pictures: The VIMSYS Model
Amarnath Gupta, Terry E. Weymouth, Ramesh Jain 0001
VLDB1