Sagar Pandit

dblp:06/3384 · also Sagar A. Pandit · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
0since 2021 · last 2015
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 54% Spatial and temporal data management · 46%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
approximate query processing
0.212013
Approximate Algorithms for Computing Spatial Distance Histograms with Accuracy Guarantees · IEEE Trans. Knowl. Data Eng. 2013
Query processing and optimization › approximate query processing
error bounds
0.212013
Approximate Algorithms for Computing Spatial Distance Histograms with Accuracy Guarantees · IEEE Trans. Knowl. Data Eng. 2013
Spatial and temporal data management › spatial indexing
quadtree
0.112009
Computing Distance Histograms Efficiently in Scientific Databases · ICDE 2009
Spatial and temporal data management
spatial indexing
0.112009
Computing Distance Histograms Efficiently in Scientific Databases · ICDE 2009
Spatial and temporal data management
spatial query processing
0.112009
Computing Distance Histograms Efficiently in Scientific Databases · ICDE 2009

Methods — techniques the papers use, named apart from their topics

mathematical error model · 0.2density map · 0.1complexity analysis · 0.1
YearPublicationVenuePosition
2015 Push-based system for molecular simulation data analysis
abstract
Many scientific fields generate, and require manipulation of big data. Known scientific data analysis systems, as well as traditional DBMSs, follow a pull-based architectural design, where the executed queries mandate the data needed. This design, while suitable for traditional transaction-based workloads where number of queries retrieve small parts of data located at various places of the database, is ill-fitted for applications involving complex analysis on most of the data. Such design involves redundant and random I/O, considerably affecting the data throughput in the system. In this paper, we design and implement a push-based type system that allows high-throughput data analysis in the process of scientific discovery. Our design improves throughput in two ways: i) it uses a sequential scan-based I/O framework that loads the data into the main memory, and then ii) the system pushes the loaded data to a number of pre-programmed queries. By this way the system lowers the unnecessary I/O overhead imposed by the randomized, index-based scan and that of a multiple data reads if each query were to be fed separately. Considering the amount of data and the number of executed queries, we believe our system provides substantial improvement over the current data analyzing systems. The efficiency of the proposed system is backed by the results of extensive experiments using real MS data. The running times of our system are compared to those of the GROMACS system. The comparison shows the advantage and the potential of using such push-based system for data system analysis.
Vladimir Grupcev, Yi-Cheng Tu, Joseph C. Fogarty, Sagar Pandit
IEEE BigData4
2013 Approximate Algorithms for Computing Spatial Distance Histograms with Accuracy Guarantees
abstract
Particle simulation has become an important research tool in many scientific and engineering fields. Data generated by such simulations impose great challenges to database storage and query processing. One of the queries against particle simulation data, the spatial distance histogram (SDH) query, is the building block of many high-level analytics, and requires quadratic time to compute using a straightforward algorithm. Previous work has developed efficient algorithms that compute exact SDHs. While beating the naive solution, such algorithms are still not practical in processing SDH queries against large-scale simulation data. In this paper, we take a different path to tackle this problem by focusing on approximate algorithms with provable error bounds. We first present a solution derived from the aforementioned exact SDH algorithm, and this solution has running time that is unrelated to the system size N. We also develop a mathematical model to analyze the mechanism that leads to errors in the basic approximate algorithm. Our model provides insights on how the algorithm can be improved to achieve higher accuracy and efficiency. Such insights give rise to a new approximate algorithm with improved time/accuracy tradeoff. Experimental results confirm our analysis.
Vladimir Grupcev, Yongke Yuan, Yi-Cheng Tu, Shaoping Chen, Sagar Pandit, Michael Weng
IEEE Trans. Knowl. Data Eng.6
2012 Parallel reactive molecular dynamics: Numerical methods and algorithmic techniques
Hasan Metin Aktulga, Joseph C. Fogarty, Sagar Pandit, Ananth Grama
Parallel Comput.3
2009 Computing Distance Histograms Efficiently in Scientific Databases
abstract
Particle simulation has become an important research tool in many scientific and engineering fields. Data generated by such simulations impose great challenges to database storage and query processing. One of the queries against particle simulation data, the spatial distance histogram (SDH) query, is the building block of many high-level analytics, and requires quadratic time to compute using a straightforward algorithm. In this paper, we propose a novel algorithm to compute SDH based on a data structure called density map, which can be easily implemented by augmenting a quad-tree index. We also show the results of rigorous mathematical analysis of the time complexity of the proposed algorithm: our algorithm runs on ominus(N3/2) for two-dimensional data and ominus(N5/3) for three-dimensional data, respectively. We also propose an approximate SDH processing algorithm whose running time is unrelated to the input size N. Experimental results confirm our analysis and show that the approximate SDH algorithm achieves very high accuracy.
Yi-Cheng Tu, Shaoping Chen, Sagar Pandit
ICDE3