Ani Thakar

dblp:t/AniThakar · also Aniruddha R. Thakar · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
0since 2021 · last 2010
0000-0002-1631-0690ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 3Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 50% High-performance computing · 29% Distributed systems · 14%
Databases, data mining, and information retrieval
3 papers
Database system architecture and tuning · 68% Distributed and cloud data management · 18% Indexing and storage engines · 14%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 64% Computing education · 36%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cloud migration
0.112010
Migrating a (large) science database to the cloud · HPDC 2010
Cloud and datacenter computing
database migration
0.112010
Migrating a (large) science database to the cloud · HPDC 2010
High-performance computing
data transfer
0.112006
Bandwidth challenge - Transporting sloan digital sky survey data using SECTOR · SC 2006
Distributed systems
distributed data processing
0.112006
Bandwidth challenge - Transporting sloan digital sky survey data using SECTOR · SC 2006
High-performance computing › data transfer
high-speed data transfer
0.112006
Bandwidth challenge - Transporting sloan digital sky survey data using SECTOR · SC 2006
Distributed and cloud data management › database middleware
database web access
0.012002
The SDSS skyserver: public access to the sloan digital sky server data · SIGMOD Conference 2002
Indexing and storage engines
multidimensional indexing
0.012000
Designing and Mining Multi-Terabyte Astronomy Archives: The Sloan Digital Sky Survey · SIGMOD Conference 2000
Database system architecture and tuning
scientific data management
0.012000
Designing and Mining Multi-Terabyte Astronomy Archives: The Sloan Digital Sky Survey · SIGMOD Conference 2000
Transport protocols and congestion control
transport protocols
0.012006
Bandwidth challenge - Transporting sloan digital sky survey data using SECTOR · SC 2006
Computing education › STEM education
science education
0.012002
The SDSS skyserver: public access to the sloan digital sky server data · SIGMOD Conference 2002

Methods — techniques the papers use, named apart from their topics

performance benchmarking · 0.2parallel data transfer · 0.1
YearPublicationVenuePosition
2010 Migrating a (large) science database to the cloud
abstract
We report on attempts to put an existing scientific (astronomical) database -- the Sloan Digital Sky Survey (SDSS) science archive [1] - in the cloud. Based on our experience, it is either very frustrating or impossible at this time to migrate an existing, complex SQL Server database into current cloud service offerings such as Amazon (EC2) and Microsoft (SQL Azure). Certainly it is impossible to migrate a large database in excess of a TB, but even with (much) smaller databases, the limitations of cloud services make it very difficult to migrate the data to the cloud without making changes to the schema and settings (for example, inability to migrate a spatial indexing library, and several other user-defined functions and stored procedures) that would invalidate performance comparisons between cloud and on-premise versions. So it is not surprising that our preliminary performance comparisons show a very large (an order of magnitude) performance discrepancy with the Amazon cloud version of the SDSS database. We have also not yet investigated the performance tweaks that could be possible within the cloud.
Ani Thakar, Alex Szalay
HPDC1
2010 A Dynamic Data Middleware Cache for Rapidly-Growing Scientific Repositories
Tanu Malik, Philip Little, Amitabh Chaudhary, Ani Thakar
Middleware5
2006 Distributing the Sloan Digital Sky Survey Using UDT and Sector
abstract
In this paper, we describe a peer-to-peer storage system called Sector that is designed to access and transport large data sets over wide area high performance networks. We also describe our recent experience using Sector to distribute the Sloan Digital Sky Survey BESTDR4 catalog data.
Yunhong Gu, Robert L. Grossman, Alex Szalay, Ani Thakar
e-Science4
2006 Bandwidth challenge - Transporting sloan digital sky survey data using SECTOR
abstract
National Center for Data Mining at UICIn our SC06 BWC entry, we will transfer SDSS (Sloan Digital Sky Survey) Data Release 5 (DR5) between the SC06 show floor in Tampa and one of the NCDM labs on the UIC campus. We will use SECTOR, our newly developed distributed data space management system, to transfer DR5 in parallel between two Linux clusters in Tampa and Chicago, respectively. SECTOR transparently manages the file locating and data moving, while it employs UDT for actual data transfer. The data transfer will be from disk to disk over a 10Gb/s shared, router link between SC06 and UIC, via StarLight. We expect to reach 5Gb/s disk-to-disk data transfer rate between the two sites.
Robert L. Grossman, Yunhong Gu, Michal Sabala, Shirley Connelly, David Hanley, Joe Mambretti, Alex Szalay, Ani Thakar, Jan vandenBerg, Alainna Wonders
SC8
2006 Data mining middleware for wide-area high-performance networks
Robert L. Grossman, Yunhong Gu, David Hanley, Michal Sabala, Joe Mambretti, Alex Szalay, Ani Thakar, Kazumi Kumazoe, Yuji Oie, Yoonjoo Kwon, Woojin Seok
Future Gener. Comput. Syst.7
2005 When Database Systems Meet the Grid
María A. Nieto-Santisteban, Jim Gray 0001, Alex Szalay, James Annis, Ani Thakar, William O'Mullane
CIDR5
2005 Batch is Back: CasJobs, Serving Multi-TB Data on the Web
abstract
The Sloan Digital Sky Survey (SDSS) science database describes over 230 million objects and is over 1.6 TB in size. The SDSS Catalog Archive Server (CAS) provides several levels of query interface to the SDSS data via the SkyServer website. Most queries execute in seconds or minutes. However, some queries can take hours or days, either because they require non-index scans of the largest tables, or because they request very large result sets, or because they represent very complex aggregations of the data. These "monster queries" not only take a long time, they also affect response times for everyone else - one or more of them can clog the entire system. To ameliorate this problem, we developed a multiserver multiqueue batch job submission, execution, and tracking system for the CAS called CasJobs. The transfer of very large result sets from queries over the network is another serious problem. Statistics suggested that much of this data transfer is unnecessary; users would prefer to store results locally in order to allow further joins and filtering. To allow local analysis, a system was developed that gives users their own personal databases (MyDB) at the server side. Users may transfer data to their MyDB, and then perform further analysis before extracting it to their own machine. MyDB tables also provide a convenient way to share results of queries with collaborators without downloading them. CasJobs is built using SOAP XML Web services and has been in operation since May 2004.
William O'Mullane, Nolan Li, María A. Nieto-Santisteban, Alex Szalay, Ani Thakar
ICWS5
2003 SkyQuery: A Web Service Approach to Federate Databases
Tanu Malik, Alex Szalay, Tamás Budavári, Ani Thakar
CIDR4
2002 The SDSS skyserver: public access to the sloan digital sky server data
abstract
The SkyServer provides Internet access to the public Sloan Digital Sky Survey (SDSS) data for both astronomers and for science education. This paper describes the SkyServer goals and architecture. It also describes our experience operating the SkyServer on the Internet. The SDSS data is public and well-documented so it makes a good test platform for research on database algorithms and performance.
Alex Szalay, Jim Gray 0001, Ani Thakar, Peter Z. Kunszt, Tanu Malik, M. Jordan Raddick, Christopher Stoughton, Jan vandenBerg
SIGMOD Conference3
2000 Designing and Mining Multi-Terabyte Astronomy Archives: The Sloan Digital Sky Survey
abstract
The next-generation astronomy digital archives will cover most of the sky at fine resolution in many wavelengths, from X-rays, through ultraviolet, optical, and infrared. The archives will be stored at diverse geographical locations. One of the first of these projects, the Sloan Digital Sky Survey (SDSS) is creating a 5-wavelength catalog over 10,000 square degrees of the sky (see http://www.sdss.org/). The 200 million objects in the multi-terabyte database will have mostly numerical attributes in a 100+ dimensional space. Points in this space have highly correlated distributions.
Alex Szalay, Peter Z. Kunszt, Ani Thakar, Jim Gray 0001, Donald R. Slutz, Robert J. Brunner
SIGMOD Conference3
1997 Generating Validation Feedback for Automatic Interpretation of Informal Requirements
Walling R. Cyre, Ani Thakar
Formal Methods Syst. Des.2