VLDB 2026 Research / reviewers in the wild / expert
Chandrika Kamath 0001
dblp:62/3419-1
· DBLP profile ↗
28ranked-venue papers
10as first author
0since 2021 · last 2019
0000-0002-0188-8174ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-authorDatabases, data management, data science and information retrieval · 9 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorSystems, architecture and hardware · 5 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Data mining · 85% Information retrieval · 15% | |
| Artificial intelligence
1 paper |
Probabilistic and Bayesian machine learning · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Computational science and engineering · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › predictive modeling
regression |
0.1 | 2 | 2004 | Effective localized regression for damage detection in large complex mechanical structures · KDD 2004 Localized Prediction of Continuous Target Variables Using Hierarchical Clustering · ICDM 2003 |
Data mining › dimensionality reduction
feature selection |
0.0 | 1 | 2004 | Feature selection in scientific applications · KDD 2004 |
Information retrieval
filtering |
0.0 | 1 | 2004 | Feature selection in scientific applications · KDD 2004 |
Data mining › dimensionality reduction › feature selection
wrapper methods |
0.0 | 1 | 2004 | Feature selection in scientific applications · KDD 2004 |
Data mining
clustering |
0.0 | 1 | 2003 | Localized Prediction of Continuous Target Variables Using Hierarchical Clustering · ICDM 2003 |
Data mining › clustering
hierarchical clustering |
0.0 | 1 | 2003 | Localized Prediction of Continuous Target Variables Using Hierarchical Clustering · ICDM 2003 |
Machine learning › Probabilistic and Bayesian machine learning
clustering |
0.0 | 1 | 2002 | Learning to Classify Galaxy Shapes Using the EM Algorithm · NIPS 2002 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model |
0.0 | 1 | 2002 | Learning to Classify Galaxy Shapes Using the EM Algorithm · NIPS 2002 |
Computational science and engineering › astronomy
astronomical data analysis |
0.0 | 1 | 2002 | Learning to Classify Galaxy Shapes Using the EM Algorithm · NIPS 2002 |
Methods — techniques the papers use, named apart from their topics
clustering · 0.1mixture model · 0.1EM algorithm · 0.1partitioning · 0.0localization · 0.0classification model · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Intelligent Exploration of Large-Scale Data: What Can We Learn in Two Passes?abstractExploring large-scale data to determine an analysis approach is often made difficult by the sheer size of the data. If the characteristics of the data, including any variations within the data set, are not taken into account, our choice of algorithms and associated parameters may not be optimal, resulting in possibly inaccurate conclusions drawn from the data. Iteratively refining the analysis approach, as is normally done with smaller data sets, becomes prohibitively expensive for large-scale data. A typical solution is to randomly subsample the data set and determine the analysis algorithms using the characteristics of this subsample. In this paper, we propose the use of an improved sampling algorithm that is modified to identify well-distributed samples in a single pass through the data set. We then describe how we can use this subsample to probe the data in a second pass. Using very simple, low-cost algorithms, we demonstrate that the additional insight gained in this second pass can improve the process of analyzing large-scale data sets. Chandrika Kamath 0001 |
IEEE BigData | 1 |
| 2018 | Regression with small data sets: a case study using code surrogates in additive manufacturing
Chandrika Kamath 0001, Ya-Ju Fan |
Knowl. Inf. Syst. | 1 |
| 2017 | Learning to Compress Unstructured Mesh Data from SimulationsabstractThe amount of data output from a computer simulation has grown to terabytes and petabytes as increasingly-complex simulations are being run on massively-parallel systems. As we approach exa-flop computing in the next decade, it is expected that the I/O subsystem will not be able to write out these large volumes of data. In this paper, we explore the use of machine learning to compress the data before it is written out. Despite the computational constraints that limit us to using very simple learning algorithms, our results show that machine learning is a viable option for compressing unstructured data. Further, by using a better sampling algorithm to generate the training set, we can obtain more accurate results compared to random sampling, at no extra cost. Chandrika Kamath 0001 |
DSAA | 1 |
| 2015 | Practical Considerations in Applying Compressed Sensing to Simulation DataabstractThe move toward exascale computing for scientific simulations is placing new demands on compression techniques. It is expected that the I/O system will not be able to support the volume of data that will be written out. To enable quantitative analysis and scientific discovery, we need techniques that can compress high-dimensional simulation data with near-perfect reconstruction. In this work, we investigate Compressed Sensing (CS) to reduce the size of the data from a fusion simulation of a tokamak in 3 dimensions (Figure (a)). The computational domain of the simulation is a toroid, composed of 32 poloidal planes (shown in blue). Each plane has nearly 600,000 grid points, arranged irregularly, and distributed across multiple processors of a massively parallel system. Since these data are analyzed to understand the behavior of coherent structures (Figure (b)) over time, it is important that these structures remain unchanged after reconstruction using CS. We conducted several experiments to understand how best to apply CS to our data set. We used several metrics to investigate the effects of preprocessing, including scaling to improve the contrast in the data and thresholding to increase the sparsity. To determine the size of the compressed data that would enable near-perfect reconstruction, we evaluated the quality of reconstruction (shown in Figures (c) and (d) using the R2 metric) as we varied the percentage of compression (m/n) for various levels of sparsity (k/n) in the data. We found that a successful application of CS is bounded by the percentage of sparsity in the data - the data have to be sparse enough for compression using CS, but not so sparse that it is more cost effective to just write out the locations and values of the non-zero data points. Our experiments also indicated that scaling the data is very helpful and thresholding helps both with compression and the coherent structure analysis performed on the data. Ya-Ju Fan, Chandrika Kamath 0001 |
DCC | 2 |
| 2015 | Identifying and Exploiting Diurnal Motifs in Wind Generation Time Series DataabstractWind energy is scheduled on the power grid using 0–6 h ahead forecasts generated from computer simulations or historical data. When the forecasts are inaccurate, control room operators use their expertise, as well as the actual generation from previous days, to estimate the amount of energy to schedule. However, this is a challenge, and it would be useful for the operators to have additional information they can exploit to make better informed decisions. In this paper, we use techniques from time series analysis to determine if there are motifs, or frequently occurring diurnal patterns in wind generation data. We compare two different representations of the data and four different ways of identifying the number of motifs. Using data from wind farms in Tehachapi Pass and mid-Columbia Basin, we describe our findings and discuss how these motifs can be used to guide scheduling decisions. Ya-Ju Fan, Chandrika Kamath 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2015 | Evaluation of connected-component labeling algorithms for distributed-memory systems
Jeremy Iverson, Chandrika Kamath 0001, George Karypis |
Parallel Comput. | 2 |
| 2014 | Incremental SVD for Insight into Wind GenerationabstractIn this paper, we formulate the problem of predicting wind generation as one of streaming data analysis. We want to understand if it is possible to use the weather data in a time window just before the current time to gain insight into how the wind generation might behave in a time interval just after the current time. Specifically, we use a singular value decomposition of the weather data, and how that the number of singular values and the largest singular value can be used to predict the magnitude of the change in the generation in the near future. The analysis uses an incremental algorithm based on a sliding window for reduced computational costs. Chandrika Kamath 0001, Ya-Ju Fan |
ICMLA | 1 |
| 2012 | Fast and Effective Lossy Compression Algorithms for Scientific Datasets
Jeremy Iverson, Chandrika Kamath 0001, George Karypis |
Euro-Par | 2 |
| 2012 | Finding Motifs in Wind Generation Time Series DataabstractWind energy is scheduled on the power grid using 0-6 hour ahead forecasts generated from computer simulations or historical data. When the forecasts are inaccurate, control room operators use their expertise, as well as the actual generation from previous days, to estimate the amount of energy to schedule. However, this is a challenge, and it would be useful for the operators to have additional information they can exploit to make better informed decisions. In this paper, we use techniques from time series analysis to determine if there are motifs, or frequently occurring diurnal patterns in wind generation data. Using data from wind farms in Tehachapi Pass and mid-Columbia Basin, we describe our findings and discuss how these motifs can be used to guide scheduling decisions. Chandrika Kamath 0001, Ya-Ju Fan |
ICMLA (2) | 1 |
| 2008 | Tracking non-rigid structures in computer simulationsabstractA key challenge in tracking moving objects is the correspondence problem, that is, the correct propagation of object labels from one time step to another. This is especially true when the objects are non-rigid structures, changing shape, and merging and splitting over time. In this work, we describe a general approach to tracking thousands of non-rigid structures in an image sequence. We show how we can minimize memory requirements and generate accurate results while working with only two frames of the sequence at a time. We demonstrate our results using data from computer simulations of a fluid-mix problem. Abel Gezahegne, Chandrika Kamath 0001 |
ICIP | 2 |
| 2007 | Estimating Missing Features to Improve Multimedia RetrievalabstractRetrieval in a multimedia database usually involves combining information from different modalities of data, such as text and images. However, all modalities of the data may not be available to form the query. The results from such a partial query are often less than satisfactory. In this paper, we present an approach to complete a partial query by estimating the missing features in the query. Our experiments with a database of images and their associated captions show that, with an initial text-only query, our completion method has similar performance to a full query with both image and text features. In addition, when we use relevance feedback, our approach outperforms the results obtained using a full query. Abraham Bagherjeiran, Nicole S. Love, Chandrika Kamath 0001 |
ICIP (2) | 3 |
| 2007 | Image Analysis for Validation of Simulations of a Fluid Mix ProblemabstractAs computer simulations gain acceptance for the modeling of complex physical phenomena, there is an increasing need to validate these simulation codes by comparing them to experiments. Currently, this is done qualitatively, using a visual approach. This is obviously very subjective and more quantitative metrics are needed, especially to identify simulations which are closer to experiments than other simulations. In this paper, we show how image processing techniques can be effectively used in such comparisons. Using an example from the problem of mixing of two fluids, we show that we can quantitatively compare experimental and simulation images by extracting higher level features to characterize the objects in the images. Chandrika Kamath 0001, Paul L. Miller |
ICIP (3) | 1 |
| 2006 | Graph-based Methods for Orbit ClassificationabstractAn important step in the quest for low-cost fusion power is the ability to perform and analyze experiments in prototype fusion reactors. One of the tasks in the analysis is the classification of orbits in Poincaré plots generated by the particles in a fusion reactor as they move within the toroidal device. In this paper, we describe the use of graph-based methods to extract features from orbits. These features are then used to classify the orbits into several categories. Our results show that existing machine learning algorithms are successful in classifying orbits with few points, a situation which can arise in data from experiments. Abraham Bagherjeiran, Chandrika Kamath 0001 |
SDM | 2 |
| 2005 | An empirical comparison of combinations of evolutionary algorithms and neural networks for classification problemsabstractThere are numerous combinations of neural networks (NNs) and evolutionary algorithms (EAs) used in classification problems. EAs have been used to train the networks, design their architecture, and select feature subsets. However, most of these combinations have been tested on only a few data sets and many comparisons are done inappropriately measuring the performance on training data or without using proper statistical tests to support the conclusions. This paper presents an empirical evaluation of eight combinations of EAs and NNs on 15 public-domain and artificial data sets. Our objective is to identify the methods that consistently produce accurate classifiers that generalize well. In most cases, the combinations of EAs and NNs perform equally well on the data sets we tried and were not more accurate than hand-designed neural networks trained with simple backpropagation. Erick Cantú-Paz, Chandrika Kamath 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2004 | Feature selection in scientific applicationsabstractNumerous applications of data mining to scientific data involve the induction of a classification model. In many cases, the collection of data is not performed with this task in mind, and therefore, the data might contain irrelevant or redundant features that affect negatively the accuracy of the induction algorithms. The size and dimensionality of typical scientific data make it difficult to use any available domain information to identify features that discriminate between the classes of interest. Similarly, exploratory data analysis techniques have limitations on the amount and dimensionality of the data they can process effectively. In this paper, we describe applications of efficient feature selection methods to data sets from astronomy, plasma physics, and remote sensing. We use variations of recently proposed filter methods as well as traditional wrapper approaches, where practical. We discuss the general challenges of feature selection in scientific datasets, the strategies for success that were common among our diverse applications, and the lessons learned in solving these problems. Erick Cantú-Paz, Shawn D. Newsam, Chandrika Kamath 0001 |
KDD | 3 |
| 2004 | Effective localized regression for damage detection in large complex mechanical structuresabstractIn this paper, we propose a novel data mining technique for the efficient damage detection within the large-scale complex mechanical structures. Every mechanical structure is defined by the set of finite elements that are called structure elements. Large-scale complex structures may have extremely large number of structure elements, and predicting the failure in every single element using the original set of natural frequencies as features is exceptionally time-consuming task. Traditional data mining techniques simply predict failure in each structure element individually using global prediction models that are built considering all data records. In order to reduce the time complexity of these models, we propose a localized clustering-regression based approach that consists of two phases: (1) building a local cluster around a data record of interest and (2) predicting an intensity of damage only in those structure elements that correspond to data records from the built cluster. For each test data record, we first build a cluster of data records from training data around it. Then, for each data record that belongs to discovered cluster, we identify corresponding structure elements and we build a localized regression model for each of these structure elements. These regression models for specific structure elements are constructed using only a specific set of relevant natural frequencies and merely those data records that correspond to the failure of that structure element. Experiments performed on the problem of damage prediction in a large electric transmission tower frame indicate that the proposed localized clustering-regression based approach is significantly more accurate and more computationally efficient than our previous hierarchical clustering approach, as well as global prediction models. Aleksandar Lazarevic, Ramdev Kanapady, Chandrika Kamath 0001 |
KDD | 3 |
| 2004 | Robust techniques for background subtraction in urban traffic videoabstractIdentifying moving objects from a video sequence is a fundamental and critical task in many computer-vision applications. A common approach is to perform background subtraction, which identifies moving objects from the portion of a video frame that differs significantly from a background model. There are many challenges in developing a good background subtraction algorithm. First, it must be robust against changes in illumination. Second, it should avoid detecting non-stationary background objects such as swinging leaves, rain, snow, and shadow cast by moving objects. Finally, its internal background model should react quickly to changes in background such as starting and stopping of vehicles. In this paper, we compare various background subtraction algorithms for detecting moving vehicles and pedestrians in urban traffic video sequences. We consider approaches varying from simple techniques such as frame differencing and adaptive median filtering, to more sophisticated probabilistic modeling techniques. While complicated techniques often produce superior performance, our experiments show that simple techniques such as adaptive median filtering can produce good results with much lower computational complexity. Sen-Ching S. Cheung, Chandrika Kamath 0001 |
VCIP | 2 |
| 2004 | Investigation of implicit active contours for scientific image segmentationabstractThe use of partial differential equations in image processing has become an active area of research in the last few years. In particular, active contours are being used for image segmentation, either explicitly as snakes, or implicitly through the level set approach. In this paper, we consider the use of the implicit active contour approach for segmenting scientific images of pollen grains obtained using a scanning electron microscope. Our goal is to better understand the pros and cons of these techniques and to compare them with the traditional approaches such as the Canny and SUSAN edge detectors. The preliminary results of our study show that the level set method is computationally expensive and requires the setting of several different parameters. However, it results in closed contours, which may be useful in separating objects from the background in an image. Sisira Weeratunga, Chandrika Kamath 0001 |
VCIP | 2 |
| 2003 | Localized Prediction of Continuous Target Variables Using Hierarchical ClusteringabstractWe propose a novel technique for the efficient prediction of multiple continuous target variables from high-dimensional and heterogeneous data sets using a hierarchical clustering approach. The proposed approach consists of three phases applied recursively: partitioning, localization and prediction. In the partitioning step, similar target variables are grouped together by a clustering algorithm. In the localization step, a classification model is used to predict which group of target variables is of particular interest. If the identified group of target variables still contains a large number of target variables, the partitioning and localization steps are repeated recursively and the identified group is further split into subgroups with more similar target variables. When the number of target variables per identified subgroup is sufficiently small, the third step predicts target variables using localized prediction models built from only those data records that correspond to the particular subgroup. Experiments performed on the problem of damage prediction in complex mechanical structures indicate that our proposed hierarchical approach is computationally more efficient and more accurate than straightforward methods of predicting each target variable individually or simultaneously using global prediction models. Aleksandar Lazarevic, Ramdev Kanapady, Chandrika Kamath 0001, Vipin Kumar 0001, Kumar K. Tamma |
ICDM | 3 |
| 2003 | Evolving neural networks to identify bent-double galaxies in the FIRST survey
Erick Cantú-Paz, Chandrika Kamath 0001 |
Neural Networks | 2 |
| 2003 | Inducing oblique decision trees with evolutionary algorithmsabstractThis paper illustrates the application of evolutionary algorithms (EAs) to the problem of oblique decision-tree (DT) induction. The objectives are to demonstrate that EAs can find classifiers whose accuracy is competitive with other oblique tree construction methods, and that, at least in some cases, this can be accomplished in a shorter time. We performed experiments with a (1+1) evolution strategy and a simple genetic algorithm on public domain and artificial data sets, and compared the results with three other oblique and one axis-parallel DT algorithms. The empirical results suggest that the EAs quickly find competitive classifiers, and that EAs scale up better than traditional methods to the dimensionality of the domain and the number of instances used in training. In addition, we show that the classification accuracy improves when the trees obtained with the EAs are combined in ensembles, and that sometimes it is possible to build the ensemble of evolutionary trees in less time than a single traditional oblique tree. Erick Cantú-Paz, Chandrika Kamath 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2002 | Evolving Neural Networks For The Classification Of Galaxies
Erick Cantú-Paz, Chandrika Kamath 0001 |
GECCO | 2 |
| 2002 | Learning to Classify Galaxy Shapes Using the EM AlgorithmabstractWe describe the application of probabilistic model-based learning to the problem of automatically identifying classes of galaxies, based on both morphological and pixel intensity characteristics. The EM algorithm can be used to learn how to spatially orient a set of galaxies so that they are geometrically aligned. We augment this “ordering-model” with a mixture model on objects, and demonstrate how classes of galaxies can be learned in an unsupervised manner using a two-level EM algorithm. The resulting models provide highly accurate classi£cation of galaxies in cross-validation experiments. 1 Introduction and Background The £eld of astronomy is increasingly data-driven as new observing instruments permit the rapid collection of massive archives of sky image data. In this paper we investigate the problem of identifying bent-double radio galaxies in the FIRST (Faint Images of the Radio Sky at Twenty-cm) Survey data set [1]. FIRST produces large numbers of radio images of the deep sky using the Very Large Array at the National Radio Astronomy Observatory. It is scheduled to cover more that 10,000 square degrees of the northern and southern caps (skies). Of particular scienti£c interest to astronomers is the identi£cation and cataloging of sky objects with a “bent-double” morphology, indicating clusters of galaxies ([8], see Figure 1). Due to the very large number of observed deep-sky radio sources, (on the order of 106 so far) it is infeasible for the astronomers to label all of them manually. The data from the FIRST Survey (http://sundog.stsci.edu/) is available in both raw image format and in the form of a catalog of features that have been automatically derived from the raw images by an image analysis program [8]. Each entry corresponds to a single detectable “blob” of bright intensity relative to the sky background: these entries are called Figure 1: 4 examples of radio-source galaxy images. The two on the left are labelled as “bent-doubles” and the two on the right are not. The con£gurations on the left have more “bend” and symmetry than the two non-bent-doubles on the right. components. The “blob” of intensities for each component is £tted with an ellipse. The ellipses and intensities for each component are described by a set of estimated features such as sky position of the centers (RA (right ascension) and Dec (declination)), peak density ¤ux and integrated ¤ux, root mean square noise in pixel intensities, lengths of the major and minor axes, and the position angle of the major axis of the ellipse counterclockwise from the north. The goal is to £nd sets of components that are spatially close and that resemble a bent-double. In the results in this paper we focus on candidate sets of components that have been detected by an existing spatial clustering algorithm [3] where each set consists of three components from the catalog (three ellipses). As of the year 2000, the catalog contained over 15,000 three-component con£gurations and over 600,000 con£gurations total. The set which we use to build and evaluate our models consists of a total of 128 examples of bent-double galaxies and 22 examples of non-bent-double con£gurations. A con£guration is labelled as a bent-double if two out of three astronomers agree to label it as such. Note that the visual identi£cation process is the bottleneck in the process since it requires signi£cant time and effort from the scientists, and is subjective and error-prone, motivating the creation of automated methods for identifying bent-doubles. Three-component bent-double con£gurations typically consist of a center or “core” com- ponent and two other side components called “lobes”. Previous work on automated classi£- cation of three-component candidate sets has focused on the use of decision-tree classi£ers using a variety of geometric and image intensity features [3]. One of the limitations of the decision-tree approach is its relative in¤exibility in handling uncertainty about the object being classi£ed, e.g., the identi£cation of which of the three components should be treated as the core of a candidate object. A bigger limitation is the £xed size of the feature vec- tor. A primary motivation for the development of a probabilistic approach is to provide a framework that can handle uncertainties in a ¤exible coherent manner. 2 Learning to Match Orderings using the EM Algorithm We denote a three-component con£guration by C = (c 1; c2; c3), where the ci’s are the components (or “blobs”) described in the previous section. Each component cx is repre- sented as a feature vector, where the speci£c features will be de£ned later. Our approach focuses on building a probabilistic model for bent-doubles: p (C) = p (c1; c2; c3), the like- lihood of the observed ci under a bent-double model where we implicitly condition (for now) on the class “bent-double.” By looking at examples of bent-double galaxies and by talking to the scientists study- ing them, we have been able to establish a number of potentially useful characteristics of the components, the primary one being geometric symmetry. In bent-doubles, two of the components will look close to being mirror images of one another with respect to a line through the third component. We will call mirror-image components lobe compo- Sergey Kirshner, Igor V. Cadez, Padhraic Smyth, Chandrika Kamath 0001 |
NIPS | 4 |
| 2002 | Approximate Splitting for Ensembles of Trees using Histogramsabstract1 Introduction Ensembles of classifiers have become an active topic of research in the data mining community. They not only provide a simple way of improving the accuracy of the classifier [3, 16, 24, 2], but also have the potential for on-line classification of large databases that do not fit into memory [4]. In addition, some approaches to the generation of ensembles can be easily parallelized, enabling a reduction in the time taken to create the classifier on a multiprocessor system [17]. There are several different ways in which ensembles can be generated and the resulting output combined to classify new instances. Implicit in many of these ensembles is the concept of randomness that is introduced either through the randomization of the training set, or the randomization of the classifier itself. Chandrika Kamath 0001, Erick Cantú-Paz, David Littau |
SDM | 1 |
| 2000 | Using Evolutionary Algorithms to Induce Oblique Decision Trees
Erick Cantú-Paz, Chandrika Kamath 0001 |
GECCO | 2 |
| 1990 | Implementation of two projection methods on a shared memory multiprocessor: DEC VAX 6240
Chandrika Kamath 0001, Sisira Weeratunga |
Parallel Comput. | 1 |
| 1989 | A projection method for solving nonsymmetric linear systems on multiprocessors
Chandrika Kamath 0001, Ahmed H. Sameh |
Parallel Comput. | 1 |
| 1982 | Implementation and performance prediction of some parallel algorithms on the Plexus microcomputer network
Chandrika Kamath 0001, Virendrakumar C. Bhavsar |
Microprocessing and Microprogramming | 1 |