Christoph F. Eick

dblp:e/CFEick · DBLP profile ↗
← Back
36ranked-venue papers in the field
9as first author
4since 2021 · last 2026
0000-0002-6798-103XORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 13 (6 first)Data Mining & Knowledge Discovery · 13 (2 first)Other / Interdisciplinary · 5Big Data, Cloud & Distributed Data Systems · 3Information Retrieval & Web Search · 1 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Exploring Topographic Data with Transitive Closure on Graphs
Adam Nelson-Archer, Christoph F. Eick, Carlos Ordonez 0001
DEXA (1)2
2026 Proteus: A system for combining paths collected using crowdsensing
Raunak Sarbajna, Christoph F. Eick, Jianyuan Ni
GeoInformatica2
2024 Norma: A Framework for Finding Threshold Associations Between Continuous Variables Using Point-wise Functions
abstract
This paper introduces Norma, a novel association analysis framework for continuous spatial variables. Unlike other point-based spatial analysis methods, Norma associates a point-wise function with each continuous variable-e.g., a function that returns poverty rates for each location in New Mexico, and then finds interesting associations between continuous variables by analyzing relationships between the respective point-wise functions. Moreover, the paper introduces a new association called Continuous Variable Threshold (CVT) pattern, aiming to identify a pair of thresholds within the domains of two continuous variables which exhibit strong associations within an observation area. For example, it may unveil a strong association between COVID-19 infection rates above 2% and poverty rates above 15% in New Mexico. To find such associations, Norma employs a novel interestingness function, which measures agreement with respect to hotspots where each point-wise function exceeds the associated threshold. To be able to compute such hotspots, the paper proposes a novel grid-based spatial hotspot-growing algorithm which computes regions of a point-wise function above a given threshold. Furthermore, Norma introduces a measure area under the curve (AUC) for assessing variable relatedness based on observed CVT associations and an algorithm to determine the AUC in 2D space. Finally, the results of a comparative case study are presented which uses county-level COVID-19 infection rates and nineteen socio-economic variables from the contiguous United States which demonstrates the merit of Norma’s association analysis approach.
Md. Mahin, Christoph F. Eick
IEEE Big Data2
2021 Learning Domain-Specific Word Embeddings from COVID-19 Tweets
abstract
The COVID-19 global pandemic has been a major catastrophic event that impacted the world’s economy. During the pandemic there was a rise in the use of social media such as Twitter by people to express their reactions and responses to the global pandemic. This drove researchers to analyze these micro-blogging texts, using natural language processing (NLP) methods, to understand information inherent in those texts. Most of these NLP tasks employ the use of word embeddings in training neural network models. These word embeddings are mainly trained on general text corpus which produce sub-optimal performance when used in domain-specific NLP tasks such as in COVID-19 related tweets. In this paper, we present a learned COVID-19 tweets domain-specific word embeddings for use in COVID-19 related tweets NLP tasks. Our evaluation results show that our domain-specific COVID-19 tweets word embeddings perform better than pretrained general word embeddings in a downstream domain-specific NLP task. Our COVID-19 tweets word embeddings are available for use by researchers who wish to perform downstream NLP tasks with pretrained domain-specific COVID-19 tweets word embeddings.
Steve Aibuedefe Aigbe, Christoph F. Eick
IEEE BigData2
2017 ST-COPOT: Spatio-temporal Clustering with Contour Polygon Trees
abstract
Nowadays, growing effort has been put to develop spatio-temporal clustering approaches that are capable of discovering interesting patterns in large spatio-temporal data streams. In this paper, we propose a 3-phase serial, density-contour based clustering algorithm called ST-COPOT, which can identify spatio-temporal cluster at multiple levels of density granularity. ST-COPOT takes the point cloud data as input and divides it into batches, next, it employs a non-parametric kernel density estimation approach and contouring algorithms to obtain spatial clusters; at last, spatio-temporal clusters are formed by identifying continuing relationships between spatial clusters in consecutive batches. Moreover, a novel data structure called contour polygon tree is introduced as a compact representation of the spatial clusters obtained for each batch for different density thresholds, and a family of novel distance functions that operate on contour polygon trees are proposed to identify continuing clusters. The experimental results on NYC taxi trips data show that ST-COPOT can effectively discover interesting spatio-temporal patterns in taxi pickup location streams.
Yongli Zhang, Christoph F. Eick
SIGSPATIAL/GIS2
2017 Supervised Taxonomies - Algorithms and Applications
abstract
This paper focuses on a new type of taxonomy called supervised taxonomy (ST). Supervised taxonomies are generated considering background information concerning class labels in addition to distance metrics, and are capable of capturing class-uniform regions in a dataset. A hierarchical, agglomerative clustering algorithm, called STAXAC that generates STs is proposed and its properties are analyzed. Experimental results are presented that show that STAXAC produces purer taxonomies than the neighbor-joining (NJ) algorithm - a very popular taxonomy generation algorithm. We introduced novel measures and algorithms that assess classification complexity, class modality, and show that STs can be used as the main input of an effective data-editing tool to enhance the accuracy of k-nearest neighbor classifiers. We demonstrated in our experimental evaluation that assessing the classification complexity of a ST provides us with a good estimate of the difficulty of the classification problem at hand. Moreover, a class modality discovery tool (CMD) has been provided that - based on a domain expert's notion of what constitutes a “note-worthy” subclass-determines if specific classes in the dataset are zero-modal, unimodal, and multi-modal.
Paul K. Amalaman, Christoph F. Eick
IEEE Trans. Knowl. Data Eng.2
2015 An optimized interestingness hotspot discovery framework for large gridded spatio-temporal datasets
abstract
We define interestingness hotspots as contiguous regions in space which are interesting based on a domain expert's notion of interestingness captured by an interestingness function. This paper centers on finding interestingness hotspots on very large gridded datasets which are quite common in scientific computing. Mining large gridded datasets with a lot of variables and measurements requires a scalable framework that can process large amounts of data in an efficient way. In our recent work, we proposed a computational framework which discovers interestingness hotspots in gridded datasets using a 3-step approach which consists of seeding, hotspot growing and post-processing steps. In this paper, we significantly improve the efficiency of the framework by utilizing parallel processing and employing more efficient data structures and algorithms. We propose a novel heap-based hotspot growing algorithm which brings down the cost of hotspot growing phase significantly. In addition, we propose a graph-based preprocessing algorithm which decreases the number of hotspots grown by merging some hotspot seeds. Other improvements to the framework involve incremental calculation of interestingness functions, and growing hotspots in parallel. The improved framework is evaluated in a case study for a very large 4-dimensional gridded air pollution dataset in which we find interestingness hotspots with respect to pollutants.
Fatih Akdag, Christoph F. Eick
IEEE BigData2
2014 A polygon-based clustering and analysis framework for mining spatial datasets
Christoph F. Eick
GeoInformatica2
2011 A framework for regional association rule mining and scoping in spatial datasets
Wei Ding 0003, Christoph F. Eick, Xiaojing Yuan, Jing Wang 0007, Jean-Philippe Nicot
GeoInformatica2
2011 Controlling patterns of geospatial phenomena
Tomasz F. Stepinski, Wei Ding 0003, Christoph F. Eick
GeoInformatica3
2011 GAC-GEO: a generic agglomerative clustering framework for geo-referenced datasets
Rachsuda Jiamthapthaksin, Christoph F. Eick, Seungchan Lee
Knowl. Inf. Syst.2
2010 Correspondence Clustering: An Approach to Cluster Multiple Related Spatial Datasets
Vadeerat Rinsurongkawong, Christoph F. Eick
PAKDD (1)2
2009 A Framework for Multi-Objective Clustering and Its Application to Co-Location Mining
Rachsuda Jiamthapthaksin, Christoph F. Eick, Ricardo Vilalta
ADMA2
2009 An architecture and algorithms for multi-run clustering
abstract
This paper addresses two main challenges for clustering which require extensive human effort: selecting appropriate parameters for an arbitrary clustering algorithm and identifying alternative clusters. We propose an architecture and a concrete system MR-CLEVER for multi-run clustering that integrates active learning with clustering algorithms. The key hypothesis of this work is that better clustering results can be obtained by combining clusters that originate from multiple runs of clustering algorithms. By defining states that represent parameter settings of a clustering algorithm, the proposed architecture actively learns a state utility function. The utility of a parameter setting is assessed based on clustering run-time, quality and novelty of the obtained clusters. Furthermore, the utility function plays an important role in guiding the clustering algorithm to seek novel solutions. Cluster novelty measures are introduced for this purpose. Finally, we also contribute a cluster summarization algorithm that assembles a final clustering as a combination of high-quality clusters originating from multiple runs. Merits of our proposed system are that it is generic and therefore can be used in conjunction with different clustering algorithms, and it reduces human effort for selecting the parameters, for comparing clustering results and for assembling clustering results. We evaluate the proposed system in conjunction with a representative based clustering algorithm namely CLEVER for a challenging data mining task involving an earthquake dataset. The obtained results demonstrate that, in comparison to the best single-run clustering, multi-run clustering discovers solutions of higher quality.
Rachsuda Jiamthapthaksin, Christoph F. Eick, Vadeerat Rinsurongkawong
CIDM2
2009 REG^2: a regional regression framework for geo-referenced datasets
abstract
Traditional regression analysis derives global relationships between variables and neglects spatial variations in variables. Hence they lack the ability to systematically discover regional relationships and to build better models that use this regional knowledge to obtain higher prediction accuracies. Since most relationships in spatial datasets are regional, there is a great need for regional regression methods that derive regional regression functions that reflect different spatial characteristics of different regions. This paper proposes a novel regional regression framework that first discovers interesting regions showing strong regional relationships between the dependent and the independent variables, and then builds a prediction model with a regional regression function associated with each region. Interesting regions are identified by running a representative-based clustering algorithm that maximizes an externally plugged in fitness function. In this work, we propose two fitness functions: an R-squared based fitness function and an AIC-based fitness function to handle overfitting better. We evaluate our framework in two case studies; (1) identifying causes of arsenic contamination in Texas water wells and (2) Boston Housing dataset determining spatially varying effects of house properties on house prices. We demonstrated that our framework effectively identifies interesting regions and builds better prediction systems that rely on regional models.
Oner Ulvi Celepcikay, Christoph F. Eick
GIS2
2009 Change Analysis in Spatial Data by Combining Contouring Algorithms with Supervised Density Functions
Chun-Sheng Chen, Vadeerat Rinsurongkawong, Christoph F. Eick, Michael D. Twa
PAKDD3
2008 Finding regional co-location patterns for sets of continuous variables in spatial datasets
abstract
This paper proposes a novel framework for mining regional co-location patterns with respect to sets of continuous variables in spatial datasets. The goal is to identify regions in which multiple continuous variables with values from the wings of their statistical distribution are co-located. A co-location mining framework is introduced that operates in the continuous domain and which views regional co-location mining as a clustering problem in which an externally given fitness function has to be maximized. Interestingness of co-location patterns is assessed using products of z-scores of the relevant continuous variables. The proposed framework is evaluated by a domain expert in a case study that analyzes Arsenic contamination in Texas water wells centering on regional co-location patterns. Our approach is able to identify known and unknown regional co-location patterns, and different sets of algorithm parameters lead to the characterization of Arsenic distribution at different scales. Moreover, inconsistent colocation sets are found for regions in South Texas and West Texas that can be clearly attributed to geological differences in the two regions, emphasizing the need for regional co-location mining techniques. Moreover, a novel, prototype-based region discovery algorithm named CLEVER is introduced that uses randomized hill climbing, and searches a variable number of clusters and larger neighborhood sizes.
Christoph F. Eick, Rachana Parmar, Wei Ding 0003, Tomasz F. Stepinski, Jean-Philippe Nicot
GIS1
2008 Discovering controlling factors of geospatial variables
abstract
Efficient means of determining factors controlling spatial distribution of an environmental class variable are of significant interest in Earth science. In this paper, we present a method for automated discovery of controlling factors by mining for emerging patterns in a database constructed from the fusion of several explanatory datasets. We introduce a new definition of pattern support to account for spatial character of the data and systematically evaluate the effectiveness of our technique using a real-world application pertaining to density of vegetation cover. Experimental results show that our method can successfully identify controlling factors for the presence of high vegetation cover.
Tomasz F. Stepinski, Wei Ding 0003, Christoph F. Eick
GIS3
2008 Towards Region Discovery in Spatial Datasets
Wei Ding 0003, Rachsuda Jiamthapthaksin, Rachana Parmar, Tomasz F. Stepinski, Christoph F. Eick
PAKDD6
2007 MOSAIC: A Proximity Graph Approach for Agglomerative Clustering
Jiyeon Choo, Rachsuda Jiamthapthaksin, Chun-Sheng Chen, Oner Ulvi Celepcikay, Christian Giusti, Christoph F. Eick
DaWaK6
2007 On supervised density estimation techniques and their application to spatial data mining
abstract
The basic idea of traditional density estimation is to model the overall point density analytically as the sum of influence functions of data points. However, traditional density estimation techniques only consider the location of a point. Supervised density estimation techniques, on the other hand, additionally consider a variable of interest that is associated with a point. Density in supervised density estimation is measured as the product of an influence function with the variable of interest. Based on this novel idea, a supervised density-based clustering named SCDE is introduced and discussed in detail. The SCDE algorithm forms clusters by associating data points with supervised density attractors which represent maxima and minima of a supervised density function.
Christoph F. Eick, Chun-Sheng Chen
GIS2
2006 A Framework for Regional Association Rule Mining in Spatial Datasets
abstract
The immense explosion of geographically referenced data calls for efficient discovery of spatial knowledge. One of the special challenges for spatial data mining is that information is usually not uniformly distributed in spatial datasets. Consequently, the discovery of regional knowledge is of fundamental importance for spatial data mining. This paper centers on discovering regional association rules in spatial datasets. In particular, we introduce a novel framework to mine regional association rules relying on a given class structure. A reward-based regional discovery methodology is introduced, and a divisive, grid-based supervised clustering algorithm is presented that identifies interesting subregions in spatial datasets. Then, an integrated approach is discussed to systematically mine regional rules. The proposed framework is evaluated in a real-world case study that identifies spatial risk patterns from arsenic in the Texas water supply.
Wei Ding 0003, Christoph F. Eick, Jing Wang 0007, Xiaojing Yuan
ICDM2
2006 Discovery of Interesting Regions in Spatial Data Sets Using Supervised Clustering
Christoph F. Eick, Banafsheh Vaezian, Jing Wang 0007
PKDD1
2005 Adaptive Clustering: Obtaining Better Clusters Using Feedback and Past Experience
abstract
Adaptive clustering uses external feedback to improve cluster quality; past experience serves to speed up execution time. An adaptive clustering environment is proposed that uses Q-learning to learn the reward values of successive data clusterings. Adaptive clustering supports the reuse of clusterings by memorizing what worked well in the past. It has the capability of exploring multiple paths in parallel when searching for good clusters. In a case study, we apply adaptive clustering to instance-based learning relying on a distance function modification approach. A distance function adaptation scheme that uses external feedback is proposed and compared with other distance function learning approaches. Experimental results indicate that the use of adaptive clustering leads to significant improvements of instance-based learning techniques, such as k-nearest neighbor classifiers. Moreover, as a by-product a new instance-based learning technique is introduced that classifies examples by solely using cluster representatives; this technique shows high promise in our experimental evaluation.
Abraham Bagherjeiran, Christoph F. Eick, Chun-Sheng Chen, Ricardo Vilalta
ICDM2
2005 A database clustering methodology and tool
Tae-Wan Ryu, Christoph F. Eick
Inf. Sci.2
2004 Using Representative-Based Clustering for Nearest Neighbor Dataset Editing
abstract
The goal of dataset editing in instance-based learning is to remove objects from a training set in order to increase the accuracy of a classifier. For example, Wilson editing removes training examples that are misclassified by a nearest neighbor classifier so as to smooth the shape of the resulting decision boundaries. This paper revolves around the use of representative-based clustering algorithms for nearest neighbor dataset editing. We term this approach supervised clustering editing. The main idea is to replace a dataset by a set of cluster prototypes. A clustering approach called supervised clustering is introduced for this purpose. Our empirical evaluation using eight UCI datasets shows that both Wilson and supervised clustering editing improve accuracy on more than 50% of the datasets tested. However, supervised clustering editing achieves four times higher compression rates than Wilson editing.
Christoph F. Eick, Nidal M. Zeidat, Ricardo Vilalta
ICDM1
2003 Class Decomposition via Clustering: A New Framework for Low-Variance Classifiers
abstract
We propose a preprocessing step to classification that applies a clustering algorithm to the training set to discover local patterns in the attribute or input space. We demonstrate how this knowledge can be exploited to enhance the predictive accuracy of simple classifiers. Our focus is mainly on classifiers characterized by high bias but low variance (e.g., linear classifiers); these classifiers experience difficulty in delineating class boundaries over the input space when a class distributes in complex ways. Decomposing classes into clusters makes the new class distribution easier to approximate and provides a viable way to reduce bias while limiting the growth in variance. Experimental results on real-world domains show an advantage in predictive accuracy when clustering is used as a preprocessing step to classification.
Ricardo Vilalta, Murali-Krishna Achari, Christoph F. Eick
ICDM3
1998 From ordered beliefs to numbers: How to elicit numbers without asking for them (doable but computationally difficult)
abstract
One of the most important parts of designing an expert system is elicitation of the expert's knowledge. This knowledge usually consists of facts and rules. Eliciting these rules and facts is relatively easy: the more complicated task is assigning weights (numerical or interval-valued degrees of belief) to different statements from the knowledge base. Experts often cannot quantify their degrees of belief, but they can order them (by suggesting which statements are more reliable). It is, therefore, reasonable to try to reconstruct the degrees of belief from such an ordering.In this paper, we analyze when such a reconstruction is possible, whether it lead to unique values of degrees of belief, and how computationally complicated the corresponding reconstruction problem can be. © 1998 John Wiley & Sons, Inc.
Brian Cloteaux, Christoph F. Eick, Bernadette Bouchon-Meunier, Vladik Kreinovich
Int. J. Intell. Syst.2
1996 Deriving Queries from Results Using Genetic Programming
Tae-Wan Ryu, Christoph F. Eick
KDD2
1993 Learning Bayesian Classification Rules through Genetic Algorithms
abstract
The paper surveys the features of an inductive learning environment named DELVAUX that learns prospector.
Christoph F. Eick, Daw Jong
CIKM1
1993 Rule-Based Consistency Enforcement for Knowledge-Based Systems
abstract
A rule-based approach for the automatic enforcement of consistency constraints is presented. In contrast to existing approaches that compile consistency checks into application programs, the approach centralizes consistency enforcement in a separate module called a knowledge-base management system. Exception handlers for constraint violations are represented as rule entities in the knowledge base. For this purpose, a new form of production rule called the activation pattern controlled rule is introduced: in contrast to classical forward chaining schemes, activation pattern controlled rules are triggered by the intent to apply a specific operation but not necessarily by the result of applying this operation. Techniques for implementing this approach are discussed, and experiments in speeding up the system performance are described. Furthermore, an argument is made for more tolerant consistency enforcement strategies, and how they can be integrated into the rule-based approach to consistency enforcement is discussed.>
Christoph F. Eick, Paul Werstein
IEEE Trans. Knowl. Data Eng.1
1991 TANGUY: Integrating Database, Rule-based and Object-Oriented Paradigms
Bogdan D. Czejdo, Christoph F. Eick, Malcolm C. Taylor
DASFAA2
1991 A Methodology for the Design and Transformation of Conceptual Schemas
Christoph F. Eick
VLDB1
1991 Toward a Formal Semantics and Inference Rules for Conceptual Data Models
Christoph F. Eick, Thomas Raupp
Data Knowl. Eng.1
1985 Acquisition of Terminological Knowledge Using Database Design Techniques
abstract
Article Free Access Share on Acquisition of terminological knowledge using database design techniques Authors: Christoph F. Eick Fakultat fur Informatik, Universitat Karlsruhe, Postfach 6380, D-7500 Karlsruhe Fakultat fur Informatik, Universitat Karlsruhe, Postfach 6380, D-7500 KarlsruheView Profile , Peter C. Lockemann Fakultat fur Informatik, Universitat Karlsruhe, Postfach 6380, D-7500 Karlsruhe Fakultat fur Informatik, Universitat Karlsruhe, Postfach 6380, D-7500 KarlsruheView Profile Authors Info & Claims SIGMOD '85: Proceedings of the 1985 ACM SIGMOD international conference on Management of dataMay 1985Pages 84–94https://doi.org/10.1145/318898.318905Published:01 May 1985Publication History 20citation379DownloadsMetricsTotal Citations20Total Downloads379Last 12 Months14Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Christoph F. Eick, Peter C. Lockemann
SIGMOD Conference1
1984 From Natural Language Requirements to good Data Base Definitions - A Data Base Design Methodology
abstract
A comprehensive and constructive design methodology for Logical data base design is proposed. A set of compatible computerized design tools is described, carrying out the following tasks: formalization of natural language information requirements; refinement, evaluation and transformation of conceptual data definitions; mapping conceptual to DBTG data definitions.
Christoph F. Eick
ICDE1