VLDB 2026 Research / reviewers in the wild / expert
Mehmet M. Dalkilic
dblp:16/4499 · also Mehmet M. Dalkiliç
· DBLP profile ↗
22ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 14 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Theory of computation · 5 · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geometric-k-means: a bound free approach to fast and eco-friendly k-meansabstractAbstract This paper introduces Geometric- k -means (or $${\mathsf{G}}k$$ -means for short), a novel approach that significantly enhances the efficiency and energy economy of the widely utilized k -means algorithm, which, despite its inception over five decades ago, remains a cornerstone in machine learning applications. The essence of $${\mathsf{G}}k$$ -means lies in its active utilization of geometric principles, specifically scalar projection, to significantly accelerate the algorithm without sacrificing solution quality. This geometric strategy enables a more discerning focus on data points that are most likely to influence cluster updates, which we call as high expressive data (HE). In contrast, low expressive data (LE), does not impact clustering outcome, is effectively bypassed, leading to considerable reductions in computational overhead. Experiments spanning synthetic, real-world and high-dimensional datasets, demonstrate $${\mathsf{G}}k$$ -means is significantly better than traditional and state of the art (SOTA) $$k$$ -means variants in runtime and distance computations (DC). Moreover, $${\mathsf{G}}k$$ -means exhibits better resource efficiency, as evidenced by its reduced energy footprint, placing it as more sustainable alternative. The software code and data for our algorithm is available at https://github.com/parichit/Geometric-k-means . Parichit Sharma, Marcin Malec, Hasan Kurban, M. Oguzhan Külekci, Mehmet M. Dalkilic |
Mach. Learn. | 5 |
| 2024 | A Novel Discrete Time Series Representation with De Bruijn Graphs for Enhanced Forecasting Using TimesNet (Extended Abstract)abstractThis paper introduces a novel method for time series forecasting using de Bruijn Graphs (dBGs) to represent discretized time series data. Our approach involves (1) encoding time series as a dBG, (2) applying both novel and existing graph encoding algorithms (like struct2vec) to extract features from dBG, and (3) integrating these features into the TimesNet model to enhance short-term univariate forecasting accuracy. Empirical results on the M4 datasets show that our method preserves the dynamics of the time series while improving forecasting performance across various datasets. Mert Onur Cakiroglu, Hasan Kurban, Elham Khorasani Buxton, Mehmet M. Dalkilic |
DSAA | 4 |
| 2024 | $p$-ClustVal: A Novel $p$-Adic Approach for Enhanced Clustering of High-Dimensional scRNASeq Data (Extended Abstract)abstractThis paper introduces$p$-ClustVal, a novel data transformation technique inspired by p-adic number theory that significantly enhances cluster discernibility in genomics data, specifically Single Cell RNA Sequencing (scRNASeq). By lever-aging p-adic-valuation,$p$-ClustVal integrates with and augments widely used clustering algorithms and dimension reduction techniques, amplifying their effectiveness in discovering meaningful structure from data. The transformation uses a data-centric heuristic to determine optimal parameters, without relying on ground truth labels, making it more user-friendly.$p$-ClustVal reduces overlap between clusters by employing alternate metric spaces inspired by p-adic-valuation, a significant shift from conventional methods. Our comprehensive evaluation spanning 30 experiments and over 1200 observations, shows that$p$-ClustVal improves performance in 91% of cases, and boosts the performance of classical and state of the art (SOTA) methods. This work contributes to data analytics and genomics by introducing a unique data transformation approach, enhancing downstream clustering algorithms, and providing empirical evidence of$p$-ClustVal's efficacy. Parichit Sharma, Sarthak Mishra, Hasan Kurban, Mehmet M. Dalkilic |
DSAA | 4 |
| 2023 | Novel NBA Fantasy League driven by Engineered Team Chemistry and Scaled Position StatisticsabstractFantasy Sports has a current market size of ${\$}$27B and is expected to grow more than ${\$}$84B in less than a decade. The intent is to create virtual teams that somehow reflect what would happen if the constituent players actually played in a team. Using individual player and team statistics, models can be trained to predict an outcome. But fans are left wanting more. To achieve a more realistic outcome, aspects of what makes live teams win need to be included: (1) transforming player statistics to reflect their relative importance with respect to a player position; (2) team chemistry (TC). In this work, we show a novel characterization of relative position statistics and a new description of TC. Drawn from the NBA’s API, we form a data set to determine whether a fantasy team makes the playoffs using almost two dozen features, including TC. Various Machine Learning models are trained on this data and the best-performing model is offered to the users through a web service. Users can not only inspect fantasy teams and their TC but can also simulate their match-ups with existing 2023 NBA teams and utilize performance visualizations to help improve their team creation process. Our web service can be accessed at https://dalkilic.luddy.indiana.edu/fantasyleague/, and the source code can be found at https://github.com/gany-15/nbafan. Ganesh Arkanath, Nishad Gupta, Hasan Kurban, Parichit Sharma, K. R. Madhavan, Elham Khorasani Buxton, Mehmet M. Dalkilic |
IEEE Big Data | 7 |
| 2022 | ccImpute: an accurate and scalable consensus clustering based algorithm to impute dropout events in the single-cell RNA-seq dataabstractBACKGROUND: In recent years, the introduction of single-cell RNA sequencing (scRNA-seq) has enabled the analysis of a cell's transcriptome at an unprecedented granularity and processing speed. The experimental outcome of applying this technology is a [Formula: see text] matrix containing aggregated mRNA expression counts of M genes and N cell samples. From this matrix, scientists can study how cell protein synthesis changes in response to various factors, for example, disease versus non-disease states in response to a treatment protocol. This technology's critical challenge is detecting and accurately recording lowly expressed genes. As a result, low expression levels tend to be missed and recorded as zero - an event known as dropout. This makes the lowly expressed genes indistinguishable from true zero expression and different than the low expression present in cells of the same type. This issue makes any subsequent downstream analysis difficult. RESULTS: To address this problem, we propose an approach to measure cell similarity using consensus clustering and demonstrate an effective and efficient algorithm that takes advantage of this new similarity measure to impute the most probable dropout events in the scRNA-seq datasets. We demonstrate that our approach exceeds the performance of existing imputation approaches while introducing the least amount of new noise as measured by clustering performance characteristics on datasets with known cell identities. CONCLUSIONS: ccImpute is an effective algorithm to correct for dropout events and thus improve downstream analysis of scRNA-seq data. ccImpute is implemented in R and is available at https://github.com/khazum/ccImpute . Marcin Malec, Hasan Kurban, Mehmet M. Dalkilic |
BMC Bioinform. | 3 |
| 2020 | Applying Class-to-Class Siamese Networks to Explain Classifications with Supportive and Contrastive Cases
Xiaomeng Ye, David B. Leake, William Huibregtse, Mehmet M. Dalkilic |
ICCBR | 4 |
| 2018 | Using Data Analytics to Optimize Public Transportation on a College CampusabstractUsing a large volume of bus data in the form of GPS coordinates (over 100 million data points) and automated passenger count data (over 1 million data points) we have developed (1) a system of analysis and prediction of future public transportation demand (2) a new model that uses concepts specific to college campuses that maximizes passenger satisfaction. Using these concepts we improve service of a model college public transportation service and more specifically the Indiana University Campus Bus Service (IUCBS). Kurt Zimmer, Hasan Kurban, Mark Jenne, Logan Keating, Perry Maull, Mehmet M. Dalkilic |
DSAA | 6 |
| 2017 | Case Study: Clustering Big Stellar Data with EMabstractWithout question, astronomy is about Big Data and clustering is a very common task over astronomy domain. The expectation-maximization algorithm is among the top 10 data mining algorithms used in scientific and industrial applications, however, we observe that astronomical community does not make use of it as a clustering algorithm. In this work, we cluster $\sim$ 1M stellar objects (simulated Galactic spectral data) via the traditional expectation-maximization algorithm for clustering (EM-T) and our extended EM-T algorithm that we call EM* and present the experimental results. Hasan Kurban, Can Kockan, Mark Jenne, Mehmet M. Dalkilic |
BDCAT | 4 |
| 2017 | A novel approach to optimization of iterative machine learning algorithms: Over heap structureabstractIterative machine learning algorithms, i.e., k-means (KM), expectation maximization (EM), become overwhelmed with big data since all data points are being continually and indiscriminately visited while a cost is being minimized. In this work, we demonstrate (1) an optimization approach to reduce training run-time complexity of iterative machine learning algorithms and (2) implementation of this framework over KM algorithm. We call this extended KM algorithm, KM*. The experimental results show that KM* outperforms KM over big real world and synthetic data sets. Lastly, we demonstrate the theoretical elements of our work. Hasan Kurban, Mehmet M. Dalkilic |
IEEE BigData | 2 |
| 2017 | Improving expectation maximization algorithm over stellar dataabstractStellar data, only a few years ago, measured in the .1M of objects. Now, sets are routinely 1M. With the launch of ESA's Gaia in 2013, we expect 1000M stellar objects measured more precisely and with more measurements. Without question, astronomy is about Big Data and clustering is a very common task over astronomy domain. The expectation-maximization algorithm is among the top 10 data mining algorithms used in scientific and industrial applications, however, we observe that astronomical community does not make use of it as a clustering algorithm. In this work, we cluster ~ 1M stellar objects (simulated Galactic spectral data) via the traditional expectation-maximization algorithm for clustering (EM-T) and our extended EM-T algorithm that we call EM* and present the experimental results. Hasan Kurban, Can Kockan, Mark Jenne, Mehmet M. Dalkilic |
IEEE BigData | 4 |
| 2016 | EM*: An EM Algorithm for Big DataabstractExisting data mining techniques, more particularly iterative learning algorithms, become overwhelmed with big data. While parallelism is an obvious and, usually, necessary strategy, we observe that both (1) continually revisiting data and (2) visiting all data are two of the most prominent problems especially for iterative, unsupervised algorithms like Expectation Maximization algorithm for clustering (EM-T). Our strategy is to embed EM-T into a non-linear hierarchical data structure(heap) that allows us to (1) separate data that needs to be revisited from data that does not and (2) narrow the iteration toward the data that is more difficult to cluster. We call this extended EM-T, EM*. We show our EM* algorithm outperform EM-T algorithm over large real world and synthetic data sets. We lastly conclude with some theoretic underpinnings that explain why EM* is successful. Hasan Kurban, Mark Jenne, Mehmet M. Dalkilic |
DSAA | 3 |
| 2014 | A new set of Random Forests with varying dynamic data reduction and voting techniquesabstractRandom forests have been used as effective models to tackle a number of classification and regression problems. In this paper, we present a new type of Random Forests (RFs) called Red(uced)-RF that adopts a new voting mechanism called Priority Vote Weighting (PV) and a new dynamic data reduction principle which improve accuracy and execution time compared to Breiman's conventional RF. Red-RF also shows that the strength of a random forest can increase without noticeably increasing correlation between the trees. We then compare performance of Red-RF, 9 new RF variants and Breiman's RF in eight experiments that involve classification problems with datasets of different sizes. Hussein Mohsen, Hasan Kurban, Mark Jenne, Mehmet M. Dalkilic |
DSAA | 4 |
| 2012 | WIGM: Discovery of Subgraph Patterns in a Large Weighted GraphabstractMany research areas have begun representing massive data sets as very large graphs. Thus, graph mining has been an active research area in recent years. Most of the graph mining research focuses on mining unweighted graphs. However, weighted graphs are actually more common. The weight on an edge may represent the likelihood or logarithmic transformation of likelihood of the existence of the edge or the strength of an edge, which is common in many biological networks. In this paper, a weighted subgraph pattern model is proposed to capture the importance of a subgraph pattern and our aim is to find these patterns in a large weighted graph. Two related problems are studied in this paper: (1) discovering all patterns with respect to a given minimum weight threshold and (2) finding k patterns with the highest weights. The weighted subgraph patterns do not possess the anti-monotonic property and in turn, most of existing subgraph mining methods could not be directly applied. Fortunately, the 1-extension property is identified so that a bounded search can be achieved. A novel weighted graph mining algorithm, namely WIGM, is devised based on the 1-extension property. Last but not least, real and synthetic data sets are used to show the effectiveness and efficiency of our proposed models and algorithms. Shirong Li, Mehmet M. Dalkilic |
SDM | 4 |
| 2007 | Using Drosophila melanogaster Data to Discover Disease-Related Protein Interactions in HumansabstractThe discovery and understanding of functional relationships amongst genes and their gene products are fundamental to our understanding of human disease. Data regarding protein-protein interactions in humans are not complete and computational techniques along with data provided from other organisms can be used to better inform relationships in humans. We demonstrate the utility of using genome scale data from Drosophila melanogaster to predict protein-protein relationships for human proteins specific to disease. To find the most likely candidates for protein-protein interaction, the predicted relationships are tested against human protein interaction and known disease data, then ranked through a support vector machine. An illustrative example related to hBrm and hSNF5 shows the validity of our approach James C. Costello, Jade E. Buchanan-Carter, Mehmet M. Dalkilic, Justen Andrews |
CIBCB | 3 |
| 2007 | A Measurement Ontology Generalizable for Emerging Domain Applications on the Semantic WebabstractThis article introduces a measurement ontology for applications to Semantic Web applications, specifically for emerging domains such as microarray analysis. The Semantic Web is the next generation Web of structured data that are automatically shared by software agents, which apply definitions and constraints organized in ontologies to correctly process data from disparate sources. One facet needed to develop Semantic Web ontologies of emerging domains is creating ontologies of concepts that are common to these domains. These general “common-sense” ontologies can be used as building blocks to develop more domain-specific ontologies. However most measurement ontologies concentrate on representing units of measurement and quantities, and not on other measurement concepts such as sampling, mean values, and evaluations of quality based on measurements. In this article, we elaborate on a measurement ontology that represents all these concepts. We present the generality of the ontology, and describe how it is developed, used for analysis and validated. Henry M. Kim, Arijit Sengupta, Mark S. Fox, Mehmet M. Dalkilic |
J. Database Manag. | 4 |
| 2006 | Trust Establishment in Data Sharing: An Incentive Model for Biodiversity Information SystemsabstractWe describe a long-felt but largely neglected problem in conservation biology, and explain how it can be addressed using incentive mechanisms inspired by techniques in computer security and cryptography. The result is a new type of database suitable for highly distributed contributions of data, in which researchers are incentivised to submit data by the guarantees extended by a conflict resolution mechanism that allows for accurate determinations of data origination. Sukamol Srikwan, Markus Jakobsson, Andrew Albrecht, Mehmet M. Dalkilic |
CollaborateCom | 4 |
| 2006 | Using Compression to Identify Classes of Inauthentic TextsabstractRecent events have made it clear that some kinds of technical texts, generated by machine and essentially meaningless, can be confused with authentic, technical texts written by humans. We identify this as a potential problem, since no existing systems for, say the web, can or do discriminate on this basis. We believe that there are subtle, short- and long-range word or even string repetitions extant in human texts, but not in many classes of computer generated texts, that can be used to discriminate based on meaning. In this paper we employ universal lossless source coding to generate features in a high-dimensional space and then apply support vector machines to discriminate between the classes of authentic and inauthentic expository texts. Compression profiles for the two kinds of text are distinct—the authentic texts being bounded by various classes of more compressible or less compressible texts that are computer generated. This in turn led to the high prediction accuracy of our models which support a conjecture that there exists a relationship between meaning and compressibility. Our results show that the learning algorithm based upon the compression profile outperformed standard term-frequency text categorization on several non-trivial classes of inauthentic texts. Availability: http://www.informatics.indiana.edu/predrag/fsi.htm. Mehmet M. Dalkilic, Wyatt Travis Clark, James C. Costello, Predrag Radivojac |
SDM | 1 |
| 2006 | COMPAM : visualization of combining pairwise alignments for multiple genomesabstractUNLABELLED: COMPAM is a tool for visualizing relationships among multiple whole genomes by combining all pairwise genome alignments. It displays shared conserved regions (blocks) and where these blocks occur (edges) as block relation graphs which can be explored interactively. An unannotated genome, e.g. can then be explored using information from well-annotated genomes, COG-based genome annotation and genes. COMPAM can run either as a stand-alone application or through an applet that is provided as service to PLATCOM, a toolset for whole genome comparative analysis, where a wide variety of genomes can be easily selected. Features provided by COMPAM include the ability to export genome relationship information into file formats that can be used by other existing tools. AVAILABILITY: http://bio.informatics.indiana.edu/projects/compam/ Do-Hoon Lee, Jeong-Hyeon Choi, Mehmet M. Dalkilic, Sun Kim |
Bioinform. | 3 |
| 2006 | High-Performance Direct Pairwise Comparison of Large Genomic SequencesabstractMany applications in comparative genomics lend themselves to implementations that take advantage of common high-performance features in modern microprocessors. However, the common suggestion that a data-parallel, multithreaded, or high-throughput implementation is possible often ignores the complexity of actually creating such software. In this paper, we present two parallel algorithms for a classic comparative genomics algorithm, the dot plot. First, we describe a data-parallel algorithm that achieves speedups of up to 14.4x over the sequential version for large genomic comparisons. Then, we use the new algorithm as the base for a coarse-grained parallel version, suitable for multiprocessor and cluster environments, that scales linearly with the number of processors. These speedups introduce the opportunity to perform full pairwise comparisons on entire genomes on a much larger scale than previously possible. We also present the experimental, model-driven approach used to develop the algorithm that allowed us to carefully study and evaluate implementation options and to fully understand the parameters affecting its performance Christopher Mueller, Mehmet M. Dalkilic, Andrew Lumsdaine |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2004 | A logic-theoretic classifier called CircleabstractWe present a novel classifier based upon principles of logic-theoretic Boolean function minimization. The classifier, called Circle, recursively produces a set of implicants (or rules). The implicant set contains information not only about the presence of features, but also about their absence in determining class values. Thus, Circle's implicant set is initially non-monotonic with respect to inserting new tuples that have feature values that were not in the training set. One important benefit of this non-monotonicity, however, is that Circle is capable of being robust in the presence of novel feature values. We have created a full implementation of Circle using Java as a host language and Oracle database backend. Because we are interested in data mining in bioinformatics, particularly genomic data, the database was borne out of necessity to both manage and effectively query the information. Mehmet M. Dalkilic, Arijit Sengupta |
ICARCV | 1 |
| 2002 | DSQL - An SQL for Structured Documents
Arijit Sengupta, Mehmet M. Dalkilic |
CAiSE | 2 |
| 2000 | Information DependenciesabstractThis paper uses the tools of information theory to examine and reason about the information content of the attributes within a relation instance. For two sets of attributes X and Y, an information dependency measure (InD measure) characterizes the uncertainty remaining about the values for the set Y when the values for the set X are known. A variety of arithmetic inequalities (InD inequalities) are shown to hold among InD measures; InD inequalities hold in any relation instance. Numeric constraints (InD constraints) on InD measures, consistent with the InD inequalities, can be applied to relation instances. Remarkably, functional and multivalued dependencies correspond to setting certain constraints to zero, with Armstrong's axioms shown to be consequences of the arithmetic inequalities applied to constraints. As an analog of completeness, for any set of constraints consistent with the inequalities, we may construct a relation instance that approximates these constraints within any positive ε. InD measures suggest many valuable applications in areas such as data mining. Mehmet M. Dalkilic, Edward L. Robertson |
PODS | 1 |