EDBT 2026 Demo / reviewers in the wild / expert
Barry L. Drake
dblp:92/2307
· DBLP profile ↗
19ranked-venue papers
1as first author
4since 2021 · last 2023
0000-0003-4087-1524ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Theory of computation · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | WellFactor: Patient Profiling using Integrative Embedding of Healthcare DataabstractIn the rapidly evolving healthcare industry, platforms now have access to not only traditional medical records, but also diverse data sets encompassing various patient interactions, such as those from healthcare web portals. To address this rich diversity of data, we introduce WellFactor: a method that derives patient profiles by integrating information from these sources. Central to our approach is the utilization of constrained low-rank approximation. WellFactor is optimized to handle the sparsity that is often inherent in healthcare data. Moreover, by incorporating task-specific label information, our method refines the embedding results, offering a more informed perspective on patients. One important feature of WellFactor is its ability to compute embeddings for new, previously unobserved patient data instantaneously, eliminating the need to revisit the entire data set or recomputing the embedding. Comprehensive evaluations on real-world healthcare data demonstrate WellFactor’s effectiveness. It produces better results compared to other existing methods in classification performance, yields meaningful clustering of patients, and delivers consistent results in patient similarity searches and predictions. Dongjin Choi, Andy Xiang, Ozgur Ozturk, Deep Shrestha, Barry L. Drake, Hamid Haidarian, Faizan Javed, Haesun Park |
IEEE Big Data | 5 |
| 2023 | Patient Clustering via Integrated Profiling of Clinical and Digital DataabstractWe introduce a novel profile-based patient clustering model designed for healthcare clinical data. By utilizing a method grounded on constrained low-rank approximation, our model takes advantage of patients' clinical data and digital interaction data, including browsing and search, to construct patient profiles. As a result of the method, nonnegative embedding vectors are generated, serving as a low-dimensional representation of the patients. Our model was assessed using real-world patient data from a healthcare web portal, with a comprehensive evaluation approach which considered clustering and recommendation capabilities. In comparison to other baselines, our approach demonstrated superior performance in terms of clustering coherence and recommendation accuracy. Dongjin Choi, Andy Xiang, Ozgur Ozturk, Deep Shrestha, Barry L. Drake, Hamid Haidarian, Faizan Javed, Haesun Park |
CIKM | 5 |
| 2023 | Co-embedding Multi-type Data for Information Fusion and Visual AnalyticsabstractThis paper proposes a novel interactive system for exploratory document search in multi-type data sets that employs a data fusion approach. The system, designed to visualize different object types collectively and clearly display their semantic proximity, utilizes a co-embedding technique for knowledge fusion of multi-type data and projects the various types of objects onto a common lower-dimensional space. This produces a more informed representation and visualization that shows both in-type and across-type semantic proximity between objects. The system enables users’ exploration of multi-type document data by providing embedding-based scatter plots to visualize semantic relations within and across object types. In addition, based on user relevance feedback of displayed objects, the system offers recommendations of relevant objects. We have demonstrated the effectiveness of the proposed system and the underlying embedding method through comparison experiments with the proposed fusion-based approach. Dongjin Choi, Barry L. Drake, Haesun Park |
FUSION | 2 |
| 2021 | ArchiText: Interactive Hierarchical Topic ModelingabstractHuman-in-the-loop topic modeling allows users to explore and steer the process to produce better quality topics that align with their needs. When integrated into visual analytic systems, many existing automated topic modeling algorithms are given interactive parameters to allow users to tune or adjust them. However, this has limitations when the algorithms cannot be easily adapted to changes, and it is difficult to realize interactivity closely supported by underlying algorithms. Instead, we emphasize the concept of tight integration, which advocates for the need to co-develop interactive algorithms and interactive visual analytic systems in parallel to allow flexibility and scalability. In this article, we describe design goals for efficiently and effectively executing the concept of tight integration among computation, visualization, and interaction for hierarchical topic modeling of text data. We propose computational base operations for interactive tasks to achieve the design goals. To instantiate our concept, we present ArchiText, a prototype system for interactive hierarchical topic modeling, which offers fast, flexible, and algorithmically valid analysis via tight integration. Utilizing interactive hierarchical topic modeling, our technique lets users generate, explore, and flexibly steer hierarchical topics to discover more informed topics and their document memberships. Hannah Kim 0001, Barry L. Drake, Alex Endert, Haesun Park |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | MEGA: Multi-View Semi-Supervised Clustering of HypergraphsabstractComplex relationships among entities can be modeled very effectively using hypergraphs. Hypergraphs model real-world data by allowing a hyperedge to include two or more entities. Clustering of hypergraphs enables us to group the similar entities together. While most existing algorithms solely consider the connection structure of a hypergraph to solve the clustering problem, we can boost the clustering performance by considering various features associated with the entities as well as auxiliary relationships among the entities. Also, we can further improve the clustering performance if some of the labels are known and we incorporate them into a clustering model. In this paper, we propose a semi-supervised clustering framework for hypergraphs that is able to easily incorporate not only multiple relationships among the entities but also multiple attributes and content of the entities from diverse sources. Furthermore, by showing the close relationship between the hypergraph normalized cut and the weighted kernel K-Means, we also develop an efficient multilevel hypergraph clustering method which provides a good initialization with our semi-supervised multi-view clustering algorithm. Experimental results show that our algorithm is effective in detecting the ground-truth clusters and significantly outperforms other state-of-the-art methods. Joyce Jiyoung Whang, Rundong Du, Sangwon Jung, Barry L. Drake, Seonggoo Kang, Haesun Park |
Proc. VLDB Endow. | 5 |
| 2019 | Hybrid clustering based on content and connection structure using joint nonnegative matrix factorization
Rundong Du, Barry L. Drake, Haesun Park |
J. Glob. Optim. | 2 |
| 2018 | TopicOnTiles: Tile-Based Spatio-Temporal Event Analytics via Exclusive Topic Modeling on Social MediaabstractDetecting anomalous events of a particular area in a timely manner is an important task. Geo-tagged social media data are useful resource for this task; however, the abundance of everyday language in them makes this task still challenging. To address such challenges, we present TopicOnTiles, a visual analytics system that can reveal information relevant to anomalous events in a multi-level tile-based map interface by using social media data. To this end, we adopt and improve a recently proposed topic modeling method that can extract spatio-temporally exclusive topics corresponding to a particular region and a time point. Furthermore, we utilize a tile-based map interface to efficiently handle large-scale data in parallel. Our user interface effectively highlights anomalous tiles using our novel glyph visualization that encodes the degree of anomaly computed by our exclusive topic modeling processes. To show the effectiveness of our system, we present several usage scenarios using real-world datasets as well as comprehensive user study results. Minsuk Choi, Dear Sungbok Shin, Jinho Choi 0005, Scott Langevin, Christopher Bethune, Philippe Horne, Nathan Kronenfeld, Ramakrishnan Kannan, Barry L. Drake, Haesun Park, Jaegul Choo |
CHI | 9 |
| 2017 | STExNMF: Spatio-Temporally Exclusive Topic Discovery for Anomalous Event DetectionabstractUnderstanding newly emerging events or topics associated with a particular region of a given day can provide deep insight on the critical events occurring in highly evolving metropolitan cities. We propose herein a novel topic modeling approach on text documents with spatio-temporal information (e.g., when and where a document was published) such as location-based social media data to discover prevalent topics or newly emerging events with respect to an area and a time point. We consider a map view composed of regular grids or tiles with each showing topic keywords from documents of the corresponding region. To this end, we present a tilebased spatio-temporally exclusive topic modeling approach called STExNMF, based on a novel nonnegative matrix factorization (NMF) technique. STExNMF mainly works based on the two following stages: (1) first running a standard NMF of each tile to obtain general topics of the tile and (2) running a spatiotemporally exclusive NMF on a weighted residual matrix. These topics likely reveal information on newly emerging events or topics of interest within a region. We demonstrate the advantages of our approach using the geo-tagged Twitter data of New York City. We also provide quantitative comparisons in terms of the topic quality, spatio-temporal exclusiveness, topic variation, and qualitative evaluations of our method using several usage scenarios. In addition, we present a fast topic modeling technique of our model by leveraging parallel computing. Dear Sungbok Shin, Minsuk Choi, Jinho Choi 0005, Scott Langevin, Christopher Bethune, Philippe Horne, Nathan Kronenfeld, Ramakrishnan Kannan, Barry L. Drake, Haesun Park, Jaegul Choo |
ICDM | 9 |
| 2017 | DC-NMF: nonnegative matrix factorization based on divide-and-conquer for fast clustering and topic modeling
Rundong Du, Da Kuang, Barry L. Drake, Haesun Park |
J. Glob. Optim. | 3 |
| 2014 | Nonlinear Adaptive Filtering with Dimension Reduction in the Wavelet DomainabstractRecent advances in adaptive filter theory and the hardware for signal acquisition have led to the realization that purely linear algorithms are often not adequate in these domains. Nonlinearities in the input space have become apparent with today's real world problems. Algorithms that process the data must keep pace with the advances in signal acquisition. Recently kernel adaptive (online) filtering algorithms have been proposed that make no assumptions regarding the linearity of the input space. Additionally, advances in wavelet data compression/dimension reduction have also led to new algorithms that are appropriate for producing a hybrid nonlinear filtering framework. In this paper we utilize a combination of wavelet dimension reduction and kernel adaptive filtering. We derive algorithms in which the dimension of the data is reduced by a wavelet transform. We follow this by kernel adaptive filtering algorithms on the reduced-domain data to find the appropriate model parameters demonstrating improved minimization of the mean-squared error (MSE). Another important feature of our methods is that the wavelet filter is also chosen based on the data, on-the-fly. In particular, it is shown that by using a few optimal wavelet coefficients from the constructed wavelet filter for both training and testing data sets as the input to the kernel adaptive filter, convergence to the near optimal learning curve (MSE) results. We demonstrate these algorithms on simulated and a real data set from food processing. Tiffany Huang, Barry L. Drake, David Aalfs, Brani Vidakovic |
DCC | 2 |
| 2010 | Supervised Raman spectra estimation based on nonnegative rank deficient least squares
Barry L. Drake, Jingu Kim, Mahendra Mallick, Haesun Park |
FUSION | 1 |
| 2009 | Comparison of Raman spectra estimation algorithms
Mahendra Mallick, Barry L. Drake, Haesun Park, Andy Register, William Dale Blair, Phil West, Ryan D. Palkki, Aaron D. Lanterman, Darren Emge |
FUSION | 2 |
| 2009 | Hierarchical Linear Discriminant Analysis for BeamformingabstractThis paper demonstrates the applicability of the recently proposed supervised dimension reduction, hierarchical linear discriminant analysis (h-LDA) to a well-known spatial localization technique in signal processing, beamforming. The main motivation of h-LDA is to overcome the drawback of LDA that each cluster is modeled as a unimodal Gaussian distribution. For this purpose, h-LDA extends the variance decomposition in LDA to the subcluster level, and modifies the definition of the within-cluster scatter matrix. In this paper, we present an efficient h-LDA algorithm for over-sampled data, where the data dimension is larger than the dimension of the data vectors. The new algorithm utilizes the Cholesky decomposition based on a generalized singular value decomposition framework. Furthermore, we analyze the data model of h-LDA by relating it to the two-way multivariate analysis of variance (MANOVA), which fits well within the context of beamforming applications. Although beamforming has been generally dealt with as a regression problem, we propose a novel way of viewing beamforming as a classification problem, and apply a supervised dimension reduction, which allows the classifier to achieve better accuracy. Our experimental results show that h-LDA outperforms several dimension reduction methods such as LDA and kernel discriminant analysis, and regression approaches such as the regularized least squares and kernelized support vector regression. Jaegul Choo, Barry L. Drake, Haesun Park |
SDM | 2 |
| 2008 | Linear discriminant analysis for data with subcluster structureabstractLinear discriminant analysis (LDA) is a widely-used feature extraction method in classification. However, the original LDA has limitations due to the assumption of a unimodal structure for each cluster, which is satisfied in many applications such as facial image data when variations such as angle and illumination can significantly influence the images of the same person. In this paper, we propose a novel method, hierarchical LDA(h-LDA), which takes into account hierarchical subcluster structures in the data sets. Our experiments show that regularized h-LDA produces better accuracy than LDA, PCA, and tensorFaces. Haesun Park, Jaegul Choo, Barry L. Drake, Jinwoo Kang |
ICPR | 3 |
| 2008 | Extracting unrecognized gene relationships from the biomedical literature via matrix factorizations
Haesun Park, Barry L. Drake |
BMC Bioinform. | 3 |
| 2007 | Extracting unrecognized gene relationships from the biomedical literature via matrix factorizationsabstractBACKGROUND: The construction of literature-based networks of gene-gene interactions is one of the most important applications of text mining in bioinformatics. Extracting potential gene relationships from the biomedical literature may be helpful in building biological hypotheses that can be explored further experimentally. Recently, latent semantic indexing based on the singular value decomposition (LSI/SVD) has been applied to gene retrieval. However, the determination of the number of factors k used in the reduced rank matrix is still an open problem. RESULTS: In this paper, we introduce a way to incorporate a priori knowledge of gene relationships into LSI/SVD to determine the number of factors. We also explore the utility of the non-negative matrix factorization (NMF) to extract unrecognized gene relationships from the biomedical literature by taking advantage of known gene relationships. A gene retrieval method based on NMF (GR/NMF) showed comparable performance with LSI/SVD. CONCLUSION: Using known gene relationships of a given gene, we can determine the number of factors used in the reduced rank matrix and retrieve unrecognized genes related with the given gene by LSI/SVD or GR/NMF. Haesun Park, Barry L. Drake |
BMC Bioinform. | 3 |
| 2007 | Multiclass classifiers based on dimension reduction with generalized LDA
Barry L. Drake, Haesun Park |
Pattern Recognit. | 2 |
| 2006 | Adaptive Nonlinear Discriminant Analysis by Regularized Minimum Squared ErrorsabstractKernelized nonlinear extensions of Fisher's discriminant analysis, discriminant analysis based on generalized singular value decomposition (LDA/GSVD), and discriminant analysis based on the minimum squared error formulation (MSE) have recently been widely utilized for handling undersampled high-dimensional problems and nonlinearly separable data sets. As the data sets are modified from incorporating new data points and deleting obsolete data points, there is a need to develop efficient updating and downdating algorithms for these methods to avoid expensive recomputation of the solution from scratch. In this paper, an efficient algorithm for adaptive linear and nonlinear kernel discriminant analysis based on regularized MSE, called adaptive KDA/RMSE, is proposed. In adaptive KDA/RMSE, updating and downdating of the computationally expensive eigenvalue decomposition (EVD) or singular value decomposition (SVD) is approximated by updating and downdating of the QR decomposition achieving an order of magnitude speed up. This fast algorithm for adaptive kernelized discriminant analysis is designed by utilizing regularization techniques and the relationship between linear and nonlinear discriminant analysis and the MSE. In addition, an efficient algorithm to compute leave-one-out cross validation is also introduced by utilizing downdating of KDA/RMSE. Barry L. Drake, Haesun Park |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2003 | A Decision Criterion for the Optimal Number of Clusters in Hierarchical Clustering
Yunjae Jung, Haesun Park, Ding-Zhu Du, Barry L. Drake |
J. Glob. Optim. | 4 |