Dongjin Choi

dblp:33/7730 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
4since 2021 · last 2023
0000-0002-7311-0644ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2023 WellFactor: Patient Profiling using Integrative Embedding of Healthcare Data
abstract
In the rapidly evolving healthcare industry, platforms now have access to not only traditional medical records, but also diverse data sets encompassing various patient interactions, such as those from healthcare web portals. To address this rich diversity of data, we introduce WellFactor: a method that derives patient profiles by integrating information from these sources. Central to our approach is the utilization of constrained low-rank approximation. WellFactor is optimized to handle the sparsity that is often inherent in healthcare data. Moreover, by incorporating task-specific label information, our method refines the embedding results, offering a more informed perspective on patients. One important feature of WellFactor is its ability to compute embeddings for new, previously unobserved patient data instantaneously, eliminating the need to revisit the entire data set or recomputing the embedding. Comprehensive evaluations on real-world healthcare data demonstrate WellFactor’s effectiveness. It produces better results compared to other existing methods in classification performance, yields meaningful clustering of patients, and delivers consistent results in patient similarity searches and predictions.
Dongjin Choi, Andy Xiang, Ozgur Ozturk, Deep Shrestha, Barry L. Drake, Hamid Haidarian, Faizan Javed, Haesun Park
IEEE Big Data1
2023 Patient Clustering via Integrated Profiling of Clinical and Digital Data
abstract
We introduce a novel profile-based patient clustering model designed for healthcare clinical data. By utilizing a method grounded on constrained low-rank approximation, our model takes advantage of patients' clinical data and digital interaction data, including browsing and search, to construct patient profiles. As a result of the method, nonnegative embedding vectors are generated, serving as a low-dimensional representation of the patients. Our model was assessed using real-world patient data from a healthcare web portal, with a comprehensive evaluation approach which considered clustering and recommendation capabilities. In comparison to other baselines, our approach demonstrated superior performance in terms of clustering coherence and recommendation accuracy.
Dongjin Choi, Andy Xiang, Ozgur Ozturk, Deep Shrestha, Barry L. Drake, Hamid Haidarian, Faizan Javed, Haesun Park
CIKM1
2023 Co-embedding Multi-type Data for Information Fusion and Visual Analytics
abstract
This paper proposes a novel interactive system for exploratory document search in multi-type data sets that employs a data fusion approach. The system, designed to visualize different object types collectively and clearly display their semantic proximity, utilizes a co-embedding technique for knowledge fusion of multi-type data and projects the various types of objects onto a common lower-dimensional space. This produces a more informed representation and visualization that shows both in-type and across-type semantic proximity between objects. The system enables users’ exploration of multi-type document data by providing embedding-based scatter plots to visualize semantic relations within and across object types. In addition, based on user relevance feedback of displayed objects, the system offers recommendations of relevant objects. We have demonstrated the effectiveness of the proposed system and the underlying embedding method through comparison experiments with the proposed fusion-based approach.
Dongjin Choi, Barry L. Drake, Haesun Park
FUSION1
2023 Real-Time External Compensation System With Error Correction Algorithm for High-Resolution Mobile Displays
abstract
This paper presents an external compensation system for QHD+ ($3040\times1224$) mobile active-matrix organic light emitting diode (AMOLED) displays at a frame rate of 60 Hz. During vertical blank periods, current sensing AFE (CS-AFE) measures OLED currents to calculate threshold voltage ($V_{TH}$) of driving thin-film transistors (TFTs). For precise$V_{TH}$calculation against panel ground noise, a differential sensing scheme with 5-bit programmable capacitor array (PCA) is employed. In addition, digital correlated double sampling (CDS) removes an offset of the CS-AFE. However, recent advances in high efficiency OLED technology have led to increase in pixel density as well as the driving TFTs to operate close to subthreshold region. Therefore, the$V_{TH}$calculation based on the quadratic model yields inaccurate results. To compensate for the modeling error, we propose an error correction algorithm, which establishes an error function using a relationship between the modeling error and calculated threshold voltage during the manufacturing process. The proposed external compensation system was verified using CMOS-modeled three transistors and one capacitor (3T1C) pixel circuit. The test chip, fabricated in a$0.18~\mu \text{m}$BCD process, comprises 26 channels. Each channel consumes 78$\mu \text{W}$and occupies 1350$\times 50\,\,\mu \text{m}^{2}$. Measurement results show that current error at$64^{\mathrm {th}}$gray level is reduced from 35.56 LSB to 6.03 LSB after error correction and four frames average.
Seunghun Oh, Dongjin Choi, Kyeonghan Shin, Haewan Cho, Franklin Bien
IEEE Trans. Circuits Syst. I Regul. Pap.3
2020 SNeCT: Scalable Network Constrained Tucker Decomposition for Multi-Platform Data Profiling
abstract
How do we integratively profile large-scale multi-platform genomic data that are high dimensional and sparse? Furthermore, how can we incorporate prior knowledge, such as the association between genes, in the analysis systematically to find better latent relationships? To solve this problem, we propose a Scalable Network Constrained Tucker decomposition method (SNeCT). SNeCT adopts parallel stochastic gradient descent approach on the proposed parallelizable network constrained optimization function. SNeCT decomposition is applied to a tensor constructed from a large scale multi-platform multi-cohort cancer data, PanCan12, constrained on a network built from PathwayCommons database. The decomposed factor matrices are applied to stratify cancers, to search for top- k similar patients given a new patient, and to illustrate how the matrices can be used to identify significant genomic patterns in each patient. In the stratification test, combined twelve-cohort data is clustered to form thirteen subclasses. The similarity of the top- k patient to the query was high for 23 clinical features, including estrogen/progesterone receptor statuses of BRCA patients with average precision value ranges from 0.72 to 0.86 and from 0.68 to 0.86, respectively. We also illustrate how the factor matrices can be used for identifying significant patterns for each patient. Resources are available at: https://github.com/leesael/SNeCT.
Dongjin Choi, Lee Sael
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 Zoom-SVD: Fast and Memory Efficient Method for Extracting Key Patterns in an Arbitrary Time Range
abstract
Given multiple time series data, how can we efficiently find latent patterns in an arbitrary time range? Singular value decomposition (SVD) is a crucial tool to discover hidden factors in multiple time series data, and has been used in many data mining applications including dimensionality reduction, principal component analysis, recommender systems, etc. Along with its static version, incremental SVD has been used to deal with multiple semi-infinite time series data and to identify patterns of the data. However, existing SVD methods for the multiple time series data analysis do not provide functionality for detecting patterns of data in an arbitrary time range: standard SVD requires data for all intervals corresponding to a time range query, and incremental SVD does not consider an arbitrary time range.
Jun-Gi Jang, Dongjin Choi, Jinhong Jung, U Kang
CIKM2
2015 Identifying the most appropriate expansion of acronyms used in wikipedia text
abstract
Summary During the past several years, many studies have been conducted to determine the expansions (the expanded forms or definitions) of acronyms used in specific texts by using linguistic and syntactical approaches to overcome disambiguation problems. An acronym is an abbreviation formed of the initial components of multiple words. Using only these initial parts can result in huge mistakes when a machine conducts experiments to determine an acronym's meaning in given texts. Finding expansions of acronyms is currently not a major concern. However, identifying how a polysemous acronym is precisely being used in a specific text is difficult. To solve disambiguation problem acronyms have, this paper proposes a method to identify the most appropriate expansions of an acronym by analysing co‐occurrence words found in articles on Wikipedia. Our goal is not to find an acronym's definition or expansion but to identify the most appropriate expansion of a given acronym as it is used in a given text. Through experiments, we demonstrate the proposed method's efficiency and performance by comparing several results. Copyright © 2014 John Wiley & Sons, Ltd.
Dongjin Choi, Pankoo Kim
Softw. Pract. Exp.1
2014 Text analysis for detecting terrorism-related articles on the web
Dongjin Choi, Byeong-Kyu Ko, Heesun Kim, Pankoo Kim
J. Netw. Comput. Appl.1
2013 Sentiment Analysis for Tracking Breaking Events: A Case Study on Twitter
Dongjin Choi, Pankoo Kim
ACIIDS (2)1
2011 Domain N-Gram Construction and Its Application to Text Editor
Myunggwon Hwang, Dongjin Choi, Hyogap Lee, Pankoo Kim
ACIIDS (1)2
2010 Least Slack Time Rate First: New Scheduling Algorithm for Multi-Processor Environment
abstract
Real-time systems have to complete the execution of a task within the predetermined time while ensuring that the execution results are logically correct. Such systems require scheduling methods that can adequately distribute the given tasks to a processor. Scheduling methods that all tasks can be executed within a predetermined deadline are called an optimal scheduling. In this paper, we propose a new and simple scheduling algorithm (LSTR: least slack time rate first) as a dynamic-priority algorithm for a multi-processor environment and demonstrate its optimal possibility through various tests.
Myunggwon Hwang, Dongjin Choi, Pankoo Kim
CISIS2