Sangyeon Lee

dblp:134/4023 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0002-3260-4285ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 59% Storage systems · 35% Distributed systems · 6%
Databases, data mining, and information retrieval
3 papers
Graph data management · 77% Data mining · 23%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › drug discovery
drug side effect prediction
0.712023
Large-scale prediction of adverse drug reactions-related proteins with network embedding · Bioinform. 2023
Parallel and multicore computing
parallel graph algorithms
0.632016
DSP-CC: I/O efficient parallel computation of connected components in billion-scale networks · ICDE 2016
DSP-CC-: I/O Efficient Parallel Computation of Connected Components in Billion-Scale Networks · IEEE Trans. Knowl. Data Eng. 2015
TurboGraph: a fast parallel graph engine handling billion-scale graphs in a single PC · KDD 2013
Storage systems › out-of-core computation
out-of-core graph processing
0.212016
DSP-CC: I/O efficient parallel computation of connected components in billion-scale networks · ICDE 2016
Graph algorithms and graph theory › graph connectivity
connected components
0.212016
DSP-CC: I/O efficient parallel computation of connected components in billion-scale networks · ICDE 2016
Graph data management › graph algorithms
connected components
0.212015
DSP-CC-: I/O Efficient Parallel Computation of Connected Components in Billion-Scale Networks · IEEE Trans. Knowl. Data Eng. 2015
Data mining › structured data mining
graph mining
0.212015
DSP-CC-: I/O Efficient Parallel Computation of Connected Components in Billion-Scale Networks · IEEE Trans. Knowl. Data Eng. 2015
Bioinformatics and computational biology › protein analysis › protein-protein interaction
protein-protein interaction network
0.212023
Large-scale prediction of adverse drug reactions-related proteins with network embedding · Bioinform. 2023
Graph data management › graph processing
graph processing systems
0.212013
TurboGraph: a fast parallel graph engine handling billion-scale graphs in a single PC · KDD 2013
Graph data management › graph processing
out-of-core graph processing
0.212013
TurboGraph: a fast parallel graph engine handling billion-scale graphs in a single PC · KDD 2013
Storage systems
flash and SSD
0.122016
DSP-CC: I/O efficient parallel computation of connected components in billion-scale networks · ICDE 2016
TurboGraph: a fast parallel graph engine handling billion-scale graphs in a single PC · KDD 2013
Distributed systems
distributed graph processing
0.112015
DSP-CC-: I/O Efficient Parallel Computation of Connected Components in Billion-Scale Networks · IEEE Trans. Knowl. Data Eng. 2015

Methods — techniques the papers use, named apart from their topics

single-target compound · 0.7network embedding · 0.7secondary storage exploitation · 0.5i/o-efficient parallel computation · 0.5sequential disk access · 0.4page-level cache · 0.4mapreduce · 0.4pin-and-slide · 0.3multi-core parallelism · 0.3i/o parallelism · 0.3
YearPublicationVenuePosition
2023 Large-scale prediction of adverse drug reactions-related proteins with network embedding
abstract
MOTIVATION: Adverse drug reactions (ADRs) are a major issue in drug development and clinical pharmacology. As most ADRs are caused by unintended activity at off-targets of drugs, the identification of drug targets responsible for ADRs becomes a key process for resolving ADRs. Recently, with the increase in the number of ADR-related data sources, several computational methodologies have been proposed to analyze ADR-protein relations. However, the identification of ADR-related proteins on a large scale with high reliability remains an important challenge. RESULTS: In this article, we suggest a computational approach, Large-scale ADR-related Proteins Identification with Network Embedding (LAPINE). LAPINE combines a novel concept called single-target compound with a network embedding technique to enable large-scale prediction of ADR-related proteins for any proteins in the protein-protein interaction network. Analysis of benchmark datasets confirms the need to expand the scope of potential ADR-related proteins to be analyzed, as well as LAPINE's capability for high recovery of known ADR-related proteins. Moreover, LAPINE provides more reliable predictions for ADR-related proteins (Value-added positive predictive value = 0.12), compared to a previously proposed method (P < 0.001). Furthermore, two case studies show that most predictive proteins related to ADRs in LAPINE are supported by literature evidence. Overall, LAPINE can provide reliable insights into the relationship between ADRs and proteomes to understand the mechanism of ADRs leading to their prevention. AVAILABILITY AND IMPLEMENTATION: The source code is available at GitHub (https://github.com/rupinas/LAPINE) and Figshare (https://figshare.com/articles/software/LAPINE/21750245) to facilitate its use. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jaesub Park, Sangyeon Lee, Kwansoo Kim, Jaegyun Jung, Doheon Lee
Bioinform.2
2019 Visualizing multifunctional PPI network with Gene Ontology annotation
abstract
Protein-protein interaction (PPI) network data contain interactions between two proteins in organisms. Some proteins are involved in more than one biological process. Those multifunctional proteins are important in the biological system because they connect various biological pathways. However, the multifunctionality of proteins increases the complexity of the PPI network and makes visualizing a challenging problem. In this research, we suggest a new method to visualize the multifunctional PPI network by integrating interaction data and Gene Ontology (GO) terms of each protein. By matching GO annotations of two proteins, we annotate edges between those two proteins with common GO terms. Also, we suggest a novel algorithm for color mapping specialized to express hierarchical relationships among GO terms based on the multidimensional scaling algorithm. Annotated edges are divided corresponding to their GO terms, and visualized by selected color, width, and opacity to represent semantic relationships among edges clearly. Those visual features are assigned to represent semantic distance among GO terms and their concreteness. This visualization method is available through the web.
Sangyeon Lee, Doheon Lee
BIBM1
2019 A Conversation with Actuators: An Exploratory Design Environment for Hybrid Materials
abstract
An exciting, expanding palette of hybrid materials is emerging that can be programmed to actuate by a range of external and internal stimuli. However, there exists a dichotomy between the physicality of the actuators and the intangible computational signal that is used to program them. For material practitioners, this lack of physical cues limits their ability to engage in a "conversation with materials" (CwM). This paper presents a creative workstation for supporting this epistemological style by bringing a stronger physicality to the computational signal and balance the conversation between physical and digital actors. The station utilizes a streaming architecture to distribute control across multiple devices and leverage the rich spatial cognition that a physical space affords. Through a formal user study, we characterize the actuation design practice supported by the CwM workstation and discuss opportunities for tangible interfaces to hybrid materials.
César Torres 0001, Molly Jane Pearce Nicholas, Sangyeon Lee, Eric Paulos
TEI3
2017 Coupling effects on turning points of infectious diseases epidemics in scale-free networks
abstract
BACKGROUND: Pandemic is a typical spreading phenomenon that can be observed in the human society and is dependent on the structure of the social network. The Susceptible-Infective-Recovered (SIR) model describes spreading phenomena using two spreading factors; contagiousness (β) and recovery rate (γ). Some network models are trying to reflect the social network, but the real structure is difficult to uncover. METHODS: We have developed a spreading phenomenon simulator that can input the epidemic parameters and network parameters and performed the experiment of disease propagation. The simulation result was analyzed to construct a new marker VRTP distribution. We also induced the VRTP formula for three of the network mathematical models. RESULTS: We suggest new marker VRTP (value of recovered on turning point) to describe the coupling between the SIR spreading and the Scale-free (SF) network and observe the aspects of the coupling effects with the various of spreading and network parameters. We also derive the analytic formulation of VRTP in the fully mixed model, the configuration model, and the degree-based model respectively in the mathematical function form for the insights on the relationship between experimental simulation and theoretical consideration. CONCLUSIONS: We discover the coupling effect between SIR spreading and SF network through devising novel marker VRTP which reflects the shifting effect and relates to entropy.
Kiseong Kim, Sangyeon Lee, Doheon Lee, Kwang Hyung Lee
BMC Bioinform.2
2016 DSP-CC: I/O efficient parallel computation of connected components in billion-scale networks
abstract
Computing connected components (CC) is a core operation on graph data. Since billion-scale graphs cannot be resident in memory of a single machine, there have been proposed a number of distributed graph processing methods. The representative ones for CC are Hash-To-Min and PowerGraph. Hash-To-Min focuses on minimizing the number of MapReduce rounds, but is still slower than in-memory methods, PowerGraph is a fast and general in-memory graph method, but requires a lot of machines for handling billion-scale graphs. We propose an ultra-fast parallel method DSP-CC, using only a single PC that exploits secondary storage like a PCI-E SSD for handling billion-scale graphs. It can compute connected components I/O efficiently using only a limited size of memory. Our experimental results show that DSP-CC significantly outperforms the representative methods including Hash-To-Min and PowerGraph.
Min-Soo Kim 0002, Sangyeon Lee, Wook-Shin Han, Himchan Park, Jeonghoon Lee 0004
ICDE2
2015 DSP-CC-: I/O Efficient Parallel Computation of Connected Components in Billion-Scale Networks
abstract
Computing connected components is a core operation on graph data. Since billion-scale graphs cannot be resident in memory of a single server, several approaches based on distributed machines have recently been proposed. The representative methods are$\mathsf{Hash\hbox{-}To\hbox{-}Min}$and$\mathsf{PowerGraph}$.$\mathsf{Hash\hbox{-}To\hbox{-}Min}$is the state-of-the artdisk-baseddistributed method which minimizes the number of MapReduce rounds.$\mathsf{PowerGraph}$is the-state-of-the-artin-memorydistributed system, which is typically faster than the disk-based distributed one, however, requires a lot of machines for handling billion-scale graphs. In this paper, we propose an I/O efficient parallel algorithm for billion-scale graphs in a single PC. We first propose theDisk-based Sequential access-oriented Parallel processing(DSP) model that exploits sequential disk access in terms of disk I/Os and parallel processing in terms of computation. We then propose an ultra-fast disk-based parallel algorithm for computing connected components,$\mathsf{DSP\hbox{-}CC}$, which largely improves the performance through sequential disk scan andpage-level cache-conscious parallel processing. Extensive experimental results show that$\mathsf{DSP\hbox{-}CC}$1) computes connected components in billion-scale graphs using the limited memory size whereas in-memory algorithms can only support medium-sized graphs with the same memory size, and 2) significantly outperforms all distributed competitors as well as a representative disk-based parallel method.
Min-Soo Kim 0002, Sangyeon Lee, Wook-Shin Han, Himchan Park, Jeonghoon Lee 0004
IEEE Trans. Knowl. Data Eng.2
2014 OPT: a new framework for overlapped and parallel triangulation in large-scale graphs
abstract
Graph triangulation, which finds all triangles in a graph, has been actively studied due to its wide range of applications in the network analysis and data mining. With the rapid growth of graph data size, disk-based triangulation methods are in demand but little researched. To handle a large-scale graph which does not fit in memory, we must iteratively load small parts of the graph. In the existing literature, achieving the ideal cost has been considered to be impossible for billion-scale graphs due to the memory size constraint. In this paper, we propose an overlapped and parallel disk-based triangulation framework for billion-scale graphs, OPT, which achieves the ideal cost by (1) full overlap of the CPU and I/O operations and (2) full parallelism of multi-core CPU and FlashSSD I/O. In OPT, triangles in memory are called the internal triangles while triangles constituting vertices in memory and vertices in external memory are called the external triangles. At the macro level, OPT overlaps the internal triangulation and the external triangulation, while it overlaps the CPU and I/O operations at the micro level. Thereby, the cost of OPT is close to the ideal cost. Moreover, OPT instantiates both vertex-iterator and edge-iterator models and benefits from multi-thread parallelism on both types of triangulation. Extensive experiments conducted on large-scale datasets showed that (1) OPT achieved the elapsed time close to that of the ideal method with less than 7% of overhead under the limited memory budget, (2) OPT achieved linear speed-up with an increasing number of CPU cores, (3) OPT outperforms the state-of-the-art parallel method by up to an order of magnitude with 6 CPU cores, and (4) for the first time in the literature, the triangulation results are reported for a billion-vertex scale real-world graph.
Jinha Kim, Wook-Shin Han, Sangyeon Lee, Kyungyeol Park, Hwanjo Yu
SIGMOD Conference3
2013 TurboGraph: a fast parallel graph engine handling billion-scale graphs in a single PC
abstract
Graphs are used to model many real objects such as social networks and web graphs. Many real applications in various fields require efficient and effective management of large-scale graph structured data. Although distributed graph engines such as GBase and Pregel handle billion-scale graphs, the user needs to be skilled at managing and tuning a distributed system in a cluster, which is a nontrivial job for the ordinary user. Furthermore, these distributed systems need many machines in a cluster in order to provide reasonable performance. In order to address this problem, a disk-based parallel graph engine called Graph-Chi, has been recently proposed. Although Graph-Chi significantly outperforms all representative (disk-based) distributed graph engines, we observe that Graph-Chi still has serious performance problems for many important types of graph queries due to 1) limited parallelism and 2) separate steps for I/O processing and CPU processing. In this paper, we propose a general, disk-based graph engine called TurboGraph to process billion-scale graphs very efficiently by using modern hardware on a single PC. TurboGraph is the first truly parallel graph engine that exploits 1) full parallelism including multi-core parallelism and FlashSSD IO parallelism and 2) full overlap of CPU processing and I/O processing as much as possible. Specifically, we propose a novel parallel execution model, called pin-and-slide. TurboGraph also provides engine-level operators such as BFS which are implemented under the pin-and-slide model. Extensive experimental results with large real datasets show that TurboGraph consistently and significantly outperforms Graph-Chi by up to four orders of magnitude! Our implementation of TurboGraph is available at ``http://wshan.net/turbograph}" as executable files.
Wook-Shin Han, Sangyeon Lee, Kyungyeol Park, Jeonghoon Lee 0004, Min-Soo Kim 0002, Jinha Kim, Hwanjo Yu
KDD2