Sunju Park

dblp:19/6392 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
1since 2021 · last 2022
0000-0002-6398-4612ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12Artificial intelligence and machine learning · 8 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4Computer networks · 2Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Graph data management · 45% Recommender systems · 20% Indexing and storage engines · 17%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 67% Performance modeling and evaluation · 33%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Graph data management › graph processing
graph processing systems
1.022022
A Data Layout With Good Data Locality for Single-Machine Based Graph Engines · IEEE Trans. Computers 2022
RealGraph: A Graph Engine Leveraging the Power-Law Distribution of Real-World Graphs · WWW 2019
Indexing and storage engines › storage management
data layout
0.612022
A Data Layout With Good Data Locality for Single-Machine Based Graph Engines · IEEE Trans. Computers 2022
Graph data management
graph processing
0.612022
A Data Layout With Good Data Locality for Single-Machine Based Graph Engines · IEEE Trans. Computers 2022
Storage systems
data layout
0.612022
A Data Layout With Good Data Locality for Single-Machine Based Graph Engines · IEEE Trans. Computers 2022
Recommender systems
collaborative filtering
0.312018
How to Impute Missing Ratings?: Claims, Solution, and Its Application to Collaborative Filtering · WWW 2018
Data integration and cleaning › missing data
missing value imputation
0.312018
How to Impute Missing Ratings?: Claims, Solution, and Its Application to Collaborative Filtering · WWW 2018
Recommender systems › collaborative filtering
rating prediction
0.312018
How to Impute Missing Ratings?: Claims, Solution, and Its Application to Collaborative Filtering · WWW 2018
Performance modeling and evaluation
workload characterization
0.212022
A Data Layout With Good Data Locality for Single-Machine Based Graph Engines · IEEE Trans. Computers 2022
Information retrieval
citation analysis
0.112010
A link-based similarity measure for scientific literature · WWW 2010
Information retrieval › similarity measure
link-based similarity
0.112010
A link-based similarity measure for scientific literature · WWW 2010
Information retrieval › document retrieval › domain-specific retrieval
scientific literature search
0.012010
A link-based similarity measure for scientific literature · WWW 2010
Information retrieval
search engines
0.012010
A link-based similarity measure for scientific literature · WWW 2010

Methods — techniques the papers use, named apart from their topics

hierarchical indicator · 0.8block-based workload allocation · 0.8deep learning · 0.3in-link and out-link transformation · 0.1
YearPublicationVenuePosition
2022 A Data Layout With Good Data Locality for Single-Machine Based Graph Engines
abstract
Graph engines have been used in many applications to handle big graphs efficiently. The majority of the research to improve their performance has focused primarily on the design of efficient graph processing. This paper claims,however,the focus should be given also to graph storage design. This is because good storage design can improve both CPU performance and I/O performance of graph engines. In this paper,we propose an efficient data layout for single-machine based graph engines. We identify the common node access pattern of the graph algorithms running on single-machine based graph engines. Based on this finding,we propose the breadth-first (BF) data layout which places the nodes processed together in the same or adjacent storage space so that they can be accessed together as much as possible. The experimental results show that the BF data layout improves both CPU and I/O performances significantly in all single-machine based graph engines.
Yong-Yeon Jo, Myung-Hwan Jang, Sang-Wook Kim, Sunju Park
IEEE Trans. Computers4
2019 RealGraph: A Graph Engine Leveraging the Power-Law Distribution of Real-World Graphs
abstract
As the size of real-world graphs has drastically increased in recent years, a wide variety of graph engines have been developed to deal with such big graphs efficiently. However, the majority of graph engines have been designed without considering the power-law degree distribution of real-world graphs seriously. Two problems have been observed when existing graph engines process real-world graphs: inefficient scanning of the sparse indicator and the delay in iteration progress due to uneven workload distribution. In this paper, we propose RealGraph, a single-machine based graph engine equipped with the hierarchical indicator and the block-based workload allocation. Experimental results on real-world datasets show that RealGraph significantly outperforms existing graph engines in terms of both speed and scalability.
Yong-Yeon Jo, Myung-Hwan Jang, Sang-Wook Kim, Sunju Park
WWW4
2018 How to Impute Missing Ratings?: Claims, Solution, and Its Application to Collaborative Filtering
abstract
Data sparsity is one of the biggest problems faced by collaborative filtering used in recommender systems. Data imputation alleviates the data sparsity problem by inferring missing ratings and imputing them to the original rating matrix. In this paper, we identify the limitations of existing data imputation approaches and suggest three new claims that all data imputation approaches should follow to achieve high recommendation accuracy. Furthermore, we propose a deep-learning based approach to compute imputed values that satisfies all three claims. Based on our hypothesis that most pre-use preferences (e.g., impressions) on items lead to their post-use preferences (e.g., ratings), our approach tries to understand via deep learning how pre-use preferences lead to post-use preferences differently depending on the characteristics of users and items. Through extensive experiments on real-world datasets, we verify our three claims and hypothesis, and also demonstrate that our approach significantly outperforms existing state-of-the-art approaches.
Youngnam Lee, Sang-Wook Kim, Sunju Park, Xing Xie 0001
WWW3
2017 Energy efficient data collection in sink-centric wireless sensor networks: A cluster-ring approach
Soo-Hoon Moon, Sunju Park, Seung-Jae Han
Comput. Commun.2
2016 C-Rank: A link-based similarity measure for scientific literature databases
Seok-Ho Yoon, Sang-Wook Kim, Sunju Park
Inf. Sci.3
2016 Credible, resilient, and scalable detection of software plagiarism using authority histograms
Dong-Kyu Chae, Jiwoon Ha, Sang-Wook Kim, Boojoong Kang, Eul-Gyu Im, Sunju Park
Knowl. Based Syst.6
2015 An analysis on information diffusion through BlogCast in a blogosphere
Jiwoon Ha, Sang-Wook Kim, Christos Faloutsos, Sunju Park
Inf. Sci.4
2015 A community-based sampling method using DPL for online social networks
Seok-Ho Yoon, Jiwon Hong, Sang-Wook Kim, Sunju Park
Inf. Sci.5
2015 Can You Trust Online Ratings? A Mutual Reinforcement Model for Trustworthy Online Rating Systems
abstract
The average of customer ratings on a product, which we call a reputation, is one of the key factors in online purchasing decisions. There is, however, no guarantee of the trustworthiness of a reputation since it can be manipulated rather easily. In this paper, we define false reputation as the problem of a reputation being manipulated by unfair ratings and design a general framework that provides trustworthy reputations. For this purpose, we propose TRUE-REPUTATION, an algorithm that iteratively adjusts a reputation based on the confidence of customer ratings. We also show the effectiveness of TRUE-REPUTATION through extensive experiments in comparisons to state-of-the-art approaches.
Hyun-Kyo Oh, Sang-Wook Kim, Sunju Park, Ming Zhou 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2014 Satisfying the target network lifetime in wireless sensor networks
Keon-Taek Lee, Hak-Jin Kim, Sunju Park, Seung-Jae Han
Comput. Networks3
2014 Accurate Approximation of the Earth Mover's Distance in Linear Time
Min-Hee Jang, Sang-Wook Kim, Christos Faloutsos, Sunju Park
J. Comput. Sci. Technol.4
2013 Trustable aggregation of online ratings
abstract
The average of the customer ratings on the product, which we call reputation, is one of the key factors in online purchasing decision of a product. There is, however, no guarantee in the trustworthiness of the reputation since it can be manipulated rather easily. In this paper, we define false reputation as the problem of the reputation to be manipulated by unfair ratings, and design a general framework that provides trustable reputation. For this purpose, we propose TRUEREPUTATION, an algorithm that iteratively adjusts the reputation based on the confidence of customer ratings.
Hyun-Kyo Oh, Sang-Wook Kim, Sunju Park, Ming Zhou 0001
CIKM3
2012 Top-N recommendation through belief propagation
abstract
The top-n recommendation focuses on finding the top-n items that the target user is likely to purchase rather than predicting his/her ratings on individual items. In this paper, we propose a novel method that provides top-n recommendation by probabilistically determining the target user's preference on items. This method models the purchasing relationships between users and items as a bipartite graph and employs Belief Propagation to compute the preference of the target user on items. We analyze the proposed method in detail by examining the changes in recommendation accuracy under different parameter settings. We also show that the proposed method is up to 40% more accurate than an existing method by comparing it with an RWR-based method via extensive experiments.
Jiwoon Ha, Soon-Hyoung Kwon, Sang-Wook Kim, Christos Faloutsos, Sunju Park
CIKM5
2012 Subject-based extraction of a latent blog community
Seok-Ho Yoon, Jung-Hwan Shin, Sang-Wook Kim, Sunju Park, Jae Bum Lee
Inf. Sci.4
2011 A linear-time approximation of the earth mover's distance
abstract
Color descriptors are one of the important features used in content-based image retrieval. The dominant color descriptor (DCD) represents a few perceptually dominant colors in an image through color quantization. For image retrieval based on DCD, the earth mover's distance and the optimal color composition distance are proposed to measure the dissimilarity between two images. Although providing good retrieval results, both methods are too time-consuming to be used in a large image database. To solve the problem, we propose a new distance function that calculates an approximate earth mover's distance in linear time. To calculate the dissimilarity in linear time, the proposed approach employs the space-filling curve for multidimensional color space. To improve the accuracy, the proposed approach uses multiple curves and adjusts the color positions. As a result, our approach achieves order-of-magnitude time improvement but incurs small errors. We have performed extensive experiments to show the effectiveness and efficiency of the proposed approach. The results reveal that our approach achieves almost the same results with the EMD in linear time.
Min-Hee Jang, Sang-Wook Kim, Christos Faloutsos, Sunju Park
CIKM4
2011 Determining Content Power Users in a Blog Network: An Approach and Its Applications
abstract
In a blog network, there are special users who induce other users to actively utilize blogs. Identifying such influential users is important when establishing business policy and business models for the blog network. This paper defines the users whose contents exhibit significant influence over other users as content power users (CPUs) and proposes a method of identifying them. We analyze the performance of the proposed method by applying it to an actual blog network and comparing its results with those of preexisting methods for determining power users. The experimental results demonstrate that the definition of CPUs is adequate to address the dynamic nature of the blogosphere and the main concerns of the blog industry. We also discuss the business models based on CPUs that could be used to stimulate user activities in a blog network.
Seung-Hwan Lim, Sang-Wook Kim, Sunju Park, Joon Ho Lee
IEEE Trans. Syst. Man Cybern. Part A3
2010 A link-based similarity measure for scientific literature
abstract
In this paper, we propose a new approach to measure sim-ilarities among academic papers based on their references. Our similarity measure uses both in-link and out-link by transforming in-link and out-link into undirected links.
Seok-Ho Yoon, Sang-Wook Kim, Sunju Park
WWW3
2009 Determining the strength of the propensities of a blog network
abstract
A blog network, composed of blogs and their relations, may exhibit two different propensities characterized by the purpose of use: an information-oriented propensity and a friendship-oriented propensity. Both propensities coexist in a blog network, and the degree of these propensities may play an important role in business and policy decisions of blog-related business. In this paper, we propose an automated method for determining the propensity values of a blog network. First, classification is used to judge the propensity values of the relation between two blogs. Then, by adding up the propensity values of all the relations in the network, one determines the propensity values of the whole network. Through extensive experiments using a large volume of real-world blog data, we demonstrate our method achieves a high level of accuracy in determining the propensity values of a relation. The results also suggest the applicability of our approach for determining the propensity values of a network.
Seok-Ho Yoon, Sang-Wook Kim, Sunju Park
CIDM3
2009 Extraction of a latent blog community based on subject
abstract
In the blogosphere, there exist posts relevant to a particular subject and blogs that show interests in the subject. In this paper, we define a set of such posts and blogs as "blog community" and propose a method for extracting the blog community associated with a particular subject. The proposed method is based on the idea that the blogs who have performed actions to the posts of a particular subject are the ones that have interests in the subject, and that the posts which have received actions from such blogs are the ones that contain the subject. The proposed method selects a small number of seed posts that contain the subject. Then, it selects the blogs that perform actions to the seed posts over some threshold and the posts that have received actions over some threshold. By repeating these two steps, it gradually expands the blog community. The experimental results show that the proposed method exhibits a higher level of accuracy than the methods proposed in prior research.
Seok-Ho Yoon, Jung-Hwan Shin, Sang-Wook Kim, Sunju Park
CIKM4
2004 Use of Markov Chains to Design an Agent Bidding Strategy for Continuous Double Auctions
abstract
As computational agents are developed for increasingly complicated e-commerce applications, the complexity of the decisions they face demands advances in artificial intelligence techniques. For example, an agent representing a seller in an auction should try to maximize the seller?s profit by reasoning about a variety of possibly uncertain pieces of information, such as the maximum prices various buyers might be willing to pay, the possible prices being offered by competing sellers, the rules by which the auction operates, the dynamic arrival and matching of offers to buy and sell, and so on. A naive application of multiagent reasoning techniques would require the seller?s agent to explicitly model all of the other agents through an extended time horizon, rendering the problem intractable for many realistically-sized problems. We have instead devised a new strategy that an agent can use to determine its bid price based on a more tractable Markov chain model of the auction process. We have experimentally identified the conditions under which our new strategy works well, as well as how well it works in comparison to the optimal performance the agent could have achieved had it known the future. Our results show that our new strategy in general performs well, outperforming other tractable heuristic strategies in a majority of experiments, and is particularly effective in a 'seller?s market', where many buy offers are available.
Sunju Park, Edmund H. Durfee, William P. Birmingham
J. Artif. Intell. Res.1
2000 Emergent Properties of a Market-based Digital Library with Strategic Agents
Sunju Park, Edmund H. Durfee, William P. Birmingham
Auton. Agents Multi Agent Syst.1