Youngchoon Park

dblp:39/6523 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-authorArtificial intelligence and machine learning · 3 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Data integration and cleaning · 92% Distributed and cloud data management · 5% Information retrieval · 4%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 100%
Software engineering, system software, and programming languages
1 paper
Services computing and microservices · 100%

Topics — the 8 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
entity matching
1.342019
Entity Matching Meets Data Science: A Progress Report from the Magellan Project · SIGMOD Conference 2019
CloudMatcher: A Hands-Off Cloud/Crowd Service for Entity Matching · Proc. VLDB Endow. 2018
Deep Learning for Entity Matching: A Design Space Exploration · SIGMOD Conference 2018
Data integration and cleaning › entity matching
deep entity matching
0.312018
Deep Learning for Entity Matching: A Design Space Exploration · SIGMOD Conference 2018
Cloud and datacenter computing
cluster resource management and scheduling
0.112017
Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services · SIGMOD Conference 2017
Natural language and speech › Information extraction and text analysis
word sense disambiguation
0.012001
Technique for eliminating irrelevant terms in term rewriting for annotated media retrieval · ACM Multimedia 2001
Information retrieval › document retrieval
concept-based retrieval
0.012001
Technique for eliminating irrelevant terms in term rewriting for annotated media retrieval · ACM Multimedia 2001
Information retrieval › search engines › semantic search
ontology-based retrieval
0.012001
Technique for eliminating irrelevant terms in term rewriting for annotated media retrieval · ACM Multimedia 2001
Multimedia analysis and retrieval
video retrieval
0.012000
MP7TV: a system for content-based querying and retrieval of digital video · ACM Multimedia 2000
Multimedia analysis and retrieval › video retrieval
video query
0.012000
MP7TV: a system for content-based querying and retrieval of digital video · ACM Multimedia 2000

Methods — techniques the papers use, named apart from their topics

crowdsourcing · 1.2interactive labeling · 0.7query optimization · 0.6operator implementation · 0.6recurrent neural network · 0.3deep learning · 0.3attention · 0.3term rewriting · 0.1discounting and redistribution model · 0.1lexicographical approach · 0.0
YearPublicationVenuePosition
2019 Entity Matching Meets Data Science: A Progress Report from the Magellan Project
abstract
Entity matching (EM) finds data instances that refer to the same real-world entity. In 2015, we started the Magellan project at UW-Madison, joint with industrial partners, to build EM systems. Most current EM systems are stand-alone monoliths. In contrast, Magellan borrows ideas from the field of data science (DS), to build a new kind of EM systems, which is an ecosystem of interoperable tools. \em This paper provides a progress report on the past 3.5 years of Magellan, focusing on the system aspects and on how ideas from the field of data science have been adapted to the EM context. We argue why EM can be viewed as a special class of DS problems, and thus can benefit from system building ideas in DS. We discuss how these ideas have been adapted to build \pymatcher\ and \cloudmatcher, EM tools for power users and lay users. These tools have been successfully used in 21 EM tasks at 12 companies and domain science groups, and have been pushed into production for many customers. We report on the lessons learned, and outline a new envisioned Magellan ecosystem, which consists of not just on-premise Python tools, but also interoperable microservices deployed, executed, and scaled out on the cloud, using tools such as Dockers and Kubernetes.
Yash Govind, Pradap Konda, Paul Suganthan G. C., Philip Martinkus, Palaniappan Nagarajan, Aravind Soundararajan, Sidharth Mudgal, Jeffrey R. Ballard, Haojun Zhang, Adel Ardalan, Sanjib Das, Derek Paulsen, Amanpreet Singh Saini, Erik Paulson 0001, Youngchoon Park, Marshall Carter, Mingju Sun, Glenn Fung, AnHai Doan
SIGMOD Conference16
2018 MatchCatcher: A Debugger for Blocking in Entity Matching
Pradap Konda, Paul Suganthan G. C., AnHai Doan, Benjamin Snyder, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Vijay Raghavendra
EDBT6
2018 Deep Learning for Entity Matching: A Design Space Exploration
abstract
Entity matching (EM) finds data instances that refer to the same real-world entity. In this paper we examine applying deep learning (DL) to EM, to understand DL's benefits and limitations. We review many DL solutions that have been developed for related matching tasks in text processing (e.g., entity linking, textual entailment, etc.). We categorize these solutions and define a space of DL solutions for EM, as embodied by four solutions with varying representational power: SIF, RNN, Attention, and Hybrid. Next, we investigate the types of EM problems for which DL can be helpful. We consider three such problem types, which match structured data instances, textual instances, and dirty instances, respectively. We empirically compare the above four DL solutions with Magellan, a state-of-the-art learning-based EM solution. The results show that DL does not outperform current solutions on structured EM, but it can significantly outperform them on textual and dirty EM. For practitioners, this suggests that they should seriously consider using DL for textual and dirty EM problems. Finally, we analyze DL's performance and discuss future research directions.
Sidharth Mudgal, Theodoros Rekatsinas, AnHai Doan, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, Vijay Raghavendra
SIGMOD Conference5
2018 CloudMatcher: A Hands-Off Cloud/Crowd Service for Entity Matching
abstract
As data science applications proliferate, more and more lay users must perform data integration (DI) tasks, which used to be done by sophisticated CS developers. Thus, it is increasingly critical that we develop hands-off DI services, which lay users can use to perform such tasks without asking for help from developers. We propose to demonstrate such a service. Specifically, we will demonstrate CloudMatcher, a hands-off cloud/crowd service for entity matching (EM). To use CloudMatcher to match two tables, a lay user only needs to upload them to the CloudMatcher's Web page then iteratively label a set of tuple pairs as match/no-match. Alternatively, the user can enlist a crowd of workers to label the pairs. In either case, the lay user can easily perform EM end-to-end without having to involve any developers. Cloud-Matcher has been used in several domain science projects at UW-Madison and at several organizations, and is scheduled to be deployed in a large company in Summer 2018. In the demonstration we will show how easy it is for lay users to perform EM (either via interactive labeling or crowdsourcing), how users can easily create and experiment with a range of EM workflows, and how CloudMatcher can scale to many concurrent users and large datasets.
Yash Govind, Erik Paulson 0001, Palaniappan Nagarajan, Paul Suganthan G. C., AnHai Doan, Youngchoon Park, Glenn Fung, Devin Conathan, Marshall Carter, Mingju Sun
Proc. VLDB Endow.6
2017 Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services
abstract
Many works have applied crowdsourcing to entity matching (EM). While promising, these approaches are limited in that they often require a developer to be in the loop. As such, it is difficult for an organization to deploy multiple crowdsourced EM solutions, because there are simply not enough developers. To address this problem, a recent work has proposed Corleone, a solution that crowdsources the entire EM workflow, requiring no developers. While promising, Corleone is severely limited in that it does not scale to large tables. We propose Falcon, a solution that scales up the hands-off crowdsourced EM approach of Corleone, using RDBMS-style query execution and optimization over a Hadoop cluster. Specifically, we define a set of operators and develop efficient implementations. We translate a hands-off crowdsourced EM workflow into a plan consisting of these operators, optimize, then execute the plan. These plans involve both machine and crowd activities, giving rise to novel optimization techniques such as using crowd time to mask machine time. Extensive experiments show that Falcon can scale up to tables of millions of tuples, thus providing a practical solution for hands-off crowdsourced EM, to build cloud-based EM services.
Sanjib Das, Paul Suganthan G. C., AnHai Doan, Jeffrey F. Naughton, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, Vijay Raghavendra, Youngchoon Park
SIGMOD Conference9
2015 Connected Smart Buildings, a New Way to Interact with Buildings
abstract
Summary form only given. Devices, people, information and software applications rarely live in isolation in modern building management. For example, networked sensors that monitor the performance of a chiller are common and collected data are delivered to building automation systems to optimize energy use. Detected possible failures are also handed to facility management staffs for repairs. Physical and cyber security services have to be incorporated to prevent improper access of not only HVAC (Heating, Ventilation, Air Conditioning) equipment but also control devices. Harmonizing these connected sensors, control devices, equipment and people is a key to provide more comfortable, safe and sustainable buildings. Nowadays, devices with embedded intelligences and communication capabilities can interact with people directly. Traditionally, few selected people (e.g., facility managers in building industry) have access and program the device with fixed operating schedule while a device has a very limited connectivity to an operating environment and context. Modern connected devices will learn and interact with users and other connected things. This would be a fundamental shift in ways in communication from unidirectional to bi-directional. A manufacturer will learn how their products and features are being accessed and utilized. An end user or a device on behalf of a user can interact and communicate with a service provider or a manufacturer without go though a distributer, almost real time basis. This will requires different business strategies and product development behaviors to serve connected customers' demands. Connected things produce enormous amount of data that result many questions and technical challenges in data management, analysis and associated services. In this talk, we will brief some of challenges that we have encountered In developing connected building solutions and services. More specifically, (1) semantic interoperability requirements among smart sensors, actuators, lighting, security and control and business applications, (2) engineering challenges in managing massively large time sensitive multi-media data in a cloud at global scale, and (3) security and privacy concerns are presented.
Youngchoon Park
IC2E1
2009 Counting People in Groups
abstract
Cameras are becoming a common tool for automated vision purposes due to their low cost. In an era of growing security concerns, camera surveillance systems have become not only important but also necessary. Algorithms for several tasks such as detecting abandoned objects and tracking people have already been successfully developed. While tracking people is relatively easy, counting people in groups is much more challenging. The mutual occlusions between people in a group make it difficult to provide an exact count. The aim of this work is to present a method of estimating the number of people in group scenarios. Several considerations for counting people are illustrated in this paper, and experimental results of the method are described and discussed.
Duc Fehr, Ravishankar Sivalingam, Vassilios Morellas, Nikolaos Papanikolopoulos, Osama A. Lotfallah, Youngchoon Park
AVSS6
2004 A Multimedia Information Repository for Cross Cultural Dance Studies
Forouzan Golshani, Pegge Vissicaro, Youngchoon Park
Multim. Tools Appl.3
2002 Towards Retrieval of Visual Information Based on the Semantic Models
Youngchoon Park, Pankoo Kim, Wonpil Kim, Jeong-Jun Song, Sethuraman Panchanathan
DEXA1
2002 A Model-Based Approach to Semantic-Based Retrieval of Visual Information
Forouzan Golshani, Youngchoon Park, Sethuraman Panchanathan
SOFSEM2
2001 Concept-Based Visual Information Management with Large Lexical Corpus
Youngchoon Park, Pankoo Kim, Forouzan Golshani, Sethuraman Panchanathan
DEXA1
2001 Technique for eliminating irrelevant terms in term rewriting for annotated media retrieval
abstract
In this paper, we present an efficient term rewriting technique that computes a degree of term to domain relevance. The proposed method resolves the problems in ontology integrated concept search. Those problems are (i) Pre-defined concept classes in ontology are not relevant to users (no proper concept class for a target annotation has not found). (ii) Too many similar concept classes are provided to a user therefore, a user may fail to choose a correct semantic class for a target annotation (ordinary users are not an expert in concept classification). The method uses sense disambiguation task for finding relevant terms for a given domain. Sense disambiguation requires term-to-term similarity measurement and term frequency measurement. For fair modeling of not observed term frequencies, discounting and redistribution model is applied. The proposed method is a compliment to our previous work presented in [13][14]. Robustness of our method is demonstrated through human judgment test that shows our method allows prediction of precise term list (overall 75% of correct prediction) that are relevant to a given domain.
Youngchoon Park, Pankoo Kim, Forouzan Golshani, Sethuraman Panchanathan
ACM Multimedia1
2001 A comprehensive curriculum for IT education and workforce development: an engineering approach
abstract
Noting the shortage of IT professionals nationally [1], we propose a comprehensive curriculum that supports a variety of programs geared to all ages from early school years to retirement and beyond. Current IT workforce development efforts are limited to training, and have not as yet focused on education and professional development. Largely, this is due to a lack of a science underpinning for IT related curricula. Without such a unified science component, a structured organization of information related concepts cannot be derived.Our proposal includes the development of a number of programs addressing the needs of a variety of learners ranging from elementary school through college and beyond. Seven programs, each with a specific emphasis for various groups, are being developed. Such essential issues as industrial-academic liaisons, workforce (re)training, promotional and awareness programs, teacher training, and IT professional role redefinition, are integral pieces of this project. All developments will be firmly founded on the scientific framework of information science and engineering [2].This work is supported by NSF grant DUE-9950168.
Forouzan Golshani, Sethuraman Panchanathan, Oris Friesen, Youngchoon Park, Jeong-Jun Song
SIGCSE4
2000 The Role of Color in Content Based Image Retrieval
abstract
Human color perception is subjective. In addition to RGB or HSV values representing color content, psychological factors, circumstantial factors, environmental factors, and physiological factors play an important role in encapsulating the color content. Typically, only the RGB or HSV values are used in indexing and retrieval of color images. In this paper, we demonstrate the superior retrieval performance of techniques which employ all of the above factors in color retrieval. We first present the variety of factors involved in human color perception and also an evaluation of the existing color indexing methods. Several interesting problems including comparison of images in the color-perceptual domain and retrieval by color affection are illustrated as potential novel image query mechanisms.
Sethuraman Panchanathan, Youngchoon Park, Pankoo Kim, Forouzan Golshani
ICIP2
2000 Efficient tools for power annotation of visual contents: a lexicographical approach
Youngchoon Park
ACM Multimedia1
2000 MP7TV: a system for content-based querying and retrieval of digital video
Youngchoon Park, Antonio Pizzarello
ACM Multimedia1
1997 ImageRoadMap: A New Content-based Image Retrieval System
Youngchoon Park, Forouzan Golshani
DEXA1