Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jaeho Shin 0001

dblp:38/1325-1 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
0since 2021 · last 2017
0000-0001-5280-3356ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Knowledge graphs · 35% Graph data management · 27% Machine learning and data management · 21%
Software engineering, system software, and programming languages
3 papers
Debugging and program repair · 36% Empirical software engineering · 36% Programming languages and type systems · 28%
Artificial intelligence
2 papers
Knowledge representation and reasoning · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge graphs
knowledge graph construction
0.732017
Incremental knowledge base construction using DeepDive · VLDB J. 2017
Incremental Knowledge Base Construction Using DeepDive · Proc. VLDB Endow. 2015
Mindtagger: A Demonstration of Data Labeling in Knowledge Base Construction · Proc. VLDB Endow. 2015
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base construction
0.522017
Incremental knowledge base construction using DeepDive · VLDB J. 2017
Extracting Databases from Dark Data with DeepDive · SIGMOD Conference 2016
Machine learning and data management
data annotation
0.212015
Mindtagger: A Demonstration of Data Labeling in Knowledge Base Construction · Proc. VLDB Endow. 2015
Graph data management › graph processing
graph processing systems
0.212015
Graft: A Debugging Tool For Apache Giraph · SIGMOD Conference 2015
Machine learning and data management
inference optimization
0.212015
Incremental Knowledge Base Construction Using DeepDive · Proc. VLDB Endow. 2015
Debugging and program repair › concurrent program debugging
distributed debugging
0.212015
Graft: A Debugging Tool For Apache Giraph · SIGMOD Conference 2015
Graph data management › graph analytics
distributed graph analysis
0.212013
Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis · Proc. VLDB Endow. 2013
Graph data management
graph analytics
0.212013
Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis · Proc. VLDB Endow. 2013
Programming languages and type systems › logic programming
datalog
0.212013
Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis · Proc. VLDB Endow. 2013
Parallel and multicore computing › parallel computing
distributed execution
0.012013
Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis · Proc. VLDB Endow. 2013
Parallel and multicore computing
parallel programming models
0.012013
Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis · Proc. VLDB Endow. 2013

Methods — techniques the papers use, named apart from their topics

statistical inference · 1.0distant supervision · 0.6probabilistic inference · 0.5approximate computation · 0.5delta stepping · 0.3variational inference · 0.2sampling · 0.2rule-based optimizer · 0.2delta-stepping · 0.2
YearPublicationVenuePosition
2017 Incremental knowledge base construction using DeepDive
Christopher De Sa, Alexander Ratner, Christopher Ré, Jaeho Shin 0001, Sen Wu 0002, Ce Zhang 0001
VLDB J.4
2016 Extracting Databases from Dark Data with DeepDive
abstract
: the mass of text, tables, and images that are widely collected and stored but which cannot be exploited by standard relational tools. If the information in dark data - scientific papers, Web classified ads, customer service notes, and so on - were instead in a relational database, it would give analysts a massive and valuable new set of "big data." DeepDive is distinctive when compared to previous information extraction systems in its ability to obtain very high precision and recall at reasonable engineering cost; in a number of applications, we have used DeepDive to create databases with accuracy that meets that of human annotators. To date we have successfully deployed DeepDive to create data-centric applications for insurance, materials science, genomics, paleontologists, law enforcement, and others. The data unlocked by DeepDive represents a massive opportunity for industry, government, and scientific researchers. DeepDive is enabled by an unusual design that combines large-scale probabilistic inference with a novel developer interaction cycle. This design is enabled by several core innovations around probabilistic training and inference.
Ce Zhang 0001, Jaeho Shin 0001, Christopher Ré, Michael J. Cafarella, Feng Niu
SIGMOD Conference2
2015 Graft: A Debugging Tool For Apache Giraph
abstract
We address the problem of debugging programs written for Pregel-like systems. After interviewing Giraph and GPS users, we developed Graft. Graft supports the debugging cycle that users typically go through: (1) Users describe programmatically the set of vertices they are interested in inspecting. During execution, Graft captures the context information of these vertices across supersteps. (2) Using Graft's GUI, users visualize how the values and messages of the captured vertices change from superstep to superstep,narrowing in suspicious vertices and supersteps. (3) Users replay the exact lines of the code vertex.compute() function that executed for the suspicious vertices and supersteps, by copying code that Graft generates into their development environments' line-by-line debuggers. Graft also has features to construct end-to-end tests for Giraph programs. Graft is open-source and fully integrated into Apache Giraph's main code base.
Semih Salihoglu, Jaeho Shin 0001, Vikesh Khanna, Ba Quan Truong, Jennifer Widom
SIGMOD Conference2
2015 Mindtagger: A Demonstration of Data Labeling in Knowledge Base Construction
abstract
End-to-end knowledge base construction systems using statistical inference are enabling more people to automatically extract high-quality domain-specific information from unstructured data. As a result of deploying DeepDive framework across several domains, we found new challenges in debugging and improving such end-to-end systems to construct high-quality knowledge bases. DeepDive has an iterative development cycle in which users improve the data. To help our users, we needed to develop principles for analyzing the system's error as well as provide tooling for inspecting and labeling various data products of the system. We created guidelines for error analysis modeled after our colleagues' best practices, in which data labeling plays a critical role in every step of the analysis. To enable more productive and systematic data labeling, we created Mindtagger, a versatile tool that can be configured to support a wide range of tasks. In this demonstration, we show in detail what data labeling tasks are modeled in our error analysis guidelines and how each of them is performed using Mindtagger.
Jaeho Shin 0001, Christopher Ré, Michael J. Cafarella
Proc. VLDB Endow.1
2015 Incremental Knowledge Base Construction Using DeepDive
abstract
Populating a database with unstructured information is a long-standing problem in industry and research that encompasses problems of extraction, cleaning, and integration. Recent names used for this problem include dealing with dark data and knowledge base construction (KBC). In this work, we describe DeepDive, a system that combines database and machine learning ideas to help develop KBC systems, and we present techniques to make the KBC process more efficient. We observe that the KBC process is iterative, and we develop techniques to incrementally produce inference results for KBC systems. We propose two methods for incremental inference, based respectively on sampling and variational techniques. We also study the tradeoff space of these methods and develop a simple rule-based optimizer. DeepDive includes all of these contributions, and we evaluate Deep-Dive on five KBC systems, showing that it can speed up KBC inference tasks by up to two orders of magnitude with negligible impact on quality.
Jaeho Shin 0001, Sen Wu 0002, Christopher De Sa, Ce Zhang 0001, Christopher Ré
Proc. VLDB Endow.1
2013 Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis
abstract
Large-scale graph analysis is becoming important with the rise of world-wide social network services. Recently in SociaLite, we proposed extensions to Datalog to efficiently and succinctly implement graph analysis programs on sequential machines. This paper describes novel extensions and optimizations of SociaLite for parallel and distributed executions to support large-scale graph analysis. With distributed SociaLite, programmers simply annotate how data are to be distributed, then the necessary communication is automatically inferred to generate parallel code for cluster of multi-core machines. It optimizes the evaluation of recursive monotone aggregate functions using a delta stepping technique. In addition, approximate computation is supported in SociaLite, allowing programmers to trade off accuracy for less time and space. We evaluated SociaLite with six core graph algorithms used in many social network analyses. Our experiment with 64 Amazon EC2 8-core instances shows that SociaLite programs performed within a factor of two with respect to ideal weak scaling. Compared to optimized Giraph, an open-source alternative of Pregel, SociaLite programs are 4 to 12 times faster across benchmark algorithms, and 22 times more succinct on average. As a declarative query language, SociaLite, with the help of a compiler that generates efficient parallel and approximate code, can be used easily to create many social apps that operate on large-scale distributed graphs.
Jiwon Seo 0002, Jongsoo Park, Jaeho Shin 0001, Monica S. Lam
Proc. VLDB Endow.3
2005 Taming False Alarms from a Domain-Unaware C Analyzer by a Bayesian Statistical Post Analysis
Yungbum Jung, Jaehwang Kim, Jaeho Shin 0001, Kwangkeun Yi
SAS3