EDBT 2026 Demo / reviewers in the wild / expert
Jaeho Shin 0001
dblp:38/1325-1
· DBLP profile ↗
7ranked-venue papers
2as first author
0since 2021 · last 2017
0000-0001-5280-3356ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Knowledge graphs · 35% Graph data management · 27% Machine learning and data management · 21% | |
| Software engineering, system software, and programming languages
3 papers |
Debugging and program repair · 36% Empirical software engineering · 36% Programming languages and type systems · 28% | |
| Artificial intelligence
2 papers |
Knowledge representation and reasoning · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge graphs
knowledge graph construction |
0.7 | 3 | 2017 | Incremental knowledge base construction using DeepDive · VLDB J. 2017 Incremental Knowledge Base Construction Using DeepDive · Proc. VLDB Endow. 2015 Mindtagger: A Demonstration of Data Labeling in Knowledge Base Construction · Proc. VLDB Endow. 2015 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base construction |
0.5 | 2 | 2017 | Incremental knowledge base construction using DeepDive · VLDB J. 2017 Extracting Databases from Dark Data with DeepDive · SIGMOD Conference 2016 |
Machine learning and data management
data annotation |
0.2 | 1 | 2015 | Mindtagger: A Demonstration of Data Labeling in Knowledge Base Construction · Proc. VLDB Endow. 2015 |
Graph data management › graph processing
graph processing systems |
0.2 | 1 | 2015 | Graft: A Debugging Tool For Apache Giraph · SIGMOD Conference 2015 |
Machine learning and data management
inference optimization |
0.2 | 1 | 2015 | Incremental Knowledge Base Construction Using DeepDive · Proc. VLDB Endow. 2015 |
Debugging and program repair › concurrent program debugging
distributed debugging |
0.2 | 1 | 2015 | Graft: A Debugging Tool For Apache Giraph · SIGMOD Conference 2015 |
Graph data management › graph analytics
distributed graph analysis |
0.2 | 1 | 2013 | Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis · Proc. VLDB Endow. 2013 |
Graph data management
graph analytics |
0.2 | 1 | 2013 | Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis · Proc. VLDB Endow. 2013 |
Programming languages and type systems › logic programming
datalog |
0.2 | 1 | 2013 | Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis · Proc. VLDB Endow. 2013 |
Parallel and multicore computing › parallel computing
distributed execution |
0.0 | 1 | 2013 | Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis · Proc. VLDB Endow. 2013 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2013 | Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis · Proc. VLDB Endow. 2013 |
Methods — techniques the papers use, named apart from their topics
statistical inference · 1.0distant supervision · 0.6probabilistic inference · 0.5approximate computation · 0.5delta stepping · 0.3variational inference · 0.2sampling · 0.2rule-based optimizer · 0.2delta-stepping · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Incremental knowledge base construction using DeepDive
Christopher De Sa, Alexander Ratner, Christopher Ré, Jaeho Shin 0001, Sen Wu 0002, Ce Zhang 0001 |
VLDB J. | 4 |
| 2016 | Extracting Databases from Dark Data with DeepDiveabstract: the mass of text, tables, and images that are widely collected and stored but which cannot be exploited by standard relational tools. If the information in dark data - scientific papers, Web classified ads, customer service notes, and so on - were instead in a relational database, it would give analysts a massive and valuable new set of "big data." DeepDive is distinctive when compared to previous information extraction systems in its ability to obtain very high precision and recall at reasonable engineering cost; in a number of applications, we have used DeepDive to create databases with accuracy that meets that of human annotators. To date we have successfully deployed DeepDive to create data-centric applications for insurance, materials science, genomics, paleontologists, law enforcement, and others. The data unlocked by DeepDive represents a massive opportunity for industry, government, and scientific researchers. DeepDive is enabled by an unusual design that combines large-scale probabilistic inference with a novel developer interaction cycle. This design is enabled by several core innovations around probabilistic training and inference. Ce Zhang 0001, Jaeho Shin 0001, Christopher Ré, Michael J. Cafarella, Feng Niu |
SIGMOD Conference | 2 |
| 2015 | Graft: A Debugging Tool For Apache GiraphabstractWe address the problem of debugging programs written for Pregel-like systems. After interviewing Giraph and GPS users, we developed Graft. Graft supports the debugging cycle that users typically go through: (1) Users describe programmatically the set of vertices they are interested in inspecting. During execution, Graft captures the context information of these vertices across supersteps. (2) Using Graft's GUI, users visualize how the values and messages of the captured vertices change from superstep to superstep,narrowing in suspicious vertices and supersteps. (3) Users replay the exact lines of the code vertex.compute() function that executed for the suspicious vertices and supersteps, by copying code that Graft generates into their development environments' line-by-line debuggers. Graft also has features to construct end-to-end tests for Giraph programs. Graft is open-source and fully integrated into Apache Giraph's main code base. Semih Salihoglu, Jaeho Shin 0001, Vikesh Khanna, Ba Quan Truong, Jennifer Widom |
SIGMOD Conference | 2 |
| 2015 | Mindtagger: A Demonstration of Data Labeling in Knowledge Base ConstructionabstractEnd-to-end knowledge base construction systems using statistical inference are enabling more people to automatically extract high-quality domain-specific information from unstructured data. As a result of deploying DeepDive framework across several domains, we found new challenges in debugging and improving such end-to-end systems to construct high-quality knowledge bases. DeepDive has an iterative development cycle in which users improve the data. To help our users, we needed to develop principles for analyzing the system's error as well as provide tooling for inspecting and labeling various data products of the system. We created guidelines for error analysis modeled after our colleagues' best practices, in which data labeling plays a critical role in every step of the analysis. To enable more productive and systematic data labeling, we created Mindtagger, a versatile tool that can be configured to support a wide range of tasks. In this demonstration, we show in detail what data labeling tasks are modeled in our error analysis guidelines and how each of them is performed using Mindtagger. Jaeho Shin 0001, Christopher Ré, Michael J. Cafarella |
Proc. VLDB Endow. | 1 |
| 2015 | Incremental Knowledge Base Construction Using DeepDiveabstractPopulating a database with unstructured information is a long-standing problem in industry and research that encompasses problems of extraction, cleaning, and integration. Recent names used for this problem include dealing with dark data and knowledge base construction (KBC). In this work, we describe DeepDive, a system that combines database and machine learning ideas to help develop KBC systems, and we present techniques to make the KBC process more efficient. We observe that the KBC process is iterative, and we develop techniques to incrementally produce inference results for KBC systems. We propose two methods for incremental inference, based respectively on sampling and variational techniques. We also study the tradeoff space of these methods and develop a simple rule-based optimizer. DeepDive includes all of these contributions, and we evaluate Deep-Dive on five KBC systems, showing that it can speed up KBC inference tasks by up to two orders of magnitude with negligible impact on quality. Jaeho Shin 0001, Sen Wu 0002, Christopher De Sa, Ce Zhang 0001, Christopher Ré |
Proc. VLDB Endow. | 1 |
| 2013 | Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph AnalysisabstractLarge-scale graph analysis is becoming important with the rise of world-wide social network services. Recently in SociaLite, we proposed extensions to Datalog to efficiently and succinctly implement graph analysis programs on sequential machines. This paper describes novel extensions and optimizations of SociaLite for parallel and distributed executions to support large-scale graph analysis. With distributed SociaLite, programmers simply annotate how data are to be distributed, then the necessary communication is automatically inferred to generate parallel code for cluster of multi-core machines. It optimizes the evaluation of recursive monotone aggregate functions using a delta stepping technique. In addition, approximate computation is supported in SociaLite, allowing programmers to trade off accuracy for less time and space. We evaluated SociaLite with six core graph algorithms used in many social network analyses. Our experiment with 64 Amazon EC2 8-core instances shows that SociaLite programs performed within a factor of two with respect to ideal weak scaling. Compared to optimized Giraph, an open-source alternative of Pregel, SociaLite programs are 4 to 12 times faster across benchmark algorithms, and 22 times more succinct on average. As a declarative query language, SociaLite, with the help of a compiler that generates efficient parallel and approximate code, can be used easily to create many social apps that operate on large-scale distributed graphs. Jiwon Seo 0002, Jongsoo Park, Jaeho Shin 0001, Monica S. Lam |
Proc. VLDB Endow. | 3 |
| 2005 | Taming False Alarms from a Domain-Unaware C Analyzer by a Bayesian Statistical Post Analysis
Yungbum Jung, Jaehwang Kim, Jaeho Shin 0001, Kwangkeun Yi |
SAS | 3 |