Bo Wu 0008

dblp:47/6534-8 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 100%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
programming by example
0.212015
An Iterative Approach to Synthesize Data Transformation Programs · IJCAI 2015
Program synthesis and code generation
programming by demonstration
0.112012
Learning Transformation Rules by Examples · AAAI 2012
Data integration and cleaning
data transformation
0.012012
Learning Transformation Rules by Examples · AAAI 2012

Methods — techniques the papers use, named apart from their topics

search algorithm · 0.3grammar space reduction · 0.3
YearPublicationVenuePosition
2016 Maximizing Correctness with Minimal User Effort to Learn Data Transformations
abstract
Data transformation often requires users to write many trivial and task-dependent programs to transform thousands of records. Recently, programming-by-example (PBE) approaches enable users to transform data without coding. A key challenge of these PBE approaches is to deliver correctly transformed results on large datasets, since these transformation programs are likely to be generated by non-expert users. To address this challenge, existing approaches aim to identify a small set of potentially incorrect records and ask users to examine these records instead of the entire dataset. However, because the transformation scenarios are highly task-dependent, existing approaches cannot capture the incorrect records for various scenarios. We present a approach that learns from past transformation scenarios to generate a meta-classifier to identify the incorrect records. Our approach color-codes these transformed records and then presents them for users to examine. The method allows users to either enter an example for a record transformed incorrectly or confirm the correctness of a transformed record. And our approach can learn from the users' labels to refine the meta-classifier to accurately identify the incorrect records. Simulation results and a user study show that our method can identify the incorrectly transformed records and reduce the user efforts in examining the results.
Bo Wu 0008, Craig A. Knoblock
IUI1
2015 An Iterative Approach to Synthesize Data Transformation Programs
Bo Wu 0008, Craig A. Knoblock
IJCAI1
2014 A system for efficient cleaning and transformation of geospatial data attributes
abstract
A significant challenge in handling geographic datasets is that the datasets can come from heterogeneous sources with various data qualities and formats. Before these datasets can be used in a Geographic Information System (GIS) for spatial analysis or to create maps, a typical task is to clean the attribute data and transform the data into a uniform format. However, conventional GIS products focus on manipulating the spatial component of geographic features and only offer basic tools for editing the attribute data (e.g., one row at a time). This limits the capability for handling large datasets in a GIS since manually editing and transforming attribute data between different formats is not practical for thousands of geographic features. In this demo, we present ArcKarma, which is built on our previous work on data transformation, to efficiently clean and transform data attributes in a GIS. ArcKarma generates transformation programs from a few user-provided examples and applies these programs to transform individual attribute columns into the desired formats. We show that ArcKarma produces accurate results and eliminates the need for laborious manual data cleaning and scripting tasks.
Yao-Yi Chiang, Bo Wu 0008, Akshay Anand, Ketan Akade, Craig A. Knoblock
SIGSPATIAL/GIS2
2014 Minimizing user effort in transforming data by example
abstract
Programming by example enables users to transform data formats without coding. To be practical, the method must synthesize the correct transformation with minimal user input. We present a method that minimizes user effort by color-coding the transformation result and recommending specific records where the user should provide examples. Simulation results and a user study show that our method significantly reduces user effort and increases the success rate for synthesizing correct transformation programs by example.
Bo Wu 0008, Pedro A. Szekely, Craig A. Knoblock
IUI1
2012 Learning Transformation Rules by Examples
abstract
This paper presents an abstract for a general data transformation approach. Using programming by demonstration technique, we learn the transformation rules through user given examples. These transformation rules are automatically generated from a predefined grammar. Due to the grammar space is huge, we propose a grammar space reduction method to reduce the search space and a sketch of search algorithm is adopted to identify the rules that are consistent with the examples. The final experimental results show our approach achieves promising results on different transformation scenarios.
Bo Wu 0008, Pedro A. Szekely, Craig A. Knoblock
AAAI1