Adrian Blue

dblp:328/6881 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Artificial intelligence
1 paper
Information extraction and text analysis · 100%
Human-computer interaction and pervasive computing
1 paper
User interface design and tools · 100%

Topics — the 2 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
DSL-based synthesis
0.612022
Blueprint: a constraint-solving approach for document extraction · Proc. VLDB Endow. 2022
Program synthesis and code generation
programming by example
0.612022
Blueprint: a constraint-solving approach for document extraction · Proc. VLDB Endow. 2022

Methods — techniques the papers use, named apart from their topics

fuzzy constraints · 1.7constraint solving · 1.7
YearPublicationVenuePosition
2022 Blueprint: a constraint-solving approach for document extraction
abstract
Blueprint is a declarative domain-specific language for document extraction. Users describe document layout using spatial, textual, semantic, and numerical fuzzy constraints, and the language runtime extracts the field-value mappings that best satisfy the constraints in a given document. We used Blueprint to develop several document extraction solutions in a commercial setting. This approach to the extraction problem proved powerful. Concise Blueprint programs were able to generate good accuracy on a broad set of use cases. However, a major goal of our work was to build a system that non-experts, and in particular non-engineers, could use effectively, and we found that writing declarative fuzzy constraint-based extraction programs was not intuitive for many users: a large up-front learning investment was required to be effective, and debugging was often challenging. To address these issues, we developed a no-code IDE for Blueprint, called Studio, as well as program synthesis functionality for automatically generating Blueprint programs from training data, which could be created by labeling document samples in our IDE. Overall, the IDE significantly improved the Blueprint development experience and the results users were able to achieve. In this paper, we discuss the design, implementation, and deployment of Blueprint and Studio. We compare our system with a state-of-the-art deep-learning based extraction tool and show that our system can achieve comparable accuracy results, with comparable development time, for appropriately-chosen use cases, while providing better interpretability and debuggability.
Andrey Mishchenko, Dominique Danco, Abhilash Jindal, Adrian Blue
Proc. VLDB Endow.4