VLDB 2026 Research / reviewers in the wild / expert
Yitzhak Mandelbaum
dblp:09/3669
· DBLP profile ↗
10ranked-venue papers
2as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
5 papers |
Programming languages and type systems · 59% Compilers and program optimization · 41% | |
| Databases, data mining, and information retrieval
1 paper |
Data models and query languages · 61% Data integration and cleaning · 30% Query processing and optimization · 9% | |
| Theoretical computer science
1 paper |
Automata and formal languages · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
parsing |
0.3 | 3 | 2010 | The next 700 data description languages · J. ACM 2010 Semantics and algorithms for data-dependent grammars · POPL 2010 The next 700 data description languages · POPL 2006 |
Programming languages and type systems › domain-specific languages
data description languages |
0.2 | 3 | 2010 | The next 700 data description languages · J. ACM 2010 PADS/ML: a functional data description language · POPL 2007 The next 700 data description languages · POPL 2006 |
Compilers and program optimization › parsing
error-correct parsing |
0.2 | 2 | 2010 | The next 700 data description languages · J. ACM 2010 The next 700 data description languages · POPL 2006 |
Programming languages and type systems
language design |
0.1 | 2 | 2010 | The next 700 data description languages · J. ACM 2010 ESP: A Language for Programmable Devices · PLDI 2001 |
Programming languages and type systems › type theory
dependent types |
0.1 | 1 | 2010 | The next 700 data description languages · J. ACM 2010 |
Programming languages and type systems
type theory |
0.1 | 1 | 2010 | The next 700 data description languages · J. ACM 2010 |
Automata and formal languages › formal grammars
context-free grammar |
0.1 | 1 | 2010 | Semantics and algorithms for data-dependent grammars · POPL 2010 |
Programming languages and type systems
functional programming |
0.1 | 1 | 2007 | PADS/ML: a functional data description language · POPL 2007 |
Data integration and cleaning
ad hoc data processing |
0.1 | 1 | 2006 | PADS: an end-to-end system for processing ad hoc data · SIGMOD Conference 2006 |
Data models and query languages
data description language |
0.1 | 1 | 2006 | PADS: an end-to-end system for processing ad hoc data · SIGMOD Conference 2006 |
Reconfigurable computing and FPGAs
programmable devices |
0.0 | 1 | 2001 | ESP: A Language for Programmable Devices · PLDI 2001 |
Query processing and optimization
XML query processing |
0.0 | 1 | 2006 | PADS: an end-to-end system for processing ad hoc data · SIGMOD Conference 2006 |
Compilers and program optimization
compiler construction |
0.0 | 1 | 2001 | ESP: A Language for Programmable Devices · PLDI 2001 |
Methods — techniques the papers use, named apart from their topics
type-theoretic semantics · 0.1parser construction · 0.1type inference · 0.1parser generation · 0.1dependent types · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | A New Method for Dependent Parsing
Trevor Jim, Yitzhak Mandelbaum |
ESOP | 2 |
| 2010 | Semantics and algorithms for data-dependent grammarsabstractWe present the design and theory of a new parsing engine, YAKKER, capable of satisfying the many needs of modern programmers and modern data processing applications. In particular, our new parsing engine handles (1) full scannerless context-free grammars with (2) regular expressions as right-hand sides for defining nonterminals. YAKKER also includes (3) facilities for binding variables to intermediate parse results and (4) using such bindings within arbitrary constraints to control parsing. These facilities allow the kind of data-dependent parsing commonly needed in systems applications, particularly those that operate over binary data. In addition, (5) nonterminals may be parameterized by arbitrary values, which gives the system good modularity and abstraction properties in the presence of data-dependent parsing. Finally, (6) legacy parsing libraries,such as sophisticated libraries for dates and times, may be directly incorporated into parser specifications. We illustrate the importance and utility of this rich collection of features by presenting its use on examples ranging from difficult programming language grammars to web server logs to binary data specification. We also show that our grammars have important compositionality properties and explain why such properties areimportant in modern applications such as automatic grammar induction. Trevor Jim, Yitzhak Mandelbaum, David Walker 0001 |
POPL | 2 |
| 2010 | The next 700 data description languagesabstractIn the spirit of Landin, we present a calculus of dependent types to serve as the semantic foundation for a family of languages called data description languages . Such languages, which include pads, datascript, and packettypes, are designed to facilitate programming with ad hoc data , that is, data not in well-behaved relational or xml formats. In the calculus, each type describes the physical layout and semantic properties of a data source. In the semantics, we interpret types simultaneously as the in-memory representation of the data described and as parsers for the data source. The parsing functions are robust, automatically detecting and recording errors in the data stream without halting parsing. We show the parsers are type-correct, returning data whose type matches the simple-type interpretation of the specification. We also prove the parsers are “error-correct,” accurately reporting the number of physical and semantic errors that occur in the returned data. We use the calculus to describe the features of various data description languages, and we discuss how we have used the calculus to improve pads. Kathleen Fisher, Yitzhak Mandelbaum, David Walker 0001 |
J. ACM | 2 |
| 2009 | Language support for processing distributed ad hoc dataabstractThis paper presents the design, theory and implementation of Gloves, a domain-specific language that allows users to specify the provenance (the derivation history starting from the origins), syntax and semantic properties of collections of distributed data sources. In particular, Gloves specifications indicate where to locate desired data, how to obtain it, when to get it or to give up trying, and what format it will be in on arrival. The Gloves system compiles such specification into a suite of data-processing tools including an archiver, a provenance tracking system, a database loading tool, an alert system, an RSS feed generator and a debugging tool. In addition, the system generates description-specific libraries so that developers can create their own applications. Gloves also provides a generic infrastructure so that advanced users can build new tools applicable to any data source with a Gloves description. We show how Gloves may be used to specify data sources from two domains: CoMon, a monitoring system for PlanetLab's 800+ nodes, and Arrakis, a monitoring system for an AT&T web hosting service. We show experimentally that our system can scale to distributed systems the size of CoMon. Finally, we provide a denotational semantics for Gloves and use this semantics to prove two important theorems. The first shows that our denotational semantics respects the typing rules for the language, while the second demonstrates that our system correctly maintains the provenance. Kenny Q. Zhu, Daniel S. Dantas, Kathleen Fisher, Limin Jia 0001, Yitzhak Mandelbaum, Vivek S. Pai, David Walker 0001 |
PPDP | 5 |
| 2008 | A Generic Programming Toolkit for PADS/ML: First-Class Upgrades for Third-Party Developers
Mary F. Fernández, Kathleen Fisher, Nate Foster, Michael Greenberg 0002, Yitzhak Mandelbaum |
PADL | 5 |
| 2007 | PADS/ML: a functional data description language
Yitzhak Mandelbaum, Kathleen Fisher, David Walker 0001, Mary F. Fernández, Artem Gleyzer |
POPL | 1 |
| 2006 | The next 700 data description languagesabstractIn the spirit of Landin, we present a calculus of dependent types to serve as the semantic foundation for a family of languages called data description languages. Such languages, which include pads, datascript, and packettypes, are designed to facilitate programming with ad hoc data, ie, data not in well-behaved relational or xml formats. In the calculus, each type describes the physical layout and semantic properties of a data source. In the semantics, we interpret types simultaneously as the in-memory representation of the data described and as parsers for the data source. The parsing functions are robust, automatically detecting and recording errors in the data stream without halting parsing. We show the parsers are type-correct, returning data whose type matches the simple-type interpretation of the specification. We also prove the parsers are "error-correct," accurately reporting the number of physical and semantic errors that occur in the returned data. We use the calculus to describe the features of various data description languages, and we discuss how we have used the calculus to improve PADS. Kathleen Fisher, Yitzhak Mandelbaum, David Walker 0001 |
POPL | 2 |
| 2006 | PADS: an end-to-end system for processing ad hoc dataabstractEnormous amounts of data exist in "well-behaved" formats such as relational tables and XML, which come equipped with extensive tool support. However, vast amounts of data also exist in non-standard or ad hoc data formats, which often lack standard or extensible tools. This deficiency forces data analysts to implement their own tools for parsing, querying, and analyzing their ad hoc data. The resulting tools typically interleave parsing, querying, and analysis, obscuring the semantics of the data format and making it nearly impossible for others to resuse the tools. This proposal describes PADS, an end-to-end system for processing ad hoc data sources. The core of PADS is a declarative language for describing ad hoc data sources and a data-description compiler that produces customizable libraries for parsing the ad hoc data. A suite of tools built around this core includes statistical data-profiling tools, a query engine that permits viewing ad hoc sources as XML and for querying them with XQuery, and an interactive front-end that helps users produce PADS descriptions quickly. Mark Daly, Yitzhak Mandelbaum, David Walker 0001, Mary F. Fernández, Kathleen Fisher, Robert Gruber |
SIGMOD Conference | 2 |
| 2003 | An effective theory of type refinementsabstractWe develop an explicit two level system that allows programmers to reason about the behavior of effectful programs. The first level is an ordinary ML-style type system, which confers standard properties on program behavior. The second level is a conservative extension of the first that uses a logic of type refinements to check more precise properties of program behavior. Our logic is a fragment of intuitionistic linear logic, which gives programmers the ability to reason locally about changes of program state. We provide a generic resource semantics for our logic as well as a sound, decidable, syntactic refinement-checking system. We also prove that refinements give rise to an optimization principle for programs. Finally, we illustrate the power of our system through a number of examples. Yitzhak Mandelbaum, David Walker 0001, Robert Harper 0001 |
ICFP | 1 |
| 2001 | ESP: A Language for Programmable DevicesabstractThis paper presents the design and implementation of Event-driven State-machines Programming (ESP)—a language for programmable devices. In traditional languages, like C, using event-driven state-machine forces a tradeoff that requires giving up ease of development and reliability to achieve high performance. ESP is designed to provide all of these three properties simultaneously. Yitzhak Mandelbaum, Kai Li 0001 |
PLDI | 2 |