Yitzhak Mandelbaum

dblp:09/3669 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
5 papers
Programming languages and type systems · 59% Compilers and program optimization · 41%
Databases, data mining, and information retrieval
1 paper
Data models and query languages · 61% Data integration and cleaning · 30% Query processing and optimization · 9%
Theoretical computer science
1 paper
Automata and formal languages · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
parsing
0.332010
The next 700 data description languages · J. ACM 2010
Semantics and algorithms for data-dependent grammars · POPL 2010
The next 700 data description languages · POPL 2006
Programming languages and type systems › domain-specific languages
data description languages
0.232010
The next 700 data description languages · J. ACM 2010
PADS/ML: a functional data description language · POPL 2007
The next 700 data description languages · POPL 2006
Compilers and program optimization › parsing
error-correct parsing
0.222010
The next 700 data description languages · J. ACM 2010
The next 700 data description languages · POPL 2006
Programming languages and type systems
language design
0.122010
The next 700 data description languages · J. ACM 2010
ESP: A Language for Programmable Devices · PLDI 2001
Programming languages and type systems › type theory
dependent types
0.112010
The next 700 data description languages · J. ACM 2010
Programming languages and type systems
type theory
0.112010
The next 700 data description languages · J. ACM 2010
Automata and formal languages › formal grammars
context-free grammar
0.112010
Semantics and algorithms for data-dependent grammars · POPL 2010
Programming languages and type systems
functional programming
0.112007
PADS/ML: a functional data description language · POPL 2007
Data integration and cleaning
ad hoc data processing
0.112006
PADS: an end-to-end system for processing ad hoc data · SIGMOD Conference 2006
Data models and query languages
data description language
0.112006
PADS: an end-to-end system for processing ad hoc data · SIGMOD Conference 2006
Reconfigurable computing and FPGAs
programmable devices
0.012001
ESP: A Language for Programmable Devices · PLDI 2001
Query processing and optimization
XML query processing
0.012006
PADS: an end-to-end system for processing ad hoc data · SIGMOD Conference 2006
Compilers and program optimization
compiler construction
0.012001
ESP: A Language for Programmable Devices · PLDI 2001

Methods — techniques the papers use, named apart from their topics

type-theoretic semantics · 0.1parser construction · 0.1type inference · 0.1parser generation · 0.1dependent types · 0.1
YearPublicationVenuePosition
2011 A New Method for Dependent Parsing
Trevor Jim, Yitzhak Mandelbaum
ESOP2
2010 Semantics and algorithms for data-dependent grammars
abstract
We present the design and theory of a new parsing engine, YAKKER, capable of satisfying the many needs of modern programmers and modern data processing applications. In particular, our new parsing engine handles (1) full scannerless context-free grammars with (2) regular expressions as right-hand sides for defining nonterminals. YAKKER also includes (3) facilities for binding variables to intermediate parse results and (4) using such bindings within arbitrary constraints to control parsing. These facilities allow the kind of data-dependent parsing commonly needed in systems applications, particularly those that operate over binary data. In addition, (5) nonterminals may be parameterized by arbitrary values, which gives the system good modularity and abstraction properties in the presence of data-dependent parsing. Finally, (6) legacy parsing libraries,such as sophisticated libraries for dates and times, may be directly incorporated into parser specifications. We illustrate the importance and utility of this rich collection of features by presenting its use on examples ranging from difficult programming language grammars to web server logs to binary data specification. We also show that our grammars have important compositionality properties and explain why such properties areimportant in modern applications such as automatic grammar induction.
Trevor Jim, Yitzhak Mandelbaum, David Walker 0001
POPL2
2010 The next 700 data description languages
abstract
In the spirit of Landin, we present a calculus of dependent types to serve as the semantic foundation for a family of languages called data description languages . Such languages, which include pads, datascript, and packettypes, are designed to facilitate programming with ad hoc data , that is, data not in well-behaved relational or xml formats. In the calculus, each type describes the physical layout and semantic properties of a data source. In the semantics, we interpret types simultaneously as the in-memory representation of the data described and as parsers for the data source. The parsing functions are robust, automatically detecting and recording errors in the data stream without halting parsing. We show the parsers are type-correct, returning data whose type matches the simple-type interpretation of the specification. We also prove the parsers are “error-correct,” accurately reporting the number of physical and semantic errors that occur in the returned data. We use the calculus to describe the features of various data description languages, and we discuss how we have used the calculus to improve pads.
Kathleen Fisher, Yitzhak Mandelbaum, David Walker 0001
J. ACM2
2009 Language support for processing distributed ad hoc data
abstract
This paper presents the design, theory and implementation of Gloves, a domain-specific language that allows users to specify the provenance (the derivation history starting from the origins), syntax and semantic properties of collections of distributed data sources. In particular, Gloves specifications indicate where to locate desired data, how to obtain it, when to get it or to give up trying, and what format it will be in on arrival. The Gloves system compiles such specification into a suite of data-processing tools including an archiver, a provenance tracking system, a database loading tool, an alert system, an RSS feed generator and a debugging tool. In addition, the system generates description-specific libraries so that developers can create their own applications. Gloves also provides a generic infrastructure so that advanced users can build new tools applicable to any data source with a Gloves description. We show how Gloves may be used to specify data sources from two domains: CoMon, a monitoring system for PlanetLab's 800+ nodes, and Arrakis, a monitoring system for an AT&T web hosting service. We show experimentally that our system can scale to distributed systems the size of CoMon. Finally, we provide a denotational semantics for Gloves and use this semantics to prove two important theorems. The first shows that our denotational semantics respects the typing rules for the language, while the second demonstrates that our system correctly maintains the provenance.
Kenny Q. Zhu, Daniel S. Dantas, Kathleen Fisher, Limin Jia 0001, Yitzhak Mandelbaum, Vivek S. Pai, David Walker 0001
PPDP5
2008 A Generic Programming Toolkit for PADS/ML: First-Class Upgrades for Third-Party Developers
Mary F. Fernández, Kathleen Fisher, Nate Foster, Michael Greenberg 0002, Yitzhak Mandelbaum
PADL5
2007 PADS/ML: a functional data description language
Yitzhak Mandelbaum, Kathleen Fisher, David Walker 0001, Mary F. Fernández, Artem Gleyzer
POPL1
2006 The next 700 data description languages
abstract
In the spirit of Landin, we present a calculus of dependent types to serve as the semantic foundation for a family of languages called data description languages. Such languages, which include pads, datascript, and packettypes, are designed to facilitate programming with ad hoc data, ie, data not in well-behaved relational or xml formats. In the calculus, each type describes the physical layout and semantic properties of a data source. In the semantics, we interpret types simultaneously as the in-memory representation of the data described and as parsers for the data source. The parsing functions are robust, automatically detecting and recording errors in the data stream without halting parsing. We show the parsers are type-correct, returning data whose type matches the simple-type interpretation of the specification. We also prove the parsers are "error-correct," accurately reporting the number of physical and semantic errors that occur in the returned data. We use the calculus to describe the features of various data description languages, and we discuss how we have used the calculus to improve PADS.
Kathleen Fisher, Yitzhak Mandelbaum, David Walker 0001
POPL2
2006 PADS: an end-to-end system for processing ad hoc data
abstract
Enormous amounts of data exist in "well-behaved" formats such as relational tables and XML, which come equipped with extensive tool support. However, vast amounts of data also exist in non-standard or ad hoc data formats, which often lack standard or extensible tools. This deficiency forces data analysts to implement their own tools for parsing, querying, and analyzing their ad hoc data. The resulting tools typically interleave parsing, querying, and analysis, obscuring the semantics of the data format and making it nearly impossible for others to resuse the tools. This proposal describes PADS, an end-to-end system for processing ad hoc data sources. The core of PADS is a declarative language for describing ad hoc data sources and a data-description compiler that produces customizable libraries for parsing the ad hoc data. A suite of tools built around this core includes statistical data-profiling tools, a query engine that permits viewing ad hoc sources as XML and for querying them with XQuery, and an interactive front-end that helps users produce PADS descriptions quickly.
Mark Daly, Yitzhak Mandelbaum, David Walker 0001, Mary F. Fernández, Kathleen Fisher, Robert Gruber
SIGMOD Conference2
2003 An effective theory of type refinements
abstract
We develop an explicit two level system that allows programmers to reason about the behavior of effectful programs. The first level is an ordinary ML-style type system, which confers standard properties on program behavior. The second level is a conservative extension of the first that uses a logic of type refinements to check more precise properties of program behavior. Our logic is a fragment of intuitionistic linear logic, which gives programmers the ability to reason locally about changes of program state. We provide a generic resource semantics for our logic as well as a sound, decidable, syntactic refinement-checking system. We also prove that refinements give rise to an optimization principle for programs. Finally, we illustrate the power of our system through a number of examples.
Yitzhak Mandelbaum, David Walker 0001, Robert Harper 0001
ICFP1
2001 ESP: A Language for Programmable Devices
abstract
This paper presents the design and implementation of Event-driven State-machines Programming (ESP)—a language for programmable devices. In traditional languages, like C, using event-driven state-machine forces a tradeoff that requires giving up ease of development and reliability to achieve high performance. ESP is designed to provide all of these three properties simultaneously.
Yitzhak Mandelbaum, Kai Li 0001
PLDI2