VLDB 2026 Research / reviewers in the wild / expert
Joshua Martinez 0003
dblp:201/8024-3
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Services computing and microservices · 100% | |
| Artificial intelligence
1 paper |
Vision and language · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Services computing and microservices
business process management |
0.8 | 1 | 2024 | WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks · NeurIPS 2024 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.2 | 1 | 2024 | WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
multimodal foundation models · 0.8multimodal foundation model · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management TasksabstractExisting ML benchmarks lack the depth and diversity of annotations needed for evaluating models on business process management (BPM) tasks. BPM is the practice of documenting, measuring, improving, and automating enterprise workflows. However, research has focused almost exclusively on one task -- full end-to-end automation using agents based on multimodal foundation models (FMs) like GPT-4. This focus on automation ignores the reality of how most BPM tools are applied today -- simply documenting the relevant workflow takes 60% of the time of the typical process optimization project. To address this gap we present WONDERBREAD, the first benchmark for evaluating multimodal FMs on BPM tasks beyond automation. Our contributions are: (1) a dataset containing 2928 documented workflow demonstrations; (2) 6 novel BPM tasks sourced from real-world applications ranging from workflow documentation to knowledge transfer to process improvement; and (3) an automated evaluation harness. Our benchmark shows that while state-of-the-art FMs can automatically generate documentation (e.g. recalling 88% of the steps taken in a video demonstration of a workflow), they struggle to re-apply that knowledge towards finer-grained validation of workflow completion (F1 < 0.3). We hope WONDERBREAD encourages the development of more "human-centered" AI tooling for enterprise applications and furthers the exploration of multimodal FMs for the broader universe of BPM tasks. We publish our dataset and experiments here: https://github.com/HazyResearch/wonderbread Michael Wornow, Avanika Narayan, Ben Viggiano, Ishan S. Khare, Tathagat Verma, Tibor Thompson, Miguel Angel Fuentes Hernandez, Sudharsan Sundar, Chloe Trujillo, Krrish Chawla, Rongfei Lu, Justin Shen, Divya Nagaraj, Joshua Martinez 0003, Vardhan Agrawal, Althea Hudson, Nigam H. Shah, Christopher Ré |
NeurIPS | 14 |