Joshua Martinez 0003

dblp:201/8024-3 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Services computing and microservices · 100%
Artificial intelligence
1 paper
Vision and language · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Services computing and microservices
business process management
0.812024
WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks · NeurIPS 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.212024
WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

multimodal foundation models · 0.8multimodal foundation model · 0.8
YearPublicationVenuePosition
2024 WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks
abstract
Existing ML benchmarks lack the depth and diversity of annotations needed for evaluating models on business process management (BPM) tasks. BPM is the practice of documenting, measuring, improving, and automating enterprise workflows. However, research has focused almost exclusively on one task -- full end-to-end automation using agents based on multimodal foundation models (FMs) like GPT-4. This focus on automation ignores the reality of how most BPM tools are applied today -- simply documenting the relevant workflow takes 60% of the time of the typical process optimization project. To address this gap we present WONDERBREAD, the first benchmark for evaluating multimodal FMs on BPM tasks beyond automation. Our contributions are: (1) a dataset containing 2928 documented workflow demonstrations; (2) 6 novel BPM tasks sourced from real-world applications ranging from workflow documentation to knowledge transfer to process improvement; and (3) an automated evaluation harness. Our benchmark shows that while state-of-the-art FMs can automatically generate documentation (e.g. recalling 88% of the steps taken in a video demonstration of a workflow), they struggle to re-apply that knowledge towards finer-grained validation of workflow completion (F1 < 0.3). We hope WONDERBREAD encourages the development of more "human-centered" AI tooling for enterprise applications and furthers the exploration of multimodal FMs for the broader universe of BPM tasks. We publish our dataset and experiments here: https://github.com/HazyResearch/wonderbread
Michael Wornow, Avanika Narayan, Ben Viggiano, Ishan S. Khare, Tathagat Verma, Tibor Thompson, Miguel Angel Fuentes Hernandez, Sudharsan Sundar, Chloe Trujillo, Krrish Chawla, Rongfei Lu, Justin Shen, Divya Nagaraj, Joshua Martinez 0003, Vardhan Agrawal, Althea Hudson, Nigam H. Shah, Christopher Ré
NeurIPS14