VLDB 2026 Research / reviewers in the wild / expert
Julian Yuya Caspary
dblp:414/6570
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0001-0606-4115ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Artificial intelligence
1 paper |
Language models and text generation · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › LLM agents › web agents
LLM-based web agents |
1.0 | 1 | 2026 | WebMall - A Multi-Shop Benchmark for Evaluating Web Agents · SIGIR 2026 |
Information retrieval › evaluation
benchmark |
1.0 | 1 | 2026 | WebMall - A Multi-Shop Benchmark for Evaluating Web Agents · SIGIR 2026 |
Information retrieval
evaluation |
1.0 | 1 | 2026 | WebMall - A Multi-Shop Benchmark for Evaluating Web Agents · SIGIR 2026 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WebMall - A Multi-Shop Benchmark for Evaluating Web AgentsabstractLLM-based web agents have the potential to automate long-running web tasks, such as searching for products in multiple e-shops and subsequently ordering the cheapest products that meet the user's needs. Benchmarks for evaluating web agents either require agents to perform tasks online using the live Web or offline using simulated environments, the latter allowing for the exact reproduction of the experimental setup. While DeepShop and ShoppingComp provide online benchmarks that require agents to perform challenging shopping tasks, existing offline benchmarks such as WebShop, WebArena, and Mind2Web cover only comparatively simple e-commerce tasks performed against a single shop containing product data from a single source. What is missing is an e-commerce benchmark that simulates multiple shops containing heterogeneous product data and requires agents to perform complex retrieval tasks. We fill this gap by introducing WebMall, the first offline multi-shop benchmark for evaluating web agents on challenging comparison shopping tasks. WebMall consists of four simulated shops populated with product data extracted from the Common Crawl. The WebMall tasks range from specific product searches and price comparisons to advanced searches for complementary or substitute products, as well as checkout processes. We validate WebMall using eight agents that differ in observation space, availability of short-term memory, and the employed LLM. The validation highlights the difficulty of the benchmark, with the best-performing agents achieving task completion rates below 65% in the task categories cheapest product search and vague product search. Ralph Peeters, Aaron Steiner, Luca Schwarz, Julian Yuya Caspary, Christian Bizer |
SIGIR | 4 |