Julian Yuya Caspary

dblp:414/6570 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0001-0606-4115ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › LLM agents › web agents
LLM-based web agents
1.012026
WebMall - A Multi-Shop Benchmark for Evaluating Web Agents · SIGIR 2026
Information retrieval › evaluation
benchmark
1.012026
WebMall - A Multi-Shop Benchmark for Evaluating Web Agents · SIGIR 2026
Information retrieval
evaluation
1.012026
WebMall - A Multi-Shop Benchmark for Evaluating Web Agents · SIGIR 2026

Methods — techniques the papers use, named apart from their topics

large language model · 2.0
YearPublicationVenuePosition
2026 WebMall - A Multi-Shop Benchmark for Evaluating Web Agents
abstract
LLM-based web agents have the potential to automate long-running web tasks, such as searching for products in multiple e-shops and subsequently ordering the cheapest products that meet the user's needs. Benchmarks for evaluating web agents either require agents to perform tasks online using the live Web or offline using simulated environments, the latter allowing for the exact reproduction of the experimental setup. While DeepShop and ShoppingComp provide online benchmarks that require agents to perform challenging shopping tasks, existing offline benchmarks such as WebShop, WebArena, and Mind2Web cover only comparatively simple e-commerce tasks performed against a single shop containing product data from a single source. What is missing is an e-commerce benchmark that simulates multiple shops containing heterogeneous product data and requires agents to perform complex retrieval tasks. We fill this gap by introducing WebMall, the first offline multi-shop benchmark for evaluating web agents on challenging comparison shopping tasks. WebMall consists of four simulated shops populated with product data extracted from the Common Crawl. The WebMall tasks range from specific product searches and price comparisons to advanced searches for complementary or substitute products, as well as checkout processes. We validate WebMall using eight agents that differ in observation space, availability of short-term memory, and the employed LLM. The validation highlights the difficulty of the benchmark, with the best-performing agents achieving task completion rates below 65% in the task categories cheapest product search and vague product search.
Ralph Peeters, Aaron Steiner, Luca Schwarz, Julian Yuya Caspary, Christian Bizer
SIGIR4