VLDB 2026 Research / reviewers in the wild / expert
Shion Sakurai
dblp:440/7419
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0001-4828-8555ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › evaluation
benchmark |
1.0 | 1 | 2026 | Preliminary Study of an Evaluation Benchmark for Vision-Language Models in Fashion E-Commerce · SIGIR 2026 |
Information retrieval
e-commerce search |
1.0 | 1 | 2026 | Preliminary Study of an Evaluation Benchmark for Vision-Language Models in Fashion E-Commerce · SIGIR 2026 |
Information retrieval
evaluation |
1.0 | 1 | 2026 | Preliminary Study of an Evaluation Benchmark for Vision-Language Models in Fashion E-Commerce · SIGIR 2026 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Preliminary Study of an Evaluation Benchmark for Vision-Language Models in Fashion E-CommerceabstractWe report an evaluation benchmark for assessing the operational suitability of Vision-Language Models (VLMs) in fashion e-commerce. General-purpose benchmarks do not adequately cover fashion-specific attributes or the structured extraction tasks common in e-commerce workflows. We define five tasks across two image streams---outfit and single-item product images---and compare six commercial and two open-source models with multiple prompt variants, including a canonical prompt and model-proposed prompts. Experiments show that the best-performing model varies by task, error patterns are more model-dependent than prompt-dependent, and model updates can improve some tasks while degrading others. These results indicate that task-specific evaluation, prompt robustness checks, and continuous monitoring are practical requirements for deploying VLMs in production fashion systems. Ryotaro Shimizu, Sai Htaung Kham, Shion Sakurai |
SIGIR | 3 |