Sai Htaung Kham

dblp:390/3397 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0005-8803-5842ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 79% Recommender systems · 21%
Artificial intelligence
1 paper
Knowledge representation and reasoning · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › evaluation
benchmark
1.012026
Preliminary Study of an Evaluation Benchmark for Vision-Language Models in Fashion E-Commerce · SIGIR 2026
Information retrieval
e-commerce search
1.012026
Preliminary Study of an Evaluation Benchmark for Vision-Language Models in Fashion E-Commerce · SIGIR 2026
Information retrieval
evaluation
1.012026
Preliminary Study of an Evaluation Benchmark for Vision-Language Models in Fashion E-Commerce · SIGIR 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
explanation generation
0.912025
Disentangling Likes and Dislikes in Personalized Generative Explainable Recommendation · WWW 2025
Recommender systems
explainable recommendation
0.912025
Disentangling Likes and Dislikes in Personalized Generative Explainable Recommendation · WWW 2025

Methods — techniques the papers use, named apart from their topics

text generation · 1.7LLM-based opinion extraction · 1.7vision-language model · 1.0
YearPublicationVenuePosition
2026 Preliminary Study of an Evaluation Benchmark for Vision-Language Models in Fashion E-Commerce
abstract
We report an evaluation benchmark for assessing the operational suitability of Vision-Language Models (VLMs) in fashion e-commerce. General-purpose benchmarks do not adequately cover fashion-specific attributes or the structured extraction tasks common in e-commerce workflows. We define five tasks across two image streams---outfit and single-item product images---and compare six commercial and two open-source models with multiple prompt variants, including a canonical prompt and model-proposed prompts. Experiments show that the best-performing model varies by task, error patterns are more model-dependent than prompt-dependent, and model updates can improve some tasks while degrading others. These results indicate that task-specific evaluation, prompt robustness checks, and continuous monitoring are practical requirements for deploying VLMs in production fashion systems.
Ryotaro Shimizu, Sai Htaung Kham, Shion Sakurai
SIGIR2
2025 Disentangling Likes and Dislikes in Personalized Generative Explainable Recommendation
abstract
Recent research on explainable recommendation generally frames the task as a standard text generation problem, and evaluates models simply based on the textual similarity between the predicted and ground-truth explanations. However, this approach fails to consider one crucial aspect of the systems: whether their outputs accurately reflect the users' (post-purchase) sentiments, i.e., whether and why they would like and/or dislike the recommended items. To shed light on this issue, we introduce new datasets and evaluation methods that focus on the users' sentiments. Specifically, we construct the datasets by explicitly extracting users' positive and negative opinions from their post-purchase reviews using an LLM, and propose to evaluate systems based on whether the generated explanations 1) align well with the users' sentiments, and 2) accurately identify both positive and negative opinions of users on the target items. We benchmark several recent models on our datasets and demonstrate that achieving strong performance on existing metrics does not ensure that the generated explanations align well with the users' sentiments. Lastly, we find that existing models can provide more sentiment-aware explanations when the users' (predicted) ratings for the target items are directly fed into the models as input. The datasets and benchmark implementation are available at: https://github.com/jchanxtarov/sent_xrec.
Ryotaro Shimizu, Takashi Wada 0001, Yu Wang 0170, Johannes Kruse 0002, Sean O'Brien, Sai Htaung Kham, Linxin Song, Yuya Yoshikawa, Yuki Saito 0002, Fugee Tsung, Masayuki Goto, Julian J. McAuley
WWW6