Ryotaro Shimizu

dblp:315/2865 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0002-4841-1824ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ZoRRO: A Zero-Weight Personalized Recommender System for Scalable News Recommendation
abstract
In this paper, we present ZoRRO (Zero-Weight Personalized Recommender System), a zero-weight and training-free framework for personalized news recommendation designed for scalable real-world deployment. We show that ZoRRO outperforms strong neural baselines in offline ranking evaluations and delivers click-through rate performance in online A/B testing that is nearly on par with a state-of-the-art deep learning model, while operating more than (600x) faster. Our experiments reveal gaps between offline and online performance, and show that models with similar click-through rate (CTR) outcomes can produce markedly different recommendation distributions, influencing the overall news flow. These findings position ZoRRO as a practical and efficient solution for large-scale news recommendation and highlight the importance of evaluating recommender systems using metrics beyond accuracy alone. Our code is available at https://github.com/johanneskruse/zorro.
Johannes Kruse 0002, Ryotaro Shimizu, Kasper Lindskow, Jon Tofteskov, Michael Riis Andersen, Julian J. McAuley, Jes Frellsen
SIGIR2
2026 Preliminary Study of an Evaluation Benchmark for Vision-Language Models in Fashion E-Commerce
abstract
We report an evaluation benchmark for assessing the operational suitability of Vision-Language Models (VLMs) in fashion e-commerce. General-purpose benchmarks do not adequately cover fashion-specific attributes or the structured extraction tasks common in e-commerce workflows. We define five tasks across two image streams---outfit and single-item product images---and compare six commercial and two open-source models with multiple prompt variants, including a canonical prompt and model-proposed prompts. Experiments show that the best-performing model varies by task, error patterns are more model-dependent than prompt-dependent, and model updates can improve some tasks while degrading others. These results indicate that task-specific evaluation, prompt robustness checks, and continuous monitoring are practical requirements for deploying VLMs in production fashion systems.
Ryotaro Shimizu, Sai Htaung Kham, Shion Sakurai
SIGIR1
2026 Optimizing pre-training for multi-label classification via generalized target-aware source data selection
abstract
While pre-trained models, such as large language models, can achieve high performance with minimal fine-tuning, the source datasets used for pre-training often contain irrelevant or blackundant data, which can degrade performance on target tasks. Domain Adaptation Information Gain (DAIG)-based source data selection improves performance by pre-training on source data selected based on rough prior knowledge obtained from target data in advance. However, DAIG’s key component, the transition matrix, lacks flexibility and is limited to handling only single-label classification tasks. To address this limitation, we propose the Generalized DAIG (GDAIG)-guided selection process, a novel framework that extends DAIG to support multi-label classification. GDAIG introduces a soft transition matrix to capture inter-label dependencies and employs binary cross-entropy loss to enable adaptation to multi-label data. By leveraging “rough prior knowledge” from initial training on target data, GDAIG actively selects informative and task-relevant source data for pre-training. Experiments on medical image and general object classification datasets demonstrate that GDAIG consistently outperforms baseline approaches, with particularly significant improvements in scenarios involving label mismatch between source and target domains (partial or no label overlap), where conventional transfer learning methods suffer from noise caused by irrelevant source labels. These results highlight GDAIG’s ability to enhance the effectiveness of pre-trained models through strategic source data selection, thereby optimizing performance for specific target tasks. Our framework goes beyond existing approaches that rely solely on pre-trained models, emphasizing the direct utilization of task-relevant source data. Furthermore, GDAIG provides a practical and effective solution for domains with scarce labeled data, such as medical image analysis. • A GDAIG-guided data selection strategy for multi-label classification is proposed. • GDAIG improves target model performance through task-relevant multi-label data selection. • A probabilistic transition matrix captures inter-label dependencies. • “Rough prior” from target data effectively guides source data pre-training. • GDAIG outperforms conventional baselines across diverse multi-label scenarios.
Kanyu Miyoshi, Ryotaro Shimizu, Linxin Song, Masayuki Goto
Neurocomputing2
2025 Static Word Embeddings for Sentence Semantic Representation
abstract
We propose new static word embeddings optimised for sentence semantic representation.We first extract word embeddings from a pretrained Sentence Transformer, and improve them with sentence-level principal component analysis, followed by either knowledge distillation or contrastive learning.During inference, we represent sentences by simply averaging word embeddings, which requires little computational cost.We evaluate models on both monolingual and cross-lingual tasks and show that our model substantially outperforms existing static models on sentence semantic tasks, and even surpasses a basic Sentence Transformer model (SimCSE) on a text embedding benchmark.Lastly, we perform a variety of analyses and show that our method successfully removes word embedding components that are not highly relevant to sentence semantics, and adjusts the vector norms based on the influence of words on sentence semantics.
Takashi Wada 0001, Yuki Hirakawa, Ryotaro Shimizu, Takahiro Kawashima, Yuki Saito 0002
EMNLP3
2025 Mastering Task Arithmetic: τJp as a Key Indicator for Weight Disentanglement
Kotaro Yoshida, Yuji Naraki, Takafumi Horie, Ryosuke Yamaki, Ryotaro Shimizu, Yuki Saito 0002, Julian J. McAuley, Hiroki Naganuma
ICLR5
2025 Attributed Synthetic Data Generation for Zero-shot Domain-specific Image Classification
abstract
Zero-shot domain-specific image classification is challenging in classifying real images without ground-truth in-domain training examples. Recent research involved knowledge from texts with a text-to-image model to generate in-domain training images in zero-shot scenarios. However, existing methods heavily rely on simple prompt strategies, limiting the diversity of synthetic training images, thus leading to inferior performance compared to real images. In this paper, we propose AttrSyn, which leverages large language models to generate attributed prompts. These prompts allow for the generation of more diverse attributed synthetic images. Experiments for zero-shot domain-specific image classification on two fine-grained datasets show that training with synthetic images generated by AttrSyn significantly outperforms CLIP’s zero-shot classification under most situations and consistently surpasses simple prompt strategies.
Shijian Wang, Linxin Song, Ryotaro Shimizu, Masayuki Goto, Hanqian Wu
ICME3
2025 Explaining Black-box Model Predictions via Two-level Nested Feature Attributions with Consistency Property
abstract
Techniques that explain the predictions of black-box machine learning models are crucial to make the models transparent, thereby increasing trust in AI systems. The input features to the models often have a nested structure that consists of high- and low-level features, and each high-level feature is decomposed into multiple low-level features. For such inputs, both high-level feature attributions (HiFAs) and low-level feature attributions (LoFAs) are important for better understanding the model's decision. In this paper, we propose a model-agnostic local explanation method that effectively exploits the nested structure of the input to estimate the two-level feature attributions simultaneously. A key idea of the proposed method is to introduce the consistency property that should exist between the HiFAs and LoFAs, thereby bridging the separate optimization problems for estimating them. Thanks to this consistency property, the proposed method can produce HiFAs and LoFAs that are both faithful to the black-box models and consistent with each other, using a smaller number of queries to the models. In experiments on image classification in multiple instance learning and text classification using language models, we demonstrate that the HiFAs and LoFAs estimated by the proposed method are accurate, faithful to the behaviors of the black-box models, and provide consistent explanations.
Yuya Yoshikawa, Masanari Kimura, Ryotaro Shimizu, Yuki Saito 0002
IJCAI3
2025 Normative Alignment of Recommender Systems via Internal Label Shift
abstract
We introduce NAILS (Normative Alignment of Recommender Systems via Internal Label Shift), a simple and scalable method for aligning recommendation outputs with target distributions over item-level attributes, such as categories. Recommender systems optimized solely for user engagement often fail to satisfy broader normative objectives, including fairness, diversity, and editorial values. NAILS modifies the user-conditional item distribution to induce a specified marginal distribution over attributes while preserving the preferences learned by an existing recommender system and requiring no model retraining. We formulate this problem as a form of label shift applied internally within a hierarchical classification framework. By adopting a stakeholder-centric perspective, NAILS enables recommendation outputs to be aligned with global normative objectives. Empirically, we show that NAILS consistently improves attribute-level alignment with minimal impact on user engagement, providing a practical mechanism for value-driven recommendation.
Johannes Kruse 0002, Kasper Lindskow, Michael Riis Andersen, Ryotaro Shimizu, Julian J. McAuley, Pierre-Alexandre Mattei, Jes Frellsen
RecSys4
2025 Disentangling Likes and Dislikes in Personalized Generative Explainable Recommendation
abstract
Recent research on explainable recommendation generally frames the task as a standard text generation problem, and evaluates models simply based on the textual similarity between the predicted and ground-truth explanations. However, this approach fails to consider one crucial aspect of the systems: whether their outputs accurately reflect the users' (post-purchase) sentiments, i.e., whether and why they would like and/or dislike the recommended items. To shed light on this issue, we introduce new datasets and evaluation methods that focus on the users' sentiments. Specifically, we construct the datasets by explicitly extracting users' positive and negative opinions from their post-purchase reviews using an LLM, and propose to evaluate systems based on whether the generated explanations 1) align well with the users' sentiments, and 2) accurately identify both positive and negative opinions of users on the target items. We benchmark several recent models on our datasets and demonstrate that achieving strong performance on existing metrics does not ensure that the generated explanations align well with the users' sentiments. Lastly, we find that existing models can provide more sentiment-aware explanations when the users' (predicted) ratings for the target items are directly fed into the models as input. The datasets and benchmark implementation are available at: https://github.com/jchanxtarov/sent_xrec.
Ryotaro Shimizu, Takashi Wada 0001, Yu Wang 0170, Johannes Kruse 0002, Sean O'Brien, Sai Htaung Kham, Linxin Song, Yuya Yoshikawa, Yuki Saito 0002, Fugee Tsung, Masayuki Goto, Julian J. McAuley
WWW1
2025 LLMOverTab: Tabular data augmentation with language model-driven oversampling
abstract
In recent years, Large Language Model (LLM) have seen significant advancements, attracting attention for their applications in various fields. These models have shown promising results in handling tabular data, especially in cases with limited datasets, by leveraging pre-trained knowledge. However, their effectiveness in addressing imbalanced data in tabular formats is less explored. To bridge this gap, our study introduces LLMOverTab, a novel approach using LLMs for oversampling in imbalanced tabular data. We conducted comprehensive experiments on diverse tabular datasets to assess the effectiveness of LLMOverTab, demonstrating its potential in improving the handling of imbalanced data. The study also explores application of LLMOverTab in zero-shot and few-shot learning contexts, providing insights into its adaptability. Additionally, we analyze the oversampled data, offering reflections on the quality of generated samples. Our research not only showcases the utility of LLMOverTab in managing imbalanced tabular data, but also opens new avenues for the application of language models in various tasks of tabular data. This study adds to the increasing interest in applying LLMs to various task domains. It provides new perspectives for the innovative use of LLMs in structured tabular data fields, highlighting their potential in a range of applications. • Introduces LLMOverTab for oversampling in imbalanced tabular data. • Surpasses traditional methods like SMOTE and other LLM approaches. • Uses prompt engineering to generate meaningful synthetic instances. • Finds LLMOverTab excels especially with LLM prediction models. • Suggests exploring different LLM architectures and prompt techniques.
Tokimasa Isomura, Ryotaro Shimizu, Masayuki Goto
Expert Syst. Appl.2
2025 Generating realistic synthetic tabular data with integrated LLM and diffusion models
abstract
Generating realistic synthetic tabular data is a crucial task for privacy-preserving data sharing, data augmentation, and learning from limited samples. However, existing methods often struggle with small sample sizes, heterogeneous feature types, and generalization in real-world scenarios. We propose TabularMDLM, a novel diffusion-based generative framework that integrates masked language modeling to synthesize high-quality tabular data. Unlike prior work, TabularMDLM applies noise only to feature values—preserving the semantic relationship between column names and values—and leverages pre-trained language models to iteratively reconstruct masked tokens during reverse diffusion. This design enhances generation quality while ensuring structural consistency. To evaluate the framework, we conduct extensive experiments on six tabular datasets of varying sizes, domains, and feature types. We compare TabularMDLM against recent baselines, including TabDDPM, CTGAN, and CTGAN+, under a privacy-conscious setting where only synthetic data is used for training and real data for testing. We also assess performance under class imbalance to validate generalization. Results show that TabularMDLM achieves consistently strong classification performance across accuracy, precision, recall, and F1 score, outperforming baselines in both balanced and imbalanced settings. In contrast to existing methods, TabularMDLM scales to diverse data types and low-resource regimes, offering practical advantages in privacy-sensitive applications. • Introduces TabularMDLM, a generative framework combining LLMs and diffusion-inspired refinement for synthetic tabular data. • Selective masking of feature values, while preserving feature names, ensures schema integrity and richer feature interactions. • Generates more diverse and representative samples than conventional oversampling methods. • Demonstrates improved predictive accuracy and robustness, even under challenging feature compositions and severe class imbalances. • Facilitates privacy-compliant data augmentation and fairer decision-making in data-sensitive and resource-constrained scenarios.
Tokimasa Isomura, Ryotaro Shimizu, Masayuki Goto
Neurocomputing2
2025 Optimizing pre-training via target-aware source data selection
abstract
• Domain adaptation information gain-based target-related data selection is proposed. • Proposed DAIG improves target model accuracy via target-related data selection. • DAIG effectively gathers relevant data from source to benefit target tasks. • DAIG extracts “rough prior” from target data for source data pre-training. • DAIG-guided selection outperforms baselines in multiple experimental settings. In recent years, due to the explosive popularity of large-scale pre-trained models such as large language models, pre-training approaches that use a massive amount of source data and can be applied to various target tasks are becoming more popular. Pre-trained models allow us to learn highly accurate target models by fine-tuning them with target data, even when their volume is insufficient. However, the source data used to train a pre-training model is generally a large and miscellaneous data set obtained in the wild without being aware of the target task, and it highly possibly contains much data that does not contribute to relearning the target task. This study defines a novel paradigm as “target-aware source data selection,” which uses the source data itself instead of a pre-training model and selects source data for pre-training and aims to increase its quality, effectiveness, and robustness. Our proposal fundamentally differs from the current studies addressing the lack of target data and conventional transfer learning approaches, improving source data quality using the novel Domain Adaptation Information Gain criteria. Specifically, the target model is pre-trained while actively selecting only informative data from the source data using the “rough-prior knowledge” obtained from the target data training before the pre-training. Finally, fine-tuning the model with the target data results in a highly accurate model for the target (downstream) task. The effectiveness of our proposed paradigm has been demonstrated through multifaceted experiments using multiple pairs of target data and source data with different strengths of their relevance.
Kanyu Miyoshi, Ryotaro Shimizu, Linxin Song, Masayuki Goto
Knowl. Based Syst.2
2025 Sparse attention is all you need for pre-training on tabular data
abstract
Abstract In the world of data-driven decision-making, tabular data reigns supreme as the most prevalent and crucial format, especially in business contexts. However, data scarcity remains a recurring challenge. In this context, transfer learning has emerged as a potent solution. This study explores the untapped potential of transfer learning in the realm of tabular data analysis, with a focus on leveraging deep learning models—especially the Transformer model—that have garnered significant recognition. Our research investigates the intricacies of tabular data and illuminates the shortcomings of conventional attention mechanisms in the Transformer model when applied to such structured datasets. This highlights the pressing requirement need for specialized solutions tailored to tabular data. We introduce an innovative transfer learning method based on series of thoroughly designed experiments across diverse business domains. This approach harnesses Transformer-based models enhanced with optimized sparse attention mechanisms, offering a groundbreaking solution for tabular data analysis. Our findings reveal the remarkable effectiveness of enhancing the attention mechanism within the Transformer in transfer learning. Specifically, pre-training with sparse attention proves increasingly powerful as data volumes increase, resulting in superior performance on large datasets. Conversely, fine-tuning with full attention becomes more impactful when data availability decreases in downstream tasks, ensuring adaptability in situations with limited data. The empirical results presented in this study provide compelling evidence of the revolutionary potential of our approach. Our optimized sparse attention model emerges as a powerful tool for researchers and practitioners seeking highly effective solutions for tabular data tasks. As tabular data remain the backbone of business operations, our study promises to revolutionize data analysis in critical domains. This work bridges the gap between limited data availability and the requirement for effective analysis in business settings, marking a significant step forward in the field of tabular data analysis.
Tokimasa Isomura, Ryotaro Shimizu, Masayuki Goto
Neural Comput. Appl.2
2023 Fashion intelligence system: An outfit interpretation utilizing images and rich abstract tags
abstract
In recent years, it has become common for consumers to familiarize themselves with the latest fashion trends through the internet and engage in their own fashion-inspired shopping activities. Therefore, making fashion-inspired shopping and browsing activities (internet surfing in the fashion domain) comfortable is essential because it leads to interactions in the fashion industry. However, fashion is a fuzzy and complex domain that contains many abstract elements, and this ambiguity and complexity can hinder users’ deep interest in the fashion industry. Therefore, we define a novel technology and domain called “fashion intelligence” and propose a system based on a visual-semantic embedding method for automatically learning and interpreting fashion and obtaining answers to users’ questions. Our proposed method can embed the abundant abstract tag information in the same projective space as outfit images. Mapping of images and tags in a projective space helps search for outfit images using fashion-specific abstract words. In addition, visually estimating the degree of relevance between images and tags helps interpret abstract words. As a result, this research helps decrease fashion-specific ambiguity and complexity and supports the marketing activities and fashion choices of both experts and non-experts.
Ryotaro Shimizu, Yuki Saito 0002, Megumi Matsutani, Masayuki Goto
Expert Syst. Appl.1
2023 Partial visual-semantic embedding: Fine-grained outfit image representation with massive volumes of tags via angular-based contrastive learning
abstract
A novel technology named fashion intelligence system, which quantifies ambiguous expressions unique to fashion, such as “casual,” “adult-casual,” and “office-casual,” was previously proposed to support users in their understanding of fashion. However, the existing visual-semantic embedding (VSE) model, which forms the basis of the system, does not support images that are composed of multiple parts, such as those containing hair, tops, trousers, skirts, and shoes. Therefore, we propose a partial VSE (PVSE) model, which enables fine-grained learning of each part of the fashion outfit. The proposed model learns embedded representations via angular-based contrastive learning. This helps in retaining three existing practical functionalities and further enables image-retrieval tasks where changes are only made to specified parts and image-reordering tasks focusing on the specified parts. In other words, the proposed model enables five types of practical functionalities, even with a simple structure. Through qualitative and quantitative experiments, we demonstrate that the proposed model is superior to conventional models, without increasing computational complexity.
Ryotaro Shimizu, Takuma Nakamura, Masayuki Goto
Knowl. Based Syst.1
2022 An explainable recommendation framework based on an improved knowledge graph attention network with massive volumes of side information
abstract
In recent years, explainable recommendation has been a topic of active study. This is because the branch of the machine learning field related to methodologies is enabling human understanding of the reasons for the outputs of recommender systems. The realization of explainable recommendation is widely expected to increase both user satisfaction and the demand for explainable recommendation systems. Explainable recommendation utilizes a wealth of side information (such as sellers, brands, user ages and genders, and bookmark information, among others) to expound the decision-making reasoning applied by recommendation models. In explainable recommendation, although learning side information containing numerous variables leads to rich interpretability, learning too many variables presents a challenge because decreases the amount of learning that a given computational resource can perform, and the accuracy of the recommendation model may be degraded. However, numerous and diverse variables are included in the side information stored by the actual companies operating massive real-world services. Hence, to realize practical applications of this valuable information, it is necessary to resolve problems such as computational cost. In this study, we propose a new framework for explainable recommendation based on an improved knowledge graph attention network model, which utilizes the side information of items and realizes high recommendation accuracy. The proposed framework enables direct interpretation by visualizing the reasons for the recommendations provided. Experimental results show that the proposed framework reduced computational time requirements by approximately 80%, while maintaining recommendation accuracy by enabling the model to learn the probabilistically given edges included in the graph structure. Moreover, the results show that the proposed framework exhibited richer interpretability than the conventional model. Finally, a multifaceted analysis suggests that the proposed framework is not only effective as an explainable recommendation model but also provides a powerful tool for planning various marketing strategies.
Ryotaro Shimizu, Megumi Matsutani, Masayuki Goto
Knowl. Based Syst.1