EDBT 2026 Demo / reviewers in the wild / expert
Ling Xiao 0001
dblp:59/4568-1
· DBLP profile ↗
16ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-4650-8841ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Language-guided frameworks for personalized video summarizationabstractExisting video summarization methods predominantly produce generic summaries and often fail to reflect user-specific preferences. To address this limitation, we explore the potential of large language models (LLMs) for video summarization and propose two language-guided frameworks for personalized video summarization. We first propose Few-Shot Video SUMmarization (FS-VSUM), a non-trainable, example-driven framework that leverages LLM-based semantic reasoning to perform annotator-personalized video summarization. By conditioning on a small number of annotated examples, FS-VSUM captures annotator-specific summarization styles and generates customized summaries without parameter updates, demonstrating the inherent capability of LLMs for controllable and personalized video summarization. We then introduce Self-Supervised Video SUMmarization (SS-VSUM), a trainable framework that formulates video summarization as a semantic textual similarity task. SS-VSUM incorporates user preferences through LLM prompts and introduces a Preserving Diversity Loss (PDL) to dynamically regulate regularization based on linguistic diversity. We further extend SS-VSUM with additional analyses and clarifications, providing a more systematic understanding of language-guided video summarization. Experimental results show that SS-VSUM achieves state-of-the-art performance on the SumMe dataset. Together, this work provides a systematic investigation of language-guided video summarization, revealing how LLMs can support both training-free personalization and trainable performance optimization. The source code for the proposed frameworks is publicly available at https://github.com/sugitomoo/VSUM. Tomoya Sugihara, Shuntaro Masuda, Ling Xiao 0001, Toshihiko Yamasaki |
Pattern Anal. Appl. | 3 |
| 2025 | ActRecognition-GPT: Utilizing Multimodal Large Language Models for Spatiotemporal Action Recognition in Nursery VideosabstractSpatiotemporal action recognition in nursery videos is essential for intelligent childcare systems that can automatically monitor children’s behaviors, detect safety risks, and generate accurate activity logs without increasing caregivers’ burden. However, occlusions, dynamic multi-person interactions, and subtle distinctions between similar actions present significant challenges. To address these challenges, we introduce ActRecognition-GPT, a multimodal framework that combines visual object detection with the reasoning capability of multimodal large language models (MLLMs). By assigning unique person IDs and incorporating both spatial and temporal context through structured prompts, the model achieves temporally consistent recognition of individual behaviors, even under visual ambiguity or occlusion. Moreover, predefined vocabularies are used to constrain generation and reduce semantic drift. Crucially, the open-vocabulary nature of MLLMs enables ActRecognition-GPT to generalize to previously unseen or ambiguous scenarios without requiring extensive labeled data, which is an essential property for modeling the diverse and evolving behaviors of children. Experimental results on a real-world nursery video dataset demonstrate a significant improvement in mean Average Precision over single-frame baselines, validating the effectiveness of combining structured visual inputs with LLM-based temporal reasoning. Kenta Watanabe, Shuntaro Masuda, Ling Xiao 0001, Toshihiko Yamasaki |
FG | 3 |
| 2025 | TourMLLM: A Retrieval-Augmented Multimodal Large Language Model for Multitask Learning in the Tourism Domain
Hiromasa Yamanishi, Ling Xiao 0001, Toshihiko Yamasaki |
ICMR | 2 |
| 2025 | LITA: LMM-Guided Image-Text Alignment for Art Assessment
Tatsumi Sunada, Kaede Shiohara, Ling Xiao 0001, Toshihiko Yamasaki |
MMM (2) | 3 |
| 2025 | Multi-level knowledge distillation for fine-grained fashion image retrieval
Ling Xiao 0001, Toshihiko Yamasaki |
Knowl. Based Syst. | 1 |
| 2024 | Improving Plasticity in Online Continual Learning via Collaborative LearningabstractOnline Continual Learning (CL) solves the problem of learning the ever-emerging new classification tasks from a continuous data stream. Unlike its offline counterpart, in online CL, the training data can only be seen once. Most existing online CL research regards catastrophic forgetting (i.e., model stability) as almost the only challenge. In this paper, we argue that the model's capability to acquire new knowledge (i.e., model plasticity) is another challenge in online CL. While replay-based strategies have been shown to be effective in alleviating catastrophic forgetting, there is a notable gap in research attention toward improving model plasticity. To this end, we propose Collaborative Continual Learning (CCL), a collaborative learning based strategy to improve the model's capability in acquiring new concepts. Additionally, we introduce Distillation Chain (DC), a collaborative learning scheme to boost the training of the models. We adapt CCL-DC to existing representative online CL works. Extensive experiments demonstrate that even if the learners are well-trained with state-of-the-art online CL methods, our strategy can still improve model plasticity dramatically, and thereby improve the overall performance by a large margin. The source code of our work is available at https://github.com/maorong-wang/CCL-DC. Maorong Wang, Nicolas Michel, Ling Xiao 0001, Toshihiko Yamasaki |
CVPR | 3 |
| 2024 | SCOMatch: Alleviating Overtrusting in Open-Set Semi-supervised Learning
Zerun Wang, Liuyu Xiang, Lang Huang 0001, Jiafeng Mao, Ling Xiao 0001, Toshihiko Yamasaki |
ECCV (51) | 5 |
| 2024 | Adversarially Robust Continual Learning with Anti-Forgetting LossabstractExisting continual learning methods focus on preventing catastrophic forgetting but often overlook the challenge of adversarial examples in image classification. In this study, we propose a novel method that balances accuracy, robustness against adversarial examples, and the prevention of forgetting. Specifically, we first theoretically and experimentally demonstrate that learning through knowledge distillation, a common strategy in continual learning, conflicts with learning through the cross-entropy loss. To resolve this conflict, we propose a novel loss function that combines an additional memory data loss with a conflict-avoiding knowledge distillation loss, effectively preventing catastrophic forgetting while ensuring robustness. Experimental results show that the proposed method outperforms existing methods by 5.17% in clean accuracy and $2.10 \%$ in robust accuracy. This method proves to be especially beneficial in scenarios where the reuse of samples from previous tasks is limited. Koki Mukai, Soichiro Kumano, Nicolas Michel, Ling Xiao 0001, Toshihiko Yamasaki |
ICIP | 4 |
| 2024 | Rethinking Momentum Knowledge Distillation in Online Continual LearningabstractOnline Continual Learning (OCL) addresses the problem of training neural networks on a continuous data stream where multiple classification tasks emerge in sequence. In contrast to offline Continual Learning, data can be seen only once in OCL, which is a very severe constraint. In this context, replay-based strategies have achieved impressive results and most state-of-the-art approaches heavily depend on them. While Knowledge Distillation (KD) has been extensively used in offline Continual Learning, it remains under-exploited in OCL, despite its high potential. In this paper, we analyze the challenges in applying KD to OCL and give empirical justifications. We introduce a direct yet effective methodology for applying Momentum Knowledge Distillation (MKD) to many flagship OCL methods and demonstrate its capabilities to enhance existing approaches. In addition to improving existing state-of-the-art accuracy by more than $10%$ points on ImageNet100, we shed light on MKD internal mechanics and impacts during training in OCL. We argue that similar to replay, MKD should be considered a central component of OCL. The code is available at https://github.com/Nicolas1203/mkd_ocl. Nicolas Michel, Maorong Wang, Ling Xiao 0001, Toshihiko Yamasaki |
ICML | 3 |
| 2024 | Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video
Tomoya Sugihara, Shuntaro Masuda, Ling Xiao 0001, Toshihiko Yamasaki |
MMAsia | 3 |
| 2024 | E-ReaRev: Adaptive Reasoning for Question Answering over Incomplete Knowledge Graphs by Edge and Meaning Extensions
Xiaotong Ye, Ling Xiao 0001, Toshihiko Yamasaki |
NLDB (2) | 2 |
| 2024 | LLaVA-Tour: A Large Multimodal Model for Japanese Tourist Spot Prediction and Review GenerationabstractTourist landmark recognition and review generation can significantly enhance travel planning by helping travelers make informed decisions, optimize experiences, and discover new destinations. These capabilities also benefit local economies and improve business efficiency in tourism. The advancements in large multimodal models have demonstrated high performance across a wide range of image processing tasks due to their extensive knowledge and reasoning capabilities. However, there are no large-scale tourism datasets and no models specifically tailored for tourism have been developed. To address these issues, we created a new dataset with over 1.3 million entries from Japanese tourism website Jalan.net, covering three types of tasks: landmark recognition, general and conditioned review generation and one support task: description generation. Using the Large Language-and-Vision Assistant (LLaVA) as the baseline, we developed instruction tuning strategies for these tasks and created a new Large Multimodal Model, LLaVA-Tour. By integrating domain-specific knowledge, our model outperformed state-of-the-art large multimodal models and review generation models in landmark recognition, and general and conditioned review generation. The code and data url are available at https://github.com/HiromasaYamanishi/LLaVATour. Hiromasa Yamanishi, Ling Xiao 0001, Toshihiko Yamasaki |
VCIP | 2 |
| 2024 | STFE-Net: A multi-stage approach to enhance statistical texture feature for defect detection on metal surfaces
Daxing Fu, Ling Xiao 0001, Jie Liu 0017, Youmin Hu, Bo Wu 0006 |
Adv. Eng. Informatics | 3 |
| 2024 | LiFSO-Net: A lightweight feature screening optimization network for complex-scale flat metal defect detection
Ling Xiao 0001, Chenhui Wan, Youmin Hu, Bo Wu 0006 |
Knowl. Based Syst. | 2 |
| 2022 | Sat: Self-Adaptive Training for Fashion Compatibility PredictionabstractThis paper presents a self-adaptive training (SAT) model for fashion compatibility prediction. It focuses on the learning of some hard items, such as those that share similar color, texture, and pattern features but are considered incompatible due to the aesthetics or temporal shifts. Specifically, we first design a method to define hard outfits and a difficulty score (DS) is defined and assigned to each outfit based on the difficulty in recommending an item for it. Then, we propose a self-adaptive triplet loss (SATL), where the DS of the outfit is considered. Finally, we propose a very simple conditional similarity network combining the proposed SATL to achieve the learning of hard items in the fashion compatibility prediction. Experiments on the publicly available Polyvore Outfits and Polyvore Outfits-D datasets demonstrate our SAT’s effectiveness in fashion compatibility prediction. Besides, our SATL can be easily extended to other conditional similarity networks to improve their performance. Ling Xiao 0001, Toshihiko Yamasaki |
ICIP | 1 |
| 2020 | OSED: Object-specific edge detection
Ling Xiao 0001, Bo Wu 0006, Youmin Hu |
J. Vis. Commun. Image Represent. | 1 |