Ling Xiao 0001

dblp:59/4568-1 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-4650-8841ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Language-guided frameworks for personalized video summarization
abstract
Existing video summarization methods predominantly produce generic summaries and often fail to reflect user-specific preferences. To address this limitation, we explore the potential of large language models (LLMs) for video summarization and propose two language-guided frameworks for personalized video summarization. We first propose Few-Shot Video SUMmarization (FS-VSUM), a non-trainable, example-driven framework that leverages LLM-based semantic reasoning to perform annotator-personalized video summarization. By conditioning on a small number of annotated examples, FS-VSUM captures annotator-specific summarization styles and generates customized summaries without parameter updates, demonstrating the inherent capability of LLMs for controllable and personalized video summarization. We then introduce Self-Supervised Video SUMmarization (SS-VSUM), a trainable framework that formulates video summarization as a semantic textual similarity task. SS-VSUM incorporates user preferences through LLM prompts and introduces a Preserving Diversity Loss (PDL) to dynamically regulate regularization based on linguistic diversity. We further extend SS-VSUM with additional analyses and clarifications, providing a more systematic understanding of language-guided video summarization. Experimental results show that SS-VSUM achieves state-of-the-art performance on the SumMe dataset. Together, this work provides a systematic investigation of language-guided video summarization, revealing how LLMs can support both training-free personalization and trainable performance optimization. The source code for the proposed frameworks is publicly available at https://github.com/sugitomoo/VSUM.
Tomoya Sugihara, Shuntaro Masuda, Ling Xiao 0001, Toshihiko Yamasaki
Pattern Anal. Appl.3
2025 ActRecognition-GPT: Utilizing Multimodal Large Language Models for Spatiotemporal Action Recognition in Nursery Videos
abstract
Spatiotemporal action recognition in nursery videos is essential for intelligent childcare systems that can automatically monitor children’s behaviors, detect safety risks, and generate accurate activity logs without increasing caregivers’ burden. However, occlusions, dynamic multi-person interactions, and subtle distinctions between similar actions present significant challenges. To address these challenges, we introduce ActRecognition-GPT, a multimodal framework that combines visual object detection with the reasoning capability of multimodal large language models (MLLMs). By assigning unique person IDs and incorporating both spatial and temporal context through structured prompts, the model achieves temporally consistent recognition of individual behaviors, even under visual ambiguity or occlusion. Moreover, predefined vocabularies are used to constrain generation and reduce semantic drift. Crucially, the open-vocabulary nature of MLLMs enables ActRecognition-GPT to generalize to previously unseen or ambiguous scenarios without requiring extensive labeled data, which is an essential property for modeling the diverse and evolving behaviors of children. Experimental results on a real-world nursery video dataset demonstrate a significant improvement in mean Average Precision over single-frame baselines, validating the effectiveness of combining structured visual inputs with LLM-based temporal reasoning.
Kenta Watanabe, Shuntaro Masuda, Ling Xiao 0001, Toshihiko Yamasaki
FG3
2025 TourMLLM: A Retrieval-Augmented Multimodal Large Language Model for Multitask Learning in the Tourism Domain
Hiromasa Yamanishi, Ling Xiao 0001, Toshihiko Yamasaki
ICMR2
2025 LITA: LMM-Guided Image-Text Alignment for Art Assessment
Tatsumi Sunada, Kaede Shiohara, Ling Xiao 0001, Toshihiko Yamasaki
MMM (2)3
2025 Multi-level knowledge distillation for fine-grained fashion image retrieval
Ling Xiao 0001, Toshihiko Yamasaki
Knowl. Based Syst.1
2024 Improving Plasticity in Online Continual Learning via Collaborative Learning
abstract
Online Continual Learning (CL) solves the problem of learning the ever-emerging new classification tasks from a continuous data stream. Unlike its offline counterpart, in online CL, the training data can only be seen once. Most existing online CL research regards catastrophic forgetting (i.e., model stability) as almost the only challenge. In this paper, we argue that the model's capability to acquire new knowledge (i.e., model plasticity) is another challenge in online CL. While replay-based strategies have been shown to be effective in alleviating catastrophic forgetting, there is a notable gap in research attention toward improving model plasticity. To this end, we propose Collaborative Continual Learning (CCL), a collaborative learning based strategy to improve the model's capability in acquiring new concepts. Additionally, we introduce Distillation Chain (DC), a collaborative learning scheme to boost the training of the models. We adapt CCL-DC to existing representative online CL works. Extensive experiments demonstrate that even if the learners are well-trained with state-of-the-art online CL methods, our strategy can still improve model plasticity dramatically, and thereby improve the overall performance by a large margin. The source code of our work is available at https://github.com/maorong-wang/CCL-DC.
Maorong Wang, Nicolas Michel, Ling Xiao 0001, Toshihiko Yamasaki
CVPR3
2024 SCOMatch: Alleviating Overtrusting in Open-Set Semi-supervised Learning
Zerun Wang, Liuyu Xiang, Lang Huang 0001, Jiafeng Mao, Ling Xiao 0001, Toshihiko Yamasaki
ECCV (51)5
2024 Adversarially Robust Continual Learning with Anti-Forgetting Loss
abstract
Existing continual learning methods focus on preventing catastrophic forgetting but often overlook the challenge of adversarial examples in image classification. In this study, we propose a novel method that balances accuracy, robustness against adversarial examples, and the prevention of forgetting. Specifically, we first theoretically and experimentally demonstrate that learning through knowledge distillation, a common strategy in continual learning, conflicts with learning through the cross-entropy loss. To resolve this conflict, we propose a novel loss function that combines an additional memory data loss with a conflict-avoiding knowledge distillation loss, effectively preventing catastrophic forgetting while ensuring robustness. Experimental results show that the proposed method outperforms existing methods by 5.17% in clean accuracy and $2.10 \%$ in robust accuracy. This method proves to be especially beneficial in scenarios where the reuse of samples from previous tasks is limited.
Koki Mukai, Soichiro Kumano, Nicolas Michel, Ling Xiao 0001, Toshihiko Yamasaki
ICIP4
2024 Rethinking Momentum Knowledge Distillation in Online Continual Learning
abstract
Online Continual Learning (OCL) addresses the problem of training neural networks on a continuous data stream where multiple classification tasks emerge in sequence. In contrast to offline Continual Learning, data can be seen only once in OCL, which is a very severe constraint. In this context, replay-based strategies have achieved impressive results and most state-of-the-art approaches heavily depend on them. While Knowledge Distillation (KD) has been extensively used in offline Continual Learning, it remains under-exploited in OCL, despite its high potential. In this paper, we analyze the challenges in applying KD to OCL and give empirical justifications. We introduce a direct yet effective methodology for applying Momentum Knowledge Distillation (MKD) to many flagship OCL methods and demonstrate its capabilities to enhance existing approaches. In addition to improving existing state-of-the-art accuracy by more than $10%$ points on ImageNet100, we shed light on MKD internal mechanics and impacts during training in OCL. We argue that similar to replay, MKD should be considered a central component of OCL. The code is available at https://github.com/Nicolas1203/mkd_ocl.
Nicolas Michel, Maorong Wang, Ling Xiao 0001, Toshihiko Yamasaki
ICML3
2024 Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video
Tomoya Sugihara, Shuntaro Masuda, Ling Xiao 0001, Toshihiko Yamasaki
MMAsia3
2024 E-ReaRev: Adaptive Reasoning for Question Answering over Incomplete Knowledge Graphs by Edge and Meaning Extensions
Xiaotong Ye, Ling Xiao 0001, Toshihiko Yamasaki
NLDB (2)2
2024 LLaVA-Tour: A Large Multimodal Model for Japanese Tourist Spot Prediction and Review Generation
abstract
Tourist landmark recognition and review generation can significantly enhance travel planning by helping travelers make informed decisions, optimize experiences, and discover new destinations. These capabilities also benefit local economies and improve business efficiency in tourism. The advancements in large multimodal models have demonstrated high performance across a wide range of image processing tasks due to their extensive knowledge and reasoning capabilities. However, there are no large-scale tourism datasets and no models specifically tailored for tourism have been developed. To address these issues, we created a new dataset with over 1.3 million entries from Japanese tourism website Jalan.net, covering three types of tasks: landmark recognition, general and conditioned review generation and one support task: description generation. Using the Large Language-and-Vision Assistant (LLaVA) as the baseline, we developed instruction tuning strategies for these tasks and created a new Large Multimodal Model, LLaVA-Tour. By integrating domain-specific knowledge, our model outperformed state-of-the-art large multimodal models and review generation models in landmark recognition, and general and conditioned review generation. The code and data url are available at https://github.com/HiromasaYamanishi/LLaVATour.
Hiromasa Yamanishi, Ling Xiao 0001, Toshihiko Yamasaki
VCIP2
2024 STFE-Net: A multi-stage approach to enhance statistical texture feature for defect detection on metal surfaces
Daxing Fu, Ling Xiao 0001, Jie Liu 0017, Youmin Hu, Bo Wu 0006
Adv. Eng. Informatics3
2024 LiFSO-Net: A lightweight feature screening optimization network for complex-scale flat metal defect detection
Ling Xiao 0001, Chenhui Wan, Youmin Hu, Bo Wu 0006
Knowl. Based Syst.2
2022 Sat: Self-Adaptive Training for Fashion Compatibility Prediction
abstract
This paper presents a self-adaptive training (SAT) model for fashion compatibility prediction. It focuses on the learning of some hard items, such as those that share similar color, texture, and pattern features but are considered incompatible due to the aesthetics or temporal shifts. Specifically, we first design a method to define hard outfits and a difficulty score (DS) is defined and assigned to each outfit based on the difficulty in recommending an item for it. Then, we propose a self-adaptive triplet loss (SATL), where the DS of the outfit is considered. Finally, we propose a very simple conditional similarity network combining the proposed SATL to achieve the learning of hard items in the fashion compatibility prediction. Experiments on the publicly available Polyvore Outfits and Polyvore Outfits-D datasets demonstrate our SAT’s effectiveness in fashion compatibility prediction. Besides, our SATL can be easily extended to other conditional similarity networks to improve their performance.
Ling Xiao 0001, Toshihiko Yamasaki
ICIP1
2020 OSED: Object-specific edge detection
Ling Xiao 0001, Bo Wu 0006, Youmin Hu
J. Vis. Commun. Image Represent.1