VLDB 2026 Research / reviewers in the wild / expert
Guang Li 0008
dblp:14/3764-8
· DBLP profile ↗
17ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0003-2898-2504ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV
Wenbo Huang 0001, Jinghui Zhang 0001, Guang Li 0008, Lei Zhang 0130, Fang Dong 0001, Takahiro Ogawa 0001, Miki Haseyama |
AAAI | 4 |
| 2025 | Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-SequenceabstractIn few-shot action recognition (FSAR), long sub-sequences of video naturally express entire actions more effectively. However, the high computational complexity of mainstream Transformer-based methods limits their application. Recent Mamba demonstrates efficiency in modeling long sequences, but directly applying Mamba to FSAR overlooks the importance of local feature modeling and alignment. Moreover, long sub-sequences within the same class accumulate intra-class variance, which adversely impacts FSAR performance. To solve these challenges, we propose a Matryoshka MAmba and CoNtrasTive LeArning framework (Manta). Firstly, the Matryoshka Mamba introduces multiple Inner Modules to enhance local feature representation, rather than directly modeling global features. An Outer Module captures dependencies of timeline between these local features for implicit temporal alignment. Secondly, a hybrid contrastive learning paradigm, combining both supervised and unsupervised methods, is designed to mitigate the negative effects of intra-class variance accumulation. The Matryoshka Mamba and the hybrid contrastive learning paradigm operate in two parallel branches within Manta, enhancing Mamba for FSAR of long sub-sequence. Manta achieves new state-of-the-art performance on prominent benchmarks, including SSv2, Kinetics, UCF101, and HMDB51. Extensive empirical studies prove that Manta significantly improves FSAR of long sub-sequence from multiple perspectives. Wenbo Huang 0001, Jinghui Zhang 0001, Guang Li 0008, Lei Zhang 0130, Shuoyuan Wang, Fang Dong 0001, Jiahui Jin 0001, Takahiro Ogawa 0001, Miki Haseyama |
AAAI | 3 |
| 2025 | Generative Dataset Distillation Based on Self-knowledge DistillationabstractDataset distillation is an effective technique for reducing the cost and complexity of model training while maintaining performance by compressing large datasets into smaller, more efficient versions. In this paper, we present a novel generative dataset distillation method that can improve the accuracy of aligning prediction logits. Our approach integrates self-knowledge distillation to achieve more precise distribution matching between the synthetic and original data, thereby capturing the overall structure and relationships within the data. To further improve the accuracy of alignment, we introduce a standardization step on the logits before performing distribution matching, ensuring consistency in the range of logits. Through extensive experiments, we demonstrate that our method outperforms existing state-of-the-art methods, resulting in superior distillation performance. Longzhen Li, Guang Li 0008, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 2 |
| 2025 | Continual Self-supervised Learning Considering Medical Domain Knowledge in Chest CT ImagesabstractWe propose a novel continual self-supervised learning method (CSSL) considering medical domain knowledge in chest CT images. Our approach addresses the challenge of sequential learning by effectively capturing the relationship between previously learned knowledge and new information at different stages. By incorporating an enhanced dark experience replay (DER) into CSSL and maintaining both diversity and representativeness within the rehearsal buffer of DER, the risk of data interference during pretraining is reduced, enabling the model to learn more richer and robust feature representations. In addition, we incorporate a mixup strategy and feature distillation to further enhance the model’s ability to learn meaningful representations. We validate our method using chest CT images obtained under two different imaging conditions, demonstrating superior performance compared to state-of-the-art methods. Ren Tasai, Guang Li 0008, Ren Togo, Minghui Tang, Takaaki Yoshimura, Hiroyuki Sugimori, Kenji Hirata, Takahiro Ogawa 0001, Kohsuke Kudo, Miki Haseyama |
ICASSP | 2 |
| 2025 | Dataset Distillation Via Vision-Language Category PrototypeabstractDataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational consumption. However, previous DD methods mainly focus on distilling information from images, often overlooking the semantic information inherent in the data. The disregard for context hinders the model's generalization ability, particularly in tasks involving complex datasets, which may result in illogical outputs or the omission of critical objects. In this study, we integrate vision-language methods into DD by introducing text prototypes to distill language information and collaboratively synthesize data with image prototypes, thereby enhancing dataset distillation performance. Notably, the text prototypes utilized in this study are derived from descriptive text information generated by an open-source large language model. This framework demonstrates broad applicability across datasets without pre-existing text descriptions, expanding the potential of dataset distillation beyond traditional image-based approaches. Compared to other methods, the proposed approach generates logically coherent images containing target objects, achieving state-of-the-art validation performance and demonstrating robust generalization. Source code and generated data are available in https://github.com/zou-yawen/Dataset-Distillation-via-Vision-Language-Category-Prototype/ Yawen Zou, Guang Li 0008, Duo Su, Jun Yu 0012, Chao Zhang 0030 |
ICCV | 2 |
| 2025 | Diversity-Driven Generative Dataset Distillation Based on Diffusion Model with Self-Adaptive MemoryabstractDataset distillation enables the training of deep neural networks with comparable performance in significantly reduced time by compressing large datasets into small and representative ones. Although the introduction of generative models has made great achievements in this field, the distributions of their distilled datasets are not diverse enough to represent the original ones, leading to a decrease in downstream validation accuracy. In this paper, we present a diversity-driven generative dataset distillation method based on a diffusion model to solve this problem. We introduce self-adaptive memory to align the distribution between distilled and real datasets, assessing the representativeness. The degree of alignment leads the diffusion model to generate more diverse datasets during the distillation process. Extensive experiments show that our method outperforms existing state-of-the-art methods in most situations, proving its ability to tackle dataset distillation tasks. Mingzhuo Li, Guang Li 0008, Jiafeng Mao, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2025 | Hyperbolic Dataset DistillationabstractTo address the computational and storage challenges posed by large-scale datasets in deep learning, dataset distillation has been proposed to synthesize a compact dataset that replaces the original while maintaining comparable model performance. Unlike optimization-based approaches that require costly bi-level optimization, distribution matching (DM) methods improve efficiency by aligning the distributions of synthetic and original data, thereby eliminating nested optimization. DM achieves high computational efficiency and has emerged as a promising solution. However, existing DM methods, constrained to Euclidean space, treat data as independent and identically distributed points, overlooking complex geometric and hierarchical relationships. To overcome this limitation, we propose a novel hyperbolic dataset distillation method, termed HDD. Hyperbolic space, characterized by negative curvature and exponential volume growth with distance, naturally models hierarchical and tree-like structures. HDD embeds features extracted by a shallow network into the Lorentz hyperbolic space, where the discrepancy between synthetic and original data is measured by the hyperbolic (geodesic) distance between their centroids. By optimizing this distance, the hierarchical structure is explicitly integrated into the distillation process, guiding synthetic samples to gravitate towards the root-centric regions of the original data distribution while preserving their underlying geometric characteristics. Furthermore, we find that pruning in hyperbolic space requires only 20\% of the distilled core set to retain model performance, while significantly improving training stability. Notably, HDD is seamlessly compatible with most existing DM methods, and extensive experiments on different datasets validate its effectiveness. To the best of our knowledge, this is the first work to incorporate the hyperbolic space into the dataset distillation process. The code is available at https://github.com/Guang000/HDD. Guang Li 0008, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
NeurIPS | 2 |
| 2025 | Cross-domain multi-step thinking: Zero-shot fine-grained traffic sign recognition in the wild
Yaozong Gan, Guang Li 0008, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
Knowl. Based Syst. | 2 |
| 2025 | LiDAR-assisted image restoration for extreme low-light conditions
Zhen Wang 0004, Yaozu Wu, Dongyuan Li, Guang Li 0008, Peide Zhu, Renhe Jiang |
Knowl. Based Syst. | 4 |
| 2024 | Cross-Domain Few-Shot In-Context Learning For Enhancing Traffic Sign RecognitionabstractIn this paper, we propose a cross-domain few-shot in-context learning method based on the multimodal large language model (MLLM) for enhancing traffic sign recognition (TSR). We first construct a traffic sign detection network based on Vision Transformer Adapter and an extraction module to extract traffic signs from the original road images. To reduce the dependence on training data and improve the performance stability of cross-country TSR, we introduce a cross-domain few-shot in-context learning method based on the MLLM. To enhance MLLM’s fine-grained recognition ability of traffic signs, the proposed method generates corresponding description texts using template traffic signs. These description texts contain key information about the shape, color, and composition of traffic signs, which can stimulate the ability of MLLM to perceive fine-grained traffic sign categories. By using the description texts, our method reduces the cross-domain differences between template and real traffic signs. Our approach requires only simple and uniform textual indications, without the need for large-scale traffic sign images and labels. We perform comprehensive evaluations on the German traffic sign recognition benchmark dataset, the Belgium traffic sign dataset, and two real-world datasets taken from Japan. The experimental results show that our method significantly enhances the TSR performance. Yaozong Gan, Guang Li 0008, Ren Togo, Keisuke Maeda, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 2 |
| 2024 | Multimodal Low-light Image Enhancement with Depth InformationabstractLow-light image enhancement has been researched several years. However, current image restoration methods predominantly focus on recovering images from RGB images, overlooking the potential of incorporating more modalities. With the advancements in personal handheld devices, we can now easily capture images with depth information using devices such as mobile phones. The integration of depth information into image restoration is a research question worthy of exploration. Therefore, in this paper, we propose a multimodal low-light image enhancement task based on depth information and establish a dataset named LED (Low-light Image Enhanced with Depth Map), consisting of 1,365 samples. Each sample in our dataset includes a low-light image, a normal-light image, and the corresponding depth map. Moreover, for the LED dataset, we design a corresponding multimodal method, which can processes the input images and depth map information simultaneously to generate the predicted normal-light images. Experimental results and detailed ablation studies proves the efficiency of our method which exceeds previous single-modal state-of-the arts methods from relevant field. Zhen Wang 0004, Dongyuan Li, Guang Li 0008, Renhe Jiang |
ACM Multimedia | 3 |
| 2024 | Handling Class Imbalance in Black-Box Unsupervised Domain Adaptation with Synthetic Minority Over-SamplingabstractBlack-box unsupervised domain adaptation (BBUDA) is a challenging task that transfers knowledge from the source domain to the target domain without access to the source data and source model, thus alleviating public concerns about data security. However, BBUDA requires the source model to function as a black-box predictor for the target data, and the pseudo-labels often exhibit class imbalance, which degrades the performance. To tackle this problem, we propose employing the synthetic minority oversampling technique (SMOTE) and adaptive sampling to rebalance data. Given that predictions often contain errors, we first select reliable high-confidence data before using SMOTE to generate synthetic samples for the minority class. Second, we incrementally select high-confidence data from the remaining low-confidence data with an adaptive sampling rate for each class, in which the minority class (with the fewest samples) is assigned a higher sampling rate and the majority class (with the most samples) is assigned a lower sampling rate. The experimental results demonstrate that our method can mitigate the class imbalance and further improve the performance of the target model. Yawen Zou, Chunzhi Gu, Guang Li 0008, Jun Yu 0012, Chao Zhang 0030 |
VCIP | 4 |
| 2024 | Importance-aware adaptive dataset distillation
Guang Li 0008, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
Neural Networks | 1 |
| 2022 | Self-Knowledge Distillation based Self-Supervised Learning for Covid-19 Detection from Chest X-Ray ImagesabstractThe global outbreak of the Coronavirus 2019 (COVID-19) has overloaded worldwide healthcare systems. Computer-aided diagnosis for COVID-19 fast detection and patient triage is becoming critical. This paper proposes a novel self-knowledge distillation based self-supervised learning method for COVID-19 detection from chest X-ray images. Our method can use self-knowledge of images based on similarities of their visual features for self-supervised learning. Experimental results show that our method achieved an HM score of 0.988, an AUC of 0.999, and an accuracy of 0.957 on the largest open COVID-19 chest X-ray dataset. Guang Li 0008, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 1 |
| 2022 | TriBYOL: Triplet BYOL for Self-Supervised Representation LearningabstractThis paper proposes a novel self-supervised learning method for learning better representations with small batch sizes. Many self-supervised learning methods based on certain forms of the siamese network have emerged and received significant attention. However, these methods need to use large batch sizes to learn good representations and require heavy computational resources. We present a new triplet network combined with a triple-view loss to improve the performance of self-supervised representation learning with small batch sizes. Experimental results show that our method can drastically outperform state-of-the-art self-supervised learning methods on several datasets in small-batch cases. Our method provides a feasible solution for self-supervised learning with real-world high-resolution images that uses small batch sizes. Guang Li 0008, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 1 |
| 2022 | Dataset complexity assessment based on cumulative maximum scaled area under Laplacian spectrum
Guang Li 0008, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
Multim. Tools Appl. | 1 |
| 2020 | Soft-Label Anonymous Gastric X-Ray Image DistillationabstractThis paper presents a soft-label anonymous gastric X-ray image distillation method based on a gradient descent approach. The sharing of medical data is demanded to construct high-accuracy computer-aided diagnosis (CAD) systems. However, the large size of the medical dataset and privacy protection are remaining problems in medical data sharing, which hindered the research of CAD systems. The idea of our distillation method is to extract the valid information of the medical dataset and generate a tiny distilled dataset that has a different data distribution. Different from model distillation, our method aims to find the optimal distilled images, distilled labels and the optimized learning rate. Experimental results show that the proposed method can not only effectively compress the medical dataset but also anonymize medical images to protect the patient's private information. The proposed approach can improve the efficiency and security of medical data sharing. Guang Li 0008, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 1 |