VLDB 2026 Research / reviewers in the wild / expert
Zongyao Li 0004
dblp:244/7588-4
· DBLP profile ↗
9ranked-venue papers
8as first author
6since 2021 · last 2024
0000-0002-3300-1806ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Parallel Transformer Framework for Video Moment RetrievalabstractIn the realm of video understanding, Video Moment Retrieval (VMR) is an important yet challenging task that aims to locate the boundary of a moment of interest within a long untrimmed video. Existing VMR methods often focus on the visual content extracted from the video only (or frame sequences), however, the rich semantic information at the object level that describes the image's content has not been explored yet. To overcome those limitations, we propose PaTF, an attention-based Parallel Transformer Framework that enriches the feature representations by exploring both low-level visual cues and high-level relational contexts of video-query pairs. Our framework consists of two parallel transformers: one for the visual-textual stream and the other for the semantic-textual stream. The visual-textual stream extracts the links between global visual features and textual information, while the semantic-textual stream emphasises the relations between objects via scene graph representations. Furthermore, our comprehensive experiment conducted on the Charades-STA dataset demonstrates that the proposed framework outperforms the state-of-the-art methods by a large margin, roughly 5% and 7% at Recall@1 with IoU = 0.5 and IoU = 0.7, respectively. Thao-Nhu Nguyen, Zongyao Li 0004, Satoshi Yamazaki, Jianquan Liu, Cathal Gurrin |
ICMR | 2 |
| 2022 | Union-Set Multi-source Model Adaptation for Semantic Segmentation
Zongyao Li 0004, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ECCV (29) | 1 |
| 2022 | Divergence-Guided Feature Alignment for Cross-Domain Object DetectionabstractDomain shift causes performance drop in cross-domain object detection. To alleviate the domain shift, a prevailing approach is global feature alignment with adversarial learning. However, such simple feature alignment has defects of unawareness of fore-ground/background regions and well-aligned/poorly-aligned regions. To remedy the defects, in this paper, we propose a novel divergence-guided feature alignment method for cross-domain object detection. Specifically, we generate source-like images of the target domain and seek cues of foreground regions and poorly-aligned regions from prediction divergence of the source-like and original images. The feature alignment is guided by the divergence maps and consequently results in adaptation performance superior to alignment unaware of the cues. Different from most previous studies focusing on two-stage object detection, this paper is devoted to adapting one-stage object detectors which have simpler and faster inference. We validated the effectiveness of our method by conducting experiments in cross-weather, cross-camera, and synthetic-to-real adaptation scenarios. Zongyao Li 0004, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 1 |
| 2022 | Improving Model Adaptation for Semantic Segmentation by Learning Model-Invariant Features with Multiple Source-Domain ModelsabstractIn this paper, we focus on a problem remaining to be studied: multi-source model adaptation, which is derived from multi-source unsupervised domain adaptation and replaces the source-domain data with source-domain pre-trained models. Pre-trained models are always easier to share than training data so that multiple source-domain models are available in many practical scenarios. Therefore, the problem setting of multi-source domain adaptation is practical in real-world applications. In this setting, we challenge the task of semantic segmentation which is difficult also in the traditional unsupervised domain adaptation due to the pixel-level knowledge transfer. Our method takes full advantage of the multiple source-domain models by learning model-invariant features, which aims to obtain target-domain features with similar distributions from the models pre-trained in different source domains. The adaptation models trained with the model-invariant feature learning benefit from the diversity of the source-domain models and can thus produce more generalizable features to the target domain. Experimental results in several adaptation settings validate the effectiveness and superiority of our method. Zongyao Li 0004, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 1 |
| 2022 | Learning intra-domain style-invariant representation for unsupervised domain adaptation of semantic segmentation
Zongyao Li 0004, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
Pattern Recognit. | 1 |
| 2021 | Semantic-Aware Unpaired Image-to-Image Translation for Urban Scene ImagesabstractUnpaired image-to-image (I2I) translation methods have been developed for several years. Present methods do not take into consideration semantic information of the original image, which may perform well on simple datasets of uncomplicated scenes, however, fail in complex datasets of scenes involving abundant objects, such as urban scenes. To tackle this problem, in this paper, we reasonably modify the previous problem setting and present a novel semantic-aware method. Specifically, in training, we use additional semantic label maps of training images, while in the test, no labels are required. We originally adopt a semantic knowledge distillation strategy to acquire semantic information from the labels and construct a particular normalization layer to introduce semantic information. Being aware of the pixel-level semantic information, our method can realize better I2I translation than the previous methods. Experiments are conducted on benchmark datasets of urban scenes to validate the effectiveness of our method. Zongyao Li 0004, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 1 |
| 2020 | Unsupervised Domain Adaptation for Semantic Segmentation with Symmetric Adaptation ConsistencyabstractUnsupervised domain adaptation, which leverages label information from other domains to solve tasks on a domain without any labels, can alleviate the problem of the scarcity of labels and expensive labeling costs faced by supervised semantic segmentation. In this paper, we utilize adversarial learning and semi-supervised learning simultaneously to solve the task of unsupervised domain adaptation in semantic segmentation. We propose a new approach that trains two segmentation models with the adversarial learning symmetrically and further introduces the consistency between the outputs of the two models into the semi-supervised learning to improve the accuracy of pseudo labels which significantly affect the final adaptation performance. We achieve state-of-the-art semantic segmentation performance on the GTA5-to-Cityscapes scenario, a widely used benchmark setting in unsupervised domain adaptation. Zongyao Li 0004, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 1 |
| 2020 | Variational Autoencoder Based Unsupervised Domain Adaptation For Semantic SegmentationabstractUnsupervised domain adaptation, which transfers supervised knowledge from a labeled domain to an unlabeled domain, remains a tough problem in the field of computer vision, especially for semantic segmentation. Some methods inspired by adversarial learning and semi-supervised learning have been developed for unsupervised domain adaptation in semantic segmentation and achieved outstanding performances. In this paper, we propose a novel method for this task. Like adversarial learning-based methods using a discriminator to align the feature distributions from different domains, we employ a variational autoencoder to get to the same destination but in a non-adversarial manner. Since the two approaches are compatible, we also integrate an adversarial loss into our method. By further introducing pseudo labels, our method can achieve state-of-the-art performances on two benchmark adaptation scenarios, GTA5-to-CITYSCAPES and SYNTHIA-to-CITYSCAPES. Zongyao Li 0004, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ICIP | 1 |
| 2019 | Semi-Supervised Learning Based on Tri-Training for Gastritis Classification using Gastric X-ray ImagesabstractThis paper presents a method of semi-supervised learning based on tri-training for gastritis classification using gastric X-ray images. The proposed method is constructed based on the tri-training architecture, and the strategies of label smoothing regularization and random erasing augmentation are utilized in the method to enhance the performance. Although the task of gastritis classification is challenging, we report that the proposed semi-supervised learning method using only a small number of labeled data achieves 0.888 harmonic mean of sensitivity and specificity on test data composed of 615 patients. Zongyao Li 0004, Ren Togo, Takahiro Ogawa 0001, Miki Haseyama |
ISCAS | 1 |