VLDB 2026 Research / reviewers in the wild / expert
Jianlong Yuan
dblp:265/2408
· DBLP profile ↗
15ranked-venue papers
5as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video GenerationabstractSora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applications, remains relatively underexplored. To bridge this gap, we propose Mask2DiT, a novel approach that establishes fine-grained, one-to-one alignment between video segments and their corresponding text annotations. Specifically, we introduce a symmetric binary mask at each attention layer within the DiT architecture, ensuring that each text annotation applies exclusively to its respective video segment while preserving temporal coherence across visual tokens. This attention mechanism enables precise segment-level textual-to-visual alignment, allowing the DiT architecture to effectively handle video generation tasks with a fixed number of scenes. To further equip the DiT architecture with the ability to generate additional scenes based on existing ones, we incorporate a segment-level conditional mask, which conditions each newly generated segment on the preceding video segments, thereby enabling auto-regressive scene extension. Both qualitative and quantitative experiments confirm that Mask2DiT excels in maintaining visual consistency across segments while ensuring semantic alignment between each segment and its corresponding text description. Our project page is https://tianhao-qi.github.io/Mask2DiTProject/. Tianhao Qi, Jianlong Yuan, Wanquan Feng, Shancheng Fang, Jiawei Liu 0001, SiYu Zhou 0002, Hongtao Xie 0001, Yongdong Zhang 0001 |
CVPR | 2 |
| 2025 | Generative Pre-trained Autoregressive Diffusion TransformerabstractIn this work, we present GPDiT, a Generative Pre-trained Autoregressive Diffusion Transformer that unifies the strengths of diffusion and autoregressive modeling for long-range video synthesis, within a continuous latent space. Instead of predicting discrete tokens, GPDiT autoregressively predicts future latent frames using a diffusion loss, enabling natural modeling of motion dynamics and semantic consistency across frames. This continuous autoregressive framework not only enhances generation quality but also endows the model with representation capabilities. Additionally, we introduce a lightweight causal attention variant and a parameter-free rotation-based time-conditioning mechanism, improving both the training and inference efficiency. Extensive experiments demonstrate that GPDiT achieves strong performance in video generation quality, video representation ability, and few-shot learning tasks, highlighting its potential as an effective framework for video modeling in continuous space. Yuan Zhang 0022, Zhiying Lu, Haoyang Huang, Jianlong Yuan, Nan Duan 0001 |
NeurIPS | 7 |
| 2024 | FAKD: Feature Augmented Knowledge Distillation for Semantic SegmentationabstractIn this work, we explore data augmentations for knowledge distillation on semantic segmentation. Due the capacity gap, small-sized student networks struggle to discover the discriminative feature space learned by a powerful teacher. Image-level augmentations allow the student to better imitate the teacher by providing extra outputs. However, existing distillation frameworks only augment a limited number of samples, which restricts the learning of a student. Inspired by the recent progress on semantic directions on feature space, this work proposes a feature-level augmented knowledge distillation (FAKD) which infinitely augments features along a semantic direction for optimal knowledge transfer. Furthermore, we introduce novel surrogate loss functions to distill the teacher’s knowledge from an infinite number of samples. The surrogate loss is an upper bound of the expected distillation loss over infinite augmented samples. Extensive experiments on four semantic segmentation benchmarks demonstrate that the proposed method boosts the performance of current knowledge distillation methods without any significant overhead. The code will be released at FAKD. Jianlong Yuan, Minh Hieu Phan, Liyang Liu |
WACV | 1 |
| 2024 | Equity in Unsupervised Domain Adaptation by Nuclear Norm MaximizationabstractNuclear norm maximization has shown the power to enhance the transferability of unsupervised domain adaptation model (UDA) in an empirical scheme. In this paper, we identify a new property termedequity, which indicates the balance degree of predicted classes, to demystify the efficacy of nuclear norm maximization for UDA theoretically. With this in mind, we offer a new discriminability-and-equity maximization paradigm built on squares loss, such that predictions are equalized explicitly. To verify its feasibility and flexibility, two new losses termed Class Weighted Squares Maximization (CWSM) and Normalized Squares Maximization (NSM), are proposed to maximize both predictive discriminability and equity, from the class level and the sample level, respectively. Importantly, we theoretically relate these two novel losses (i.e., CWSM and NSM) to the equity maximization under mild conditions, and empirically suggest the importance of the predictive equity in UDA. Moreover, it is very efficient to realize the equity constraints in both losses. Experiments of cross-domain image classification on three popular benchmark datasets show that both CWSM and NSM contribute to outperforming the corresponding counterparts. Mengzhu Wang, Shanshan Wang 0008, Xun Yang 0001, Jianlong Yuan, Wenju Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Inter-Class and Inter-Domain Semantic Augmentation for Domain GeneralizationabstractThe domain generalization approach seeks to develop a universal model that performs well on unknown target domains with the aid of diverse source domains. Data augmentation has proven to be an effective method to enhance domain generalization in computer vision. Recently, semantic-level based data augmentation has yielded remarkable results. However, these methods focus on sampling semantic directions on feature space from intra-class and intra-domain, limiting the diversity of the source domain. To address this issue, we propose a novel approach called Inter-Class and Inter-Domain Semantic Augmentation (CDSA) for domain generalization. We first introduce a sampling-based method called CrossSmooth to obtain semantic directions from inter-class. Then, CrossVariance obtains the styles of different domains by sampling semantic directions. Our experiments on four well-known domain generalization benchmark datasets (Digits-DG, PACS, Office-Home, and DomainNet) demonstrate the effectiveness of our approach. We also validate our approach on commonly-used semantic segmentation datasets, namely GTAV, SYNTHIA, Cityscapes, Mapillary, and BDDS which also show significant improvements. Mengzhu Wang, Yuehua Liu, Jianlong Yuan, Shanshan Wang 0008, Zhibin Wang 0004, Wei Wang 0335 |
IEEE Trans. Image Process. | 3 |
| 2023 | Efficient Mask Correction for Click-Based Interactive Image SegmentationabstractThe goal of click-based interactive image segmentation is to extract target masks with the input of positive/negative clicks. Every time a new click is placed, existing methods run the whole segmentation network to obtain a corrected mask, which is inefficient since several clicks may be needed to reach satisfactory accuracy. To this end, we propose an efficient method to correct the mask with a lightweight mask correction network. The whole network remains a low computational cost from the second click, even if we have a large backbone. However, a simple correction network with limited capacity is not likely to achieve comparable performance with a classic segmentation network. Thus, we propose a click-guided self-attention module and a click-guided correlation module to effectively exploits the click information to boost performance. First, several tem-plates are selected based on the semantic similarity with click features. Then the self-attention module propagates the template information to other pixels, while the correlation module directly uses the templates to obtain target out-lines. With the efficient architecture and two click-guided modules, our method shows preferable performance and efficiency compared to existing methods. The code will be released at https://github.com/feiaxyt/EMC-Click. Jianlong Yuan, Zhibin Wang 0004, Fan Wang 0019 |
CVPR | 2 |
| 2023 | Foundation Model Drives Weakly Incremental Learning for Semantic SegmentationabstractModern incremental learning for semantic segmentation methods usually learn new categories based on dense annotations. Although achieve promising results, pixel-by-pixel labeling is costly and time-consuming. Weakly incremental learning for semantic segmentation (WILSS) is a novel and attractive task, which aims at learning to segment new classes from cheap and widely available image-level labels. Despite the comparable results, the image-level labels can not provide details to locate each segment, which limits the performance of WILSS. This inspires us to think how to improve and effectively utilize the supervision of new classes given image-level labels while avoiding forgetting old ones. In this work, we propose a novel and data-efficient frame-work for WILSS, named FMWISS. Specifically, we propose pre-training based co-segmentation to distill the knowledge of complementary foundation models for generating dense pseudo labels. We further optimize the noisy pseudo masks with a teacher-student architecture, where a plug-in teacher is optimized with a proposed dense contrastive loss. Moreover, we introduce memory-based copy-paste augmentation to improve the catastrophic forgetting problem of old classes. Extensive experiments on Pascal VOC and COCO datasets demonstrate the superior performance of our framework, e.g., FMWISS achieves 70.7% and 73.3% in the 15–5 VOC setting, outperforming the state-of-the-art method by 3.4% and 6.1%, respectively. Chaohui Yu, Qiang Zhou 0001, Jingliang Li, Jianlong Yuan, Zhibin Wang 0004, Fan Wang 0019 |
CVPR | 4 |
| 2023 | K2NN: Self-Supervised Learning with Hierarchical Nearest Neighbors for Remote SensingabstractSelf-supervised learning aims to learn applicable pre-trained models from massive unlabeled data. Besides image-level pretext tasks, many recent pixel-level studies have been pro-posed to learn dense information in each image. However, most of those methods focus on obtaining pair of matched patches from the same image with different augmentation. At the same time, little effort is devoted to exploiting matched patches from different images. In this work, we develop a novel pixel-level task that leverages an ensemble of nearest neighbors from multiple images to explore diverse objects in each image, especially for remote sensing data. Besides, a sampling strategy with a submodular function is adopted to efficiently update the memory bank consisting of patches. The extensive experiments on remote sensing data confirm the effectiveness of our method. Jianlong Yuan, Yuanhong Xu |
ICASSP | 1 |
| 2023 | UniNeXt: Exploring A Unified Architecture for Vision RecognitionabstractVision Transformers have shown great potential in computer vision tasks. Most recent works have focused on elaborating the spatial token mixer for performance gains. However, we observe that a well-designed general architecture can significantly improve the performance of the entire backbone, regardless of which spatial token mixer is equipped. In this paper, we propose UniNeXt, an improved general architecture for the vision backbone. To verify its effectiveness, we instantiate the spatial token mixer with various typical and modern designs, including both convolution and attention modules. Compared with the architecture in which they are first proposed, our UniNeXt architecture can steadily boost the performance of all the spatial token mixers, and narrows the performance gap among them. Surprisingly, our UniNeXt equipped with naive local window attention even outperforms the previous state-of-the-art. Interestingly, the ranking of these spatial token mixers also changes under our UniNeXt, suggesting that an excellent spatial token mixer may be stifled due to a suboptimal general architecture, which further shows the importance of the study on the general architecture of vision backbone. Code is available at UniNeXt. Fangjian Lin, Jianlong Yuan, Sitong Wu, Fan Wang 0019, Zhibin Wang 0004 |
ACM Multimedia | 2 |
| 2023 | Mixture-of-Experts Learner for Single Long-Tailed Domain GeneralizationabstractDomain generalization (DG) refers to the task of training a model on multiple source domains and test it on a different target domain with different distribution. In this paper, we address a more challenging and realistic scenario known as Single Long-Tailed Domain Generalization, where only one source domain is available and the minority class in this domain has an abundance of instances in other domains. To tackle this task, we propose a novel approach called Mixture-of-Experts Learner for Single Long-Tailed Domain Generalization (MoEL), which comprises two key strategies. The first strategy is a simple yet effective data augmentation technique that leverages saliency maps to identify important regions on the original images and preserves these regions during augmentation. The second strategy is a new skill-diverse expert learning approach that trains multiple experts from a single long-tailed source domain and leverages mutual learning to aggregate their learned knowledge for the unknown target domain. We evaluate our method on various benchmark datasets, including Digits-DG, CIFAR-10-C, PACS, and DomainNet, and demonstrate its superior performance compared to previous single domain generalization methods. Additionally, the ablation study is also conducted to illustrate the inner workings of our approach. Mengzhu Wang, Jianlong Yuan, Zhibin Wang 0004 |
ACM Multimedia | 2 |
| 2023 | Semi-supervised Semantic Segmentation with Mutual Knowledge DistillationabstractConsistency regularization has been widely studied in recent semi- supervised semantic segmentation methods, and promising per- formance has been achieved. In this work, we propose a new con- sistency regularization framework, termed mutual knowledge dis- tillation (MKD), combined with data and feature augmentation. We introduce two auxiliary mean-teacher models based on consis- tency regularization. More specifically, we use the pseudo-labels generated by a mean teacher to supervise the student network to achieve a mutual knowledge distillation between the two branches. In addition to using image-level strong and weak augmentation, we also discuss feature augmentation. This involves considering various sources of knowledge to distill the student network. Thus, we can significantly increase the diversity of the training samples. Experiments on public benchmarks show that our framework out- performs previous state-of-the-art (SOTA) methods under various semi-supervised settings. Code is available at https://github.com/jianlong-yuan/semi-mmseg. Jianlong Yuan, Jinchao Ge, Zhibin Wang 0004, Yifan Liu 0001 |
ACM Multimedia | 1 |
| 2022 | Semantic Data Augmentation based Distance Metric Learning for Domain GeneralizationabstractDomain generalization (DG) aims to learn a model on one or more different but related source domains that could be generalized into an unseen target domain. Existing DG methods try to prompt the diversity of source domains for the model's generalization ability, while they may have to introduce auxiliary networks or striking computational costs. On the contrary, this work applies the implicit semantic augmentation in feature space to capture the diversity of source domains. Concretely, an additional loss function of distance metric learning (DML) is included to optimize the local geometry of data distribution. Besides, the logits from cross entropy loss with infinite augmentations is adopted as input features for the DML loss in lieu of the deep features. We also provide a theoretical analysis to show that the logits can approximate the distances defined on original features well. Further, we provide an in-depth analysis of the mechanism and rational behind our approach, which gives us a better understanding of why leverage logits in lieu of features can help domain generalization. The proposed DML loss with the implicit augmentation is incorporated into a recent DG method, that is, Fourier Augmented Co-Teacher framework (FACT). Meanwhile, our method also can be easily plugged into various DG methods. Extensive experiments on three benchmarks (Digits-DG, PACS and Office-Home) have demonstrated that the proposed method is able to achieve the state-of-the-art performance. Mengzhu Wang, Jianlong Yuan, Qi Qian 0001, Zhibin Wang 0004, Hao Li 0030 |
ACM Multimedia | 2 |
| 2022 | Contrastive Haze-Aware Learning for Dynamic Remote Sensing Image DehazingabstractImage dehazing methods aim to recover a clear image from its hazy counterpart. While various dehazing methods have been proposed, their performance on real-world remote sensing (RS) images remains unsatisfying. A key reason is that the complex weather and imaging conditions (e.g., large fields of view) cause the haze condition to dramatically change in different images, while most existing methods fail to flexibly adapt their dehazing model to the specific haze condition in each image. To mitigate this problem, we present a contrastive haze-aware learning based dynamic dehazing method which demonstrates two aspects of advantage. On one hand, a contrastive clustering scheme is utilized to learn the image-wise haze representation using a set of real-world hazy images in an unsupervised manner, which enables identifying and discriminating the specific haze condition in each given hazy image. On the other hand, with the learned haze representation, a parameter generator can produce haze-aware parameters to dynamically construct a dehazing model for the given hazy image, which empowers us to adaptively dehaze the image based on its specific haze condition and thus improves the generalization ability. In addition, a new contrastive loss defined based on the learned haze representation is further utilized for model training and leads to better performance. To demonstrate the effectiveness of the proposed method, we evaluate it on two benchmark RS image datasets including various real-world hazy images, and observe obviously superiority over other state-of-the-art competitors. Jiangtao Nie, Wei Wei 0008, Lei Zhang 0054, Jianlong Yuan, Zhibin Wang 0004, Hao Li 0030 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | A Simple Baseline for Semi-supervised Semantic Segmentation with Strong Data Augmentation*abstractRecently, significant progress has been made on semantic segmentation. However, the success of supervised semantic segmentation typically relies on a large amount of labeled data, which is time-consuming and costly to obtain. Inspired by the success of semi-supervised learning methods for image classification, here we propose a simple yet effective semi-supervised learning framework for semantic segmentation. We demonstrate that the devil is in the details: a set of simple designs and training techniques can collectively improve the performance of semi-supervised semantic segmentation significantly. Previous works [3], [25] fail to effectively employ strong augmentation in pseudo-label learning, as the large distribution disparity caused by strong augmentation harms the batch nor-malization statistics. We design a new batch normalization, namely distribution-specific batch normalization (DSBN) to address this problem and show the importance of strong augmentation for semantic segmentation. Moreover, we design a self-correction loss, which is effective in terms of noise resistance. We conduct a series of ablation studies to show the effectiveness of each component. Our method achieves state-of-the-art results in the semi-supervised settings on the Cityscapes and Pascal VOC datasets. Jianlong Yuan, Yifan Liu 0001, Chunhua Shen, Zhibin Wang 0004, Hao Li 0030 |
ICCV | 1 |
| 2020 | Multi Receptive Field Network for Semantic SegmentationabstractSemantic segmentation is one of the key tasks in computer vision, which is to assign a category label to each pixel in an image. Despite significant progress achieved recently, most existing methods still suffer from two challenging issues: l)the size of objects and stuff in an image can be very diverse, demanding for incorporating multi-scale features into the fully convolutional networks (FCNs); 2) the pixels close to or at the boundaries of object/stuff are hard to classify due to the intrinsic weakness of convolutional networks. To address the first issue, we propose a new Multi-Receptive Field Module (MRFM), explicitly taking multi-scale features into account. For the second issue, we design an edge-aware loss which is effective in distinguishing the boundaries of object/stuff. With these two designs, our Multi Receptive Field Network achieves new state-of-the-art results on two widely-used semantic segmentation benchmark datasets. Specifically, we achieve a mean IoU of 83.0% on the Cityscapes dataset and 88.4% mean IoU on the Pascal VOC2012 dataset. Jianlong Yuan, Zelu Deng, Zhenbo Luo |
WACV | 1 |