EDBT 2026 Demo / reviewers in the wild / expert
Yutian Lin
dblp:198/1146
· DBLP profile ↗
24ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-0643-0533ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 9 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Condition-Guided Diffusion for Multi-Modal Pedestrian Trajectory Prediction Incorporating Intention and Interaction PriorsabstractPedestrian behavior exhibits inherent multi-modality, necessitating predictions that balance accuracy and diversity to adapt effectively to various complex scenarios. However, conventional noise addition in diffusion models is often aimless and unguided, leading to redundant noise reduction steps and the generation of uncontrollable samples. To address these issues, we propose a Prior Condition-Guided Diffusion Model (CGD-TraP) for multi-modal pedestrian trajectory prediction. Instead of directly adding Gaussian noise to trajectories at each timestep during the forward process, our approach leverages internal intention and external interaction to guide noise estimation. Specifically, we design two specialized modules to extract and aggregate intention and interaction features. These features are then adaptively fused through a spatial-temporal fusion based on selective state space, which estimates a controllable noisy trajectory distribution. By optimizing the noise addition process in a more controlled and efficient manner, our method ensures that the denoising process is effectively guided, resulting in predictions that are both accurate and diverse. Extensive experiments on the ETH-UCY, SDD, and NBA datasets demonstrate that CGD-TraP surpasses state-of-the-art diffusion-based and other generative methods, achieving superior efficiency, accuracy, and diversity. Yanghong Liu, Xingping Dong, Yutian Lin, Mang Ye, Kaihao Zhang, Bo Du 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | MAC: A Benchmark for Multiple Attributes Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to learn semantic primitives (attributes and objects) from seen compositions and recognize unseen attribute-object compositions. Existing CZSL datasets focus on single attributes, neglecting the fact that objects naturally exhibit multiple interrelated attributes. Their narrow attribute scope and single attribute labeling introduce annotation biases, misleading the learning of attribute and causing inaccurate evaluation. To address these issues, we introduce the Multi-Attribute Composition (MAC) dataset, encompassing 22,838 images and 17,627 compositions with comprehensive attribute annotations. MAC shows a complex relationship between attributes and objects, with each attribute type linked to an average of 82.2 object classes, and each object type associated with 31.4 attribute classes. Based on MAC, we propose multi-attribute compositional zero-shot learning that requires deeper semantic understanding and advanced attribute associations, establishing a more realistic and challenging benchmark for CZSL. We propose Multi-attribute Visual-Primitive Integrator (MVP-Integrator), a robust baseline for multi-attribute CZSL, which disentangles semantic primitives and performs effective visual-primitive association. Experiments demonstrate that MVP-Integrator significantly outperforms existing CZSL methods on MAC with improved efficiency. The dataset will be released publicly upon publication. Yutian Lin, Sibei Yang, Yu Wu 0011, Bo Du 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Penalizing Boundary Activation for Object Completeness in Diffusion Models
Haoyang Xu, Sibei Yang, Yutian Lin |
ICCV | 4 |
| 2025 | Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillationabstractVideo-to-Audio (V2A) Generation achieves significant progress and plays a crucial role in film and video post-production. However, current methods overlook the cinematic language, a critical component of artistic expression in filmmaking. As a result, their performance deteriorates in scenarios where Foley targets are only partially visible. To address this challenge, we propose a simple self-distillation approach to extend V2A models to cinematic language scenarios. By simulating the cinematic language variations, the student model learns to align the video features of training pairs with the same audio-visual correspondences, enabling it to effectively capture the associations between sounds and partial visual information. Our method not only achieves impressive improvements under partial visibility across all evaluation metrics, but also enhances performance on the large-scale V2A dataset, VGGSound. Feizhen Huang, Yu Wu 0011, Yutian Lin, Bo Du 0001 |
IJCAI | 3 |
| 2025 | Accident Anticipation via Temporal Occurrence PredictionabstractAccident anticipation aims to predict potential collisions in an online manner, enabling timely alerts to enhance road safety. Existing methods typically predict frame-level risk scores as indicators of hazard. However, these approaches rely on ambiguous binary supervision—labeling all frames in accident videos as positive—despite the fact that risk varies continuously over time, leading to unreliable learning and false alarms. To address this, we propose a novel paradigm that shifts the prediction target from current-frame risk scoring to directly estimating accident scores at multiple future time steps (e.g., 0.1s–2.0s ahead), leveraging precisely annotated accident timestamps as supervision. Our method employs a snippet-level encoder to jointly model spatial and temporal dynamics, and a Transformer-based temporal decoder that predicts accident scores for all future horizons simultaneously using dedicated temporal queries. Furthermore, we introduce a refined evaluation protocol that reports Time-to-Accident (TTA) and recall—evaluated at multiple pre-accident intervals (0.5s, 1.0s, and 1.5s)—only when the false alarm rate (FAR) remains within an acceptable range, ensuring practical relevance. Experiments show that our method achieves superior performance in both recall and TTA under realistic FAR constraints. Project page: https://happytianhao.github.io/TOP/ Yiyang Zou, Zihao Mao, Peilun Xiao, Hongda Yang, Tracy Li, Yutian Lin |
NeurIPS | 10 |
| 2025 | Visible-infrared person re-identification via patch-mixed cross-modality learning
Yutian Lin, Bo Du 0001 |
Pattern Recognit. | 2 |
| 2025 | MixIR: Mixing Input and Representations for Contrastive LearningabstractRecently, contrastive learning has shown significant progress in learning visual representations from unlabeled data. The core idea is training the backbone to be invariant to different augmentations of an instance. While most methods only maximize the feature similarity between two augmented data, we further generate more challenging training samples and force the model to keep predicting aggregated representation on these hard samples. In this article, we propose MixIR, a mixture-based approach upon the traditional Siamese network. On the one hand, we input two augmented images of an instance to the backbone and obtain the aggregated representation by performing an elementwise maximum of two features. On the other hand, we take the mixture of these augmented images as input and expect the model prediction to be close to the aggregated representation. In this way, the model could access more variant data samples of an instance and keep predicting invariant representations for them. Thus, the learned model is more discriminative compared with previous contrastive learning methods. Extensive experiments on large-scale datasets show that MixIR steadily improves the baseline and achieves competitive results with state-of-the-art methods. Our code is available at https://github.com/happytianhao/MixIR. Yutian Lin, Bo Du 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Omni-Q: Omni-Directional Scene Understanding for Unsupervised Visual GroundingabstractUnsupervised visual grounding methods alleviate the issue of expensive manual annotation of image-query pairs by generating pseudo-queries. However, existing methods are prone to confusing the spatial relationships between objects and rely on designing complex prompt modules to gener-ate query texts, which severely impedes the ability to gener-ate accurate and comprehensive queries due to ambiguous spatial relationships and manually-defined fixed templates. To tackle these challenges, we propose a omni-directional language query generation approach for unsupervised visual grounding named Omni-Q. Specifically, we develop a 3D spatial relation module to extend the 2D spatial representation to 3D, thereby utilizing 3D location information to accurately determine the spatial position among objects. Besides, we introduce a spatial graph module, leveraging the power of graph structures to establish accurate and diverse object relationships and thus enhancing the flexibility of query generation. Extensive experiments on five public benchmark datasets demonstrate that our method significantly outperforms existing state-of-the-art unsupervised methods by up to 16.17%. In addition, when applied in the supervised setting, our method can freely save up to 60% human annotations without a loss of performance. Yutian Lin, Yu Wu 0011 |
CVPR | 2 |
| 2024 | Improving Bird's Eye View Semantic Segmentation by Task DecompositionabstractSemantic segmentation in bird's eye view (BEV) plays a crucial role in autonomous driving. Previous methods usually follow an end-to-end pipeline, directly predicting the BEV segmentation map from monocular RGB inputs. However, the challenge arises when the RGB inputs and BEV targets from distinct perspectives, making the direct point-to-point predicting hard to optimize. In this paper, we decompose the original BEV segmentation task into two stages, namely BEV map reconstruction and RGB-BEV feature alignment. In the first stage, we train a BEV autoencoder to reconstruct the BEV segmentation maps given cor-rupted noisy latent representation, which urges the decoder to learn fundamental knowledge of typical BEV patterns. The second stage involves mapping RGB input images into the BEV latent space of the first stage, directly optimizing the correlations between the two views at the feature level. Our approach simplifies the complexity of combining perception and generation into distinct steps, equipping the model to handle intricate and challenging scenes effectively. Besides, we propose to transform the BEV segmentation map from the Cartesian to the polar coordinate system to establish the column-wise correspondence between RGB images and BEV maps. Moreover, our method requires neither multi-scale features nor camera intrinsic parameters for depth estimation and saves computational overhead. Extensive experiments on nuScenes and Argoverse show the effectiveness and efficiency of our method. Code is available at https://github.com/happytianhao/TaDe. Yongcan Chen, Yu Wu 0011, Bo Du 0001, Peilun Xiao, Hongda Yang, Guozhen Li, Yi Yang 0001, Yutian Lin |
CVPR | 11 |
| 2024 | DifTraj: Diffusion Inspired by Intrinsic Intention and Extrinsic Interaction for Multi-Modal Trajectory Prediction
Yanghong Liu, Xingping Dong, Yutian Lin, Mang Ye |
IJCAI | 3 |
| 2024 | Toward Real Ultra Image Segmentation: Leveraging Surrounding Context to Cultivate General Segmentation ModelabstractExisting ultra image segmentation methods suffer from two major challenges, namely the scalability issue (i.e. they lack the stability and generality of standard segmentation models, as they are tailored to specific datasets), and the architectural issue (i.e. they are incompatible with real-world ultra image scenes, as they compromise between image size and computing resources).
To tackle these issues, we revisit the classic sliding inference framework, upon which we propose a Surrounding Guided Segmentation framework (SGNet) for ultra image segmentation.
The SGNet leverages a larger area around each image patch to refine the general segmentation results of local patches.
Specifically, we propose a surrounding context integration module to absorb surrounding context information and extract specific features that are beneficial to local patches. Note that, SGNet can be seamlessly integrated to any general segmentation model.
Extensive experiments on five datasets demonstrate that SGNet achieves competitive performance and consistent improvements across a variety of general segmentation models, surpassing the traditional ultra image segmentation methods by a large margin. Yutian Lin, Yu Wu 0011, Bo Du 0001 |
NeurIPS | 2 |
| 2024 | Promote knowledge mining towards open-world semi-supervised learning
Yutian Lin, Yu Wu 0011, Bo Du 0001 |
Pattern Recognit. | 2 |
| 2024 | Reliable Cross-Camera Learning in Random Camera Person Re-IdentificationabstractMost existing person re-identification (Re-ID) methods rely on high-cost manual annotations. To overcome the applicable issue, we focus on a novel semi-supervised Re-ID without cross-camera annotations, which we call random camera supervised person Re-ID (RCS). It is beneficial to real-world application, since a short-time and cheap annotation is conductive to rapid deployment of person re-ID. But only a small proportion of identities are annotated under a random camera, which is extremely challenging for Re-ID cross-camera pairing without cross-camera labeling or a labeled image for each identity. Towards reliable cross-camera learning, we propose a random camera guided framework (RCG) that can fully make use of the few labeled images with promising performance. RCG has two components: 1) Different from other complex methods to improve the clustering accuracy, Random camera guided clustering is adopted to mine cross-camera images of each identity, where the few labeled data helps to guide the simple but effective cluster split and combination. 2) Network learning under RCS is conducted with cluster-wise and camera-wise contrastive learning, where we deal with the camera variance in subgroup unit innovatively and further emphasis the importance of the labeled images. Extensive experiments on three large-scale Re-ID datasets show that our proposed approach not only outperforms state-of-the-art methods by a large margin, but achieves better performance with less annotation and more flexible RCS setting. Zhengqi Liu, Yutian Lin, Bo Du 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Revisit Weakly-Supervised Audio-Visual Video Parsing from the Language PerspectiveabstractWe focus on the weakly-supervised audio-visual video parsing task (AVVP), which aims to identify and locate all the events in audio/visual modalities. Previous works only concentrate on video-level overall label denoising across modalities, but overlook the segment-level label noise, where adjacent video segments (i.e., 1-second video clips) may contain different events. However, recognizing events on the segment is challenging because its label could be any combination of events that occur in the video. To address this issue, we consider tackling AVVP from the language perspective, since language could freely describe how various events appear in each segment beyond fixed labels. Specifically, we design language prompts to describe all cases of event appearance for each video. Then, the similarity between language prompts and segments is calculated, where the event of the most similar prompt is regarded as the segment-level label. In addition, to deal with the mislabeled segments, we propose to perform dynamic re-weighting on the unreliable segments to adjust their labels. Experiments show that our simple yet effective approach outperforms state-of-the-art methods by a large margin. Yu Wu 0011, Bo Du 0001, Yutian Lin |
NeurIPS | 4 |
| 2023 | Privacy-Protected Person Re-Identification via Virtual SamplesabstractMost person re-identification (re-ID) approaches are based on representation learning of pedestrian images, which assume that the person’s appearance captured by cameras in the target is fully available. However, the exposure of appearance could cause serious privacy leakages. To address this issue, we focus on a new privacy-protected person re-ID task where the person’s appearance is unavailable for training. We first overcome the dilemma of lacking real person images by utilizing the virtual pedestrian samples (e.g., PersonX). Then, we introduce a composition of data augmentations to simulate real conditions, where the learned model is transferred to the real target in a black way. Specifically, the background behind, surrounding illumination, pose, scale, and attributes of pedestrians from the target scene, irrelevant to privacy, are utilized to generate virtual images. With above privacy-irrelevant information, we propose a Translation-Rendering-Sampling (TRS) framework to generate images toward the distribution of the real-world dataset (including background, pose, attribute,etc). Extensive experiments are conducted on several realistic re-ID datasets excluding the person’s appearance. The experiments show that our method outperforms the baseline significantly as well as some transfer learning methods. Yutian Lin, Zheng Wang 0007, Bo Du 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Unsupervised Person Re-Identification With Stochastic Training StrategyabstractUnsupervised person re-identification (re-ID) has attracted increasing research interests because of its scalability and possibility for real-world applications. State-of-the-art unsupervised re-ID methods usually follow a clustering-based strategy, which generates pseudo labels by clustering and maintains a memory to store instance features and represent the centroid of the clusters for contrastive learning. This approach suffers two problems. First, the centroid generated by unsupervised learning may not be a perfect prototype. Forcing images to get closer to the centroid emphasizes the result of clustering, which could accumulate clustering errors during iterations. Second, previous instance memory based methods utilize features updated at different training iterations to represent one centroid, these features are inconsistent due to the change of encoder. To this end, we propose an unsupervised re-ID approach with a stochastic learning strategy. Specifically, we adopt a stochastic updated memory, where a random instance from a cluster is used to update the cluster-level memory for contrastive learning. In this way, the relationship between randomly selected pair of images are learned to avoid the training bias caused by unreliable pseudo labels. By picking a sole last seen sample to directly update each cluster center, the stochastic memory is also always up-to-date for classifying to keep the consistency. Besides, to relieve the issue of camera variance, a unified distance matrix is proposed during clustering, where the distance bias from different camera domains is reduced and the variances of identities are emphasized. Our proposed method outperforms the state-of-the-arts in all the common unsupervised and UDA re-ID tasks. The code will be available at https://github.com/lithium770/Unsupervised-Person-re-ID-with-Stochastic-Training-Strategy. Yutian Lin, Bo Du 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Dual-Regularization Complementary Learning for Image ClassificationabstractDeep supervised learning has achieved great success in image classification. However, most existing methods require high-quality labeled images, which are not easy to obtain. In this paper, we focus on learning with the complementary label specifying classes an image does not belong to, which is easier to obtain. Previous methods produce much lower performance than learning with true labels due to the inherent ambiguity of complementary labels. To deal with the ambiguity of complementary labels, we propose a new complementary learning method called Dual-regularization Complementary Learning (DRCL). Specially, we train two deep neural networks simultaneously and enforce them to regularize the outputs of each other. In this way, the two networks can learn from each other and tend to generate consistent outputs even when supervised by the ambiguous complementary labels. Experiments on four datasets demonstrate the superiority of our approach over state-of-the-art methods. Lingjuan Ge, Mingming Gong, Yutian Lin, Bo Du 0001 |
ICME | 3 |
| 2020 | Unsupervised Person Re-Identification via Softened Similarity LearningabstractPerson re-identification (re-ID) is an important topic in computer vision. This paper studies the unsupervised setting of re-ID, which does not require any labeled information and thus is freely deployed to new scenarios. There are very few studies under this setting, and one of the best approach till now used iterative clustering and classification, so that unlabeled images are clustered into pseudo classes for a classifier to get trained, and the updated features are used for clustering and so on. This approach suffers two problems, namely, the difficulty of determining the number of clusters, and the hard quantization loss in clustering. In this paper, we follow the iterative training mechanism but discard clustering, since it incurs loss from hard quantization, yet its only product, image-level similarity, can be easily replaced by pairwise computation and a softened classification task. With these improvements, our approach becomes more elegant and is more robust to hyper-parameter changes. Experiments on two image-based and video-based datasets demonstrate state-of-the-art performance under the unsupervised re-ID setting. Yutian Lin, Lingxi Xie, Yu Wu 0011, Chenggang Yan 0001, Qi Tian 0001 |
CVPR | 1 |
| 2020 | Bayesian query expansion for multi-camera person re-identification
Yutian Lin, Zhedong Zheng, Chenqiang Gao, Yi Yang 0001 |
Pattern Recognit. Lett. | 1 |
| 2020 | Unsupervised Person Re-identification via Cross-Camera Similarity ExplorationabstractMost person re-identification (re-ID) approaches are based on supervised learning, which requires manually annotated data. However, it is not only resource-intensive to acquire identity annotation but also impractical for large-scale data. To relieve this problem, we propose a cross-camera unsupervised approach that makes use of unsupervised style-transferred images to jointly optimize a convolutional neural network (CNN) and the relationship among the individual samples for person re-ID. Our algorithm considers two fundamental facts in the re- ID task, i.e., variance across diverse cameras and similarity within the same identity. In this paper, we propose an iterative framework which overcomes the camera variance and achieves across-camera similarity exploration. Specifically, we apply an unsupervised style transfer model to generate style-transferred training images with different camera styles. Then we iteratively exploit the similarity within the same identity from both the original and the style-transferred data. We start with considering each training image as a different class to initialize the Convolutional Neural Network (CNN) model. Then we measure the similarity and gradually group similar samples into one class, which increases similarity within each identity. We also introduce a diversity regularization term in the clustering to balance the cluster distribution. The experimental results demonstrate that our algorithm is not only superior to state-of-the-art unsupervised re-ID approaches, but also performs favorably compared with other competing unsupervised domain adaptation methods (UDA) and semi-supervised learning methods. Yutian Lin, Yu Wu 0011, Chenggang Yan 0001, Mingliang Xu 0001, Yi Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | A Bottom-Up Clustering Approach to Unsupervised Person Re-IdentificationabstractMost person re-identification (re-ID) approaches are based on supervised learning, which requires intensive manual annotation for training data. However, it is not only resourceintensive to acquire identity annotation but also impractical to label the large-scale real-world data. To relieve this problem, we propose a bottom-up clustering (BUC) approach to jointly optimize a convolutional neural network (CNN) and the relationship among the individual samples. Our algorithm considers two fundamental facts in the re-ID task, i.e., diversity across different identities and similarity within the same identity. Specifically, our algorithm starts with regarding individual sample as a different identity, which maximizes the diversity over each identity. Then it gradually groups similar samples into one identity, which increases the similarity within each identity. We utilizes a diversity regularization term in the bottom-up clustering procedure to balance the data volume of each cluster. Finally, the model achieves an effective trade-off between the diversity and similarity. We conduct extensive experiments on the large-scale image and video re-ID datasets, including Market-1501, DukeMTMCreID, MARS and DukeMTMC-VideoReID. The experimental results demonstrate that our algorithm is not only superior to state-of-the-art unsupervised re-ID approaches, but also performs favorably than competing transfer learning and semi-supervised learning methods. Yutian Lin, Xuanyi Dong, Liang Zheng 0001, Yan Yan 0002, Yi Yang 0001 |
AAAI | 1 |
| 2019 | Improving person re-identification by attribute and identity learning
Yutian Lin, Liang Zheng 0001, Zhedong Zheng, Yu Wu 0011, Zhilan Hu, Chenggang Yan 0001, Yi Yang 0001 |
Pattern Recognit. | 1 |
| 2019 | Progressive Learning for Person Re-Identification With One ExampleabstractIn this paper, we focus on the one-example person re-identification (re-ID) task, where each identity has only one labeled example along with many unlabeled examples. We propose a progressive framework which gradually exploits the unlabeled data for person re-ID. In this framework, we iteratively (1) update the Convolutional Neural Network (CNN) model and (2) estimate pseudo labels for the unlabeled data. We split the training data into three parts, i.e., labeled data, pseudo-labeled data, and indexlabeled data. Initially, the re-ID model is trained using the labeled data. For the subsequent model training, we update the CNN model by the joint training on the three data parts. The proposed joint training method can optimize the model by both the data with labels (or pseudo labels) and the data without any reliable labels. For the label estimation step, instead of using a static sampling strategy, we propose a progressive sampling strategy to increase the number of the selected pseudo-labeled candidates step by step. We select a few candidates with most reliable pseudo labels from unlabeled examples as the pseudo-labeled data, and keep the rest as index-labeled data by assigning them with the data indexes. During iterations, the index-labeled data are dynamically transferred to pseudo-labeled data. Notably, the rank-1 accuracy of our method outperforms the state-of-the-art method by 21.6 points (absolute, i.e., 62.8% vs. 41.2%) on MARS, and 16.6 points on DukeMTMC-VideoReID. Extended to the few-example setting, our approach with only 20% labeled data surprisingly achieves comparable performance to the supervised state-of-the-art method with 100% labeled data. Yu Wu 0011, Yutian Lin, Xuanyi Dong, Yan Yan 0006, Yi Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Exploit the Unknown Gradually: One-Shot Video-Based Person Re-Identification by Stepwise LearningabstractWe focus on the one-shot learning for video-based person re-Identification (re-ID). Unlabeled tracklets for the person re-ID tasks can be easily obtained by preprocessing, such as pedestrian detection and tracking. In this paper, we propose an approach to exploiting unlabeled tracklets by gradually but steadily improving the discriminative capability of the Convolutional Neural Network (CNN) feature representation via stepwise learning. We first initialize a CNN model using one labeled tracklet for each identity. Then we update the CNN model by the following two steps iteratively: 1. sample a few candidates with most reliable pseudo labels from unlabeled tracklets; 2. update the CNN model according to the selected data. Instead of the static sampling strategy applied in existing works, we propose a progressive sampling method to increase the number of the selected pseudo-labeled candidates step by step. We systematically investigate the way how we should select pseudo-labeled tracklets into the training set to make the best use of them. Notably, the rank-1 accuracy of our method outperforms the state-of-the-art method by 21.46 points (absolute, i.e., 62.67% vs. 41.21%) on the MARS dataset, and 16.53 points on the DukeMTMC-VideoReID dataset. Yu Wu 0011, Yutian Lin, Xuanyi Dong, Yan Yan 0006, Wanli Ouyang, Yi Yang 0001 |
CVPR | 2 |