EDBT 2026 Demo / reviewers in the wild / expert
Tianyu Yang 0003
dblp:120/8076-3
· DBLP profile ↗
20ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-9674-5220ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReEx-SQL: Reasoning with Execution-Aware Reinforcement Learning for Text-to-SQLabstractYaxun Dai, Wenxuan Xie, Xialie Zhuang, Tianyu Yang, Ziyi Liu, Haiqin Yang, Yiying Yang, Yuhang Zhao, Pingfu Chao, Wenhao Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yaxun Dai, Wenxuan Xie, Xialie Zhuang, Tianyu Yang 0003, Ziyi Liu 0005, Haiqin Yang, Pingfu Chao |
ACL (1) | 4 |
| 2025 | StableDepth: Scene-Consistent and Scale-Invariant Monocular Depth
Lihe Yang, Tianyu Yang 0003, Chaohui Yu, Yixing Lao, Hengshuang Zhao |
ICCV | 3 |
| 2024 | A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video EditingabstractThis paper presents a video inversion approach for zero-shot video editing, which models the input video with low-rank representation during the inversion process. The existing video editing methods usually apply the typical 2D DDIM inversion or naï ve spatial-temporal DDIM inversion before editing, which leverages time-varying representation for each frame to derive noisy latent. Unlike most existing approaches, we propose a Spatial-Temporal Expectation-Maximization (STEM) inversion, which formulates the dense video feature under an expectation-maximization manner and iteratively estimates a more compact basis set to represent the whole video. Each frame applies the fixed and global representation for inversion, which is more friendly for temporal consistency during reconstruction and editing. Extensive qualitative and quantitative experiments demonstrate that our STEM inversion can achieve consistent improvement on two state-of-the-art video editing methods. Project page: https://steminv.github.io/page/. Maomao Li, Yu Li 0003, Tianyu Yang 0003, Yunfei Liu 0001, Dongxu Yue, Zhihui Lin |
CVPR | 3 |
| 2024 | OMG: Occlusion-Friendly Personalized Multi-concept Generation in Diffusion Models
Zhe Kong, Yong Zhang 0034, Tianyu Yang 0003, Tao Wang 0052, Kaihao Zhang, Bizhu Wu, Guanying Chen, Wei Liu 0005, Wenhan Luo |
ECCV (31) | 3 |
| 2024 | AddMe: Zero-Shot Group-Photo Synthesis by Inserting People Into Scenes
Dongxu Yue, Maomao Li, Yunfei Liu 0001, Ailing Zeng, Tianyu Yang 0003, Yu Li 0003 |
ECCV (20) | 6 |
| 2024 | Compress3D: A Compressed Latent Space for 3D Generation from a Single Image
Tianyu Yang 0003, Yu Li 0003, Lei Zhang 0001, Xi Zhao 0002 |
ECCV (18) | 2 |
| 2024 | Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic PromptsabstractRecent text-to-3D generation methods achieve impressive 3D content creation capacity thanks to the advances in image diffusion models and optimizing strategies. However, current methods struggle to generate correct 3D content for a complex prompt in semantics, i.e., a prompt describing multiple interacted objects binding with different attributes. In this work, we propose a general framework named Progressive3D, which decomposes the entire generation into a series of locally progressive editing steps to create precise 3D content for complex prompts, and we constrain the content change to only occur in regions determined by user-defined region prompts in each editing step. Furthermore, we propose an overlapped semantic component suppression technique to encourage the optimization process to focus more on the semantic differences between prompts. Extensive experiments demonstrate that the proposed Progressive3D framework generates precise 3D content for prompts with complex semantics through progressive editing steps and is general for various text-to-3D methods driven by different 3D representations. Xinhua Cheng, Tianyu Yang 0003, Yu Li 0003, Lei Zhang 0001, Jian Zhang 0018, Li Yuan 0007 |
ICLR | 2 |
| 2024 | TOSS: High-quality Text-guided Novel View Synthesis from a Single ImageabstractIn this paper, we present TOSS, which introduces text to the task of novel view synthesis (NVS) from just a single RGB image.
While Zero123 has demonstrated impressive zero-shot open-set NVS capabilities, it treats NVS as a pure image-to-image translation problem. This approach suffers from the challengingly under-constrained nature of single-view NVS: the process lacks means of explicit user control and often result in implausible NVS generations.
To address this limitation, TOSS uses text as high-level semantic information to constrain the NVS solution space.
TOSS fine-tunes text-to-image Stable Diffusion pre-trained on large-scale text-image pairs and introduces modules specifically tailored to image and camera pose conditioning, as well as dedicated training for pose correctness and preservation of fine details.
Comprehensive experiments are conducted with results showing that our proposed TOSS outperforms Zero123 with higher-quality NVS results and faster convergence. We further support these results with comprehensive ablations that underscore the effectiveness and potential of
the introduced semantic guidance and architecture design. Yukai Shi, He Cao, Boshi Tang, Xianbiao Qi, Tianyu Yang 0003, Shilong Liu 0004, Lei Zhang 0001, Harry Shum |
ICLR | 6 |
| 2023 | DropMAE: Masked Autoencoders with Spatial-Attention Dropout for Tracking TasksabstractIn this paper, we study masked autoencoder (MAE) pretraining on videos for matching-based downstream tasks, including visual object tracking (VOT) and video object segmentation (VOS). A simple extension of MAE is to randomly mask out frame patches in videos and reconstruct the frame pixels. However, we find that this simple baseline heavily relies on spatial cues while ignoring temporal relations for frame reconstruction, thus leading to sub-optimal temporal matching representations for VOT and VOS. To alleviate this problem, we propose DropMAE, which adaptively performs spatial-attention dropout in the frame reconstruction to facilitate temporal correspondence learning in videos. We show that our DropMAE is a strong and efficient temporal matching learner, which achieves better finetuning results on matching-based tasks than the ImageNet-based MAE with$2\times$faster pre-training speed. Moreover, we also find that motion diversity in pre-training videos is more important than scene diversity for improving the performance on VOT and VOS. Our pre-trained DropMAE model can be directly loaded in existing ViT-based trackers for fine-tuning without further modifications. Notably, DropMAE sets new state-of-the-art performance on 8 out of 9 highly competitive video tracking and segmentation datasets. Our code and pre-trained models are available at https://github.com/jimmy-dq/DropMAE.git. Qiangqiang Wu, Tianyu Yang 0003, Ziquan Liu, Baoyuan Wu, Ying Shan, Antoni B. Chan |
CVPR | 2 |
| 2023 | Scalable Video Object Segmentation with Simplified FrameworkabstractThe current popular methods for video object segmentation (VOS) implement feature matching through several hand-crafted modules that separately perform feature extraction and matching. However, the above hand-crafted designs empirically cause insufficient target interaction, thus limiting the dynamic target-aware feature learning in VOS. To tackle these limitations, this paper presents a scalable Simplified VOS (SimVOS) framework to perform joint feature extraction and matching by leveraging a single transformer backbone. Specifically, SimVOS employs a scalable ViT backbone for simultaneous feature extraction and matching between query and reference features. This design enables SimVOS to learn better target-ware features for accurate mask prediction. More importantly, SimVOS could directly apply well-pretrained ViT backbones (e.g., MAE [21]) for VOS, which bridges the gap between VOS and large-scale self-supervised pre-training. To achieve a better performance-speed trade-off, we further explore within-frame attention and propose a new token refinement module to improve the running speed and save computational cost. Experimentally, our SimVOS achieves state-of-the-art results on popular video object segmentation benchmarks, i.e., DAVIS-2017 (88.0% $\mathcal{J}\& \mathcal{F}$), DAVIS-2016 (92.9% $\mathcal{J}\& \mathcal{F}$) and YouTube-VOS 2019 (84.2% $\mathcal{J}\& \mathcal{F}$), without applying any synthetic video or BL30K pre-training used in previous VOS approaches. Our code and models are available at https://github.com/jimmy-dq/SimVOS.git. Qiangqiang Wu, Tianyu Yang 0003, Antoni B. Chan |
ICCV | 2 |
| 2022 | Motion-aware Contrastive Video Representation Learning via Foreground-background MergingabstractIn light of the success of contrastive learning in the image domain, current self-supervised video representation learning methods usually employ contrastive loss to facilitate video representation learning. When naively pulling two augmented views of a video closer, the model however tends to learn the common static background as a shortcut but fails to capture the motion information, a phenomenon dubbed as background bias. Such bias makes the model suffer from weak generalization ability, leading to worse performance on downstream tasks such as action recognition. To alleviate such bias, we propose Foreground-background Merging (FAME) to deliberately compose the moving foreground region of the selected video onto the static background of others. Specifically, without any off-the-shelf detector, we extract the moving fore-ground out of background regions via the frame difference and color statistics, and shuffle the background regions among the videos. By leveraging the semantic consistency between the original clips and the fused ones, the model focuses more on the motion patterns and is debiased from the background shortcut. Extensive experiments demonstrate that FAME can effectively resist background cheating and thus achieve the state-of-the-art performance on downstream tasks across UCF101, HMDB51, and Diving48 datasets. The code and configurations are released at https://github.com/Mark12Ding/FAME. Shuangrui Ding, Maomao Li, Tianyu Yang 0003, Rui Qian 0001, Haohang Xu, Qingyi Chen, Jue Wang 0001, Hongkai Xiong |
CVPR | 3 |
| 2022 | Exploring Denoised Cross-video Contrast for Weakly-supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization aims to localize actions in untrimmed videos with only video-level labels. Most existing methods address this problem with a “localization-by-classification” pipeline that localizes action regions based on snippet-wise classification sequences. Snippet-wise classifications are unfortunately error prone due to the sparsity of video-level labels. Inspired by recent success in unsupervised contrastive representation learning, we propose a novel denoised cross-video contrastive algorithm, aiming to enhance the feature discrimination ability of video snippets for accurate temporal action localization in the weakly-supervised setting. This is enabled by three key designs: 1) an effective pseudo-label denoising module to alleviate the side effects caused by noisy contrastive features, 2) an efficient region-level feature contrast strategy with a region-level memory bank to capture “global” contrast across the entire dataset, and 3) a diverse contrastive learning strategy to enable action-background separation as well as intra-class compactness & inter-class separability. Extensive experiments on THUMOS14 and ActivityNet v1.3 demonstrate the superior performance of our approach. Tianyu Yang 0003, Wei Ji 0011, Jue Wang 0001, Li Cheng 0001 |
CVPR | 2 |
| 2022 | SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-MaximizationabstractMatching-based methods, especially those based on space-time memory, are significantly ahead of other solutions in semi-supervised video object segmentation (VOS). However, continuously growing and redundant template features lead to an inefficient inference. To alleviate this, we propose a novel Sequential Weighted Expectation-Maximization (SWEM) network to greatly reduce the redundancy of memory features. Different from the previous methods which only detect feature redundancy between frames, SWEM merges both intra-frame and inter-frame similar features by leveraging the sequential weighted EM algorithm. Further, adaptive weights for frame features endow SWEM with the flexibility to represent hard samples, improving the discrimination of templates. Besides, the proposed method maintains a fixed number of template features in memory, which ensures the stable inference complexity of the VOS system. Extensive experiments on commonly used DAVIS and YouTube-VOS datasets verify the high efficiency (36 FPS) and high performance (84.3% J&F on DAVIS 2017 validation dataset) of SWEM. Zhihui Lin, Tianyu Yang 0003, Maomao Li, Chun Yuan 0003, Wei Liu 0005 |
CVPR | 2 |
| 2022 | Unsupervised Pre-training for Temporal Action Localization TasksabstractUnsupervised video representation learning has made remarkable achievements in recent years. However, most existing methods are designed and optimized for video classification. These pretrained models can be sub-optimal for temporal localization tasks due to the inherent discrepancy between video-level classification and clip-level localization. To bridge this gap, we make the first attempt to propose a self-supervised pretext task, coined as Pseudo Action Localization (PAL) to Unsupervisedly Pre-train feature encoders for Temporal Action Localization tasks (UP-TAL). Specifically, we first randomly select temporal regions, each of which contains multiple clips, from one video as pseudo actions and then paste them onto different temporal positions of the other two videos. The pretext task is to align the features of pasted pseudo action regions from two synthetic videos and maximize the agreement between them. Compared to the existing unsupervised video representation learning approaches, our PAL adapts better to downstream TAL tasks by introducing a temporal equivariant contrastive learning paradigm in a temporally dense and scale-aware manner. Extensive experiments show that PAL can utilize large-scale unlabeled video data to significantly boost the performance of existing TAL methods. Our codes and models will be made publicly available at https://github.com/zhang-can/UP-TAL. Can Zhang 0001, Tianyu Yang 0003, Junwu Weng, Meng Cao 0002, Jue Wang 0001, Yuexian Zou |
CVPR | 2 |
| 2022 | LocVTP: Video-Text Pre-training for Temporal Localization
Meng Cao 0002, Tianyu Yang 0003, Junwu Weng, Can Zhang 0001, Jue Wang 0001, Yuexian Zou |
ECCV (26) | 2 |
| 2021 | VideoMoCo: Contrastive Video Representation Learning With Temporally Adversarial ExamplesabstractMoCo [11] is effective for unsupervised image representation learning. In this paper, we propose VideoMoCo for unsupervised video representation learning. Given a video sequence as an input sample, we improve the temporal feature representations of MoCo from two perspectives. First, we introduce a generator to drop out several frames from this sample temporally. The discriminator is then learned to encode similar feature representations regardless of frame removals. By adaptively dropping out different frames during training iterations of adversarial learning, we augment this input sample to train a tempo-rally robust encoder. Second, we use temporal decay to model key attenuation in the memory queue when computing the contrastive loss. As the momentum encoder updates after keys enqueue, the representation ability of these keys degrades when we use the current input sample for contrastive learning. This degradation is reflected via temporal decay to attend the input sample to recent keys in the queue. As a result, we adapt MoCo to learn video representations without empirically designing pretext tasks. By empowering the temporal robustness of the encoder and modeling the temporal decay of the keys, our VideoMoCo improves MoCo temporally based on contrastive learning. Experiments on benchmark datasets including UCF101 and HMDB51 show that VideoMoCo stands as a state-of-the-art video representation learning method. Tian Pan 0003, Yibing Song, Tianyu Yang 0003, Wei Liu 0005 |
CVPR | 3 |
| 2021 | Visual Tracking via Dynamic Memory NetworksabstractTemplate-matching methods for visual tracking have gained popularity recently due to their good performance and fast speed. However, they lack effective ways to adapt to changes in the target object's appearance, making their tracking accuracy still far from state-of-the-art. In this paper, we propose a dynamic memory network to adapt the template to the target's appearance variations during tracking. The reading and writing process of the external memory is controlled by an LSTM network with the search feature map as input. A spatial attention mechanism is applied to concentrate the LSTM input on the potential target as the location of the target is at first unknown. To prevent aggressive model adaptivity, we apply gated residual template learning to control the amount of retrieved memory that is used to combine with the initial template. In order to alleviate the drift problem, we also design a "negative" memory unit that stores templates for distractors, which are used to cancel out wrong responses from the object template. To further boost the tracking performance, an auxiliary classification loss is added after the feature extractor part. Unlike tracking-by-detection methods where the object's information is maintained by the weight parameters of neural networks, which requires expensive online fine-tuning to be adaptable, our tracker runs completely feed-forward and adapts to the target's appearance changes by updating the external memory. Moreover, the capacity of our model is not determined by the network size as with other trackers - the capacity can be easily enlarged as the memory requirements of a task increase, which is favorable for memorizing long-term object information. Extensive experiments on the OTB and VOT datasets demonstrate that our trackers perform favorably against state-of-the-art tracking methods while retaining real-time speed. Tianyu Yang 0003, Antoni B. Chan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | ROAM: Recurrently Optimizing Tracking ModelabstractIn this paper, we design a tracking model consisting of response generation and bounding box regression, where the first component produces a heat map to indicate the presence of the object at different positions and the second part regresses the relative bounding box shifts to anchors mounted on sliding-window locations. Thanks to the resizable convolutional filters used in both components to adapt to the shape changes of objects, our tracking model does not need to enumerate different sized anchors, thus saving model parameters. To effectively adapt the model to appearance variations, we propose to offline train a recurrent neural optimizer to update tracking model in a meta-learning setting, which can converge the model in a few gradient steps. This improves the convergence speed of updating the tracking model while achieving better performance. We extensively evaluate our trackers, ROAM and ROAM++, on the OTB, VOT, LaSOT, GOT-10K and TrackingNet benchmark and our methods perform favorably against state-of-the-art algorithms. Tianyu Yang 0003, Pengfei Xu 0013, Runbo Hu, Antoni B. Chan |
CVPR | 1 |
| 2019 | Density-Preserving Hierarchical EM Algorithm: Simplifying Gaussian Mixture Models for Approximate InferenceabstractWe propose an algorithm for simplifying a finite mixture model into a reduced mixture model with fewer mixture components. The reduced model is obtained by maximizing a variational lower bound of the expected log-likelihood of a set of virtual samples. We develop three applications for our mixture simplification algorithm: recursive Bayesian filtering using Gaussian mixture model posteriors, KDE mixture reduction, and belief propagation without sampling. For recursive Bayesian filtering, we propose an efficient algorithm for approximating an arbitrary likelihood function as a sum of scaled Gaussian. Experiments on synthetic data, human location modeling, visual tracking, and vehicle self-localization show that our algorithm can be widely used for probabilistic data analysis, and is more accurate than other mixture simplification methods. Lei Yu 0013, Tianyu Yang 0003, Antoni B. Chan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Learning Dynamic Memory Networks for Object Tracking
Tianyu Yang 0003, Antoni B. Chan |
ECCV (9) | 1 |