Qingyi Tao

dblp:179/0983 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0003-4575-612XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 MatAnyone: Stable Video Matting with Consistent Memory Propagation
abstract
Auxiliary-free human video matting methods, which rely solely on input frames, often struggle with complex or ambiguous backgrounds. To address this, we propose MatAnyone, a robust framework tailored for target-assigned video matting. Specifically, building on a memory-based paradigm, we introduce a consistent memory propagation module via region-adaptive memory fusion, which adaptively integrates memory from the previous frame. This ensures semantic stability in core regions while preserving fine-grained details along object boundaries. For robust training, we present a larger, high-quality, and diverse dataset for video matting. Additionally, we incorporate a novel training strategy that efficiently leverages large-scale segmentation data, boosting matting stability. With this new network design, dataset, and training strategy, MatAnyone delivers robust and accurate video matting results in diverse real-world scenarios, outperforming existing methods.
Peiqing Yang 0001, Shangchen Zhou, Jixin Zhao, Qingyi Tao, Chen Change Loy
CVPR4
2025 SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer
abstract
Photorealistic style transfer (PST) enables real-world color grading by adapting reference image colors while preserving content structure. Existing methods mainly follow either approaches: generation-based methods that prioritize stylistic fidelity at the cost of content integrity and efficiency, or global color transformation methods such as LUT, which preserve structure but lack local adaptability. To bridge this gap, we propose Spatial Adaptive 4D Look-Up Table (SA-LUT), combining LUT efficiency with neural network adaptability. SA-LUT features: (1) a Style-guided 4D LUT Generator that extracts multi-scale features from the style image to predict a 4D LUT, and (2) a Context Generator using content-style cross-attention to produce a context map. This context map enables spatially-adaptive adjustments, allowing our 4D LUT to apply precise color transformations while preserving structural integrity. To establish a rigorous evaluation framework for photorealistic style transfer, we introduce PST50, the first benchmark specifically designed for PST assessment. Experiments demonstrate that SA-LUT substantially outperforms state-of-the-art methods, achieving a 66.7% reduction in LPIPS score compared to 3D LUT approaches, while maintaining real-time performance at 16 FPS for video stylization. Our code and benchmark are available at https://github.com/Ry3nG/SA-LUT
Zerui Gong, Qingyi Tao, Qinyue Li, Chen Change Loy
ICCV3
2025 Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
abstract
Unifying visual understanding and generation within a single multimodal framework remains a significant challenge, as the two inherently heterogeneous tasks require representations at different levels of granularity. Current approaches that utilize vector quantization (VQ) or variational autoencoders (VAE) for unified visual representation prioritize intrinsic imagery features over semantics, compromising understanding performance. In this work, we take inspiration from masked image modelling (MIM) that learns rich semantics via a mask-and-reconstruct pre-training and its successful extension to masked autoregressive (MAR) image generation. A preliminary study on the MAR encoder's representation reveals exceptional linear probing accuracy and precise feature response to visual concepts, which indicates MAR's potential for visual understanding tasks beyond its original generation role. Based on these insights, we present \emph{Harmon}, a unified autoregressive framework that harmonizes understanding and generation tasks with a shared MAR encoder. Through a three-stage training procedure that progressively optimizes understanding and generation capabilities, Harmon achieves state-of-the-art image generation results on the GenEval, MJHQ30K and WISE benchmarks while matching the performance of methods with dedicated semantic encoders (e.g., Janus) on image understanding benchmarks. Our code and models will be available at https://github.com/wusize/Harmon.
Size Wu, Lumin Xu, Sheng Jin 0007, Qingyi Tao, Wentao Liu 0002, Wei Li 0319, Chen Change Loy
ICCV6
2024 Modeling Continuous Motion for 3D Point Cloud Object Tracking
abstract
The task of 3D single object tracking (SOT) with LiDAR point clouds is crucial for various applications, such as autonomous driving and robotics. However, existing approaches have primarily relied on appearance matching or motion modeling within only two successive frames, thereby overlooking the long-range continuous motion property of objects in 3D space. To address this issue, this paper presents a novel approach that views each tracklet as a continuous stream: at each timestamp, only the current frame is fed into the network to interact with multi-frame historical features stored in a memory bank, enabling efficient exploitation of sequential information. To achieve effective cross-frame message passing, a hybrid attention mechanism is designed to account for both long-range relation modeling and local geometric feature extraction. Furthermore, to enhance the utilization of multi-frame features for robust tracking, a contrastive sequence enhancement strategy is proposed, which uses ground truth tracklets to augment training sequences and promote discrimination against false positives in a contrastive manner. Extensive experiments demonstrate that the proposed method outperforms the state-of-the-art method by significant margins on multiple benchmarks.
Gongjie Zhang, Changqing Zhou, Qingyi Tao, Lewei Lu, Shijian Lu
AAAI5
2023 Towards Robust and Expressive Whole-body Human Pose and Shape Estimation
abstract
Whole-body pose and shape estimation aims to jointly predict different behaviors (e.g., pose, hand gesture, facial expression) of the entire human body from a monocular image. Existing methods often exhibit suboptimal performance due to the complexity of in-the-wild scenarios. We argue that the prediction accuracy of these models is significantly affected by the quality of the _bounding box_, e.g., scale, alignment. The natural discrepancy between the ideal bounding box annotations and model detection results is particularly detrimental to the performance of whole-body pose and shape estimation. In this paper, we propose a novel framework to enhance the robustness of whole-body pose and shape estimation. Our framework incorporates three new modules to address the above challenges from three perspectives: (1) a **Localization Module** enhances the model's awareness of the subject's location and semantics within the image space; (2) a **Contrastive Feature Extraction Module** encourages the model to be invariant to robust augmentations by incorporating a contrastive loss and positive samples; (3) a **Pixel Alignment Module** ensures the reprojected mesh from the predicted camera and body model parameters are more accurate and pixel-aligned. We perform comprehensive experiments to demonstrate the effectiveness of our proposed framework on body, hands, face and whole-body benchmarks.
Hui En Pang, Zhongang Cai, Lei Yang 0059, Qingyi Tao, Tianwei Zhang 0004, Ziwei Liu 0002
NeurIPS4
2023 PGDiff: Guiding Diffusion Models for Versatile Face Restoration via Partial Guidance
abstract
Exploiting pre-trained diffusion models for restoration has recently become a favored alternative to the traditional task-specific training approach. Previous works have achieved noteworthy success by limiting the solution space using explicit degradation models. However, these methods often fall short when faced with complex degradations as they generally cannot be precisely modeled. In this paper, we introduce $\textit{partial guidance}$, a fresh perspective that is more adaptable to real-world degradations compared to existing works. Rather than specifically defining the degradation process, our approach models the desired properties, such as image structure and color statistics of high-quality images, and applies this guidance during the reverse diffusion process. These properties are readily available and make no assumptions about the degradation process. When combined with a diffusion prior, this partial guidance can deliver appealing results across a range of restoration tasks. Additionally, our method can be extended to handle composite tasks by consolidating multiple high-quality image properties, achieved by integrating the guidance from respective tasks. Experimental results demonstrate that our method not only outperforms existing diffusion-prior-based approaches but also competes favorably with task-specific models.
Peiqing Yang 0001, Shangchen Zhou, Qingyi Tao, Chen Change Loy
NeurIPS3
2023 Medical visual question answering: A survey
Donghao Zhang 0004, Qingyi Tao, Danli Shi, Gholamreza Haffari, Qi Wu 0001, Mingguang He, ZongYuan Ge
Artif. Intell. Medicine3
2023 ReCasNet: Improving consistency within the two-stage mitosis detection framework
Chawan Piansaddhayanon, Sakun Santisukwongchote, Shanop Shuangshoti, Qingyi Tao, Sira Sriswasdi, Ekapol Chuangsuwanich
Artif. Intell. Medicine4
2023 Contrastive pre-training and linear interaction attention-based transformer for universal medical reports generation
Donghao Zhang 0004, Danli Shi, Renjing Xu, Qingyi Tao, Lin Wu 0001, Mingguang He, ZongYuan Ge
J. Biomed. Informatics5
2021 Retrospective Class Incremental Learning
abstract
Existing works study the Class Incremental learning (CIL) problem with the assumption that the data for previous classes are absent, or only a small subset of samples (known as exemplars) are accessible. Differently, we propose a new and practical setting called retrospective CIL, where all the previous data are accessible, but with bounded training budgets for old data replay. Since only a small subset of old samples can be replayed, it brings a new research problem, i.e., dynamically sampling old data along the incremental training process. As incremental learning particularly suffers from catastrophic forgetting, we propose to use the forgettability of the old samples as the sampling priorities to favour the forgotten samples during the dynamic sampling process. To achieve this, we introduce a forgetting rate metric with graph- based propagation to estimate the sample forgettability. The proposed method brings improvements on two benchmark datasets.
Qingyi Tao, Chen Change Loy, Jianfei Cai 0001, ZongYuan Ge, Simon See
ICME1
2020 Exploring Bottom-Up and Top-Down Cues With Attentive Learning for Webly Supervised Object Detection
abstract
Fully supervised object detection has achieved great success in recent years. However, abundant bounding boxes annotations are needed for training a detector for novel classes. To reduce the human labeling effort, we propose a novel webly supervised object detection (WebSOD) method for novel classes which only requires the web images without further annotations. Our proposed method combines bottom-up and top-down cues for novel class detection. Within our approach, we introduce a bottom-up mechanism based on the well-trained fully supervised object detector (i.e. Faster RCNN) as an object region estimator for web images by recognizing the common objectiveness shared by base and novel classes. With the estimated regions on the web images, we then utilize the top-down attention cues as the guidance for region classification. Furthermore, we propose a residual feature refinement (RFR) block to tackle the domain mismatch between web domain and the target domain. We demonstrate our proposed method on PASCAL VOC dataset with three different novel/base splits. Without any target-domain novel-class images and annotations, our proposed webly supervised object detection model is able to achieve promising performance for novel classes. Moreover, we also conduct transfer learning experiments on large scale ILSVRC 2013 detection dataset and achieve state-of-the-art performance.
Qingyi Tao, Guosheng Lin, Jianfei Cai 0001
CVPR2
2019 Improving Deep Lesion Detection Using 3D Contextual and Spatial Attention
Qingyi Tao, ZongYuan Ge, Jianfei Cai 0001, Jianxiong Yin, Simon See
MICCAI (6)1
2019 M2E-Try On Net: Fashion from Model to Everyone
abstract
Most existing virtual try-on applications require clean clothes images. Instead, we present a novel virtual Try-On network, M2E-Try On Net, which transfers the clothes from a model image to a person image without the need of any clean product images. To obtain a realistic image of person wearing the desired model clothes, we aim to solve the following challenges: 1) non-rigid nature of clothes - we need to align poses between the model and the user; 2) richness in textures of fashion items - preserving the fine details and characteristics of the clothes is critical for photo-realistic transfer; 3) variation of identity appearances - it is required to fit the desired model clothes to the person identity seamlessly. To tackle these challenges, we introduce three key components, including the pose alignment network (PAN), the texture refinement network (TRN) and the fitting network (FTN). Since it is unlikely to gather image pairs of input person image and desired output image (i.e. person wearing the desired clothes), our framework is trained in a self-supervised manner to gradually transfer the poses and textures of the model's clothes to the desired appearance. In the experiments, we verify on the Deep Fashion dataset and MVC dataset that our method can generate photo-realistic images for the person to try-on the model clothes. Furthermore, we explore the model capability for different fashion items, including both upper and lower garments.
Guosheng Lin, Qingyi Tao, Jianfei Cai 0001
ACM Multimedia3
2019 Exploiting Web Images for Weakly Supervised Object Detection
abstract
In recent years, the performance of object detection has advanced significantly with the evolution of deep convolutional neural networks. However, the state-of-the-art object detection methods still rely on accurate bounding box annotations that require extensive human labeling. Object detection without bounding box annotations, that is, weakly supervised detection methods, are still lagging far behind. As weakly supervised detection only uses image level labels and does not require the ground truth of bounding box location and label of each object in an image, it is generally very difficult to distill knowledge of the actual appearances of objects. Inspired by curriculum learning, this paper proposes an easy-to-hard knowledge transfer scheme that incorporates easy web images to provide prior knowledge of object appearance as a good starting point. While exploiting large-scale free web imagery, we introduce a sophisticated labor-free method to construct a web dataset with good diversity in object appearance. After that, semantic relevance and distribution relevance are introduced and utilized in the proposed curriculum training scheme. Our end-to-end learning with the constructed web data achieves remarkable improvement across most object classes, especially for the classes that are often considered hard in other works.
Qingyi Tao, Hao Yang 0033, Jianfei Cai 0001
IEEE Trans. Multim.1
2018 VQA-E: Explaining, Elaborating, and Enhancing Your Answers for Visual Questions
Qing Li 0003, Qingyi Tao, Shafiq R. Joty, Jianfei Cai 0001, Jiebo Luo 0001
ECCV (7)2
2018 Zero-Annotation Object Detection with Web Knowledge Transfer
Qingyi Tao, Hao Yang 0033, Jianfei Cai 0001
ECCV (11)1
2015 Efficient image retrieval based mobile indoor localization
abstract
Vision based localization has been investigated for many years. The existing Structure from Motion (SfM) technique can reconstruct the 3D models based on the input images. The image retrieval and feature matching allow us to find the correspondence between the query image and the 3D model. According to these, the location can be easily calculated. In mobile scenarios, the limited CPU speed, memory storage and network latency bring in new challenges. The state-of-the-art solution can not be easily adopted due to the complicated calculation and large resource consumption. In this paper, we leverage the techniques developed during the MPEG-7 Compact Descriptors for Visual Search (CDVS) standardization, which aims to provide high performance and low complexity compact descriptors. We show that these techniques are suitable for mobile device and can achieve state-of-the-art retrieval performance in indoor environment. Besides, we propose additional components including blur measurement and result smoothing to improve the performance of the location calculation process. Based on these techniques, a whole system which enables fast vision based localization on mobile device is developed. We present experiments on the real world situation, showing that the system can strike a balance between accuracy and efficiency.
Ruoyun He, Qingyi Tao, Jianfei Cai 0001, Ling-Yu Duan
VCIP3