Qingsong Xie

dblp:56/6042 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 SCott: Accelerating Diffusion Models with Stochastic Consistency Distillation
abstract
The iterative sampling procedure employed by diffusion models (DMs) often leads to significant latency. To address this, we propose Stochastic Consistency Distillation (SCott) to enable accelerated text-to-image generation, where high-quality generations can be achieved with just 2-4 sampling steps or even1 step, and further improvements can be obtained by additional cost, e.g., 4 steps. In contrast to vanilla consistency distillation (CD) which distills the ordinary differential equation solvers-based sampling process of a pre-trained teacher model into a student, SCott explores the possibility and validates the efficacy of integrating stochastic differential equation (SDE) solvers into CD to fully unleash the potential of the teacher. SCott is augmented with elaborate strategies to control the noise strength and sampling process of the SDE solver. An adversarial loss is further incorporated to strengthen the sample quality with rare sampling steps. Empirically, on the MSCOCO-2017 5K dataset with a Stable Diffusion-V1.5 teacher, SCott achieves an FID of 21.9, surpassing that of the 1-step InstaFlow (23.4) and the 4-step UFOGen (22.1). Moreover, SCott can yield more diverse samples than other consistency models for high-resolution image generation, with up to 16% improvement in a qualified metric.
Hongjian Liu, Qingsong Xie, Tianxiang Ye, Zhijie Deng, Chen Chen 0015, Shixiang Tang, Xueyang Fu, Haonan Lu, Zhengjun Zha
AAAI2
2025 MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
abstract
Masked Image Modeling (MIM) with Vector Quantization (VQ) has achieved great success in both self-supervised pre-training and image generation. However, most existing methods struggle to address the trade-off in shared latent space for generation quality vs. representation learning and efficiency. To push the limits of this paradigm, we propose MergeVQ, which incorporates token merging techniques into VQ-based generative models to bridge the gap between image generation and visual representation learning in a unified architecture. During pre-training, MergeVQ decouples top-k semantics from latent space with the token merge module after self-attention blocks in the encoder for subsequent Look-up Free Quantization (LFQ) and global alignment and recovers their fine-grained details through cross-attention in the decoder for reconstruction. As for second-stage generation, we introduce MergeAR, which performs KV Cache compression for efficient raster-order prediction. Extensive experiments on ImageNet verify that MergeVQ as an AR generative model achieves competitive performance in both visual representation learning and image generation tasks while maintaining favorable token efficiency and inference speed. Code and model will be available at https://apexgen-x.github.io/MergeVQ.
Siyuan Li 0002, Luyuan Zhang, Zedong Wang, Juanxi Tian, Cheng Tan 0012, Zicheng Liu 0006, Chang Yu 0001, Qingsong Xie, Haonan Lu, Haoqian Wang, Zhen Lei 0001
CVPR8
2025 Advancing Text-to-3D Generation with Linearized Lookahead Variational Score Distillation
abstract
Text-to-3D generation based on score distillation of pre-trained 2D diffusion models has gained increasing interest, with variational score distillation (VSD) as a remarkable example. VSD proves that vanilla score distillation can be improved by introducing an extra score-based model, which characterizes the distribution of images rendered from 3D models, to correct the distillation gradient. Despite the theoretical foundations, VSD, in practice, is likely to suffer from slow and sometimes ill-posed convergence. In this paper, we perform an in-depth investigation of the interplay between the introduced score model and the 3D model, and find that there exists a mismatching problem between LoRA and 3D distributions in practical implementation. We can simply adjust their optimization order to improve the generation quality. By doing so, the score model looks ahead to the current 3D state and hence yields more reasonable corrections. Nevertheless, naive lookahead VSD may suffer from unstable training in practice due to the potential over-fitting. To address this, we propose to use a linearized variant of the model for score distillation, giving rise to the Linearized Lookahead Variational Score Distillation ($L^2$-VSD). $L^2$-VSD can be realized efficiently with forward-mode autodiff functionalities of existing deep learning libraries. Extensive experiments validate the efficacy of $L^2$-VSD, revealing its clear superiority over prior score distillation-based methods. We also show that our method can be seamlessly incorporated into any other VSD-based text-to-3D framework.
Bingde Liu, Qingsong Xie, Haonan Lu, Zhijie Deng
ICCV3
2025 Decomposition of Graphic Design with Unified Multimodal Model
abstract
We propose Layer Decomposition of Graphic Designs (LDGD), a novel vision task that converts composite graphic design (e.g., posters) into structured representations comprising ordered RGB-A layers and metadata. By transforming visual content into structured data, LDGD facilitates precise image editing and offers significant advantages for digital content creation, management, and reuse. This task presents two core challenges: (1) predicting the attribute information (metadata) of each layer, and (2) recovering the occluded regions within overlapping layers to enable high-fidelity image reconstruction. To address this, we present the Decompose Layer Model (DeaM), a large unified multimodal model that integrates a conjoined visual encoder, a language model, and a condition-aware RGB-A decoder. DeaM adopts a two-stage processing pipeline: first generates layer-specific metadata containing information such as spatial coordinates and quantized encodings, and then reconstructs pixel-accurate layer images using a condition-aware RGB-A decoder. Beyond full decomposition, the model supports interactive decomposition via textual or point-based prompts. Extensive experiments demonstrate the effectiveness of the proposed method. The code is accessed at https://github.com/witnessai/DeaM.
Hui Nie 0001, Yutao Cheng, Maoke Yang, Gonglei Shi, Qingsong Xie
ICML6
2025 LOVECon: text-driven training-free long video editing with ControlNet
Zhenyi Liao, Qingsong Xie, Zhijie Deng
Sci. China Inf. Sci.2
2025 A survey of binary code representation technology
abstract
Binary analysis, as an important foundational technology, provides support for numerous applications in the fields of software engineering and security research. With the continuous expansion of software scale and the complex evolution of software architecture, binary analysis technology is facing new challenges. To break through existing bottlenecks, researchers have applied artificial intelligence (AI) technology to the understanding and analysis of binary code. The core lies in characterizing binary code, i.e., how to use intelligent methods to generate representation vectors containing semantic information for binary code, and apply them to multiple downstream tasks of binary analysis. In this paper, we provide a comprehensive survey of recent advances in binary code representation technology, and introduce the workflow of existing research in two parts, i.e., binary code feature selection methods and binary code feature embedding methods. The feature selection section includes mainly two parts: definition and classification of features, and feature construction. First, the abstract definition and classification of features are systematically explained, and second, the process of constructing specific representations of features is introduced in detail. In the feature embedding section, based on the different intelligent semantic understanding models used, the embedding methods are classified into four categories based on the usage of text-embedding models and graph-embedding models. Finally, we summarize the overall development of existing research and provide prospects for some potential research directions related to binary code representation technology.
Taiyan Wang, Qingsong Xie, Zulie Pan, Min Zhang 0054
Frontiers Inf. Technol. Electron. Eng.2
2025 Instruct-ReID++: Towards Universal Purpose Instruction-Guided Person Re-Identification
abstract
Recently, person re-identification (ReID) has witnessed fast development due to its broad practical applications and proposed various settings, e.g., traditional ReID, clothes-changing ReID, and visible-infrared ReID. However, current studies primarily focus on single specific tasks, which limits model applicability in real-world scenarios. This paper aims to address this issue by introducing a novel instruct-ReID task that unifies 6 existing ReID tasks in one model and retrieves images based on provided visual or textual instructions. Instruct-ReID is the first exploration of a general ReID setting, where 6 existing ReID tasks can be viewed as special cases by assigning different instructions. To facilitate research in this new instruct-ReID task, we propose a large-scale OmniReID++ benchmark equipped with diverse data and comprehensive evaluation methods, e.g., task-specific and task-free evaluation settings. In the task-specific evaluation setting, gallery sets are categorized according to specific ReID tasks. We propose a novel baseline model, IRM, with an adaptive triplet loss to handle various retrieval tasks within a unified framework. For task-free evaluation setting, where target person images are retrieved from task-agnostic gallery sets, we further propose a new method called IRM++ with novel memory bank-assisted learning. Extensive evaluations of IRM and IRM++ on OmniReID++ benchmark demonstrate the superiority of our proposed methods, achieving state-of-the-art performance on 10 test sets.
Weizhen He, Yiheng Deng, Yunfeng Yan, Feng Zhu 0006, Yizhou Wang 0007, Lei Bai 0001, Qingsong Xie, Rui Zhao 0001, Donglian Qi, Wanli Ouyang, Shixiang Tang
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Instruct-ReID: A Multi-Purpose Person Re-Identification Task with Instructions
abstract
Human intelligence can retrieve any person according to both visual and language descriptions. However, the current computer vision community studies specific person re-identification (ReID) tasks in different scenarios separately, which limits the applications in the real world. This paper strives to resolve this problem by proposing a new instruct-ReID task that requires the model to retrieve images according to the given image or language instructions. Our instruct-ReID is a more general ReID setting, where existing 6 ReID tasks can be viewed as special cases by designing different instructions. We propose a large-scale OmniReID benchmark and an adaptive triplet loss as a baseline method to facilitate research in this new setting. Experimental results show that the proposed multi-purpose ReID model, trained on our OmniReID benchmark without finetuning, can improve +0.5%, +0.6%, +7.7% mAP on Market1501, MSMT17, CUHK03 for traditional ReID, +6.4%, +7.1%, +11.2% mAP on PRCC, VC-Clothes, LTCC for clothes-changing ReID, +11.7% mAP on COCAS+ real2 for clothes template based clothes-changing ReID when using only RGB images, +24.9% mAP on COCAS+ real2 for our newly defined language-instructed ReID, +4.3% on LLCM for visible-infrared ReID, +2.6% on CUHK-PEDES for text-to-image ReID. The datasets, the model, and code are available at https://github.com/hwz-zju/Instruct-ReID.
Weizhen He, Yiheng Deng, Shixiang Tang, Qihao Chen, Qingsong Xie, Yizhou Wang 0007, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001, Wanli Ouyang, Donglian Qi, Yunfeng Yan
CVPR5
2024 PEA-Diffusion: Parameter-Efficient Adapter with Knowledge Distillation in Non-english Text-to-Image Generation
Jian Ma 0010, Chen Chen 0015, Qingsong Xie, Haonan Lu
ECCV (68)3
2024 Unsupervised Domain Adaptation for Medical Image Segmentation by Disentanglement Learning and Self-Training
abstract
Unsupervised domain adaption (UDA), which aims to enhance the segmentation performance of deep models on unlabeled data, has recently drawn much attention. In this paper, we propose a novel UDA method (namely DLaST) for medical image segmentation via disentanglement learning and self-training. Disentanglement learning factorizes an image into domain-invariant anatomy and domain-specific modality components. To make the best of disentanglement learning, we propose a novel shape constraint to boost the adaptation performance. The self-training strategy further adaptively improves the segmentation performance of the model for the target domain through adversarial learning and pseudo label, which implicitly facilitates feature alignment in the anatomy space. Experimental results demonstrate that the proposed method outperforms the state-of-the-art UDA methods for medical image segmentation on three public datasets, i.e., a cardiac dataset, an abdominal dataset and a brain dataset. The code will be released soon.
Qingsong Xie, Yuexiang Li, Nanjun He, Munan Ning, Kai Ma 0002, Guoxing Wang, Yong Lian 0001, Yefeng Zheng 0001
IEEE Trans. Medical Imaging1
2024 Deep recurrent residual channel attention network for single image super-resolution
Yepeng Liu 0003, Dezhi Yang, Fan Zhang 0045, Qingsong Xie, Caiming Zhang 0001
Vis. Comput.4
2023 HumanBench: Towards General Human-Centric Perception with Projector Assisted Pretraining
abstract
Human-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the metaverse. It is desirable to have a general pretrain model for versatile human-centric downstream tasks. This paper forges ahead along this path from the aspects of both benchmark and pretraining methods. Specifically, we propose a HumanBench based on existing datasets to comprehensively evaluate on the common ground the generalization abilities of different pretraining methods on 19 datasets from 6 diverse downstream tasks, including person ReID, pose estimation, human parsing, pedestrian attribute recognition, pedestrian detection, and crowd counting. To learn both coarse-grained and fine-grained knowledge in human bodies, we further propose a Projector AssisTed Hierarchical pretraining method (PATH) to learn diverse knowledge at different granularity levels. Comprehensive evaluations on HumanBench show that our PATH achieves new state-of-the-art results on 17 downstream datasets and on-par results on the other 2 datasets. The code will be publicly at https://github.com/OpenGVLab/HumanBench.
Shixiang Tang, Qingsong Xie, Meilin Chen, Yizhou Wang 0007, Yuanzheng Ci, Lei Bai 0001, Feng Zhu 0006, Haiyang Yang, Rui Zhao 0001, Wanli Ouyang
CVPR3
2023 Topic identification of text-based expert stock comments using multi-level information fusion
abstract
Abstract Stock investment is an important mode of asset allocation and a crucial means of financial management. How to grasp the movement of stock price and predict its trend have been the focus of investors and investment companies. Since expert stock comments contain abundant essential information for investment decisions, how to identify the topic of expert stock comments with high precision and efficiency is an important research topic. However, the existing methods usually employ single feature selection strategies for topic identification of stock comments, which may lead to low accuracy. Thus, to deal with this limitation, we propose a multi‐level information fusion method to construct a topic identification system of stock comments. Specifically, we firstly fuse various complementary feature selection methods via a multi‐view learning framework which can comprehensively represent text‐based topics. In addition, regarding the decision process, we propose a fusion strategy based on belief value which can further improve the classification performance. The experimental results indicate that the proposed multi‐level information fusion method is not only superior to other methods in terms of classification, it is also able to accurately capture topics of expert stock comments.
Feng Zhao 0006, Xiaofeng Zhang 0003, Qingsong Xie
Expert Syst. J. Knowl. Eng.5
2023 Learning the model update with local trusted templates for visual tracking
abstract
Abstract Existing Siamese trackers usually do not update templates or adopt single‐updating strategies. However, historical information cannot be effectively utilized when using these strategies, and model drift from complex tracking challenges cannot be addressed. To address this issue, a novel tracking framework that learns the model update with local trusted templates is proposed in this paper. The authors propose a complementary confidence evaluation method to select local trusted templates in a sliding window. This provides high‐confidence historical information. The authors also propose a method including linear learning and deep learning to learn to model updates. Different from traditional update strategies, the authors’ method combines non‐linear and linear updates to obtain reliable templates with the most abundant historical information, which solves the complex tracking challenges to a certain extent. Finally, the adaptive fusion response maps of the two strategies determine the final tracking based on the confidence evaluation. Experimental results on NFS, UAVDT, UAV123, UAV20L and VOT2016 show that our method performs favourably when compared with current state‐of‐the‐art methods.
Zhiyong An, Ximin Zhang, Zhuhai Wang, Qingsong Xie
IET Image Process.5
2020 Discrete Biorthogonal Wavelet Transform Based Convolutional Neural Network for Atrial Fibrillation Diagnosis from Electrocardiogram
abstract
For the problem of early detection of atrial fibrillation (AF) from electrocardiogram (ECG), it is difficult to capture subject-invariant discriminative features from ECG signals, due to the high variation in ECG morphology across subjects and the noise in ECG. In this paper, we propose an Discrete Biorthogonal Wavelet Transform (DBWT) Based Convolutional Neural Network (CNN) for AF detection, shortly called DBWT-AFNet. In DBWT-AFNet, rather than directly feeding ECG into CNN, DBWT is used to separate sub-signals in frequency band of heart beat from ECG, whose output is fed to CNN for AF diagnosis. Such sub-signals are better than the raw ECG for subject-invariant CNN representation learning because noisy information irrelevant to human beat has been largely filtered out. To strengthen the generalization ability of CNN to discover subject-invariant pattern in ECG, skip connection is exploited to propagate information well in neural network and channel attention is designed to adaptively highlight informative channel-wise features. Experiments show that the proposed DBWT-AFNet outperforms the state-of- the-art methods, especially for ECG segments classification across different subjects, where no data from testing subjects have been used in training.
Qingsong Xie, Shikui Tu, Guoxing Wang, Yong Lian 0001, Lei Xu 0001
IJCAI1
2020 A digital signal processor (DSP)-based system for embedded continuous-time cuffless blood pressure monitoring using single-channel PPG signal
Qirui Zhang 0001, Qingsong Xie, Kefeng Duan, Min Wang 0014, Guoxing Wang
Sci. China Inf. Sci.2