Xueqing Deng

dblp:209/9919 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
abstract
Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses dense, pixel-precise segmentation tasks, which typically involve category-guided or referral-based object segmentation. Although both directions are essential for developing models with human-level video comprehension, they have largely evolved separately, with distinct benchmarks and architectures. This paper aims to unify these efforts by introducing ViCaS, a new dataset containing thousands of challenging videos, each annotated with detailed, human-written captions and temporally consistent, pixel-accurate masks for multiple objects with phrase grounding. Our benchmark evaluates models on both holistic/high-level understanding and language-guided, pixel-precise segmentation. We also present carefully validated evaluation measures and propose an effective model architecture that can tackle our benchmark.
Ali Athar, Xueqing Deng, Liang-Chieh Chen
CVPR2
2025 Leveraging Panoptic Scene Graph for Evaluating Fine-Grained Text-to-Image Generation
Xueqing Deng, Qihang Yu, Liang-Chieh Chen
ICCV1
2025 Randomized Autoregressive Visual Generation
abstract
This paper presents Randomized AutoRegressive modeling (RAR) for visual generation, which sets a new state-of-the-art performance on the image generation task while maintaining full compatibility with language modeling frameworks. The proposed RAR is simple: during a standard autoregressive training process with a next-token prediction objective, the input sequence-typically ordered in raster form-is randomly permuted into different factorization orders with a probability r, where r starts at 1 and linearly decays to 0 over the course of training. This annealing training strategy enables the model to learn to maximize the expected likelihood over all factorization orders and thus effectively improve the model's capability of modeling bidirectional contexts. Importantly, RAR preserves the integrity of the autoregressive modeling framework, ensuring full compatibility with language modeling while significantly improving performance in image generation. On the ImageNet-256 benchmark, RAR achieves an FID score of 1.48, not only surpassing prior state-of-the-art autoregressive image generators but also outperforming leading diffusion-based and masked transformer-based methods. Code and models will be made available at https://github.com/bytedance/1d-tokenizer
Qihang Yu, Ju He, Xueqing Deng, Xiaohui Shen, Liang-Chieh Chen
ICCV3
2025 COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
abstract
This paper introduces the COCONut-PanCap dataset, created to enhance panoptic segmentation and grounded image captioning. Building upon the COCO dataset with advanced COCONut panoptic masks, this dataset aims to overcome limitations in existing image-text datasets that often lack detailed, scene-comprehensive descriptions. The COCONut-PanCap dataset incorporates fine-grained, region-level captions grounded in panoptic segmentation masks, ensuring consistency and improving the detail of generated captions.Through human-edited, densely annotated descriptions, COCONut-PanCap supports improved training of vision-language models (VLMs) for image understanding and generative models for text-to-image tasks.Experimental results demonstrate that COCONut-PanCap significantly boosts performance across understanding and generation tasks, offering complementary benefits to large-scale datasets. This dataset sets a new benchmark for evaluating models on joint panoptic segmentation and grounded captioning tasks, addressing the need for high-quality, detailed image-text annotations in multi-modal learning.
Xueqing Deng, Qihang Yu, Ali Athar, Xiaojie Jin 0004, Xiaohui Shen, Liang-Chieh Chen
NeurIPS1
2025 WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
abstract
Generative video modeling has made significant strides, yet ensuring structural and temporal consistency over long sequences remains a challenge. Current methods predominantly rely on RGB signals, leading to accumulated errors in object structure and motion over extended durations. To address these issues, we introduce WorldWeaver, a robust framework for long video generation that jointly models RGB frames and perceptual conditions within a unified long-horizon modeling scheme. Our training framework offers three key advantages. First, by jointly predicting perceptual conditions and color information from a unified representation, it significantly enhances temporal consistency and motion dynamics. Second, by leveraging depth cues, which we observe to be more resistant to drift than RGB, we construct a memory bank that preserves clearer contextual information, improving quality in long-horizon video generation. Third, we employ segmented noise scheduling for training prediction groups, which further mitigates drift and reduces computational cost. Extensive experiments on both diffusion and rectified flow-based models demonstrate the effectiveness of WorldWeaver in reducing temporal drift and improving the fidelity of generated videos.
Xueqing Deng, Shoufa Chen, Angtian Wang, Qiushan Guo, Mingfei Han 0003, Zeyue Xue, Mengzhao Chen, Ping Luo 0002
NeurIPS2
2024 COCONut: Modernizing COCO Segmentation
abstract
In recent decades, the vision community has witnessed remarkable progress in visual recognition, partially owing to advancements in dataset benchmarks. Notably, the established COCO benchmark has propelled the development of modern detection and segmentation systems. How-ever, the COCO segmentation benchmark has seen compar-atively slow improvement over the last decade. Originally equipped with coarse polygon annotations for ‘thing’ in-stances, it gradually incorporated coarse superpixel anno-tations for ‘stuff’ regions, which were subsequently heuris-tically amalgamated to yield panoptic segmentation anno-tations. These annotations, executed by different groups of raters, have resulted not only in coarse segmentation masks but also in inconsistencies between segmentation types. In this study, we undertake a comprehensive reeval-uation of the COCO segmentation annotations. By enhancing the annotation quality and expanding the dataset to encompass 383K images with more than 5.18M panoptic masks, we introduce COCONut, the COCO Next Universal segmenTation dataset. COCONut harmonizes segmentation annotations across semantic, instance, and panoptic segmentation with meticulously crafted high-quality masks, and establishes a robust benchmark for all segmentation tasks. To our knowledge, COCONut stands as the inaugural large-scale universal segmentation dataset, verified by hu-man raters. We anticipate that the release of COCONut will significantly contribute to the community's ability to assess the progress of novel neural networks.
Xueqing Deng, Qihang Yu, Xiaohui Shen, Liang-Chieh Chen
CVPR1
2024 MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
abstract
State-of-the-art video-text retrieval (VTR) methods typically involve fully fine-tuning a pre-trained model (e.g. CLIP) on specific datasets. However, this can result in significant storage costs in practical applications as a separate model per task must be stored. To address this issue, we present our pioneering work that enables parameter-efficient VTR using a pre-trained model, with only a small number of tunable parameters during training. Towards this goal, we propose a new method dubbed Multimodal Video Adapter (MV-Adapter) for efficiently transferring the knowledge in the pre-trained CLIP from image-text to video-text. Specifically, MV-Adapter utilizes bottleneck structures in both video and text branches, along with two novel components. The first is a Temporal Adaptation Module that is incorporated in the video branch to introduce global and local temporal contexts. We also train weights calibrations to adjust to dynamic variations across frames. The second is Cross Modality Tying that generates weights for video/text branches through sharing cross modality factors, for better aligning between modalities. Thanks to above innovations, MV-Adapter can achieve comparable or better performance than standard full fine-tuning with negligible parameters overhead. Notably, MV-Adapter consistently outperforms various competing methods in V2T/T2V tasks with large margins on five widely used VTR benchmarks (MSR-VTT, MSVD, LSMDC, DiDemo, and ActivityNet). Codes will be released.
Xiaojie Jin 0004, Weibo Gong, Xueqing Deng, Peng Wang 0037, Zhao Zhang 0001, Xiaohui Shen, Jiashi Feng
CVPR5
2024 Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences
abstract
Leveraging multi-view diffusion models as priors for 3D optimization have alleviated the problem of 3D consistency, e.g., the Janus face problem or the content drift problem, in zero-shot text-to-3D models. However, the 3D geomet-ric fidelity of the output remains an unresolved issue; albeit the rendered 2D views are realistic, the underlying geom-etry may contain errors such as unreasonable concavities. In this work, we propose CorrespondentDream, an effective method to leverage annotation-free, cross-view corre- spondences yielded from the diffusion U-Net to provide additional 3D prior to the NeRF optimization process. We find that these correspondences are strongly consistent with hu-man perception, and by adopting it in our loss design, we are able to produce NeRF models with geometries that are more coherent with common sense, e.g., more smoothed ob-ject surface, yielding higher 3D fidelity. We demonstrate the efficacy of our approach through various comparative qualitative results and a solid user study.
Kejie Li, Xueqing Deng, Yichun Shi, Minsu Cho
CVPR3
2024 An Image is Worth 32 Tokens for Reconstruction and Generation
abstract
Recent advancements in generative models have highlighted the crucial role of image tokenization in the efficient synthesis of high-resolution images. Tokenization, which transforms images into latent representations, reduces computational demands compared to directly processing pixels and enhances the effectiveness and efficiency of the generation process. Prior methods, such as VQGAN, typically utilize 2D latent grids with fixed downsampling factors. However, these 2D tokenizations face challenges in managing the inherent redundancies present in images, where adjacent regions frequently display similarities. To overcome this issue, we introduce **T**ransformer-based 1-D**i**mensional **Tok**enizer (TiTok), an innovative approach that tokenizes images into 1D latent sequences. TiTok provides a more compact latent representation, yielding substantially more efficient and effective representations than conventional techniques. For example, a 256 × 256 × 3 image can be reduced to just **32** discrete tokens, a significant reduction from the 256 or 1024 tokens obtained by prior methods. Despite its compact nature, TiTok achieves competitive performance to state-of-the-art approaches. Specifically, using the same generator framework, TiTok attains **1.97** gFID, outperforming MaskGIT baseline significantly by 4.21 at ImageNet 256 × 256 benchmark. The advantages of TiTok become even more significant when it comes to higher resolution. At ImageNet 512 × 512 benchmark, TiTok not only outperforms state-of-the-art diffusion model DiT-XL/2 (gFID 2.74 vs. 3.04), but also reduces the image tokens by 64×, leading to **410× faster** generation process. Our best-performing variant can significantly surpasses DiT-XL/2 (gFID **2.13** vs. 3.04) while still generating high-quality samples **74× faster**. Codes and models are available at https://github.com/bytedance/1d-tokenizer
Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, Liang-Chieh Chen
NeurIPS3
2023 Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP
abstract
Open-vocabulary segmentation is a challenging task requiring segmenting and recognizing objects from an open set of categories in diverse environments. One way to address this challenge is to leverage multi-modal models, such as CLIP, to provide image and text features in a shared embedding space, which effectively bridges the gap between closed-vocabulary and open-vocabulary recognition. Hence, existing methods often adopt a two-stage framework to tackle the problem, where the inputs first go through a mask generator and then through the CLIP model along with the predicted masks. This process involves extracting features from raw images multiple times, which can be ineffective and inefficient. By contrast, we propose to build everything into a single-stage framework using a _shared **F**rozen **C**onvolutional **CLIP** backbone_, which not only significantly simplifies the current two-stage pipeline, but also remarkably yields a better accuracy-cost trade-off. The resulting single-stage system, called FC-CLIP, benefits from the following observations: the _frozen_ CLIP backbone maintains the ability of open-vocabulary classification and can also serve as a strong mask generator, and the _convolutional_ CLIP generalizes well to a larger input resolution than the one used during contrastive image-text pretraining. Surprisingly, FC-CLIP advances state-of-the-art results on various benchmarks, while running practically fast. Specifically, when training on COCO panoptic data only and testing in a zero-shot manner, FC-CLIP achieve 26.8 PQ, 16.8 AP, and 34.1 mIoU on ADE20K, 18.2 PQ, 27.9 mIoU on Mapillary Vistas, 44.0 PQ, 26.8 AP, 56.2 mIoU on Cityscapes, outperforming the prior art under the same setting by +4.2 PQ, +2.4 AP, +4.2 mIoU on ADE20K, +4.0 PQ on Mapillary Vistas and +20.1 PQ on Cityscapes, respectively. Additionally, the training and testing time of FC-CLIP is 7.5x and 6.6x significantly faster than the same prior art, while using 5.9x fewer total model parameters. Meanwhile, FC-CLIP also sets a new state-of-the-art performance across various open-vocabulary semantic segmentation datasets. Code and models are available at https://github.com/bytedance/fc-clip
Qihang Yu, Ju He, Xueqing Deng, Xiaohui Shen, Liang-Chieh Chen
NeurIPS3
2022 NightLab: A Dual-level Architecture with Hardness Detection for Segmentation at Night
abstract
The semantic segmentation of nighttime scenes is a challenging problem that is key to impactful applications like self-driving cars. Yet, it has received little attention compared to its daytime counterpart. In this paper, we propose NightLab, a novel nighttime segmentation framework that leverages multiple deep learning models imbued with night-aware features to yield State-of-The-Art (SoTA) performance on multiple night segmentation benchmarks. Notably, NightLab contains models at two levels of granularity, i.e. image and regional, and each level is composed of light adaptation and segmentation modules. Given a nighttime image, the image level model provides an initial segmentation estimate while, in parallel, a hardness detection module identifies regions and their surrounding context that need further analysis. A regional level model focuses on these difficult regions to provide a significantly improved segmentation. All the models in NightLab are trained end-to-end using a set of proposed night-aware losses without handcrafted heuristics. Extensive experiments on the NightCity [44] and BDD100K [59] datasets show NightLab achieves SoTA performance compared to concurrent methods. Code and dataset are available at https://github.com/xdeng7/NightLab.
Xueqing Deng, Xiaochen Lian, Shawn D. Newsam
CVPR1
2022 DistPro: Searching a Fast Knowledge Distillation Process via Meta Optimization
Xueqing Deng, Shawn D. Newsam
ECCV (34)1
2021 Scale Aware Adaptation for Land-Cover Classification in Remote Sensing Imagery
abstract
Land-cover classification using remote sensing imagery is an important Earth observation task. Recently, land cover classification has benefited from the development of fully connected neural networks for semantic segmentation. The benchmark datasets available for training deep segmentation models in remote sensing imagery tend to be small, however, often consisting of only a handful of images from a single location with a single scale. This limits the models' ability to generalize to other datasets. Domain adaptation has been proposed to improve the models' generalization but we find these approaches are not effective for dealing with the scale variation commonly found between remote sensing image collections. We therefore propose a scale aware adversarial learning framework to perform joint cross-location and cross-scale land-cover classification. The framework has a dual discriminator architecture with a standard feature discriminator as well as a novel scale discriminator. We also introduce a scale attention module which produces scale-enhanced features. Experimental results show that the proposed frame-work outperforms state-of-the-art domain adaptation methods by a large margin. The open-sourced codes are available on Github: https://github.com/xdeng7/scale-aware_da.
Xueqing Deng, Yi Zhu 0001, Shawn D. Newsam
WACV1
2021 aPRBind: protein-RNA interface prediction by combining sequence and I-TASSER model-based structural features learned with convolutional neural networks
abstract
MOTIVATION: Protein-RNA interactions play a critical role in various biological processes. The accurate prediction of RNA-binding residues in proteins has been one of the most challenging and intriguing problems in the field of computational biology. The existing methods still have a relatively low accuracy especially for the sequence-based ab-initio methods. RESULTS: In this work, we propose an approach aPRBind, a convolutional neural network-based ab-initio method for RNA-binding residue prediction. aPRBind is trained with sequence features and structural ones (particularly including residue dynamics information and residue-nucleotide propensity developed by us) that are extracted from the predicted structures by I-TASSER. The analysis of feature contributions indicates the sequence features are most important, followed by dynamics information, and the sequence and structural features are complementary in binding site prediction. The performance comparison of our method with other peer ones on benchmark dataset shows that aPRBind outperforms some state-of-the-art ab-initio methods. Additionally, aPRBind can give a better prediction for the modeled structures with TM-score≥0.5, and meanwhile since the structural features are not very sensitive to the refined 3D structures, aPRBind has only a marginal dependence on the accuracy of the structure model, which allows aPRBind to be applied to the RNA-binding site prediction for the modeled or unbound structures. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/ChunhuaLiLab/aPRbind. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Weikang Gong, Yanpeng Zhao, Xueqing Deng, Chunhua Li 0001
Bioinform.4
2020 Cross-Time and Orientation-Invariant Overhead Image Geolocalization Using Deep Local Features
abstract
Overhead image geolocalization is becoming increasingly important due to the growing collection of drone imagery without location information. In this paper, we perform large-scale overhead image geolocalization by matching a query image to wide-area reference imagery with known location. We use deep local features so that the query image need not align with but only overlap the tiled reference imagery. We further address two key challenges. For when the query and reference imagery are from different dates, we perform cross-time geolocalization using time invariant features learned using a Siamese network. For when the query and reference imagery are oriented differently, we introduce an orientation normalization network. We demonstrate our contributions on two new high-resolution overhead image datasets. Our method significantly outperforms strong baselines on cross-time geolocalization and is shown to exhibit promising orientation invariance.
Xueqing Deng, Yi Zhu 0001, Shawn D. Newsam
WACV2
2020 Analyses on clustering of the conserved residues at protein-RNA interfaces and its application in binding site identification
abstract
BACKGROUND: The maintenance of protein structural stability requires the cooperativity among spatially neighboring residues. Previous studies have shown that conserved residues tend to occur clustered together within enzyme active sites and protein-protein/DNA interfaces. It is possible that conserved residues form one or more local clusters in protein tertiary structures as it can facilitate the formation of functional motifs. In this work, we systematically investigate the spatial distributions of conserved residues as well as hot spot ones within protein-RNA interfaces. RESULTS: The analysis of 191 polypeptide chains from 160 complexes shows the polypeptides interacting with tRNAs evolve relatively rapidly. A statistical analysis of residues in different regions shows that the interface residues are often more conserved, while the most conserved ones are those occurring at protein interiors which maintain the stability of folded polypeptide chains. Additionally, we found that 77.8% of the interfaces have the conserved residues clustered within the entire interface regions. Appling the clustering characteristics to the identification of the real interface, there are 31.1% of cases where the real interfaces are ranked in top 10% of 1000 randomly generated surface patches. In the conserved clusters, the preferred residues are the hydrophobic (Leu, Ile, Met), aromatic (Tyr, Phe, Trp) and interestingly only one positively charged Arg residues. For the hot spot residues, 51.5% of them are situated in the conserved residue clusters, and they are largely consistent with the preferred residue types in the conserved clusters. CONCLUSIONS: The protein-RNA interface residues are often more conserved than non-interface surface ones. The conserved interface residues occur more spatially clustered relative to the entire interface residues. The high consistence of hot spot residue types and the preferred residue types in the conserved clusters has important implications for the experimental alanine scanning mutagenesis study. This work deepens the understanding of the residual organization at protein-RNA interface and is of potential applications in the identification of binding site and hot spot residues.
Xueqing Deng, Weikang Gong, Chunhua Li 0001
BMC Bioinform.2
2019 Large Scale Unsupervised Domain Adaptation of Segmentation Networks with Adversarial Learning
abstract
Most current state-of-the-art methods for semantic segmentation on remote sensing imagery require large labeled data, which is scarcely available. Due to the distribution shifting phenomenon inherent in remote sensing imagery, the reuse of pre-trained models on new areas of interest rarely yield satisfactory results. In this paper, we approach this problem from an adversarial learning perspective toward unsupervised domain adaptation. The core concept is to infuse fully convolutional neural networks and adversarial networks for semantic segmentation assuming the structures in the scene and objects of interest are similar in two set of images. Models are trained on a source dataset where ground truth is available and adapted to new target dataset iteratively via a adversarial loss on unlabeled samples. We use two real large scale datasets to validate the framework: 1) cross city road extraction and 2) cross country building extraction. The preliminary results show the usefulness of considering adversarial learning for indirect re-use of the pre-trained models. Experimental validation suggests significant benefits over models without adaptation.
Xueqing Deng, Hsiuhan Lexie Yang, Nikhil Makkar, Dalton D. Lunga
IGARSS1
2019 Fine-Grained Land Use Classification at the City Scale Using Ground-Level Images
abstract
Multimedia researchers have exploited large collections of community-contributed geo-referenced images to better understand a particular image, such as its subject matter or where it was taken, as well as to better understand a geographic location, such as the most visited tourist spots in a city or what the local cuisine is like. The goal of this paper is to better understand location. In particular, we use geo-referenced image collections to better understand what occurs in different parts of a city at fine spatial and activity class scales. This problem is known as land use mapping in the geographical sciences. We propose a novel framework to perform fine-grained land use mapping at the city scale using ground-level images. Mapping land use is considerably more difficult than mapping land cover and is generally not possible using overhead imagery as it requires close-up views and seeing inside buildings. We postulate that the growing collections of geo-referenced, ground-level images suggest an alternate approach to this geographic knowledge discovery problem. We develop a general framework that uses Flickr images to map 45 different land-use classes for the city of San Francisco, CA, USA. Individual images are classified using a novel convolutional neural network containing two streams: one for recognizing objects and another for recognizing scenes. This network is trained in an end-to-end manner directly on the labeled training images. We propose several novel strategies to overcome the noisiness of our user-generated data including search-based training set augmentation and online adaptive training. We derive a ground truth map of San Francisco in order to evaluate our method. We demonstrate the effectiveness of our approach through geovisualization and quantitative analysis. Our framework achieves over 29% recall at the individual land parcel level that represents a strong baseline for the challenging 45-way land use classification problem, especially given the noisiness of the image data.
Yi Zhu 0001, Xueqing Deng, Shawn D. Newsam
IEEE Trans. Multim.2
2018 What is it like down there?: generating dense ground-level views and image features from overhead imagery using conditional generative adversarial networks
abstract
This paper investigates conditional generative adversarial networks (cGANs) to overcome a fundamental limitation of using geotagged media for geographic discovery, namely its sparse and uneven spatial distribution. We train a cGAN to generate ground-level views of a location given overhead imagery. We show the "fake" ground-level images are natural looking and are structurally similar to the real images. More significantly, we show the generated images are representative of the locations and that the representations learned by the cGANs are informative. In particular, we show that dense feature maps generated using our framework are more effective for land-cover classification than approaches which spatially interpolate features extracted from sparse ground-level images. To our knowledge, ours is the first work to use cGANs to generate ground-level views given overhead imagery in order to explore the benefits of the learned representations.
Xueqing Deng, Yi Zhu 0001, Shawn D. Newsam
SIGSPATIAL/GIS1
2018 Spatial Morphing Kernel Regression for Feature Interpolation
abstract
In recent years, geotagged social media has become popular as a novel source for geographic knowledge discovery. Ground-level images and videos provide a different perspective than overhead imagery and can be applied to a range of applications such as land use mapping, activity detection, pollution mapping, etc. The sparse and uneven distribution of this data presents a problem, however, for generating dense maps. We therefore investigate the problem of spatially interpolating the high-dimensional features extracted from sparse social media to enable dense labeling using standard classifiers. Further, we show how prior knowledge about region boundaries can be used to improve the interpolation through spatial morphing kernel regression. We show that an interpolate-then-classify framework can produce dense maps from sparse observations but that care must be taken in choosing the interpolation method. We also show that the spatial morphing kernel improves the results.
Xueqing Deng, Yi Zhu 0001, Shawn D. Newsam
ICIP1