VLDB 2026 Research / reviewers in the wild / expert
Xuying Zhang
dblp:276/3115
· DBLP profile ↗
14ranked-venue papers
5as first author
13since 2021 · last 2025
0009-0009-3070-8767ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TAR3D: Creating High-Quality 3D Assets Via Next-Part PredictionabstractWe present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQ-VAE) and a Generative Pre-trained Transformer (GPT) to generate high-quality 3D assets. The core insight of this work is to migrate the multimodal unification and promising learning capabilities of the next-token prediction paradigm to conditional 3D object generation. To achieve this, the 3D VQ-VAE first encodes a wide range of 3D shapes into a compact triplane latent space and utilizes a set of discrete representations from a trainable codebook to reconstruct fine-grained geometries under the supervision of query point occupancy. Then, the 3D GPT, equipped with a custom triplane position embedding called TriPE, predicts the codebook index sequence with prefilling prompt tokens in an autoregressive manner so that the composition of 3D geometries can be modeled part by part. Extensive experiments on ShapeNet and Objaverse demonstrate that TAR3D can achieve superior generation quality over existing methods in text-to-3D and image-to-3D tasks Xuying Zhang, Yangguang Li 0001, Renrui Zhang, Kai Wang 0001, Wanli Ouyang, Zhiwei Xiong, Peng Gao 0007, Qibin Hou, Ming-Ming Cheng |
ICCV | 1 |
| 2025 | AR-1-to-3: Single Image to Consistent 3D Object via Next-View Prediction
Xuying Zhang, Yupeng Zhou, Kai Wang 0001, Zhen Li 0031, Shaohui Jiao, Daquan Zhou, Qibin Hou, Ming-Ming Cheng |
ICCV | 1 |
| 2025 | OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic SegmentationabstractRecent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for multiple visual modalities remains unexplored. In this paper, we propose a novel multi-modal learning framework, termed OmniSegmentor. It has two key innovations: 1) Based on ImageNet, we assemble a large-scale dataset for multi-modal pretraining, called OmniSegmentor, which contains five popular visual modalities; 2) We provide an efficient pretraining manner to endow the model with the capacity to encode different modality information in the OmniSegmentor. For the first time, we introduce a universal multi-modal pretraining framework that consistently amplifies the model's perceptual capabilities across various scenarios, regardless of the arbitrary combination of the involved modalities. Remarkably, our OmniSegmentor achieves new state-of-the-art records on a wide range of multi-modal semantic segmentation datasets, including NYU Depthv2, EventScape, MFNet, DeLiVER, SUNRGBD, and KITTI-360. Data, model checkpoints, and source code will be made publicly available: https://github.com/VCIP-RGBD/DFormer. Bowen Yin, Jiao-Long Cao, Xuying Zhang, Ming-Ming Cheng, Qibin Hou |
NeurIPS | 3 |
| 2025 | Camouflaged Object Detection with Adaptive Partition and Background Retrieval
Bowen Yin, Xuying Zhang, Li Liu 0002, Ming-Ming Cheng, Yongxiang Liu, Qibin Hou |
Int. J. Comput. Vis. | 2 |
| 2025 | Referring Camouflaged Object DetectionabstractWe consider the problem of referring camouflaged object detection (Ref-COD), a new task that aims to segment specified camouflaged objects based on a small set of referring images with salient target objects. We first assemble a large-scale dataset, called R2C7K, which consists of 7 K images covering 64 object categories in real-world scenarios. Then, we develop a simple but strong dual-branch framework, dubbed R2CNet, with a reference branch embedding the common representations of target objects from referring images and a segmentation branch identifying and segmenting camouflaged objects under the guidance of the common representations. In particular, we design a Referring Mask Generation module to generate pixel-level prior mask and a Referring Feature Enrichment module to enhance the capability of identifying specified camouflaged objects. Extensive experiments show the superiority of our Ref-COD methods over their COD counterparts in segmenting specified camouflaged objects and identifying the main body of target objects. Xuying Zhang, Bowen Yin, Zheng Lin 0005, Qibin Hou, Deng-Ping Fan, Ming-Ming Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | TeMO: Towards Text-Driven 3D Stylization for Multi-Object MeshesabstractRecent progress in the text-driven 3D stylization of a single object has been considerably promoted by CLIP-based methods. However, the stylization of multi-object 3D scenes is still impeded in that the image-text pairs used for pre-training CLIP mostly consist of an object. Meanwhile, the local details of multiple objects may be susceptible to omission due to the existing supervision manner primarily relying on coarse-grained contrast of image-text pairs. To overcome these challenges, we present a novel framework, dubbed TeMO, to parse multi-object 3D scenes and edit their styles under the contrast supervision at multiple levels. We first propose a Decoupled Graph Attention (DGA) module to distinguishably reinforce the features of 3D surface points. Particularly, a cross-modal graph is constructed to align the object points accurately and noun phrases decoupled from the 3D mesh and textual description. Then, we develop a Cross-Grained Contrast (CGC) supervision system, where a fine-grained loss between the words in the textual description and the randomly rendered images are constructed to complement the coarse-grained loss. Extensive experiments show that our method can synthesize high-quality stylized content and outperform the existing methods over a wide range of multi-object 3D meshes. Xuying Zhang, Bowen Yin, Zheng Lin 0005, Qibin Hou, Ming-Ming Cheng |
CVPR | 1 |
| 2024 | DFormer: Rethinking RGBD Representation Learning for Semantic SegmentationabstractWe present DFormer, a novel RGB-D pretraining framework to learn transferable representations for RGB-D segmentation tasks. DFormer has two new key innovations: 1) Unlike previous works that encode RGB-D information with RGB pretrained backbone, we pretrain the backbone using image-depth pairs from ImageNet-1K, and thus the DFormer is endowed with the capacity to encode RGB-D representations; 2) DFormer comprises a sequence of RGB-D blocks, which are tailored for encoding both RGB and depth information through a novel building block design. DFormer avoids the mismatched encoding of the 3D geometry relationships in depth maps by RGB pretrained backbones, which widely lies in existing methods but has not been resolved. We finetune the pretrained DFormer on two popular RGB-D tasks, i.e., RGB-D semantic segmentation and RGB-D salient object detection, with a lightweight decoder head. Experimental results show that our DFormer achieves new state-of-the-art performance on these two tasks with less than half of the computational cost of the current best methods on two RGB-D semantic segmentation datasets and five RGB-D salient object detection datasets. Code will be made publicly available. Bowen Yin, Xuying Zhang, Zhongyu Li 0006, Li Liu 0002, Ming-Ming Cheng, Qibin Hou |
ICLR | 2 |
| 2024 | No-Reference Segmentation Annotation Quality AssessmentabstractImage segmentation tasks aim to separate the image into masks that represent different objects or regions, where deep-learning-based methods have become mainstream. In the common practice, researchers utilize large-scale datasets including images along with their annotations to train their models, and evaluate the predictions with evaluation metrics. However, to our knowledge, no metrics have been proposed to assess the quality of the segmentation annotations, which will bring benefits to both the labeling and experimental process. In this paper, we fill this research gap and propose the first no-reference segmentation annotation quality assessment named SAQ. Based on our observation, we utilize the normal gradients of pixels on the annotation contours to represent the degree of fitting the real contours, which reflect the annotation accuracy. To alleviate the image differences, we adopt the gradient ranking score rather than directly using the gradient value. The multi-scale strategy is introduced to accommodate annotations of objects with different structures. Extensive experiments on datasets for various segmentation tasks have demonstrated the rationality of our proposed SAQ, and the assessment results of their annotation quality can serve as significant references for researchers. Zheng Lin 0005, Zheng-Peng Duan, Xuying Zhang, Luojun Lin |
ICME | 3 |
| 2024 | Wideband Power Spectrum Sensing: A Fast Practical Solution for Nyquist Folding ReceiverabstractThe limited availability of spectrum resources has been growing into a critical problem in wireless communications, remote sensing, and electronic surveillance, etc. To address the high-speed sampling bottleneck of wideband spectrum sensing, a fast and practical solution of power spectrum estimation for Nyquist folding receiver (NYFR) is proposed in this paper. The NYFR architecture can theoretically achieve full-band signal sensing with a hundred percent probability of intercept. However, the existing algorithm is difficult to realize in real-time due to its high complexity and complicated calculations. By exploring the sub-sampling principle inherent in NYFR, a computationally efficient method is introduced with compressive covariance sensing. That can be efficiently implemented via only the non-uniform fast Fourier transform, fast Fourier transform, and some simple multiplication operations. Meanwhile, the state-of-the-art power spectrum reconstruction model for NYFR of time-domain and frequency-domain is constructed in this paper as a comparison. Furthermore, the computational complexity of the proposed method scales linearly with the Nyquist-rate sampled number of samples and the sparsity of spectrum occupancy. Simulation results and discussion demonstrate that the low complexity in sampling and computation is a more practical solution to meet real-time wideband spectrum sensing applications. The proposed fast power spectrum sensing method for NYFR enables an increase in precision from 10-3 to 10-4 and a decrease in execution time from 10-1 to 10-3 seconds. Dechang Wang, Kailun Tian, Han Cong Feng, Yuxin Zhao 0006, Sen Cao, Xuying Zhang, Junyu Yuan, Bin Tang 0001 |
IEEE Internet Things J. | 8 |
| 2024 | CamoFormer: Masked Separable Attention for Camouflaged Object DetectionabstractHow to identify and segment camouflaged objects from the background is challenging. Inspired by the multi-head self-attention in Transformers, we present a simple masked separable attention (MSA) for camouflaged object detection. We first separate the multi-head self-attention into three parts, which are responsible for distinguishing the camouflaged objects from the background using different mask strategies. Furthermore, we propose to capture high-resolution semantic representations progressively based on a simple top-down decoder with the proposed MSA to attain precise segmentation results. These structures plus a backbone encoder form a new model, dubbed CamoFormer. Extensive experiments show that CamoFormer achieves new state-of-the-art performance on three widely-used camouflaged object detection benchmarks. To better evaluate the performance of the proposed CamoFormer around the border regions, we propose to use two new metrics, i.e., BR-M and BR-F. There are on average ∼ 5% relative improvements over previous methods in terms of S-measure and weighted F-measure. Bowen Yin, Xuying Zhang, Deng-Ping Fan, Shaohui Jiao, Ming-Ming Cheng, Luc Van Gool, Qibin Hou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Distributed UAV Swarm Augmented Wideband Spectrum Sensing Using Nyquist Folding ReceiverabstractDistributed unmanned aerial vehicle (UAV) swarms are formed by multiple UAVs with increased portability, higher levels of sensing capabilities, and more powerful autonomy. These features make them attractive for many applications, potentially increasing the shortage of spectrum resources. In this paper, wideband spectrum sensing augmented technology is discussed for distributed UAV swarms to improve the utilization of spectrum using Nyquist folding receiver (NYFR). Compared to the other sub-sampling schemes, the NYFR has low hardware complexity, power consumption, and high recovery efficiency for non-strictly sparse conditions. In particular, there is a focus discussion on the sensing model of two multichannel scenarios for the distributed UAV swarms, one with a complete functional receiver for the UAV swarm with reconfigurable intelligent surface (RIS), and another with a decentralized UAV swarm equipped with a complete functional receiver for each UAV element. Moreover, the property for multiple pulse reconstruction is analyzed through the Gershgorin circle theorem, especially for very short pulses. And the block sparse recovery property is analyzed for wide bandwidth signals. Experiment results show augmented spectrum sensing efficiency under non-strictly sparse conditions, and improve the processing capability for multiple signals and wide bandwidth signals while reducing interference from folded noise and subsampled harmonics. Kailun Tian, Han Cong Feng, Yuxin Zhao 0006, Dechang Wang, Sen Cao, Xuying Zhang, Junyu Yuan, Bin Tang 0001 |
IEEE Trans. Wirel. Commun. | 8 |
| 2022 | DIFNet: Boosting Visual Information Flow for Image CaptioningabstractCurrent Image Captioning (IC) methods predict textual words sequentially based on the input visual information from the visual feature extractor and the partially generated sentence information. However, for most cases, the partially generated sentence may dominate the target word prediction due to the insufficiency of visual information, making the generated descriptions irrelevant to the content of the given image. In this paper, we propose a Dual Information Flow Network (DIFNet11Source code is available at: https://github.com/mrwu-mac/DIFNet) to address this issue, which takes segmentation feature as another visual information source to enhance the contribution of visual information for prediction. To maximize the use of two information flows, we also propose an effective feature fusion module termed Iterative Independent Layer Normalization (IILN) which can condense the most relevant inputs while retraining modality-specific information in each flow. Experiments show that our method is able to enhance the dependence of prediction on visual information, making word prediction more focused on the visual content, and thus achieves new state-of-the-art performance on the MSCOCO dataset, e.g., 136.2 CIDEr on COCO Karpathy test split. Mingrui Wu, Xuying Zhang, Xiaoshuai Sun, Yiyi Zhou, Chao Chen 0026, Jiaxin Gu, Xing Sun 0001, Rongrong Ji |
CVPR | 2 |
| 2021 | RSTNet: Captioning With Adaptive Attention on Visual and Non-Visual WordsabstractRecent progress on visual question answering has explored the merits of grid features for vision language tasks. Meanwhile, transformer-based models have shown remarkable performance in various sequence prediction problems. However, the spatial information loss of grid features caused by flattening operation, as well as the defect of the transformer model in distinguishing visual words and non visual words, are still left unexplored. In this paper, we first propose Grid-Augmented (GA) module, in which relative geometry features between grids are incorporated to enhance visual representations. Then, we build a BERT-based language model to extract language context and propose Adaptive-Attention (AA) module on top of a transformer decoder to adaptively measure the contribution of visual and language cues before making decisions for word prediction. To prove the generality of our proposals, we apply the two modules to the vanilla transformer model to build our Relationship-Sensitive Transformer (RSTNet) for image captioning task. The proposed model is tested on the MSCOCO benchmark, where it achieves new state-of-art results on both the Karpathy test split and the online test server. Source code is available at GitHub1. Xuying Zhang, Xiaoshuai Sun, Yunpeng Luo, Jiayi Ji, Yiyi Zhou, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
CVPR | 1 |
| 2020 | Exploring Language Prior for Mode-Sensitive Visual Attention ModelingabstractModeling human visual attention mechanism is a fundamental problem for the understanding of human vision, which has also been demonstrated as an important module for various multimedia applications such as image captioning and visual question answering. In this paper, we propose a new probabilistic framework for attention, and introduce the concept ofmode to model the flexibility and adaptability of attention modulation in complex environments. We characterize the correlations between the visual input, the activated mode, the saliency and the spatial allocation of attention via a graphical model representation, based on which we explore the lingual guidance from captioning data for the implementation of a mode-sensitive attention (MSA) model. The proposed framework explicitly justifies the usage of center bias for fixation prediction and can convert an arbitrary learning-based backbone attention model to a more robust multi-mode version. Experimental results on the York120, MIT1003 and PASCAL datasets demonstrate the effectiveness of the proposed method. Xiaoshuai Sun, Xuying Zhang, Liujuan Cao, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
ACM Multimedia | 2 |