EDBT 2026 Demo / reviewers in the wild / expert
Xinhua Cheng
dblp:260/2943
· DBLP profile ↗
26ranked-venue papers
7as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 360Explorer: Exploring 4D Controllable World in Panoramic VideosabstractWe present 360Explorer, a novel approach for generating 4D controllable panoramic videos conditioned on user-provided 3D instructions for exploring and manipulating dynamic worlds. Compared to existing perspective-based methods struggle to address spatial consistency during camera rotation in place, we introduce the panoramic view in controllable video generation models to inherently maintain the view recall consistency. By introducing dynamic point clouds as the 4D scene representations, 360Explorer unifies the modeling of camera transformations and object movements as incomplete renders to describe precise control instructions in 3D worlds. To tackle the data limitation in acquiring multi-viewpoint panoramic videos, we further propose a reverse warping strategy to construct the training dataset on easily accessible monocular panoramic videos. Extensive experiments demonstrate that 360Explorer achieves superior performance in creating 4D controllable panoramic videos with camera transformation and object movements aligned with diverse provided instructions. Xinhua Cheng, Haiyang Zhou, Wangbo Yu, Tanghui Jia, Bin Lin 0014, Yunyang Ge, Li Yuan 0007 |
AAAI | 1 |
| 2026 | NeuralGS: Bridging Neural Fields and 3D Gaussian Splatting for Compact 3D Representationsabstract3D Gaussian Splatting (3DGS) achieves impressive quality and rendering speed, but with millions of 3D Gaussians and significant storage and transmission costs. In this paper, we aim to develop a simple yet effective method called NeuralGS that compresses the original 3DGS into a compact representation. Our observation is that neural fields like NeRF can represent complex 3D scenes with Multi-Layer Perceptron (MLP) neural networks using only a few megabytes. Thus, NeuralGS effectively adopts the neural field representation to encode the attributes of 3D Gaussians with MLPs, only requiring a small storage size even for a large-scale scene. To achieve this, we adopt a clustering strategy and fit the Gaussians within each cluster using different tiny MLPs, based on importance scores of Gaussians as fitting weights. We experiment on multiple datasets, achieving a 91$\times$ average model size reduction without harming the visual quality. Zhenyu Tang 0004, Chaoran Feng 0001, Xinhua Cheng, Wangbo Yu, Junwu Zhang, Yuan Liu 0025, Xiaoxiao Long, Wenping Wang 0001, Li Yuan 0007 |
AAAI | 3 |
| 2026 | HoloDreamer: Holistic 3D Panoramic Scene Generation From Text Descriptionsabstract3D scene generation is in high demand across various domains, including virtual reality, gaming, and the film industry. Owing to the powerful generative capabilities of text-to-image diffusion models that provide reliable priors, creating 3D scenes using only text prompts has become viable, thereby significantly advancing research in text-driven 3D scene generation. Prevailing methods typically employ the diffusion model to generate an initial local image, followed by iteratively outpainting the local image to gradually generate scenes. Nevertheless, these outpainting-based approaches are prone to producing globally inconsistent results with low completeness, restricting their broader applications. To tackle these problems, we introduce HoloDreamer, a framework that begins by generating a high-definition panorama to holistically initialize the full scene, and leverages 3D Gaussian Splatting (3D-GS) for rapid 3D scene reconstruction, thereby facilitating the creation of view-consistent and fully enclosed 3D scenes. Specifically, we propose Stylized Equirectangular Panorama Generation, a pipeline that combines multiple diffusion models to enable stylized and detailed equirectangular panorama generation from complex text prompts. Subsequently, Enhanced Two-Stage Panorama Reconstruction is introduced, conducting a two-stage optimization of 3D-GS to inpaint the missing region and enhance the integrity of the scene. Comprehensive experiments demonstrated that our method outperforms prior works in terms of overall visual consistency and harmony, as well as reconstruction quality and rendering robustness when generating fully enclosed scenes. Haiyang Zhou, Xinhua Cheng, Wangbo Yu, Yonghong Tian 0001, Li Yuan 0007 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger ScenesabstractCompared to frame-based methods, computational neuromorphic imaging using event cameras offers significant advantages, such as minimal motion blur, enhanced temporal resolution, and high dynamic range. The multi-view consistency of Neural Radiance Fields combined with the unique benefits of event cameras, has spurred recent research into reconstructing NeRF from data captured by moving event cameras. While showing impressive performance, existing methods rely on ideal conditions with the availability of uniform and high-quality event sequences and accurate camera poses, and mainly focus on the object level reconstruction, thus limiting their practical applications. In this work, we propose AE-NeRF to address the challenges of learning event-based NeRF from non-ideal conditions, including non-uniform event sequences, noisy poses, and various scales of scenes. Our method exploits the density of event streams and jointly learn a pose correction module with an event-based NeRF (e-NeRF) framework for robust 3D reconstruction from inaccurate camera poses. To generalize to larger scenes, we propose hierarchical event distillation with a proposal e-NeRF network and a vanilla e-NeRF network to resample and refine the reconstruction process. We further propose an event reconstruction loss and a temporal loss to improve the view consistency of the reconstructed scene. We established a comprehensive benchmark that includes large-scale scenes to simulate practical non-ideal conditions, incorporating both synthetic and challenging real-world event datasets. The experimental results show that our method achieves a new state-of-the-art in event-based 3D reconstruction. Chaoran Feng 0001, Wangbo Yu, Xinhua Cheng, Zhenyu Tang 0004, Junwu Zhang, Li Yuan 0007, Yonghong Tian 0001 |
AAAI | 3 |
| 2025 | Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction CycleabstractRecent 3D large reconstruction models typically employ a two-stage process, including first generate multi-view images by a multi-view diffusion model, and then utilize a feed-forward model to reconstruct images to 3D content. However, multi-view diffusion models often produce low-quality and inconsistent images, adversely affecting the quality of the final 3D reconstruction. To address this issue, we propose a unified 3D generation framework called Cycle3D, which cyclically utilizes a 2D diffusion-based generation module and a feed-forward 3D reconstruction module during the multi-step diffusion process. Concretely, 2D diffusion model is applied for generating high-quality texture, and the reconstruction model guarantees multi-view consistency. Moreover, 2D diffusion model can further control the generated content and inject reference-view information for unseen views, thereby enhancing the diversity and texture consistency of 3D generation during the denoising process. Extensive experiments demonstrate the superior ability of our method to create 3D content with high-quality and consistency compared with state-of-the-art baselines. Zhenyu Tang 0004, Junwu Zhang, Xinhua Cheng, Wangbo Yu, Chaoran Feng 0001, Yatian Pang, Bin Lin 0014, Li Yuan 0007 |
AAAI | 3 |
| 2025 | RoomPainter: View-Integrated Diffusion for Consistent Indoor Scene TexturingabstractIndoor scene texture synthesis has garnered significant interest due to its important potential applications in virtual reality, digital media and creative arts. Existing diffusion-model-based researches either rely on per-view inpainting techniques, which are plagued by severe cross-view inconsistencies and conspicuous seams, or adopt optimization-based approaches that involve substantial computational overhead. In this work, we present RoomPainter, a frame-work that seamlessly integrates efficiency and consistency to achieve high-fidelity texturing of indoor scenes. The core of RoomPainter features a zero-shot technique that effectively adapts a 2D diffusion model for 3D-consistent texture synthesis, along with a two-stage generation strategy that ensures both global and local consistency. Specifically, we introduce Attention-Guided Multi-View Integrated Sampling (MVIS) combined with a neighbor-integrated attention mechanism for zero-shot texture map generation. Using the MVIS, we firstly generate texture map for the entire room to ensure global consistency, then adopt its variant, namely Attention-Guided Multi-View Integrated Repaint Sampling (MVRS) to repaint individual instances within the room, thereby further enhancing local consistency and addressing the occlusion problem. Experiments demonstrate that RoomPainter achieves superior performance for indoor scene texture synthesis in visual quality, global consistency and generation efficiency. Zhipeng Huang 0001, Wangbo Yu, Xinhua Cheng, ChengShu Zhao, Yunyang Ge, Mingyi Guo, Li Yuan 0007, Yonghong Tian 0001 |
CVPR | 3 |
| 2025 | WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion ModelabstractVideo Variational Autoencoder (VAE) encodes videos into a low-dimensional latent space, becoming a key component of most Latent Video Diffusion Models (LVDMs) to reduce model training costs. However, as the resolution and duration of generated videos increase, the encoding cost of Video VAEs becomes a limiting bottleneck in training LVDMs. Moreover, the block-wise inference method adopted by most LVDMs can lead to discontinuities of latent space when processing long-duration videos. The key to addressing the computational bottleneck lies in decomposing videos into distinct components and efficiently encoding the critical information. Wavelet transform can decompose videos into multiple frequency-domain components and improve the efficiency significantly, we thus propose Wavelet Flow VAE (WF-VAE), an autoencoder that leverages multi-level wavelet transform to facilitate low-frequency energy flow into latent representation. Furthermore, we introduce a method called Causal Cache, which maintains the integrity of latent space during block-wise inference. Compared to state-of-the-art video VAEs, WF-VAE demonstrates superior performance in both PSNR and LPIPS metrics, achieving 2× higher throughput and 4× lower memory consumption while maintaining competitive reconstruction quality. Our code and models are available at https://github.com/PKU-YuanGroup/WF-VAE. Zongjian Li, Bin Lin 0014, Liuhan Chen, Xinhua Cheng, Shenghai Yuan 0002, Li Yuan 0007 |
CVPR | 5 |
| 2025 | OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion ModelabstractVariational Autoencoder (VAE), compressing videos into latent representations, is a crucial preceding component of Latent Video Diffusion Models (LVDMs). However, most LVDMs utilize 2D image VAE, which only compresses video spatially. This will lead to temporally redundant representations, reducing the efficiency of LVDMs. To eliminate this issue, we propose an omni-dimension compression VAE, named OD-VAE, which can temporally and spatially compress videos based on 3D-Causal-CNN architecture. To obtain a better trade-off between video reconstruction quality and compression speed, we further introduce and analyze four model variants of OD-VAE. In addition, a novel initialization method is designed to train our OD-VAE more efficiently, and a novel inference strategy is proposed to enable OD-VAE to handle videos of arbitrary length with limited GPU memory. Comprehensive experiments on video reconstruction and LVDM-based video generation demonstrate the effectiveness and efficiency of our proposed methods. The source code and models are available at here. Liuhan Chen, Zongjian Li, Bin Lin 0014, Qian Wang 0062, Shenghai Yuan 0002, Xinhua Cheng, Li Yuan 0007 |
ICME | 8 |
| 2025 | HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation
Haiyang Zhou, Wangbo Yu, Jiawen Guan, Xinhua Cheng, Yonghong Tian 0001, Li Yuan 0007 |
ACM Multimedia | 4 |
| 2025 | MagicTime: Time-Lapse Video Generation Models as Metamorphic SimulatorsabstractRecent advances in text-to-video generation (T2V) have achieved remarkable success in synthesizing high-quality general videos from textual descriptions. A largely overlooked problem in T2V is that existing models have not adequately encoded physical knowledge of the real world, thus generated videos tend to have limited motion and poor variations. In this paper, we propose MagicTime, a metamorphic time-lapse video generation model, which learns real-world physics knowledge from time-lapse videos and implements metamorphic generation. First, we design a simple yet effective two-stage Magic Adaptive Strategy, encode more physical knowledge from metamorphic videos, and transform pre-trained T2V models to generate metamorphic videos. Second, we introduce a Dynamic Frames Extraction strategy to adapt to metamorphic time-lapse videos, which have a wider variation range and cover dramatic object metamorphic processes, thus embodying more physical knowledge than general videos. Finally, we introduce a Magic Text-Encoder to improve the understanding of metamorphic video prompts. Furthermore, we create a time-lapse video-text dataset called ChronoMagic, specifically curated to unlock the metamorphic video generation ability. Extensive experiments demonstrate the superiority and effectiveness of MagicTime for generating high-quality and dynamic metamorphic videos, suggesting time-lapse video generation is a promising path toward building metamorphic simulators of the physical world. Shenghai Yuan 0002, Jinfa Huang, Yujun Shi, Yongqi Xu, Rui-Jie Zhu 0003, Bin Lin 0014, Xinhua Cheng, Li Yuan 0007, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion ModelabstractPanorama video recently attracts more interest in both study and application, courtesy of its immersive experience. Due to the expensive cost of capturing$360^{\circ}$panoramic videos, generating desirable panorama videos by prompts is urgently required. Lately, the emerging text-to-video (T2V) diffusion methods demonstrate notable effectiveness in standard video generation. However, due to the significant gap in content and motion patterns between panoramic and standard videos, these methods encounter challenges in yielding satisfactory$360^{\circ}$panoramic videos. In this paper, we propose a pipeline named 360-Degree Video Diffusion model (360DVD) for generating$360^{\circ}$panoramic videos based on the given prompts and motion conditions. Specifically, we introduce a lightweight 360-Adapter accompanied by 360 Enhancement Techniques to transform pre-trained T2V models for panorama video generation. We further propose a new panorama dataset named WEB360 consisting of panoramic video-text pairs for training 360DVD, addressing the absence of captioned panoramic video datasets. Extensive experiments demonstrate the superior-ity and effectiveness of 360DVD for panorama video gen-eration. Our project page is at https: / /akaneqwq. github. io/360DVD/. Chong Mou, Xinhua Cheng, Jian Zhang 0018 |
CVPR | 4 |
| 2024 | Repaint123: Fast and High-Quality One Image to 3D Generation with Progressive Controllable Repainting
Junwu Zhang, Zhenyu Tang 0004, Yatian Pang, Xinhua Cheng, Peng Jin 0001, Yida Wei, Munan Ning, Li Yuan 0007 |
ECCV (25) | 4 |
| 2024 | Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic PromptsabstractRecent text-to-3D generation methods achieve impressive 3D content creation capacity thanks to the advances in image diffusion models and optimizing strategies. However, current methods struggle to generate correct 3D content for a complex prompt in semantics, i.e., a prompt describing multiple interacted objects binding with different attributes. In this work, we propose a general framework named Progressive3D, which decomposes the entire generation into a series of locally progressive editing steps to create precise 3D content for complex prompts, and we constrain the content change to only occur in regions determined by user-defined region prompts in each editing step. Furthermore, we propose an overlapped semantic component suppression technique to encourage the optimization process to focus more on the semantic differences between prompts. Extensive experiments demonstrate that the proposed Progressive3D framework generates precise 3D content for prompts with complex semantics through progressive editing steps and is general for various text-to-3D methods driven by different 3D representations. Xinhua Cheng, Tianyu Yang 0003, Yu Li 0003, Lei Zhang 0001, Jian Zhang 0018, Li Yuan 0007 |
ICLR | 1 |
| 2024 | ResVR: Joint Rescaling and Viewport Rendering of Omnidirectional ImagesabstractWith the advent of virtual reality technology, omnidirectional image (ODI) rescaling techniques are increasingly embraced to reduce transmitted and stored file sizes while preserving high image quality. Despite this progress, current ODI rescaling methods predominantly focus on enhancing the quality of images in equirectangular projection (ERP) format, which overlooks the fact that the content viewed on head-mounted displays (HMDs) is actually a rendered viewport instead of an ERP image. In this work, we emphasize that focusing solely on ERP quality results in inferior viewport visual experiences for users. Thus, we propose ResVR, which is the first comprehensive framework for the joint Rescaling and Viewport Rendering of ODIs. ResVR allows obtaining LR ERP images for transmission while rendering high-quality viewports for users to watch on HMDs. In our ResVR, a novel discrete pixel sampling strategy is developed to tackle the complex mapping between the viewport and ERP, enabling end-to-end training of the ResVR pipeline. Furthermore, a spherical pixel shape representation technique is innovatively derived from spherical differentiation to significantly improve the visual quality of rendered viewports. Extensive experiments demonstrate that our ResVR outperforms existing methods in viewport rendering tasks across different fields of view, resolutions, and view directions while keeping a low transmission overhead. Code is available at https://github.com/lwq20020127/ResVR. Shijie Zhao 0001, Bin Chen 0006, Xinhua Cheng, Li Zhang 0006, Jian Zhang 0018 |
ACM Multimedia | 4 |
| 2024 | Prompt2Poster: Automatically Artistic Chinese Poster Creation from Prompt OnlyabstractAs a critical component in graphic design, artistic posters are widely applied in the advertising and entertainment industry, thus the automatic poster creation from user-provided prompts has become increasingly desired recently. Although existing Text2Image methods create impressive images aligned with given prompts, they fail to generate ideal artistic posters, especially with Chinese texts. To create desired artistic Chinese posters including an aligned background, reasonable layouts, and stylized graphical texts from given prompts only, we propose an automatic poster creation framework, named Prompt2Poster. Our framework utilizes the capacity of the powerful Large Language Model (LLM) to extract user intention from provided prompts and generate the aligned background. Although only taking a user prompt as the input, linguistic, visual, and geometrical information is fully utilized in the framework, bringing the ability to fit different distributions. To achieve the use of multi-modal information in the framework, two carefully designed modules, Controllable Layout Generator (CLG) and Graphical Text Generator (GTG) are proposed, leading to accurate and pleasurable visual results. Comprehensive experiments demonstrate that our Prompt2Poster achieves superior performance, especially in text quality and visual harmony. Yunyang Ge, Liuhan Chen, Haiyang Zhou, Qian Wang 0062, Xinhua Cheng, Li Yuan 0007 |
ACM Multimedia | 6 |
| 2024 | OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary UnderstandingabstractThis paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) that possesses the capability for 3D point-level open vocabulary understanding. Our primary motivation stems from observing that existing 3DGS-based open vocabulary methods mainly focus on 2D pixel-level parsing. These methods struggle with 3D point-level tasks due to weak feature expressiveness and inaccurate 2D-3D feature associations. To ensure robust feature presentation and 3D point-level understanding, we first employ SAM masks without cross-frame associations to train instance features with 3D consistency. These features exhibit both intra-object consistency and inter-object distinction. Then, we propose a two-stage codebook to discretize these features from coarse to fine levels. At the coarse level, we consider the positional information of 3D points to achieve location-based clustering, which is then refined at the fine level.
Finally, we introduce an instance-level 3D-2D feature association method that links 3D points to 2D masks, which are further associated with 2D CLIP features. Extensive experiments, including open vocabulary-based 3D object selection, 3D point cloud understanding, click-based 3D object selection, and ablation studies, demonstrate the effectiveness of our proposed method. The source code is available at our project page https://3d-aigc.github.io/OpenGaussian. Yanmin Wu, Jiarui Meng, Haijie Li, Chenming Wu, Yahao Shi, Xinhua Cheng, Chen Zhao 0011, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Jian Zhang 0018 |
NeurIPS | 6 |
| 2024 | ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video GenerationabstractWe propose a novel text-to-video (T2V) generation benchmark, ChronoMagic-Bench, to evaluate the temporal and metamorphic knowledge skills in time-lapse video generation of the T2V models (e.g. Sora and Lumiere). Compared to existing benchmarks that focus on visual quality and text relevance of generated videos, ChronoMagic-Bench focuses on the models’ ability to generate time-lapse videos with significant metamorphic amplitude and temporal coherence. The benchmark probes T2V models for their physics, biology, and chemistry capabilities, in a free-form text control. For these purposes, ChronoMagic-Bench introduces 1,649 prompts and real-world videos as references, categorized into four major types of time-lapse videos: biological, human creation, meteorological, and physical phenomena, which are further divided into 75 subcategories. This categorization ensures a comprehensive evaluation of the models’ capacity to handle diverse and complex transformations. To accurately align human preference on the benchmark, we introduce two new automatic metrics, MTScore and CHScore, to evaluate the videos' metamorphic attributes and temporal coherence. MTScore measures the metamorphic amplitude, reflecting the degree of change over time, while CHScore assesses the temporal coherence, ensuring the generated videos maintain logical progression and continuity. Based on the ChronoMagic-Bench, we conduct comprehensive manual evaluations of eighteen representative T2V models, revealing their strengths and weaknesses across different categories of prompts, providing a thorough evaluation framework that addresses current gaps in video generation research. More encouragingly, we create a large-scale ChronoMagic-Pro dataset, containing 460k high-quality pairs of 720p time-lapse videos and detailed captions. Each caption ensures high physical content and large metamorphic amplitude, which have a far-reaching impact on the video generation community. The source data and code are publicly available on https://pku-yuangroup.github.io/ChronoMagic-Bench. Shenghai Yuan 0002, Jinfa Huang, Yongqi Xu, Yaoyang Liu, Shaofeng Zhang, Yujun Shi, Rui-Jie Zhu 0003, Xinhua Cheng, Jiebo Luo 0001, Li Yuan 0007 |
NeurIPS | 8 |
| 2023 | Semi-attention Partition for Occluded Person Re-identificationabstractThis paper proposes a Semi-Attention Partition (SAP) method to learn well-aligned part features for occluded person re-identification (re-ID). Currently, the mainstream methods employ either external semantic partition or attention-based partition, and the latter manner is usually better than the former one. Under this background, this paper explores a potential that the weak semantic partition can be a good teacher for the strong attention-based partition. In other words, the attention-based student can substantially surpass its noisy semantic-based teacher, contradicting the common sense that the student usually achieves inferior (or comparable) accuracy. A key to this effect is: the proposed SAP encourages the attention-based partition of the (transformer) student to be partially consistent with the semantic-based teacher partition through knowledge distillation, yielding the so-called semi-attention. Such partial consistency allows the student to have both consistency and reasonable conflict with the noisy teacher. More specifically, on the one hand, the attention is guided by the semantic partition from the teacher. On the other hand, the attention mechanism itself still has some degree of freedom to comply with the inherent similarity between different patches, thus gaining resistance against noisy supervision. Moreover, we integrate a battery of well-engineered designs into SAP to reinforce their cooperation (e.g., multiple forms of teacher-student consistency), as well as to promote reasonable conflict (e.g., mutual absorbing partition refinement and a supervision signal dropout strategy). Experimental results confirm that the transformer student achieves substantial improvement after this semi-attention learning scheme, and produces new state-of-the-art accuracy on several standard re-ID benchmarks. Mengxi Jia, Yifan Sun 0003, Yunpeng Zhai, Xinhua Cheng, Yi Yang 0001, Ying Li 0012 |
AAAI | 4 |
| 2023 | Panoptic Compositional Feature Field for Editable Scene Rendering with Network-Inferred Labels via Metric LearningabstractDespite neural implicit representations demonstrating impressive high-quality view synthesis capacity, decom-posing such representations into objects for instance-level editing is still challenging. Recent works learn object-compositional representations supervised by ground truth instance annotations and produce promising scene editing results. However, ground truth annotations are manually labeled and expensive in practice, which limits their usage in real-world scenes. In this work, we attempt to learn an object-compositional neural implicit representation for editable scene rendering by leveraging labels inferred from the off-the-shelf 2D panoptic segmentation networks instead of the ground truth annotations. We propose a novel framework named Panoptic Compositional Feature Field (PCFF), which introduces an instance quadruplet metric learning to build a discriminating panoptic feature space for reliable scene editing. In addition, we propose semantic-related strategies to further exploit the correlations between semantic and appearance attributes for achieving better rendering results. Experiments on multiple scene datasets including ScanNet, Replica, and ToyDesk demonstrate that our proposed method achieves superior performance for novel view synthesis and produces convincing real-world scene editing results. Xinhua Cheng, Yanmin Wu, Mengxi Jia, Jian Zhang 0018 |
CVPR | 1 |
| 2023 | EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Groundingabstract3D visual grounding aims to find the object within point clouds mentioned by free-form natural language descriptions with rich semantic cues. However, existing methods either extract the sentence-level features coupling all words or focus more on object names, which would lose the word-level information or neglect other attributes. To alleviate these issues, we present EDA that Explicitly Decouples the textual attributes in a sentence and conducts Dense Alignment between such fine-grained language and point cloud objects. Specifically, we first propose a text decoupling module to produce textual features for every semantic component. Then, we design two losses to supervise the dense matching between two modalities: position alignment loss and semantic alignment loss. On top of that, we further introduce a new visual grounding task, locating objects without object names, which can thoroughly evaluate the model's dense alignment capacity. Through experiments, we achieve state-of-the-art performance on two widely-adopted 3D visual grounding datasets, ScanRefer and SR3D/NR3D, and obtain absolute leadership on our newly-proposed task. The source code is available at https://github.com/yanmin-wu/EDA. Yanmin Wu, Xinhua Cheng, Renrui Zhang, Zesen Cheng, Jian Zhang 0018 |
CVPR | 2 |
| 2023 | Null-Space Diffusion Sampling for Zero-Shot Point Cloud CompletionabstractPoint cloud completion aims at estimating the complete data of objects from degraded observations. Despite existing completion methods achieving impressive performances, they rely heavily on degraded-complete data pairs for supervision. In this work, we propose a novel framework named Null-Space Diffusion Sampling (NSDS) to solve the point cloud completion task in a zero-shot manner. By leveraging a pre-trained point cloud diffusion model as the off-the-shelf generator, our sampling approach can generate desired completion outputs with the guidance of the observed degraded data without any extra training. Furthermore, we propose a tolerant loop mechanism to improve the quality of completion results for hard cases. Experimental results demonstrate our zero-shot framework achieves superior completion performance than unsupervised methods and comparable performance to supervised methods in various degraded situations. Xinhua Cheng, Nan Zhang 0015, Jiwen Yu, Yinhuai Wang, Ge Li 0002 |
IJCAI | 1 |
| 2023 | Learning Disentangled Representation Implicitly Via Transformer for Occluded Person Re-IdentificationabstractPerson re-IDentification (re-ID) under various occlusions has been a long-standing challenge as person images with different types of occlusions often suffer from misalignment in image matching and ranking. Most existing methods tackle this challenge by aligning spatial features of body parts according to external semantic cues or feature similarities but this alignment approach is complicated and sensitive to noises. We design DRL-Net, a disentangled representation learning network that handles occluded re-ID without requiring strict person image alignment or any additional supervision. Leveraging transformer architectures, DRL-Net achieves alignment-free re-ID via global reasoning of local features of occluded person images. It measures image similarity by automatically disentangling the representation of undefined semantic components, e.g., human body parts or obstacles, under the guidance of semantic preference object queries in the transformer. In addition, we design a decorrelation constraint in the transformer decoder and impose it over object queries for better focus on different semantic components. To better eliminate interference from occlusions, we design a contrast feature learning technique (CFL) for better separation of occlusion features and discriminative ID features. Extensive experiments over occluded and holistic re-ID benchmarks show that the DRL-Net achieves superior re-ID performance consistently and outperforms the state-offi-the-art by large margins for occluded re-ID dataset. Mengxi Jia, Xinhua Cheng, Shijian Lu, Jian Zhang 0018 |
IEEE Trans. Multim. | 2 |
| 2022 | More is better: Multi-source Dynamic Parsing Attention for Occluded Person Re-identificationabstractOccluded person re-identification (re-ID) has been a long-standing challenge in surveillance systems. Most existing methods tackle this challenge by aligning spatial features of human parts according to external semantic cues, which are inferred from the off-the-shelf semantic models (e.g. human parsing and pose estimation). However, there is a significant domain gap between the images in re-ID datasets and the images used for training the semantic models, such that inevitably making those semantic cues unreliable and deteriorating the re-ID performance. Multi-source knowledge ensemble has been proved to be effective for domain adaptation. Inspired by this, we propose a multi-source dynamic parsing attention (MSDPA) mechanism that leverages knowledge learned from different source datasets to generate reliable semantic cues and dynamically integrate and adapt them in a self-supervised manner by attention mechanism. Specifically, we first design a parsing embedding module (PEM) to integrate and embed the multi-source semantic cues into the patch tokens through a voting procedure. To further exploit correlations among body parts with similar semantics, we design a dynamic parsing attention block (DPAB) to guide the patch sequences aggregation by prior attentions which are dynamically generated from human parsing results. Extensive experiments over occluded, partial, and holistic re-ID datasets show that the MSDPA achieves superior re-ID performance consistently and outperforms the state-of-the-art methods by large margins on occluded datasets. Xinhua Cheng, Mengxi Jia, Jian Zhang 0018 |
ACM Multimedia | 1 |
| 2022 | A Simple Visual-Textual Baseline for Pedestrian Attribute RecognitionabstractPedestrian attribute recognition (PAR), which aims to identify attributes of the pedestrians captured in video surveillance, is a challenging task due to the poor quality of images and diverse spatial distribution among attributes. Existing methods usually model PAR as a multi-label classification problem and manually map attributes to an ordered list corresponding to the outputs of classifiers or sequential models. However, the inherent textual information among attribute annotations is largely neglected in these visual-only methods. In this paper, we first alleviate this issue by proposing a novel visual-textual baseline (VTB) for PAR which introduces an additional textual modality to explore the textual semantic correlations from attribute annotations by pre-trained textual encoders instead of human definitions. VTB encodes pedestrian images and attribute annotations into visual and textual features respectively, interacts with information across modalities, and predicts recognition results independently to remove the influence of attribute orders. Furthermore, we introduce transformer encoder as the cross-modal fusion module in VTB for sufficient intra-modal and cross-modal correlations exploration. Our method achieves superior performance over most existing visual-only methods on two widely used datasets including RAP and PA-100K, demonstrating the effectiveness of utilizing textual modality to PAR. Our method is expected to serve as a multimodal PAR baseline and inspire new insights for multimodal fusion in future PAR research. Our code is available athttps://github.com/cxh0519/VTB. Xinhua Cheng, Mengxi Jia, Jian Zhang 0018 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Matching on Sets: Conquer Occluded Person Re-identification Without AlignmentabstractOccluded person re-identification (re-ID) is a challenging task as different human parts may become invisible in cluttered scenes, making it hard to match person images of different identities. Most existing methods address this challenge by aligning spatial features of body parts according to semantic information (e.g. human poses) or feature similarities but this approach is complicated and sensitive to noises. This paper presents Matching on Sets (MoS), a novel method that positions occluded person re-ID as a set matching task without requiring spatial alignment. MoS encodes a person image by a pattern set as represented by a `global vector’ with each element capturing one specific visual pattern, and it introduces Jaccard distance as a metric to compute the distance between pattern sets and measure image similarity. To enable Jaccard distance over continuous real numbers, we employ minimization and maximization to approximate the operations of intersection and union, respectively. In addition, we design a Jaccard triplet loss that enhances the pattern discrimination and allows to embed set matching into deep neural networks for end-to-end training. In the inference stage, we introduce a conflict penalty mechanism that detects mutually exclusive patterns in the pattern union of image pairs and decreases their similarities accordingly. Extensive experiments over three widely used datasets (Market1501, DukeMTMC and Occluded-DukeMTMC) show that MoS achieves superior re-ID performance. Additionally, it is tolerant of occlusions and outperforms the state-of-the-art by large margins for Occluded-DukeMTMC. Mengxi Jia, Xinhua Cheng, Yunpeng Zhai, Shijian Lu, Siwei Ma 0001, Yonghong Tian 0001, Jian Zhang 0018 |
AAAI | 2 |
| 2020 | Detection Features as Attention (Defat): A Keypoint-Free Approach to Amur Tiger Re-IdentificationabstractAutomatically identifying animals in camera-trap images has attracted increasing attention due to its valuable potential in wildlife conservation. A typical pipeline of existing methods includes separated animal detection and re-identification modules, and state-of-the-art re-identification methods either use annotated keypoints of animals to extract robust features, or employ extra branches to learn multiple features. In this paper, in contrast, we propose a keypoint-free approach to Amur tiger re-identification by exploiting the feature maps extracted by the detection module to help the re-identification module learn more effective features. We devise a detection-features-as-attention (DeFAt) module, which generates an additive mask for the input image based on the detection feature maps. We experimentally show that using the masked image the re-identification module lays more attention on the tiger region in the image, while the distraction by the messy background is removed to some extent. Our evaluation results prove that the proposed DeFAt module can effectively improve the Amur tiger re-identification accuracy when key-point annotations are not available. Xinhua Cheng, Jianing Zhu, Qijun Zhao |
ICIP | 1 |