EDBT 2026 Demo / reviewers in the wild / expert
Haozhe Xie
dblp:185/0576
· DBLP profile ↗
24ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0001-9596-5179ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CrowdMoGen: Event-Driven Collective Human Motion Generation
Haozhe Xie, Chenyang Gu, Ziwei Liu 0002 |
Int. J. Comput. Vis. | 4 |
| 2026 | Compositional Generative Model of Unbounded 4D Citiesabstract3D scene generation has garnered growing attention in recent years and has made significant progress. Generating 4D cities is more challenging than 3D scenes due to the presence of structurally complex, visually diverse objects like buildings and vehicles, and heightened human sensitivity to distortions in urban environments. To tackle these issues, we propose CityDreamer4D, a compositional generative model specifically tailored for generating unbounded 4D cities. Our main insights are 1) 4D city generation should separate dynamic objects (e.g., vehicles) from static scenes (e.g., buildings and roads), and 2) all objects in the 4D scene should be composed of different types of neural fields for buildings, vehicles, and background stuff. Specifically, we propose Traffic Scenario Generator and Unbounded Layout Generator to produce dynamic traffic scenarios and static city layouts using a highly compact BEV representation. Objects in 4D cities are generated by combining stuff-oriented and instance-oriented neural fields for background stuff, buildings, and vehicles. To suit the distinct characteristics of background stuff and instances, the neural fields employ customized generative hash grids and periodic positional embeddings as scene parameterizations. Furthermore, we offer a comprehensive suite of datasets for city generation, including OSM, GoogleEarth, and CityTopia. The OSM dataset provides a variety of real-world city layouts, while the Google Earth and CityTopia datasets deliver large-scale, high-quality city imagery complete with 3D instance annotations. Leveraging its compositional design, CityDreamer4D supports a range of downstream applications, such as instance editing, city stylization, and urban simulation, while delivering state-of-the-art performance in generating realistic 4D cities. Haozhe Xie, Zhaoxi Chen 0009, Fangzhou Hong, Ziwei Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Multi-view Consistent 3D Panoptic Scene Understandingabstract3D panoptic scene understanding seeks to create novel view images with 3D-consistent panoptic segmentation, which is crucial for many vision and robotics applications. Mainstream methods (e.g., Panoptic Lifting) directly use machine-generated 2D panoptic segmentation masks as training labels. However, these generated masks often exhibit multi-view inconsistencies, leading to ambiguities during the optimization process. To address this, we present Multi-view Consistent 3D Panoptic Scene Understanding (MVC-PSU), featuring two key components: 1) Probabilistic Semantic Aligner, which associates semantic information of corresponding pixels across multiple views by probabilistic alignment to ensure that predicted panoptic segmentation masks are consistent across different views. 2) Geometric Consistency Enforcer, which uses multi-view projection and monocular depth consistency to ensure that the geometry of the reconstructed scene is accurate and consistent across different views. Experimental results demonstrate that the proposed MVC-PSU surpasses state-of-the-art methods on the ScanNet, Replica, and HyperSim datasets. Xianzhu Liu, Xin Sun 0003, Haozhe Xie, Zonglin Li 0004, Ru Li 0002, Shengping Zhang |
AAAI | 3 |
| 2025 | 3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive DiffusionabstractThe increasing demand for high-quality 3D assets across various industries necessitates efficient and automated 3D content creation. Despite recent advancements in 3D generative models, existing methods still face challenges with optimization speed, geometric fidelity, and the lack of assets for physically based rendering (PBR). In this paper, we introduce 3DTopia-XL, a scalable native 3D generative model designed to overcome these limitations. 3DTopia-XL leverages a novel primitive-based 3D representation, PrimX, which encodes detailed shape, albedo, and material field into a compact tensorial format, facilitating the modeling of high-resolution geometry with PBR assets. On top of the novel representation, we propose a generative framework based on Diffusion Transformer (DiT), which comprises 1) Primitive Patch Compression, 2) and Latent Primitive Diffusion. 3DTopia-XL learns to generate high-quality 3D assets from textual or visual inputs. Extensive qualitative and quantitative evaluations are conducted to demonstrate that 3DTopia-XL significantly outperforms existing methods in generating high-quality 3D assets with fine-grained textures and materials, efficiently bridging the quality gap between generative models and real-world applications. Zhaoxi Chen 0009, Jiaxiang Tang, Yuhao Dong, Ziang Cao, Fangzhou Hong, Yushi Lan, Tengfei Wang 0002, Haozhe Xie, Shunsuke Saito, Liang Pan, Dahua Lin, Ziwei Liu 0002 |
CVPR | 8 |
| 2025 | Generative Gaussian Splatting for Unbounded 3D City Generationabstract3D city generation with NeRF-based methods shows promising generation results but is computationally inefficient. Recently 3D Gaussian splatting (3D-GS) has emerged as a highly efficient alternative for object-level 3D generation. However, adapting 3D-GS from finite-scale 3D objects and humans to infinite-scale 3D cities is non-trivial. Unbounded 3D city generation entails significant storage overhead (out-of-memory issues), arising from the need to expand points to billions, often demanding hundreds of Gigabytes of VRAM for a city scene spanning 10km2. In this paper, we propose GaussianCity, a generative Gaussian splatting framework dedicated to efficiently synthesizing unbounded 3D cities with a single feed-forward pass. Our key insights are two-fold: 1) Compact 3D Scene Representation: We introduce BEV-Point as a highly compact intermediate representation, ensuring that the growth in VRAM usage for unbounded scenes remains constant, thus enabling unbounded city generation. 2) Spatial-aware Gaussian Attribute Decoder: We present spatial-aware BEV-Point decoder to produce 3D Gaussian attributes, which leverages Point Serializer to integrate the structural and contextual characteristics of BEV points. Extensive experiments demonstrate that GaussianCity achieves state-of-the-art results in both drone-view and street-view 3D city generation. Notably, compared to CityDreamer, GaussianCity exhibits superior performance with a speedup of 60 times (10.72 FPS v.s. 0.18 FPS). Haozhe Xie, Zhaoxi Chen 0009, Fangzhou Hong, Ziwei Liu 0002 |
CVPR | 1 |
| 2025 | DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic ScenesabstractUrban scene generation has been developing rapidly recently. However, existing methods primarily focus on generating static and single-frame scenes, overlooking the inherently dynamic nature of real-world driving environments. In this work, we introduce DynamicCity, a novel 4D occupancy generation framework capable of generating large-scale, high-quality dynamic 4D scenes with semantics. DynamicCity mainly consists of two key models. **1)** A VAE model for learning HexPlane as the compact 4D representation. Instead of using naive averaging operations, DynamicCity employs a novel **Projection Module** to effectively compress 4D features into six 2D feature maps for HexPlane construction, which significantly enhances HexPlane fitting quality (up to **12.56** mIoU gain). Furthermore, we utilize an **Expansion & Squeeze Strategy** to reconstruct 3D feature volumes in parallel, which improves both network training efficiency and reconstruction accuracy than naively querying each 3D point (up to **7.05** mIoU gain, **2.06x** training speedup, and **70.84\%** memory reduction). **2)** A DiT-based diffusion model for HexPlane generation. To make HexPlane feasible for DiT generation, a **Padded Rollout Operation** is proposed to reorganize all six feature planes of the HexPlane as a squared 2D feature map. In particular, various conditions could be introduced in the diffusion or sampling process, supporting **versatile 4D generation applications**, such as trajectory- and command-driven generation, inpainting, and layout-conditioned generation. Extensive experiments on the CarlaSC and Waymo datasets demonstrate that DynamicCity significantly outperforms existing state-of-the-art 4D occupancy generation methods across multiple metrics. The code and models have been released to facilitate future research. Hengwei Bian, Lingdong Kong, Haozhe Xie, Liang Pan, Yu Qiao 0001, Ziwei Liu 0002 |
ICLR | 3 |
| 2025 | 2D Semantic-Guided Semantic Scene Completion
Xianzhu Liu, Haozhe Xie, Shengping Zhang, Hongxun Yao, Rongrong Ji, Liqiang Nie, Dacheng Tao |
Int. J. Comput. Vis. | 2 |
| 2025 | Patch-Based Spatio-Temporal Deformable Attention BiRNN for Video DeblurringabstractSuccessful video deblurring relies on effectively using sharp pixels from other frames to recover the blurry pixels of the current frame. However, mainstream methods only use estimated optical flows to align and fuse features from adjacent frames without considering the pixel-wise blur levels, leading to the introduction of blurry pixels from adjacent frames. Furthermore, these methods fail to effectively exploit information from the entire input video. To address these limitations, we propose STDANet++, which redesigns the state-of-the-art method STDANet by introducing patch-based spatio-temporal deformable attention (PSTDA) module and long-term frame fusion (LTFF) module to the BiRNN-based structure. By effectively utilizing sharp information across the entire video, the proposed method outperforms state-of-the-art methods on the GoPro, DVD and BSD datasets, according to our experimental results. The source code is available at https://github.com/huicongzhang/STDANetPP. Huicong Zhang, Haozhe Xie, Shengping Zhang, Hongxun Yao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | CityDreamer: Compositional Generative Model of Unbounded 3D Citiesabstract3D city generation is a desirable yet challenging task, since humans are more sensitive to structural distortions in urban environments. Additionally, generating 3D cities is more complex than 3D natural scenes since buildings, as objects of the same class, exhibit a wider range of appear-ances compared to the relatively consistent appearance of objects like trees in natural scenes. To address these challenges, we propose CityDreamer, a compositional generative model designed specifically for unbounded 3D cities. Our key insight is that 3D city generation should be a com-position of different types of neural fields: 1) various building instances, and 2) background stuff, such as roads and green lands. Specifically, we adopt the bird's eye view scene representation and employ a volumetric render for both instance-oriented and stuff-oriented neural fields. The generative hash grid and periodic positional embedding are tailored as scene parameterization to suit the distinct characteristics of building instances and background stuff. Furthermore, we contribute a suite of CityGen Datasets, including OSM and GoogleEarth, which comprises a vast amount of real-world city imagery to enhance the realism of the generated 3D cities both in their layouts and appear-ances. CityDreamer achieves state-of-the-art performance not only in generating realistic 3D cities but also in local-ized editing within the generated cities. Haozhe Xie, Zhaoxi Chen 0009, Fangzhou Hong, Ziwei Liu 0002 |
CVPR | 1 |
| 2024 | Blur-Aware Spatio-Temporal Sparse Transformer for Video DeblurringabstractVideo deblurring relies on leveraging information from other frames in the video sequence to restore the blurred regions in the current frame. Mainstream approaches employ bidirectional feature propagation, spatio-temporal Transformers, or a combination of both to extract information from the video sequence. However, limitations in memory and computational resources constraints the temporal window length of the spatio-temporal transformer, preventing the extraction of longer temporal contextual information from the video sequence. Additionally, bidirectional feature propagation is highly sensitive to inaccurate optical flow in blurry frames, leading to error accumulation during the propagation process. To address these issues, we propose BSSTNet, Blur-aware Spatio-temporal Sparse Transformer Network. It introduces the blur map, which converts the originally dense attention into a sparse form, enabling a more extensive utilization of information throughout the entire video sequence. Specifically, BSSTNet (1) uses a longer temporal window in the transformer, lever-aging information from more distant frames to restore the blurry pixels in the current frame. (2) introduces bidirectional feature propagation guided by blur maps, which reduces error accumulation caused by the blur frame. The experimental results demonstrate the proposed BSSTNet out-performs the state-of-the-art methods on the GoPro and DVD datasets. Huicong Zhang, Haozhe Xie, Hongxun Yao |
CVPR | 2 |
| 2023 | Learning Geometric Transformation for Point Cloud Completion
Shengping Zhang, Xianzhu Liu, Haozhe Xie, Liqiang Nie, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001 |
Int. J. Comput. Vis. | 3 |
| 2022 | Spatio-Temporal Deformable Attention Network for Video Deblurring
Huicong Zhang, Haozhe Xie, Hongxun Yao |
ECCV (16) | 2 |
| 2022 | BacklitNet: A dataset and network for backlit image enhancement
Xiaoqian Lv, Shengping Zhang, Qinglin Liu, Haozhe Xie, Bineng Zhong 0001, Huiyu Zhou 0001 |
Comput. Vis. Image Underst. | 4 |
| 2021 | Efficient Regional Memory Network for Video Object SegmentationabstractRecently, several Space-Time Memory based networks have shown that the object cues (e.g. video frames as well as the segmented object masks) from the past frames are useful for segmenting objects in the current frame. However, these methods exploit the information from the memory by global-to-global matching between the current and past frames, which lead to mismatching to similar objects and high computational complexity. To address these problems, we propose a novel local-to-local matching solution for semi-supervised VOS, namely Regional Memory Network (RMNet). In RMNet, the precise regional memory is constructed by memorizing local regions where the target objects appear in the past frames. For the current query frame, the query regions are tracked and predicted based on the optical flow estimated from the previous frame. The proposed local-to-local matching effectively alleviates the ambiguity of similar objects in both memory and query frames, which allows the information to be passed from the regional memory to the query region efficiently and effectively. Experimental results indicate that the proposed RM-Net performs favorably against state-of-the-art methods on the DAVIS and YouTube-VOS datasets. Haozhe Xie, Hongxun Yao, Shangchen Zhou, Shengping Zhang, Wenxiu Sun |
CVPR | 1 |
| 2021 | Single-View 3D Object Reconstruction From Shape Priors in MemoryabstractExisting methods for single-view 3D object reconstruction directly learn to transform image features into 3D representations. However, these methods are vulnerable to images containing noisy backgrounds and heavy occlusions because the extracted image features do not contain enough information to reconstruct high-quality 3D shapes. Humans routinely use incomplete or noisy visual cues from an image to retrieve similar 3D shapes from their memory and reconstruct the 3D shape of an object. Inspired by this, we propose a novel method, named Mem3D, that explicitly constructs shape priors to supplement the missing information in the image. Specifically, the shape priors are in the forms of "image-voxel" pairs in the memory network, which is stored by a well-designed writing strategy during training. We also propose a voxel triplet loss function that helps to retrieve the precise 3D shapes that are highly related to the input image from shape priors. The LSTM-based shape encoder is introduced to extract information from the retrieved 3D shapes, which are useful in recovering the 3D shape of an object that is heavily occluded or in complex environments. Experimental results demonstrate that Mem3D significantly improves reconstruction quality and performs favorably against state-of-the-art methods on the ShapeNet and Pix3D datasets. Shuo Yang 0006, Min Xu 0001, Haozhe Xie, Stuart W. Perry, Jiahao Xia 0001 |
CVPR | 3 |
| 2021 | Long-Range Feature Propagating for Natural Image MattingabstractNatural image matting estimates the alpha values of unknown regions in the trimap. Recently, deep learning based methods propagate the alpha values from the known regions to unknown regions according to the similarity between them. However, we find that more than 50% pixels in the unknown regions cannot be correlated to pixels in known regions due to the limitation of small effective reception fields of common convolutional neural networks, which leads to inaccurate estimation when the pixels in the unknown regions cannot be inferred only with pixels in the reception fields. To solve this problem, we propose Long-Range Feature Propagating Network (LFPNet), which learns the long-range context features outside the reception fields for alpha matte estimation. Specifically, we first design the propagating module which extracts the context features from the downsampled image. Then, we present Center-Surround Pyramid Pooling (CSPP) that explicitly propagates the context features from the surrounding context image patch to the inner center image patch. Finally, we use the matting module which takes the image, trimap and context features to estimate the alpha matte. Experimental results demonstrate that the proposed method performs favorably against the state-of-the-art methods on the AlphaMatting and Adobe Image Matting datasets. Qinglin Liu, Haozhe Xie, Shengping Zhang, Bineng Zhong 0001, Rongrong Ji |
ACM Multimedia | 2 |
| 2021 | Toward 3D object reconstruction from stereo images
Haozhe Xie, Hongxun Yao, Shangchen Zhou, Shengping Zhang, Xiaojun Tong, Wenxiu Sun |
Neurocomputing | 1 |
| 2020 | GRNet: Gridding Residual Network for Dense Point Cloud Completion
Haozhe Xie, Hongxun Yao, Shangchen Zhou, Jiageng Mao, Shengping Zhang, Wenxiu Sun |
ECCV (9) | 1 |
| 2020 | PRF-Ped: Multi-scale Pedestrian Detector with Prior-based Receptive FieldabstractMulti-scale feature representation is a common strategy to handle the scale variation in pedestrian detection. Existing methods simply utilize the convolutional pyramidal features for multi-scale representation. However, they rarely pay attention to the differences among different feature scales and extract multi-scale features from a single feature map, which may make the detectors sensitive to scale-variance in multi-scale pedestrian detection. In this paper, we introduce a bidirectional feature enhancement module (BFEM) to augment the semantic information of low-level features and the localization information of high-level features. In addition, we propose a prior-based receptive field block (PRFB) for multi-scale pedestrian feature extraction, where the receptive field is closer to the aspect ratio of the pedestrian target. Consequently, it is less affected by the surrounding background when extracting features. Experimental results indicate that the proposed method outperforms the state-of-the-art methods on the CityPersons and Caltech datasets. Yuzhi Tan, Hongxun Yao, Haoran Li 0012, Xiusheng Lu, Haozhe Xie |
ICPR | 5 |
| 2020 | Pix2Vox++: Multi-scale Context-aware 3D Object Reconstruction from Single and Multiple Images
Haozhe Xie, Hongxun Yao, Shengping Zhang, Shangchen Zhou, Wenxiu Sun |
Int. J. Comput. Vis. | 1 |
| 2019 | DAVANet: Stereo Deblurring With View AggregationabstractNowadays stereo cameras are more commonly adopted in emerging devices such as dual-lens smartphones and unmanned aerial vehicles. However, they also suffer from blurry images in dynamic scenes which leads to visual discomfort and hampers further image processing. Previous works have succeeded in monocular deblurring, yet there are few studies on deblurring for stereoscopic images. By exploiting the two-view nature of stereo images, we propose a novel stereo image deblurring network with Depth Awareness and View Aggregation, named DAVANet. In our proposed network, 3D scene cues from the depth and varying information from two views are incorporated, which help to remove complex spatially-varying blur in dynamic scenes. Specifically, with our proposed fusion network, we integrate the bidirectional disparities estimation and deblurring into a unified framework. Moreover, we present a large-scale multi-scene dataset for stereo deblurring, containing 20,637 blurry-sharp stereo image pairs from 135 diverse sequences and their corresponding bidirectional disparities. The experimental results on our dataset demonstrate that DAVANet outperforms state-of-the-art methods in terms of accuracy, speed, and model size. Shangchen Zhou, Jiawei Zhang 0002, Wangmeng Zuo, Haozhe Xie, Jinshan Pan, Jimmy S. J. Ren |
CVPR | 4 |
| 2019 | Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesabstractRecovering the 3D representation of an object from single-view or multi-view RGB images by deep neural networks has attracted increasing attention in the past few years. Several mainstream works (e.g., 3D-R2N2) use recurrent neural networks (RNNs) to fuse multiple feature maps extracted from input images sequentially. However, when given the same set of input images with different orders, RNN-based approaches are unable to produce consistent reconstruction results. Moreover, due to long-term memory loss, RNNs cannot fully exploit input images to refine reconstruction results. To solve these problems, we propose a novel framework for single-view and multi-view 3D reconstruction, named Pix2Vox. By using a well-designed encoder-decoder, it generates a coarse 3D volume from each input image. Then, a context-aware fusion module is introduced to adaptively select high-quality reconstructions for each part (e.g., table legs) from different coarse 3D volumes to obtain a fused 3D volume. Finally, a refiner further refines the fused 3D volume to generate the final output. Experimental results on the ShapeNet and Pix3D benchmarks indicate that the proposed Pix2Vox outperforms state-of-the-arts by a large margin. Furthermore, the proposed method is 24 times faster than 3D-R2N2 in terms of backward inference time. The experiments on ShapeNet unseen 3D categories have shown the superior generalization abilities of our method. Haozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou, Shengping Zhang |
ICCV | 1 |
| 2019 | Spatio-Temporal Filter Adaptive Network for Video DeblurringabstractVideo deblurring is a challenging task due to the spatially variant blur caused by camera shake, object motions, and depth variations, etc. Existing methods usually estimate optical flow in the blurry video to align consecutive frames or approximate blur kernels. However, they tend to generate artifacts or cannot effectively remove blur when the estimated optical flow is not accurate. To overcome the limitation of separate optical flow estimation, we propose a Spatio-Temporal Filter Adaptive Network (STFAN) for the alignment and deblurring in a unified framework. The proposed STFAN takes both blurry and restored images of the previous frame as well as blurry image of the current frame as input, and dynamically generates the spatially adaptive filters for the alignment and deblurring. We then propose the new Filter Adaptive Convolutional (FAC) layer to align the deblurred features of the previous frame with the current frame and remove the spatially variant blur from the features of the current frame. Finally, we develop a reconstruction network which takes the fusion of two transformed features to restore the clear frames. Both quantitative and qualitative evaluation results on the benchmark datasets and real-world videos demonstrate that the proposed algorithm performs favorably against state-of-the-art methods in terms of accuracy, speed as well as model size. Shangchen Zhou, Jiawei Zhang 0002, Jinshan Pan, Wangmeng Zuo, Haozhe Xie, Jimmy S. J. Ren |
ICCV | 5 |
| 2016 | A network-based pathway-expanding approach for pathway analysisabstractBACKGROUND: Pathway analysis combining multiple types of high-throughput data, such as genomics and proteomics, has become the first choice to gain insights into the pathogenesis of complex diseases. Currently, several pathway analysis methods have been developed to study complex diseases. However, these methods did not take into account the interaction between internal and external genes of the pathway and between pathways. Hence, these approaches still face some challenges. Here, we propose a network-based pathway-expanding approach that takes the topological structures of biological networks into account. RESULTS: First, two weighted gene-gene interaction networks (tumor and normal) are constructed integrating protein-protein interaction(PPI) information, gene expression data and pathway databases. Then, they are used to identify significant pathways through testing the difference of topological structures of expanded pathways in the two weighted networks. The proposed method is employed to analyze two breast cancer data. As a result, the top 15 pathways identified using the proposed method are supported by biological knowledge from the published literatures and other methods. In addition, the proposed method is also compared with other methods, such as GSEA and SPIA, and estimated using the classification performance of the top 15 expanded pathways. CONCLUSIONS: A novel network-based pathway-expanding approach is proposed to avoid the limitations of existing pathway analysis approaches. Experimental results indicate that the proposed method can accurately and reliably identify significant pathways which are related to the corresponding disease. Jie Li 0055, Haozhe Xie, Hanqing Xue, Yadong Wang 0001 |
BMC Bioinform. | 3 |