Zifan Shi

dblp:136/5965 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Discovering interpretable blast Loading equations from Black-Box Machine learning models
abstract
Boiling Liquid Expanding Vapour Explosion (BLEVE) is a high-energy event characterised by intense blast waves that pose serious safety risks. Accurate prediction of the resulting overpressure wave is essential for knowledge-intensive engineering analysis and decision support. While empirical methods are available for predicting overpressure in simple BLEVE scenarios, they fail to capture nonlinear relationships in multi-feature and complex conditions. Computational Fluid Dynamics (CFD) methods offer high accuracy in overpressure wave prediction but are computationally intensive, expensive to use and difficult to integrate into automated or real-time engineering workflows. Machine learning models offer a promising alternative for rapid predictions, but their limited interpretability, particularly in deep learning architectures, poses a significant barrier to integration into real-world engineering systems. This study proposes a systematic approach combining machine learning, explainable artificial intelligence, and symbolic regression for BLEVE overpressure prediction. A feedforward neural network model is developed and interpreted using SHapley Additive exPlanations (SHAP). Global SHAP analysis identified nine features with the most significant contributions, which were subsequently used to train a global surrogate model via symbolic regression. This approach yielded an explicit mathematical expression that approximates the behaviour of the original neural network. The derived equation achieved a relative error of 15.73% on simulated data and 35.45% on experimental data, outperforming existing empirical formulas. This research demonstrates the potential of combining black-box machine learning models with xAI techniques to develop interpretable and reliable equations for blast load prediction. More importantly, it introduces a novel data-driven methodology of data-model-interpretation-equation that formalises engineering knowledge by transforming black-box models into explicit and interpretable computational representations.
Zifan Shi, Yanda Shao, Hong Hao
Adv. Eng. Informatics1
2025 Learning Naturally Aggregated Appearance for Efficient 3D Editing
abstract
Neural radiance fields, which represent a 3D scene as a color field and a density field, have demonstrated great progress in novel view synthesis yet are unfavorable for editing due to the implicitness. This work studies the task of efficient 3D editing, where we focus on editing speed and user interactivity. To this end, we propose to learn the color field as an explicit 2D appearance aggregation, also called canonical image, with which users can easily customize their 3D editing via 2D image processing. We complement the canonical image with a projection field that maps 3D points onto 2D pixels for texture query. This field is initialized with a pseudo canonical camera model and optimized with offset regularity to ensure the naturalness of the canonical image. Extensive experiments on different datasets suggest that our representation, dubbed AGAP, well supports various ways of 3D editing (e.g., stylization, instance segmentation, and interactive drawing). Our approach demonstrates remarkable efficiency by being at least 20× faster per edit compared to existing NeRF-based editing methods. Project page is available at h ttps: //felixcheng97.github.io/AGAP/.
Ka Leong Cheng, Qiuyu Wang, Zifan Shi, Kecheng Zheng, Yinghao Xu 0001, Hao Ouyang, Qifeng Chen 0001, Yujun Shen
3DV3
2025 Exploring Sparse MoE in GANs for Text-conditioned Image Synthesis
abstract
Due to the difficulty in scaling up, generative adversarial networks (GANs) seem to be falling out of grace with the task of text-conditioned image synthesis. Sparsely activated mixture-of-experts (MoE) has recently been demonstrated as a valid solution to training large-scale models with limited resources. Inspired by this, we present Aurora, a GAN-based text-to-image generator that employs a collection of experts to learn feature processing, together with a sparse router to adaptively select the most suitable expert for each feature point. We adopt a two-stage training strategy, which first learns a base model at 64 × 64 resolution followed by an upsampler to produce 512 × 512 images. Trained with only public data, our approach encouragingly closes the performance gap between GANs and industry-level diffusion models, maintaining a fast inference speed. We release the code and checkpoints here to facilitate the community for further development.
Jiapeng Zhu 0001, Ceyuan Yang, Kecheng Zheng, Yinghao Xu 0001, Zifan Shi, Qifeng Chen 0001, Yujun Shen
CVPR5
2025 Data-Driven BLEVE Overpressure Prediction Using Explainable Machine Learning
Zifan Shi, Yanda Shao, Hong Hao
EANN (2)1
2025 MASTTN: A Multi-Feature Adaptive Spatio-Temporal Trend Network for Intelligent Cellular Traffic Prediction
abstract
With the popularization of 5G technology and the growth of data traffic demands, cellular traffic prediction has become crucial for network management. However, existing methods face the following challenges: (1) Most studies only consider normal traffic and neglect the dynamic impact of external environmental changes; (2) Due to model complexity, existing methods rely on short-term data and struggle to capture long-term traffic trends. To address these issues, this paper proposes a Multi-Feature Adaptive Spatiotemporal Trend Network (MASTTN), which includes feature embedding, mask reconstruction, and spatiotemporal trend extraction modules. First, a gated unit dynamically integrates temporal, spatial, and external features to capture the multidimensional correlations of traffic. Next, a random mask reconstruction approach extracts new sequences containing trend information from longer historical time series. Finally, dilated causal convolution and short-term spatiotemporal feature extraction modules are used to refine both long- and short-term features. Experimental results show that MASTTN achieves a minimum improvement of 5.63% and a maximum improvement of 16.78% in long-term prediction over 120 minutes, compared to baseline models, and the effectiveness of each module is verified through perturbation experiments.
Mengrui Guo, Xuanli Wu, Zifan Shi
IWCMC4
2025 RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets
abstract
We present RigAnything , a novel autoregressive transformer-based model, which makes 3D assets rig-ready by probabilistically generating joints and skeleton topologies and assigning skinning weights in a template-free manner. Unlike most existing auto-rigging methods, which rely on predefined skeleton templates and are limited to specific categories like humanoid, RigAnything approaches the rigging problem in an autoregressive manner, iteratively predicting the next joint based on the global input shape and the previous prediction. While autoregressive models are typically used to generate sequential data, RigAnything extends its application to effectively learn and represent skeletons, which are inherently tree structures. To achieve this, we organize the joints in a breadth-first search (BFS) order, enabling the skeleton to be defined as a sequence of 3D locations and the parent index. Furthermore, our model improves the accuracy of position prediction by leveraging diffusion modeling, ensuring precise and consistent placement of joints within the hierarchy. This formulation allows the autoregressive model to efficiently capture both spatial and hierarchical relationships within the skeleton. Trained end-to-end on both RigNet and Objaverse datasets, RigAnything demonstrates state-of-the-art performance across diverse object types, including humanoids, quadrupeds, marine creatures, insects, and many more, surpassing prior methods in quality, robustness, generalizability, and efficiency. It achieves significantly faster performance than existing auto-rigging methods, completing rigging in under a few seconds per shape.
Isabella Liu, Wang Yifan 0001, Hao Tan 0002, Zexiang Xu, Xiaolong Wang 0004, Hao Su 0001, Zifan Shi
ACM Trans. Graph.8
2024 Gaussian Shell Maps for Efficient 3D Human Generation
abstract
Efficient generation of 3D digital humans is important in several industries, including virtual reality, social media, and cinematic production. 3D generative adversarial net-works (GANs) have demonstrated state-of-the-art (SOTA) quality and diversity for generated assets. Current 3D GAN architectures, however, typically rely on volume representations, which are slow to render, thereby hampering the GAN training and requiring multi- view-inconsistent 2D upsam-plers. Here, we introduce Gaussian Shell Maps (GSMs) as a framework that connects SOTA generator network archi-tectures with emerging 3D Gaussian rendering primitives using an articulable multi shell-based scaffold. In this set-ting, a CNN generates a 3D texture stack with features that are mapped to the shells. The latter represent inflated and deflated versions of a template surface of a digital human in a canonical body pose. Instead of rasterizing the shells directly, we sample 3D Gaussians on the shells whose at-tributes are encoded in the texture features. These Gaus-sians are efficiently and differentiably rendered. The ability to articulate the shells is important during GAN training and, at inference time, to deform a body into arbitrary user-defined poses. Our efficient rendering scheme bypasses the need for view-inconsistent upsamplers and achieves high-quality multi-view consistent renderings at a native resolution of 512 ×512 pixels. We demonstrate that GSMs suc-cessfully generate 3D humans when trained on single-view datasets, including SHHQ and DeepFashion. Project Page: rameenabdal.github.io/GaussianShellMaps
Rameen Abdal, Wang Yifan 0001, Zifan Shi, Yinghao Xu 0001, Ryan Po, Zhengfei Kuang, Qifeng Chen 0001, Dit-Yan Yeung, Gordon Wetzstein
CVPR3
2024 Real-Time 3D-Aware Portrait Editing from a Single Image
Qingyan Bai, Zifan Shi, Yinghao Xu 0001, Hao Ouyang, Qiuyu Wang, Ceyuan Yang, Xuan Wang 0009, Gordon Wetzstein, Yujun Shen, Qifeng Chen 0001
ECCV (51)2
2024 GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation
Yinghao Xu 0001, Zifan Shi, Wang Yifan 0001, Hansheng Chen 0001, Ceyuan Yang, Sida Peng, Yujun Shen, Gordon Wetzstein
ECCV (15)2
2024 DMV3D: Denoising Multi-view Diffusion Using 3D Large Reconstruction Model
abstract
We propose DMV3D, a novel 3D generation approach that uses a transformer-based 3D large reconstruction model to denoise multi-view diffusion. Our reconstruction model incorporates a triplane NeRF representation and, functioning as a denoiser, can denoise noisy multi-view images via 3D NeRF reconstruction and rendering, achieving single-stage 3D generation in the 2D diffusion denoising process. We train DMV3D on large-scale multi-view image datasets of extremely diverse objects using only image reconstruction losses, without accessing 3D assets. We demonstrate state-of-the-art results for the single-image reconstruction problem where probabilistic modeling of unseen object parts is required for generating diverse reconstructions with sharp textures. We also show high-quality text-to-3D generation results outperforming previous 3D diffusion models. Our project website is at: https://dmv3d.github.io/.
Yinghao Xu 0001, Hao Tan 0002, Fujun Luan, Sai Bi, Peng Wang 0099, Zifan Shi, Kalyan Sunkavalli, Gordon Wetzstein, Zexiang Xu, Kai Zhang 0045
ICLR7
2023 Learning 3D-Aware Image Synthesis with Unknown Pose Distribution
abstract
Existing methods for 3D-aware image synthesis largely depend on the 3D pose distribution pre-estimated on the training set. An inaccurate estimation may mislead the model into learning faulty geometry. This work proposes PoF3D that frees generative radiance fields from the requirements of 3D pose priors. We first equip the generator with an efficient pose learner, which is able to infer a pose from a latent code, to approximate the underlying true pose distribution automatically. We then assign the discriminator a task to learn pose distribution under the supervision of the generator and to differentiate real and synthesized images with the predicted pose as the condition. The pose-free generator and the pose-aware discriminator are jointly trained in an adversarial manner. Extensive results on a couple of datasets confirm that the performance of our approach, regarding both image quality and geometry quality, is on par with state of the art. To our best knowledge, PoF3D demonstrates the feasibility of learning high-quality 3D-aware image synthesis without using 3D pose priors for the first time. Project page can be found here.
Zifan Shi, Yujun Shen, Yinghao Xu 0001, Sida Peng, Yiyi Liao, Qifeng Chen 0001, Dit-Yan Yeung
CVPR1
2023 DisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene Synthesis
abstract
Existing 3D-aware image synthesis approaches mainly focus on generating a single canonical object and show limited capacity in composing a complex scene containing a variety of objects. This work presents DisCoScene: a 3D-aware generative model for high-quality and controllable scene synthesis. The key ingredient of our method is a very abstract object-level representation (i.e., 3D bounding boxes without semantic annotation) as the scene layout prior, which is simple to obtain, general to describe various scene contents, and yet informative to disentangle objects and background. Moreover, it serves as an intuitive user control for scene editing. Based on such a prior, the proposed model spatially disentangles the whole scene into object-centric generative radiance fields by learning on only 2D images with the global-local discrimination. Our model obtains the generation fidelity and editing flexibility of individual objects while being able to efficiently compose objects and the background into a complete scene. We demonstrate state-of-the-art performance on many scene datasets, including the challenging Waymo outdoor dataset. Project page can be found here.
Yinghao Xu 0001, Menglei Chai, Zifan Shi, Sida Peng, Ivan Skorokhodov, Aliaksandr Siarohin, Ceyuan Yang, Yujun Shen, Hsin-Ying Lee 0001, Bolei Zhou, Sergey Tulyakov
CVPR3
2023 LinkGAN: Linking GAN Latents to Pixels for Controllable Image Synthesis
abstract
This work presents an easy-to-use regularizer for GAN training, which helps explicitly link some axes of the latent space to a set of pixels in the synthesized image. Establishing such a connection facilitates a more convenient local control of GAN generation, where users can alter the image content only within a spatial area simply by partially resampling the latent code. Experimental results confirm four appealing properties of our regularizer, which we call LinkGAN. (1) The latent-pixel linkage is applicable to either a fixed region (i.e., same for all instances) or a particular semantic category (i.e., varying across instances), like the sky. (2) Two or multiple regions can be independently linked to different latent axes, which further supports joint control. (3) Our regularizer can improve the spatial controllability of both 2D and 3D-aware GAN models, barely sacrificing the synthesis performance. (4) The models trained with our regularizer are compatible with GAN inversion techniques and maintain editability on real images. Project page can be found here.
Jiapeng Zhu 0001, Ceyuan Yang, Yujun Shen, Zifan Shi, Bo Dai 0002, Deli Zhao, Qifeng Chen 0001
ICCV4
2023 Benchmarking and Analyzing 3D-aware Image Synthesis with a Modularized Codebase
abstract
Despite the rapid advance of 3D-aware image synthesis, existing studies usually adopt a mixture of techniques and tricks, leaving it unclear how each part contributes to the final performance in terms of generality. Following the most popular and effective paradigm in this field, which incorporates a neural radiance field (NeRF) into the generator of a generative adversarial network (GAN), we builda well-structured codebase through modularizing the generation process. Such a design allows researchers to develop and replace each module independently, and hence offers an opportunity to fairly compare various approaches and recognize their contributions from the module perspective. The reproduction of a range of cutting-edge algorithms demonstrates the availability of our modularized codebase. We also perform a variety of in-depth analyses, such as the comparison across different types of point feature, the necessity of the tailing upsampler in the generator, the reliance on the camera pose prior, etc., which deepen our understanding of existing methods and point out some further directions of the research work. Code and models will be made publicly available to facilitate the development and evaluation of this field.
Qiuyu Wang, Zifan Shi, Kecheng Zheng, Yinghao Xu 0001, Sida Peng, Yujun Shen
NeurIPS2
2023 Design of Coded Caching Schemes With Linear Subpacketizations Based on Injective Arc Coloring of Regular Digraphs
abstract
Coded caching is an effective technique to decongest the amount of traffic in the backhaul link. In such a scheme, each file hosted in the server is divided into a number of packets to pursue a low broadcasting rate based on the designed placements at each user’s cache. However, the implementation complexity of this scheme increases with the number of packets. It is important to design a scheme with a small subpacketization level and a relatively low transmission rate. Recently, placement delivery array (PDA) was proposed to address the subpacketization bottleneck of coded caching. This paper investigates the design of PDA from a new perspective, i.e., the injective arc coloring of regular digraphs. It is shown that the injective arc coloring of a regular digraph can yield a PDA with the same number of rows and columns. Based on this, a new class of regular digraphs are defined and the upper bounds on the injective chromatic index of such digraphs are derived. Consequently, four new coded caching schemes with a linear subpacketization level and a relatively small transmission rate are proposed, one of which generalizes the existing scheme for the scenario with a more flexible number of users.
Xianzhang Wu, Minquan Cheng, Li Chen 0013, Congduan Li, Zifan Shi
IEEE Trans. Commun.5
2022 3D-Aware Indoor Scene Synthesis with Depth Priors
Zifan Shi, Yujun Shen, Jiapeng Zhu 0001, Dit-Yan Yeung, Qifeng Chen 0001
ECCV (16)1
2022 Improving 3D-aware Image Synthesis with A Geometry-aware Discriminator
abstract
3D-aware image synthesis aims at learning a generative model that can render photo-realistic 2D images while capturing decent underlying 3D shapes. A popular solution is to adopt the generative adversarial network (GAN) and replace the generator with a 3D renderer, where volume rendering with neural radiance field (NeRF) is commonly used. Despite the advancement of synthesis quality, existing methods fail to obtain moderate 3D shapes. We argue that, considering the two-player game in the formulation of GANs, only making the generator 3D-aware is not enough. In other words, displacing the generative mechanism only offers the capability, but not the guarantee, of producing 3D-aware images, because the supervision of the generator primarily comes from the discriminator. To address this issue, we propose GeoD through learning a geometry-aware discriminator to improve 3D-aware GANs. Concretely, besides differentiating real and fake samples from the 2D image space, the discriminator is additionally asked to derive the geometry information from the inputs, which is then applied as the guidance of the generator. Such a simple yet effective design facilitates learning substantially more accurate 3D shapes. Extensive experiments on various generator architectures and training datasets verify the superiority of GeoD over state-of-the-art alternatives. Moreover, our approach is registered as a general framework such that a more capable discriminator (i.e., with a third task of novel view synthesis beyond domain classification and geometry extraction) can further assist the generator with a better multi-view consistency. Project page can be found at https://vivianszf.github.io/geod.
Zifan Shi, Yinghao Xu 0001, Yujun Shen, Deli Zhao, Qifeng Chen 0001, Dit-Yan Yeung
NeurIPS1
2021 Neural Camera Simulators
abstract
We present a controllable camera simulator based on deep neural networks to synthesize raw image data under different camera settings, including exposure time, ISO, and aperture. The proposed simulator includes an exposure module that utilizes the principle of modern lens designs for correcting the luminance level. It also contains a noise module using the noise level function and an aperture module with adaptive attention to simulate the side effects on noise and defocus blur. To facilitate the learning of a simulator model, we collect a dataset of the 10,000 raw images of 450 scenes with different exposure settings. Quantitative experiments and qualitative comparisons show that our approach outperforms relevant baselines in raw data synthesize on multiple cameras. Furthermore, the camera simulator enables various applications, including large-aperture enhancement, HDR, auto exposure, and data augmentation for training local feature detectors. Our work represents the first attempt to simulate a camera sensor’s behavior leveraging both the advantage of traditional raw sensor features and the power of data-driven deep learning. The code and the dataset are available at https://github.com/ken-ouyang/neural_image_simulator.
Hao Ouyang, Zifan Shi, Chenyang Lei, Ka Lung Law, Qifeng Chen 0001
CVPR2
2021 Stereo Waterdrop Removal with Row-wise Dilated Attention
abstract
Existing vision systems for autonomous driving or robots are sensitive to waterdrops adhered to windows or camera lenses. Most recent waterdrop removal approaches take a single image as input and often fail to recover the missing content behind waterdrops faithfully. Thus, we propose a learning-based model for waterdrop removal with stereo images. To better detect and remove waterdrops from stereo images, we propose a novel row-wise dilated attention module to enlarge attention’s receptive field for effective information propagation between the two stereo images. In addition, we propose an attention consistency loss between the ground-truth disparity map and attention scores to enhance the left-right consistency in stereo images. Because of related datasets’ unavailability, we collect a real-world dataset that contains stereo images with and without waterdrops. Extensive experiments on our dataset suggest that our model outperforms state-of-the-art methods both quantitatively and qualitatively. Our source code and the stereo waterdrop dataset are available at https://github.com/VivianSZF/Stereo-Waterdrop-Removal.
Zifan Shi, Na Fan 0002, Dit-Yan Yeung, Qifeng Chen 0001
IROS1