Zhihong Pan 0001

dblp:59/4950-1 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0003-0866-762XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Gradient-Free Classifier Guidance for Diffusion Model Sampling
abstract
Unguided sampling in diffusion models is known to generate images of wide variations, albeit with a trade-off in image fidelity. Guided sampling methods, such as classifier guidance (CG) and classifier-free guidance (CFG), focus on sampling in well-learned high-probability regions to generate images of high fidelity, but each has its limitations. CG is computationally expensive due to the costly classifier gradient descent process, while CFG, being gradient-free, is more efficient but compromises class label alignment compared to CG. In this work, we introduce Gradient-Free Classifier Guidance (GFCG), a novel method utilizing a pre-trained classifier solely in inference mode, entirely avoiding gradient calculations. GFCG introduces two innovative mechanisms: (1) adaptive reference class selection, dynamically determining an appropriate undesired reference class; and (2) adaptive guidance scaling, dynamically adjusting the guidance strength based on classifier confidence during sampling. Experiments on both class-conditioned and text-to-image generation demonstrate that the proposed GFCG method consistently improves class prediction accuracy while preserving diversity. We also show that GFCG is complementary to other guided sampling methods. When combined with Autoguidance (ATG), GFCG attains record performance on ImageNet 512×512, with a FDDINOv2of 23.39, while maintaining a high Precision of 94.0% compared to ATG’s 89.9%.
Rahul Shenoy, Zhihong Pan 0001, Kaushik Balakrishnan, Qisen Cheng, Yongmoon Jeon, Heejune Yang
WACV2
2024 GBSD: Generative Bokeh with Stage Diffusion
abstract
The bokeh effect is an artistic technique that blurs out-of-focus areas in a photograph and has gained interest due to recent developments in text-to-image synthesis and the ubiquity of smartphone cameras and photo sharing apps. Prior work on rendering bokeh effects have focused on manipulating photographs using classical computer graphics or neural rendering techniques, but have either depth discontinuity artifacts or are restricted to reproducing bokeh effects that are present in the training data. In this paper, we present generative bokeh with stage diffusion (GBSD), the first generative text-to-image model that synthesizes photorealistic images with a bokeh style. Motivated by how image synthesis occurs progressively in diffusion models, our approach combines latent diffusion models with a 2-stage conditioning algorithm to render bokeh effects on semantically defined objects. Since GBSD focuses the blurring effect on objects, this semantic bokeh effect is more versatile than classical rendering techniques. We evaluate GBSD both quantitatively and qualitatively and demonstrate its ability to be applied in both text-to-image and image-to-image settings.
Jieren Deng, Xin Zhou 0017, Hao Tian 0005, Zhihong Pan 0001, Derek Aguiar
ICASSP4
2023 Diffusion Motion: Generate Text-Guided 3D Human Motion by Diffusion Model
abstract
We propose a simple and novel method for generating 3D human motion from complex natural language sentences, which describe different velocity, direction and composition of all kinds of actions. Different from existing methods that use classical generative architecture, we apply the Denoising Diffusion Probabilistic Model to this task, synthesizing diverse motion results under the guidance of texts. The diffusion model converts white noise into structured 3D motion by a Markov process with a series of denoising steps and is efficiently trained by optimizing a variational lower bound. To achieve the goal of text-conditioned image synthesis, we use the classifier-free guidance strategy to add text embedding into the model during training. Our experiments demonstrate that our model achieves competitive results on HumanML3D test set quantitatively and can generate more visually natural and diverse examples. We also show with experiments that our model is capable of zero-shot generation of motions for unseen text guidance.
Zhihong Pan 0001, Xin Zhou 0017
ICASSP2
2023 Raising The Limit of Image Rescaling Using Auxiliary Encoding
abstract
Normalizing flow models using invertible neural networks (INN) have been widely investigated for successful generative image super-resolution (SR) by learning the transformation between the normal distribution of latent variable z and the conditional distribution of high-resolution (HR) images gave a low-resolution (LR) input. Recently, image rescaling models like IRN utilize the bidirectional nature of INN to push the performance limit of image upscaling by optimizing the downscaling and upscaling steps jointly. While the random sampling of latent variable z is useful in generating diverse photo-realistic images, it is not desirable for image rescaling when accurate restoration of the HR image is more important. Hence, in places of random sampling of z, we propose auxiliary encoding modules to further push the limit of image rescaling performance. Two options to store the encoded latent variables in downscaled LR images, both readily supported in existing image file format, are proposed. One is saved as the alpha-channel, the other is saved as meta-data in the image header, and the corresponding modules are denoted as suffixes -A and -M respectively. Optimal network architectural changes are investigated for both options to demonstrate their effectiveness in raising the rescaling performance limit on different baseline models including IRN and DLV-IRN.
Chenzhong Yin, Zhihong Pan 0001, Xin Zhou 0017, Paul Bogdan
ICASSP2
2023 Effective Real Image Editing with Accelerated Iterative Diffusion Inversion
abstract
Despite all recent progress, it is still challenging to edit and manipulate natural images with modern generative models. When using Generative Adversarial Network (GAN), one major hurdle is in the inversion process mapping a real image to its corresponding noise vector in the latent space, since it is necessary to be able to reconstruct an image to edit its contents. Likewise for Denoising Diffusion Implicit Models (DDIM), the linearization assumption in each inversion step makes the whole deterministic inversion process unreliable. Existing approaches that have tackled the problem of inversion stability often incur in significant trade-offs in computational efficiency. In this work we propose an Accelerated Iterative Diffusion Inversion method, dubbed AIDI, that significantly improves reconstruction accuracy with minimal additional overhead in space and time complexity. By using a novel blended guidance technique, we show that effective results can be obtained on a large range of image editing tasks without large classifier-free guidance in inversion. Furthermore, when compared with other diffusion inversion based works, our proposed process is shown to be more robust for fast image editing in the 10 and 20 diffusion steps’ regimes.
Zhihong Pan 0001, Riccardo Gherardi, Xiufeng Xie, Stephen Huang
ICCV1
2023 HollowNeRF: Pruning Hashgrid-Based NeRFs with Trainable Collision Mitigation
abstract
Neural radiance fields (NeRF) have garnered significant attention, with recent works such as Instant-NGP accelerating NeRF training and evaluation through a combination of hashgrid-based positional encoding and neural networks. However, effectively leveraging the spatial sparsity of 3D scenes remains a challenge. To cull away unnecessary regions of the feature grid, existing solutions rely on prior knowledge of object shape or periodically estimate object shape during training by repeated model evaluations, which are costly and wasteful. To address this issue, we propose HollowNeRF, a novel compression solution for hashgrid-based NeRF which automatically sparsifies the feature grid during the training phase. Instead of directly compressing dense features, HollowNeRF trains a coarse 3D saliency mask that guides efficient feature pruning, and employs an alternating direction method of multipliers (ADMM) pruner to sparsify the 3D saliency mask during training. By exploiting the sparsity in the 3D scene to redistribute hash collisions, HollowNeRF improves rendering quality while using a fraction of the parameters of comparable state-of-the-art solutions, leading to a better cost-accuracy trade-off. Our method delivers comparable rendering quality to Instant-NGP, while utilizing just 31% of the parameters. In addition, our solution can achieve a PSNR accuracy gain of up to 1dB using only 56% of the parameters.
Xiufeng Xie, Riccardo Gherardi, Zhihong Pan 0001, Stephen Huang
ICCV3
2023 Smooth and Stepwise Self-Distillation for Object Detection
abstract
Distilling the structured information captured in feature maps has contributed to improved results for object detection tasks, but requires careful selection of baseline architectures and substantial pre-training. Self-distillation addresses these limitations and has recently achieved state-of-the-art performance for object detection despite making several simplifying architectural assumptions. Building on this work, we propose Smooth and Stepwise Self-Distillation (SSSD) for object detection. Our SSSD architecture forms an implicit teacher from object labels and a feature pyramid network backbone to distill label-annotated feature maps using Jensen-Shannon distance, which is smoother than distillation losses used in prior work. We additionally add a distillation coefficient that is adaptively configured based on the learning rate. We extensively benchmark SSSD against a baseline and two state-of-the-art object detector architectures on the COCO dataset by varying the coefficients and backbone and detector networks. We demonstrate that SSSD achieves higher average precision in most experimental settings, is robust to a wide range of coefficients, and benefits from our stepwise distillation procedure.
Jieren Deng, Xin Zhou 0017, Hao Tian 0005, Zhihong Pan 0001, Derek Aguiar
ICIP4
2023 Effective Invertible Arbitrary Image Rescaling
abstract
Great successes have been achieved using deep learning techniques for image super-resolution (SR) with fixed scales. To increase its real world applicability, numerous models have also been proposed to restore SR images with arbitrary scale factors, including asymmetric ones where images are resized to different scales along horizontal and vertical directions. Though most models are only optimized for the unidirectional upscaling task while assuming a predefined downscaling kernel for low-resolution (LR) inputs, recent models based on Invertible Neural Networks (INN) are able to increase upscaling accuracy significantly by optimizing the downscaling and upscaling cycle jointly. However, limited by the INN architecture, it is constrained to fixed integer scale factors and requires one model for each scale. Without increasing model complexity, a simple and effective invertible arbitrary rescaling network (IARN) is proposed to achieve arbitrary image rescaling by training only one model in this work. Using innovative components like position-aware scale encoding and preemptive channel splitting, the network is optimized to convert the non-invertible rescaling cycle to an effectively invertible process. It is shown to achieve a state-of-the-art (SOTA) performance in bidirectional arbitrary rescaling without compromising perceptual quality in LR outputs. It is also demonstrated to perform well on tests with asymmetric scales using the same network architecture.
Zhihong Pan 0001, Baopu Li, Dongliang He, Errui Ding
WACV1
2023 Arbitrary Style Guidance for Enhanced Diffusion-Based Text-to-Image Generation
abstract
Diffusion-based text-to-image generation models like GLIDE and DALLE-2 have gained wide success recently for their superior performance in turning complex text inputs into images of high quality and wide diversity. In particular, they are proven to be very powerful in creating graphic arts of various formats and styles. Although current models supported specifying style formats like oil painting or pencil drawing, fine-grained style features like color distributions and brush strokes are hard to specify as they are randomly picked from a conditional distribution based on the given text input. Here we propose a novel style guidance method to support generating images using arbitrary style guided by a reference image. The generation method does not require a separate style transfer model to generate desired styles while maintaining image quality in generated content as controlled by the text input. Additionally, the guidance method can be applied without a style reference, denoted as self style guidance, to generate images of more diverse styles. Comprehensive experiments prove that the proposed method remains robust and effective in a wide range of conditions, including diverse graphic art forms, image content types and diffusion models.
Zhihong Pan 0001, Xin Zhou 0017, Hao Tian 0005
WACV1
2023 Bidirectional Translation Between UHD-HDR and HD-SDR Videos
abstract
With the popularization of ultra high definition (UHD) high dynamic range (HDR) displays, recent works focus onupgradinghigh definition (HD) standard dynamic range (SDR) videos to UHD-HDR versions, aiming to provides richer details and higher contrasts on advanced modern displays. However, joint considering theupgrading&downgradingtranslations between two types of videos, which is practical in real applications, is generally neglected. On the one hand,downgradingtranslation is the key to showing UHD-HDR videos on HD-SDR displays. On the other hand, considering both translations enables joint optimization and results in high quality translation. To this end, we propose the bidirectional translation network (BiT-Net), which jointly considers two translations in one network for the first time. In brief, BiT-Net is elaborately designed in aninvertiblefashion that can be efficiently inferred along forward and backward directions fordowngradingandupgradingtasks, respectively. Based on this framework, we divide each direction into three sub-tasks,i.e., decomposition, structure-guided translation, and synthesis, to effectively translate the dynamic range and the high-frequency details. Benefiting from the dedicated architecture, our BiT-Net can work on 1) downgrading UHD-HDR videos, 2) upgrading existing HD-SDR videos, and 3) synthesizing UHD-HDR versions from the downgraded HD-SDR videos. Experiments show that the proposed method achieves state-of-the-art performances on all these three tasks.
Mingde Yao, Dongliang He, Xin Li 0106, Zhihong Pan 0001, Zhiwei Xiong
IEEE Trans. Multim.4
2022 Towards Bidirectional Arbitrary Image Rescaling: Joint Optimization and Cycle Idempotence
abstract
Deep learning based single image super-resolution models have been widely studied and superb results are achieved in upscaling low-resolution images with fixed scale factor and downscaling degradation kernel. To improve real world applicability of such models, there are growing interests to develop models optimized for arbitrary upscaling factors. Our proposed method is the first to treat arbitrary rescaling, both upscaling and downscaling, as one unified process. Using joint optimization of both directions, the proposed model is able to learn upscaling and downscaling simultaneously and achieve bidirectional arbitrary image rescaling. It improves the performance of current arbitrary upscaling models by a large margin while at the same time learns to maintain visual perception quality in downscaled images. The proposed model is further shown to be robust in cycle idempotence test, free of severe degradations in reconstruction accuracy when the downscaling-to-upscaling cycle is applied repetitively. This robustness is beneficial for image rescaling in the wild when this cycle could be applied to one image for multiple times. It also performs well on tests with arbitrary large scales and asymmetric scales, even when the model is not trained with such tasks. Extensive experiments are conducted to demonstrate the superior performance of our model.
Zhihong Pan 0001, Baopu Li, Dongliang He, Mingde Yao, Xin Li 0106, Errui Ding
CVPR1
2022 Learning Adjustable Image Rescaling with Joint Optimization of Perception and Distortion
abstract
The performance of image super-resolution (SR) have been greatly advanced by deep learning techniques recently. Most models are only optimized for the ill-posed upscaling task while assuming a predefined downscaling kernel for low-resolution (LR) inputs. Additionally, there exists a conflict between the objective and perceptual qualities of upscaled outputs for optimizing these models. To achieve an effective trade-off between these two qualities, the current methods are either inflexible as the model is optimized for a fixed tradeoff, or inefficient as it needs to interpolate weights or images from two separately trained models. Based on the invertible rescaling net (IRN) which learns image downscaling and upscaling together, we propose a joint optimization method to train just one model that could achieve adjustable trade-off between perception and distortion for upscaling at inference time. Additionally, it’s shown in experiments that this jointly optimized model could produce results with better accuracy while maintaining high perceptual quality compared to one optimized for perceptual quality only.
Zhihong Pan 0001
ICASSP1
2022 Enhancing Image Rescaling using Dual Latent Variables in Invertible Neural Network
abstract
Normalizing flow models have been used successfully for generative image super-resolution (SR) by approximating complex distribution of natural images to simple tractable distribution in latent space through Invertible Neural Networks (INN). These models can generate multiple realistic SR images from one low-resolution (LR) input using randomly sampled points in the latent space, simulating the ill-posed nature of image upscaling where multiple high-resolution (HR) images correspond to the same LR. Lately, the invertible process in INN has also been used successfully by bidirectional image rescaling models like IRN and HCFlow for joint optimization of downscaling and inverse upscaling, resulting in significant improvements in upscaled image quality. While they are optimized for image downscaling too, the ill-posed nature of image downscaling, where one HR image could be downsized to multiple LR images depending on different interpolation kernels and resampling methods, is not considered. A new downscaling latent variable, in addition to the original one representing uncertainties in image upscaling, is introduced to model variations in the image downscaling process. This dual latent variable enhancement is applicable to different image rescaling models and it is shown in extensive experiments that it can improve image upscaling accuracy consistently without sacrificing image quality in downscaled LR images. It is also shown to be effective in enhancing other INN-based models for image restoration applications like image hiding.
Min Zhang 0030, Zhihong Pan 0001, Xin Zhou 0017, C.-C. Jay Kuo
ACM Multimedia2
2021 Real Image Super-Resolution Using Token Based Contextual Attention
Zhihong Pan 0001, Baopu Li
ICASSP1
2021 Cursor-based Adaptive Quantization for Deep Convolutional Neural Network
Baopu Li, Yanwen Fan, Zhihong Pan 0001, Zhiyu Cheng
IJCNN3
2021 Automatic Channel Pruning with Hyper-parameter Search and Dynamic Masking
abstract
Modern deep neural network models tend to be large and computationally intensive. One typical solution to this issue is model pruning. However, most current model pruning algorithms depend on hand crafted rules or need to input the pruning ratio beforehand. To overcome this problem, we propose a learning based automatic channel pruning algorithm for deep neural network, which is inspired by recent automatic machine learning (Auto ML). A two objectives' pruning problem that aims for the weights and the remaining channels for each layer is first formulated. An alternative optimization approach is then proposed to derive the channel numbers and weights simultaneously. In the process of pruning, we utilize a searchable hyper-parameter, remaining ratio, to denote the number of channels in each convolution layer, and then a dynamic masking process is proposed to describe the corresponding channel evolution. To adjust the trade-off between accuracy of a model and the pruning ratio of floating point operations, a new loss function is further introduced. Extensive experimental results on benchmark datasets demonstrate that our scheme achieves competitive results for neural network pruning.
Baopu Li, Yanwen Fan, Zhihong Pan 0001, Yuchen Bian
ACM Multimedia3
2020 Deep Residual Network for MSFA Raw Image Denoising
abstract
Multispectral filter arrays (MSFA) is increasingly used in multispectral imaging. While many previous works studied the denoising algorithms for CFA based cameras, denoising MSFA raw images is little discussed. The major challenges for denoising MSFA data include 1) more channels than CFA and no predominant channel; 2) compatibility between denoising and the subsequent demosaicking process. To overcome these challenges, we propose a new deep residual network designed to account for the uniqueness of MSFA mosaic patterns. First, a split and stride convolution layer is innovated to match the mosaic pattern of the MSFA raw image. Then, data augmentation using MSFA shifting and dynamic noise is proposed to make the model robust to different noise levels. In addition, a new network optimization criteria is also suggested by using the noise standard deviation to normalize the L1 loss function. Comprehensive experiments demonstrate that the proposed deep residual network outperforms the state-of-the-art denoising algorithms in MSFA field.
Zhihong Pan 0001, Baopu Li, Hsuchun Cheng, Sid Ying-Ze Bao
ICASSP1
2003 Face Recognition in Hyperspectral Images
abstract
Hyperspectral cameras provide useful discriminants for human face recognition that cannot be obtained by other imaging methods. We examine the utility of using near-infrared hyperspectral images for the recognition of faces over a database of 200 subjects. The hyperspectral images were collected using a CCD camera equipped with a liquid crystal tunable filter. Spectral measurements over the near-infrared allow the sensing of subsurface tissue structure, which is significantly different from person to person but relatively stable over time. The local spectral properties of human tissue are nearly invariant to face orientation and expression, which allows hyperspectral discriminants to be used for recognition over a large range of poses and expressions. We describe a face recognition algorithm that exploits spectral measurements for multiple facial tissue types. We demonstrate experimentally that this algorithm can be used to recognize faces over time in the presence of changes in facial pose and expression.
Zhihong Pan 0001, Glenn Healey, Manish Prasad, Bruce J. Tromberg
CVPR (1)1
2003 Face Recognition in Hyperspectral Images
abstract
Hyperspectral cameras provide useful discriminants for human face recognition that cannot be obtained by other imaging methods. We examine the utility of using near-infrared hyperspectral images for the recognition of faces over a database of 200 subjects. The hyperspectral images were collected using a CCD camera equipped with a liquid crystal tunable filter to provide 31 bands over the near-infrared (0.7 /spl mu/m-1.0 /spl mu/m). Spectral measurements over the near-infrared allow the sensing of subsurface tissue structure which is significantly different from person to person, but relatively stable over time. The local spectral properties of human tissue are nearly invariant to face orientation and expression which allows hyperspectral discriminants to be used for recognition over a large range of poses and expressions. We describe a face recognition algorithm that exploits spectral measurements for multiple facial tissue types. We demonstrate experimentally that this algorithm can be used to recognize faces over time in the presence of changes in facial pose and expression.
Zhihong Pan 0001, Glenn Healey, Manish Prasad, Bruce J. Tromberg
IEEE Trans. Pattern Anal. Mach. Intell.1