Qiang Tang 0002

dblp:17/2212-2 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 ArchitectHead: Continuous Level of Detail Control for 3D Gaussian Head Avatars
abstract
3D Gaussian Splatting (3DGS) has enabled photorealistic and real-time rendering of 3D head avatars. Existing 3DGS-based avatars typically rely on tens of thousands of 3D Gaussian points (Gaussians), with the number of Gaussians fixed after training. However, many practical applications require adjustable levels of detail (LOD) to balance rendering efficiency and visual quality. In this work, we propose "ArchitectHead", the first framework for creating 3D Gaussian head avatars that support continuous control over LOD. Our key idea is to parameterize the Gaussians in a 2D UV feature space and propose a UV feature field composed of multi-level learnable feature maps to encode their latent features. A lightweight neural network-based decoder then transforms these latent features into 3D Gaussian attributes for rendering. ArchitectHead controls the number of Gaussians by dynamically resampling feature maps from the UV feature field at the desired resolutions. This method enables efficient and continuous control of LOD without retraining. Experimental results show that ArchitectHead achieves state-of-the-art (SOTA) quality in self and cross-identity reenactment tasks at the highest LOD, while maintaining near SOTA performance at lower LODs. At the lowest LOD, our method uses only 6.2% of the Gaussians while the quality degrades moderately (L1 Loss +7.9%, PSNR −0.97%, SSIM −0.6%, LPIPS Loss +24.1%), and the rendering speed nearly doubles. Project homepage: https://peizhiyan.github.io/docs/architect/.
Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001
WACV3
2025 One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild
abstract
Current Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they often falter when applied to in-the-wild samples, primarily due to the distribution gap between the training datasets and real-world test samples. While some researchers aim to enhance model generalizability through sophisticated training procedures, advanced architectures, or by creating more diverse datasets, we adopt the test-time fine-tuning paradigm to customize a pre-trained Text2Image (T2I) model. However, naively applying test-time tuning results in inconsistencies in facial identities and appearance attributes. To address this, we introduce a Visual Consistency Module (VCM), which enhances appearance consistency by combining the face, text, and image embedding. Our approach, named OnePoseTrans, requires only a single source image to generate high-quality pose transfer results, offering greater stability than state-of-the-art data-driven methods. For each test case, OnePoseTrans customizes a model in around 48 seconds with an NVIDIA V100 GPU.
Dongqi Fan, Rui Ma 0011, Qiang Tang 0002, Zili Yi
ICASSP5
2025 Estimating Virtual Camera FOV to Reduce Perspective Shape Distortion in 2D-to-3D Face Reconstruction
abstract
Existing image-based 3D face reconstruction methods rely on a virtual camera to project the reconstructed 3D face onto the 2D image plane for comparison with the input image, a crucial step for accurate results. To simplify the reconstruction process, these methods often use fixed camera intrinsics and assume minimal perspective distortion, overlooking the varying distortion levels in "in-the-wild" images and leading to inaccuracies in reconstructed 3D face shapes. To address this issue, we propose estimating the virtual camera’s optimal field-of-view (FOV) for a given image, enabling consistent 3D face reconstruction across varying distortion levels. We introduce two synthetic datasets: one to train our FOV estimation network (FOV-Net) and another to evaluate its performance and reconstruction accuracy. We use the FOV-Net predicted FOV to initialize the camera, which is used in the fitting-based reconstruction process. Experiments show that our approach significantly improves reconstruction consistency under different levels of perspective distortion.
Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001
ICIP3
2025 FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
abstract
Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining, limiting flexibility and generalization. Inversion-based approaches offer stronger alignment but incur high inference cost due to dual-path denoising. We present \textbf{FreeControl}, a training-free framework for semantic structural control in diffusion models. Unlike prior methods that extract attention across multiple timesteps, FreeControl performs \textit{one-step attention extraction} from a single, optimally chosen timestep and reuses it throughout denoising. This enables efficient structural guidance without inversion or retraining. To further improve quality and stability, we introduce \textit{Latent-Condition Decoupling (LCD)}: a principled separation of the timestep condition and the noised latent used in attention extraction. LCD provides finer control over attention quality and eliminates structural artifacts. FreeControl also supports compositional control via reference images assembled from multiple sources, enabling intuitive scene layout design and stronger prompt alignment. FreeControl introduces a new paradigm for test-time control—enabling structurally and semantically aligned, visually coherent generation directly from raw images, with the flexibility for intuitive compositional design and compatibility with modern diffusion models at ~5\% additional cost.
Jiang Lin, Zhiqiu Zhang, Jizhi Zhang, Qiang Tang 0002, Zili Yi
NeurIPS7
2025 Towards Secure and Usable 3D Assets: A Novel Framework for Automatic Visible Watermarking
abstract
3D models, particularly AI-generated ones, have wit-nessed a recent surge across various industries such as en-tertainment. Hence, there is an alarming need to protect the intellectual property and avoid the misuse of these valuable assets. As a viable solution to address these concerns, we rigorously define the novel task of automated 3D visible wa-termarking in terms of two competing aspects: watermark quality and asset utility. Moreover, we propose a method of embedding visible watermarks that automatically deter-mines the right location, orientation, and number of wa-termarks to be placed on arbitrary 3D assets for high wa-termark quality and asset utility. Our method is based on a novel rigid-body optimization that uses back-propagation to automatically learn transforms for ideal watermark place-ment. In addition, we propose a novel curvature-matching method for fusing the watermark into the 3D model that further improves readability and security. Finally, we provide a detailed experimental analysis on two benchmark 3D datasets validating the superior performance of our approach in comparison to baselines. Code and demo are available here11https://developer.huaweicloud.com/develop/aigallery/notebook/detail?id=15adbaaa-2583-4ec3-804a-61c29f001e03.
Gursimran Singh, Tianxi Hu, Qiang Tang 0002, Yong Zhang 0004
WACV4
2025 Gaussian Déjà-vu: Creating Controllable 3D Gaussian Head-Avatars with Enhanced Generalization and Personalization Abilities
abstract
Recent advancements in 3D Gaussian Splatting (3DGS) have unlocked significant potential for modeling 3D head avatars, providing greater flexibility than mesh-based methods and more efficient rendering compared to NeRF-based approaches. Despite these advancements, the creation of controllable 3DGS-based head avatars remains time-intensive, often requiring tens of minutes to hours. To expedite this process, we here introduce the “Gaussian Déjà-vu” framework, which first obtains a generalized model of the head avatar and then personalizes the result. The generalized model is trained on large 2D (synthetic and real) image datasets. This model provides a well-initialized 3D Gaussian head that is further refined using a monocular video to achieve the personalized head avatar. For personalizing, we propose learnable expression-aware rectification blendmaps to correct the initial 3D Gaussians, ensuring rapid convergence without the reliance on neural networks. Experiments demonstrate that the proposed method meets its objectives. It outperforms state-of-the-art 3D Gaussian head avatars in terms of photorealistic quality as well as reduces training time consumption to at least a quarter of the existing methods, producing the avatar in minutes. Project homepage: https://peizhiyan.github.io/docs/dejavu
Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001
WACV3
2025 Neural 3D Face Shape Stylization Based on Single Style Template via Weakly Supervised Learning
abstract
3D Face shape stylization refers to transforming a realistic 3D face shape into a different style, such as a cartoon face style. To solve this problem, this paper proposes modeling this task as a deformation transfer problem. This approach significantly reduces labor costs, as the artists would only need to create a single template for each face style. Realistic facial features of the original 3D face e.g. the nose or chin shape, would thus be automatically transferred to those in the style template. Deformation transfer methods, however, have two drawbacks. They are slow and they require re-optimization for every new input face. To address these weaknesses, we propose a neural network-based 3D face shape stylization method. This method is trained through weakly supervised learning, and its template's structure is preserved using our novel template-guided mesh smoothing regularization. Our method is the first learning-based deformation transfer method for 3D face shape stylization. Its employment offers the useful and practical benefit of not requiring paired training data. The experiments show that the quality of the stylized faces obtained by our method is comparable to that of the traditional deformation transfer method, achieving an average Chamfer Distance of approximately 0.01 mm. However, our approach significantly boosts the processing speed, achieving a rate approximately 3,000 times faster than the traditional deformation transfer.
Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001
IEEE Trans. Vis. Comput. Graph.3
2023 Learning Disentangled Features for Nerf-Based Face Reconstruction
abstract
The 3D-aware parametric face model named HeadNeRF achieved advantages in rendering photo-realistic face images. However, it has two limitations: (1) it uses single-image fitting reconstruction that is slow and prone to overfitting; (2) it lacks explicit 3D geometry information, making using semantic facial-parts-based loss challenging. This paper presents a 3D-aware face reconstruction learning framework tailored for HeadNeRF to address the limitations. We train a face encoder network that can directly learn the disentangled features for facial reconstruction to address the first limitation. For the second limitation, we introduce a lightweight semantic face segmentation network and facial-parts-based loss function to improve the reconstruction accuracy and quality. Our experiments show that the proposed method achieves a low reconstruction time consumption and enhanced reconstruction accuracy. Project page: https://peizhiyan.github.io/docs/headnerf+
Peizhi Yan, Rabab K. Ward, Dan Wang 0011, Qiang Tang 0002, Shan Du 0001
ICIP4
2022 NEO-3DF: Novel Editing-Oriented 3D Face Creation and Reconstruction
Peizhi Yan, James Gregson, Qiang Tang 0002, Rabab K. Ward, Shan Du 0001
ACCV (1)3
2021 Spatial-Temporal Residual Aggregation for High Resolution Video Inpainting
Vishnu Sanjay Ramiya Srinivasan, Rui Ma 0011, Qiang Tang 0002, Zili Yi
BMVC3
2020 Contextual Residual Aggregation for Ultra High-Resolution Image Inpainting
abstract
Recently data-driven image inpainting methods have made inspiring progress, impacting fundamental image editing tasks such as object removal and damaged image repairing. These methods are more effective than classic approaches, however, due to memory limitations they can only handle low-resolution inputs, typically smaller than 1K. Meanwhile, the resolution of photos captured with mobile devices increases up to 8K. Naive up-sampling of the low-resolution inpainted result can merely yield a large yet blurry result. Whereas, adding a high-frequency residual image onto the large blurry image can generate a sharp result, rich in details and textures. Motivated by this, we propose a Contextual Residual Aggregation (CRA) mechanism that can produce high-frequency residuals for missing contents by weighted aggregating residuals from contextual patches, thus only requiring a low-resolution prediction from the network. Since convolutional layers of the neural network only need to operate on low-resolution inputs and outputs, the cost of memory and computing power is thus well suppressed. Moreover, the need for high-resolution training datasets is alleviated. In our experiments, we train the proposed model on small images with resolutions 512 × 512 and perform inference on high-resolution images, achieving compelling inpainting quality. Our model can inpaint images as large as 8K with considerable hole sizes, which is intractable with previous learning-based approaches. We further elaborate on the light-weight design of the network architecture, achieving real-time performance on 2K images on a GTX 1080 Ti GPU. Codes are available at: https://github. com/Ascend-Huawei/Ascend-Canada/tree/ master/Models/Research_HiFIll_Model.
Zili Yi, Qiang Tang 0002, Shekoofeh Azizi, Daesik Jang
CVPR2
2020 Animating Through Warping: An Efficient Method for High-Quality Facial Expression Animation
abstract
Advances in deep neural networks have considerably improved the art of animating a still image without operating in 3D domain. Whereas, prior arts can only animate small images (typically no larger than 512x512) due to memory limitations, difficulty of training and lack of high-resolution (HD) training datasets, which significantly reduce their potential for applications in movie production and interactive systems. Motivated by the idea that HD images can be generated by adding high-frequency residuals to low-resolution results produced by a neural network, we propose a novel framework known as Animating Through Warping (ATW) to enable efficient animation of HD images.
Zili Yi, Qiang Tang 0002, Vishnu Sanjay Ramiya Srinivasan
ACM Multimedia2
2010 Fast block-size partitioning using empirical rate-distortion models for MPEG-2 to H.264/AVC transcoding
abstract
We present an efficient H.264/AVC block-size partitioning prediction method, which is based on our proposed empirical rate and distortion models. Compared to other state-of-the-art transcoding methods, and for the same rate-distortion performance, our proposed algorithm requires the least computational complexity, reaching a 73% reduction in variable block-size motion estimation for SDTV sequences, and 71% reduction for CIF sequences.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ISCAS1
2010 Efficient Motion Re-Estimation With Rate-Distortion Optimization for MPEG-2 to H.264/AVC Transcoding
abstract
One objective in MPEG-2 to H.264/advanced video coding transcoding is to improve the H.264/AVC compression ratio by using more advanced macroblock encoding modes. The motion re-estimation process is by far the most time-consuming process in this type of video transcoding. In this paper, we present an efficient H.264/AVC block size partitioning prediction algorithm for MPEG-2 to H.264/AVC transcoding applications. Our algorithm uses rate-distortion optimization techniques and predicted initial motion vectors to estimate block size partitioning. It is also shown that using block size partitioning smaller than 8 × 8 (i.e., 8 × 4, 4 × 8, and 4 × 4) results in negligible compression improvements, and thus these sizes should be avoided in transcoding. Experimental results show that, compared to the state-of-the-art transcoding scheme, our transcoder yields similar rate-distortion performance, while the computational complexity is significantly reduced, requiring an average of 29% of the computations. Compared to the full-search scheme, our proposed algorithm reduces the computational complexity by about 99.47% for standard-definition television sequences and 98.66% for common intermediate format sequences. Compared to UMHexagonS, the fast motion estimation algorithm used in H.264/AVC, the experimental results show that our proposed algorithm is a better trade-off between computational complexity and picture quality.
Qiang Tang 0002, Panos Nasiopoulos
IEEE Trans. Circuits Syst. Video Technol.1
2009 Efficient motion vector re-estimation for MPEG-2 TO H.264/AVC transcoding with arbitrary down-sizing ratios
abstract
As for down-sizing MPEG-2 to H.264/AVC transcoding, an efficient algorithm of estimating initial H.264/AVC motion vectors is proposed. By using the estimated initial motion vectors, only a small range of motion vector refinement is sufficient to find the final motion vector for each partition. Experimental results show that our proposed algorithm achieves average 0.08 dB improvement (maximum 0.24 dB) in the picture quality compared to the other state-of-art method. At the same time, the computational complexity of estimating the initial motion vectors is less than that of the other state-of-art technique (average 36% reduction).
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ICIP1
2008 Fast block size prediction for MPEG-2 to H.264/AVC transcoding
abstract
One objective in MPEG-2 to H.264 transcoding is to improve the H.264 compression ratio by using more accurate H.264 motion vectors. Motion re-estimation is by far the most time consuming process in video transcoding, and improving the searching speed is a challenging problem. We introduce a new transcoding scheme that uses the MPEG-2 DCT coefficients to predict the block size partitioning for H.264. Performance evaluations have shown that, for the same rate-distortion performance, our proposed scheme achieves an impressive reduction in the computational complexity of more than 82% compared to the full range motion estimation used by H.264.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ICASSP1
2008 Compensation of Requantization and Interpolation Errors in MPEG-2 to H.264 Transcoding
abstract
Implementing MPEG-2 to H.264 transcoding schemes in the pixel domain introduces a high degree of computational complexity. In the transform domain, this transcoding is more computationally efficient, and several methods have been developed to address that approach. However, incompatibilities between the two standards, such as the mismatches between the MPEG-2 and H.264 motion compensation processes, cause several distortions that may affect the overall picture quality. In this study, we address the main distortions that result from requantization errors: luminance half-pixel and chrominance quarter/three-quarter interpolation errors. Then, we propose algorithms that compensate for these errors. The traditional requantization error compensation algorithm for DCT coefficients is updated so that it can be applied to the H.264 integer transform coefficients. Equations that compensate for the luminance half-pixel and chrominance quarter/three-quarter pixel interpolation errors are derived. To remove the interpolation errors, the previous H.264 frame is needed. Thus, the compensation scheme includes a closed-loop H.264 motion compensation process, which is implemented in the pixel domain. To evaluate the performance of the proposed compensation algorithms in terms of picture quality, our scheme is compared with two different cascaded pixel-domain transcoding structures. The first structure reuses the MPEG-2 motion vectors, and the other implements plusmn2 pixels motion vector refinement, but each one has an H.264 deblocking filter. The experimental results show that the proposed compensation algorithms achieve 5-dB quality improvement over the open-loop transform-domain-based transcoding and almost the same picture quality (0.3-0.6 dB) as the cascaded structures. An additional advantage is the reduction in computational complexity that ranges from 13% to 69% compared with the two cascaded methods.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
IEEE Trans. Circuits Syst. Video Technol.1
2007 Efficient Chrominance Compensation for MPEG2 to H.264 Transcoding
abstract
Although open-loop transcoding is known as the most computational efficient transcoding structure, it is also known to introduce many distortions in the transcoded video. This paper addresses the chrominance distortions resulting from the open-loop MPEG2 to H.264 transcoding structure and proposes algorithms to compensate for the chrominance distortions. The open-loop structure is replaced by a closed-loop transcoding structure, which provides high-quality video by removing the chrominance distortions, resulting in an average of 6 dB picture quality improvement.
Qiang Tang 0002, Panos Nasiopoulos, Rabab K. Ward
ICASSP (1)1
2006 An Efficient MPEG2 to H.264 Half-Pixel Motion Compensation Transcoding
abstract
An efficient MPEG2 to H.264 half-pixel motion compensation transcoding method is proposed. The Inter macroblock transcoding is implemented in the transform domain. An algorithm is designed to compensate for the errors which arise because of the different half-pixel interpolation procedures used by MPEG2 and H.264/AVC. The experimental results show that the PSNR values of the transcoded H.264 streams result in significant improvement (average 5.5 dB) after we reduce the half-pixel interpolation errors.
Qiang Tang 0002, Rabab K. Ward, Panos Nasiopoulos
ICIP1