EDBT 2026 Demo / reviewers in the wild / expert
Hao-Zhi Huang 0001
dblp:170/1729 · also Haozhi Huang 0001
· DBLP profile ↗
20ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-3925-0237ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion modelabstractDiffusion-based video generation techniques have significantly improved zero-shot talking-head avatar generation, enhancing the naturalness of both head motion and facial expressions. However, existing methods suffer from poor controllability, making them less applicable to real-world scenarios such as filmmaking and live streaming for e-commerce. To address this limitation, we propose FLAP, a novel approach that integrates explicit 3D intermediate parameters (head poses and facial expressions) into the diffusion model for end-to-end generation of realistic portrait videos. The proposed architecture allows the model to generate vivid portrait videos from audio while simultaneously incorporating additional control signals, such as head rotation angles and eye-blinking frequency. Furthermore, the decoupling of head pose and facial expression allows for independent control of each, offering precise manipulation of both the avatar's pose and facial expressions. We also demonstrate its flexibility in integrating with existing 3D head generation methods, bridging the gap between 3D model-based approaches and end-to-end diffusion techniques. Extensive experiments show that our method outperforms recent audio-driven portrait video models in both naturalness and controllability. Lingzhou Mu, Baiji Liu, Guiming Mo, Jiawei Jin, Kai Zhang 0012, Hao-Zhi Huang 0001 |
ACM Multimedia | 7 |
| 2025 | AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial AnimationabstractSpeech-driven 3D facial animation aims at generating facial movements that are synchronized with the driving speech, which has been widely explored recently. Existing works mostly neglect the person-specific talking style in generation, including facial expression and head pose styles. Several works intend to capture the personalities by fine-tuning modules. However, limited training data leads to the lack of vividness. In this work, we proposeAdaMesh, a novel adaptive speech-driven facial animation approach, which learns the personalized talking style from a reference video of about 10 seconds and generates vivid facial expressions and head poses. Specifically, we propose mixture-of-low-rank adaptation (MoLoRA) to fine-tune the expression adapter, which efficiently captures the facial expression style. For the personalized pose style, we propose a pose adapter by building a discrete pose prior and retrieving the appropriate style embedding with a semantic-aware pose style matrix without fine-tuning. Extensive experimental results show that our approach outperforms state-of-the-art methods, preserves the talking style in the reference video, and generates vivid facial animation. Liyang Chen, Weihong Bao, Shun Lei, Boshi Tang, Zhiyong Wu 0001, Shiyin Kang, Hao-Zhi Huang 0001, Helen M. Meng |
IEEE Trans. Multim. | 7 |
| 2024 | FusionDreamer: Consistent Images Generation from Sparse-view ImagesabstractWe introduce FusionDreamer, a diffusion-based approach that can generate high-quality multiview-consistent images from any number of input images. Recent diffusion-based methods like Zero123 and SyncDreamer showcase the capability to generate novel views from a single input view through the utilization of pre-trained large-scale 2D diffusion models, but they are constrained by the restriction to a solitary view. This limitation results in the underutilization of valuable information in potential multiview scenarios, leading to challenges in accurately predicting intricate details and complex shapes. To address this limitation, we propose a Geometer-aware Pixel-wise Attention Mechanism, extending these methods to utilize multiview information when generating novel views. Experiments show that a simple extension of the single view generation model is not enough, and FusionDreamer can consider multiple views and obtain novel views with higher consistency and better quality. Risheng Huang, Hao-Zhi Huang 0001, Zongqing Lu 0001 |
ICME | 3 |
| 2023 | Is Bigger Always Better? An Empirical Study on Efficient Architectures for Style Transfer and BeyondabstractNetwork architecture plays a pivotal role in style transfer. Most existing algorithms use VGG19 as the feature extractor, which incurs a high computational cost. In this work, we conduct an empirical study on the popular network architectures and find that some more efficient networks can replace VGG19 while having comparable style transfer performance. Beyond that, we show that an efficient network can be further accelerated by removing its empty channels via a simple channel pruning method tweaked for style transfer. To prevent the potential performance drop due to using a more lightweight network and obtain better style transfer results, we introduce a more accurate deep feature alignment strategy to improve existing style transfer modules. Taking GoogLeNet as an exemplary efficient network, the pruned GoogLeNet with the improved style transfer module is 2.3 ~ 107.4× faster than the state-of-the-art approaches and can achieve 68.03 FPS on 512×512 images. Extensive experiments demonstrate that VGG19 can be replaced by a more lightweight network with significantly improved efficiency and comparable style transfer quality. Jie An 0002, Tao Li 0040, Hao-Zhi Huang 0001, Jinwen Ma, Jiebo Luo 0001 |
WACV | 3 |
| 2023 | Robust Pose Transfer With Dynamic Details Using Neural Video RenderingabstractPose transfer of human videos aims to generate a high-fidelity video of a target person imitating actions of a source person. A few studies have made great progress either through image translation with deep latent features or neural rendering with explicit 3D features. However, both of them rely on large amounts of training data to generate realistic results, and the performance degrades on more accessible Internet videos due to insufficient training frames. In this paper, we demonstrate that the dynamic details can be preserved even when trained from short monocular videos. Overall, we propose a neural video rendering framework coupled with an image-translation-based dynamic details generation network (D$^{2}$G-Net), which fully utilizes both the stability of explicit 3D features and the capacity of learning components. To be specific, a novel hybrid texture representation is presented to encode both the static and pose-varying appearance characteristics, which is then mapped to the image space and rendered as a detail-rich frame in the neural rendering stage. Through extensive comparisons, we demonstrate that our neural human video renderer is capable of achieving both clearer dynamic details and more robust performance even on accessible short videos with only 2 k$\sim$4 k frames, as illustrated in Fig. 1. Yang-Tian Sun, Hao-Zhi Huang 0001, Xuan Wang 0009, Yukun Lai, Wei Liu 0005, Lin Gao 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | High-Fidelity 3D Digital Human Head Creation from RGB-D SelfiesabstractWe present a fully automatic system that can produce high-fidelity, photo-realistic three-dimensional (3D) digital human heads with a consumer RGB-D selfie camera. The system only needs the user to take a short selfie RGB-D video while rotating his/her head and can produce a high-quality head reconstruction in less than 30 s. Our main contribution is a new facial geometry modeling and reflectance synthesis procedure that significantly improves the state of the art. Specifically, given the input video a two-stage frame selection procedure is first employed to select a few high-quality frames for reconstruction. Then a differentiable renderer-based 3D Morphable Model (3DMM) fitting algorithm is applied to recover facial geometries from multiview RGB-D data, which takes advantages of a powerful 3DMM basis constructed with extensive data generation and perturbation. Our 3DMM has much larger expressive capacities than conventional 3DMM, allowing us to recover more accurate facial geometry using merely linear basis. For reflectance synthesis, we present a hybrid approach that combines parametric fitting andConvolutional Neural Networks (CNNs)to synthesize high-resolution albedo/normal maps with realistic hair/pore/wrinkle details. Results show that our system can produce faithful 3D digital human faces with extremely realistic details. The main code and the newly constructed 3DMM basis is publicly available. Linchao Bao, Xiangkai Lin, Haoxian Zhang, Xuefei Zhe, Hao-Zhi Huang 0001, Xinwei Jiang, Jue Wang 0001, Dong Yu 0001, Zhengyou Zhang |
ACM Trans. Graph. | 8 |
| 2021 | UniFaceGAN: A Unified Framework for Temporally Consistent Facial Video EditingabstractRecent research has witnessed advances in facial image editing tasks including face swapping and face reenactment. However, these methods are confined to dealing with one specific task at a time. In addition, for video facial editing, previous methods either simply apply transformations frame by frame or utilize multiple frames in a concatenated or iterative fashion, which leads to noticeable visual flickers. In this paper, we propose a unified temporally consistent facial video editing framework termed UniFaceGAN. Based on a 3D reconstruction model and a simple yet efficient dynamic training sample selection mechanism, our framework is designed to handle face swapping and face reenactment simultaneously. To enforce the temporal consistency, a novel 3D temporal loss constraint is introduced based on the barycentric coordinate interpolation. Besides, we propose a region-aware conditional normalization layer to replace the traditional AdaIN or SPADE to synthesize more context-harmonious results. Compared with the state-of-the-art facial image editing methods, our framework generates video portraits that are more photo-realistic and temporally smooth. Meng Cao 0002, Hao-Zhi Huang 0001, Hao Wang 0050, Xuan Wang 0009, Li Shen 0008, Linchao Bao, Zhifeng Li 0001, Jiebo Luo 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Quantized Adam with Error FeedbackabstractIn this article, we present a distributed variant of an adaptive stochastic gradient method for training deep neural networks in the parameter-server model. To reduce the communication cost among the workers and server, we incorporate two types of quantization schemes, i.e., gradient quantization and weight quantization, into the proposed distributed Adam. In addition, to reduce the bias introduced by quantization operations, we propose an error-feedback technique to compensate for the quantized gradient. Theoretically, in the stochastic nonconvex setting, we show that the distributed adaptive gradient method with gradient quantization and error feedback converges to the first-order stationary point, and that the distributed adaptive gradient method with weight quantization and error feedback converges to the point related to the quantized level under both the single-worker and multi-worker modes. Last, we apply the proposed distributed adaptive gradient methods to train deep neural networks. Experimental results demonstrate the efficacy of our methods. Congliang Chen, Li Shen 0008, Hao-Zhi Huang 0001, Wei Liu 0005 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2021 | Aesthetic-guided outward image croppingabstractImage cropping is a commonly used post-processing operation for adjusting the scene composition of an input photography, therefore improving its aesthetics. Existing automatic image cropping methods are all bounded by the image border, thus have very limited freedom for aesthetics improvement if the original scene composition is far from ideal, e.g. the main object is too close to the image border. In this paper, we propose a novel, aesthetic-guided outward image cropping method. It can go beyond the image border to create a desirable composition that is unachievable using previous cropping methods. Our method first evaluates the input image to determine how much the content of the image should be extrapolated by a field of view (FOV) evaluation model. We then synthesize the image content in the extrapolated region, and seek an optimal aesthetic crop within the expanded FOV, by jointly considering the aesthetics of the cropped view, and the local image quality of the extrapolated image content. Experimental results show that our method can generate more visually pleasing image composition in cases that are difficult for previous image cropping tools due to the border constraint, and can also automatically degrade to an inward method when high quality image extrapolation is infeasible. Feng-Heng Li, Hao-Zhi Huang 0001, Yong Zhang 0034, Shao-Ping Lu, Jue Wang 0001 |
ACM Trans. Graph. | 3 |
| 2020 | Adaptive Activation Network and Functional Regularization for Efficient and Flexible Deep Multi-Task LearningabstractMulti-task learning (MTL) is a common paradigm that seeks to improve the generalization performance of task learning by training related tasks simultaneously. However, it is still a challenging problem to search the flexible and accurate architecture that can be shared among multiple tasks. In this paper, we propose a novel deep learning model called Task Adaptive Activation Network (TAAN) that can automatically learn the optimal network architecture for MTL. The main principle of TAAN is to derive flexible activation functions for different tasks from the data with other parameters of the network fully shared. We further propose two functional regularization methods that improve the MTL performance of TAAN. The improved performance of both TAAN and the regularization methods is demonstrated by comprehensive experiments. Yingru Liu, Dongliang Xie, Xin Wang 0001, Li Shen 0008, Hao-Zhi Huang 0001, Niranjan Balasubramanian |
AAAI | 6 |
| 2020 | Temporally Coherent Video Harmonization Using Adversarial NetworksabstractCompositing is one of the most important editing operations for images and videos. The process of improving the realism of composite results is often called harmonization. Previous approaches for harmonization mainly focus on images. In this paper, we take one step further to attack the problem of video harmonization. Specifically, we train a convolutional neural network in an adversarial way, exploiting a pixel-wise disharmony discriminator to achieve more realistic harmonized results and introducing a temporal loss to increase temporal consistency between consecutive harmonized frames. Thanks to the pixel-wise disharmony discriminator, we are also able to relieve the need of input foreground masks. Since existing video datasets which have ground-truth foreground masks and optical flows are not sufficiently large, we propose a simple yet efficient method to build up a synthetic dataset supporting supervised training of the proposed adversarial network. The experiments show that training on our synthetic dataset generalizes well to the real-world composite dataset. In addition, our method successfully incorporates temporal consistency during training and achieves more harmonious visual results than previous methods. Hao-Zhi Huang 0001, Sen-Zhe Xu 0001, Junxiong Cai, Wei Liu 0005, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Pose2Seg: Detection Free Human Instance SegmentationabstractThe standard approach to image instance segmentation is to perform the object detection first, and then segment the object from the detection bounding-box. More recently, deep learning methods like Mask R-CNN perform them jointly. However, little research takes into account the uniqueness of the "human" category, which can be well defined by the pose skeleton. Moreover, the human pose skeleton can be used to better distinguish instances with heavy occlusion than using bounding-boxes. In this paper, we present a brand new pose-based instance segmentation framework for humans which separates instances based on human pose, rather than proposal region detection. We demonstrate that our pose-based framework can achieve better accuracy than the state-of-art detection-based approach on the human instance segmentation problem, and can moreover better handle occlusion. Furthermore, there are few public datasets containing many heavily occluded humans along with comprehensive annotations, which makes this a challenging problem seldom noticed by researchers. Therefore, in this paper we introduce a new benchmark "Occluded Human (OCHuman)", which focuses on occluded humans with comprehensive annotations including bounding-box, human pose and instance masks. This dataset contains 8110 detailed annotated human instances within 4731 images. With an average 0.67 MaxIoU for each person, OCHuman is the most complex and challenging dataset related to human instance segmentation. Through this dataset, we want to emphasize occlusion as a challenging problem for researchers to study. Song-Hai Zhang, Ruilong Li, Paul L. Rosin, Zixi Cai, Dingcheng Yang, Hao-Zhi Huang 0001, Shi-Min Hu 0001 |
CVPR | 8 |
| 2018 | Neural Stereoscopic Image Style Transfer
Xinyu Gong, Hao-Zhi Huang 0001, Lin Ma 0002, Fumin Shen, Wei Liu 0005, Tong Zhang 0001 |
ECCV (5) | 2 |
| 2018 | Unsupervised Image-to-Image Translation with Stacked Cycle-Consistent Adversarial Networks
Minjun Li, Hao-Zhi Huang 0001, Lin Ma 0002, Wei Liu 0005, Tong Zhang 0001, Yu-Gang Jiang 0001 |
ECCV (9) | 2 |
| 2017 | Real-Time Neural Style Transfer for VideosabstractRecent research endeavors have shown the potential of using feed-forward convolutional neural networks to accomplish fast style transfer for images. In this work, we take one step further to explore the possibility of exploiting a feed-forward network to perform style transfer for videos and simultaneously maintain temporal consistency among stylized video frames. Our feed-forward network is trained by enforcing the outputs of consecutive frames to be both well stylized and temporally consistent. More specifically, a hybrid loss is proposed to capitalize on the content information of input frames, the style information of a given style image, and the temporal information of consecutive frames. To calculate the temporal loss during the training stage, a novel two-frame synergic training mechanism is proposed. Compared with directly applying an existing image style transfer method to videos, our proposed method employs the trained network to yield temporally consistent stylized videos which are much more visually pleasant. In contrast to the prior video style transfer method which relies on time-consuming optimization on the fly, our method runs in real time while generating competitive visual results. Hao-Zhi Huang 0001, Hao Wang 0050, Wenhan Luo, Lin Ma 0002, Zhifeng Li 0001, Wei Liu 0005 |
CVPR | 1 |
| 2017 | Practical automatic background substitution for live videoabstractIn this paper we present a novel automatic background substitution approach for live video. The objective of background substitution is to extract the foreground from the input video and then combine it with a new background. In this paper, we use a color line model to improve the Gaussian mixture model in the background cut method to obtain a binary foreground segmentation result that is less sensitive to brightness differences. Based on the high quality binary segmentation results, we can automatically create a reliable trimap for alpha matting to refine the segmentation boundary. To make the composition result more realistic, an automatic foreground color adjustment step is added to make the foreground look consistent with the new background. Compared to previous approaches, our method can produce higher quality binary segmentation results, and to the best of our knowledge, this is the first time such an automatic and integrated background substitution system has been proposed which can run in real time, which makes it practical for everyday applications. Hao-Zhi Huang 0001, Xiaonan Fang 0001, Yufei Ye 0001, Song-Hai Zhang, Paul L. Rosin |
Comput. Vis. Media | 1 |
| 2016 | Efficient, Edge-Aware, Combined Color Quantization and DitheringabstractIn this paper, we present a novel algorithm to simultaneously accomplish color quantization and dithering of images. This is achieved by minimizing a perception-based cost function, which considers pixel-wise differences between filtered versions of the quantized image and the input image. We use edge aware filters in defining the cost function to avoid mixing colors on the opposite sides of an edge. The importance of each pixel is weighted according to its saliency. To rapidly minimize the cost function, we use a modified multi-scale iterative conditional mode (ICM) algorithm, which updates one pixel a time while keeping other pixels unchanged. As ICM is a local method, careful initialization is required to prevent termination at a local minimum far from the global one. To address this problem, we initialize ICM with a palette generated by a modified median-cut method. Compared with previous approaches, our method can produce high-quality results with a fewer visual artifacts but also requires significantly less computational effort. Hao-Zhi Huang 0001, Kun Xu 0003, Ralph R. Martin, Fei-Yue Huang, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Faithful Completion of Images of Scenic Landmarks Using Internet ImagesabstractPrevious works on image completion typically aim to produce visually plausible results rather than factually correct ones. In this paper, we propose an approach to faithfully complete the missing regions of an image. We assume that the input image is taken at a well-known landmark, so similar images taken at the same location can be easily found on the Internet. We first download thousands of images from the Internet using a text label provided by the user. Next, we apply two-step filtering to reduce them to a small set of candidate images for use as source images for completion. For each candidate image, a co-matching algorithm is used to find correspondences of both points and lines between the candidate image and the input image. These are used to find an optimal warp relating the two images. A completion result is obtained by blending the warped candidate image into the missing region of the input image. The completion results are ranked according to combination score, which considers both warping and blending energy, and the highest ranked ones are shown to the user. Experiments and results demonstrate that our method can faithfully complete images. Zhe Zhu, Hao-Zhi Huang 0001, Kun Xu 0003, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Active Exploration of Large 3D Model RepositoriesabstractWith broader availability of large-scale 3D model repositories, the need for efficient and effective exploration becomes more and more urgent. Existing model retrieval techniques do not scale well with the size of the database since often a large number of very similar objects are returned for a query, and the possibilities to refine the search are quite limited. We propose an interactive approach where the user feeds an active learning procedure by labeling either entire models or parts of them as "like" or "dislike" such that the system can automatically update an active set of recommended models. To provide an intuitive user interface, candidate models are presented based on their estimated relevance for the current query. From the methodological point of view, our main contribution is to exploit not only the similarity between a query and the database models but also the similarities among the database models themselves. We achieve this by an offline pre-processing stage, where global and local shape descriptors are computed for each model and a sparse distance metric is derived that can be evaluated efficiently even for very large databases. We demonstrate the effectiveness of our method by interactively exploring a repository containing over 100 K models. Lin Gao 0004, Yan-Pei Cao 0001, Yukun Lai, Hao-Zhi Huang 0001, Leif Kobbelt, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2014 | Learning Natural Colors for Image RecoloringabstractAbstract We present a data‐driven method for automatically recoloring a photo to enhance its appearance or change a viewer's emotional response to it. A compact representation called a RegionNet summarizes color and geometric features of image regions, and geometric relationships between them. Correlations between color property distributions and geometric features of regions are learned from a database of well‐colored photos. A probabilistic factor graph model is used to summarize distributions of color properties and generate an overall probability distribution for color suggestions. Given a new input image, we can generate multiple recolored results which unlike previous automatic results, are both natural and artistic, and compatible with their spatial arrangements. Hao-Zhi Huang 0001, Song-Hai Zhang, Ralph R. Martin, Shi-Min Hu 0001 |
Comput. Graph. Forum | 1 |