Zhewei Huang

dblp:211/8001 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2026 PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
abstract
Jingcheng Hu, Yinmin Zhang, Shijie Shang, Xiaobo Yang, Yue Peng, Zhewei Huang, Hebin Zhou, Xin Wu, Jie Cheng, Fanqi Wan, Xiangwen Kong, Chengyuan Yao, Kaiwen Yan, Ailin Huang, Hongyu Zhou, Qi Han, Zheng Ge, Xiangyu Zhang, Heung-Yeung Shum. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jingcheng Hu, Yinmin Zhang, Shijie Shang, Zhewei Huang, Hebin Zhou, Fanqi Wan, Xiangwen Kong, Chengyuan Yao, Kaiwen Yan, Ailin Huang, Zheng Ge, Xiangyu Zhang 0005, Harry Shum
ACL (1)6
2025 Learning temporal-aware representation for controllable interventional radiology imaging
Wei Si, Zhaolin Zheng, Zhewei Huang, Ximing Xu 0002, Ruijue Wang, Ji-Gang Bao, Xiantong Zhen, Jun Xu 0019
Comput. Vis. Image Underst.3
2025 Advancing video self-supervised learning via image foundation models
Jingwei Wu, Zhewei Huang, Chang Liu 0047
Pattern Recognit. Lett.2
2024 Scale-Adaptive Feature Aggregation for Efficient Space-Time Video Super-Resolution
abstract
The Space-Time Video Super-Resolution (STVSR) task aims to enhance the visual quality of videos, by simultaneously performing video frame interpolation (VFI) and video super-resolution (VSR). However, facing the challenge of the additional temporal dimension and scale inconsistency, most existing STVSR methods are complex and inflexible in dynamically modeling different motion amplitudes. In this work, we find that choosing an appropriate processing scale achieves remarkable benefits in flow-based feature propagation. We propose a novel Scale-Adaptive Feature Aggregation (SAFA) network that adaptively selects sub-networks with different processing scales for individual samples. Experiments on four public STVSR benchmarks demonstrate that SAFA achieves state-of-the-art performance. Our SAFA network outperforms recent state-of-the-art methods such as TMNet [83] and VideoINR [10] by an average improvement of over 0.5dB on PSNR, while requiring less than half the number of parameters and only 1/3 computational costs.
Zhewei Huang, Ailin Huang, Xiaotao Hu, Jun Xu 0019, Shuchang Zhou 0001
WACV1
2023 A Dynamic Multi-Scale Voxel Flow Network for Video Prediction
abstract
The performance of video prediction has been greatly boosted by advanced deep neural networks. However, most of the current methods suffer from large model sizes and require extra inputs, e.g., semantic/depth maps, for promising performance. For efficiency consideration, in this paper, we propose a Dynamic Multi-scale Voxel Flow Network (DMVFN) to achieve better video prediction performance at lower computational costs with only RGB images, than previous methods. The core of our DMVFN is a differentiable routing module that can effectively perceive the motion scales of video frames. Once trained, our DMVFN selects adaptive sub-networks for different inputs at the inference stage. Experiments on several benchmarks demonstrate that our DMVFN is an order of magnitude faster than Deep Voxel Flow [35] and surpasses the state-of-the-art iterative-based OPT [63] on generated image quality.
Xiaotao Hu, Zhewei Huang, Ailin Huang, Jun Xu 0019, Shuchang Zhou 0001
CVPR2
2023 Collaborative Neural Rendering Using Anime Character Sheets
abstract
Drawing images of characters with desired poses is an essential but laborious task in anime production. Assisting artists to create is a research hotspot in recent years. In this paper, we present the Collaborative Neural Rendering (CoNR) method, which creates new images for specified poses from a few reference images (AKA Character Sheets). In general, the diverse hairstyles and garments of anime characters defies the employment of universal body models like SMPL, which fits in most nude human shapes. To overcome this, CoNR uses a compact and easy-to-obtain landmark encoding to avoid creating a unified UV mapping in the pipeline. In addition, the performance of CoNR can be significantly improved when referring to multiple reference images, thanks to feature space cross-view warping in a carefully designed neural network. Moreover, we have collected a character sheet dataset containing over 700,000 hand-drawn and synthesized images of diverse poses to facilitate research in this area. The code and dataset is available at https://github.com/megvii-research/IJCAI2023-CoNR.
Zuzeng Lin, Ailin Huang, Zhewei Huang
IJCAI3
2022 Real-Time Intermediate Flow Estimation for Video Frame Interpolation
Zhewei Huang, Wen Heng, Boxin Shi, Shuchang Zhou 0001
ECCV (14)1
2022 Perceptual Conversational Head Generation with Regularized Driver and Enhanced Renderer
abstract
This paper reports our solution for ACM Multimedia ViCo 2022 Conversational Head Generation Challenge, which aims to generate vivid face-to-face conversation videos based on audio and reference images. Our solution focuses on training a generalized audio-to-head driver using regularization and assembling a high visual quality renderer. We carefully tweak the audio-to-behavior model and post-process the generated video using our foreground-background fusion module. We get first place in the listening head generation track and second place in the talking head generation track in the official leaderboard. Our code is available at https://github.com/megvii-research/MM2022-ViCoPerceptualHeadGeneration.
Ailin Huang, Zhewei Huang, Shuchang Zhou 0001
ACM Multimedia2
2019 Learning to Paint With Model-Based Deep Reinforcement Learning
abstract
We show how to teach machines to paint like human painters, who can use a small number of strokes to create fantastic paintings. By employing a neural renderer in model-based Deep Reinforcement Learning (DRL), our agents learn to determine the position and color of each stroke and make long-term plans to decompose texture-rich images into strokes. Experiments demonstrate that excellent visual effects can be achieved using hundreds of strokes. The training process does not require the experience of human painters or stroke tracking data. The code is available at https://github.com/hzwer/ICCV2019-LearningToPaint.
Zhewei Huang, Shuchang Zhou 0001, Wen Heng
ICCV1