Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jinhua Hao

dblp:372/0042 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
6 papers
Image and video processing · 76% Multimedia systems and quality of experience · 18% Image and video coding · 6%
Artificial intelligence
4 papers
Generative modeling · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › super-resolution
image super-resolution
1.622025
Visual Autoregressive Modeling for Image Super-Resolution · ICML 2025
XPSR: Cross-Modal Priors for Diffusion-Based Image Super-Resolution · ECCV (11) 2024
Machine learning › Generative modeling › autoregressive model
autoregressive image generation
0.912025
Visual Autoregressive Modeling for Image Super-Resolution · ICML 2025
Machine learning › Generative modeling
autoregressive model
0.912025
Visual Autoregressive Modeling for Image Super-Resolution · ICML 2025
Image and video processing › super-resolution › image super-resolution
generative image super-resolution
0.912025
Visual Autoregressive Modeling for Image Super-Resolution · ICML 2025
Image and video processing › image resampling
image rescaling
0.912025
Plug-and-Play Tri-Branch Invertible Block for Image Rescaling · AAAI 2025
Image and video processing
image restoration
0.912025
Plug-and-Play Tri-Branch Invertible Block for Image Rescaling · AAAI 2025
Image and video processing › video enhancement
compressed video quality enhancement
0.812024
CPGA: Coding Priors-Guided Aggregation Network for Compressed Video Quality Enhancement · CVPR 2024
Image and video processing › super-resolution › image super-resolution › generative image super-resolution
diffusion-based super-resolution
0.812024
XPSR: Cross-Modal Priors for Diffusion-Based Image Super-Resolution · ECCV (11) 2024
Image and video processing › image restoration › compression artifact removal
JPEG artifact removal
0.812024
OAPT: Offset-Aware Partition Transformer for Double JPEG Artifacts Removal · ECCV (22) 2024
Multimedia systems and quality of experience › video quality assessment
no-reference video quality assessment
0.812024
PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the Wild · CVPR 2024
Multimedia systems and quality of experience
video quality assessment
0.812024
PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the Wild · CVPR 2024
Machine learning › Generative modeling
diffusion model
0.522025
Visual Autoregressive Modeling for Image Super-Resolution · ICML 2025
XPSR: Cross-Modal Priors for Diffusion-Based Image Super-Resolution · ECCV (11) 2024
Machine learning › Generative modeling › diffusion model › diffusion model inference
diffusion-based refinement
0.312025
Visual Autoregressive Modeling for Image Super-Resolution · ICML 2025
Image and video coding
image compression
0.312025
Plug-and-Play Tri-Branch Invertible Block for Image Rescaling · AAAI 2025
Image and video coding
lossy compression
0.312025
Plug-and-Play Tri-Branch Invertible Block for Image Rescaling · AAAI 2025
Machine learning › Generative modeling › diffusion model
image restoration
0.212024
OAPT: Offset-Aware Partition Transformer for Double JPEG Artifacts Removal · ECCV (22) 2024

Methods — techniques the papers use, named apart from their topics

rotary positional encoding · 1.7next-scale prediction · 1.7classifier-free guidance · 1.7autoregressive modeling · 1.7pseudo clustering · 1.5pretrained model selection · 1.5contrastive learning · 1.5luminance-chrominance decomposition · 0.9invertible neural network · 0.9transformer · 0.8inter-frame temporal aggregation · 0.8diffusion model · 0.8cross-modal priors · 0.8
YearPublicationVenuePosition
2025 Plug-and-Play Tri-Branch Invertible Block for Image Rescaling
abstract
High-resolution (HR) images are commonly downscaled to low-resolution (LR) to reduce bandwidth, followed by upscaling to restore their original details. Recent advancements in image rescaling algorithms have employed invertible neural networks (INNs) to create a unified framework for downscaling and upscaling, ensuring a one-to-one mapping between LR and HR images. Traditional methods, utilizing dual-branch based vanilla invertible blocks, process high-frequency and low-frequency information separately, often relying on specific distributions to model high-frequency components. However, processing the low-frequency component directly in the RGB domain introduces channel redundancy, limiting the efficiency of image reconstruction. To address these challenges, we propose a plug-and-play tri-branch invertible block (T-InvBlocks) that decomposes the low- frequency branch into luminance (Y) and chrominance (CbCr) components, reducing redundancy and enhancing feature processing. Additionally, we adopt an all-zero mapping strategy for high-frequency components during upscaling, focusing essential rescaling information within the LR image. Our T-InvBlocks can be seamlessly integrated into existing rescaling models, improving performance in both general rescaling tasks and scenarios involving lossy compression. Extensive experiments confirm that our method advances the state of the art in HR image reconstruction.
Jingwei Bao, Jinhua Hao, Ming Sun 0008, Chao Zhou 0003, Shuyuan Zhu
AAAI2
2025 Visual Autoregressive Modeling for Image Super-Resolution
abstract
Image Super-Resolution (ISR) has seen significant progress with the introduction of remarkable generative models. However, challenges such as the trade-off issues between fidelity and realism, as well as computational complexity, have also posed limitations on their application. Building upon the tremendous success of autoregressive models in the language domain, we propose VARSR, a novel visual autoregressive modeling for ISR framework with the form of next-scale prediction. To effectively integrate and preserve semantic information in low-resolution images, we propose using prefix tokens to incorporate the condition. Scale-aligned Rotary Positional Encodings are introduced to capture spatial structures and the diffusion refiner is utilized for modeling quantization residual loss to achieve pixel-level fidelity. Image-based Classifier-free Guidance is proposed to guide the generation of more realistic images. Furthermore, we collect large-scale data and design a training process to obtain robust generative priors. Quantitative and qualitative results show that VARSR is capable of generating high-fidelity and high-realism images with more efficiency than diffusion-based methods. Our codes are released at https://github.com/quyp2000/VARSR.
Yunpeng Qu, Kun Yuan 0003, Jinhua Hao, Kai Zhao 0011, Qizhi Xie, Ming Sun 0008, Chao Zhou 0003
ICML3
2024 PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the Wild
abstract
Video quality assessment (VQA) is a challenging problem due to the numerous factors that can affect the perceptual quality of a video, e.g., content attractiveness, distortion type, motion pattern, and level. However, annotating the Mean opinion score (MOS) for videos is expensive and time-consuming, which limits the scale of VQA datasets, and poses a significant obstacle for deep learning-based methods. In this paper, we propose a VQA method named PTM-VQA, which leverages PreTrained Models to transfer knowledge from models pretrained on various pre-tasks, enabling benefits for VQA from different aspects. Specifically, we extract features of videos from different pretrained models with frozen weights and integrate them to generate representation. Since these models possess var-ious fields of knowledge and are often trained with labels irrelevant to quality, we propose an Intra-Consistency and Inter-Divisibility (ICID) loss to impose constraints on features extracted by multiple pretrained models. The intra-consistency constraint ensures that features extracted by different pretrained models are in the same unified quality-aware latent space, while the inter-divisibility introduces pseudo clusters based on the annotation of samples and tries to separate features of samples from different clusters. Furthermore, with a constantly growing number of pretrained models, it is crucial to determine which models to use and how to use them. To address this problem, we propose an efficient scheme to select suitable candidates. Models with better clustering performance on VQA datasets are chosen to be our candidates. Extensive experiments demonstrate the effectiveness of the proposed method.
Kun Yuan 0003, Mading Li, Muyi Sun, Ming Sun 0008, Jiachao Gong, Jinhua Hao, Chao Zhou 0003, Yansong Tang
CVPR7
2024 CPGA: Coding Priors-Guided Aggregation Network for Compressed Video Quality Enhancement
abstract
Recently, numerous approaches have achieved notable success in compressed video quality enhancement (VQE). However, these methods usually ignore the utilization of valuable coding priors inherently embedded in compressed videos, such as motion vectors and residual frames, which carry abundant temporal and spatial information. To remedy this problem, we propose the Coding Priors-Guided Aggregation (CPGA) network to utilize temporal and spatial information from coding priors. The CPGA mainly consists of an inter-frame temporal aggregation (ITA) module and a multi-scale non-local aggregation (MNA) module. Specifically, the ITA module aggregates temporal information from consecutive frames and coding priors, while the MNA module globally captures spatial information guided by residual frames. In addition, to facilitate research in VQE task, we newly construct the Video Coding Priors (VCP) dataset, comprising 300 videos with various coding priors extracted from corresponding bitstreams. It remedies the shortage of previous datasets on the lack of coding information. Experimental results demonstrate the superiority of our method compared to existing state-of-the-art methods. The code and dataset will be released at https://github.com/VQE-CPGA/CPGA.
Jinhua Hao, Yukang Ding, Yu Liu 0091, Qiao Mo, Ming Sun 0008, Chao Zhou 0003, Shuyuan Zhu
CVPR2
2024 OAPT: Offset-Aware Partition Transformer for Double JPEG Artifacts Removal
Qiao Mo, Yukang Ding, Jinhua Hao, Ming Sun 0008, Chao Zhou 0003, Feiyu Chen 0001, Shuyuan Zhu
ECCV (22)3
2024 XPSR: Cross-Modal Priors for Diffusion-Based Image Super-Resolution
Yunpeng Qu, Kun Yuan 0003, Kai Zhao 0011, Qizhi Xie, Jinhua Hao, Ming Sun 0008, Chao Zhou 0003
ECCV (11)5