Hengyuan Zhao

dblp:260/3042 · also Henry Hengyuan Zhao · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-8047-4465ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
abstract
Retrieval Augmented Generation enhances the response accuracy of Large Language Models (LLMs) by integrating retrieval and generation modules with external knowledge, demonstrating particular strength in real-time queries and Visual Question Answering tasks. However, the effectiveness of RAG is frequently hindered by the precision of the retriever: many retrieved samples fed into the generation phase are irrelevant or misleading, posing a critical bottleneck to LLMs’ performance. To address this challenge, we introduce \textbf{VaccineRAG}, a novel Chain-of-Thought-based retrieval-augmented generation dataset. On one hand, VaccineRAG employs a benchmark to evaluate models using data with varying positive/negative sample ratios, systematically exposing inherent weaknesses in current LLMs. On the other hand, it enhances models’ sample-discrimination capabilities by prompting LLMs to generate explicit Chain-of-Thought (CoT) analysis for each sample before producing final answers. Furthermore, to enhance the model’s ability to learn long-sequence complex CoT content, we propose \textbf{Partial-GRPO}. By modeling the outputs of LLMs as multiple components rather than a single whole, our model can make more informed preference selections for complex sequences, thereby enhancing its capacity to learn complex CoT. Comprehensive evaluations and ablation studies on VaccineRAG validate the effectiveness of the proposed scheme.
Qixin Sun, Ziqin Wang, Hengyuan Zhao, Kaiyou Song, Si Liu 0001, Xiaolin Hu 0001, Qingpei Guo, Linjiang Huang
AAAI3
2026 From Charts to Code: A Hierarchical Benchmark for Multimodal Models
abstract
Jiahao Tang, Henry Hengyuan Zhao, Lijian Wu, Zijian Zhang, Yifei Tao, Dongxing Mao, Yang Wan, Jingru Tan, Min Zeng, Min Li, Alex Jinpeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiahao Tang, Hengyuan Zhao, Lijian Wu, Yifei Tao, Dongxing Mao, Yang Wan, Jingru Tan, Min Zeng 0004, Min Li 0007, Alex Jinpeng Wang
ACL (1)2
2024 GENIXER: Empowering Multimodal Large Language Model as a Powerful Data Generator
Hengyuan Zhao, Pan Zhou 0002, Zheng Shou 0001
ECCV (23)1
2024 Reference Prompted Model Adaptation for Referring Camouflaged Object Detection
abstract
The goal of referring camouflaged object detection is to identify and segment the specified object hidden in the surroundings given text or images as references. The previous method still faces limitations in learning discriminative object features and comprehensively exploiting reference information due to coarse reference-image fusion upon disunified network components. In this paper, we propose a novel Reference Prompted Model Adaptation (RPMA) pipeline that employs rich and fine-grained semantic knowledge in a generic segmentation network to enhance the Ref-COD model’s capability. Within RPMA, we design a Cross Reference Adapter (CRA) to integrate reference information into the generic segmentation network to prompt reference-relevant camouflaged image features, and also devise a Reference-guided Dynamic Convolution (RDC) for foreground-background segmentation via reference-generated kernels. Extensive experiments on the Ref-COD benchmark show that our method achieves new state-of-the-art performance.
Xuewei Liu, Shaofei Huang 0001, Ruipu Wu, Hengyuan Zhao, Xiaoming Wei, Jizhong Han, Si Liu 0001
ICME4
2024 LOVA3: Learning to Visual Question Answering, Asking and Assessment
abstract
Question answering, asking, and assessment are three innate human traits crucial for understanding the world and acquiring knowledge. By enhancing these capabilities, humans can more effectively utilize data, leading to better comprehension and learning outcomes. However, current Multimodal Large Language Models (MLLMs) primarily focus on question answering, often neglecting the full potential of questioning and assessment skills. In this study, we introduce LOVA3, an innovative framework named ``Learning tO Visual Question Answering, Asking and Assessment,'' designed to equip MLLMs with these additional capabilities. Our approach involves the creation of two supplementary training tasks GenQA and EvalQA, aiming at fostering the skills of asking and assessing questions in the context of images. To develop the questioning ability, we compile a comprehensive set of multimodal foundational tasks. For assessment, we introduce a new benchmark called EvalQABench, comprising 64,000 training samples (split evenly between positive and negative samples) and 5,000 testing samples. We posit that enhancing MLLMs with the capabilities to answer, ask, and assess questions will enhance their multimodal comprehension, ultimately improving overall performance. To validate this hypothesis, we train MLLMs using the LOVA3 framework and evaluate them on a range of multimodal datasets and benchmarks. Our results demonstrate consistent performance gains, underscoring the critical role of these additional tasks in fostering comprehensive intelligence in MLLMs.
Hengyuan Zhao, Pan Zhou 0002, Difei Gao, Zechen Bai, Zheng Shou 0001
NeurIPS1
2024 Temporally consistent video colorization with deep feature propagation and self-regularization learning
abstract
Video colorization is a challenging and highly ill-posed problem. Although recent years have witnessed remarkable progress in single image colorization, there is relatively less research effort on video colorization, and existing methods always suffer from severe flickering artifacts (temporal inconsistency) or unsatisfactory colorization. We address this problem from a new perspective, by jointly considering colorization and temporal consistency in a unified framework. Specifically, we propose a novel temporally consistent video colorization (TCVC) framework. TCVC effectively propagates frame-level deep features in a bidirectional way to enhance the temporal consistency of colorization. Furthermore, TCVC introduces a self-regularization learning (SRL) scheme to minimize the differences in predictions obtained using different time steps. SRL does not require any ground-truth color videos for training and can further improve temporal consistency. Experiments demonstrate that our method can not only provide visually pleasing colorized video, but also with clearly better temporal consistency than state-of-the-art methods. A video demo is provided at https://www.youtube.com/watch?v=c7dczMs-olE , while code is available at https://github.com/lyh-18/TCVC-Temporally-Consistent-Video-Colorization .
Yihao Liu 0001, Hengyuan Zhao, Kelvin C. K. Chan, Xintao Wang 0002, Chen Change Loy, Yu Qiao 0001, Chao Dong 0005
Comput. Vis. Media2
2024 SCT: A Simple Baseline for Parameter-Efficient Fine-Tuning via Salient Channels
Hengyuan Zhao, Pichao Wang, Hao Luo 0004, Fan Wang 0019, Zheng Shou 0001
Int. J. Comput. Vis.1
2023 Evaluating the Generalization Ability of Super-Resolution Networks
abstract
Performance and generalization ability are two important aspects to evaluate the deep learning models. However, research on the generalization ability of Super-Resolution (SR) networks is currently absent. Assessing the generalization ability of deep models not only helps us to understand their intrinsic mechanisms, but also allows us to quantitatively measure their applicability boundaries, which is important for unrestricted real-world applications. To this end, we make the first attempt to propose a Generalization Assessment Index for SR networks, namely SRGA. SRGA exploits the statistical characteristics of the internal features of deep networks to measure the generalization ability. Specially, it is a non-parametric and non-learning metric. To better validate our method, we collect a patch-based image evaluation set (PIES) that includes both synthetic and real-world images, covering a wide range of degradations. With SRGA and PIES dataset, we benchmark existing SR models on the generalization ability. This work provides insights and tools for future research on model generalization in low-level vision.
Yihao Liu 0001, Hengyuan Zhao, Jinjin Gu, Yu Qiao 0001, Chao Dong 0005
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Very Lightweight Photo Retouching Network With Conditional Sequential Modulation
abstract
Photo retouching aims at improving the aesthetic visual quality of images that suffer from photographic defects, especially for poor contrast, over/under exposure, and inharmonious saturation. In practice, photo retouching can be accomplished by a series of image processing operations. As most commonly-used retouching operations are pixel-independent, i.e., the manipulation on one pixel is uncorrelated with its neighboring pixels, we can take advantage of this property and design a specialized algorithm for efficient global photo retouching. We analyze these global operations and find that they can be mathematically formulated by a Multi-Layer Perceptron (MLP). Based on this observation, we propose an extremely lightweight framework – Conditional Sequential Retouching Network (CSRNet). Benefiting from the utilization of$1\times 1$convolution, CSRNet only contains less than 37 K trainable parameters, which are orders of magnitude smaller than existing learning-based methods. Experiments show that our method achieves state-of-the-art performance on the benchmark MIT-Adobe FiveK dataset quantitively and qualitatively. In addition to achieve global photo retouching, the proposed framework can be easily extended to learn local enhancement effects. The extended model, namely CSRNet-L, also achieves competitive results in various local enhancement tasks.
Yihao Liu 0001, Jingwen He, Xiangyu Chen 0006, Zhengwen Zhang, Hengyuan Zhao, Chao Dong 0005, Yu Qiao 0001
IEEE Trans. Multim.5
2021 ClassSR: A General Framework to Accelerate Super-Resolution Networks by Data Characteristic
abstract
We aim at accelerating super-resolution (SR) networks on large images (2K-8K). The large images are usually decomposed into small sub-images in practical usages. Based on this processing, we found that different image regions have different restoration difficulties and can be processed by networks with different capacities. Intuitively, smooth areas are easier to super-solve than complex textures. To utilize this property, we can adopt appropriate SR networks to process different sub-images after the decomposition. On this basis, we propose a new solution pipeline – ClassSR that combines classification and SR in a unified framework. In particular, it first uses a Class-Module to classify the subimages into different classes according to restoration difficulties, then applies an SR-Module to perform SR for different classes. The Class-Module is a conventional classification network, while the SR-Module is a network container that consists of the to-be-accelerated SR network and its simplified versions. We further introduce a new classification method with two losses – Class-Loss and Average-Loss to produce the classification results. After joint training, a majority of sub-images will pass through smaller networks, thus the computational cost can be significantly reduced. Experiments show that our ClassSR can help most existing methods (e.g., FSRCNN, CARN, SRResNet, RCAN) save up to 50% FLOPs on DIV8K datasets. This general framework can also be applied in other low-level vision tasks.
Xiangtao Kong, Hengyuan Zhao, Yu Qiao 0001, Chao Dong 0005
CVPR2