Zheng Hui

dblp:211/5794 · DBLP profile ↗
← Back
27ranked-venue papers
10as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Value of Information: A Framework for Human-Agent Communication
abstract
Yijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang, Ivan Vulić, Andreea Bobu, Nigel Collier. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang, Ivan Vulic, Andreea Bobu, Nigel Collier
ACL (1)3
2026 Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
abstract
Zheng Hui, Yijiang River Dong, Sanhanat Sivapiromrat, Ehsan Shareghi, Nigel Collier. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zheng Hui, Yijiang River Dong, Sanhanat Sivapiromrat, Ehsan Shareghi, Nigel Collier
ACL (1)1
2025 VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models
abstract
Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V) still lag far behind in frame quality and text alignment, owing to insufficient quality and quantity of training videos. In this paper, we introduce VideoElevator, a training-free and plug-and-play method, which elevates the performance of T2V using superior capabilities of T2I. Different from conventional T2V sampling (i.e., temporal and spatial modeling), VideoElevator explicitly decomposes each sampling step into temporal motion refining and spatial quality elevating. Specifically, temporal motion refining uses encapsulated T2V to enhance temporal consistency, followed by inverting to the noise distribution required by T2I. Then, spatial quality elevating harnesses inflated T2I to directly predict less noisy latent, adding more photo-realistic details. We have conducted experiments in extensive prompts under the combination of various T2V and T2I. The results show that VideoElevator not only improves the performance of T2V baselines with foundational T2I, but also facilitates stylistic video synthesis with personalized T2I. Please watch all videos in supplementary materials for better view.
Yabo Zhang, Yuxiang Wei 0001, Xianhui Lin, Zheng Hui, Peiran Ren, Xuansong Xie, Wangmeng Zuo
AAAI4
2025 Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation
abstract
Multimodal large language models have experienced rapid growth, and numerous different models have emerged. The interpretability of LVLMs remains an under-explored area. Especially when faced with more complex tasks such as chain-of-thought reasoning, its internal mechanisms still resemble a black box that is difficult to decipher. By studying the interaction and information flow between images and text, we noticed that in models such as LLaVA1.5, image tokens that are semantically related to text are more likely to have information flow convergence in the LLM decoding layer, and these image tokens receive higher attention scores. However, those image tokens that are less relevant to the text do not have information flow convergence, and they only get very small attention scores. To efficiently utilize the image information, we propose a new image token reduction method, Simignore, which aims to improve the complex reasoning ability of LVLMs by computing the similarity between image and text embeddings and ignoring image tokens that are irrelevant and unimportant to the text. Through extensive experiments, we demonstrate the effectiveness of our method for complex reasoning tasks.
Fanshuo Zeng, Yihao Quan, Zheng Hui, Jiawei Yao
AAAI4
2025 PropaInsight: Toward Deeper Understanding of Propaganda in Terms of Techniques, Appeals, and Intent
abstract
Propaganda plays a critical role in shaping public opinion and fueling disinformation. While existing research primarily focuses on identifying propaganda techniques, it lacks the ability to capture the broader motives and the impacts of such content. To address these challenges, we introduce PropaInsight, a conceptual framework grounded in foundational social science research, which systematically dissects propaganda into techniques, arousal appeals, and underlying intent. PropaInsight offers a more granular understanding of how propaganda operates across different contexts. Additionally, we present PropaGaze, a novel dataset that combines human-annotated data with high-quality synthetic data generated through a meticulously designed pipeline. Our experiments show that off-the-shelf LLMs struggle with propaganda analysis, but PropaGaze significantly improves performance. Fine-tuned Llama-7B-Chat achieves 203.4% higher text span IoU in technique identification and 66.2% higher BertScore in appeal analysis compared to 1-shot GPT-4-Turbo. Moreover, PropaGaze complements limited human-annotated data in data-sparse and cross-domain scenarios, demonstrating its potential for comprehensive and generalizable propaganda analysis.
Jiateng Liu, Lin Ai, Zizhou Liu, Payam Karisani, Zheng Hui, Yi R. Fung 0001, Preslav Nakov, Julia Hirschberg, Heng Ji 0001
COLING5
2025 QuantCache: Adaptive Importance-Guided Quantization with Hierarchical Latent and Layer Caching for Video Generation
Zhiteng Li, Zheng Hui, Yulun Zhang 0001, Linghe Kong, Xiaokang Yang 0001
ICCV3
2025 Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
abstract
Large language models (LLMs) show potential as computer agents, enhancing productivity and software accessibility in multi-modal tasks. However, measuring agent performance in sufficiently realistic and complex environments becomes increasingly challenging as: (i) most benchmarks are limited to specific modalities/domains (e.g., text-only, web navigation, Q&A) and (ii) full benchmark evaluations are slow (on order of magnitude of multiple hours/days) given the multi-step sequential nature of tasks. To address these challenges, we introduce Windows Agent Arena: a general environment focusing exclusively on the Windows operating system (OS) where agents can operate freely within a real OS to use the same applications and tools available to human users when performing tasks. We create 150+ diverse tasks across representative domains that require agentic abilities in planning, screen understanding, and tool usage. Our benchmark is scalable and can be seamlessly parallelized for a full benchmark evaluation in as little as $20$ minutes. Our work not only speeds up the development and evaluation cycle of multi-modal agents, but also highlights and analyzes existing shortfalls in the agentic abilities of several multimodal LLMs as agents within the Windows computing environment---with the best achieving only a 19.5\% success rate compared to a human success rate of 74.5\%.
Rogerio Bonatti, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, Lawrence Jang, Zheng Hui
ICML12
2025 Language Knowledge-Assisted Representation Learning for Skeleton-Based Action Recognition
abstract
How humans understand and recognize the actions of others is a complex neuroscientific problem that involves a combination of cognitive mechanisms and neural networks. Research has shown that humans have brain areas that recognize actions that process top-down attentional information, such as the temporoparietal association area. Also, humans have brain regions dedicated to understanding the minds of others and analyzing their intentions, such as the medial prefrontal cortex of the temporal lobe. Skeleton-based action recognition creates mappings for the complex connections between the human skeleton movement patterns and behaviors. Although existing studies encoded meaningful node relationships and synthesized action representations for classification with good results, few of them considered incorporating a priori knowledge to aid potential representation learning for better performance. LA-GCN proposes a graph convolution network using large-scale language models (LLM) knowledge assistance. First, the LLM knowledge is mapped into a priori global relationship (GPR) topology and a priori category relationship (CPR) topology between nodes. The GPR guides the generation of new “bone” representations, aiming to emphasize essential node information from the data level. The CPR mapping simulates category prior knowledge in human brain regions, encoded by the PC-AC module and used to add additional supervision—forcing the model to learn class-distinguishable features. In addition, to improve information transfer efficiency in topology modeling, we propose multi-hop attention graph convolution. It aggregates each node's k-order neighbor simultaneously to speed up model convergence. LA-GCN reaches state-of-the-art on NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets.
Yan Gao 0025, Zheng Hui, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Multim.3
2024 Improving Diffusion-Based Image Restoration with Error Contraction and Error Correction
abstract
Generative diffusion prior captured from the off-the-shelf denoising diffusion generative model has recently attained significant interest. However, several attempts have been made to adopt diffusion models to noisy inverse problems either fail to achieve satisfactory results or require a few thousand iterations to achieve high-quality reconstructions. In this work, we propose a diffusion-based image restoration with error contraction and error correction (DiffECC) method. Two strategies are introduced to contract the restoration error in the posterior sampling process. First, we combine existing CNN-based approaches with diffusion models to ensure data consistency from the beginning. Second, to amplify the error contraction effects of the noise, a restart sampling algorithm is designed. In the error correction strategy, the estimation-correction idea is proposed on both the data term and the prior term. Solving them iteratively within the diffusion sampling framework leads to superior image generation results. Experimental results for image restoration tasks such as super-resolution (SR), Gaussian deblurring, and motion deblurring demonstrate that our approach can reconstruct high-quality images compared with state-of-the-art sampling-based diffusion models.
Qiqi Bao 0001, Zheng Hui, Rui Zhu 0006, Peiran Ren, Xuansong Xie, Wenming Yang
AAAI2
2024 Auto-Prox: Training-Free Vision Transformer Architecture Search via Automatic Proxy Discovery
abstract
The substantial success of Vision Transformer (ViT) in computer vision tasks is largely attributed to the architecture design. This underscores the necessity of efficient architecture search for designing better ViTs automatically. As training-based architecture search methods are computationally intensive, there’s a growing interest in training-free methods that use zero-cost proxies to score ViTs. However, existing training-free approaches require expert knowledge to manually design specific zero-cost proxies. Moreover, these zero-cost proxies exhibit limitations to generalize across diverse domains. In this paper, we introduce Auto-Prox, an automatic proxy discovery framework, to address the problem. First, we build the ViT-Bench-101, which involves different ViT candidates and their actual performance on multiple datasets. Utilizing ViT-Bench-101, we can evaluate zero-cost proxies based on their score-accuracy correlation. Then, we represent zero-cost proxies with computation graphs and organize the zero-cost proxy search space with ViT statistics and primitive operations. To discover generic zero-cost proxies, we propose a joint correlation metric to evolve and mutate different zero-cost proxy candidates. We introduce an elitism-preserve strategy for search efficiency to achieve a better trade-off between exploitation and exploration. Based on the discovered zero-cost proxy, we conduct a ViT architecture search in a training-free manner. Extensive experiments demonstrate that our method generalizes well to different datasets and achieves state-of-the-art results both in ranking correlation and final accuracy. Codes can be found at https://github.com/lilujunai/Auto-Prox-AAAI24.
Zimian Wei, Peijie Dong, Zheng Hui, Anggeng Li, Lujun Li 0001, Menglong Lu, Hengyue Pan, Dongsheng Li 0001
AAAI3
2024 Enhancing Pre-Trained Generative Language Models with Question Attended Span Extraction on Machine Reading Comprehension
abstract
Machine Reading Comprehension (MRC) poses a significant challenge in the field of Natural Language Processing (NLP).While mainstream MRC methods predominantly leverage extractive strategies using encoder-only models such as BERT, generative approaches face the issue of out-of-control generation -a critical problem where answers generated are often incorrect, irrelevant, or unfaithful to the source text.To address these limitations in generative models for extractive MRC, we introduce the Question-Attended Span Extraction (QASE) module.Integrated during the finetuning phase of pre-trained generative language models (PLMs), QASE significantly enhances their performance, allowing them to surpass the extractive capabilities of advanced Large Language Models (LLMs) such as GPT-4 in few-shot settings.Notably, these gains in performance do not come with an increase in computational demands.The efficacy of the QASE module has been rigorously tested across various datasets, consistently achieving or even surpassing state-of-the-art (SOTA) results, thereby bridging the gap between generative and extractive models in extractive MRC tasks.Our code is available at this GitHub repository.
Lin Ai, Zheng Hui, Zizhou Liu, Julia Hirschberg
EMNLP2
2024 Defending Against Social Engineering Attacks in the Age of LLMs
abstract
Lin Ai, Tharindu Sandaruwan Kumarage, Amrita Bhattacharjee, Zizhou Liu, Zheng Hui, Michael S. Davinroy, James Cook, Laura Cassani, Kirill Trapeznikov, Matthias Kirchner, Arslan Basharat, Anthony Hoogs, Joshua Garland, Huan Liu, Julia Hirschberg. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Lin Ai, Tharindu Kumarage, Amrita Bhattacharjee, Zizhou Liu, Zheng Hui, Michael Davinroy, James Cook, Laura Cassani, Kirill Trapeznikov, Matthias Kirchner, Arslan Basharat, Anthony Hoogs, Joshua Garland, Huan Liu 0001, Julia Hirschberg
EMNLP5
2024 TweetIntent@Crisis: A Dataset Revealing Narratives of Both Sides in the Russia-Ukraine Crisis
abstract
This paper introduces TweetIntent@Crisis, a novel Twitter dataset centered on the Russia-Ukraine crisis. Comprising over 17K tweets from government-affiliated accounts of both nations, the dataset is meticulously annotated to identify underlying intents and detailed intent-related information. Our analysis demonstrates the dataset's capability in revealing fine-grained intents and nuanced narratives within the tweets from both parties involved in the crisis. We aim for TweetIntent@Crisis to provide the research community with a valuable tool for understanding and analyzing granular media narratives and their impact in this geopolitical conflict.
Lin Ai, Sameer Gupta, Shreya Oak, Zheng Hui, Zizhou Liu, Julia Hirschberg
ICWSM4
2024 LL-Diff: Low-Light Image Enhancement Utilizing Langevin Sampling Diffusion
abstract
In this paper, we propose a new algorithm called LL-Diff, which is innovative compared to traditional augmentation methods in that it introduces the sampling method of Langevin dynamics. This sampling approach simulates the motion of particles in complex environments and can better handle noise and details in low-light conditions. We also incorporate a causal attention mechanism to achieve causality and address the issue of confounding effects. This attention mechanism enables us to better capture local information while avoiding over-enhancement. We have conducted experiments on the LOL-V1 and LOL-V2 datasets, and the results show that LL-Diff significantly improves computational speed and several evaluation metrics, demonstrating the superiority and effectiveness of our method for low-light image enhancement tasks. The code will be released on GitHub when the paper has been accepted.
Boren Ding, Xiaofeng Zhang 0006, Zekun Yu, Zheng Hui
Int. J. Pattern Recognit. Artif. Intell.4
2022 Adaptive Modulation and Rectangular Convolutional Network for Stereo Image Super-Resolution
Xiumei Wang 0002, Tianmeng Li, Zheng Hui, Peitao Cheng
Pattern Recognit. Lett.3
2021 Learning the Non-Differentiable Optimization for Blind Super-Resolution
abstract
Previous convolutional neural network (CNN) based blind super-resolution (SR) methods usually adopt an iterative optimization way to approximate the ground-truth (GT) step-by-step. This solution always involves more computational costs to bring about time-consuming inference. At present, most blind SR algorithms are dedicated to obtaining high-fidelity results; their loss function generally employs L1 loss. To further improve the visual quality of SR results, perceptual metric, such as NIQE, is necessary to guide the network optimization. However, due to the non-differentiable property of NIQE, it cannot be as the loss function. Towards these issues, we propose an adaptive modulation network (AMNet) for multiple degradations SR, which is composed of the pivotal adaptive modulation layer (AMLayer). It is an efficient yet lightweight fusion layer between blur kernel and image features. Equipped with the blur kernel predictor, we naturally upgrade the AMNet to the blind SR model. Instead of considering iterative strategy, we make the blur kernel predictor trainable in the whole blind SR model, in which AMNet is well-trained. Also, we fit deep reinforcement learning into the blind SR model (AMNet-RL) to tackle the non-differentiable optimization problem. Specifically, the blur kernel predictor will be the actor to estimate the blur kernel from the input low-resolution (LR) image. The reward is designed by the pre-defined differentiable or non-differentiable metric. Extensive experiments show that our model can outperform state-of-the-art methods in both fidelity and perceptual metrics.
Zheng Hui, Jie Li 0001, Xiumei Wang 0002, Xinbo Gao 0001
CVPR1
2021 Progressive perception-oriented network for single image super-resolution
Zheng Hui, Jie Li 0001, Xinbo Gao 0001, Xiumei Wang 0002
Inf. Sci.1
2020 Lightweight Image Super-resolution with Local Attention Enhancement
Yunchu Yang, Xiumei Wang 0002, Xinbo Gao 0001, Zheng Hui
PRCV (1)4
2020 Lightweight image super-resolution with feature enhancement residual network
Zheng Hui, Xinbo Gao 0001, Xiumei Wang 0002
Neurocomputing1
2020 A novel high payload steganography scheme based on absolute moment block truncation coding
Zheng Hui
Multim. Tools Appl.1
2020 A Deep Neural Network Application for Improved Prediction of $\text{HbA}_{\text{1c}}$ in Type 1 Diabetes
abstract
[Formula: see text] is a primary marker of long-term average blood glucose, which is an essential measure of successful control in type 1 diabetes. Previous studies have shown that [Formula: see text] estimates can be obtained from 5-12 weeks of daily blood glucose measurements. However, these methods suffer from accuracy limitations when applied to incomplete data with missing periods of measurements. The aim of this article is to overcome these limitations improving the accuracy and robustness of [Formula: see text] prediction from time series of blood glucose. A novel data-driven [Formula: see text] prediction model based on deep learning and convolutional neural networks is presented. The model focuses on the extraction of behavioral patterns from sequences of self-monitored blood glucose readings on various temporal scales. Assuming that subjects who share behavioral patterns have also similar capabilities for diabetes control and resulting [Formula: see text], it becomes possible to infer the [Formula: see text] of subjects with incomplete data from multiple observations of similar behaviors. Trained and validated on a dataset, containing 1543 real world observation epochs from 759 subjects, the model has achieved the mean absolute error of 4.80 [Formula: see text] mmol/mol, median absolute error of 3.81 [Formula: see text] mmol/mol and [Formula: see text] of 0.71 ± 0.09 on average during the 10 fold cross validation. Automatic behavioral characterization via extraction of sequential features by the proposed convolutional neural network structure has significantly improved the accuracy of [Formula: see text] prediction compared to the existing methods.
Aleksandr Zaitcev, Mohammad R. Eissa, Zheng Hui, Tim Good, Jackie Elliott, Mohammed Benaissa
IEEE J. Biomed. Health Informatics3
2019 A Novel Robust Blind Digital Image Watermarking Scheme Against JPEG2000 Compression
Zheng Hui
ICIG (3)1
2019 Lightweight Image Super-Resolution with Information Multi-distillation Network
abstract
In recent years, single image super-resolution (SISR) methods using deep convolution neural network (CNN) have achieved impressive results. Thanks to the powerful representation capabilities of the deep networks, numerous previous ways can learn the complex non-linear mapping between low-resolution (LR) image patches and their high-resolution (HR) versions. However, excessive convolutions will limit the application of super-resolution technology in low computing power devices. Besides, super-resolution of any arbitrary scale factor is a critical issue in practical applications, which has not been well solved in the previous approaches. To address these issues, we propose a lightweight information multi-distillation network (IMDN) by constructing the cascaded information multi-distillation blocks (IMDB), which contains distillation and selective fusion parts. Specifically, the distillation module extracts hierarchical features step-by-step, and fusion module aggregates them according to the importance of candidate features, which is evaluated by the proposed contrast-aware channel attention mechanism. To process real images with any sizes, we develop an adaptive cropping strategy (ACS) to super-resolve block-wise image patches using the same well-trained model. Extensive experiments suggest that the proposed method performs favorably against the state-of-the-art SR algorithms in term of visual quality, memory footprint, and inference time. Code is available at \urlhttps://github.com/Zheng222/IMDN.
Zheng Hui, Xinbo Gao 0001, Yunchu Yang, Xiumei Wang 0002
ACM Multimedia1
2019 Dual residual attention module network for single image super resolution
Xiumei Wang 0002, Yanan Gu, Xinbo Gao 0001, Zheng Hui
Neurocomputing4
2018 Fast and Accurate Single Image Super-Resolution via Information Distillation Network
abstract
Recently, deep convolutional neural networks (CNNs) have been demonstrated remarkable progress on single image super-resolution. However, as the depth and width of the networks increase, CNN-based super-resolution methods have been faced with the challenges of computational complexity and memory consumption in practice. In order to solve the above questions, we propose a deep but compact convolutional network to directly reconstruct the high resolution image from the original low resolution image. In general, the proposed model consists of three parts, which are feature extraction block, stacked information distillation blocks and reconstruction block respectively. By combining an enhancement unit with a compression unit into a distillation block, the local long and short-path features can be effectively extracted. Specifically, the proposed enhancement unit mixes together two different types of features and the compression unit distills more useful information for the sequential blocks. In addition, the proposed network has the advantage of fast execution due to the comparatively few numbers of filters per layer and the use of group convolution. Experimental results demonstrate that the proposed method is superior to the state-of-the-art methods, especially in terms of time performance. Code is available at https://github.com/Zheng222/IDN-Caffe.
Zheng Hui, Xiumei Wang 0002, Xinbo Gao 0001
CVPR1
2018 Two-Stage Convolutional Network for Image Super-Resolution
abstract
Deep convolutional neural networks (DCNN) have recently advanced the state-of-the-art on the issue of single image super-resolution (SR). In this work, we propose a two-stage convolutional network (TSCN) to estimate the desired high-resolution (HR) image from the corresponding low-resolution (LR) image. Specifically, we propose the multi-path information fusion (MIF) module that collects abundant information from feature maps of the input, output and intermediary in a module and distills primary information therein. Several cascaded MIF modules are used to progressively extract features desired by reconstruction and the output of each module is gathered for rebuilding the HR image. In addition, we introduce a refinement network with local residual topology architecture as the second stage so as to further restore the high-frequency details of HR image produced by the first stage. Due to less number of filters, the compact model achieves fast inference time and brings about state-of-the-art SR results on four benchmark datasets simultaneously. Code is available at https://github.com/Zheng222/TSCN.
Zheng Hui, Xiumei Wang 0002, Xinbo Gao 0001
ICPR1
2017 Deep Networks for Single Image Super-Resolution with Multi-context Fusion
Zheng Hui, Xiumei Wang 0002, Xinbo Gao 0001
ICIG (1)1