EDBT 2026 Demo / reviewers in the wild / expert
Zongyu Guo
dblp:247/4138
· DBLP profile ↗
23ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-3770-3442ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent TrainingabstractGUI agents that interact with graphical interfaces on behalf of users are a promising direction for practical AI assistants, yet training them is hindered by scarce suitable environments. We present InfiniteWeb, a system that automatically generates functional web environments at scale for GUI agent training. While LLMs perform well on generating a single webpage, building a realistic and functional website with many interconnected pages faces challenges. We address these challenges through unified specification, task-centric test-driven development, and combining website seed variation with reference design images. Our system also generates verifiable task evaluators enabling dense reward signals for reinforcement learning. Experiments show that our system surpasses commercial coding agents at realistic website construction, and GUI agents trained on our generated environments achieve significant performance improvements on OSWorld and Online-Mind2Web, demonstrating the effectiveness of the proposed system. Zezhou Wang, Zongyu Guo, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ACL (1) | 4 |
| 2026 | Versatile learned video compression
Runsen Feng, Zongyu Guo, Zhizheng Zhang 0004, Weiping Li 0003, Zhibo Chen 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | GTPC-SSCD: Gate-guided Two-level Perturbation Consistency-based Semi-Supervised Change DetectionabstractSemi-supervised change detection (SSCD) utilizes partially labeled data and abundant unlabeled data to detect differences between multi-temporal remote sensing images. The mainstream SSCD methods based on consistency regularization have limitations. They perform perturbations mainly at a single level, restricting the utilization of unlabeled data and failing to fully tap its potential. In this paper, we introduce a novel Gate-guided Two-level Perturbation Consistency regularization-based SSCD method (GTPC-SSCD). It simultaneously maintains strong-to-weak consistency at the image level and perturbation consistency at the feature level, enhancing the utilization efficiency of unlabeled data. Moreover, we develop a hardness analysis-based gating mechanism to assess the training complexity of different samples and determine the necessity of performing feature perturbations for each sample. Through this differential treatment, the network can explore the potential of unlabeled data more efficiently. Extensive experiments conducted on six benchmark CD datasets demonstrate the superiority of our GTPC-SSCD over seven state-of-the-art methods. Qi'ao Xu, Zongyu Guo, Rui Huang 0006, Yuxiang Zhang 0003 |
ICME | 3 |
| 2025 | Deep Video Discovery: Agentic Search with Tool Use for Long-form Video UnderstandingabstractLong-form video understanding presents significant challenges due to extensive temporal-spatial complexity and the difficulty of question answering under such extended contexts.
While Large Language Models (LLMs) have demonstrated considerable advancements in video analysis capabilities and long context handling, they continue to exhibit limitations when processing information-dense hour-long videos.
To overcome such limitations, we propose the $\textbf{D}eep \ \textbf{V}ideo \ \textbf{D}iscovery \ (\textbf{DVD})$ agent to leverage an $\textit{agentic search}$ strategy over segmented video clips. Different from previous video agents manually designing a rigid workflow, our approach emphasizes the autonomous nature of agents.
By providing a set of search-centric tools on multi-granular video database,
our DVD agent leverages the advanced reasoning capability of LLM to plan on its current observation state, strategically selects tools to orchestrate adaptive workflow for different queries in light of the gathered information.
We perform comprehensive evaluation on multiple long video understanding benchmarks that demonstrates our advantage.
Our DVD agent achieves state-of-the-art performance on the challenging LVBench dataset, reaching an accuracy of $\textbf{74.2\%}$, which substantially surpasses all prior works, and further improves to $\textbf{76.0\%}$ with transcripts. Zhaoyang Jia, Zongyu Guo, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001 |
NeurIPS | 3 |
| 2024 | PAPnet: A Plug-and-play Virus Network for Backdoor AttackabstractMost existing backdoor attacks focus on designing various trigger injection methods and fine-tuning victim networks, which are difficult to deploy in real-world applications. In this paper, we propose a plug-and-play virus network, dubbed PAPnet, for backdoor attack. PAPnet is a lightweight network with the same dimensional output as the victim network. In the training stage, we only need the output of the victim network and train PAPnet to learn from poisoned data and clean data. This makes PAPnet easier to learn than fine-tunebased backdoor attack methods. Besides, PAPnet can be easily attached to the different classification network models without modifying the architecture of the victim network and fine-tuning processing. We have conducted various experiments on four datasets with four classical classification networks. Experimental results demonstrate the superiority of our proposed method. Rui Huang 0006, Zongyu Guo, Qingyi Zhao, Wei Fan 0001 |
CSCWD | 2 |
| 2024 | Conditional Neural Video Coding with Spatial-Temporal Super-ResolutionabstractThis fact sheet describes our proposed method for the video track of Challenge on Learned Image Compression (CLIC) 2024. Our scheme follows the typical hybrid coding framework with advanced techniques in motion estimation, context mining, and spatial-temporal super-resolution to enhance rate-distortion performance, particularly at low bitrates. Henan Wang, Xiaohan Pan, Runsen Feng, Zongyu Guo, Zhibo Chen 0001 |
DCC | 4 |
| 2024 | SPY-Watermark: Robust Invisible Watermarking for Backdoor AttackabstractBackdoor attack aims to deceive a victim model when facing backdoor instances while maintaining its performance on benign data. Current methods use manual patterns or special perturbations as triggers, while they often overlook the robustness against data corruption, making backdoor attacks easy to defend in practice. To address this issue, we propose a novel backdoor attack method named Spy-Watermark, which remains effective when facing data collapse and backdoor defense. Therein, we introduce a learnable watermark embedded in the latent domain of images, serving as the trigger. Then, we search for a watermark that can withstand collapse during image decoding, cooperating with several anti-collapse operations to further enhance the resilience of our trigger against data corruption. Extensive experiments are conducted on CIFAR10, GTSRB, and ImageNet datasets, demonstrating that Spy-Watermark overtakes ten state-of-the-art methods in terms of robustness and stealthiness. Ruofei Wang, Renjie Wan, Zongyu Guo, Qing Guo 0005, Rui Huang 0006 |
ICASSP | 3 |
| 2024 | PBIM: Paired Backdoor Injection Method for Change Detection
Rui Huang 0006, Mengjia Hao, Zongyu Guo |
ICIC (3) | 3 |
| 2024 | RECOMBINER: Robust and Enhanced Compression with Bayesian Implicit Neural RepresentationsabstractCOMpression with Bayesian Implicit NEural Representations (COMBINER) is a recent data compression method that addresses a key inefficiency of previous Implicit Neural Representation (INR)-based approaches: it avoids quantization and enables direct optimization of the rate-distortion performance. However, COMBINER still has significant limitations: 1) it uses factorized priors and posterior approximations that lack flexibility; 2) it cannot effectively adapt to local deviations from global patterns in the data; and 3) its performance can be susceptible to modeling choices and the variational parameters' initializations. Our proposed method, Robust and Enhanced COMBINER (RECOMBINER), addresses these issues by 1) enriching the variational approximation while retaining a low computational cost via a linear reparameterization of the INR weights, 2) augmenting our INRs with learnable positional encodings that enable them to adapt to local details and 3) splitting high-resolution data into patches to increase robustness and utilizing expressive hierarchical priors to capture dependency across patches. We conduct extensive experiments across several data modalities, showcasing that RECOMBINER achieves competitive results with the best INR-based methods and even outperforms autoencoder-based codecs on low-resolution images at low bitrates. Our PyTorch implementation is available at https://github.com/cambridge-mlg/RECOMBINER/. Jiajun He 0003, Gergely Flamich, Zongyu Guo, José Miguel Hernández-Lobato |
ICLR | 3 |
| 2024 | Exploring the rate-distortion-complexity optimization in neural image compression
Runsen Feng, Zongyu Guo, Zhibo Chen 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | NVTC: Nonlinear Vector Transform CodingabstractIn theory, vector quantization (VQ) is always better than scalar quantization (SQ) in terms of rate-distortion (RD) performance [33]. Recent state-of-the-art methods for neural image compression are mainly based on nonlinear transform coding (NTC) with uniform scalar quantization, overlooking the benefits of VQ due to its exponentially increased complexity. In this paper, we first investigate on some toy sources, demonstrating that even if modern neural networks considerably enhance the compression performance of SQ with nonlinear transform, there is still an insurmountable chasm between SQ and VQ. Therefore, revolving around VQ, we propose a novel framework for neural image compression named Nonlinear Vector Transform Coding (NVTC). NVTC solves the critical complexity issue of VQ through (1) a multi-stage quantization strategy and (2) nonlinear vector transforms. In addition, we apply entropy-constrained VQ in latent space to adaptively determine the quantization boundaries for joint rate-distortion optimization, which improves the performance both theoretically and experimentally. Compared to previous NTC approaches, NVTC demonstrates superior rate-distortion performance, faster decoding speed, and smaller model size. Our code is available at https://github.com/USTC-IMCL/NVTC. Runsen Feng, Zongyu Guo, Weiping Li 0003, Zhibo Chen 0001 |
CVPR | 2 |
| 2023 | Versatile Neural Processes for Learning Implicit Neural Representations
Zongyu Guo, Cuiling Lan, Zhizheng Zhang 0004, Yan Lu 0001, Zhibo Chen 0001 |
ICLR | 1 |
| 2023 | Compression with Bayesian Implicit Neural RepresentationsabstractMany common types of data can be represented as functions that map coordinates to signal values, such as pixel locations to RGB values in the case of an image. Based on this view, data can be compressed by overfitting a compact neural network to its functional representation and then encoding the network weights. However, most current solutions for this are inefficient, as quantization to low-bit precision substantially degrades the reconstruction quality. To address this issue, we propose overfitting variational Bayesian neural networks to the data and compressing an approximate posterior weight sample using relative entropy coding instead of quantizing and entropy coding it. This strategy enables direct optimization of the rate-distortion performance by minimizing the $\beta$-ELBO, and target different rate-distortion trade-offs for a given network architecture by adjusting $\beta$. Moreover, we introduce an iterative algorithm for learning prior weight distributions and employ a progressive refinement process for the variational posterior that significantly enhances performance. Experiments show that our method achieves strong performance on image and audio compression while retaining simplicity. Zongyu Guo, Gergely Flamich, Jiajun He 0003, Zhibo Chen 0001, José Miguel Hernández-Lobato |
NeurIPS | 1 |
| 2023 | Learning Cross-Scale Weighted Prediction for Efficient Neural Video CompressionabstractNeural video codecs have demonstrated great potential in video transmission and storage applications. Existing neural hybrid video coding approaches rely on optical flow or Gaussian-scale flow for prediction, which cannot support fine-grained adaptation to diverse motion content. Towards more content-adaptive prediction, we propose a novel cross-scale prediction module that achieves more effective motion compensation. Specifically, on the one hand, we produce a reference feature pyramid as prediction sources and then transmit cross-scale flows that leverage the feature scale to control the precision of prediction. On the other hand, for the first time, a weighted prediction mechanism is introduced even if only a single reference frame is available, which can help synthesize a fine prediction result by transmitting cross-scale weight maps. In addition to the cross-scale prediction module, we further propose a multi-stage quantization strategy, which improves the rate-distortion performance with no extra computational penalty during inference. We show the encouraging performance of our efficient neural video codec (ENVC) on several benchmark datasets. In particular, the proposed ENVC can compete with the latest coding standard H.266/VVC in terms of sRGB PSNR on UVG dataset for the low-latency mode. We also analyze in detail the effectiveness of the cross-scale prediction module in handling various video content, and provide a comprehensive ablation study to analyze those important components. Test code is available at https://github.com/USTC-IMCL/ENVC. Zongyu Guo, Runsen Feng, Zhizheng Zhang 0004, Xin Jin 0014, Zhibo Chen 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | Image Coding for Machines with Omnipotent Feature Learning
Ruoyu Feng 0001, Xin Jin 0014, Zongyu Guo, Runsen Feng, Tianyu He, Zhizheng Zhang 0004, Simeng Sun, Zhibo Chen 0001 |
ECCV (37) | 3 |
| 2022 | Causal Contextual Prediction for Learned Image CompressionabstractOver the past several years, we have witnessed impressive progress in the field of learned image compression. Recent learned image codecs are commonly based on autoencoders, that first encode an image into low-dimensional latent representations and then decode them for reconstruction purposes. To capture spatial dependencies in the latent space, prior works exploit hyperprior and spatial context model to build an entropy model, which estimates the bit-rate for end-to-end rate-distortion optimization. However, such an entropy model is suboptimal from two aspects: (1) It fails to capture global-scope spatial correlations among the latents. (2) Cross-channel relationships of the latents remain unexplored. In this paper, we propose the concept of separate entropy coding to leverage a serial decoding process for causal contextual entropy prediction in the latent space. Acausal context modelis proposed that separates the latents across channels and makes use of channel-wise relationships to generate highly informative adjacent contexts. Furthermore, we propose acausal global prediction modelto find global reference points for accurate predictions of undecoded points. Both these two models facilitate entropy estimation without the transmission of overhead. In addition, we further adopt a new group-separated attention module to build more powerful transform networks. Experimental results demonstrate that our full image compression model outperforms standard VVC/H.266 codec on Kodak dataset in terms of both PSNR and MS-SSIM, yielding the state-of-the-art rate-distortion performance. Zongyu Guo, Zhizheng Zhang 0004, Runsen Feng, Zhibo Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Soft then Hard: Rethinking the Quantization in Neural Image CompressionabstractQuantization is one of the core components in lossy image compression. For neural image compression, end-to-end optimization requires differentiable approximations of quantization, which can generally be grouped into three categories: additive uniform noise, straight-through estimator and soft-to-hard annealing. Training with additive uniform noise approximates the quantization error variationally but suffers from the train-test mismatch. The other two methods do not encounter this mismatch but, as shown in this paper, hurt the rate-distortion performance since the latent representation ability is weakened. We thus propose a novel soft-then-hard quantization strategy for neural image compression that first learns an expressive latent space softly, then closes the train-test mismatch with hard quantization. In addition, beyond the fixed integer-quantization, we apply scaled additive uniform noise to adaptively control the quantization granularity by deriving a new variational upper bound on actual rate. Experiments demonstrate that our proposed methods are easy to adopt, stable to train, and highly effective especially on complex compression models. Zongyu Guo, Zhizheng Zhang 0004, Runsen Feng, Zhibo Chen 0001 |
ICML | 1 |
| 2021 | Accelerate Neural Image Compression with Channel-Adaptive Arithmetic CodingabstractWe have witnessed the revolutionary progress of learned image compression despite a short history of this field. Some challenges still remain such as computational complexity that prevent the practical application of learning-based codecs. In this paper, we address the issue of heavy time complexity from the view of arithmetic coding. Prevalent learning-based image compression scheme first maps the natural image into latent representations and then conduct arithmetic coding on quantized latent maps. Previous arithmetic coding schemes define the start and end value of the arithmetic codebook as the minimum and maximum of the whole latent maps, ignoring the fact that the value ranges in most channels are shorter. Hence, we propose to use a channel-adaptive codebook to accelerate arithmetic coding. We find that the latent channels have different frequency-related characteristics, which are verified by experiments of neural frequency filtering. Further, the value ranges of latent maps are different across channels which are relatively image-independent. The channel-adaptive characteristics allow us to establish efficient prior codebooks that cover more appropriate ranges to reduce the runtime. Experimental results demonstrate that both the arithmetic encoding and decoding can be accelerated while preserving the rate-distortion performance of compression model. Zongyu Guo, Jun Fu 0007, Runsen Feng, Zhibo Chen 0001 |
ISCAS | 1 |
| 2021 | Analyzing Time Complexity of Practical Learned Image Compression ModelsabstractWe have witnessed the rapid development of learned image compression (LIC). The latest LIC models have outperformed almost all traditional image compression standards in terms of rate-distortion (RD) performance. However, the time complexity of LIC model is still underdiscovered, limiting the practical applications in industry. Even with the acceleration of GPU, LIC models still struggle with long coding time, especially on the decoder side. In this paper, we analyze and test a few prevailing and representative LIC models, and compare their complexity with traditional codecs including H.265/HEVC intra and H.266/VVC intra. We provide a comprehensive analysis on every module in the LIC models, and investigate how bitrate changes affect coding time. We observe that the time complexity bottleneck mainly exists in entropy coding and context modelling. Although this paper pay more attention to experimental statistics, our analysis reveals some insights for further acceleration of LIC model, such as model modification for parallel computing, model pruning and a more parallel context model. Xiaohan Pan, Zongyu Guo, Zhibo Chen 0001 |
VCIP | 2 |
| 2020 | Region Normalization for Image InpaintingabstractFeature Normalization (FN) is an important technique to help neural network training, which typically normalizes features across spatial dimensions. Most previous image inpainting methods apply FN in their networks without considering the impact of the corrupted regions of the input image on normalization, e.g. mean and variance shifts. In this work, we show that the mean and variance shifts caused by full-spatial FN limit the image inpainting network training and we propose a spatial region-wise normalization named Region Normalization (RN) to overcome the limitation. RN divides spatial pixels into different regions according to the input mask, and computes the mean and variance in each region for normalization. We develop two kinds of RN for our image inpainting network: (1) Basic RN (RN-B), which normalizes pixels from the corrupted and uncorrupted regions separately based on the original inpainting mask to solve the mean and variance shift problem; (2) Learnable RN (RN-L), which automatically detects potentially corrupted and uncorrupted regions for separate normalization, and performs global affine transformation to enhance their fusion. We apply RN-B in the early layers and RN-L in the latter layers of the network respectively. Experiments show that our method outperforms current state-of-the-art methods quantitatively and qualitatively. We further generalize RN to other inpainting networks and achieve consistent performance improvements. Tao Yu 0012, Zongyu Guo, Xin Jin 0014, Shilin Wu, Zhibo Chen 0001, Weiping Li 0003, Zhizheng Zhang 0004, Sen Liu 0001 |
AAAI | 2 |
| 2019 | Progressive Image Inpainting with Full-Resolution Residual NetworkabstractRecently, learning-based algorithms for image inpainting achieve remarkable progress dealing with squared or irregular holes. However, they fail to generate plausible textures inside damaged area because there lacks surrounding information. A progressive inpainting approach would be advantageous for eliminating central blurriness, i.e., restoring well and then updating masks. In this paper, we propose full-resolution residual network (FRRN) to fill irregular holes, which is proved to be effective for progressive image inpainting. We show that well-designed residual architecture facilitates feature integration and texture prediction. Additionally, to guarantee completion quality during progressive inpainting, we adopt N Blocks, One Dilation strategy, which assigns several residual blocks for one dilation step. Correspondingly, a step loss function is applied to improve the performance of intermediate restorations. The experimental results demonstrate that the proposed FRRN framework for image inpainting is much better than previous methods both quantitatively and qualitatively. Zongyu Guo, Zhibo Chen 0001, Tao Yu 0012, Jiale Chen 0001, Sen Liu 0001 |
ACM Multimedia | 1 |
| 2019 | Deep Scalable Image Compression via Hierarchical Feature DecorrelationabstractScalable image compression allows reconstructing complete images through partially decoding. It plays an important role for image transmission and storage. In this paper, we study the problem of feature decorrelation for Deep Neural Network (DNN) based image codec. Inspired by self-attention mechanism [1], we design a transformer-based decorrelation unit (DU) and adopt it in our scalable image compression framework to reduce the redundancy of feature representations at different levels. Experimental results demonstrate that proposed framework outperforms the state-of-the-art DNN-based scalable image codec and conventional scalable image codecs in terms of MS-SSIM. We also conduct ablation experiments which explicitly verify the effectiveness of decorrelation unit in our scheme. Zongyu Guo, Zhizheng Zhang 0004, Zhibo Chen 0001 |
PCS | 1 |
| 2019 | Beyond Coding: Detection-driven Image Compression with Semantically Structured Bit-streamabstractWith the development of 5G and edge computing, it is increasingly important to offload intelligent media computing to edge device. Traditional media coding scheme codes the media into one binary stream without a semantic structure, which prevents many important intelligent applications from operating directly in bit-stream level, including semantic analysis, parsing specific content, media editing, etc. Therefore, in this paper, we propose a learning based Semantically Structured Coding (SSC) framework to generate Semantically Structured Bit-stream (SSB), where each part of bit-stream represents a certain object and can be directly used for aforementioned tasks. Specifically, we integrate an object detection module in our compression framework to locate and align the object in feature domain. After applying quantization and entropy coding, the features are re-organized according to detected and aligned objects to form a bit-stream. Besides, different from existing learning-based compression schemes that individually train models for specific bit-rate, we share most of model parameters among various bit-rates to significantly reduce model size for variable-rate compression. Experimental results demonstrate that only at the cost of negligible overhead, objects can be completely reconstructed from partial bit-stream. We also verified that classification and pose estimation can be directly performed on partial bit-stream without performance degradation. Tianyu He, Simeng Sun, Zongyu Guo, Zhibo Chen 0001 |
PCS | 3 |