Zhaoyang Jia

dblp:195/7077 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
15since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Towards Practical Real-Time Neural Video Compression
abstract
We introduce a practical real-time neural video codec (NVC) designed to deliver high compression ratio, low latency and broad versatility. In practice, the coding speed of NVCs depends on 1) computational costs, and 2) non-computational operational costs, such as memory I/O and the number of function calls. While most efficient NVCs prioritize reducing computational cost, we identify operational cost as the primary bottleneck to achieving higher coding speed. Leveraging this insight, we introduce a set of efficiency-driven design improvements focused on minimizing operational costs. Specifically, we employ implicit temporal modeling to eliminate complex explicit motion modules, and use single low-resolution latent representations rather than progressive downsampling. These innovations significantly accelerate NVC without sacrificing compression quality. Additionally, we implement model integerization for consistent cross-device coding and a module-bank-based rate control scheme to improve practical adaptability. Experiments show our proposed DCVC-RT achieves an impressive average encoding/decoding speed at 125.2/112.8 fps (frames per second) for 1080p video, while saving an average of 21% in bitrate compared to H.266/VTM. The code is available at https://github.com/microsoft/DCVC.
Zhaoyang Jia, Bin Li 0012, Jiahao Li 0001, Wenxuan Xie, Houqiang Li, Yan Lu 0001
CVPR1
2025 DLF: Extreme Image Compression with Dual-Generative Latent Fusion
Naifu Xue, Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001
ICCV2
2025 One-Step Diffusion-Based Image Compression with Semantic Distillation
abstract
While recent diffusion-based generative image codecs have shown impressive performance, their iterative sampling process introduces unpleasant latency. In this work, we revisit the design of a diffusion-based codec and argue that multi-step sampling is not necessary for generative compression. Based on this insight, we propose OneDC, a One-step Diffusion-based generative image Codec—that integrates a latent compression module with a one-step diffusion generator. Recognizing the critical role of semantic guidance in one-step diffusion, we propose using the hyperprior as a semantic signal, overcoming the limitations of text prompts in representing complex visual content. To further enhance the semantic capability of the hyperprior, we introduce a semantic distillation mechanism that transfers knowledge from a pretrained generative tokenizer to the hyperprior codec. Additionally, we adopt a hybrid pixel- and latent-domain optimization to jointly enhance both reconstruction fidelity and perceptual realism. Extensive experiments demonstrate that OneDC achieves SOTA perceptual quality even with one-step generation, offering over 39% bitrate reduction and 20× faster decoding compared to prior multi-step diffusion-based codecs. Project: https://onedc-codec.github.io/
Naifu Xue, Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001
NeurIPS2
2025 Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
abstract
Long-form video understanding presents significant challenges due to extensive temporal-spatial complexity and the difficulty of question answering under such extended contexts. While Large Language Models (LLMs) have demonstrated considerable advancements in video analysis capabilities and long context handling, they continue to exhibit limitations when processing information-dense hour-long videos. To overcome such limitations, we propose the $\textbf{D}eep \ \textbf{V}ideo \ \textbf{D}iscovery \ (\textbf{DVD})$ agent to leverage an $\textit{agentic search}$ strategy over segmented video clips. Different from previous video agents manually designing a rigid workflow, our approach emphasizes the autonomous nature of agents. By providing a set of search-centric tools on multi-granular video database, our DVD agent leverages the advanced reasoning capability of LLM to plan on its current observation state, strategically selects tools to orchestrate adaptive workflow for different queries in light of the gathered information. We perform comprehensive evaluation on multiple long video understanding benchmarks that demonstrates our advantage. Our DVD agent achieves state-of-the-art performance on the challenging LVBench dataset, reaching an accuracy of $\textbf{74.2\%}$, which substantially surpasses all prior works, and further improves to $\textbf{76.0\%}$ with transcripts.
Zhaoyang Jia, Zongyu Guo, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001
NeurIPS2
2025 OMKar: Optical Map Based Automated Karyotyping of Genomes to Identify Constitutional Disorders
Siavash Raeisi Dehkordi, Zhaoyang Jia, Joey Estabrook, Jen Hauenstein, Neil Miller, Naz Güleray-Lafci, Jürgen Neesen, Alex Hastie, Andy Wing Chun Pang, Paul Dremsek, Vineet Bafna
RECOMB2
2025 C³shartMark: A Chart Watermarking Scheme With Consecutive-Encoding and Concurrent-Decoding
abstract
Chart images are widely employed as the intuitive form to express information, which renders them highly valuable. Consequently, there is an urgent demand to develop a watermarking algorithm for copyright protection and leakage prevention of chart images. Nevertheless, existing chart watermarking methods fail to thoroughly consider the chart image’s special characteristics and simply rely on the previous natural image-based watermarking framework. Compared to natural images, the chart image generally exhibits relatively simple layouts and textures, containing fewer complex texture regions that watermarks are typically embedded in. Therefore, the embedding locations of watermarks for different distortions can be relatively dispersed in natural images, while for chart images, watermark embedding regions under various distortion conditions tend to be relatively concentrated and share more overlaps. Inspired by the above special characteristics of chart images, to sufficiently leverage them and design a better framework, this paper proposes C3hartMark, a chart watermarking scheme with consecutive-encoding and concurrent-decoding. Instead of using the combined noise layer as existing methods to ensure multiple robustness, a novel consecutive training framework is introduced in this paper, which efficiently utilizes the overlapping of embedded watermark features in chart images, and simultaneously, mitigates the poor convergence brought by the combined noise layer. During the extraction stage, multiple concurrent decoders are introduced to extract the potential embedded watermarks for different distortions independently. Moreover, we also incorporate two special noise layers, namely Captioning and Fusion, to address the corresponding realistic distortions in chart images, and an agnostic noise layer to accommodate potential channel transmission distortions unknown during training. Through extensive experiments, we demonstrate that with the better visual quality, C3hartMark simultaneously outperforms existing state-of-the-art (SOTA) watermarking methods in terms of robustness, achieving 99.57% extraction accuracy under JPEG compression (QF=60).
Linfeng Ma, Han Fang 0004, Zehua Ma, Zhaoyang Jia, Weiming Zhang 0001, Nenghai Yu
IEEE Trans. Circuits Syst. Video Technol.4
2025 Generative Latent Coding for Ultra-Low Bitrate Image and Video Compression
abstract
Most existing approaches for image and video compression perform transform coding in the pixel space to reduce redundancy. However, due to the misalignment between the pixel-space distortion and human perception, such schemes often face the difficulties in achieving both high-realism and high-fidelity at ultra-low bitrate. To solve this problem, we propose Generative Latent Coding (GLC) models for image and video compression, termed GLC-image and GLC-Video. The transform coding of GLC is conducted in the latent space of a generative vector-quantized variational auto-encoder (VQ-VAE). Compared to the pixel-space, such a latent space offers greater sparsity, richer semantics and better alignment with human perception, and show its advantages in achieving high-realism and high-fidelity compression. To further enhance performance, we improve the hyper prior by introducing a spatial categorical hyper module in GLC-image and a spatio-temporal categorical hyper module in GLC-video. Additionally, the code-prediction-based loss function is proposed to enhance the semantic consistency. Experiments demonstrate that our scheme shows high visual quality at ultra-low bitrate for both image and video compression. For image compression, GLC-image achieves an impressive bitrate of less than 0.04 bpp, achieving the same FID as previous SOTA model MS-ILLM while using 45% fewer bitrate on the CLIC 2020 test set. For video compression, GLC-video achieves 65.3% bitrate saving over PLVC in terms of DISTS.
Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Generative Latent Coding for Ultra-Low Bitrate Image Compression
abstract
Most existing image compression approaches perform transform coding in the pixel space to reduce its spatial re-dundancy. However, they encounter difficulties in achieving both high-realism and high-fidelity at low bitrate, as the pixel-space distortion may not align with human perception. To address this issue, we introduce a Generative Latent Coding (GLC) architecture, which performs transform coding in the latent space of a generative vector-quantized variational auto-encoder (VQ- VAE), instead of in the pixel space. The generative latent space is characterized by greater sparsity, richer semantic and better alignment with human perception, rendering it advantageous for achieving high-realism and high-fidelity compression. Additionally, we introduce a categorical hyper module to reduce the bit cost of hyper-information, and a code-prediction-based su-pervision to enhance the semantic consistency. Experiments demonstrate that our GLC maintains high visual quality with less than 0.04 bpp on natural images and less than 0.01 bpp on facial images. On the CLIC2020 test set, we achieve the same FID as MS-ILLM with 45% fewer bits. Furthermore, the powerful generative latent space enables various applications built on our GLC pipeline, such as image restoration and style transfer.
Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001
CVPR1
2024 Long-Term Temporal Context Gathering for Neural Video Compression
Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001
ECCV (66)2
2024 Exploring Neighbor Correspondence Matching for Multiple-hypotheses Video Frame Synthesis
abstract
Video frame synthesis, which consists of interpolation and extrapolation , is an essential video processing technique that can be applied to various scenarios. However, most existing methods cannot handle small objects or large motion well, especially in high-resolution videos such as 4K videos. To eliminate such limitations, we introduce a neighbor correspondence matching (NCM) algorithm for flow-based frame synthesis. Since the current frame is not available in video frame synthesis, NCM is performed in a current-frame-agnostic fashion to establish multi-scale correspondences in the spatial-temporal neighborhoods of each pixel. Based on the powerful motion representation capability of NCM, we propose a heterogeneous coarse-to-fine scheme for intermediate flow estimation. The coarse-scale and fine-scale modules are trained progressively, making NCM computationally efficient and robust to large motions. We further explore the mechanism of NCM and find that neighbor correspondence is powerful, since it provides multiple-hypotheses motion information for synthesis. Based on this analysis, we introduce a multiple-hypotheses estimation process for video frame extrapolation, resulting in a more robust framework, NCM-MH. Experimental results show that NCM and NCM-MH achieve 31.63 and 28.08 dB for interpolation and extrapolation on the most challenging X4K1000FPS benchmark, outperforming all the other state-of-the-art methods that use two reference frames as input.
Zhaoyang Jia, Yan Lu 0001, Houqiang Li
ACM Trans. Multim. Comput. Commun. Appl.1
2023 De-END: Decoder-Driven Watermarking Network
abstract
Deep-learning-based watermarking technique is being extensively studied. Most existing approaches adopt a similar encoder-driven scheme which we name END (Encoder-NoiseLayer-Decoder) architecture. In this paper, we revamp the architecture and creatively design a decoder-driven watermarking network dubbed De-END which greatly outperforms the existing END-based methods. The motivation for designing De-END originated from the potential drawback we discovered in END architecture: The encoder may embed redundant features that are not necessary for decoding, limiting the performance of the whole network. We conducted a detailed analysis and found that such limitations are caused by unsatisfactory coupling between the encoder and decoder in END. De-END addresses such drawbacks by adopting a Decoder -Encoder-Noiselayer-Decoder architecture. In De-END, the host image is firstly processed by the decoder to generate a latent feature map instead of being directly fed into the encoder. This latent feature map is concatenated to the original watermark message and then processed by the encoder. This change in design is crucial as it makes the feature of encoder and decoder directly shared thus the encoder and decoder are better coupled. We conducted extensive experiments and the results show that this framework outperforms the existing state-of-the-art (SOTA) END-based deep learning watermarking both in visual quality and robustness. On the premise of the same decoder structure, the visual quality (measured by PSNR) of De-END improves by 1.6dB (45.16dB to 46.84dB), and extraction accuracy after JPEG compression (QF=50) distortion outperforms more than 4% (94.9% to 99.1%).
Han Fang 0004, Zhaoyang Jia, Yupeng Qiu, Jiyi Zhang, Weiming Zhang 0001, Ee-Chien Chang
IEEE Trans. Multim.2
2023 Encoded Feature Enhancement in Watermarking Network for Distortion in Real Scenes
abstract
Deep-learning based watermarking framework has been extensively studied recently. The main structure of such framework is an encoder, a noise layer and a decoder. By training with different distortion sets in the noise layer, the whole network can realize different robustness. However, such framework has a huge drawback that the noise layer must be differentiable, otherwise it cannot be trained end-to-end. But for practical use, much distortions are non-differentiable, so such framework cannot be applied. To address such limitations, this paper propose a triple-phase watermarking framework for practical distortions. The proposed framework consists of three phases including a noise-free initial phase, a mask-guided frequency enhancement phase and an adversarial-training phase. Phase 1 aims to initialize an encoder to embed watermark with high visual quality and a decoder to extract the watermark. In order to generate high quality watermarked image, we design the just noticeable difference (JND)-mask image loss in phase 1 to guide the encoder. At phase 2, based on the investigation of the encoded features and distortions, we propose a mask-guided frequency enhancement algorithm to enhance the encoded feature which ensures the survival of such features after distortion, so that there will be enough features to be learned in phase 3. And phase 3 aims to train a stronger decoder to extract the watermark from the image after practical distortions. The combination of these 3 phases can well handle the non-differentiable problems and make the whole network trainable. Various experiments indicate the superior performance of the proposed scheme in the view of traditional differentiable image processing distortion robustness and practical non-differentiable distortion robustness.
Han Fang 0004, Zhaoyang Jia, Hang Zhou 0007, Zehua Ma, Weiming Zhang 0001
IEEE Trans. Multim.2
2022 PIMoG: An Effective Screen-shooting Noise-Layer Simulation for Deep-Learning-Based Watermarking Network
abstract
With the omnipresence of camera phone and digital display, capturing digitally displayed image with camera phone are getting widely practiced. In the context of watermarking, this brings forth the issue of screen-shooting robustness. The key to acquiring screen-shooting robustness is designing a good noise layer that could represent screen-shooting distortions in a deep-learning-based watermarking framework. However, it is very difficult to quantitatively formulate the screen-shooting distortion since the screen-shooting process is too complex. In order to design an effective noise layer for screen-shooting robustness, we propose new insight in this paper, that is, it is not necessary to quantitatively simulate the overall procedure in the screen-shooting noise layer, only including the most influenced distortions is enough to generate an effective noise layer with strong robustness. To verify this insight, we propose a screen-shooting noise layer dubbed PIMoG. Specifically, we summarize the most influenced distortions of screen-shooting process into three parts (p erspective distortion, i llumination distortion and mo iré distortion) and further simulate them in a differentiable way. For the rest distortion, we utilize the G aussian noise to approximate the main part of them. As a result, the whole network can be trained end-to-end with such noise layer. Extensive experiments illustrate the superior performance of the proposed PIMoG noise layer. In addition to the noise layer design, we also propose a gradient mask-guided image loss and an edge mask-guided image loss to further improve the robustness and invisibility of the whole network respectively. Based on the proposed loss and PIMoG noise layer, the whole framework outperforms the SOTA watermarking method with at least 5% in extraction accuracy and achieves more than 97% accuracy in different screen-shooting conditions.
Han Fang 0004, Zhaoyang Jia, Zehua Ma, Ee-Chien Chang, Weiming Zhang 0001
ACM Multimedia2
2022 Neighbor Correspondence Matching for Flow-based Video Frame Synthesis
abstract
Video frame synthesis, which consists of interpolation and extrapolation, is an essential video processing technique that can be applied to various scenarios. However, most existing methods cannot handle small objects or large motion well, especially in high-resolution videos such as 4K videos. To eliminate such limitations, we introduce a neighbor correspondence matching (NCM) algorithm for flow-based frame synthesis. Since the current frame is not available in video frame synthesis, NCM is performed in a current-frame-agnostic fashion to establish multi-scale correspondences in the spatial-temporal neighborhoods of each pixel. Based on the powerful motion representation capability of NCM, we further propose to estimate intermediate flows for frame synthesis in a heterogeneous coarse-to-fine scheme. Specifically, the coarse-scale module is designed to leverage neighbor correspondences to capture large motion, while the fine-scale module is more computationally efficient to speed up the estimation process. Both modules are trained progressively to eliminate the resolution gap between training dataset and real-world videos. Experimental results show that NCM achieves state-of-the-art performance on several benchmarks. In addition, NCM can be applied to various practical scenarios such as video compression to achieve better performance.
Zhaoyang Jia, Yan Lu 0001, Houqiang Li
ACM Multimedia1
2021 MBRS: Enhancing Robustness of DNN-based Watermarking by Mini-Batch of Real and Simulated JPEG Compression
abstract
Based on the powerful feature extraction ability of deep learning architecture, recently, deep-learning based watermarking algorithms have been widely studied. The basic framework of such algorithm is the auto-encoder like end-to-end architecture with an encoder, a noise layer and a decoder. The key to guarantee robustness is the adversarial training with the differential noise layer. However, we found that none of the existing framework can well ensure the robustness against JPEG compression, which is non-differential but is an essential and important image processing operation. To address such limitations, we proposed a novel end-to-end training architecture, which utilizes Mini-Batch of Real and Simulated JPEG compression (MBRS) to enhance the JPEG robustness. Precisely, for different mini-batches, we randomly choose one of real JPEG, simulated JPEG and noise-free layer as the noise layer. Besides, we suggest to utilize the Squeeze-and-Excitation blocks which can learn better feature in embedding and extracting stage, and propose a "message processor" to expand the message in a more appreciate way. Meanwhile, to improve the robustness against crop attack, we propose an additive diffusion block into the network. The extensive experimental results have demonstrated the superior performance of the proposed scheme compared with the state-of-the-art algorithms. Under the JPEG compression with quality factor $Q=50$, our models achieve a bit error rate less than 0.01% for extracted messages, with PSNR larger than 36 for the encoded images, which shows the well-enhanced robustness against JPEG attack. Besides, under many other distortions such as Gaussian filter, crop, cropout and dropout, the proposed framework also obtains strong robustness. The code implemented by PyTorch is avaiable in https://github.com/jzyustc/MBRS.
Zhaoyang Jia, Han Fang 0004, Weiming Zhang 0001
ACM Multimedia1
2020 Comprehensive Verification and Analysis of Multi-Scale Remote Sensing Products for Surface Freezing- Thawing Status on the Qinghai-Tibet Plateau
abstract
The surface soil freezing and thawing (F/T) status plays important roles in water and energy exchanges, hydrology cycle and terrestrial ecosystem. Passive microwave remote sensing tends to be one of the most effective ways of monitoring global surface F/T status, due to its strong sensitivity to the changes of liquid water and its short revisiting period. Many surface F/T products have been produced based on the observations from different passive microwave sensors. However, due to the differences in spatial scales, satellite data sources, transit times etc., the accuracy and the consistency of different F/T products need to be well considered in front of the comprehensive utilization. In this study, the surface F/T products with resolutions of 36 km, 25 km, 25 km and 6 km derived from SMAP, SSMIS, AMSR-2, AMSR-2 (downscaling) were validated separately based on the in-situ observations in Ngari prefecture and Maqu on the Qinghai-Tibet Plateau. Then a comparison analysis was conducted on the F/T products derived from SMAP and SSMIS to figure out the consistency between F/T products derived from different sensors. The results show that: 1) the verification accuracy in Maqu is higher than that in Ngari prefecture, which reflects the limited capability of passive microwave in identifying the soil surface F/T status in the arid region. 2) The F/T products with resolutions of 6 km and 25 km, which were both derived from AMSR-2 and have the same transit time, showed similar validation accuracies. 3) There is a good agreement between the F/T products derived from SMAP and that from SSMIS. This study will be conducive to the comprehensive application of multi-scale F/T products in the fields of climate change and ecological environment.
Xiaokang Kou, Zhaoyang Jia, Shuang Yan, Mengjie Jin 0003, Tianliang Wang
IGARSS2
2019 Study on Digital Image Inpainting Method Based on Multispectral Image Decomposition Synthesis
abstract
The paper analyzes the image inpainting problem of damaged Painting Arts for high fidelity images reproduction, and a digital image inpainting method based on multispectral image decomposition synthesis is proposed. Firstly, multi-channel images of Painting Arts are obtained by multispectral technology. Then, a polynomial regression method based on principal component is used to reconstruct the spectral image. The reconstructed image is decomposed by VO image decomposition model. During the inpainting process, the channel correlation of the structure image and the texture image of multispectral image is effectively removed. The digital image inpainting is performed respectively. Finally, the digital inpainted image is obtained by synthesis. The experimental results show that the digital image inpainting based on multispectral image decomposition synthesis reduces the problem of low image inpainting accuracy caused by the correlation between the color components in the traditional digital image inpainting process, and reduces the mismatch of the inpainting image. Appearance of pseudo color of inpainting image is reduced. MSE of multispectral images inpainting qualities is 2.7951 and PSNR of multispectral images inpainting qualities is 44.1681, so it is superior to traditional image inpainting algorithm. It provides a reliable basis for digital inpainting, digital archives and high fidelity replication of defective Painting Arts.
Zhaoyang Jia, Guangxue Chen
Int. J. Pattern Recognit. Artif. Intell.1
2017 An Image Stitching Algorithm Based on Histogram Matching and SIFT Algorithm
abstract
Image stitching among images that have significant illumination changes will lead to unnatural mosaic image. An image stitching algorithm based on histogram matching and scale-invariant feature transform (SIFT) algorithm is brought out to solve the problem in this paper. First, histogram matching is used for image adjustment, so that the images to be stitched are at the same level of illumination, then the paper adopts SIFT algorithm to extract the key points of the images and performs the rough matching process, followed by RANSAC algorithm for fine matches, and finally calculates the appropriate mathematical mapping model between two images and according to the mapping relationship, a simple weighted average algorithm is used for image blending. The experimental results show that the algorithm is effective.
Guangxue Chen, Zhaoyang Jia
Int. J. Pattern Recognit. Artif. Intell.3