Jia Wang 0004

dblp:58/6299-4 · DBLP profile ↗
← Back
88ranked-venue papers
16as first author
37since 2021 · last 2026
0000-0002-1714-3206ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 5 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 12 · 11 since 2021Theory of computation · 8 · 6 first-authorDatabases, data management, data science and information retrieval · 7 · 2 first-authorComputer networks · 5 · 1 since 2021Security and privacy · 3 · 3 since 2021Systems, architecture and hardware · 2
YearPublicationVenuePosition
2026 Bidirectional Noise Injection: Enhancing Diffusion Models via Coordinated Input-Output Perturbation
abstract
Diffusion models have demonstrated remarkable success in image generation, yet a persistent challenge remains: the bias between model predictions and the target distribution. In this paper, we propose a Bidirectional Noise Injection framework for enhancing diffusion models, implemented via Coordinated Input-Output Perturbation (CIOP). Our approach mitigates this bias by randomly applying synchronized noise injection to both the model inputs and the prediction targets during the training stage. This stochastic, synchronized noise injected acts as a smoothing mechanism that effectively reduces the 2-Wasserstein distance between the predicted and target distributions, as substantiated by our theoretical analysis based on optimal transport theory. Extensive experiments on multiple benchmark datasets and various generative tasks demonstrate that our method improves generation quality and training efficiency without incurring additional computational cost. Furthermore, the design of CIOP enables seamless integration with existing diffusion model improvements and advanced frameworks, thereby broadening its applicability. These results highlight the potential of Bidirectional Noise Injection via CIOP to alleviate bias in diffusion-based generative models across a wide range of settings.
Tianyi Zheng 0001, Jiayang Gao, Peng-Tao Jiang, Fengxiang Yang, Ben Wan, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0130
AAAI8
2026 Preference-guided debiasing for no-reference enhancement image quality assessment
Shiqi Gao, Zitong Xu, Huiyu Duan, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
Image Vis. Comput.6
2026 Gradient flow-based iterative pruning for efficient and high-quality lightweight diffusion models
Ben Wan, Tianyi Zheng 0001, Yuxiao Wang 0004, Jia Wang 0004
Neural Networks4
2026 Bidirectional Beta-Tuned Diffusion Model
abstract
Diffusion models have gained significant attention in the field of generative modeling due to their capability to produce high-quality samples. However, recent studies show that applying a uniform treatment to all distributions during the training of diffusion models is sub-optimal. In this paper, we present a comprehensive theoretical analysis of the forward process in diffusion models. Our findings indicate that distribution variations are not uniform throughout the diffusion process, with the sharpest changes occurring during the initial stages. Moreover, we observe that the initial distribution converges to a Gaussian distribution at an exponential rate, indicating that different initial distributions rapidly become quite similar during the forward diffusion process. Consequently, employing a uniform timestep sampling strategy does not effectively capture these dynamics, potentially leading to sub-optimal training outcomes for diffusion models. To remedy this, we introduce the Bidirectional Beta-Tuned Diffusion Model (BB-TDM). The BB-TDM leverages the Beta distribution to design the timestep sampling distribution and enhance the separation between different initial distributions during the diffusion process. By selecting appropriate parameters, the BB-TDM ensures that the timestep sampling distribution is aligned with the properties of the forward diffusion process and moderates the convergence speed of different initial distributions. Extensive experiments across various benchmark datasets on different diffusion models confirm the efficacy of the proposed BB-TDM.
Tianyi Zheng 0001, Jiayang Zou, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0115
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Security Problem in Cluster Distributed Storage Systems: Regenerating Code Against Two General Types of Active Adversaries
Tinghan Wang, Chenhao Ying 0001, Jia Wang 0004, Yuan Luo 0003
IEEE Trans. Inf. Forensics Secur.3
2026 Multi-Dimensional Quality Assessment for Single-Image-to-3D Contents: Dataset and Model
abstract
The rapid advancement of AI generation technologies has led to the widespread use of AI-generated multimedia content, including images, videos, and 3D contents, across various applications. While significant progress has been made in quality evaluation for 2D content, evaluating the quality of 3D content synthesized from single image remains an underexplored problem. To bridge this gap, we introduce the first comprehensive subjective evaluation database tailored for assessing the quality of 3D content generated from single image. Our database, named AIGC-SI23DCQA, includes three distinct categories of input images, i.e., realistic images, AI-generated images, and computer graphic (CG) images, with 100 images in each category. Using five representative single-image-to-3D algorithms, we produce 1,500 3D contents and collect 94,500 annotations across three quality dimensions, including texture fidelity, shape accuracy, and overall quality. Based on the constructed database, we first benchmark and evaluate the performance of existing quality assessment methods revealing their limitations in addressing this novel task. Thus, we further propose a novel objective quality assessment method, termed I3DQA, for effective single-image-to-3D content quality assessment. Specifically, I3DQA first extracts the reference features from the source image, and the multi-modal features from the generated 3D content, including the projected video, patches, and large-multimodal model (LMM) features. These features are integrated through symmetric transformer blocks, enabling effective quality-related feature fusion and score prediction. Extensive experiments demonstrate the superior performance of our method and validate the effectiveness of its components. This work provides a foundational resource and a robust framework for advancing research in this emerging field, and our database and model are released at https://github.com/ZedFu/SI23DCQA.
Huiyu Duan, Jing Liu 0002, Yun Liu 0009, Xiaohong Liu 0001, Jia Wang 0004, Xiongkuo Min, Patrick Le Callet, Guangtao Zhai
IEEE Trans. Image Process.7
2025 Pruning for Sparse Diffusion Models Based on Gradient Flow
abstract
Diffusion Models (DMs) have impressive capabilities among generation models, but are limited to slower inference speeds and higher computational costs. Previous works utilize one-shot structure pruning to derive lightweight DMs from pre-trained ones, but this approach often leads to a significant drop in generation quality and may result in the removal of crucial weights. Thus we propose a iterative pruning method based on gradient flow, including the gradient flow pruning process and the gradient flow pruning criterion. We employ a progressive soft pruning strategy to maintain the continuity of the mask matrix and guide it along the gradient flow of the energy function based on the pruning criterion in sparse space, thereby avoiding the sudden information loss typically caused by one-shot pruning. Gradient-flow based criterion prune parameters whose removal increases the gradient norm of loss function and can enable fast convergence for a pruned model in iterative pruning stage. Our extensive experiments on widely used datasets demonstrate that our method achieves superior performance in efficiency and consistency with pre-trained models.
Ben Wan, Tianyi Zheng 0001, Zhaoyu Chen 0001, Yuxiao Wang 0004, Jia Wang 0004
ICASSP5
2025 Information Density Principle for MLLM Benchmarks
abstract
With the emergence of Multimodal Large Language Models (MLLMs), hundreds of benchmarks have been developed to ensure the reliability of MLLMs in downstream tasks. However, the evaluation mechanism itself may not be reliable. For developers of MLLMs, questions remain about which benchmark to use and whether the test results meet their requirements. Therefore, we propose a critical principle of Information Density, which examines how much insight a benchmark can provide for the development of MLLMs. We characterize it from four key dimensions: (1) Fallacy, (2) Difficulty, (3) Redundancy, (4) Diversity. Through a comprehensive analysis of more than 10,000 samples, we measured the information density of 19 MLLM benchmarks. Experiments show that using the latest benchmarks in testing can provide more insight compared to previous ones, but there is still room for improvement in their information density. We hope this principle can promote the development and application of future MLLM benchmarks. Project page: https://github.com/lcysyzxdxc/bench4bench
Chunyi Li 0001, Xiaozhe Li, Yuan Tian 0017, Ziheng Jia, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Haodong Duan, Kai Chen 0026, Guangtao Zhai
ICCV8
2025 SI23DCQA: Perceptual Quality Assessment of Single Image-to-3D Content
abstract
In recent years, significant efforts have been dedicated to advancing 3D content generation. However, existing quality assessment research predominantly focuses on evaluating Text-to-3D Content (T23DC) while ignoring Single Image-to-3D Content (SI23DC). In this paper, we establish the first Single Image-to-3D Content Quality Assessment (SI23DCQA) database to comprehensively study the perceptual quality of SI23DCs. The database contains 1500 SI23DCs, which are generated by 5 common SI23DC algorithms from 300 images including realistic images, AI generated images, and model rendered images. Afterward, we carry out a well-designed subjective experiment to collect subjective quality ratings for SI23DCs from three perspectives including overall, color, and shape. Additionally, a benchmark experiment is conducted with the state-of-the-art no reference image quality assessment (NR-IQA), no reference video quality assessment (NR-VQA), and no reference 3D quality assessment (NR-3DQA) and the experimental results show that current quality assessment methods are limited in evaluating the perceptual loss of SI23DCs. The database is released on https://github.com/ZedFu/SI23DCQA.
Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
ICME6
2025 Differential Properties of Information in Jump-Diffusion Channels
abstract
We propose a channel modeling using jumpdiffusion processes, and study the differential properties of entropy and mutual information. By utilizing the Kramers-Moyal and Kolmogorov equations, we express the mutual information between the input and the output in series and integral forms, presented by Fisher-type information and mismatched KL divergence. We extend de Bruijn's identity and the I-MMSE relation to encompass general Markov processes.
Luyao Fan, Jiayang Zou, Jiayang Gao, Jia Wang 0004
ISIT4
2025 Convexity of Mutual Information Along the Fokker-Planck Flow
abstract
We study the convexity of mutual information as a function of time along the Fokker-Planck flow. The results are generalizations of that along heat flow and Ornstein-Ulenbeck flow, which were established by A. Wibisono and V. Jog. We prove the existence and uniqueness of the classical solutions to a class of Fokker-Planck equations and then we obtain the second derivative of mutual information along the Fokker-Planck equation. If the initial distribution is sufficiently strongly logconcave compared to the steady state, then mutual information always preserves convexity under suitable conditions. In particular, if there exists some time point at which the distribution is sufficiently strongly log-concave, then mutual information will preserve convexity after that time.
Jiayang Zou, Luyao Fan, Jiayang Gao, Jia Wang 0004
ISIT4
2025 LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
abstract
The rapid advancement of Text-guided Image Editing (TIE) enables image modifications through text prompts. However, current TIE models still struggle to balance image quality, editing alignment, and consistency with the original image, limiting their practical applications. Existing TIE evaluation benchmarks and metrics have limitations on scale or alignment with human perception. To this end, we introduce EBench-18K, the first large-scale image Editing Benchmark including 18K edited images with fine-grained human preference annotations for evaluating TIE. Specifically, EBench-18K includes 1,080 source images with corresponding editing prompts across 21 tasks, 18K+ edited images produced by 17 state-of-the-art TIE models, 55K+ mean opinion scores (MOSs) assessed from three evaluation dimensions, and 18K+ question-answering (QA) pairs. Based on EBench-18K, we employ outstanding LMMs to assess edited images, while the evaluation results, in turn, provide insights into assessing the alignment between the LMMs' understanding ability and human preferences. Then, we propose LMM4Edit, a LMM-based metric for evaluating image Editing models from perceptual quality, editing alignment, attribute preservation, and task-specific QA accuracy in an all-in-one manner. Extensive experiments show that LMM4Edit achieves outstanding performance and aligns well with human preference. Zero-shot validation on the other datasets also shows the generalization ability of our model. The dataset and code are available at https://github.com/IntMeGroup/LMM4Edit.
Zitong Xu, Huiyu Duan, Bingnan Liu, Guangji Ma, Shiqi Gao, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai, Weisi Lin
ACM Multimedia9
2025 GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
abstract
The rapid evolution of Multi-modality Large Language Models (MLLMs) is driving significant advancements in visual understanding and generation. Nevertheless, a comprehensive assessment of their capabilities, concerning the fine-grained physical principles especially in geometric optics, remains underexplored. To address this gap, we introduce GOBench, the first benchmark to systematically evaluate MLLMs' ability across two tasks: 1) Generating Optically Authentic Imagery and 2) Understanding Underlying Optical Phenomena. We curate high-quality prompts of geometric optical scenarios and use MLLMs to construct the GOBench-Gen-1k dataset. We then organize subjective experiments to assess the generated imagery based on Optical Authenticity, Aesthetic Quality, and Instruction Fidelity, revealing MLLMs' generation flaws that violate optical principles. For the understanding task, we apply crafted evaluation instructions to test the optical understanding ability of eleven prominent MLLMs. The experimental results demonstrate that current models face significant challenges in both optical generation and understanding. The top-performing generative model, GPT-4o-Image, cannot perfectly complete all generation tasks, and the best-performing MLLM model, Gemini-2.5Pro, attains a mere 37.35% accuracy in optical understanding. Database and codes are publicly available at: https://github.com/aiben-ch/GOBench.
Xiaorong Zhu, Ziheng Jia, Haodong Duan, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
ACM Multimedia7
2025 A Light-Aware Quality Assessment Method for Relighted Human Heads Based on Multi-task Learning
Farong Wen, Yingjie Zhou 0003, Xiaohong Liu 0001, Jia Wang 0004, Jiezhang Cao, Yu Wang 0002, Guangtao Zhai
PRCV (12)5
2025 Subjective and Objective Quality-of-Experience Evaluation Study for Live Video Streaming
abstract
In recent years, live video streaming has gained widespread popularity across various social media platforms. Quality of experience (QoE), which reflects end-users’ satisfaction and overall experience, plays a critical role for media service providers to optimize large-scale live compression and transmission strategies to achieve perceptually optimal rate-distortion trade-off. Although many QoE metrics for video-on-demand (VoD) have been proposed, there remain significant challenges in developing QoE metrics for live video streaming. To bridge this gap, we conduct a comprehensive study of subjective and objective QoE evaluations for live video streaming. For the subjective QoE study, we introduce the first live video streaming QoE dataset, TaoLive QoE, which consists of 42 source videos collected from real live broadcasts and 1, 155 corresponding distorted ones degraded due to a variety of streaming distortions, including conventional streaming distortions such as compression, stalling, as well as live streaming-specific distortions like frame skipping, variable frame rate, etc. Subsequently, a human study was conducted to derive subjective QoE scores of videos in the TaoLive QoE dataset. For the objective QoE study, we benchmark existing QoE models on the TaoLive QoE dataset as well as publicly available QoE datasets for VoD scenarios, highlighting that current models struggle to accurately assess video QoE, particularly for live content. Hence, we propose an end-to-end QoE evaluation model, Tao-QoE, which integrates multi-scale semantic features and optical flow-based motion features to predicting a retrospective QoE score, eliminating reliance on statistical quality of service (QoS) features. Extensive experiments demonstrate that Tao-QoE outperforms other models on the TaoLive QoE dataset and five publicly available QoE datasets, showcasing the effectiveness and feasibility of Tao-QoE.
Zehao Zhu, Wei Sun 0029, Jun Jia, Jia Wang 0004, Guangtao Zhai
VCIP4
2025 Situation-adaptive neural network for fast pre-computing image enhancement
Xinyue Li 0001, Huiyu Duan, Jia Wang 0004, Xiaohong Liu 0001, Guangtao Zhai
Sci. China Inf. Sci.3
2025 Exploring Bidirectional Bounds for Minimax-Training of Energy-Based Models
Cong Geng, Jia Wang 0004, Li Chen 0021, Jes Frellsen, Søren Hauberg
Int. J. Comput. Vis.2
2025 Enhancing the accuracy of Generative Adversarial Networks with Fokker-Planck Equations
Ben Wan, Tianyi Zheng 0001, Zhaoyu Chen 0001, Jia Wang 0004
Neurocomputing4
2025 EnfoMax: Domain entropy and mutual information maximization for domain generalized face anti-spoofing
Tianyi Zheng 0001, Bo Li 0115, Shuang Wu 0001, Ben Wan, Guodong Mu, Shice Liu, Shouhong Ding, Jia Wang 0004
Neurocomputing8
2025 EBM-WGF: Training energy-based models with Wasserstein gradient flow
Ben Wan, Cong Geng, Tianyi Zheng 0001, Jia Wang 0004
Neural Networks4
2025 Information-Theoretic Security Problem in Cluster Distributed Storage Systems: Regenerating Code Against Two General Types of Eavesdroppers
abstract
In recent years, there has been growing interest in heterogeneous distributed storage systems (DSSs), such as clustered DSSs, which are widely used in practice. However, research regarding information-theoretic security in heterogeneous DSSs remains limited. Furthermore, unlike traditional DSSs, the heterogeneous DSSs face eavesdropper with diverse operating patterns, complicating the secrecy models. In this paper, we aim to investigate the secrecy capacity and code constructions for clustered DSSs (CDSSs), a type of heterogeneous DSSs in which the system is divided into clusters with an equal number of nodes and different repair bandwidths for intra-cluster and cross-cluster against two types of eavesdroppers: the occupying-type eavesdropper and the osmotic-type eavesdropper. We construct two CDSS secrecy models tailored to these aforementioned eavesdroppers, derive the upper bounds on adjustable secrecy capacities, and explore the relationships between the upper bounds of perfect secrecy capacities and the number of compromised nodes. Notably, the upper bounds obtained in this paper generalize those of the traditional DSS model. Additionally, we propose three repair-by-transfer code constructions that achieve the secrecy capacity under both eavesdropper scenarios. These codes are based on nested MDS code and represent a generalized form of the minimum bandwidth regenerating (MBR) codes in traditional DSSs.
Tinghan Wang, Chenhao Ying 0001, Jia Wang 0004, Yuan Luo 0003
IEEE Trans. Inf. Forensics Secur.3
2025 Multi-Dimensional Quality Assessment for Text-to-3D Assets: Dataset and Model
abstract
Recent advancements in text-to-image (T2I) generation have spurred the development of text-to-3D asset (T23DA) generation, leveraging pretrained 2D text-to-image diffusion models for text-to-3D asset synthesis. Despite the growing popularity of text-to-3D asset generation, its evaluation has not been well considered and studied. However, given the significant quality discrepancies among various text-to-3D assets, there is a pressing need for quality assessment models aligned with human subjective judgments. To tackle this challenge, we conduct a comprehensive study to explore the T23DA quality assessment (T23DAQA) problem in this work from both subjective and objective perspectives. Given the absence of corresponding databases, we first establish the largest text-to-3D asset quality assessment database to date, termed the AIGC-T23DAQA database. This database encompasses 969 validated 3D assets generated from 170 prompts via 6 popular text-to-3D asset generation models, and corresponding subjective quality ratings for these assets from the perspectives of quality, authenticity, and text-asset correspondence, respectively. Subsequently, we establish a comprehensive benchmark based on the AIGC-T23DAQA database, and devise an effective T23DAQA model to evaluate the generated 3D assets from the aforementioned three perspectives, respectively. Specifically, the proposed method utilizes the projection videos of text-to-3D assets to extract 3D shape, texture and text-asset correspondence features, then fuses them to calculate the final three preference scores respectively. Extensive experimental results demonstrate the effectiveness of the proposed T23DAQA method in evaluating the quality of AI generated 3D asset, which is more consistent with human perception. To the best of our knowledge, this is the first work that studies the problem of text-guided 3D generation quality assessment, and our database and codes will be released to facilitate future research.
Huiyu Duan, Xiaohong Liu 0001, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
IEEE Trans. Multim.6
2025 Quality-Guided Skin Tone Enhancement for Portrait Photography
abstract
In recent years, learning-based color and tone enhancement methods for photos have become increasingly popular. However, most learning-based image enhancement methods just learn a mapping from one distribution to another based on one dataset, lacking the ability to adjust images continuously and controllably. It is important to enable the learning-based enhancement models to adjust an image continuously, since in many cases we may want to get a slighter or stronger enhancement effect rather than one fixed adjusted result. In this paper, we propose a quality-guided image enhancement paradigm that enables image enhancement models to learn the distribution of images with various quality ratings. By learning this distribution, image enhancement models can associate image features with their corresponding perceptual qualities, which can be used to adjust images continuously according to different quality scores. To validate the effectiveness of our proposed method, a subjective quality assessment experiment is first conducted, focusing on skin tone adjustment in portrait photography. Guided by the subjective quality ratings obtained from this experiment, our method can adjust the skin tone corresponding to different quality requirements. Furthermore, an experiment conducted on 10 natural raw images corroborates the effectiveness of our model in situations with fewer subjects and fewer shots, and also demonstrates its general applicability to natural images.
Shiqi Gao, Huiyu Duan, Xinyue Li 0001, Yicong Peng, Qihang Xu, Yuanyuan Chang, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Multim.8
2024 Beta-Tuned Timestep Diffusion Model
Tianyi Zheng 0001, Peng-Tao Jiang, Ben Wan, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0115
ECCV (3)6
2024 AttentionLUT: Attention Fusion-Based Canonical Polyadic LUT for Real-Time Image Enhancement
abstract
Recently, many algorithms have employed image-adaptive lookup tables (LUTs) to achieve real-time image enhancement. Nonetheless, a prevailing trend among existing methods has been the employment of linear combinations of basic LUTs to formulate image-adaptive LUTs, which limits the generalization ability of these methods. To address this limitation, we propose a novel framework named AttentionLut for real-time image enhancement, which utilizes the attention mechanism to generate image-adaptive LUTs. Our proposed framework consists of three lightweight modules. We begin by employing the global image context feature module to extract image-adaptive features. Subsequently, the attention fusion module integrates the image feature with the priori attention feature obtained during training to generate image-adaptive canonical polyadic tensors. Finally, the canonical polyadic reconstruction module is deployed to reconstruct image-adaptive residual 3DLUT, which is subsequently utilized for enhancing input images. Experiments on the benchmark MIT-Adobe FiveK dataset demonstrate that the proposed method achieves better enhancement performance quantitatively and qualitatively than the state-of-the-art methods.
Yicong Peng, Qihang Xu, Xiaohong Liu 0001, Jia Wang 0004, Guangtao Zhai
ICASSP6
2024 FAMIM: A Novel Frequency-Domain Augmentation Masked Image Model Framework for Domain Generalizable Face Anti-Spoofing
abstract
While existing face anti-spoofing (FAS) methods have achieved high performance on in-domain datasets, good generalization is crucial for their real-world application. Previous domain generalizable FAS methods have attempted to identify common features of live samples from different domains in the spatial domain, but finding such features is challenging. To address this issue, we propose a solution to the FAS problem in the frequency domain. Specifically, we propose a novel Frequency-domain Augmentation Masked Image Model Framework (FAMIM) for domain generalizable face anti-spoofing. Our approach involves mixing different style information into the input face image in the frequency domain and then reconstructing the original face image. This reconstruction task is designed to make our encoder insensitive to the domain style information of the input face image. Additionally, we design a label fusion module in FAMIM to ensure that the encoder extracts liveness or spoof features in the masked image. Our extensive experiments on four widely used domain generalizable FAS datasets demonstrate that our method achieves state-of-the-art performance.
Tianyi Zheng 0001, Qinji Yu, Zhaoyu Chen 0001, Jia Wang 0004
ICASSP4
2024 Non-uniform Timestep Sampling: Towards Faster Diffusion Model Training
abstract
Diffusion models have garnered significant success in generative tasks, emerging as the predominant model in this domain. Despite their success, the substantial computational resources required for training diffusion models restrict their practical applications. In this paper, we resort to the optimal transport theory to accelerate the training of diffusion models, providing an in-depth analysis of the forward diffusion process. It shows that the upper bound on the Wasserstein distance of the distribution between any two timesteps in the diffusion process is an exponential decrease of the initial distance by a factor of times. This finding suggests that the state distribution of the diffusion model has a non-uniform rate of change at different points in time, thus highlighting the different importance of the diffusion timestep. To this end, we propose a novel non-uniform timestep sampling method based on the Bernoulli distribution, which favors more frequent sampling in significant timestep intervals. The key idea is to make the model focus on timesteps with larger differences, thus accelerating the training of the diffusion model. Experiments on benchmark datasets reveal that the proposed method significantly reduces the computational overhead while improving the quality of the generated images.
Tianyi Zheng 0001, Cong Geng, Peng-Tao Jiang, Ben Wan, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0115
ACM Multimedia7
2024 Perceptual Skin Tone Color Difference Measurement for Portrait Photography
abstract
In portrait photography, measuring the perceptual color differences (CDs) of skin tone is significant. Many studies have documented that the perception of skin tone is characteristically different from that of other colors. However, most existing CD measures are proposed based on psychophysical data of uniform color patches or natural images, and do not generalize well to the measurement of skin tone. In this paper, we construct the first large-scale portrait dataset for perceptual skin tone CD assessment and conduct psychophysical experiments to collect 160,000 perceptual CD judgments for 40,000 image triplets. Based on this dataset, we propose a deep skin tone CD measure for portrait photography. Extensive experiments demonstrate that our measure substantially outperforms existing CD measures on the problem of assessing skin tone CDs. The constructed dataset and code will be released to facilitate future research.
Shiqi Gao, Huiyu Duan, Qihang Xu, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet
VCIP4
2024 ReLI-QA: A Multidimensional Quality Assessment Dataset for Relighted Human Heads
abstract
Lighting conditions significantly affect the quality of both real and AI-generated images. Facial images are particularly sensitive to lighting due to their detailed nature and the importance of facial features in conveying identity. Poor lighting can easily obscure these critical details. To address this issue, various portrait relighting methods have been developed to adjust the lighting in improperly exposed images. However, these methods often encounter challenges such as overexposure, underexposure, and detail loss in the relighted portraits. Consequently, there is a need for effective quality assessment and control of relighted human heads (RHHs). In this study, one proposed simple baseline and three typical relighting methods are applied to six selected human head (HH) images, resulting in the creation of a quality assessment dataset named ReLI-QA, which comprises 840 RHHs. A multidimensional subjective quality assessment method based on visual guidance is proposed to accurately evaluate the visual quality of each RHH in the dataset. By analyzing the results of subjective experiments, the quality of RHHs is shown to be affected by multiple factors. Finally, based on ReLI-QA, some typical image quality assessment (IQA) methods are selected for benchmark experiments. The experimental results show the limitations of the existing methods in RHH quality assessment. The dataset and code for this research has been released at https://github.com/zyj-2000/ReLI-QA.
Yingjie Zhou 0003, Farong Wen, Jun Jia, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai
VCIP6
2024 MFAE: Masked Frequency Autoencoders for Domain Generalization Face Anti-Spoofing
abstract
The generalizable face anti-spoofing (FAS) has attracted much attention recently. Even though many existing methods perform well under intra-domain settings, the model’s performance in the unseen domain is not satisfying. In this paper, we shift our attention to the frequency domain to seek a solution. Specifically, we examine the characteristics of different frequency band components of FAS images and observe that the model’s cross-domain performance is very sensitive to low-frequency features. To alleviate this sensitivity and improve the model’s performance in FAS cross-domain tasks, we propose a new approach called Masked Frequency Autoencoders (MFAE). MFAE randomly masks a portion of frequencies on the low-frequency spectrum of the image and then reconstructs the image from the resulting embedding. This innovative Masked Image Modeling (MIM) strategy can be used as a self-supervised task for pre-training vision transformers (ViTs), which can reduce the ViT encoder’s sensitivity to domain shifts. Additionally, we add an auxiliary content-regularization decoder in our MFAE to encourage the encoder to be insensitive to low-frequency features. The results show that the model insensitive to low-frequency features performs well on extensive public datasets and outperforms other state-of-the-art methods in cross-domain FAS tasks.
Tianyi Zheng 0001, Bo Li 0115, Shuang Wu 0001, Ben Wan, Guodong Mu, Shice Liu, Shouhong Ding, Jia Wang 0004
IEEE Trans. Inf. Forensics Secur.8
2023 Fundamental Limits of Communication Efficiency for Model Aggregation in Distributed Learning: A Rate-Distortion Approach
abstract
One of the main focuses in distributed learning is communication efficiency, since model aggregation at each round of training can consist of millions to billions of parameters. Several model compression methods, such as gradient quantization and sparsification, have been proposed to improve the communication efficiency of model aggregation. However, the information-theoretic minimum communication cost for a given distortion of gradient estimators is still unknown. In this paper, we study the fundamental limit of communication cost of model aggregation in distributed learning from a rate-distortion perspective. By formulating the model aggregation as a vector Gaussian CEO problem, we derive the rate region bound and sum-rate-distortion function for the model aggregation problem, which reveals the minimum communication rate at a particular gradient distortion upper bound. We also analyze the communication cost at each iteration and total communication cost based on the sum-rate-distortion function with the gradient statistics of real-world datasets. It is found that the communication gain by exploiting the correlation between worker nodes is significant for SignSGD, and a high distortion of gradient estimator can achieve low total communication cost in gradient compression.
Naifu Zhang, Meixia Tao, Jia Wang 0004, Fan Xu 0001
IEEE Trans. Commun.3
2022 Calculation of ophthalmic diagnostic parameters on a single eye image based on deep neural network
Xuefei Song, Xiongkuo Min, Huifang Zhou, Wei Sun 0029, Jia Wang 0004, Guangtao Zhai
Multim. Tools Appl.6
2021 Sum-Rate-Distortion Function for Indirect Multiterminal Source Coding in Federated Learning
abstract
One of the main focus in federated learning (FL) is the communication efficiency since a large number of participating edge devices send their updates to the edge server at each round of the model training. Existing works reconstruct each model update from edge devices and implicitly assume that the local model updates are independent over edge devices. In FL, however, the model update is an indirect multi-terminal source coding problem, also called as the CEO problem where each edge device cannot observe directly the gradient that is to be reconstructed at the decoder, but is rather provided only with a noisy version. The existing works do not leverage the redundancy in the information transmitted by different edges. This paper studies the rate region for the indirect multiterminal source coding problem in FL. The goal is to obtain the minimum achievable rate at a particular upper bound of gradient variance. We obtain the rate region for the quadratic vector Gaussian CEO problem under unbiased estimator and derive an explicit formula of the sum-rate-distortion function in the special case where gradient are identical over edge device and dimension. Finally, we analyse communication efficiency of convex Mini-batched SGD and non-convex Minibatched SGD based on the sum-rate-distortion function, respectively.
Naifu Zhang, Meixia Tao, Jia Wang 0004
ISIT3
2021 Bounds all around: training energy-based models with bidirectional bounds
abstract
Energy-based models (EBMs) provide an elegant framework for density estimation, but they are notoriously difficult to train. Recent work has established links to generative adversarial networks, where the EBM is trained through a minimax game with a variational value function. We propose a bidirectional bound on the EBM log-likelihood, such that we maximize a lower bound and minimize an upper bound when solving the minimax game. We link one bound to a gradient penalty that stabilizes training, thereby provide grounding for best engineering practice. To evaluate the bounds we develop a new and efficient estimator of the Jacobi-determinant of the EBM generator. We demonstrate that these developments stabilize training and yield high-quality density estimation and sample generation.
Cong Geng, Jia Wang 0004, Jes Frellsen, Søren Hauberg
NeurIPS2
2021 Psycho-visual modulation based information display: introduction and survey
Zhongpai Gao, Jia Wang 0004, Guangtao Zhai
Frontiers Comput. Sci.3
2021 Fine localization and distortion resistant detection of multi-class barcode in complex environments
Xiongkuo Min, Jun Jia, Zehao Zhu, Jia Wang 0004, Guangtao Zhai
Multim. Tools Appl.5
2021 Group Re-Identification With Group Context Graph Neural Networks
abstract
Group re-identification aims to match groups of people across disjoint cameras. In this task, the contextual information from neighbor individuals can be exploited for re-identifying each individual within the group as well as the entire group. However, compared with single person re-identification, it brings new challenges including group layout and group membership changes. Motivated by the observation that individuals who are close together are more likely to keep in the same group under different cameras than those who are far apart, we propose to model each group as a spatial K-nearest neighbor graph (SKNNG) and design a group context graph neural network (GCGNN) for graph representation learning. Specifically, for each node in the graph, the proposed GCGNN learns an embedding which aggregates the contextual information from neighbor nodes. We design multiple weighting kernels for neighborhood aggregation based on the graph properties including node in-degrees and spatial relationship attributes. We compute the similarity scores between node embeddings of two graphs for group member association and obtain the matching score between the two graphs by summing up the similarity scores of all linked node pairs. Experimental results on three public datasets show that our approach performs favorably against state-of-the-art methods and achieves high efficiency.
Ji Zhu 0002, Hua Yang 0001, Weiyao Lin, Nian Liu 0002, Jia Wang 0004, Wenjun Zhang 0001
IEEE Trans. Multim.5
2020 Generalized Gaussian Multiterminal Source Coding: The Symmetric Case
abstract
Consider a generalized multiterminal source coding system, where (mℓ) encoders, each observing a distinct size-m subset of I (ℓ ≥ 2) zero-mean unit-variance exchangeable Gaussian sources with correlation coefficient p, compress their observations in such a way that a joint decoder can reconstruct the sources within a prescribed mean squared error distortion based on the compressed data. The optimal rate-distortion performance of this system was previously known only for the two extreme cases m = ℓ (the centralized case) and m = 1 (the distributed case), and except when ρ = 0, the centralized system can achieve strictly lower compression rates than the distributed system under all non-trivial distortion constraints. Somewhat surprisingly, it is established in the present paper that the optimal rate-distortion performance of the afore-described generalized multiterminal source coding system with m ≥ 2 coincides with that of the centralized system for all distortions when ρ ≤ 0 and for distortions below an explicit positive threshold (depending on m) when ρ > 0. Moreover, when ρ > 0, the minimum achievable rate of generalized multiterminal source coding subject to an arbitrary positive distortion constraint d is shown to be within a finite gap (depending on m and d) from its centralized counterpart in the large I limit except for possibly the critical distortion d = 1 - ρ.
Jun Chen 0005, Yameng Chang, Jia Wang 0004, Yizhong Wang
IEEE Trans. Inf. Theory4
2019 Understanding VAEs in Fisher-Shannon Plane
abstract
In information theory, Fisher information and Shannon information (entropy) are respectively used to quantify the uncertainty associated with the distribution modeling and the uncertainty in specifying the outcome of given variables. These two quantities are complementary and are jointly applied to information behavior analysis in most cases. The uncertainty property in information asserts a fundamental trade-off between Fisher information and Shannon information, which enlightens us the relationship between the encoder and the decoder in variational auto-encoders (VAEs). In this paper, we investigate VAEs in the Fisher-Shannon plane, and demonstrate that the representation learning and the log-likelihood estimation are intrinsically related to these two information quantities. Through extensive qualitative and quantitative experiments, we provide with a better comprehension of VAEs in tasks such as high-resolution reconstruction, and representation learning in the perspective of Fisher information and Shannon information. We further propose a variant of VAEs, termed as Fisher auto-encoder (FAE), for practical needs to balance Fisher information and Shannon information. Our experimental results have demonstrated its promise in improving the reconstruction accuracy and avoiding the noninformative latent code as occurred in previous works.
Huangjie Zheng, Jiangchao Yao, Ya Zhang 0002, Ivor W. Tsang, Jia Wang 0004
AAAI5
2018 Symmetric Generalized Gaussian Multiterminal Source Coding
abstract
Consider a generalized multiterminal source coding system, where (ℓ : m) encoders, each observing a distinct size-m subset of ℓ(ℓ ≥ 2) zero-mean unit-variance symmetrically correlated Gaussian sources with correlation coefficient ρ, compress their observations in such a way that a joint decoder can reconstruct the sources within a prescribed mean squared error distortion based on the compressed data. The optimal rate-distortion performance of this system was previously known only for the two extreme cases m = ℓ (the centralized case) and m = 1 (the distributed case), and except when ρ = 0, the centralized system can achieve strictly lower compression rates than the distributed system under all non-trivial distortion constraints. Somewhat surprisingly, it is established in the present paper that the optimal rate-distortion performance of the afore-described generalized multiterminal source coding system with m ≥ 2 coincides with that of the centralized system for all distortions when ρ ≤ 0 and for distortions below an explicit positive threshold (depending on m) when ρ > 0.
Jun Chen 0005, Yameng Chang, Jia Wang 0004, Yizhong Wang
ISIT4
2018 Hardware Trojan Detection in Third-Party Digital Intellectual Property Cores by Multilevel Feature Analysis
abstract
In modern integrated circuit (IC) designs, intellectual property (IP) cores are often outsourced and designed by third-party vendors, resulting in the partial relinquishment of the control over the IC design flow. Thus, reliable verifications are required to mitigate the threat of hardware Trojans (HTs) which may be inserted into IP cores by malicious vendors. Existing trustiness verification methods cannot take the merit of high efficiency and accuracy at the same time. In this paper, we propose a multilevel fast trustiness verification framework based on feature analysis to detect HTs in third-party digital IP cores. The proposed framework combines flip-flop level and combinational logic level feature analysis to achieve both high efficiency and accuracy. Experimental results demonstrate that both explicitly and implicitly triggered HTs can be detected in very short time with a negligible false positive rate. More importantly, our framework has the unique advantage of being scalable to defend against future and stealthier HTs by adding new features into the framework.
Xiaoming Chen 0003, Qiaoyi Liu, Jia Wang 0004, Qiang Xu 0001, Yu Wang 0002, Yongpan Liu, Huazhong Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2017 Portable information security display system via Spatial PsychoVisual Modulation
abstract
With the rapid development of visual media, people prefer to pay more attention to privacy protection in public situations. Currently, most existing researches on information security such as cryptography and steganography mainly concern transmission and yet little has been done to keep the information displayed on screens from reaching eyes of the bystanders. At the same time, the reporter just stands in front of the screen during traditional meetings. Limited time, screen area and report forms, which inevitably leads to limited information. As a result, we design a portable screen for assisting the reporter to present any information in a new form. In some public occasions, the reporter can show private content with important information to authorized audience while the others can not see that. In this paper, we propose a new Spatial PsychoVisual Modulation (SPVM) based solution to the privacy problem. This system uses two synchronized projectors with linear polarization filters and polarization glasses, and a camera with linear polarization filter the metallic screen. It can guarantee the system shows private information synchronously. We have implemented the system and experimental results demonstrate the effectiveness and robustness of the proposed information security display system.
Guangtao Zhai, Jia Wang 0004, Ke Gu 0001
VCIP3
2016 Quality assessment for dual-view display system
abstract
Spatial psychovisual modulation (SPVM) is a new information display technology, which aims to generate multiple visual percepts for different viewers on a single display simultaneously. After the proposal of SPVM, lots of efforts have been made and several applications (i.e., dual-view display system) have been implemented based on this technology. The dual-view display (DVD) system is considered as an effective digital image hiding system based on SPVM theory, but little work has been dedicated to the perceptual quality assessment of DVD system. Up to now, there is no clear and standard method to evaluate the performance of the dual-view display system. It is important for the viewers to see a clear and non-aliasing image when they are front of the screen. Therefore, in this paper, we will build a DVD database and carry out a subjective experiment to evaluate the performance of the DVD system, and then we investigate and analyze the performance of prevailing no-reference (NR) image quality metrics on the particular DVD system. We have a sufficient belief that this paper can supply the guideline for the performance on the DVD system and serve as a good testing bed for future research of SPVM technology.
Yuanchun Chen, Guangtao Zhai, Ke Gu 0001, Jia Wang 0004, Zhongpai Gao, Yucheng Zhu
VCIP5
2016 Color guided thermal image super resolution
abstract
IR imaging is playing an important role in industry these days. However, the IR cameras only support low resolution for technical reasons. In this paper, we propose to use the visible camera as a guidance to super resolution the IR images. We setup a prototype of the IR-Color multi-sensor imaging system and construct a dataset including videos collected from different scenes for further research. Then we present our color guided algorithm which is suitable for this kind of super resolution problem. We test our algorithm on our dataset. The experiment results indicate that our algorithm performs well, and avoids the over-texture problems that traditional solutions suffer from. In further work, we will apply our multi-sensor system to more applications in computer vision.
Guangtao Zhai, Jia Wang 0004, Chunjia Hu, Yuanchun Chen
VCIP3
2015 FASTrust: Feature analysis for third-party IP trust verification
abstract
Third-party intellectual property (3PIP) cores are widely used in integrated circuit designs. It is essential and important to ensure their trustworthiness. Existing hardware trust verification techniques suffer from high computational complexity, low extensibility, and inability to detect implicitly-triggered hardware trojans (HTs). To tackle the above problems, in this paper, we present a novel 3PIP trust verification framework, named FASTrust, which conducts HT feature analysis on the flip-flop level control-data flow graph (CDFG) of the circuit. FASTrust is not only able to identify existing explicitly-triggered and implicitly-triggered HTs appeared in the literature in an efficient and effective manner, but more importantly, it also has the unique advantage of being scalable to defend against future and more stealthy HTs by adding new features to the system.
Xiaoming Chen 0003, Jie Zhang 0046, Qiaoyi Liu, Jia Wang 0004, Qiang Xu 0001, Yu Wang 0002, Huazhong Yang
ITC5
2015 Hierarchical video summarization with loitering indication
abstract
In this paper, a hierarchical and informative summarization framework is proposed, which facilitates rapid video browsing. Moreover, a method for loitering detection is exploited to indicate potential abnormal behaviors. The hierarchical framework includes two levels: a holistic-level and an object-level. The holistic-level summarization provides viewers with a comprehensive and compact representation of the original video, while the object-level summarization extracts the narrative information of each object, including trajectory, direction, time, changes of appearance and indication of the loitering behavior. The two summarizations are formulated as two different energy minimization problems, which are solved by the proposed heuristic algorithms. Our framework is evaluated on two publicly available datasets. Experimental results demonstrate that the proposed method performs favourably in providing holistic- and object-level information, fast browsing, and loitering detection.
Ruipeng Lu, Hua Yang 0001, Ji Zhu 0002, Shuang Wu 0001, Jia Wang 0004, David Bull 0001
VCIP5
2014 Large Discriminative Structured Set Prediction Modeling With Max-Margin Markov Network for Lossless Image Coding
abstract
Inherent statistical correlation for context-based prediction and structural interdependencies for local coherence is not fully exploited in existing lossless image coding schemes. This paper proposes a novel prediction model where the optimal correlated prediction for a set of pixels is obtained in the sense of the least code length. It not only exploits the spatial statistical correlations for the optimal prediction directly based on 2D contexts, but also formulates the data-driven structural interdependencies to make the prediction error coherent with the underlying probability distribution for coding. Under the joint constraints for local coherence, max-margin Markov networks are incorporated to combine support vector machines structurally to make max-margin estimation for a correlated region. Specifically, it aims to produce multiple predictions in the blocks with the model parameters learned in such a way that the distinction between the actual pixel and all possible estimations is maximized. It is proved that, with the growth of sample size, the prediction error is asymptotically upper bounded by the training error under the decomposable loss function. Incorporated into the lossless image coding framework, the proposed model outperforms most prediction schemes reported.
Wenrui Dai, Hongkai Xiong, Jia Wang 0004, Yuan F. Zheng
IEEE Trans. Image Process.3
2014 Vector Gaussian Multiterminal Source Coding
abstract
We derive an outer bound of the rate region of the vector Gaussian L -terminal CEO problem by establishing a lower bound on each supporting hyperplane of the rate region. To this end, we prove a new extremal inequality by exploiting the connection between differential entropy and Fisher information as well as some fundamental estimation-theoretic inequalities. It is shown that the outer bound matches the Berger-Tung inner bound in the high-resolution regime. We then derive a lower bound on each supporting hyperplane of the rate region of the direct vector Gaussian L -terminal source coding problem by coupling it with the CEO problem through a limiting argument. The tightness of this lower bound in the high-resolution regime and the weak-dependence regime is also proved.
Jia Wang 0004, Jun Chen 0005
IEEE Trans. Inf. Theory1
2014 On Two-Stage Sequential Coding of Correlated Sources
abstract
We study the problem of two-stage sequential coding (TSSC), which is an extension of sequential coding of correlated sources. Let X and Y be dependent random variables. The network contains two encoders and two decoders: 1) a Y encoder with input Y; 2) an X encoder with inputs X and Y; 3) a Y decoder that reconstructs Y; and 4) an X decoder that reconstructs X. The first stage is traditional sequential coding, where the Y encoder describes Y to both the X decoder and Y decoder, and the X encoder describes X and Y to the X decoder. At the second stage, the Y encoder refines the description of Y, and the X encoder refines the description of X. The TSSC model is a theoretical abstraction of scalable video coding; here, Y and X represent successive frames of a video sequence, and the two stages together give an embedded description that allows the video to be decoded at two distinct rates. We give an inner bound on the rate distortion region for this TSSC model. The tight bound on the rate distortion region is derived when Y must be reconstructed losslessly (in the usual Shannon sense) in the second stage. We also study the minimum total rate of the TSSC model and show that the minimum total rate of one-stage sequential coding cannot be achieved at both stages for jointly Gaussian sources. This theoretical result can shed light on the rate-distortion performance behavior of scalable video coding widely noted by practitioners.
Jia Wang 0004, Xiaolin Wu 0001, Jun Sun 0005, Songyu Yu
IEEE Trans. Inf. Theory1
2013 A dual-model approach to blind quality assessment of noisy images
abstract
Physiological and psychological evidences exist that the human visual system (HVS) has different behavioral patterns under low and high noise/artifact levels. We propose in this paper a dual-model approach to blind or no-reference (NR) image quality assessment (IQA) of noisy images through differentiating near-threshold and suprathreshold noise conditions. The underlying assumption for the proposed dual-model method is that for images with low level near-threshold noise, HVS tries to gauge the strength of the noise, so the image quality can be well approximated via measuring strength of the noise. And for images with their contents overwhelmed by high level suprathreshold noise, the HVS tries to recover meaningful structure from the noisy pixels using past experiences and prior knowledge encoded into an internal generative model of the brain. So image quality is closely related to the agreement between the noisy observation and the internal generative model explainable part of the image. More specifically, under the near-threshold noise condition, a noise level estimation algorithm based on natural image statistics is used, while under suprathreshold condition, an active inference model based on the free energy principle is adopted. The near-and suprathreshold models can be seamlessly integrated through a transformation between both estimates. The proposed dual-model algorithm has been tested on additive Gaussian noise contaminated images. Experimental results and comparative studies suggest that although being a no-reference approach, the proposed algorithm has prediction accuracy comparable to some of the best full-reference (FR) IQA methods.
Guangtao Zhai, André Kaup, Jia Wang 0004, Xiaokang Yang 0001
PCS3
2013 Retina model inspired image quality assessment
abstract
We proposed in this paper a retina model based approach for image quality assessment. The retinal model is consisted of an optical modulation transfer module and an adaptive low-pass filtering module. We treat the model as a black box and design the adaptive filter using an information theoretical approach. Since the information rate of visual signals is far beyond the processing power of the human visual system, there must be an effective data reduction stage in human visual brain. Therefore, the underlying assumption for the retina model is that the retina reduces the data amount of the visual scene while retaining as much useful information as possible. For full reference image quality assessment, the original and distorted images pass through the retinal filter before some kind of distance is calculated between the images. Retina filtering can serve as a general preprocessing stage for most existing image quality metrics. We show in this paper that retina model based MSE/PSNR, though being straightforward, has already state of the art performance on several image quality databases.
Guangtao Zhai, André Kaup, Jia Wang 0004, Xiaokang Yang 0001
VCIP3
2013 On the Generalization of Natural Type Selection to Multiple Description Coding
abstract
Natural type selection was originally proposed by Zamir and Rose for universal single description coding. In this paper, we generalize this principle to universal multiple description coding (MDC). Two schemes based on random codebooks are proposed: one is of fixed distortion and the other is of fixed weight. Their operational sum-rate-distortion functions are derived, which coincide with the EGC (El Gamal-Cover) sum-rate bound if the parameters of the schemes are optimized. It is also shown that in both schemes the joint type of reconstruction codewords can be used to improve the rate-distortion (R-D) performance. Based on our theoretical results, a practical universal scheme is proposed by leveraging the MDC methods based on low-density generator matrix (LDGM) codes. The performance of this scheme is compared experimentally with the EGC bound, which shows its effectiveness.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005, Jun Chen 0005
IEEE Trans. Commun.2
2013 Gaussian Robust Sequential and Predictive Coding
abstract
We introduce two new source coding problems: robust sequential coding and robust predictive coding. For the Gauss-Markov source model with the mean squared error distortion measure, we characterize certain supporting hyperplanes of the rate region of these two coding problems. Our investigation also reveals an information-theoretic minimax theorem and the associated extremal inequalities.
Lin Song 0003, Jun Chen 0005, Jia Wang 0004, Tie Liu 0002
IEEE Trans. Inf. Theory3
2013 Vector Gaussian Two-Terminal Source Coding
abstract
We derive a lower bound on each supporting line of the rate region of the vector Gaussian two-terminal CEO problem, which is a special case of the indirect vector Gaussian two-terminal source coding problem. The key technical ingredient is a new extremal inequality. It is shown that the lower bound coincides with the Berger-Tung upper bound in the high-resolution regime. Similar results are derived for the direct vector Gaussian two-terminal source coding problem.
Jia Wang 0004, Jun Chen 0005
IEEE Trans. Inf. Theory1
2012 Gaussian robust sequential and predictive coding
abstract
We introduce two new source coding problems: robust sequential coding and robust predictive coding. For the Gauss-Markov source model, we characterize certain supporting hyperplanes of the rate region of these two coding problems. Our investigation also reveals a class of extremal inequalities and minimax theorems.
Lin Song 0003, Jun Chen 0005, Jia Wang 0004, Tie Liu 0002
ISIT3
2012 On the vector Gaussian L-terminal CEO problem
abstract
We derive an outer bound of the rate region of the vector Gaussian L-terminal CEO problem by establishing a lower bound on each supporting hyperplane of the rate region. To this end we prove a new extremal inequality by exploiting the connection between differential entropy and Fisher information as well as some fundamental estimation-theoretic inequalities. It is shown that the outer bound matches the Berger-Tung inner bound in the high-resolution regime.
Jia Wang 0004, Jun Chen 0005
ISIT1
2012 On vector Gaussian multiterminal source coding
abstract
We derive a lower bound on each supporting hyperplane of the rate region of the vector Gaussian multiterminal source coding problem by coupling it with the CEO problem through a limiting argument. The tightness of this lower bound in the high-resolution regime and the weak-dependence regime is proved.
Jia Wang 0004, Jun Chen 0005
ITW1
2012 Multiple Description Image Coding Based on Delta-Sigma Quantization With Rate-Distortion Optimization
abstract
Recently, Østergaard and Zamir revealed the connection between multiple description coding and delta-sigma quantization (DSQ). The principle has been applied to image coding, with main focus on the framework where each block is processed separately. In this brief, we propose a two-channel multiple description image coding scheme that performs inter-block processing. The source image is first rearranged into a block sequence. Then, vector DSQ is performed with a bank of noise-shaping filters. Their coefficients as well as the quantization steps are chosen by a rate-distortion optimization algorithm. A post-processing algorithm is proposed for side decoding. Experiment results show the improvement achieved by the proposed scheme in terms of both peak signal-to-noise ratio values and subjective quality.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005
IEEE Trans. Image Process.2
2011 On Natural Type Selection in Universal Multiple Description Coding
abstract
In this paper, we generalize the concept of natural type selection, initially proposed by R.Zamir et.al., to multiple description coding. We prove results that are parallel to those of adaptive single description coding. A universal multiple description coding scheme is then proposed. The scheme codes a source vector at a time, and updates its codebooks based on the joint type of the codewords reconstructed previously. Based on our theoretical results, it can be shown that the rate-distortion performance of the proposed scheme gradually improves as coding proceeds.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005
DCC2
2011 Vector Gaussian multiple description coding with individual and central distortion constraints
abstract
We characterize the rate region of vector Gaussian multiple description coding with individual and central covariance distortion constraints. Specifically, we derive a lower bound and an upper bound for each supporting hyperplane of the rate region and show that these two bounds coincide. The rate region of vector Gaussian multiple description coding with individual and central trace distortion constraints as well as its extension to the stationary Gaussian process setup is also characterized.
Jun Chen 0005, Jia Wang 0004
ISIT2
2011 On the vector Gaussian CEO problem
abstract
A lower bound on each supporting line of the rate region of the vector Gaussian CEO problem is derived. The key technical ingredient is a new extremal inequality. It is shown that the lower bound coincides with the Berger-Tung upper bound in the high-resolution regime. The application of the new bounding technique to the vector Gaussian multiterminal source coding problem is also discussed.
Jun Chen 0005, Jia Wang 0004
ISIT2
2011 On Computation of Performance Bounds of Optimal Index Assignment
abstract
Channel-optimized index assignment of source codewords is arguably the simplest way of improving transmission error resilience, while keeping the source and/or channel codes intact. But optimal design of index assignment is an instance of quadratic assignment problem (QAP), one of the hardest optimization problems in the NP-complete class. In this work we make a progress in the research of index assignment optimization. We apply some recent results of QAP research to compute the strongest lower bounds so far for channel distortion of BSC among all index assignments. The strength of the resulting lower bounds is validated by comparing them against the upper bounds produced by heuristic index assignment algorithms.
Xiaolin Wu 0001, Hans D. Mittelmann, Jia Wang 0004
IEEE Trans. Commun.4
2011 Distributed Multiple Description Video Coding on Packet Loss Channels
abstract
In this paper, we are to solve the drift problem of multiple description video coding on packet loss channels by using state-of-the-art distributed techniques. We first present an asymptotically optimal code design of multiple descriptions in the Wyner-Ziv (MDWZ) setting. Then we propose a distributed multiple description video coding (DMDVC) scheme, which performs MDWZ coding on each nonintra coded frame. Instead of the prediction loops used in traditional multiple description video coding, Slepian-Wolf based coding is used to exploit interframe correlations. A bitplane extraction scheme is proposed to improve the balance between two descriptions, so that side informations can be interchanged between the side decoders of DMDVC with negligible quality degradation, which is crucial to robust transmission over packet loss channels. Experiment results demonstrate the robustness of our scheme, especially at high packet loss rates.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005
IEEE Trans. Image Process.2
2011 On the Role of the Refinement Layer in Multiple Description Coding and Scalable Coding
abstract
We clarify the relationship among several existing achievable multiple description rate-distortion regions by investigating the role of refinement layer in multiple description coding. Specifically, we show that the refinement layer in the El Gamal-Cover (EGC) scheme and the Venkataramani-Kramer-Goyal (VKG) scheme can be removed; as a consequence, the EGC region is equivalent to the EGC* region (an antecedent version of the EGC region) while the VKG region (when specialized to the 2-description case) is equivalent to the Zhang-Berger (ZB) region. Moreover, we prove that for multiple description coding with individual and hierarchical distortion constraints, the number of layers in the VKG scheme can be significantly reduced when only certain weighted sum rates are concerned. The role of refinement layer in scalable coding (a special case of multiple description coding) is also studied.
Jia Wang 0004, Jun Chen 0005, Paul W. Cuff, Haim H. Permuter
IEEE Trans. Inf. Theory1
2010 Information Flows in Video Coding
abstract
We study information theoretical performance of common video coding methodologies at the frame level. Via an abstraction of consecutive video frames as correlated random variables, many existing video coding techniques, including the baseline of MPEG-x and H.26x, the scalable coding and the distributed video coding, can have corresponding information theoretical models. The theoretical achievable rate distortion regions have been completely solved for some systems while for others remain open. We show that the achievable rate region of sequential coding equals to that of predictive coding for Markov sources. We give a theoretical analysis of the coding efficiency of B frames in the popular hybrid video coding architecture, bringing new understanding of the current practice. We also find that distributed sequential video coding generally incurs a performance loss if the source is not Markov.
Jia Wang 0004, Xiaolin Wu 0001
DCC1
2010 On Computation of Performance Bounds of Optimal Index Assignment
abstract
Channel-optimized index assignment of source codewords is arguably the simplest way of improving transmission error resilience, while keeping the source and/or channel codes intact. But optimal design of index assignment is an instance of quadratic assignment problem (QAP), one of the hardest optimization problems in the NP-complete class. In this paper we make a progress in the research of index assignment optimization. We apply some recent results of QAP research to compute the strongest lower bounds so far for channel distortion of BSC among all index assignments. The strength of the resulting lower bounds is validated by comparing them against the upper bounds produced by heuristic index assignment algorithms.
Xiaolin Wu 0001, Hans D. Mittelmann, Jia Wang 0004
DCC4
2010 Backward Adaptive Pre/Post-Filtered DPCM with Near-Optimal Rate-Distortion Performance
abstract
In this paper, we propose two backward adaptive coding algorithms based on a recently invented pre/post-filtered DPCM (Differential Pulse-Coded Modulation) codec. The pre/post-filters and the predictor are adapted jointly. Source statistics are assumed unknown a priori. One of the algorithms is based on power spectrum estimation, and the other gradient descent. Some properties of the algorithms are analyzed. Experiment results show that both algorithms achieve near-optimal rate-distortion performance, significantly outperforming the adaptive DPCM without adaptive pre/post-filtering.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005, Rong Xie 0004
ICC2
2010 A multiple description codec based on combinatorial optimization and its application to image coding
abstract
We propose a general N-channel MDC (Multiple Description Coding) framework which can integrate the advantages of various low-dimensional MDC schemes. Given the operational rate-distortion functions of low-dimensional codecs, we show how to optimize the proposed MDC framework. We prove that the optimization problem can be reduced to a combinatorial problem which in certain cases admits solutions. We then apply the proposed optimization algorithm to multiple description image coding. Experiment results show the effectiveness of our approach.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005, Cheng Zhi
ICIP2
2010 Robust multiresolution coding with Hamming distortion measure
abstract
In multiresolution coding a source sequence is encoded into a base layer and a refinement layer. The refinement layer, constructed using a conditional codebook, is in general not decodable without the correct reception of the base layer. By relating multiresolution coding with multiple description coding, we show that it is in fact possible to construct multiresolution codes in certain ways so that the refinement layer alone can be used to reconstruct the source to achieve a nontrivial distortion. As a consequence, one can improve the robustness of the existing multiresolution coding schemes without sacrificing the efficiency. Specifically, we obtain an explicit expression of the minimum distortion achievable by the refinement layer for arbitrary finite alphabet sources with Hamming distortion measure. Experimental results show that the information-theoretic limits can be approached using practical robust multiresolution coding schemes based on low-density generator matrix codes.
Jun Chen 0005, Sorina Dumitrescu, Jia Wang 0004
ISIT4
2010 On the sum rate of vector Gaussian multiterminal source coding
abstract
Upper and lower bounds on the minimum sum rate of vector Gaussian multiterminal source coding are derived. It is shown that the two bounds coincide in the high-resolution regime and the weak-dependence regime. The two-terminal case is investigated in detail.
Jia Wang 0004, Jun Chen 0005
ISIT1
2010 Robust Multiresolution Coding
abstract
In multiresolution coding a source sequence is encoded into a base layer and a refinement layer. The refinement layer, constructed using a conditional codebook, is in general not decodable without the correct reception of the base layer. By relating multiresolution coding with multiple description coding, we show that it is in fact possible to construct multiresolution codes in certain ways so that the refinement layer alone can be used to reconstruct the source to achieve a nontrivial distortion. As a consequence, one can improve the robustness of the existing multiresolution coding schemes without sacrificing the efficiency. Specifically, we obtain an explicit expression of the minimum distortion achievable by the refinement layer for arbitrary finite alphabet sources with Hamming distortion measure. Experimental results show that the information-theoretic limits can be approached using a practical robust multiresolution coding scheme based on low-density generator matrix codes.
Jun Chen 0005, Sorina Dumitrescu, Jia Wang 0004
IEEE Trans. Commun.4
2010 On the sum rate of Gaussian multiterminal source coding: new proofs and results
abstract
We show that the lower bound on the sum rate of the direct and indirect Gaussian multiterminal source coding problems can be derived in a unified manner by exploiting the semidefinite partial order of the distortion covariance matrices associated with the minimum mean squared error (MMSE) estimation and the so-called reduced optimal linear estimation, through which an intimate connection between the lower bound and the Berger-Tung upper bound is revealed. We give a new proof of the minimum sum rate of the indirect Gaussian multiterminal source coding problem (i.e., the Gaussian CEO problem). For the direct Gaussian multiterminal source coding problem, we derive a general lower bound on the sum rate and establish a set of sufficient conditions under which the lower bound coincides with the Berger-Tung upper bound. We show that the sufficient conditions are satisfied for a class of sources and distortion constraints; in particular, they hold for arbitrary positive definite source covariance matrices in the high-resolution regime. In contrast with the existing proofs, the new method does not rely on Shannon's entropy power inequality.
Jia Wang 0004, Jun Chen 0005, Xiaolin Wu 0001
IEEE Trans. Inf. Theory1
2009 Model-Guided Adaptive Recovery of Compressive Sensing
abstract
For the new signal acquisition methodology of compressive sensing (CS) a challenge is to find a space in which the signal is sparse and hence recoverable faithfully. Given the nonstationarity of many natural signals such as images, the sparse space is varying in time or spatial domain. As such, CS recovery should be conducted in locally adaptive, signal-dependent spaces to counter the fact that the CS measurements are global and irrespective of signal structures. On the contrary existing CS reconstruction methods use a fixed set of bases (e.g., wavelets, DCT, and gradient spaces) for the entirety of a signal. To rectify this problem we propose a new model-based framework to facilitate the use of adaptive bases in CS recovery. In a case study we integrate a piecewise stationary autoregressive model into the recovery process for CS-coded images, and are able to increase the reconstruction quality by 2 ~ 7dB over existing methods. The new CS recovery framework can readily incorporate prior knowledge to boost reconstruction quality.
Xiaolin Wu 0001, Xiangjun Zhang, Jia Wang 0004
DCC3
2009 On the minimum sum rate of Gaussian multiterminal source coding: New proofs
abstract
We show that the minimum sum rate of the Gaussian multiterminal source coding problems can be derived in a unified manner by exploiting the semidefinite partial order of the distortion covariance matrices associated with the MMSE estimation and the so-called reduced optimal linear estimation. In contrast to the existing proofs, the new method does not rely on Shannon's entropy power inequality. Furthermore, this new method leads to a direct proof of the minimum sum rate of the Gaussian two-terminal source coding problem without coupling it to a Gaussian CEO problem.
Jia Wang 0004, Jun Chen 0005, Xiaolin Wu 0001
ISIT1
2008 A Novel Multiple Description Video Codec Based on Slepian-Wolf Coding
abstract
One major task in multiple description video coding is to prevent drift on packet loss channels, where transmission errors occur in each description. We propose a distributed multiple description video coding (DMDVC) scheme excluding any prediction loops. The new codec suffers from no drift problem. In the two-channel mode of symmetry side informations (SI), one side decoder can use the SI of the other without any decoding quality degradation. Thereby, DMDVC achieves high robustness on packet loss channels.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005, Peng Wang 0026, Songyu Yu
DCC2
2008 A projection method for derivation of non-Shannon-type information inequalities
abstract
In 1998, Zhang and Yeung found the first unconditional non-Shannon-type information inequality. Recently, Dougherty, Freiling and Zeger gave six new unconditional non-Shannon-type information inequalities. This work generalizes their work and provides a method to systematically derive non-Shannon-type information inequalities. An application of this method reveals new 4-variable non-Shannon-type information inequalities.
Weidong Xu, Jia Wang 0004, Jun Sun 0005
ISIT2
2007 On Multi-Stage Sequential Coding of Correlated Sources
abstract
We study the problem of multi-stage sequential coding (MSSC), which is an extension of sequential coding of correlated sources. Consider two correlated random variables X and Y to be coded in two stages. The first stage is sequential coding as referred to in the existing literature. At the second stage, the Y encoder refines the information of Y without any knowledge of X, and X encoder refines the information of X with the knowledge of Y, while all previous outputs are known at the decoder. As the sequential coding problem provides a theoretical abstraction of video coding, the MSSC model is a theoretical abstraction of scalable video coding, which is an important application of network communications. We give an achievable region for the MSSC system. The given achievable region is tight when Y is required to be reconstructed perfectly in the usual Shannon sense at the second stage. We also study the minimum total rate MSSC problem, and derive the minimum total rate for Gaussian sources. This result disproves the possibility that the minimum total rate of one stage sequential coding can be achieved at both stages even for correlated Gaussian sources. Thus we offer a theoretical explanation for the performance loss of scalable video coding widely noted by practitioners
Jia Wang 0004, Xiaolin Wu 0001, Jun Sun 0005, Songyu Yu
DCC1
2007 The Wyner-Ziv Rate-Distortion Function of Multivariate Gaussian Sources and Its Application in Distributed Video Coding
abstract
Wyner-Ziv coding is presented in this paper. It is extended to the scenario of multivariate source and side information, whose rate-distortion function is obtained by a reverse water-filling method for the joint quadratic-Gaussian case.
Peng Wang 0026, Jia Wang 0004, Songyu Yu, Erkang Chen, Xiaokang Yang 0001
DCC2
2007 Spatial Error Concealment Technique using Verge Points
abstract
This paper presents a spatial error concealment technique using "verge points" which are the high-curvature points on image surfaces. Associated with curvature information, the verge points provide important characteristics in reconstructing broken edges. In our method, only a single layer of verge points surrounding the missing area are extracted. The broken edges are reconstructed by pairing these verge points with respect to both the curvature information and their neighboring pixel values. Simulations show that with our pairing criterion, most of the object edges are recovered and our concealment technique yields satisfactory results especially in terms of subjective perception.
Yizhi Gao, Yunqiang Liu, Xiaokang Yang 0001, Jia Wang 0004
ICASSP (1)5
2007 Rate-distortion Optimized Trellis-Coded Quantization
abstract
In this paper, a rate-distortion optimized trellis-coded quantization (RDOTCQ) algorithm is presented. Based on rate-distortion criterion, the proposed algorithm improves coding efficiency by adaptively optimizing quantization levels of small signals. In addition, it has no overhead and is fully compatible with the standard decoder. The proposed algorithm can be applied to any TCQ-based systems. In particular, the algorithm has been verified on the platform of the quadtree classified and trellis coded quantized (QTCQ) wavelet image compression system. Experimental results show that the proposed algorithm has a better rate-distortion performance than the QTCQ and some other compression methods.
Jun Sun 0005, Jia Wang 0004
ICME3
2007 A Rate-distortion Based Quantization Level Adjustment Algorithm in Block-based Video Compression
abstract
In this paper, a rate-distortion based quantization level adjustment (RDQLA) algorithm is presented. Based on rate-distortion criterion, the quantization level adjustment algorithm effectively improves coding efficiency by adaptively optimizing quantization levels of the signals near the boundaries of quantization cells and adjusting quantization levels per block. The proposed algorithm can be applied in any block based image and video coding method. In addition, it has no overhead and is fully compatible with the existing compression standards. The algorithm has been verified on the platform of H.263 and H.264. Experimental results show that the proposed algorithm improves the performance substantially. It is shown that the proposed algorithm has a gain of 1 dB comparing with the newest H.264 reference code and more than 1 dB comparing with H.263 standard for high bit rates.
Jun Sun 0005, Jia Wang 0004, Xiaokang Yang 0001
ICME3
2007 Multiple Descriptions with Side Informations Also Known At the Encoder
abstract
We propose a new scheme of multiple descriptions with side information (SI). The two side decoders of the system use two different SI streams. Both SI streams are available to the central decoder and to the encoder. We give an inner bound for this system for general source and SI. The tight bound is obtained for the quadratic Gaussian case. This result is compared with our previous result of the MDWZ (multiple descriptions in the Wyner-Ziv setting) problem in which none of the side information is known at the encoder. It is shown that when side information is absent at the encoder, there is a performance loss. The proposed scheme and its achievable region have practical significance. It offers theoretical insight into the multiple description video coding (MDVC) and suggests an optimal coding strategy.
Jia Wang 0004, Xiaolin Wu 0001, Songyu Yu, Jun Sun 0005
ISIT1
2006 Object Geometry Based Error Resilient Video Coding
abstract
We present an object geometry based error resilience method. We extract object geometry through analyzing motion vector field at the encoder using the iterative self-organizing data analysis technique algorithm. The extracted object geometry, in the form of an index map, is then embedded in the bitstream. At the decoder side, this information is exploited to recover motion vectors of corrupted macroblocks. The proposed method has been validated in MPEG2 codec in a standard-compliant way. Experimental results carried out on several sequences have shown that the proposed method outperforms two conventional temporal error concealment methods by an average PSNR gain of 0.9 dB and 1.5 dB, respectively.
Yizhi Gao, Daqing Zhang 0001, Xiaokang Yang 0001, Jia Wang 0004
ICIP6
2006 Multiple Descriptions in the Wyner-Ziv Setting
abstract
We propose a new scheme of multiple descriptions in the Wyner-Ziv setting (MD-WZ). The two side decoders of MD-WZ use two different side information (SI) streams. Both SI streams are available to the central decoder, but none to the encoder. We derive an achievable region (inner bound) for this MD-WZ system for general source and SI. If the source and SI are correlated Gaussian and for quadratic distortion metric, the tight bound is obtained. Our result is an extension of Ozarow's result on multiple descriptions of Gaussian source without SI. The MD-WZ coding scheme is shown to have a property of practical significance. For symmetric case where the joint distributions of the source and the two SI are the same and the two channels are balanced, interchanging the two channels causes no performance loss for Gaussian source. Considering that the existing multi-description video coding methods suffer from the notorious drifting problem induced by channels interchange, this work lends a theoretical support to distributed multi-description video coding in the Wyner-Ziv setting
Jia Wang 0004, Xiaolin Wu 0001, Songyu Yu, Jun Sun 0005
ISIT1
2006 Transform domain transcoding from MPEG-2 to H.264 with interpolation drift-error compensation
abstract
With the increasingly extensive applications of the new emerging video coding standard H.264, it inspires an urgent need to transcode the widely available MPEG-2 compressed video to H.264 format. In this paper, we investigate the issues on transcoding MPEG-2 into H.264 in transform domain with consideration of drift error due to the mismatch of motion compensation, and propose a transform domain solution to transcode MPEG-2 into H.264. We first analyze two major kinds of drifting error resulting from the mismatch of motion compensation: interpolation error and quantization error. The former is caused by the difference between the interpolation filters adopted in these two standards, thus very unique to the transform domain transcoding from MPEG-2 to H.264. Furthermore, it is identified as the dominant factor of the video quality degradation, especially in the case of small quantization parameter at high bit-rate, by extensive experimental results. As a major contribution of this paper, the close form of interpolation error is derived from transform domain. We then proposed the transcoding scheme based on quantization error drifting compensation and Interpolation error drifting compensation. Experimental results show that the proposed transform domain transcoding scheme achieves very promising performance in terms of low computational complexity and high transcoded video quality. Most importantly, its peak signal-to-noise ratio is very close to the cascaded transcoding architecture with time-consuming decoding and recoding process in pixel domain.
Tuanjie Qian, Jun Sun 0005, Xiaokang Yang 0001, Jia Wang 0004
IEEE Trans. Circuits Syst. Video Technol.5
2003 1-D and 2-D transforms from integers to integers
abstract
Substituting a real valued linear transform with an integer-to-integer mapping has become very important in lots of applications. This paper introduces a new kind of matrix decomposition method called lifting-like factorization, which leads to a theorem: every 2/sup n/-order real matrix with determinant norm 1 can be expressed as the product of one permutation matrix and at most three unit triangular matrices. Rounding error of this method is analyzed. Realization of 2D integer transform is also studied and it is shown that a 2D integer-to-integer transform cannot be realized by performing two 1D integer transforms separately. Left and right permutation matrices are introduced to reduce rounding error and an application of this method to intDCT is discussed.
Jia Wang 0004, Jun Sun 0005, Songyu Yu
ICASSP (2)1
2002 Modified wavelet coding of arbitrarily shaped objects based on extrapolation and reflection (EAR)
Jia Wang 0004, Jun Sun 0005, Songyu Yu
VCIP1
2002 Wavelet image coding based on directional dilation
Jia Wang 0004, Songyu Yu, Jun Sun 0005
VCIP1