Jun Chen 0005

dblp:85/5901-5 · DBLP profile ↗
← Back
169ranked-venue papers
35as first author
66since 2021 · last 2026
0000-0002-8084-9332ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 51 · 10 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 49 · 14 first-author · 13 since 2021Computer networks · 34 · 8 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 2 first-author · 20 since 2021Artificial intelligence and machine learning · 21 · 2 first-author · 20 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Zero-Reference Joint Low-Light Enhancement and Deblurring via Visual Autoregressive Modeling with VLM-Derived Modulation
abstract
Real-world dark images commonly exhibit not only low visibility and contrast but also complex noise and blur, posing significant restoration challenges. Existing methods often rely on paired data or fail to model dynamic illumination and blur characteristics, leading to poor generalization. To tackle this, we propose a generative framework based on visual autoregressive (VAR) modeling, guided by perceptual priors from the vision-language model (VLM). Specifically, to supply informative conditioning cues for VAR models, we deploy an adaptive curve estimation scheme to modulate the diverse illumination based on VLM-derived visibility scores. In addition, we integrate dynamic and spatial-frequency-aware Rotary Positional Encodings (SF-RoPE) into VAR to enhance its ability to model structures degraded by blur. Furthermore, we propose a recursive phase-domain modulation strategy that mitigates blur-induced artifacts in the phase domain via bounded iterative refinement guided by VLM-assessed blur scores. Our framework is fully unsupervised and achieves state-of-the-art performance on benchmark datasets.
Wei Dong 0011, Han Zhou 0003, Junwei Lin, Jun Chen 0005
AAAI4
2026 Listening Like Humans: Semantics-Guided Noise-Robust Multimodal Speech Recognition
abstract
Severe acoustic degradation is often caused by overlapping noise, disfluencies, and environmental distortions.This phenomenon results in the dissolution of linguistic structures and the generation of unreliable ASR outputs.Inspired by human speech comprehension, we propose Speech-MLM, a novel multimodal framework that reframes ASR as semantics-guided speech reconstruction.This perspective introduces three core challenges: (C1) collapse of linguistic structure under acoustic degradation, (C2) semantic ambiguity under noise, and (C3) misalignment across modalities.To address these issues, we propose Speech-MLM, a multimodal ASR framework that integrates speech, spectrogram-derived visual cues, and textual variants to enhance robustness.It consists of: (i) Cognitive Structure Extractor that recovers prosodic structure from visualized acoustic features, (ii) Semantic Weaver that learns semantic equivalence across varied textual forms, and (iii) Retrieval-Guided Fusion Learner that unifies modalities within a shared semantic space.Experiments on multiple real-world noisy datasets demonstrate that Speech-MLM achieves an average 38.85% reduction in WER, while also attaining 98.71% BERTScore and 96.7% USE, over advanced baselines, demonstrating substantial gains in semantic robustness and generalization across domains.
Jun Chen 0005, Yian Yao, Shuxin Zhong, Kaishun Wu
ACL (1)2
2026 RetinexDual: Retinex-Based Dual Nature Approach for Generalized Ultra-high-Definition Image Restoration
Mohab Kishawy, Ali Abdellatif Hussein, Jun Chen 0005
ICPR (1)3
2026 StormMind: Disentangled Layerwise Modeling for Convective Weather Systems
abstract
Timely nowcasting is critical for public safety during fast-evolving storms, where even short delays can trigger cascading failures—as in the October 2024 Spain flash flood that claimed over 90 lives within minutes. While radar offers reliable real-time sensing of atmospheric structure, models that collapse 3D volumes into 2D slices inevitably discard vertical information essential for capturing storm growth, phase transitions, and collapse. We introduce StormMind, a physically grounded framework that forecasts convective evolution by modeling causal interactions across stratified atmospheric layers. StormMind addresses two fundamental challenges:(1) the nonlinear, asynchronous coupling between low-, mid-, and high-level processes; and (2) reflectivity uncertainty, where storms with distinct vertical structures may appear deceptively similar on radar, masking their true phase and intensity. To tackle these issues, StormMind designs: i) a Convection Dynamics Extractor that models storm evolution from two complementary perspectives—horizontal morphology, capturing the spatial organization of physical processes within individual atmospheric layers, and vertical coupling, modeling energy exchanges across layers; and ii) a Convection Manifestation Reconstructor that adaptively fuses intra- and inter-layer signals, conditioned on the evolving storm state, to infer phase transitions (e.g., initiation, intensification, dissipation). Evaluated on the large-scale 3D-NEXRAD dataset (2020–2022, U.S.), StormMind outperforms strong baselines, achieving a 14.71% gain in CSI40. In real-world deployment with the Guangzhou Meteorological Bureau (Mar–May 2025), it improves CSI40 by 9.39% and boosts early-warning accuracy (98.33%)
Jun Chen 0005, Minghui Qiu, Lin Chen 0020, Shuxin Zhong, Binghong Chen, Kaishun Wu
KDD (1)1
2026 Wandatch: Infrastructure-Free Point-to-Command with Smartwatches and Speakers
Lin Chen 0020, Yandao Huang, Minghui Qiu, Shuxin Zhong, Jun Chen 0005, Kaishun Wu
PerCom5
2026 CODA: A Continuous Online Evolve Framework for Deploying HAR Sensing Systems
abstract
In always-on HAR deployments, model accuracy erodes silently as domain shift accumulates over time. Addressing this challenge requires moving beyond one-off updates toward instance-driven adaptation from streaming data. However, continuous adaptation exposes a fundamental tension: systems must selectively learn from informative instances while actively forgetting obsolete ones under long-term, non-stationary drift. To address them, we propose CODA, a continuous online adaptation framework for mobile sensing. CODA introduces two synergistic components: (i) Cache-based Selective Assimilation, which prioritizes informative instances likely to enhance system performance under sparse supervision, and (ii) an Adaptive Temporal Retention Strategy, which enables the system to gradually forget obsolete instances as sensing conditions evolve. By treating adaptation as a principled cache evolution rather than parameter-heavy retraining, CODA maintains high accuracy without model reconfiguration. We conduct extensive evaluations on four heterogeneous datasets spanning phone, watch, and multi-sensor configurations. Results demonstrate that CODA consistently outperforms one-off adaptation under non-stationary drift, remains robust against imperfect feedback, and incurs negligible on-device latency.
Minghui Qiu, Jun Chen 0005, Lin Chen 0020, Shuxin Zhong, Yandao Huang, Lu Wang 0002, Kaishun Wu
SECON2
2026 Med2ECG: Medical-Guided BCG-To-ECG Reconstruction for Diverse Populations
abstract
Continuous ECG monitoring is vital for early detection of arrhythmias and other cardiac abnormalities—especially during sleep, when symptoms often go unnoticed—yet existing solutions remain expensive, obtrusive, and impractical for long-term daily use. Ballistocardiography (BCG)—a passive, contactless modality that captures cardiac-induced body motion—offers a compelling alternative. However, prior efforts treat ECG reconstruction as waveform regression, leading to overfitting to individual-specific or posture-dependent artifacts. Inspired by the fact that ECG and BCG reflect parallel structures in cardiac event sequences (e.g., P/QRS/T-waves vs. I/J-waves), we design Med2ECG: a structure-aligned system that reconstructs ECG from high-fidelity BCG by explicitly aligning their latent physiological events. Med2ECG incorporates three key designs: (i) Multi-Scale Feature Extractor captures hierarchical temporal dynamics, preserving clinically relevant fine-grained features; (ii) Shared-Personalized Experts employs a Mixture-of-Experts (MoE) to adaptively disentangle signal variations due to individual and environmental factors; (iii) Medical-Informed Strategies introduces a diagnostic-driven multi-objective loss, integrating structural alignment, morphological fidelity, and landmark-aware supervision to preserve clinically critical intervals. Experiments across public and self-collected in-hospital datasets (20 healthy individuals and 10 patients with diverse cardiovascular conditions), Med2ECG achieves > 0.92 Pearson correlation, < 15% amplitude error, and precise PR/QRS/QT/RR interval estimation within 5–20 ms—demonstrating strong generalization across subjects, postures, and environments.
Lin Chen 0020, Yandao Huang, Chenggao Li, Jun Chen 0005, Shuxin Zhong, Minghui Qiu, Chunzhen Guo, Qian Zhang 0001, Kaishun Wu
SenSys4
2026 LEAP: LLM-Enhanced E-commerce Demand Prediction under Emergent Events
Shuxin Zhong, Jun Chen 0005, Kaishun Wu
WWW4
2026 GCFed: Exploiting Gradient Correlation for Client Selection and Rate Allocation in Federated Learning
Yangyi Liu, Stefano Rini, Jun Chen 0005
IEEE Internet Things J.3
2026 Structured Grouping Collaborative Decorrelated Regularization for Model Pruning in Infrared Small Target Detection
Yonghao Li, Jun Chen 0005, Boyang Li 0007, Yulan Guo, Longguang Wang, Siyi Deng
Pattern Recognit.2
2026 On the Fundamental Limits of Integrated Sensing and Communications Under Logarithmic Loss
abstract
We study a unified information-theoretic framework for integrated sensing and communications (ISAC), applicable to both monostatic and bistatic sensing scenarios. Special attention is given to the case where the sensing receiver (Rx) is required to produce a “soft" estimate of the state sequence, with logarithmic loss serving as the performance metric. We derive lower and upper bounds on the capacity-distortion function, which delineates the fundamental tradeoff between communication rate and sensing distortion. These bounds coincide when the channel between the ISAC transmitter (Tx) and the communication Rx is degraded with respect to the channel between the ISAC Tx and the sensing Rx, or vice versa. Furthermore, we provide a complete characterization of the capacity-distortion function for an ISAC system that simultaneously transmits information over a binary-symmetric channel and senses additive Bernoulli states through another binary-symmetric channel. The Gaussian counterpart of this problem is also explored, which, together with a state-splitting trick, fully determines the capacity-distortion-power function under the squared error distortion measure.
Jun Chen 0005, Lei Yu 0003, Yonglong Li, Wuxian Shi, Yiqun Ge, Wen Tong
IEEE Trans. Commun.1
2026 Gaussian Rate-Distortion-Perception Coding and Entropy-Constrained Scalar Quantization
abstract
This paper investigates the tightness of existing bounds on the quadratic Gaussian distortion-rate-perception functions with limited common randomness and the i.i.d. output constraint, under perception measures based on the Kullback–Leibler divergence and the squared Wasserstein-2 distance. For the squared Wasserstein-2 distance-based perception measure, we improve the best-known lower bound by introducing a tunable parameter. Moreover, via the connection between rate-distortion-perception coding and entropy-constrained scalar quantization, it is revealed that all existing bounds, including the improved one, are generally not tight in the weak perception constraint regime. Our findings shed light on the information-theoretic performance limits of rate-distortion-perception coding and offer guidelines for developing practical schemes.
Liangyan Li, Jun Chen 0005, Lei Yu 0003, Zhongshan Zhang
IEEE Trans. Commun.3
2026 Channel-Aware Optimal Transport: A Theoretical Framework for Generative Communication
Xiqiang Qu, Ruibin Li, Jun Chen 0005, Lei Yu 0003, Xinbing Wang
IEEE Trans. Inf. Theory3
2025 Low-Light Image Enhancement via Generative Perceptual Priors
abstract
Although significant progress has been made in enhancing visibility, retrieving texture details, and mitigating noise in Low-Light (LL) images, the challenge persists in applying current Low-Light Image Enhancement (LLIE) methods to real-world scenarios, primarily due to the diverse illumination conditions encountered. Furthermore, the quest for generating enhancements that are visually realistic and attractive remains an underexplored realm. In response to these challenges, we present a novel LLIE framework with the guidance of Generative Perceptual Priors (GPP-LLIE) derived from vision-language models (VLMs). Specifically, we first propose a pipeline that guides VLMs to assess multiple visual attributes of the LL image and quantify the assessment to output the global and local perceptual priors. Subsequently, to incorporate these generative perceptual priors to benefit LLIE, we introduce a transformer-based backbone in the diffusion process, and develop a new layer normalization (GPP-LN) and an attention mechanism (LPP-Attn) guided by global and local perceptual priors. Extensive experiments demonstrate that our model outperforms current SOTA methods on paired LL datasets and exhibits superior generalization on real-world data.
Han Zhou 0003, Wei Dong 0011, Xiaohong Liu 0010, Yulun Zhang 0001, Guangtao Zhai, Jun Chen 0005
AAAI6
2025 A Hierarchical Geometry-Guided Transformer for Histological Subtyping of Primary Liver Cancer
abstract
Primary liver malignancies are widely recognized as the most heterogeneous and prognostically diverse cancers of the digestive system. Among these, hepatocellular carcinoma (HCC) and intrahepatic cholangiocarcinoma (ICC) emerge as the two principal histological subtypes, demonstrating significantly greater complexity in tissue morphology and cellular architecture than other common tumors. The intricate representation of features in Whole Slide Images (WSIs) encompasses abundant crucial information for liver cancer histological subtyping, regarding hierarchical pyramid structure, tumor microenvironment (TME), and geometric representation. However, recent approaches have not adequately exploited these indispensable effective descriptors, resulting in a limited understanding of histological representation and suboptimal subtyping performance. To mitigate these limitations, A hieRarchical Geometry-gUided tranSformer (ARGUS) is proposed to advance histological subtyping in liver cancer by capturing the macro-meso-micro hierarchical information within the TME. Extensive experiments on public and private cohorts demonstrate that our ARGUS achieves state-of-the-art (SOTA) performance in histological subtyping of liver cancer, which provide an effective diagnostic tool for primary liver malignancies in clinical practice. Related code will be available to public.
Anwen Lu, Yiping Jiao, Geyang Xu, Hongyi Gong, Chengfei Cai, Jun Chen 0005, Jun Xu 0005
BIBM7
2025 Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents
abstract
Large multimodal models (LMMs) have achieved impressive progress in vision-language understanding, yet they face limitations in real-world applications requiring complex reasoning over a large number of images. Existing benchmarks for multi-image question-answering are limited in scope, each question is paired with only up to 30 images, which does not fully capture the demands of large-scale retrieval tasks encountered in the real-world usages. To reduce these gaps, we introduce two document haystack benchmarks, dubbed DocHaystack and InfoHaystack, designed to evaluate LMM performance on large-scale visual document retrieval and understanding. Additionally, we propose V-RAG, a novel, vision-centric retrieval-augmented generation (RAG) framework that leverages a suite of multimodal vision encoders, each optimized for specific strengths, and a dedicated question-document relevance module. V-RAG sets a new standard, with a 9% and 11% improvement in Recall@1 on the challenging DocHaystack-1000 and InfoHaystack-1000 benchmarks, respectively, compared to the previous best baseline models. Additionally, integrating V-RAG with LMMs enables them to efficiently operate across thousands of images, yielding significant improvements on our DocHaystack and InfoHaystack benchmarks. Our code and datasets are available at https://github.com/Vision-CAIR/dochaystacks
Jun Chen 0005, Dannong Xu, Junjie Fei, Chun-Mei Feng 0001, Mohamed Elhoseiny 0001
CVPR1
2025 LITA-GS: Illumination-Agnostic Novel View Synthesis via Reference-Free 3D Gaussian Splatting and Physical Priors
abstract
Directly employing 3D Gaussian Splatting (3DGS) on images with adverse illumination conditions exhibits considerable difficulty in achieving high-quality, normally-exposed representations due to: (1) The limited Structure from Motion (SfM) points estimated in adverse illumination scenarios fail to capture sufficient scene details; (2) Without ground-truth references, the intensive information loss, significant noise, and color distortion pose substantial challenges for 3DGS to produce high-quality results; (3) Combining existing exposure correction methods with 3DGS does not achieve satisfactory performance due to their individual enhancement processes, which lead to the illumination inconsistency between enhanced images from different viewpoints. To address these issues, we propose LITA-GS, a novel illumination-agnostic novel view synthesis method via reference-free 3DGS and physical priors. Firstly, we introduce an illumination-invariant physical prior extraction pipeline. Secondly, based on the extracted robust spatial structure prior, we develop the lighting-agnostic structure rendering strategy, which facilitates the optimization of the scene structure and object appearance. Moreover, a progressive denoising module is introduced to effectively mitigate the noise within the light-invariant representation. We adopt the unsupervised strategy for the training of LITA-GS and extensive experiments demonstrate that LITA-GS surpasses the state-of-the-art (SOTA) NeRF-based method while enjoying faster inference speed and costing reduced training time. The code is released at https://github.com/LowLevelAI/LITA-GS.
Han Zhou 0003, Wei Dong 0011, Jun Chen 0005
CVPR3
2025 WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation
Zhongyu Yang, Jun Chen 0005, Dannong Xu, Junjie Fei, Xiaoqian Shen, Liangbing Zhao, Chun-Mei Feng 0001, Mohamed Elhoseiny 0001
ICCV2
2025 Source-Channel Separation Theorems for Distortion Perception Coding
abstract
It is well known that separation between lossy source coding and channel coding is asymptotically optimal under classical additive distortion measures. Recently, coding under a new class of quality considerations, often referred to as perception or realism, has attracted significant attention due to its close connection to neural generative models and semantic communications. In this work, we revisit source-channel separation under the consideration of distortion-perception. We show that when the perception quality is measured on the block level, i.e., in the strong sense, the optimality of separation still holds when common randomness is shared between the encoder and the decoder; however, separation is no longer optimal when such common randomness is not available. In contrast, when the perception quality is the average per-symbol measure, i.e., in the weak sense, the optimality of separation holds regardless of the availability of common randomness.
Chao Tian 0002, Jun Chen 0005, Krishna Narayanan 0001
ISIT2
2025 Temporal Model-Based Federated Active Medical Image Classification
Yunlu Yan, Chun-Mei Feng 0001, Yuexiang Li, Jinheng Xie, Jun Chen 0005, Mohamed Elhoseiny 0001, Kaishun Wu, Lei Zhu 0003
MICCAI (14)5
2025 AU-IQA: A Benchmark Dataset for Perceptual Quality Assessment of AI-Enhanced User-Generated Content
abstract
AI-based image enhancement techniques have been widely adopted in various visual applications, significantly improving the perceptual quality of user-generated content (UGC). However, the lack of specialized quality assessment models has become a significant limiting factor in this field, limiting user experience and hindering the advancement of enhancement methods. While perceptual quality assessment methods have shown strong performance on UGC and AIGC individually, their effectiveness on AI-enhanced UGC (AI-UGC) which blends features from both-remains largely unexplored. To address this gap, we construct AU-IQA, a benchmark dataset comprising 4,800 AI-UGC images produced by three representative enhancement types which include super-resolution, low-light enhancement, and denoising. On this dataset, we further evaluate a range of existing quality assessment models, including traditional IQA methods and large multimodal models. Finally, we provide a comprehensive analysis of how well current approaches perform in assessing the perceptual quality of AI-UGC. The access link to the AU-IQA is https://github.com/WNNGGU/AU-IQA-Dataset.
Shushi Wang, Chunyi Li 0001, Han Zhou 0003, Wei Dong 0011, Jun Chen 0005, Guangtao Zhai, Xiaohong Liu 0001
ACM Multimedia6
2025 Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning
abstract
The rapid spread of multimodal misinformation on social media has raised growing concerns, while research on video misinformation detection remains limited due to the lack of large-scale, diverse datasets. Existing methods often overfit to rigid templates and lack deep reasoning over deceptive content. To address these challenges, we introduce FakeVV, a large-scale benchmark comprising over 100,000 video-text pairs with fine-grained, interpretable annotations. In addition, we further propose Fact-R1, a novel framework that integrates deep reasoning with collaborative rule-based reinforcement learning. Fact-R1 is trained through a three-stage process: (1) misinformation long-Chain-of-Thought (CoT) instruction tuning, (2) preference alignment via Direct Preference Optimization (DPO), and (3) Group Relative Policy Optimization (GRPO) using a novel verifiable reward function. This enables Fact-R1 to exhibit emergent reasoning behaviors comparable to those observed in advanced text-based reinforcement learning systems, but in the more complex multimodal misinformation setting. Our work establishes a new paradigm for misinformation detection, bridging large-scale video understanding, reasoning-guided alignment, and interpretable verification.
Fanrui Zhang, Qiang Zhang 0051, Jun Chen 0005, Sinbadliu, Junxiong Lin, Jiahong Yan, Jiawei Liu 0001, Zhengjun Zha
NeurIPS4
2025 Guest Editorial: Rethinking the Information Identification, Representation, and Transmission Pipeline: New Approaches to Data Compression and Communication
Jun Chen 0005, Alexandros G. Dimakis, Yong Fang 0001, Ashish Khisti, Ayfer Özgür, Nir Shlezinger
IEEE J. Sel. Areas Commun.1
2025 Information Compression in the AI Era: Recent Advances and Future Challenges
abstract
This survey article focuses on the emerging connections between machine learning and data compression. While the fundamental limits of classical (lossy) data compression are well-established through rate-distortion theory, recent advancements have uncovered new theoretical analyses and application areas inspired by machine learning. We review recent works on task-based and goal-oriented compression, rate-distortion-perception theory, and compression for estimation and inference. Deep learning-based approaches have provided natural, data-driven methods for compression. Accordingly, we survey recent efforts in applying deep learning techniques to task-based or goal-oriented compression, as well as image/video compression and transmission. Additionally, we discuss the potential use of large language models for text compression. Finally, we outline future research directions in this promising field.
Jun Chen 0005, Yong Fang 0001, Ashish Khisti, Ayfer Özgür, Nir Shlezinger
IEEE J. Sel. Areas Commun.1
2025 Performance Analysis and Enhancement of P-LDPC Codes for Lossy Compression of Binary Sources
abstract
Due to the duality between source coding and channel coding, high-performance channel codes are often adopted for addressing source coding problems. In this paper, we investigate the design and analysis of protograph low-density parity-check (P-LDPC) codes in lossy source coding systems. First, we propose a novel lossy source coding architecture based on P-LDPC codes, and empirically verify that although high-performance channel codes perform exceptionally well in their intended domain, they are not inherently optimal for lossy compression. To facilitate analysis, we introduce the lossy compression protograph extrinsic information transfer (LC-PEXIT) algorithm, which aids in evaluating the rate-distortion (RD) performance of P-LDPC codes. Additionally, the LC-PEXIT algorithm allows for the prediction of key parameters, such as the prior coefficientPand the reinforcement rateR, both of which critically influence system performance. We further propose a design algorithm for lossy P-LDPC codes and provide two sets of design examples. Experimental results show strong alignment between the predicted and actual values ofPandRf, with the designed P-LDPC codes exhibiting superior RD performance compared to classical P-LDPC and state-of-the-art ultra-sparse LDPC (US-LDPC) codes. In particular, the designed P-LDPC codes over benchmark codes increases the system coding efficiency of up to approximate 70% and achieves a performance gain of 1.69 ~ 3.26 dB, making it highly effective for lossy compression applications.
Sanya Liu, Xinfeng Wu, Yi Fang 0005, Jun Chen 0005, Chen Chen 0060, Lin Zhou 0011
IEEE Trans. Commun.4
2025 Output-Constrained Lossy Source Coding With Application to Rate-Distortion-Perception Theory
abstract
The distortion-rate function of output-constrained lossy source coding with limited common randomness is analyzed for the special case of the squared error distortion measure. An explicit expression is obtained when both the source and reconstruction distributions are Gaussian. This further leads to a partial characterization of the information-theoretic limit of quadratic Gaussian rate-distortion-perception coding, with the perception measure given by either the Kullback-Leibler divergence or the squared quadratic Wasserstein distance, from which Wagner’s result for the perfect realism setting and Zhang et al.’s result for the unlimited common randomness setting can be recovered as special cases.
Liangyan Li, Jun Chen 0005, Zhongshan Zhang
IEEE Trans. Commun.3
2025 Rate-Distortion-Perception Theory for the Quadratic Wasserstein Space
abstract
We derive a single-letter characterization of the fundamental distortion-rate-perception tradeoff with limited common randomness under the squared error distortion measure and the squared Wasserstein-2 perception measure. This characterization is further shown to admit an explicit evaluation in the case of Gaussian sources. In addition, we clarify two different notions of universal representation. As a byproduct, soft-covering lemmas with respect to the Wasserstein-2 distance are established.
Xiqiang Qu, Jun Chen 0005, Lei Yu 0003, Xiangyu Xu 0002
IEEE Trans. Inf. Theory2
2025 Universal Rate-Distortion-Perception Representations for Lossy Compression
abstract
In the context of lossy compression, Blau & Michaeli [1] adopt a mathematical notion of perceptual quality and define the information rate-distortion-perception function, generalizing the classical rate-distortion tradeoff. We consider the notion of universal representations in which one may fix a rate and an encoder then vary the decoder to achieve any point within a collection of distortion and perception constraints. We prove that the corresponding information-theoretic universal rate-distortion-perception function is operationally achievable in an approximate sense. Under MSE distortion, we show that the entire distortion-perception tradeoff of a Gaussian source can be achieved by a single encoder of the same rate asymptotically. We then characterize the achievable distortion-perception region for a fixed representation in the case of arbitrary distributions, and identify conditions under which the aforementioned results continue to hold approximately. Finally, we extend our notion of universality to the case where the rate is no longer fixed and additional bits can be sent at a second stage, generalizing the classical theory of successive refinement [2] with perception constraints. This motivates the study of practical constructions that are approximately universal across the RDP tradeoff, thereby alleviating the need to design a new encoder for each objective. We provide experimental results on MNIST and SVHN suggesting that on image compression tasks, the operational tradeoffs achieved by machine learning models with a fixed encoder suffer only a small penalty when compared to their variable encoder counterparts.
Jingjing Qian, Jun Chen 0005, Ashish Khisti
IEEE Trans. Inf. Theory3
2025 Two-Component GMM Source Coding and Optimization
abstract
A protograph low-density parity-check (P-LDPC) code-based lossy coding system is proposed for compressing the Gaussian mixture model source, as a particular case of shipping transportation systems. The compression performance is benchmarked against the rate-distortion bounds for this source. Some effective methods are developed to improve both the encoder and the decoder to reduce the compression distortion. Experimental results demonstrate that the proposed methods achieve good performance while maintaining low complexity.
Dan Song 0008, Jinkai Ren, Lin Wang 0003, Huihui Wu, Jun Chen 0005, Guanrong Chen
IEEE Trans. Intell. Transp. Syst.5
2024 GLARE: Low Light Image Enhancement via Generative Latent Feature Based Codebook Retrieval
Han Zhou 0003, Wei Dong 0011, Xiaohong Liu 0001, Shuaicheng Liu, Xiongkuo Min, Guangtao Zhai, Jun Chen 0005
ECCV (48)7
2024 Rate-Limited Optimal Transport for Quantum Gaussian Observables
abstract
The rate-limited optimal transport problem is introduced for the continuous-variable quantum measurement systems in the form of output-constrained rate-distortion coding. The main coding theorem provides a single-letter characterization of the achievable rate region for lossy quantum-to-classical source coding that transforms a sufficiently large tensor product of IID continuous-variable quantum states from a quantum source to a sequence of IID samples from a classical continuous destination distribution with a prescribed distortion level. The evaluation of rate region is performed for the systems with quantum Gaussian source and Gaussian destination distribution. We establish a Gaussian observable optimality theorem for such systems and provide an analytical formulation of the rate-limited quantum-classical Wasserstein distance in the case of isotropic and one-mode Gaussian quantum systems.11The proofs of the theorems and the details of the results are provided in the extended online version [1] available at https://arxiv.org/abs/2305.10004 for further reference. This work was supported in part by NSF grants CCF 2007878 and CCF 2132815.
Hafez M. Garmaroudi, S. Sandeep Pradhan, Jun Chen 0005
ISIT3
2024 Rate-Distortion-Perception Tradeoff for Lossy Compression Using Conditional Perception Measure
abstract
This paper studies the rate-distortion-perception (RDP) tradeoff for a memoryless source model in the asymptotic limit of large block-lengths. The perception measure is based on a divergence between the distributions of the source and reconstruction sequences conditioned on the encoder output, first proposed by Mentzer et al. We consider the case when there is no shared randomness between the encoder and the decoder. For the case of discrete memoryless sources we derive a single-letter characterization of the RDP function, in contrast to the marginal-distribution metric case (introduced by Blau and Michaeli), whose RDP characterization remains open when there is no shared randomness. The achievability scheme is based on lossy source coding with a posterior reference map. For the case of continuous valued sources under squared error distortion measure and squared quadratic Wasserstein perception measure we also derive a single-letter characterization and show that a noise-adding mechanism at the decoder suffices to achieve the optimal representation. Interestingly, the RDP function characterized for the case of zero perception loss coincides with that of the marginal metric and further zero perception loss can be achieved with a 3-dB penalty in minimum distortion. Finally we specialize to the case of Gaussian sources, and derive the RDP function for Gaussian vector case and propose a waterfilling like solution. We also partially characterize the RDP function for a mixture of Gaussian vector sources.
Sadaf Salehkalaibar, Jun Chen 0005, Ashish Khisti, Wei Yu 0001
ISIT2
2024 ECMamba: Consolidating Selective State Space Model with Retinex Guidance for Efficient Multiple Exposure Correction
abstract
Exposure Correction (EC) aims to recover proper exposure conditions for images captured under over-exposure or under-exposure scenarios. While existing deep learning models have shown promising results, few have fully embedded Retinex theory into their architecture, highlighting a gap in current methodologies. Additionally, the balance between high performance and efficiency remains an under-explored problem for exposure correction task. Inspired by Mamba which demonstrates powerful and highly efficient sequence modeling, we introduce a novel framework based on \textbf{Mamba} for \textbf{E}xposure \textbf{C}orrection (\textbf{ECMamba}) with dual pathways, each dedicated to the restoration of reflectance and illumination map, respectively. Specifically, we firstly derive the Retinex theory and we train a Retinex estimator capable of mapping inputs into two intermediary spaces, each approximating the target reflectance and illumination map, respectively. This setup facilitates the refined restoration process of the subsequent \textbf{E}xposure \textbf{C}orrection \textbf{M}amba \textbf{M}odule (\textbf{ECMM}). Moreover, we develop a novel \textbf{2D S}elective \textbf{S}tate-space layer guided by \textbf{Retinex} information (\textbf{Retinex-SS2D}) as the core operator of \textbf{ECMM}. This architecture incorporates an innovative 2D scanning strategy based on deformable feature aggregation, thereby enhancing both efficiency and effectiveness. Extensive experiment results and comprehensive ablation studies demonstrate the outstanding performance and the importance of each component of our proposed ECMamba. Code is available at \url{https://github.com/LowlevelAI/ECMamba}.
Wei Dong 0011, Han Zhou 0003, Yulun Zhang 0001, Xiaohong Liu 0001, Jun Chen 0005
NeurIPS5
2024 Minimum Entropy Coupling with Bottleneck
abstract
This paper investigates a novel lossy compression framework operating under logarithmic loss, designed to handle situations where the reconstruction distribution diverges from the source distribution. This framework is especially relevant for applications that require joint compression and retrieval, and in scenarios involving distributional shifts due to processing. We show that the proposed formulation extends the classical minimum entropy coupling framework by integrating a bottleneck, allowing for controlled variability in the degree of stochasticity in the coupling. We explore the decomposition of the Minimum Entropy Coupling with Bottleneck (MEC-B) into two distinct optimization problems: Entropy-Bounded Information Maximization (EBIM) for the encoder, and Minimum Entropy Coupling (MEC) for the decoder. Through extensive analysis, we provide a greedy algorithm for EBIM with guaranteed performance, and characterize the optimal solution near functional mappings, yielding significant theoretical insights into the structural complexity of this problem. Furthermore, we illustrated the practical application of MEC-B through experiments in Markov Coding Games (MCGs) under rate limits. These games simulate a communication scenario within a Markov Decision Process, where an agent must transmit a compressed message from a sender to a receiver through its actions. Our experiments highlighted the trade-offs between MDP rewards and receiver accuracy across various compression rates, showcasing the efficacy of our method compared to conventional compression baseline.
M. Reza Ebrahimi, Jun Chen 0005, Ashish Khisti
NeurIPS2
2024 M22: A Communication-Efficient Algorithm for Federated Learning Inspired by Rate-Distortion
abstract
In federated learning (FL), the communication constraint between the remote clients and the Parameter Server (PS) is a crucial bottleneck. For this reason, model updates must be compressed so as to minimize the loss in accuracy resulting from the communication constraint. This paper proposes “M-magnitude weighted L2 distortion + 2 degrees of freedom” (M22) algorithm, a rate-distortion inspired approach to gradient compression for federated training of deep neural networks (DNNs). In particular, we propose a family of distortion measures between the original gradient and the reconstruction we referred to as “$M$-magnitude weighted$L_{2}$” distortion, and we assume that gradient updates follow an i.i.d. distribution – generalized normal or Weibull, which have two degrees of freedom. In both the distortion measure and the gradient distribution, there is one free parameter for each that can be fitted as a function of the iteration number. Given a choice of gradient distribution and distortion measure, we design the quantizer to minimize the expected distortion in gradient reconstruction. To measure the gradient compression performance under a communication constraint, we define the per-bit accuracy as the optimal improvement in accuracy that one bit of communication brings to the centralized model over the training period. Using this performance measure, we systematically benchmark the choice of gradient distribution and distortion measure. We provide substantial insights on the role of these choices and argue that significant performance improvements can be attained using such a rate-distortion inspired compressor.
Yangyi Liu, Stefano Rini, Sadaf Salehkalaibar, Jun Chen 0005
IEEE Trans. Commun.4
2024 Rate-Limited Quantum-to-Classical Optimal Transport in Finite and Continuous-Variable Quantum Systems
abstract
We consider the rate-limited quantum-to-classical optimal transport in terms of output-constrained rate-distortion coding for both finite-dimensional and continuous-variable quantum-to-classical systems with limited classical common randomness. The main coding theorem provides a single-letter characterization of the achievable rate region of a lossy quantum measurement source coding for an exact construction of the destination distribution (or the equivalent quantum state) while maintaining a threshold of distortion from the source state according to a generally defined distortion observable. The constraint on the output space fixes the output distribution to an IID predefined probability mass function. Therefore, this problem can also be viewed as information-constrained optimal transport which finds the optimal cost of transporting the source quantum state to the destination classical distribution via a quantum measurement with limited communication rate and common randomness. We develop a coding framework for continuous-variable quantum systems by employing a clipping projection and a dequantization block and using our finite-dimensional coding theorem. Moreover, for the Gaussian quantum systems, we derive an analytical solution for rate-limited Wasserstein distance of order 2, along with a Gaussian optimality theorem, showing that Gaussian measurement optimizes the rate in a system with Gaussian quantum source and Gaussian destination distribution. The results further show that in contrast to the classical Wasserstein distance of Gaussian distributions, which corresponds to an infinite transmission rate, in the Quantum Gaussian measurement system, the optimal transport is achieved with a finite transmission rate due to the inherent noise of the quantum measurement imposed by Heisenberg’s uncertainty principle.
Hafez M. Garmaroudi, S. Sandeep Pradhan, Jun Chen 0005
IEEE Trans. Inf. Theory3
2024 Rate-Distortion-Perception Tradeoff Based on the Conditional-Distribution Perception Measure
abstract
This paper studies the rate-distortion-perception (RDP) tradeoff for a memoryless source model in the asymptotic limit of large block-lengths. The perception measure is based on a divergence between the distributions of the source and reconstruction sequences conditioned on the encoder output, first proposed by Mentzer et al. We consider the case when there is no shared randomness between the encoder and the decoder and derive a single-letter characterization of the RDP function, for the case of discrete memoryless sources. This is in contrast to the marginal-distribution metric case (introduced by Blau and Michaeli), whose RDP characterization remains open when there is no shared randomness. The achievability scheme is based on lossy source coding with a posterior reference map. For the case of continuous valued sources under the squared error distortion measure and the squared quadratic Wasserstein perception measure, we also derive a single-letter characterization and show that the decoder can be restricted to a noise-adding mechanism. Interestingly, the RDP function characterized for the case of zero perception loss coincides with that of the marginal metric, and further zero perception loss can be achieved with a 3-dB penalty in minimum distortion. Finally we specialize to the case of Gaussian sources, and derive the RDP function for Gaussian vector case and propose a reverse water-filling type solution. We also partially characterize the RDP function for a mixture of Gaussian vector sources.
Sadaf Salehkalaibar, Jun Chen 0005, Ashish Khisti, Wei Yu 0001
IEEE Trans. Inf. Theory2
2024 New Proofs of Gaussian Extremal Inequalities With Applications
abstract
The conventional enhancement-and-perturbation approach to establishing Gaussian extremal inequalities is refined via a novel monotone path argument in the product probability space. This refined approach is illustrated with simplified/corrected proofs of the Liu-Viswanath extremal inequality and a vector generalization of Costa’s entropy power inequality. The power of this refinement is further demonstrated by characterizing two information-theoretic limits, namely, the capacity region of the multiple-input multiple-output (MIMO) Gaussian broadcast channel with private and common messages and the rate-distortion-equivocation function of vector Gaussian secure source coding, which have previously resisted the attack of the conventional approach.
Yinfei Xu, Jun Chen 0005, Shi Jin 0002
IEEE Trans. Inf. Theory3
2024 Graphs of Joint Types, Noninteractive Simulation, and Stronger Hypercontractivity
abstract
In this paper, we study the type graph, namely, a bipartite graph induced by a joint type. We investigate the maximum edge density of induced bipartite subgraphs of this graph having a number of vertices on each side on an exponential scale in the length$n$of the type. This can be seen as an isoperimetric problem. We provide asymptotically sharp bounds for the exponent of the maximum edge density as the length of the type goes to infinity. We also study the biclique rate region of the type graph, which is defined as the set of$(R_{1},R_{2})$such that there exists a biclique of the type graph which has respectively$2^{nR_{1}}$and$2^{nR_{2}}$vertices on the two sides. We provide asymptotically sharp bounds for the biclique rate region as well. We then discuss the connections of these results to noninteractive simulation and hypercontractivity inequalities. Furthermore, as an application of our results, a new outer bound for the zero-error capacity region of the binary adder channel is provided, which improves the previously best known bound, due to Austrin, Kaski, Koivisto, and Nederlof. Our proofs in this paper are based on the method of types and linear algebra.
Lei Yu 0003, Venkat Anantharam, Jun Chen 0005
IEEE Trans. Inf. Theory3
2023 Contrastive Semi-Supervised Learning for Underwater Image Restoration via Reliable Bank
abstract
Despite the remarkable achievement of recent underwater image restoration techniques, the lack of labeled data has become a major hurdle for further progress. In this work, we propose a mean-teacher based Semi-supervised Underwater Image Restoration (Semi-UIR) framework to incorporate the unlabeled data into network training. However, the naive mean-teacher method suffers from two main problems: (1) The consistency loss used in training might become ineffective when the teacher's prediction is wrong. (2) Using L1 distance may cause the network to overfit wrong labels, resulting in confirmation bias. To address the above problems, we first introduce a reliable bank to store the “best-ever” outputs as pseudo ground truth. To assess the quality of outputs, we conduct an empirical analysis based on the monotonicity property to select the most trustworthy NR-IQA method. Besides, in view of the confirmation bias problem, we incorporate contrastive regularization to prevent the overfitting on wrong labels. Experimental results on both full-reference and non-reference underwater benchmarks demonstrate that our algorithm has obvious improvement over SOTA methods quantitatively and qualitatively. Code has been released at https://github.com/Huang-ShiRui/Semi-UIR.
Shirui Huang, Huan Liu 0014, Jun Chen 0005, Yunsong Li 0001
CVPR4
2023 M22: Rate-Distortion Inspired Gradient Compression
abstract
In federated learning (FL), the communication constraint between the remote users and the Parameter Server (PS) is a crucial bottleneck. This paper proposes M22, a rate-distortion inspired approach to model update compression for distributed training of deep neural networks (DNNs). In particular, (i) we propose a family of distortion measures referred to as "M-magnitude weighted L2" norm, and (ii) we assume that gradient updates follow an i.i.d. distribution with two degrees of freedom – generalized normal and Weibull distributions. To measure the gradient compression performance under a communication constraint, we define the per-bit accuracy as the optimal improvement in accuracy that a bit of communication brings to the centralized model over the training period. Using this performance measure, we systematically benchmark the choice of gradient distributions and the distortion measure. We provide substantial insights on the role of these choices and argue that significant performance improvements can be attained using such a rate-distortion inspired compressor.
Yangyi Liu, Sadaf Salehkalaibar, Stefano Rini, Jun Chen 0005
ICASSP4
2023 The Cardinality Bound on the Information Bottleneck Representations is Tight
abstract
The information bottleneck (IB) method aims to find compressed representations of a variable X that retain the most relevant information about a target variable Y. We show that for a wide family of distributions – namely, when Y is generated by X through a Hamming channel, under mild conditions – the optimal IB representations require an alphabet strictly larger than that of X. This implies that, despite several recent works, the cardinality bound first identified by Witsenhausen and Wyner in 1975 is tight. At the core of our finding is the observation that the IB function in this setting is not strictly concave, similar to the deterministic case, even though the joint distribution of X and Y is of full support. Finally, we provide a complete characterization of the IB function, as well as of the optimal representations for the Hamming case.
Etam Benger, Shahab Asoodeh, Jun Chen 0005
ISIT3
2023 Rate-limited Quantum-to-Classical Optimal Transport: A Lossy Source Coding Perspective
abstract
We consider the rate-limited quantum-to-classical optimal transport in terms of output-constrained rate-distortion coding for discrete quantum measurement systems with limited classical common randomness. The main coding theorem provides the achievable rate region of a lossy measurement source coding for an exact construction of the destination distribution (or the equivalent quantum state) while maintaining a threshold of distortion from the source state according to a generally defined distortion observable. The constraint on the output space fixes the output distribution to an i.i.d. predefined probability mass function. Therefore, this problem can also be viewed as information-constrained optimal transport which finds the optimal cost of transporting the source quantum state to the destination state via an entanglement-breaking channel with limited communication rate and common randomness.1
Hafez M. Garmaroudi, S. Sandeep Pradhan, Jun Chen 0005
ISIT3
2023 On the choice of Perception Loss Function for Learned Video Compression
abstract
We study causal, low-latency, sequential video compression when the output is subjected to both a mean squared-error (MSE) distortion loss as well as a perception loss to target realism. Motivated by prior approaches, we consider two different perception loss functions (PLFs). The first, PLF-JD, considers the joint distribution (JD) of all the video frames up to the current one, while the second metric, PLF-FMD, considers the framewise marginal distributions (FMD) between the source and reconstruction. Using information theoretic analysis and deep-learning based experiments, we demonstrate that the choice of PLF can have a significant effect on the reconstruction, especially at low-bit rates. In particular, while the reconstruction based on PLF-JD can better preserve the temporal correlation across frames, it also imposes a significant penalty in distortion compared to PLF-FMD and further makes it more difficult to recover from errors made in the earlier output frames. Although the choice of PLF decisively affects reconstruction quality, we also demonstrate that it may not be essential to commit to a particular PLF during encoding and the choice of PLF can be delegated to the decoder. In particular, encoded representations generated by training a system to minimize the MSE (without requiring either PLF) can be {\em near universal} and can generate close to optimal reconstructions for either choice of PLF at the decoder. We validate our results using (one-shot) information-theoretic analysis, detailed study of the rate-distortion-perception tradeoff of the Gauss-Markov source model as well as deep-learning based experiments on moving MNIST and KTH datasets.
Sadaf Salehkalaibar, Buu Phan, Jun Chen 0005, Wei Yu 0001, Ashish Khisti
NeurIPS3
2023 Meta-Auxiliary Learning for Future Depth Prediction in Videos
abstract
We consider a new problem of future depth prediction in videos. Given a sequence of observed frames in a video, the goal is to predict the depth map of a future frame that has not been observed yet. Depth estimation plays a vital role for scene understanding and decision-making in intelligent systems. Predicting future depth maps can be valuable for autonomous vehicles to anticipate the behaviours of their surrounding objects. Our proposed model for this problem has a two-branch architecture. One branch is for the primary task of future depth prediction. The other branch is for an auxiliary task of image reconstruction. The auxiliary branch can act as a regularization. Inspired by some recent work on test-time adaption, we use the auxiliary task during testing to adapt the model to a specific test video. We also propose a novel meta-auxiliary learning that learns the model specifically for the purpose of effective test-time adaptation. Experimental results demonstrate that our proposed approach outperforms other alternative methods.
Huan Liu 0014, Zhixiang Chi, Yuanhao Yu, Yang Wang 0003, Jun Chen 0005, Jin Tang 0005
WACV5
2023 Feature Graph Convolution Network With Attentive Fusion for Large-Scale Point Clouds Semantic Segmentation
abstract
Unstructured nature of 3D point clouds in large scenes is a challenging problem to effectively learn local geometric structures for point cloud semantic segmentation. To address this issue, we proposed a Feature Graph Convolution Network with Attentive Fusion (FGC-AFNet) in this letter. Our method takes large point clouds as input and uses the Feature Graph Convoluton (FGC) module to construct a graph of the central point with its neighboring points to extract local features. Then, we reduced the number of points using Random Sampling (RS) to expand the receptive field gradually to obtain multi-level features. The network also employs a dual Attention Fusion (AF) mechanism for efficient feature aggregation. One is at different levels for semantic feature fusion, another is for narrowing the semantic feature gap between the encoder and decoder. Compared to state-of-the-art methods on the S3DIS and Toronto3D datasets, our method obtained competitive results, with an overall accuracy of 88.6% and 96.58%, and a mean intersection over union of 71.2% and 81.92% on S3DIS and Toronto3D, respectively.
Jun Chen 0005, Yiping Chen 0002, Cheng Wang 0003
IEEE Geosci. Remote. Sens. Lett.1
2023 Constrained Secrecy Capacity of Finite-Input Intersymbol Interference Wiretap Channels
abstract
We consider reliable and secure communication over intersymbol interference wiretap channels (ISI-WTCs). In particular, we first derive an achievable secure rate for ISI-WTCs without imposing any constraints on the input distribution. Afterwards, we focus on the setup where the input distribution of the ISI-WTC is constrained to be a time-invariant finite-order Markov chain. Optimizing the parameters of this Markov chain toward maximizing the achievable secure rates is a computationally intractable problem in general, and so, toward finding a local maximum, we propose an iterative algorithm that at every iteration replaces the secure rate function with a suitable surrogate function whose maximum can be found efficiently. Although the secure rates achieved in the unconstrained setup are potentially larger than the secure rates achieved in the constrained setup, the latter setup has the advantage of leading to efficient algorithms for estimating and optimizing the achievable secure rates, and also has the benefit of being the basis of efficient coding schemes.
Aria Nouri, Reza Asvadi, Jun Chen 0005, Pascal O. Vontobel
IEEE Trans. Commun.3
2023 GridDehazeNet+: An Enhanced Multi-Scale Network With Intra-Task Knowledge Transfer for Single Image Dehazing
abstract
Adverse weather conditions such as haze can deteriorate the performance of autonomous driving and intelligent transport systems. As a potential remedy, we propose an enhanced multi-scale network, dubbed GridDehazeNet+, for single image dehazing. The proposed dehazing method does not rely on the Atmosphere Scattering Model (ASM), and an explanation as to why it is not necessarily performing the dimension reduction offered by this model is provided. GridDehazeNet+ consists of three modules: pre-processing, backbone, and post-processing. The trainable pre-processing module can generate learned inputs with better diversity and more pertinent features as compared to those derived inputs produced by hand-selected pre-processing methods. The backbone module implements multi-scale estimation with two major enhancements: 1) a novel grid structure that effectively alleviates the bottleneck issue via dense connections across different scales; 2) a spatial-channel attention block that can facilitate adaptive fusion by consolidating dehazing-relevant features. The post-processing module helps to reduce the artifacts in the final output. Due to domain shift, the model trained on synthetic data may not generalize well on real data. To address this issue, we shape the distribution of synthetic data to match that of real data, and use the resulting translated data to finetune our network. We also propose a novel intra-task knowledge transfer mechanism that can memorize and take advantage of synthetic domain knowledge to assist the learning process on the translated data. Experimental results demonstrate that the proposed method outperforms the state-of-the-art on several synthetic dehazing datasets, and achieves the superior performance on real-world hazy images after finetuning.
Xiaohong Liu 0001, Zhihao Shi, Jun Chen 0005, Guangtao Zhai
IEEE Trans. Intell. Transp. Syst.4
2023 Enabling Trimap-Free Image Matting With a Frequency-Guided Saliency-Aware Network via Joint Learning
abstract
This paper presents a strategic approach to tackling trimap-free natural image matting. Specifically, to address the false detection issue of existing trimap-free matting algorithms when the foreground object is not uniquely defined, we design a novel tangled structure (TangleNet) to handle foreground detection and matting prediction simultaneously. TangleNet enables information exchange between foreground segmentation and alpha prediction, producing high-quality alpha mattes for the most salient foreground object based on RGB inputs alone. TangleNet boosts network performance with a frequency-guided attention mechanism utilizing wavelet data. Additionally, we pretrain for salient object detection to aid in the foreground segmentation. Experimental results demonstrate that TangleNet is on par with the state-of-the-art matting methods requiring additional inputs, and outperforms all previous trimap-free algorithms in terms of both qualitative and quantitative results.
Linhui Dai, Xiaohong Liu 0001, Chengqi Li, Zhihao Shi, Jun Chen 0005, Martin Brooks
IEEE Trans. Multim.6
2023 Fast Human Pose Estimation in Compressed Videos
abstract
Current approaches for human pose estimation in videos can be categorized into per-frame and warping-based methods. Both approaches have their pros and cons. For example, per-frame methods are generally more accurate, but they are often slow. Warping-based approaches are more efficient, but the performance is usually not good. To bridge the gap, in this paper, we propose a novel fast framework for human pose estimation to meet the real-time inference with controllable accuracy degradation in compressed video domain. Our approach takes advantage of the motion representation (called “motion vector”) that is readily available in a compressed video. Pose joints in a frame are obtained by directly warping the pose joints from the previous frame using the motion vectors. We also propose modules to correct possible errors introduced by the pose warping when needed. Extensive experimental results demonstrate the effectiveness of our proposed framework for accelerating the speed of top-down human pose estimation in videos.
Huan Liu 0014, Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jun Chen 0005, Jin Tang 0005
IEEE Trans. Multim.6
2022 Towards Multi-domain Single Image Dehazing via Test-time Training
abstract
Recent years have witnessed significant progress in the area of single image dehazing, thanks to the employment of deep neural networks and diverse datasets. Most of the existing methods perform well when the training and testing are conducted on a single dataset. However, they are not able to handle different types of hazy images using a dehazing model trained on a particular dataset. One possible remedy is to perform training on multiple datasets jointly. However, we observe that this training strategy tends to compromise the model performance on individual datasets. Motivated by this observation, we propose a test-time training method which leverages a helper network to assist the dehazing model in better adapting to a domain of interest. Specifically, during the test time, the helper network evaluates the quality of the dehazing results, then directs the dehazing network to improve the quality by adjusting its parameters via self-supervision. Nevertheless, the inclusion of the helper network does not automatically ensure the desired performance improvement. For this reason, a metalearning approach is employed to make the objectives of the dehazing and helper networks consistent with each other. We demonstrate the effectiveness of the proposed method by providing extensive supporting experiments.
Huan Liu 0014, Liangyan Li, Sadaf Salehkalaibar, Jun Chen 0005
CVPR5
2022 Video Frame Interpolation Transformer
abstract
Existing methods for video interpolation heavily rely on deep convolution neural networks, and thus suffer from their intrinsic limitations, such as content-agnostic kernel weights and restricted receptive field. To address these issues, we propose a Transformer-based video interpolation framework that allows content-aware aggregation weights and considers long-range dependencies with the self-attention operations. To avoid the high computational cost of global self-attention, we introduce the concept of local attention into video interpolation and extend it to the spatial-temporal domain. Furthermore, we propose a space-time separation strategy to save memory usage, which also improves performance. In addition, we develop a multi-scale frame synthesis scheme to fully realize the potential of Transformers. Extensive experiments demonstrate the proposed model performs favorably against the state-of-the-art methods both quantitatively and qualitatively on a variety of benchmark datasets. The code and models are released at https://github.com/zhshi0816/Video-Frame-Interpolation-Transformer.
Zhihao Shi, Xiangyu Xu 0002, Xiaohong Liu 0001, Jun Chen 0005, Ming-Hsuan Yang 0001
CVPR4
2022 Few-Shot Class-Incremental Learning via Entropy-Regularized Data-Free Replay
Huan Liu 0014, Li Gu, Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jun Chen 0005, Jin Tang 0005
ECCV (24)6
2022 Lossy Compression with Distribution Shift as Entropy Constrained Optimal Transport
Huan Liu 0014, Jun Chen 0005, Ashish Khisti
ICLR3
2022 On Distributed Lossy Coding of Symmetrically Correlated Gaussian Sources
abstract
A distributed lossy compression network with$L$encoders and a decoder is considered. Each encoder observes a source and sends a compressed version to the decoder. The decoder produces a joint reconstruction of target signals with the mean squared error distortion below a given threshold. It is assumed that the observed sources can be expressed as the sum of target signals and corruptive noises which are independently generated from two symmetric multivariate Gaussian distributions. The minimum compression rate of this network versus the distortion threshold is referred to as the rate-distortion function, for which an explicit lower bound is established by solving a minimization problem. Our lower bound matches the well-known Berger-Tung upper bound for some values of the distortion threshold. The asymptotic gap between the upper and lower bounds is characterized in the large$L$limit.
Siyao Zhou 0002, Sadaf Salehkalaibar, Jingjing Qian, Jun Chen 0005, Wuxian Shi, Yiqun Ge, Wen Tong
IEEE Trans. Commun.4
2022 PSCC-Net: Progressive Spatio-Channel Correlation Network for Image Manipulation Detection and Localization
abstract
To defend against manipulation of image content, such as splicing, copy-move, and removal, we develop a Progressive Spatio-Channel Correlation Network (PSCC-Net) to detect and localize image manipulations. PSCC-Net processes the image in a two-path procedure: a top-down path that extracts local and global features and a bottom-up path that detects whether the input image is manipulated, and estimates its manipulation masks at multiple scales, where each mask is conditioned on the previous one. Different from the conventional encoder-decoder and no-pooling structures, PSCC-Net leverages features at different scales with dense cross-connections to produce manipulation masks in a coarse-to-fine fashion. Moreover, a Spatio-Channel Correlation Module (SCCM) captures both spatial and channel-wise correlations in the bottom-up path, which endows features with holistic cues, enabling the network to cope with a wide range of manipulation attacks. Thanks to the light-weight backbone and progressive mechanism, PSCC-Net can process$1,080\text{P}$images at 50+FPS. Extensive experiments demonstrate the superiority of PSCC-Net over the state-of-the-art methods on both detection and localization. Codes and models are available athttps://github.com/proteus1991/PSCC-Net.
Xiaohong Liu 0001, Yaojie Liu, Jun Chen 0005, Xiaoming Liu 0002
IEEE Trans. Circuits Syst. Video Technol.3
2022 Dense Point Cloud Completion Based on Generative Adversarial Network
abstract
Point cloud completion aims to reconstruct complete point clouds from partial point clouds, which is widely used in various fields such as autonomous driving and robotics. Most existing methods are sparse point cloud completion, where the number of point clouds after completion is relatively small and the details are insufficient. This article proposes a novel end-to-end generative adversarial network-based dense point cloud completion architecture (DPCG-Net). We design two generative adversarial network (GAN)-based modules that translate point cloud completion into mapping between global feature distributions obtained by encoding partial point clouds and ground truth, respectively. The first designed generator module proposes skip connections to fully connected layer-based network for regenerating global feature and changing the global feature distribution derived from the encoder module to approximate the ground truth global feature distribution. The second proposed discriminator module divides high-dimensional global feature vectors into several smaller batches for judgment to guarantee the similarity between the regenerated global feature and the ground truth. We perform quantitative and qualitative experiments on the ShapeNet and KITTI datasets. Experiments on ShapeNet demonstrate that our model outperforms other models in cases where the lack of a large proportion of point clouds results in a large loss of spatial structure, especially when 80% of point clouds are missing. Moreover, KITTI experiments reveal that it is also valid for realistic situations. In addition, application in classification shows that the classification accuracy of point clouds completed with DPCG-Net is as high as 86.5% under the condition of 80% missing point clouds.
Ming Cheng 0002, Guoyan Li, Yiping Chen 0002, Jun Chen 0005, Cheng Wang 0003, Jonathan Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Video Frame Interpolation via Generalized Deformable Convolution
abstract
Video frame interpolation aims at synthesizing intermediate frames from nearby source frames while maintaining spatial and temporal consistencies. The existing deep-learning-based video frame interpolation methods can be roughly divided into two categories: flow-based methods and kernel-based methods. The performance of flow-based methods is often jeopardized by the inaccuracy of flow map estimation due to oversimplified motion models, while that of kernel-based methods tends to be constrained by the rigidity of kernel shape. To address these performance-limiting issues, a novel mechanism named generalized deformable convolution is proposed, which can effectively learn motion information in a data-driven manner and freely select sampling points in space-time. We further develop a new video frame interpolation method based on this mechanism. Our extensive experiments demonstrate that the new method performs favorably against the state-of-the-art, especially when dealing with complex motions. Code is available athttps://github.com/zhshi0816/GDConvNet.
Zhihao Shi, Xiaohong Liu 0001, Kangdi Shi, Linhui Dai, Jun Chen 0005
IEEE Trans. Multim.5
2022 On Optimal Power Control for Energy Harvesting Communications With Lookahead
abstract
Consider the problem of power control for an energy harvesting communication system, where the transmitter is equipped with a finite-sized rechargeable battery and is able to look ahead to observe a fixed number of future energy arrivals. An implicit characterization of the maximum average throughput over an additive white Gaussian noise channel and the associated optimal power control policy is provided via the Bellman equation under the assumption that the energy arrival process is stationary and memoryless. A more explicit characterization is obtained for the case of Bernoulli energy arrivals by means of asymptotically tight upper and lower bounds on both the maximum average throughput and the optimal power control policy. Apart from their pivotal role in deriving the desired analytical results, such bounds are highly valuable from a numerical perspective as they can be efficiently computed using convex optimization solvers.
Ali Zibaeenejad, Shengtian Yang, Jun Chen 0005
IEEE Trans. Wirel. Commun.3
2021 Type Graphs and Small-Set Expansion
abstract
In this paper, we study the type graph, namely a bipartite graph induced by a joint type. We study the maximum edge density of induced bipartite subgraphs of this graph having a number of vertices on each side on an exponential scale. This can be seen as an isoperimetric problem. We provide asymptotically sharp bounds for the exponent of the maximum edge density as the blocklength goes to infinity. We also study the biclique rate region of the type graph, which is defined as the set of ($R_{1}, R_{2}$) such that there exists a biclique of the type graph which has respectively$e^{nR_{1}}$and$e^{nR_{2}}$vertices on the two sides. We provide asymptotically sharp bounds for the biclique rate region as well. We also apply similar techniques to strengthen small-set expansion theorems.
Lei Yu 0003, Venkat Anantharam, Jun Chen 0005
ISIT3
2021 Finite-Input Intersymbol Interference Wiretap Channels
abstract
We consider reliable and secure communication over intersymbol interference wiretap channels (ISI-WTCs). In particular, we first examine the setup where the source at the input of an ISI-WTC is unconstrained and then, based on a general achievability result for arbitrary wiretap channels, we derive an achievable secure rate for this ISI-WTC. Afterwards, we examine the setup where the source at the input of an ISI-WTC is constrained to be a finite-state machine source (FSMS) of a certain order and structure. optimizing the parameters of this FSMS toward maximizing the secure rate is a computationally intractable problem in general, and so, toward finding a local maximum, we propose an iterative algorithm that at every iteration replaces the secure rate function by a suitable surrogate function whose maximum can be found efficiently.
Aria Nouri, Reza Asvadi, Jun Chen 0005, Pascal O. Vontobel
ITW3
2021 Universal Rate-Distortion-Perception Representations for Lossy Compression
abstract
In the context of lossy compression, Blau &amp; Michaeli (2019) adopt a mathematical notion of perceptual quality and define the information rate-distortion-perception function, generalizing the classical rate-distortion tradeoff. We consider the notion of universal representations in which one may fix an encoder and vary the decoder to achieve any point within a collection of distortion and perception constraints. We prove that the corresponding information-theoretic universal rate-distortion-perception function is operationally achievable in an approximate sense. Under MSE distortion, we show that the entire distortion-perception tradeoff of a Gaussian source can be achieved by a single encoder of the same rate asymptotically. We then characterize the achievable distortion-perception region for a fixed representation in the case of arbitrary distributions, and identify conditions under which the aforementioned results continue to hold approximately. This motivates the study of practical constructions that are approximately universal across the RDP tradeoff, thereby alleviating the need to design a new encoder for each objective. We provide experimental results on MNIST and SVHN suggesting that on image compression tasks, the operational tradeoffs achieved by machine learning models with a fixed encoder suffer only a small penalty when compared to their variable encoder counterparts.
Jingjing Qian, Jun Chen 0005, Ashish Khisti
NeurIPS3
2021 Decoding Polar Codes for a Generalized Gilbert-Elliott Channel With Unknown Parameter
abstract
Decoding of polar codes, a class of capacity-achieving channel codes, typically requires the perfect knowledge of channel parameter in advance. This paper aims to investigate how to decode polar codes when channel parameter is unknown. Specifically, we study a generalized Gilbert-Elliott channel model, which assumes that the channel switches between a finite number of states. On the platform of Soft CANcellation (SCAN), which is a low-complexity iterative decoding algorithm of polar codes superior to the widely-used Successive Cancellation (SC) decoder, we propose three adaptive algorithms,i.e., Sliding-Window SCAN (SWSCAN), Weighted-Window SCAN (W2SCAN), and Linear-Weighting SCAN (LWSCAN). These adaptive SCAN decoders are seeded with a coarse estimate of channel state, and after each SCAN iteration, the decoders progressively refine the estimate of channel state. Experimental results demonstrate that the proposed adaptive SCAN decoders outperform the original SCAN decoder and other competitors.
Yong Fang 0001, Jun Chen 0005
IEEE Trans. Commun.2
2021 Exploit Camera Raw Data for Video Super- Resolution via Hidden Markov Model Inference
abstract
To the best of our knowledge, the existing deep-learning-based Video Super-Resolution (VSR) methods exclusively make use of videos produced by the Image Signal Processor (ISP) of the camera system as inputs. Such methods are 1) inherently suboptimal due to information loss incurred by non-invertible operations in ISP, and 2) inconsistent with the real imaging pipeline where VSR in fact serves as a pre-processing unit of ISP. To address this issue, we propose a new VSR method that can directly exploit camera sensor data, accompanied by a carefully built Raw Video Dataset (RawVD) for training, validation, and testing. This method consists of a Successive Deep Inference (SDI) module and a reconstruction module, among others. The SDI module is designed according to the architectural principle suggested by a canonical decomposition result for Hidden Markov Model (HMM) inference; it estimates the target high-resolution frame by repeatedly performing pairwise feature fusion using deformable convolutions. The reconstruction module, built with elaborately designed Attention-based Residual Dense Blocks (ARDBs), serves the purpose of 1) refining the fused feature and 2) learning the color information needed to generate a spatial-specific transformation for accurate color correction. Extensive experiments demonstrate that owing to the informativeness of the camera raw data, the effectiveness of the network architecture, and the separation of super-resolution and color correction processes, the proposed method achieves superior VSR results compared to the state-of-the-art and can be adapted to any specific camera-ISP. Code and dataset are available at https://github.com/proteus1991/RawVSR.
Xiaohong Liu 0001, Kangdi Shi, Zhe Wang 0033, Jun Chen 0005
IEEE Trans. Image Process.4
2021 On the Optimality of the Greedy Policy for Battery Limited Energy Harvesting Communications
abstract
Consider a battery limited energy harvesting communication system with online power control. Assuming independent and identically distributed (i.i.d.) energy arrivals and the harvest-store-use architecture, it is shown that the greedy policy achieves the maximum throughput if and only if the battery capacity is below a certain positive threshold that admits a precise characterization. Simple lower and upper bounds on this threshold are established. The asymptotic relationship between the threshold and the mean of the energy arrival process is analyzed for several examples.
Ali Zibaeenejad, Yaohui Jing, Jun Chen 0005
IEEE Trans. Inf. Theory4
2021 Vector Gaussian Successive Refinement With Degraded Side Information
abstract
We investigate the problem of the successive refinement for Wyner-Ziv coding with degraded side information and obtain a complete characterization of the rate region for the quadratic vector Gaussian case. The achievability part is based on the evaluation of the Tian-Diggavi inner bound that involves Gaussian auxiliary random vectors. For the converse part, a matching outer bound is obtained with the aid of a new extremal inequality. Herein, the proof of this extremal inequality depends on the integration of the monotone path argument and the doubling trick as well as information-estimation relations.
Yinfei Xu, Xuan Guang, Jun Chen 0005
IEEE Trans. Inf. Theory4
2020 Energy-Awareness Dynamic Trajectory Planning for UAV-Enabled Data Collection in mMTC Networks
abstract
Massive machine-type communications (mMTC) is a key enabling technology for Internet of Things (IoT) services in 5G and beyond. Efficient data collection from massive machine-type communication devices (MTCDs) performing sensing tasks is an important part of the service. In this paper, we consider an unmanned aerial vehicle (UAV) being deployed to facilitate data collection from MTCDs. Taking into account the limited energy of battery-powered MTCDs, the UAV trajectory is optimized to improve the energy efficiency of data collection. By fixing the starting and ending points of the UAV trajectory, a globally optimal (GO) trajectory can be obtained, based on the assumption that the UAV's serving radius and access capacity (number of served MTCDs) are unlimited. Interestingly, it is shown that the optimal trajectory always exists as long as the UAV flying height is greater than its service radius multiplied by a constant. However, the increase of the UAV flying height deteriorates the channel, leading to reduced efficiency of energy consumption. Alternatively, a greedy dynamic (GD) trajectory optimization scheme with limited UAV service radius and access capacity is then investigated, resulting in the optimal service location of the UAV being at a lower flying height, and the energy consumption for accomplishing the data collection task being reduced. Specifically, the UAV sorts the MTCDs within its serving radius based on the distance and selects its closest serving MTCD set. In a serving MTCD set, there is an optimal hovering location that maximizes the data collection efficiency. The UAV dynamically adjusts its service set and the optimal data collection location when MTCDs finish their data transmission and exit the service set. The process continues until all the MTCDs are served and the UAV arrives at the ending point of the trajectory. Simulation results show that both the GO and GD algorithm can improve the efficiency of overall energy consumption. In particular, the online dynamic trajectory optimization scheme is less restrictive and achieves higher efficiency.
Lingfeng Shen, Ning Wang 0004, Jun Chen 0005, Xiaomin Mu, Kon Max Wong
VTC Fall3
2020 End-To-End Trainable Video Super-Resolution Based on a New Mechanism for Implicit Motion Estimation and Compensation
abstract
Video super-resolution aims at generating a high-resolution video from its low-resolution counterpart. With the rapid rise of deep learning, many recently proposed video super-resolution methods use convolutional neural networks in conjunction with explicit motion compensation to capitalize on statistical dependencies within and across low-resolution frames. Two common issues of such methods are noteworthy. Firstly, the quality of the final reconstructed HR video is often very sensitive to the accuracy of motion estimation. Secondly, the warp grid needed for motion compensation, which is specified by the two flow maps delineating pixel displacements in horizontal and vertical directions, tends to introduce additional errors and jeopardize the temporal consistency across video frames. To address these issues, we propose a novel dynamic local filter network to perform implicit motion estimation and compensation by employing, via locally connected layers, sample-specific and position-specific dynamic local filters that are tailored to the target pixels. We also propose a global refinement network based on ResBlock and autoencoder structures to exploit non-local correlations and enhance the spatial consistency of super-resolved frames. The experimental results demonstrate that the proposed method outperforms the state-of-the-art, and validate its strength in terms of local transformation handling, temporal consistency as well as edge sharpness.
Xiaohong Liu 0001, Lingshi Kong, Jiying Zhao, Jun Chen 0005
WACV5
2020 Lattice-Based Robust Distributed Source Coding for Three Correlated Sources
Sorina Dumitrescu, Dania Elzouki, Jun Chen 0005
IEEE Trans. Commun.3
2020 Joint Component Design for the JSCC System Based on DP-LDPC Codes
abstract
The joint base matrix BJof the joint source-channel coding (JSCC) system based on double protograph low-density parity-check (DP-LDPC) codes consists of four components, namely, the source code Bs, the channel code Bc, the type1 connection edge BL1and the type-2 connection edge BL2, each having a non-negligible influence on the system performance. Different from the traditional component-specific design approach, we propose a joint design and optimization algorithm based on the idea of multi-objective differential evolution (MODE). Specifically, we consider the optimization of the DP-LDPC JSCC system through joint design of three components Bs, Bc, BL1and all four components Bs, Bc, BL1, BL2, respectively. The proposed algorithm has low search complexity due to the reduction in size and element value of base matrices. The joint protograph extrinsic information transfer (JPEXIT) analyses and the simulation results demonstrate that the resulting JSCC system is free from a high error floor, requires fewer number of iterations for reaching the same bit error rate (BER) and achieves significant coding gains as compared to the state-of-the-art. Our DP-LDPC JSCC system is also shown to outperform its separation-based counterpart by a wide margin.
Sanya Liu, Lin Wang 0003, Jun Chen 0005, Shaohua Hong
IEEE Trans. Commun.3
2020 Generalized Gaussian Multiterminal Source Coding in the High-Resolution Regime
abstract
A conjectural expression of the asymptotic gap between the rate-distortion function of an arbitrary generalized Gaussian multiterminal source coding system and that of its centralized counterpart in the high-resolution regime is proposed. The validity of this expression is verified when the number of sources is no more than 3.
Xiaolan Tu, Siyao Zhou 0002, Jun Chen 0005
IEEE Trans. Commun.4
2020 Generalized Gaussian Multiterminal Source Coding: The Symmetric Case
abstract
Consider a generalized multiterminal source coding system, where (mℓ) encoders, each observing a distinct size-m subset of I (ℓ ≥ 2) zero-mean unit-variance exchangeable Gaussian sources with correlation coefficient p, compress their observations in such a way that a joint decoder can reconstruct the sources within a prescribed mean squared error distortion based on the compressed data. The optimal rate-distortion performance of this system was previously known only for the two extreme cases m = ℓ (the centralized case) and m = 1 (the distributed case), and except when ρ = 0, the centralized system can achieve strictly lower compression rates than the distributed system under all non-trivial distortion constraints. Somewhat surprisingly, it is established in the present paper that the optimal rate-distortion performance of the afore-described generalized multiterminal source coding system with m ≥ 2 coincides with that of the centralized system for all distortions when ρ ≤ 0 and for distortions below an explicit positive threshold (depending on m) when ρ > 0. Moreover, when ρ > 0, the minimum achievable rate of generalized multiterminal source coding subject to an arbitrary positive distortion constraint d is shown to be within a finite gap (depending on m and d) from its centralized counterpart in the large I limit except for possibly the critical distortion d = 1 - ρ.
Jun Chen 0005, Yameng Chang, Jia Wang 0004, Yizhong Wang
IEEE Trans. Inf. Theory1
2020 A Maximin Optimal Online Power Control Policy for Energy Harvesting Communications
abstract
A general theory of online power control for discrete-time battery limited energy harvesting communications is developed, which leads to, among other things, an explicit characterization of a maximin optimal policy. This policy only requires the knowledge of the (effective) mean of the energy arrival process and maximizes the minimum asymptotic expected average reward (with the minimization taken over all energy arrival distributions of a given (effective) mean). Moreover, it is universally near optimal and has a strictly better worst-case performance as well as a strictly improved lower multiplicative factor in comparison with the fixed fraction policy proposed by Shaviv and Özgür when the objective is to maximize the throughput over an additive white Gaussian noise channel. The competitiveness of this maximin optimal policy is also demonstrated via numerical examples.
Shengtian Yang, Jun Chen 0005
IEEE Trans. Wirel. Commun.2
2019 Capacity-Achieving Private Information Retrieval Codes with Optimal Message Size and Upload Cost
abstract
We propose a new capacity-achieving code for the private information retrieval (PIR) problem, and show that it has the minimum message size (being one less than the number of servers) and the minimum upload cost (being roughly linear in the number of messages) among a general class of capacity-achieving codes, and in particular, among all capacity-achieving linear codes. Different from existing code constructions, the proposed code is asymmetric, and this asymmetry appears to be the key factor leading to the optimal message size and the optimal upload cost. The converse results on the message size and the upload cost are obtained by a strategic analysis of the information theoretic proof of the PIR capacity, from which a set of critical properties of any capacity-achieving code in the code class of interest is extracted.
Chao Tian 0002, Hua Sun 0001, Jun Chen 0005
ICC3
2019 GridDehazeNet: Attention-Based Multi-Scale Network for Image Dehazing
abstract
We propose an end-to-end trainable Convolutional Neural Network (CNN), named GridDehazeNet, for single image dehazing. The GridDehazeNet consists of three modules: pre-processing, backbone, and post-processing. The trainable pre-processing module can generate learned inputs with better diversity and more pertinent features as compared to those derived inputs produced by hand-selected pre-processing methods. The backbone module implements a novel attention-based multi-scale estimation on a grid network, which can effectively alleviate the bottleneck issue often encountered in the conventional multi-scale approach. The post-processing module helps to reduce the artifacts in the final output. Experimental results indicate that the GridDehazeNet outperforms the state-of-the-arts on both synthetic and real-world images. The proposed hazing method does not rely on the atmosphere scattering model, and we provide an explanation as to why it is not necessarily beneficial to take advantage of the dimension reduction offered by the atmosphere scattering model for image dehazing, even if only the dehazing results on synthetic images are concerned.
Xiaohong Liu 0001, Yongrui Ma, Zhihao Shi, Jun Chen 0005
ICCV4
2019 The Optimal Power Control Policy for an Energy Harvesting System with Look-Ahead: Bernoulli Energy Arrivals
abstract
We study power control for an energy harvesting communication system with independent and identically distributed Bernoulli energy arrivals. It is assumed that the transmitter is equipped with a finite-sized rechargeable battery and is able to look ahead to observe a fixed number of future arrivals. A complete characterization is provided for the optimal power control policy that achieves the maximum long-term average throughput over an additive white Gaussian noise channel.
Ali Zibaeenejad, Jun Chen 0005
ISIT2
2019 Single Image Dehazing with a Generic Model-Agnostic Convolutional Neural Network
abstract
A simple convolutional neural network is proposed in this letter and is trained end-to-end to restore clear images from hazy inputs. The proposed network is generic and agnostic in the sense that it is not designed specifically for image dehazing and, in particular, it has no knowledge of the atmosphere scattering model. Remarkably, this network achieves record-breaking dehazing performance on several standard data sets that are synthesized using the atmosphere scattering model. This surprising finding suggests that there might be a need to rethink the predominant plug-in approach to image dehazing.
Zheng Liu 0010, Botao Xiao, Muhammad Alrabeiah, Jun Chen 0005
IEEE Signal Process. Lett.5
2019 Successive Wyner-Ziv Coding for the Binary CEO Problem Under Logarithmic Loss
abstract
The$L$-link binary Chief Executive Officer (CEO) problem under logarithmic loss is investigated in this paper. A quantization splitting technique is applied to convert the problem under consideration to a$(2L-1)$-step successive Wyner-Ziv (WZ) problem, for which a practical coding scheme is proposed. In the proposed scheme, Low-Density Generator-Matrix (LDGM) codes are used for binary quantization while Low-Density Parity-Check (LDPC) codes are used for syndrome generation; the decoder performs successive decoding based on the received syndromes and produces a soft reconstruction of the remote source. The simulation results indicate that the rate-distortion performance of the proposed scheme can approach the theoretical inner bound based on binary-symmetric test-channel models.
Mahdi Nangir, Reza Asvadi, Jun Chen 0005, Mahmoud Ahmadian-Attari, Tadashi Matsumoto 0001
IEEE Trans. Commun.3
2019 Robust Distributed Compression of Symmetrically Correlated Gaussian Sources
abstract
Consider a lossy compression system with I-distributed encoders and a centralized decoder. Each encoder compresses its observed source and forwards the compressed data to the decoder for joint reconstruction of the target signals under the mean-squared-error distortion constraint. It is assumed that the observed sources can be expressed as the sum of the target signals and the corruptive noises, which are generated independently from two symmetric multivariate Gaussian distributions. Depending on the parameters of such distributions, the rate-distortion limit of this system is characterized either completely or at least for sufficiently low distortions. The results are further extended to the robust distributed compression setting, where the outputs of a subset of encoders may also be used to produce a non-trivial reconstruction of the corresponding target signals. In particular, we obtain in the high-resolution regime a precise characterization of the minimum achievable reconstruction distortion based on the outputs of k + 1 or more encoders when every k out of all I encoders are operated collectively in the same mode that is greedy in the sense of minimizing the distortion incurred by the reconstruction of the corresponding k target signals with respect to the average rate of these k encoders.
Yizhong Wang, Jun Chen 0005
IEEE Trans. Commun.4
2019 Lattice-Based Robust Distributed Source Coding
abstract
In this paper, we propose a lattice-based robust distributed source coding system for two correlated sources and provide a detailed performance analysis under the high resolution assumption. It is shown, among other things, that, in the asymptotic regime where: 1) the side distortion approaches 0 and 2) the ratio between the central and side distortions approaches 0, our scheme is capable of achieving the information-theoretic limit of quadratic multiple description coding when the two sources are identical, whereas a variant of the random coding scheme by Chen and Berger with Gaussian codes has a performance loss of 0.5 bits relative to this limit.
Dania Elzouki, Sorina Dumitrescu, Jun Chen 0005
IEEE Trans. Inf. Theory3
2019 Capacity-Achieving Private Information Retrieval Codes With Optimal Message Size and Upload Cost
abstract
We propose a new capacity-achieving code for the private information retrieval (PIR) problem, and show that it has the minimum message size (being one less than the number of servers) and the minimum upload cost (being roughly linear in the number of messages) among a general class of capacity-achieving codes, and in particular, among all capacity-achieving linear codes. Different from existing code constructions, the proposed code is asymmetric, and this asymmetry appears to be the key factor leading to the optimal message size and the optimal upload cost. The converse results on the message size and the upload cost are obtained by an analysis of the information theoretic proof of the PIR capacity, from which a set of critical properties of any capacity-achieving code in the code class of interest is extracted. The symmetry structure of the PIR problem is then analyzed, which allows us to construct symmetric codes from asymmetric ones, yielding a meaningful bridge between the proposed code and existing ones in the literature.
Chao Tian 0002, Hua Sun 0001, Jun Chen 0005
IEEE Trans. Inf. Theory3
2019 Intrinsic Capacity
abstract
Every channel can be expressed as a convex combination of deterministic channels with each deterministic channel corresponding to one particular intrinsic state. Such convex combinations are, in general, not unique, each giving rise to a specific intrinsic-state distribution. In this paper, we study the maximum and minimum capacities of a channel when the realization of its intrinsic state is causally available at the encoder and/or the decoder. Several conclusive results are obtained for binary-input channels and binary-output channels. By-products of our investigation include a generalization of the Birkhoff-von Neumann theorem and a condition on the uselessness of causal state information at the encoder.
Shengtian Yang, Jun Chen 0005, Jian-Kang Zhang 0002
IEEE Trans. Inf. Theory3
2018 Symmetric Generalized Gaussian Multiterminal Source Coding
abstract
Consider a generalized multiterminal source coding system, where (ℓ : m) encoders, each observing a distinct size-m subset of ℓ(ℓ ≥ 2) zero-mean unit-variance symmetrically correlated Gaussian sources with correlation coefficient ρ, compress their observations in such a way that a joint decoder can reconstruct the sources within a prescribed mean squared error distortion based on the compressed data. The optimal rate-distortion performance of this system was previously known only for the two extreme cases m = ℓ (the centralized case) and m = 1 (the distributed case), and except when ρ = 0, the centralized system can achieve strictly lower compression rates than the distributed system under all non-trivial distortion constraints. Somewhat surprisingly, it is established in the present paper that the optimal rate-distortion performance of the afore-described generalized multiterminal source coding system with m ≥ 2 coincides with that of the centralized system for all distortions when ρ ≤ 0 and for distortions below an explicit positive threshold (depending on m) when ρ > 0.
Jun Chen 0005, Yameng Chang, Jia Wang 0004, Yizhong Wang
ISIT1
2018 A Shannon-Theoretic Approach to the Storage-Retrieval Tradeoff in PIR Systems
abstract
We consider the storage-retrieval rate tradeoff in private information retrieval systems using a Shannon-theoretic approach. Our focus is on the canonical two-message two-database case, for which a coding scheme based on random codebook generation, joint typicality encoding, and the binning technique is proposed. It is first shown that when the retrieval rate is kept optimal, the proposed non-linear scheme uses less storage than the optimal linear scheme. Since the other extreme point corresponding to using the minimum storage requires both messages to be retrieved, the performance through space-sharing of the two points can also be achieved. However, using the proposed scheme, further improvement can be achieved over this simple strategy. Although the random-coding based scheme has a diminishing but nonzero probability of error, the coding error can be eliminated if variable-length codes are allowed. Novel outer bounds are finally provided and used to establish the superiority of the non-linear codes over linear codes.
Chao Tian 0002, Hua Sun 0001, Jun Chen 0005
ISIT3
2018 A Geometric Property of Relative Entropy and the Universal Threshold Phenomenon for Binary-Input Channels with Noisy State Information at the Encoder
abstract
Tight lower and upper bounds on the ratio of relative entropies of two probability distributions with respect to a common third one are established, where the three distributions are collinear in the standard (n - 1)-simplex. These bounds are leveraged to analyze the capacity of an arbitrary binary-input channel with noisy causal state information (provided by a side channel) at the encoder and perfect state information at the decoder, and in particular to determine the exact universal threshold on the noise measure of the side channel, above which the capacity is the same as that with no encoder side information.
Shengtian Yang, Jun Chen 0005
ITW2
2018 Index Mapping for Bit-Error Resilient Multiple Description Lattice Vector Quantizer
abstract
In conventional multiple description coding (MDC), two descriptions of a source are generated and sent over ON/OFF channels. In this paper, we are interested in exploiting the redundancy built in MDC to additionally confer robustness against other channel errors. In particular, we consider a multiple description lattice vector quantizer (MDLVQ) whose output (a pair of side lattice points) is mapped to a pair of binary indexes and each index is sent over a binary channel. One channel is noiseless, while the other is noisy. Thus, at the decoder, one description is received error-free, while the other may carry bit errors. Then the decoder uses the error-free description as side information to improve the reconstruction. The effectiveness of the decoder in alleviating the impact of bit errors depends on the mapping -y of side lattice points to binary indexes. We propose the design of a structured bit-error resilient mapping -y. For this, the set of side lattice points is first partitioned using Voronoi regions of an appropriate coarse lattice. Next a good linear channel code is selected, each Voronoi region is assigned a coset of this channel code, and the side lattice points within each Voronoi region are mapped to binary sequences in the corresponding coset. In addition, we argue that the performance of -y is improved by assigning cosets close in Hamming distance to neighboring Voronoi regions, and propose a technique to achieve this goal. We derive a lower bound on the error correction performance of the proposed mapping -y in terms of the performance of the channel code C used in its construction. Interestingly, we prove that, as the rate of the MDLVQ grows to infinity, the mapping -y becomes as good as the code C in correcting bit errors. Simulation results show the significant superiority of the proposed index mapping versus random mappings.
Sorina Dumitrescu, Jun Chen 0005
IEEE Trans. Commun.3
2018 Analysis and Code Design for the Binary CEO Problem Under Logarithmic Loss
abstract
In this paper, we propose an efficient coding scheme for the binary Chief Executive Officer (CEO) problem under logarithmic loss criterion. Courtade and Weissman obtained the exact rate-distortion bound for a two-link binary CEO problem under this criterion. We find optimal parameters of the binary symmetric test-channel model for the encoder of each link by using the given bound. Furthermore, an efficient coding scheme based on compound low-density generator matrix (LDGM)-low-density parity-check (LDPC) codes is presented to achieve the theoretical rates. In the proposed encoding scheme, a binary quantizer using LDGM codes and a syndrome generator using LDPC codes are applied. The proposed decoder employs a sum-product algorithm and a soft estimator to produce an approximate a posteriori distribution of the source bits given the data received through both links. Our numerical examples verify a close performance of the proposed coding scheme to the theoretical bound in several cases.
Mahdi Nangir, Reza Asvadi, Mahmoud Ahmadian-Attari, Jun Chen 0005
IEEE Trans. Commun.4
2018 Caching and Delivery via Interference Elimination
abstract
We propose a new coded caching scheme where linear combinations of the file segments are cached at the users, for the cases where the number of files is no greater than the number of users. When a user requests a certain file in the delivery phase, the other file segments in the cached linear combinations can be viewed as interference. The proposed scheme combines rank-metric codes and maximum distance-separable codes to facilitate the decoding and elimination of the interference and also to simultaneously deliver useful contents to the intended users. The performance of the proposed scheme can be explicitly evaluated, and we show that it can achieve improvement over known memory-rate tradeoff achievable results in the literature in some regime; for certain special cases, the new memory-rate tradeoff points can be shown to be optimal.
Chao Tian 0002, Jun Chen 0005
IEEE Trans. Inf. Theory2
2017 Generalized Gaussian multiterminal source coding and probabilistic graphical models
abstract
The sum-rate distortion function of generalized Gaussian multiterminal source coding is shown to coincide with that of joint encoding in the high-resolution regime if and only if the source-encoder bipartite graph and the undirected graphical model (also known as Gaussian Markov network or Gaussian Markov random field) of the source distribution satisfy a certain condition.
Jun Chen 0005, Farrokh Etezadi, Ashish Khisti
ISIT1
2017 Index mapping for bit-error resilient multiple description lattice vector quantizer
abstract
This work addresses the construction of bit-error resilient multiple description lattice vector quantizers (MDLVQ) by proposing the design of a structured mapping γ of side lattice points to binary indexes. We assume that the first description is correct while the second description may carry bit errors. To design the mapping γ the set of side lattice points is first partitioned into Voronoi regions of an appropriate coarse lattice. Next a good channel code C is selected, each Voronoi region is assigned a coset of this channel code and the side lattice points within each Voronoi region are mapped to binary sequences in the corresponding coset. We derive a lower bound on the error correction performance of the mapping γ in terms of the performance of the code C and we show that, as the rate of the MDLVQ grows to i, the mapping γ becomes as good as the code C. Simulation results show the significant superiority of the proposed mapping versus random mappings.
Sorina Dumitrescu, Jun Chen 0005
ISIT3
2017 Intrinsic capacity
abstract
Every channel can be expressed as a convex combination of deterministic channels with each deterministic channel corresponding to one particular intrinsic state. Such convex combinations are in general not unique, each giving rise to a specific intrinsic-state distribution. In this paper we study the maximum and the minimum capacities of a channel when the realization of its intrinsic state is causally available at the encoder and/or the decoder. Several conclusive results are obtained for binary-input channels and binary-output channels. Byproducts of our investigation include a generalization of the Birkhoff-von Neumann theorem and a condition on the uselessness of causal state information at the encoder.
Shengtian Yang, Jun Chen 0005, Jian-Kang Zhang 0002
ISIT3
2017 A Truncated Prediction Framework for Streaming Over Erasure Channels
abstract
We propose a new coding technique for sequential transmission of a stream of Gauss-Markov sources over erasure channels under a zero decoding delay constraint. Our proposed scheme is a combination (hybrid) of predictive coding with truncated memory, and quantization-and-binning. We study the optimality of our proposed scheme using an information theoretic model. In our setup, the encoder observes a stream of source vectors that are spatially independent and identically distributed (i.i.d.) and temporally sampled from a first-order stationary Gauss-Markov process. The channel introduces an erasure burst of a certain maximum length B, starting at an arbitrary time, not known to the transmitter. The reconstruction of each source vector at the destination must be with zero delay and satisfy a quadratic distortion constraint with an average distortion of D. The decoder is not required to reconstruct those source vectors that belong to the period spanning the erasure burst and a recovery window of length W following it. We study the minimum compression rate R(B, W, D) in this setup. As our main result, we establish upper and lower bounds on the compression rate. The upper bound (achievability) is based on our hybrid scheme. It achieves significant gains over baseline schemes such as (leaky) predictive coding, memoryless binning, a separation-based scheme, and a group of pictures-based scheme. The lower bound is established by observing connection to a network source coding problem. The bounds simplify in the high resolution regime, where we provide explicit expressions whenever possible, and identify conditions when the proposed scheme is close to optimal. We finally discuss the interplay between the parameters of our burst erasure channel and the statistical channel models and explain how the bounds in the former model can be used to derive insights into the simulation results involving the latter. In particular, our proposed scheme outperforms the baseline schemes over the i.i.d. erasure channel and the Gilbert-Elliott channel, and achieves performance close to a lower bound in some regimes.
Farrokh Etezadi, Ashish Khisti, Jun Chen 0005
IEEE Trans. Inf. Theory3
2017 Matched Multiuser Gaussian Source Channel Communications via Uncoded Schemes
abstract
We investigate whether uncoded schemes are optimal for Gaussian sources on multiuser Gaussian channels. Particularly, we consider two problems: the first is to send correlated Gaussian sources on a Gaussian broadcast channel where each receiver is interested in reconstructing only one source component (or one specific linear function of the sources) under the mean squared error distortion measure; the second is to send correlated Gaussian sources on a Gaussian multiple-access channel, where each transmitter observes a noisy combination of the sources, and the receiver wishes to reconstruct the individual source components (or individual linear functions) under the mean squared error distortion measure. It is shown that when the channel parameters satisfy certain general conditions, the induced distortion tuples are on the boundary of the achievable distortion region, and thus optimal. Instead of following the conventional approach of attempting to characterize the achievable distortion region, we ask the question whether and how a match can be effectively determined. This decision problem formulation helps to circumvent the difficult optimization problem often embedded in region characterization problems, and it also leads us to focus on the critical conditions in the outer bounds that make the inequalities become equalities, which effectively decouple the overall problem into several simpler sub-problems. Optimality results previously unknown in the literature are obtained using this novel approach. Explicit and novel outer bounds are derived for the two problems as the byproducts of our investigation.
Chao Tian 0002, Jun Chen 0005, Suhas N. Diggavi, Shlomo Shamai
IEEE Trans. Inf. Theory2
2017 The Sum Rate of Vector Gaussian Multiple Description Coding With Tree-Structured Covariance Distortion Constraints
abstract
A single-letter lower bound on the sum rate of multiple description coding with tree-structured distortion constraints is established by generalizing Ozarow's celebrated converse argument through the introduction of auxiliary random variables that form a Markov tree. For the quadratic vector Gaussian case, this lower bound is shown to be achievable by an extended El Gamal-Cover scheme, yielding a complete characterization of the minimum sum rate.
Yinfei Xu, Jun Chen 0005
IEEE Trans. Inf. Theory2
2017 When is Noisy State Information at the Encoder as Useless as No Information or as Good as Noise-Free State?
abstract
For any binary-input channel with perfect state information at the decoder, if the mutual information between the noisy state observation at the encoder and the true channel state is below a positive threshold determined solely by the state distribution, then the capacity is the same as that with no encoder side information. A complementary phenomenon is revealed for the generalized probing capacity. Extensions beyond binary-input channels are developed.
Jun Chen 0005, Tsachy Weissman, Jian-Kang Zhang 0002
IEEE Trans. Inf. Theory2
2016 Orbit-entropy cones and extremal pairwise orbit-entropy inequalities
abstract
The notion of orbit-entropy cone is introduced. Specifically, orbit-entropy cone equation is the projection of equation induced by G, where equation is the closure of entropy region for n random variables and G is a permutation group over {0; 1;...; n-1}. For symmetric group Sn(with arbitrary n) and cyclic group Cn(with n ≤ 5), the associated orbit-entropy cones are shown to be characterized by the Shannon type inequalities. Moreover, the extremal pairwise relationship between orbit-entropies is determined completely for partitioned symmetric groups and partially for cyclic groups.
Jun Chen 0005, Amir Salimi, Tie Liu 0002, Chao Tian 0002
ISIT1
2016 Cyclically symmetric entropy inequalities
abstract
A cyclically symmetric entropy inequality is of the form hO≥chO′, where hOand hO′are two cyclic orbit entropy terms. A computational approach is formulated for bounding the extremal value of c̄, which is denoted by c̄O,O′. For two non-empty orbits O and O′ of a cyclic group, it is said that O dominates O′ if c̄O,O′= 1. Special attention is paid to characterizing such dominance relationship, and a graphical method is developed for that purpose.
Jun Chen 0005, Chao Tian 0002, Tie Liu 0002, Zhiqing Xiao
ISIT1
2016 Caching and delivery via interference elimination
abstract
We propose a new caching scheme where linear combinations of the file segments are cached at the users, for the scenarios where the number of files is no greater than the number of users. When a user requests a certain file in the delivery phase, the other file segments in the cached linear combinations can be viewed as interferences. The proposed scheme combines rank metric codes and maximum distance separable codes to facilitate the decoding and elimination of these interferences, and also to simultaneously deliver useful contents to the intended users. The performance of the proposed scheme can be explicitly evaluated, and we show that the new scheme can strictly improve existing tradeoff inner bounds in the literature; for certain cases, the new tradeoff points are in fact optimal.
Chao Tian 0002, Jun Chen 0005
ISIT2
2016 When is noisy state information at the encoder as useless as no information or as good as noise-free state?
abstract
For any binary-input channel with perfect state information at the decoder, if the mutual information between the noisy state observation at the encoder and the true channel state is below a positive threshold determined solely by the state distribution, then the capacity is the same as that with no encoder side information. A complementary phenomenon is revealed for a similarly defined quantity.
Jun Chen 0005, Tsachy Weissman, Jian-Kang Zhang 0002
ISIT2
2016 A Source-Channel Separation Theorem With Application to the Source Broadcast Problem
abstract
A converse method is developed for the source broadcast problem. Specifically, it is shown that the separation architecture is optimal for a variant of the source broadcast problem, and the associated source-channel separation theorem can be leveraged, via a reduction argument, to establish a necessary condition for the original problem, which unifies several existing results in the literature. Somewhat surprisingly, this method, albeit based on the source-channel separation theorem, can be used to prove the optimality of non-separation-based schemes and determine the performance limits in certain scenarios where the separation architecture is suboptimal.
Kia Khezeli, Jun Chen 0005
IEEE Trans. Inf. Theory2
2016 On the Optimal Fronthaul Compression and Decoding Strategies for Uplink Cloud Radio Access Networks
abstract
This paper investigates the compress-and-forward scheme for an uplink cloud radio access network (C-RAN) model, where multi-antenna base stations (BSs) are connected to a cloud-computing-based central processor (CP) via capacity-limited fronthaul links. The BSs compress the received signals with Wyner-Ziv coding and send the representation bits to the CP; the CP performs the decoding of all the users' messages. Under this setup, this paper makes progress toward the optimal structure of the fronthaul compression and CP decoding strategies for the compress-and-forward scheme in the C-RAN. On the CP decoding strategy design, this paper shows that under a sum fronthaul capacity constraint, a generalized successive decoding strategy of the quantization and user message codewords that allows arbitrary interleaved order at the CP achieves the same rate region as the optimal joint decoding. Furthermore, it is shown that a practical strategy of successively decoding the quantization codewords first, then the user messages, achieves the same maximum sum rate as joint decoding under individual fronthaul constraints. On the joint optimization of user transmission and BS quantization strategies, this paper shows that if the input distributions are assumed to be Gaussian, then under joint decoding, the optimal quantization scheme for maximizing the achievable rate region is Gaussian. Moreover, Gaussian input and Gaussian quantization with joint decoding achieve to within a constant gap of the capacity region of the Gaussian multiple-input multiple-output (MIMO) uplink C-RAN model. Finally, this paper addresses the computational aspect of optimizing uplink MIMO C-RAN by showing that under fixed Gaussian input, the sum rate maximization problem over the Gaussian quantization noise covariance matrices can be formulated as convex optimization problems, thereby facilitating its efficient solution.
Yinfei Xu, Wei Yu 0001, Jun Chen 0005
IEEE Trans. Inf. Theory4
2015 Price of perfection: Limited prediction for streaming over erasure channels
abstract
We study sequential transmission of Gauss-Markov sources over erasure channels under a zero decoding delay constraint. A two-stage coding scheme which can be described as a hybrid between predictive coding with limited past and quantization & binning is proposed. This scheme can achieve significant performance gains over baseline schemes in simulations involving i.i.d. erasure channels, and in certain regimes can attain performance close to a fundamental lower bound. We consider an information theoretic model for streaming that explains the weakness of baseline schemes (e.g., predictive coding, memoryless binning, etc.) and illustrates the advantage of our proposed hybrid scheme over these. Techniques from multi-terminal source coding are used to derive a new lower bound on the compression rate and identify cases when the hybrid coding scheme is close to optimal. We discuss qualitatively the interplay between the parameters of our information theoretic model and the statistical models used in simulations.
Farrokh Etezadi, Ashish Khisti, Jun Chen 0005
ISIT3
2015 Matched multiuser Gaussian source-channel communications via uncoded schemes
abstract
We investigate whether uncoded schemes are optimal for Gaussian sources on multiuser Gaussian channels. Particularly, we consider two problems: the first is to send correlated Gaussian sources on a Gaussian broadcast channel where each receiver is interested in reconstructing only one source component (or one specific linear function of the sources) under the mean squared error distortion measure; the second is to send correlated Gaussian sources on a Gaussian multiple-access channel, where each transmitter observes a noisy combination of the source, and the receiver wishes to reconstruct the individual source components (or individual linear functions) under the mean squared error distortion measure. It is shown that when the channel parameters match certain general conditions, the induced distortion tuples are on the boundary of the achievable distortion region, and thus optimal. Instead of following the conventional approach of attempting to characterize the achievable distortion region, we ask the question whether and how a match can be effectively determined. This decision problem formulation helps to circumvent the difficult optimization problem often embedded in region characterization problems, and it also leads us to focus on the critical conditions in the outer bounds that make the inequalities become equalities, which effectively decouples the overall problem into several simpler sub-problems.
Chao Tian 0002, Jun Chen 0005, Suhas N. Diggavi, Shlomo Shamai
ISIT2
2015 On the sum rate of multiple description coding with tree-structured distortion constraints
abstract
A single-letter lower bound on the sum rate of multiple description coding with tree-structured distortion constraints is established and is shown to be tight in the quadratic Gaussian case.
Yinfei Xu, Jun Chen 0005
ISIT2
2015 Optimality of gaussian fronthaul compression for uplink MIMO cloud radio access networks
abstract
This paper investigates the compress-and-forward scheme for an uplink cloud radio access network (C-RAN) model, where multi-antenna base-stations (BSs) are connected to a cloudcomputing based central processor (CP) via capacity-limited fronthaul links. The BSs perform Wyner-Ziv coding to compress and send the received signals to the CP; the CP performs either joint decoding of both the quantization codewords and the user messages at the same time, or the more practical successive decoding of the quantization codewords first, then the user messages. Under this setup, this paper makes progress toward the optimization of the fronthaul compression scheme by proving two results. First, it is shown that if the input distributions are assumed to be Gaussian, then under joint decoding, the optimal Wyner-Ziv quantization scheme for maximizing the achievable rate region is Gaussian. Second, for fixed Gaussian input, under a sum fronthaul capacity constraint and assuming Gaussian quantization, this paper shows that successive decoding and joint decoding achieve the same maximum sum rate. In this case, the optimization of Gaussian quantization noise covariance matrices for maximizing sum rate can be formulated as a convex optimization problem, therefore can be solved efficiently.
Yinfei Xu, Jun Chen 0005, Wei Yu 0001
ISIT3
2015 Outer Bounds on the Admissible Source Region for Broadcast Channels With Correlated Sources
abstract
Two outer bounds on the admissible source region for broadcast channels with correlated sources are presented: the first one is strictly tighter than the existing outer bound by Gohari and Anantharam, while the second one provides a complete characterization of the admissible source region in the case where the two sources are conditionally independent given the common part. These outer bounds are deduced from the general necessary conditions established for the lossy source broadcast problem via suitable comparisons between the virtual broadcast channel (induced by the source and the reconstructions) and the physical broadcast channel.
Kia Khezeli, Jun Chen 0005
IEEE Trans. Inf. Theory2
2015 Polar Codes for Multiple Descriptions
abstract
A polar coding scheme is proposed for the multiple description coding (MDC) problem and is shown to be able to achieve a certain rate pair on the dominant line of the achievable rate region determined by El Gamal and Cover. This scheme is an adaptation of the one developed by ŗaşoğlu et al. for the multiple access channel (MAC) to the MDC setting. The analysis of the proposed scheme contains two new ingredients: 1) a certain MDC-MAC duality and 2) an auxiliary random process that involves both the mutual information and the Bhattacharyya parameter. The decorrelation effect of the polar transform is also investigated.
Qi Shi 0005, Lin Song 0003, Chao Tian 0002, Jun Chen 0005, Sorina Dumitrescu
IEEE Trans. Inf. Theory4
2015 Broadcasting Correlated Vector Gaussians
abstract
The problem of sending two correlated vector Gaussian sources over a bandwidth-matched two-user scalar Gaussian broadcast channel is studied in this paper, where each receiver wishes to reconstruct its target source under a covariance distortion constraint. We derive a lower bound on the optimal tradeoff between the transmit power and the achievable reconstruction distortion pair. Our derivation is based on a new bounding technique which involves the introduction of appropriate remote sources. Furthermore, it is shown that this lower bound is achievable by a class of hybrid schemes for the special case, where the weak receiver wishes to reconstruct a scalar source under the mean squared error distortion constraint.
Lin Song 0003, Jun Chen 0005, Chao Tian 0002
IEEE Trans. Inf. Theory2
2015 Distributed Multilevel Diversity Coding
abstract
In distributed multilevel diversity coding, K correlated sources (each with K components) are encoded in a distributed manner such that, given the outputs from any α encoders, the decoder can reconstruct the first α components of each of the corresponding α sources. For this problem, the optimality of a multilayer Slepian-Wolf coding scheme based on binning and superposition is established when K ≤ 3. The same conclusion is shown to hold for general K under a certain symmetry condition, which generalizes a celebrated result by Yeung and Zhang.
Zhiqing Xiao, Jun Chen 0005, Yunzhou Li, Jing Wang 0001
IEEE Trans. Inf. Theory2
2014 An improved outer bound on the admissible source region for broadcast channels with correlated sources
abstract
A general necessary condition is established for the lossy source broadcast problem and is further leveraged to derive an outer bound on the admissible source region for broadcast channels with correlated sources. It is shown that the new outer bound is strictly tighter than the existing one by Gohari and Anantharam.
Kia Khezeli, Jun Chen 0005
ISIT2
2014 A source-channel separation theorem with application to the source broadcast problem
abstract
A source-channel separation theorem is established for a certain joint source-channel problem. It is shown that this separation theorem can be used in conjunction with a simple reduction argument to establish a general converse result for the source broadcast problem. Somewhat surprisingly, this converse result, albeit based on the source-channel separation theorem, can be used to prove the optimality of non-separation based schemes (e.g., hybrid coding schemes) and determine the performance limits in certain scenarios where the separation architecture is suboptimal.
Kia Khezeli, Jun Chen 0005
ISIT2
2014 Polar codes for multiple descriptions
abstract
A coding scheme based on polar codes is proposed for the multiple description problem. This scheme is an adaptation of the one developed by ŗaşoğlu et al. for the multiple access channel to the multiple description setting. Specifically, it is shown that the proposed scheme is able to achieve a certain rate pair on the dominant line of the achievable rate region determined by El Gamal and Cover. Due to the dependence between the descriptions, a complication arises in the performance analysis. To resolve this difficulty, a novel technique is introduced which converts dependent descriptions to independent descriptions without affecting the coding rates.
Qi Shi 0005, Lin Song 0003, Chao Tian 0002, Jun Chen 0005, Sorina Dumitrescu
ISIT4
2014 Capacity-Achieving Distributions for the Discrete-Time Poisson Channel - Part I: General Properties and Numerical Techniques
abstract
Despite being an accepted model for a wide variety of optical channels, few general results on optimal signalling for discrete-time Poisson (DTP) channels are known. Among the most significant is that under simultaneous peak and average constraints, the capacity-achieving distributions are discrete with a finite number of mass points. In this paper, several fundamental properties of capacity-achieving distributions for DTP channels are established. In particular, we demonstrate that all capacity-achieving distributions of the DTP channel have zero as a mass point. In the case of only a peak constraint, it is further shown that the optimal distribution always has a mass point at the maximum amplitude. Finally, under solely an average power constraint, it is shown that a finite number of mass points are insufficient to achieve the capacity. In addition to these analytical results, a numerical algorithm based on deterministic annealing is presented which can efficiently compute both the channel capacity and the associated optimal input distribution under peak and average power constraints. Numerical lower bounds based on the envelope of information rates induced by the maxentropic distributions are also shown to be extremely close to the capacity, especially in the low power regime.
Jihai Cao, Steve Hranilovic, Jun Chen 0005
IEEE Trans. Commun.3
2014 Capacity-Achieving Distributions for the Discrete-Time Poisson Channel - Part II: Binary Inputs
abstract
Discrete-time Poisson (DTP) channels exist in many scenarios including space laser communication systems which operate over long distances and which can be corrupted by reflected and scattered light. Through simulation, binary-input distributions have been observed to be optimal in many cases, however, little analytical work exists on conditions for optimality or the form of optimal signalling. In this second part, the general properties of Part I are extended to the case of DTP channels where binary-inputs are optimal. Necessary and sufficient conditions on the optimality of binary (i.e. two mass point) distributions are presented by leveraging the general properties of DTP capacity-achieving distributions. Closed-form expressions of the capacity-achieving distributions are derived in several important special cases including zero dark current and for high dark current. Numerical results are presented to elucidate the developed analytical work.
Jihai Cao, Steve Hranilovic, Jun Chen 0005
IEEE Trans. Commun.3
2014 A Lower Bound on the Sum Rate of Multiple Description Coding With Symmetric Distortion Constraints
abstract
We derive a single-letter lower bound on the minimum sum rate of multiple description coding with symmetric distortion constraints. For the binary uniform source with the erasure distortion measure or Hamming distortion measure, this lower bound can be evaluated with the aid of certain minimax theorems. A similar minimax theorem is established in the quadratic Gaussian setting, which is further leveraged to analyze the special case where the minimum sum rate subject to two levels of distortion constraints (with the second level imposed on the complete set of descriptions) is attained; in particular, we determine the minimum achievable distortions at the intermediate levels.
Lin Song 0003, Jun Chen 0005
IEEE Trans. Inf. Theory3
2014 Optimality and Approximate Optimality of Source-Channel Separation in Networks
abstract
We consider the source-channel separation architecture for lossy source coding in communication networks. It is shown that the separation approach is optimal in two general scenarios and is approximately optimal in a third scenario. The two scenarios for which separation is optimal complement each other: the first is when the memoryless sources at source nodes are arbitrarily correlated, each of which is to be reconstructed at possibly multiple destinations within certain distortions, but the channels in this network are synchronized, orthogonal, and memoryless point-to-point channels; the second is when the memoryless sources are mutually independent, each of which is to be reconstructed only at one destination within a certain distortion, but the channels are general, including multi-user channels, such as multiple access, broadcast, interference, and relay channels, possibly with feedback. The third scenario, for which we demonstrate approximate optimality of source-channel separation, generalizes the second scenario by allowing each source to be reconstructed at multiple destinations with different distortions. For this case, the loss from optimality using the separation approach can be upper-bounded when a difference distortion measure is taken, and in the special case of quadratic distortion measure, this leads to universal constant bounds.
Chao Tian 0002, Jun Chen 0005, Suhas N. Diggavi, Shlomo Shamai
IEEE Trans. Inf. Theory2
2014 Vector Gaussian Multiterminal Source Coding
abstract
We derive an outer bound of the rate region of the vector Gaussian L -terminal CEO problem by establishing a lower bound on each supporting hyperplane of the rate region. To this end, we prove a new extremal inequality by exploiting the connection between differential entropy and Fisher information as well as some fundamental estimation-theoretic inequalities. It is shown that the outer bound matches the Berger-Tung inner bound in the high-resolution regime. We then derive a lower bound on each supporting hyperplane of the rate region of the direct vector Gaussian L -terminal source coding problem by coupling it with the CEO problem through a limiting argument. The tightness of this lower bound in the high-resolution regime and the weak-dependence regime is also proved.
Jia Wang 0004, Jun Chen 0005
IEEE Trans. Inf. Theory2
2013 On symmetric multiple description coding
abstract
We derive a single-letter lower bound on the minimum sum rate of multiple description coding with symmetric distortion constraints. For the binary uniform source with the Hamming distortion measure, this lower bound can be evaluated with the aid of a certain minimax theorem. A similar minimax theorem is established in the quadratic Gaussian setting, which is further leveraged to analyze the special case where the minimum sum rate subject to two levels of distortion constraints (with the second level imposed on the complete set of descriptions) is attained; in particular, we determine the minimum achievable distortions at the intermediate levels.
Lin Song 0003, Jun Chen 0005
ITW3
2013 On the Generalization of Natural Type Selection to Multiple Description Coding
abstract
Natural type selection was originally proposed by Zamir and Rose for universal single description coding. In this paper, we generalize this principle to universal multiple description coding (MDC). Two schemes based on random codebooks are proposed: one is of fixed distortion and the other is of fixed weight. Their operational sum-rate-distortion functions are derived, which coincide with the EGC (El Gamal-Cover) sum-rate bound if the parameters of the schemes are optimized. It is also shown that in both schemes the joint type of reconstruction codewords can be used to improve the rate-distortion (R-D) performance. Based on our theoretical results, a practical universal scheme is proposed by leveraging the MDC methods based on low-density generator matrix (LDGM) codes. The performance of this scheme is compared experimentally with the EGC bound, which shows its effectiveness.
Yuhua Fan, Jia Wang 0004, Jun Sun 0005, Jun Chen 0005
IEEE Trans. Commun.4
2013 Gaussian Robust Sequential and Predictive Coding
abstract
We introduce two new source coding problems: robust sequential coding and robust predictive coding. For the Gauss-Markov source model with the mean squared error distortion measure, we characterize certain supporting hyperplanes of the rate region of these two coding problems. Our investigation also reveals an information-theoretic minimax theorem and the associated extremal inequalities.
Lin Song 0003, Jun Chen 0005, Jia Wang 0004, Tie Liu 0002
IEEE Trans. Inf. Theory2
2013 Vector Gaussian Two-Terminal Source Coding
abstract
We derive a lower bound on each supporting line of the rate region of the vector Gaussian two-terminal CEO problem, which is a special case of the indirect vector Gaussian two-terminal source coding problem. The key technical ingredient is a new extremal inequality. It is shown that the lower bound coincides with the Berger-Tung upper bound in the high-resolution regime. Similar results are derived for the direct vector Gaussian two-terminal source coding problem.
Jia Wang 0004, Jun Chen 0005
IEEE Trans. Inf. Theory2
2012 Broadcast correlated Gaussians: The vector-scalar case
abstract
The problem of sending a set of correlated Gaussian sources over a bandwidth-matched two-user scalar Gaussian broadcast channel is studied in this work, where the strong receiver wishes to reconstruct several source components (i.e., a vector source) under a distortion covariance matrix constraint and the weak receiver wishes to reconstruct a single source component (i.e., a scalar source) under the mean squared error distortion constraint. We provide a complete characterization of the optimal tradeoff between the transmit power and the achievable reconstruction distortion pair for this problem. The converse part is based on a new bounding technique which involves the introduction of an appropriate remote source. The forward part is based on a hybrid scheme where the digital portion uses dirty paper channel code and Wyner-Ziv source code. This scheme is different from the optimal scheme proposed by Tian et al. in a recent work for the scalar-scalar case, which implies that the optimal scheme for the scalar-scalar case is in fact not unique.
Lin Song 0003, Jun Chen 0005, Chao Tian 0002
ISIT2
2012 Gaussian robust sequential and predictive coding
abstract
We introduce two new source coding problems: robust sequential coding and robust predictive coding. For the Gauss-Markov source model, we characterize certain supporting hyperplanes of the rate region of these two coding problems. Our investigation also reveals a class of extremal inequalities and minimax theorems.
Lin Song 0003, Jun Chen 0005, Jia Wang 0004, Tie Liu 0002
ISIT2
2012 On the vector Gaussian L-terminal CEO problem
abstract
We derive an outer bound of the rate region of the vector Gaussian L-terminal CEO problem by establishing a lower bound on each supporting hyperplane of the rate region. To this end we prove a new extremal inequality by exploiting the connection between differential entropy and Fisher information as well as some fundamental estimation-theoretic inequalities. It is shown that the outer bound matches the Berger-Tung inner bound in the high-resolution regime.
Jia Wang 0004, Jun Chen 0005
ISIT2
2012 On vector Gaussian multiterminal source coding
abstract
We derive a lower bound on each supporting hyperplane of the rate region of the vector Gaussian multiterminal source coding problem by coupling it with the CEO problem through a limiting argument. The tightness of this lower bound in the high-resolution regime and the weak-dependence regime is proved.
Jia Wang 0004, Jun Chen 0005
ITW2
2012 Gaussian Multiple Description Coding with Low-Density Generator Matrix Codes
abstract
It is shown that the coding problem for an arbitrary point on the dominant face of an L-description El Gamal-Cover (EGC) region can be converted to that for a vertex of a K-description EGC region for some K≤ 2L-1, where the latter problem can be solved via successive coding. A practical scheme is proposed for the quadratic Gaussian case by reducing each step in successive coding to a Gaussian quantization operation and implementing such an operation using low-density generator matrix codes. The effectiveness of this scheme is verified through extensive simulation experiments.
Jun Chen 0005, Sorina Dumitrescu
IEEE Trans. Commun.1
2012 LDGM-Based Multiple Description Coding for Finite Alphabet Sources
abstract
This work presents an LDGM-based practical successive coding scheme for the multiple description (MD) problem for finite alphabet sources. The scheme, which targets the Zhang-Berger (ZB) rate-distortion region, is shown to be asymptotically optimal with joint typicality encoding, while as a practical encoding solution a message passing algorithm is adopted. We further discuss in more detail the application of the coding scheme in three cases of the MD problem with the Hamming distortion measure: 1) no excess sum-rate for binary sources, 2) successive refinement, and 3) no excess marginal rate for the uniform binary source. In the no excess sum-rate case some progress is made in the characterization of fundamental limits by deriving the analytical expression of the distortion region for general binary sources, and of the auxiliary variables needed to achieve its boundary. The exact expression of the Zhang-Berger upper bound to the central distortion is also provided for the case of no excess marginal rate for the uniform binary source. The proposed LDGM-based coding scheme is tested in practice for all three aforementioned cases. The experimental results show very good performance, demonstrating its ability to approach the theoretical rate-distortion limits or the available upper bounds.
Sorina Dumitrescu, Jun Chen 0005, Zhibin Sun
IEEE Trans. Commun.3
2011 Vector Gaussian multiple description coding with individual and central distortion constraints
abstract
We characterize the rate region of vector Gaussian multiple description coding with individual and central covariance distortion constraints. Specifically, we derive a lower bound and an upper bound for each supporting hyperplane of the rate region and show that these two bounds coincide. The rate region of vector Gaussian multiple description coding with individual and central trace distortion constraints as well as its extension to the stationary Gaussian process setup is also characterized.
Jun Chen 0005, Jia Wang 0004
ISIT1
2011 On the vector Gaussian CEO problem
abstract
A lower bound on each supporting line of the rate region of the vector Gaussian CEO problem is derived. The key technical ingredient is a new extremal inequality. It is shown that the lower bound coincides with the Berger-Tung upper bound in the high-resolution regime. The application of the new bounding technique to the vector Gaussian multiterminal source coding problem is also discussed.
Jun Chen 0005, Jia Wang 0004
ISIT1
2011 On the Role of the Refinement Layer in Multiple Description Coding and Scalable Coding
abstract
We clarify the relationship among several existing achievable multiple description rate-distortion regions by investigating the role of refinement layer in multiple description coding. Specifically, we show that the refinement layer in the El Gamal-Cover (EGC) scheme and the Venkataramani-Kramer-Goyal (VKG) scheme can be removed; as a consequence, the EGC region is equivalent to the EGC* region (an antecedent version of the EGC region) while the VKG region (when specialized to the 2-description case) is equivalent to the Zhang-Berger (ZB) region. Moreover, we prove that for multiple description coding with individual and hierarchical distortion constraints, the number of layers in the VKG scheme can be significantly reduced when only certain weighted sum rates are concerned. The role of refinement layer in scalable coding (a special case of multiple description coding) is also studied.
Jia Wang 0004, Jun Chen 0005, Paul W. Cuff, Haim H. Permuter
IEEE Trans. Inf. Theory2
2010 Robust multiresolution coding with Hamming distortion measure
abstract
In multiresolution coding a source sequence is encoded into a base layer and a refinement layer. The refinement layer, constructed using a conditional codebook, is in general not decodable without the correct reception of the base layer. By relating multiresolution coding with multiple description coding, we show that it is in fact possible to construct multiresolution codes in certain ways so that the refinement layer alone can be used to reconstruct the source to achieve a nontrivial distortion. As a consequence, one can improve the robustness of the existing multiresolution coding schemes without sacrificing the efficiency. Specifically, we obtain an explicit expression of the minimum distortion achievable by the refinement layer for arbitrary finite alphabet sources with Hamming distortion measure. Experimental results show that the information-theoretic limits can be approached using practical robust multiresolution coding schemes based on low-density generator matrix codes.
Jun Chen 0005, Sorina Dumitrescu, Jia Wang 0004
ISIT1
2010 Optimality and approximate optimality of source-channel separation in networks
abstract
We consider the optimality of source-channel separation in networks, and show that such a separation approach is optimal or approximately optimal for a large class of scenarios. More precisely, for lossy coding of memoryless sources in a network, when the sources are mutually independent, and each source is needed only at one destination (or at multiple destinations at the same distortion level), the separation approach is optimal; for the same setting but each source is needed at multiple destinations under a restricted class of distortion measures, the separation approach is approximately optimal, in the sense that the loss from optimum can be upper-bounded. The communication channels in the network are general, including various multiuser channels with finite memory and feedback, the sources and channels can have different bandwidths, and the sources can be present at multiple nodes.
Chao Tian 0002, Jun Chen 0005, Suhas N. Diggavi, Shlomo Shamai
ISIT2
2010 On the sum rate of vector Gaussian multiterminal source coding
abstract
Upper and lower bounds on the minimum sum rate of vector Gaussian multiterminal source coding are derived. It is shown that the two bounds coincide in the high-resolution regime and the weak-dependence regime. The two-terminal case is investigated in detail.
Jia Wang 0004, Jun Chen 0005
ISIT2
2010 Robust Multiresolution Coding
abstract
In multiresolution coding a source sequence is encoded into a base layer and a refinement layer. The refinement layer, constructed using a conditional codebook, is in general not decodable without the correct reception of the base layer. By relating multiresolution coding with multiple description coding, we show that it is in fact possible to construct multiresolution codes in certain ways so that the refinement layer alone can be used to reconstruct the source to achieve a nontrivial distortion. As a consequence, one can improve the robustness of the existing multiresolution coding schemes without sacrificing the efficiency. Specifically, we obtain an explicit expression of the minimum distortion achievable by the refinement layer for arbitrary finite alphabet sources with Hamming distortion measure. Experimental results show that the information-theoretic limits can be approached using a practical robust multiresolution coding scheme based on low-density generator matrix codes.
Jun Chen 0005, Sorina Dumitrescu, Jia Wang 0004
IEEE Trans. Commun.1
2010 Achieving the rate-distortion bound with low-density generator matrix codes
abstract
It is shown that binary low-density generator matrix codes can achieve the rate-distortion bound of discrete memoryless sources with general distortion measure via multilevel quantization. A practical encoding scheme based on the survey-propagation algorithm is proposed. The effectiveness of the proposed scheme is verified through simulation.
Zhibin Sun, Mingkai Shao, Jun Chen 0005, Kon Max Wong, Xiaolin Wu 0001
IEEE Trans. Commun.3
2010 LDPC code design for asynchronous Slepian-Wolf coding
abstract
We consider asynchronous Slepian-Wolf coding where the two encoders may not have completely accurate timing information to synchronize their individual block code boundaries, and propose LDPC code design in this scenario. A new information-theoretic coding scheme based on source splitting is provided, which can achieve the entire asynchronous Slepian-Wolf rate region. Unlike existing methods based on source splitting, the proposed scheme does not require common randomness at the encoder and the decoder, or the construction of super-letter from several individual symbols. We then design LDPC codes based on this new scheme, by applying the recently discovered source-channel code correspondence. Experimental results validate the effectiveness of the proposed method.
Zhibin Sun, Chao Tian 0002, Jun Chen 0005, Kon Max Wong
IEEE Trans. Commun.3
2010 Tighter bounds on the capacity of finite-state channels via Markov set-chains
abstract
The theory of Markov set-chains is applied to derive upper and lower bounds on the capacity of finite-state channels that are tighter than the classic bounds by Gallager. The new bounds coincide and yield single-letter capacity characterizations for a class of channels with the state process known at the receiver, including channels whose long-term marginal state distribution is independent of the input process. Analogous results are established for finite-state multiple access channels.
Jun Chen 0005, Haim H. Permuter, Tsachy Weissman
IEEE Trans. Inf. Theory1
2010 New coding schemes for the symmetric K -description problem
abstract
We propose novel coding schemes for the K-description problem with symmetric rates and symmetric distortion constraints. There are two main new ingredients in these schemes: the first one is akin to the method seen in the well-known butterfly network of network coding literature, and systematic erasure channel codes are applied on certain carefully chosen source coding component; the second approach is built on the quantization splitting technique which was previously proven useful in the Gaussian CEO problem. We first focus on a special case of the three description problem, where any two descriptions are rate-distortion optimal jointly, referred to as the no two description excess rate case. For this special case and the quadratic Gaussian source, we show that the two aforementioned approaches lead to rate-distortion points outside the achievable region based on the source-channel erasure codes, previously proposed by Pradhan, Puri, and Ramchandran. Interestingly, though only the symmetric problem is considered in our work, the proposed schemes in fact benefit from time-sharing several asymmetric rate-distortion points. The insights gained through the no two description excess rate case lead to strategic combination of the new ingredients with the existing coding scheme, yielding new coding schemes for the symmetric K -description problem.
Chao Tian 0002, Jun Chen 0005
IEEE Trans. Inf. Theory2
2010 On the sum rate of Gaussian multiterminal source coding: new proofs and results
abstract
We show that the lower bound on the sum rate of the direct and indirect Gaussian multiterminal source coding problems can be derived in a unified manner by exploiting the semidefinite partial order of the distortion covariance matrices associated with the minimum mean squared error (MMSE) estimation and the so-called reduced optimal linear estimation, through which an intimate connection between the lower bound and the Berger-Tung upper bound is revealed. We give a new proof of the minimum sum rate of the indirect Gaussian multiterminal source coding problem (i.e., the Gaussian CEO problem). For the direct Gaussian multiterminal source coding problem, we derive a general lower bound on the sum rate and establish a set of sufficient conditions under which the lower bound coincides with the Berger-Tung upper bound. We show that the sufficient conditions are satisfied for a class of sources and distortion constraints; in particular, they hold for arbitrary positive definite source covariance matrices in the high-resolution regime. In contrast with the existing proofs, the new method does not rely on Shannon's entropy power inequality.
Jia Wang 0004, Jun Chen 0005, Xiaolin Wu 0001
IEEE Trans. Inf. Theory2
2009 Capacity Region of Reversely Degraded Gaussian MIMO Broadcast Channel
abstract
We consider the problem of broadcasting a common message and two individual messages to two users on a product channel of two reversely degraded Gaussian multiple-input multiple-output (MIMO) broadcast channels. Though El Gamal provided a single letter characterization for the general discrete memoryless problem in 1980, this characterization in fact does not include a channel cost constraint, and thus does not apply directly to the Gaussian MIMO setting. We show El GamaFs single letter characterization can indeed be generalized to include channel cost constraints, however special care has to be taken and the characterization holds only with certain class of cost functions. This characterization has an equivalent form, and by utilizing this form, as well as the enhancement technique and an extremal inequality which were only discovered recently, we show that indeed Gaussian codebooks are optimal for this MIMO setting.
Jun Chen 0005, Chao Tian 0002
GLOBECOM1
2009 Gaussian multiple description coding with individual and central distortion constraints
abstract
The rate region of Gaussian multiple description coding with individual and central distortion constraints is completely characterized. It is shown that each supporting hyperplane of the rate region is associated with a saddle point of a min-max problem.
Jun Chen 0005
ISIT1
2009 Asynchronous Slepian-Wolf code design
abstract
We consider asynchronous Slepian-Wolf coding where the two encoders may not have completely accurate timing information to synchronize their individual block code boundaries, and propose LDPC code design in this scenario. A new information-theoretic coding scheme based on source splitting is provided, which can achieve the entire asynchronous Slepian-Wolf rate region. Unlike existing methods based on source splitting, the proposed scheme does not require common randomness at the encoder and the decoder, or constructing super-letter from several individual symbols. Furthermore, we show that linear codes are sufficient for each coding step of this scheme. We subsequently design LDPC codes based on this new scheme, by applying the recently discovered source-channel code correspondence. Experimental results validate the effectiveness of the proposed method.
Zhibin Sun, Chao Tian 0002, Jun Chen 0005, Kon Max Wong
ISIT3
2009 On the minimum sum rate of Gaussian multiterminal source coding: New proofs
abstract
We show that the minimum sum rate of the Gaussian multiterminal source coding problems can be derived in a unified manner by exploiting the semidefinite partial order of the distortion covariance matrices associated with the MMSE estimation and the so-called reduced optimal linear estimation. In contrast to the existing proofs, the new method does not rely on Shannon's entropy power inequality. Furthermore, this new method leads to a direct proof of the minimum sum rate of the Gaussian two-terminal source coding problem without coupling it to a Gaussian CEO problem.
Jia Wang 0004, Jun Chen 0005, Xiaolin Wu 0001
ISIT2
2009 Quantization splitting for symmetric K-channel multiple descriptions
abstract
We propose a new coding scheme for the symmetric K-channel multiple description problem based on the quantization splitting technique, which was previously successfully applied to the Gaussian CEO problem. Unlike a coding scheme we discovered earlier, the scheme proposed here can provide performance better than the one by Pradhan, Puri and Ramchandran in a component-wise manner. Though the method is conceptually straightforward once the analogy to the Gaussian CEO coding scheme is made, the general coding scheme requires constraining the space of the splitting random variables in a much more delicate way. We provide a set of conditions for a specific choice of splitting structure to yield valid splitting random variables.
Chao Tian 0002, Jun Chen 0005
ITW2
2009 The equivalence between slepian-wolf coding and channel coding under density evolution
abstract
We consider Slepian-Wolf code design based on low density parity-check (LDPC) coset codes. The density evolution formula for Slepian-Wolf coding is derived. An intimate connection between Slepian-Wolf coding and channel coding is then established. Specifically we show that, under density evolution, each Slepian-Wolf coding problem is equivalent to a channel coding problem for a binary-input output-symmetric channel.
Jun Chen 0005, Dake He, Ashish Jagmohan
IEEE Trans. Commun.1
2009 Rate region of Gaussian multiple description coding with individual and central distortion constraints
abstract
The rate region of Gaussian multiple description coding with individual and central distortion constraints is completely characterized. Specifically, a lower bound and an upper bound are derived for each supporting hyperplane of the rate region, where the lower bound is associated with a max-min game while the upper bound is associated with a min-max game; furthermore, it is shown that these two bounds coincide due to the existence of a saddle point.
Jun Chen 0005
IEEE Trans. Inf. Theory1
2009 On the duality between Slepian-Wolf coding and channel coding under mismatched decoding
abstract
In this paper, Slepian-Wolf coding with a mismatched decoding metric is studied. Two different dualities between Slepian-Wolf coding and channel coding under mismatched decoding are established. These two dualities provide a systematic framework for comparing linear Slepian-Wolf codes, nonlinear Slepian-Wolf codes, and variable-rate Slepian-Wolf codes. In contrast with the fact that linear codes suffice to achieve the Slepian-Wolf limit under matched decoding, the minimum rate achievable with nonlinear Slepian-Wolf codes under mismatched decoding can be strictly lower than that achievable with linear Slepian-Wolf codes.
Jun Chen 0005, Dake He, Ashish Jagmohan
IEEE Trans. Inf. Theory1
2009 On the linear codebook-level duality between Slepian-Wolf coding and channel coding
abstract
In this paper, it is shown that each Slepian-Wolf coding problem is related to a dual channel coding problem in the sense that the sphere packing exponents, random coding exponents, and correct decoding exponents in these two problems are mirror-symmetrical to each other. This mirror symmetry is interpreted as a manifestation of the linear codebook-level duality between Slepian-Wolf coding and channel coding. Furthermore, this duality, in conjunction with a systematic analysis of the expurgated exponents, reveals that nonlinear Slepian-Wolf codes can strictly outperform linear Slepian-Wolf codes in terms of rate-error tradeoff at high rates. The linear codebook-level duality is also established for general sources and channels.
Jun Chen 0005, Dake He, Ashish Jagmohan, Luis A. Lastras, En-Hui Yang
IEEE Trans. Inf. Theory1
2009 Multiple description coding for stationary Gaussian sources
abstract
We consider the problem of multiple description coding for stationary Gaussian sources under the squared error distortion measure. The rate region is characterized for the 2-description case. It is shown that each supporting line of the rate region is achievable with a transform lattice quantization scheme. We show the optimal coding scheme has a natural spectral domain coding interpretation, which yields a reverse water-filling solution with a frequency-dependent water level instead of the flat water level as in the conventional single description case.
Jun Chen 0005, Chao Tian 0002, Suhas N. Diggavi
IEEE Trans. Inf. Theory1
2009 On the redundancy of Slepian--Wolf coding
abstract
In this paper, the redundancy of both variable and fixed rate Slepian–Wolf coding is considered. Given any jointly memoryless source-side information pair$\{(X_i, Y_i)\}_{i=1}^{\infty}$with finite alphabet, the redundancy$R^n(\epsilon_n)$of variable rate Slepian–Wolf coding of$X_1^n$with decoder only side information$Y_1^n$depends on both the block length$n$and the decoding block error probability$\epsilon_n$, and is defined as the difference between the minimum average compression rate of order$n$variable rate Slepian–Wolf codes having the decoding block error probability less than or equal to$\epsilon_n$, and the conditional entropy$H(X\vert Y)$, where$H(X\vert Y)$is the conditional entropy rate of the source given the side information. The redundancy of fixed rate Slepian–Wolf coding of$X_1^n$with decoder only side information$Y_1^n$is defined similarly and denoted by$R^n_F(\epsilon_n)$. It is proved that under mild assumptions about$\epsilon_n,$$R^n(\epsilon_n) = d_v \sqrt{-\log\epsilon_n/n} + o(\sqrt{-\log \epsilon_n/n})$and$R^n_{F}(\epsilon_n) = d_f \sqrt{- \log \epsilon_n / n} + o(\sqrt{-\log \epsilon_n/n})$, where$d_f$and$d_v$are two constants completely determined by the joint distribution of the source-side information pair. Since$d_v$is generally smaller than$d_f$, our results show that variable rate Slepian–Wolf coding is indeed more efficient than fixed rate Slepian–Wolf coding.
Dake He, Luis A. Lastras, En-Hui Yang, Ashish Jagmohan, Jun Chen 0005
IEEE Trans. Inf. Theory5
2009 Capacity region of the finite-state multiple-access channel with and without feedback
abstract
The capacity region of the finite-state multiple-access channel (FS-MAC) with feedback that may be an arbitrary time-invariant function of the channel output samples is considered. We characterize both an inner and an outer bound for this region, using Massey's directed information. These bounds are shown to coincide, and hence yield the capacity region, of indecomposable FS-MACs without feedback and of stationary and indecomposable FS-MACs with feedback, where the state process is not affected by the inputs. Though “multiletter” in general, our results yield explicit conclusions when applied to specific scenarios of interest. For example, our results allow us to do the following.Identify a large class of FS-MACs, that includes the additive$\bmod \,2$noise MAC where the noise may have memory, for which feedback does not enlarge the capacity region.
Haim H. Permuter, Tsachy Weissman, Jun Chen 0005
IEEE Trans. Inf. Theory3
2009 Remote vector Gaussian source coding with decoder side information under mutual information and distortion constraints
abstract
Let X , Y , Z be zero-mean, jointly Gaussian random vectors of dimensions nx, ny, and nz, respectively. Let P be the set of random variables W such that W harr Y harr (X, Z) is a Markov string. We consider the following optimization problem:WisinPminI(Y; Z) subject to one of the following two possible constraints: 1) I(X; W|Z) ges RI, and 2) the mean squared error between X and Xcirc = E(X|W, Z) is less than d . The problem under the first kind of constraint is motivated by multiple-input multiple-output (MIMO) relay channels with an oblivious transmitter and a relay connected to the receiver through a dedicated link, while for the second case, it is motivated by source coding with decoder side information where the sensor observation is noisy. In both cases, we show that jointly Gaussian solutions are optimal. Moreover, explicit water filling interpretations are given for both cases, which suggest transform coding approaches performed in different transform domains, and that the optimal solution for one problem is, in general, suboptimal for the other.
Chao Tian 0002, Jun Chen 0005
IEEE Trans. Inf. Theory2
2008 On Universal Variable-Rate Slepian-Wolf Coding
abstract
Lower and upper bounds on the reliability region of universal variable-rate Slepian-Wolf coding are derived.
Jun Chen 0005, Dake He, Ashish Jagmohan, Luis A. Lastras
ICC1
2008 Slepian-Wolf coding with a mismatched decoder
abstract
Slepian-Wolf coding with a mismatched decoding metric is studied. Two different dualities between Slepian-Wolf coding and channel coding under mismatched decoding are established. These two dualities provide a systematic framework for comparing linear Slepian-Wolf codes, nonlinear Slepian-Wolf codes, and variable-rate Slepian-Wolf codes. In contrast with the fact that linear codes suffice to achieve the Slepian-Wolf limit under matched decoding, the minimum rate achievable with linear Slepian-Wolf codes under mismatched decoding can be strictly higher than that achievable with nonlinear Slepian-Wolf codes.
Jun Chen 0005, Dake He, Ashish Jagmohan
ISIT1
2008 On the capacity of finite-state channels
abstract
New upper and lower bounds on the capacity of finite-state channels are established. For a class of channels, these bounds yield a single-letter capacity formula, which is previously unknown in the literature.
Jun Chen 0005, Haim H. Permuter, Tsachy Weissman
ISIT1
2008 New bounds for the capacity region of the Finite-State Multiple Access Channel
abstract
The capacity region of the finite-state multiple access channel (FS-MAC) with feedback that may be an arbitrary time-invariant function of the channel output samples is considered. We provided a sequence of inner and outer bounds for this region. These bounds are shown to coincide, and hence yield the capacity region for two cases of FS-MACs: (1) when the state process is stationary and ergodic and not affected by the inputs; (2) an indecomposable FS-MAC without feedback. Though the capacity region is "multi-letter" in general, our results yield explicit conclusions when applied to specific scenarios of interest.
Haim H. Permuter, Tsachy Weissman, Jun Chen 0005
ISIT3
2008 A novel coding scheme for symmetric multiple description coding
abstract
We propose a novel coding scheme for multiple description coding with K descriptions, with symmetric rate and symmetric distortion constraints. In this new coding scheme, channel codes are introduced on top of some carefully chosen source coding component. The channel coding component for k = 3 is akin to the network coding idea. We first show that for the quadratic Gaussian source, adding a simple channel coding component is able to achieve rate- distortion point outside the achievable region based on (n, k) source-channel erasure codes (SCEC), previously proposed by Puri et al, when Gaussian codebook is assumed to be optimal for that scheme. The channel coding component and (n, k) SCEC can be strategically combined, and for the special case of quadratic Gaussian three description case where any two of them are rate-distortion optimal jointly, the resulting scheme can achieve performance better than a direct time sharing between the new operating point and the (n, k) SCEC-based scheme. Moreover, this combination can be generalized naturally, which leads to a new coding scheme for the general if-description problem.
Chao Tian 0002, Jun Chen 0005
ISIT2
2008 Successive Wyner-Ziv Coding Scheme and Its Application to the Quadratic Gaussian CEO Problem
abstract
In this paper, we introduce a distributed source coding scheme called successive Wyner-Ziv coding. We show that every point in the rate region of the quadratic Gaussian CEO problem can be achieved via successive Wyner-Ziv coding. The concept of successive refinement in single source coding is generalized to the distributed source coding scenario, which we refer to as distributed successive refinement. For the quadratic Gaussian CEO problem, we establish a necessary and sufficient condition for distributed successive refinement, where the successive Wyner-Ziv coding scheme plays an important role.
Jun Chen 0005, Toby Berger
IEEE Trans. Inf. Theory1
2008 Robust Distributed Source Coding
abstract
We consider a distributed source coding system in which several observations must be encoded separately and communicated to the decoder by using limited transmission rate. We introduce a robust distributed coding scheme which flexibly trades off between system robustness and compression efficiency. The optimality of this coding scheme is proved for various special cases.
Jun Chen 0005, Toby Berger
IEEE Trans. Inf. Theory1
2008 Successive Refinement for Hypothesis Testing and Lossless One-Helper Problem
abstract
We investigate two closely related successive refinement (SR) coding problems: 1) In the hypothesis testing (HT) problem, bivariate hypothesis$H_{0}:P_{XY}$against$H_{1}: P_{X}P_{Y}$, i.e., test against independence is considered. One remote sensor collects data stream$X$and sends summary information, constrained by SR coding rates, to a decision center which observes data stream$Y$directly. 2) In the one-helper (OH) problem,$X$and$Y$are encoded separately and the receiver seeks to reconstruct$Y$losslessly. Multiple levels of coding rates are allowed at the two sensors, and the transmissions are performed in an SR manner. We show that the SR-HT rate-error-exponent region and the SR-OH rate region can be reduced to essentially the same entropy characterization form. Single-letter solutions are thus provided in a unified fashion, and the connection between them is discussed. These problems are also related to the information bottleneck (IB) problem, and through this connection we provide a straightforward operational meaning for the IB method. Connection to the pattern recognition problem, the notion of successive refinability, and two specific sources are also discussed. A strong converse for the SR-HT problem is proved by generalizing the image size characterization method, which shows the optimal type-two error exponents under constant type-one error constraints are independent of the exact values of those constants.
Chao Tian 0002, Jun Chen 0005
IEEE Trans. Inf. Theory2
2008 Multiuser Successive Refinement and Multiple Description Coding
abstract
In this correspondence, we consider the multiuser successive refinement (MSR) problem, where the users are connected to a central server via links with different noiseless capacities, and each user wishes to reconstruct in a successive-refinement fashion. An achievable region is given for the two-user two-layer case and it provides the complete rate-distortion region for the Gaussian source under the MSE distortion measure. The key observation is that this problem includes the multiple description (MD) problem (with two descriptions) as a subsystem, and the techniques useful in the MD problem can be extended to this case. It is shown that the coding scheme based on the universality of random binning is suboptimal, because multiple Gaussian side informations only at the decoders do incur performance loss, in contrast to the case of single side information at the decoder. It is further shown that unlike the single user case, when there are multiple users, the loss of performance by a multistage coding approach can be unbounded for the Gaussian source. The result suggests that in such a setting, the benefit of using successive refinement is not likely to justify the accompanying performance loss. The MSR problem is also related to the source coding problem where each decoder has its individual side information, while the encoder has the complete set of the side informations. The MSR problem further includes several variations of the MD problem, for which the specialization of the general result is investigated and the implication is discussed.
Chao Tian 0002, Jun Chen 0005, Suhas N. Diggavi
IEEE Trans. Inf. Theory2
2007 Multiple Description Coding for Stationary and Ergodic Sources
abstract
We consider the problem of multiple description (MD) coding for stationary sources with the squared error distortion measure. The MD rate region is derived for the stationary and ergodic Gaussian sources, and is shown to be achievable with a practical transform lattice quantization scheme. Moreover, the proposed scheme is asymptotically optimal at high resolution for all stationary sources with finite differential entropy rate
Jun Chen 0005, Chao Tian 0002, Suhas N. Diggavi
DCC1
2007 On the Redundancy-Error Tradeoff in Slepian-Wolf Coding and Channel Coding
abstract
We characterize the redundancy-error tradeoff in Slepian-Wolf coding. Similar results are derived for a class of cyclic-symmetric channels. Through the linear codebook-level duality between Slepian-Wolf coding and channel coding, we show that, in Slepian-Wolf coding, linear codes are optimal in terms of redundancy-error tradeoff at rate close to the Slepian-Wolf limit but suboptimal at high rate.
Jun Chen 0005, Dake He, Ashish Jagmohan, Luis A. Lastras
ISIT1
2007 Hypothesis Testing Under Successive Refinement Communication Constraints
abstract
We investigate the distributed successive refinement (SR) hypothesis testing problem. Bivariate hypothesis Ho : Pxy against Hi : PxPy, i.e., test against independence is considered, when a remote sensor sends compressed information about data stream X subject to SR coding constraints to a decision site, where data stream Y can be observed directly. We show that this problem is closely related to the SR lossless one-helper problem and the SR pattern recognition problem. More precisely, the rate-type-two-error-exponent region of the SR hypothesis testing problem, the rate region of the SR one-helper problem and that of the SR pattern recognition problem can be reduced to essentially the same entropy characterization form, up to an isometry. Single letter solution is subsequently given in this unified framework, and this connection is further explored. Strong converse result is proved by generalizing the image size characterization technique, which shows the optimal type-two error exponents for SR hypothesis testing under fixed type-one error constraints are independent of the exact values of those constraints. The notion of successive refinability is defined, and somewhat surprisingly for large value of type-one error constraints, a source is always successive refinable for hypothesis testing.
Chao Tian 0002, Jun Chen 0005
ISIT2
2007 Capacity Results for Block-Stationary Gaussian Fading Channels With a Peak Power Constraint
abstract
A peak-power-limited single-antenna block-stationary Gaussian fading channel is studied, where neither the transmitter nor the receiver knows the channel state information, but both know the channel statistics. This model subsumes most previously studied Gaussian fading models. The asymptotic channel capacity in the high signal-to-noise ratio (SNR) regime is first computed, and it is shown that the behavior of the channel capacity depends critically on the channel model. For the special case where the fading process is symbol-by-symbol stationary, it is shown that the codeword length must scale at least logarithmically with SNR in order to guarantee that the communication rate can grow logarithmically with SNR with decoding error probability bounded away from one. An expression for the capacity per unit energy is also derived. Furthermore, it is shown that the capacity per unit energy is achievable using temporal ON-OFF signaling with optimally allocated ON symbols, where the optimal ON-symbol allocation scheme may depend on the peak power constraint.
Jun Chen 0005, Venugopal V. Veeravalli
IEEE Trans. Inf. Theory1
2006 Slepian-Wolf Code Design via Source-Channel Correspondence
abstract
We consider Slepian-Wolf code design based on LDPC (low-density parity-check) coset codes for memoryless source-side information pairs. A density evolution formula, equipped with a concentration theorem, is derived for Slepian-Wolf coding based on LDPC coset codes. As a consequence, an intimate connection between Slepian-Wolf coding and channel coding is established. Specifically we show that, under density evolution, design of binary LDPC coset codes for Slepian-Wolf coding of an arbitrary memoryless source-side information pair reduces to design of binary LDPC codes for binary-input output-symmetric channels without loss of optimality. With this connection, many classic results in channel coding can be easily translated into the Slepian-Wolf setting
Jun Chen 0005, Dake He, Ashish Jagmohan
ISIT1
2006 Multiple Description Quantization Via Gram-Schmidt Orthogonalization
abstract
The multiple description (MD) problem has received considerable attention as a model of information transmission over unreliable channels. A general framework for designing efficient MD quantization schemes is proposed in this paper. We provide a systematic treatment of the El Gamal-Cover (EGC) achievable MD rate-distortion region, and show it can be decomposed into a simplified-EGC (SEGC) region and a superimposed refinement operation. Furthermore, any point in the SEGC region can be achieved via a successive quantization scheme along with quantization splitting. For the quadratic Gaussian case, the proposed scheme has an intrinsic connection with the Gram-Schmidt orthogonalization, which implies that the whole Gaussian MD rate-distortion region is achievable with a sequential dithered lattice-based quantization scheme as the dimension of the (optimal) lattice quantizers becomes large. Moreover, this scheme is shown to be universal for all independent and identically distributed (i.i.d.) smooth sources with performance no worse than that for an i.i.d. Gaussian source with the same variance and asymptotically optimal at high resolution. A class of MD scalar quantizers in the proposed general framework is also constructed and is illustrated geometrically; the performance is analyzed in the high-resolution regime, which exhibits a noticeable improvement over the existing MD scalar quantization schemes
Jun Chen 0005, Chao Tian 0002, Toby Berger, Sheila S. Hemami
IEEE Trans. Inf. Theory1
2005 When is Bit Allocation for Predictive Video Coding Easy?
abstract
This paper addresses the problem of bit allocation among frames in a predictively encoded video sequence. Finding optimal solutions to this problem potentially requires making an exponential number of calls to the encoder. To better understand the structure of the rate-distortion data output by video encoders, a simple model of a sequentially encoded autoregressive Gaussian random field is theoretically investigated. The rate-distortion data for the model exhibits an additive-separability property, i.e. the rate can be decomposed into a sum of independent functions of single distortion variables. This property implies the near-optimal behavior of a non-backtracking steepest-descent (SD) based bit allocation algorithm. The SD algorithm when applied to video coding produces near-optimal solutions by making a linear number of calls to the encoder. Results are presented for MPEG-2 encoding of standard video sequences.
Yegnaswamy Sermadevi, Jun Chen 0005, Sheila S. Hemami, Toby Berger
DCC2
2005 A new class of universal multiple description lattice quantizers
abstract
We propose a new class of universal multiple description lattice quantizers based on the method of quantization splitting. For Gaussian sources and squared error distortion measure, our scheme can achieve the whole multiple description rate-distortion region, as the dimension of the (optimal) lattice quantizers becomes large
Jun Chen 0005, Chao Tian 0002, Toby Berger, Sheila S. Hemami
ISIT1