EDBT 2026 Demo / reviewers in the wild / expert
Chang Wen Chen
dblp:29/4638 · also Changwen Chen
· DBLP profile ↗
413ranked-venue papers
27as first author
78since 2021 · last 2026
0000-0002-6720-234XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 277 · 23 first-author · 51 since 2021Computer networks · 81 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 45 · 3 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 5 since 2021Systems, architecture and hardware · 15 · 2 since 2021Databases, data management, data science and information retrieval · 4Human-computer interaction and ubiquitous computing · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BLADE: Adaptive Wi-Fi Contention Control for Next-Generation Real-Time Communication
Fengqian Guo, Longwei Jiang, Congcong Miao, Chenren Xu, Hancheng Lu, Chang Wen Chen, Yaxiong Xie |
NSDI | 8 |
| 2026 | A Survey on Video Temporal Grounding With Multimodal Large Language ModelabstractThe recent advancement in video temporal grounding (VTG) has significantly enhanced fine-grained video understanding, primarily driven by multimodal large language models (MLLMs). With superior multimodal comprehension and reasoning abilities, VTG approaches based on MLLMs (VTG-MLLMs) are gradually surpassing traditional fine-tuned methods. They not only achieve competitive performance but also excel in generalization across zero-shot, multi-task, and multi-domain settings. Despite extensive surveys on general video-language understanding, comprehensive reviews specifically addressing VTG-MLLMs remain scarce. To fill this gap, this survey systematically examines current research on VTG-MLLMs through a three-dimensional taxonomy: 1) the functional roles of MLLMs, highlighting their architectural significance; 2) training paradigms, analyzing strategies for temporal reasoning and task adaptation; and 3) video feature processing techniques, which determine spatiotemporal representation effectiveness. We further discuss benchmark datasets, evaluation protocols, and summarize empirical findings. Finally, we identify existing limitations and propose promising research directions. Jianlong Wu, Wei Liu 0005, Ye Liu 0002, Meng Liu 0006, Liqiang Nie, Zhouchen Lin, Chang Wen Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | An Efficient Neural Rate Control for JPEG-AIabstractRate control (RC) is a critical component in learned image compression (LIC), particularly in the emerging JPEG-AI standard, which enables adaptive bitrate achievement to meet diverse bandwidth constraints. JPEG-AI default RC employs an iterative optimization process, wherein a pre-trained RC model is selected and the (generated) latent representations are adjusted based on the mismatch between actual and target bitrates. Despite satisfactory results, such a trial-and-error paradigm necessitates multiple processing cycles, resulting in inevitable computational overhead. We propose an efficient neural rate control framework for JPEG-AI to address this limitation. Our idea is to train a ResNet-based neural control (NRC) to learn the mapping from the input images and target bitrates to the optimal coding parameters. The trained NRC can then be applied to predict the coding parameters based on the new input images and target bitrates directly. Experimental results on DIV2K and MSCOCO datasets show that our NRC achieves comparable rate-distortion performance while reducing encoding time by about 5× compared to JPEG-AI default RC. Guanchen Ding, Zhenzhong Chen 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | De-LightSAM: Modality-Decoupled Lightweight SAM for Generalizable Medical SegmentationabstractThe universality of deep neural networks across different modalities and their generalization capabilities to unseen domains play an essential role in medical image segmentation. The recent segment anything model (SAM) has demonstrated strong adaptability across diverse natural scenarios. However, the huge computational costs, demand for manual annotations as prompts and conflict-prone decoding process of SAM degrade its generalization capabilities in medical scenarios. To address these limitations, we propose a modality-decoupled lightweight SAM for domain-generalized medical image segmentation, named De-LightSAM. Specifically, we first devise a lightweight domain-controllable image encoder (DC-Encoder) that produces discriminative visual features for diverse modalities. Further, we introduce the self-patch prompt generator (SP-Generator) to automatically generate high-quality dense prompt embeddings for guiding segmentation decoding. Finally, we design the query-decoupled modality decoder (QM-Decoder) that leverages a one-to-one strategy to provide an independent decoding channel for every modality, preventing mutual knowledge interference of different modalities. Moreover, we design a multi-modal decoupled knowledge distillation (MDKD) strategy to leverage robust common knowledge to complement domain-specific medical feature representations. Extensive experiments indicate that De-LightSAM outperforms state-of-the-arts in diverse medical imaging segmentation tasks, displaying superior modality universality and generalization capabilities. Especially, De-LightSAM uses only 2.0% parameters compared to SAM-H. The source code is available at https://github.com/xq141839/De-LightSAM. Qing Xu 0014, Xiangjian He, Chenxin Li, Fiseha B. Tesema, Wenting Duan, Zhen Chen 0013, Rong Qu, Jonathan M. Garibaldi, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2026 | MoGenVD: A Motion-Centered Quality Assessment Benchmark for Text-to-Video Generation
Yingxue Zhang 0004, Zike Yang, Zihang Su, Yaosi Hu, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | MotionPrior: Exploring Efficient Learning of Motion Concepts for Few-Shot Video GenerationabstractThe diffusion-based text-to-image generation has achieved remarkable progress and realistic content generation performance, greatly promoting the development in text-to-video generation. Although equipped with powerful image diffusion models, video generation modeling still requires massive labeled data and a high training resource cost. Recent, work has been focused on cost-effective video generation in a one-shot or few-shot manner based on the image diffusion model with minimum demand for video data and computing resources. However, these video generation models only support the generation of one single motion pattern/concept. This raises an important question: Can we improve generation freedom with a light training burden? In this paper, we explore a cost-effective video generation scheme for adaptive motion concepts by learning motion priors from a small set of video data. Specifically, we construct a learnable bank for motion concepts and propose the Dual-Semantic-guided Motion Attention module to locate the corresponding motion elements from the bank with the guidance of textual semantic and visual semantic. The extracted motion elements are inserted into video latents via lightweight motion injection layer, which is capable of integrating motion semantic effectively with much fewer parameters compared to the conventional temporal attention layer. In addition, we introduce a temporal-aware noise prior and an inter-frame consistency constraint to strengthen the learning of temporal dependency and improve video smoothness. Extensive experiments validate that the proposed method can learn motion priors adaptively from a small set of training videos to generate smooth videos that involve either single or multiple motion concepts. The results demonstrate that the proposed scheme achieves superior performance compared to existing few-shot video generation methods and even some large-scale video generation models. More information and results are available at https://youncy-hu.github.io/motionprior/. Yaosi Hu, Chang Wen Chen |
IEEE Trans. Image Process. | 2 |
| 2025 | SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject ControlabstractAutonomous driving progress relies on large-scale annotated datasets. In this work, we explore the potential of generative models to produce vast quantities of freely-labeled data for autonomous driving applications and present SubjectDrive, the first model proven to scale generative data production in a way that could continuously improve autonomous driving applications. We investigate the impact of scaling up the quantity of generative data on the performance of downstream perception models and find that enhancing data diversity plays a crucial role in effectively scaling generative data production. Therefore, we have developed a novel model equipped with a subject control mechanism, which allows the generative model to leverage diverse external data sources for producing varied and useful data. Extensive evaluations confirm SubjectDrive's efficacy in generating scalable autonomous driving training data, marking a significant step toward revolutionizing data production methods in this field. Binyuan Huang, Yuqing Wen, Yaosi Hu, Yingfei Liu, Fan Jia 0006, Weixin Mao, Tiancai Wang, Chi Zhang 0026, Chang Wen Chen, Zhenzhong Chen 0001, Xiangyu Zhang 0005 |
AAAI | 10 |
| 2025 | Removing Out-of-Focus Reflective Flares via Color Alignment
Fengbo Lan, Chang Wen Chen |
ICCV | 2 |
| 2025 | Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
Zeyu Xi, Haoying Sun, Yaofei Wu, Junchi Yan, Lifang Wu, Liang Wang 0001, Chang Wen Chen |
ICCV | 8 |
| 2025 | Prohibited Items Segmentation via Occlusion-aware Bilayer ModelingabstractInstance segmentation of prohibited items in security X-ray images is a critical yet challenging task. This is mainly caused by the significant appearance gap between prohibited items in X-ray images and natural objects, as well as the severe overlapping among objects in X-ray images. To address these issues, we propose an occlusion-aware instance segmentation pipeline designed to identify prohibited items in X-ray images. Specifically, to bridge the representation gap, we integrate the Segment Anything Model (SAM) into our pipeline, taking advantage of its rich priors and zero-shot generalization capabilities. To address the overlap between prohibited items, we design an occlusion-aware bilayer mask decoder module that explicitly models the occlusion relationships. To supervise occlusion estimation, we manually annotated occlusion areas of prohibited items in two large-scale X-ray image segmentation datasets, PIDray and PIXray. We then reorganized these additional annotations together with the original information as two occlusion-annotated datasets, PIDray-A and PIXray-A. Extensive experimental results on these occlusion-annotated datasets demonstrate the effectiveness of our proposed method. The datasets and codes are available at: https://github.com/Ryh1218/Occ. Yunhan Ren, Ruihuang Li, Lingbo Liu, Chang Wen Chen |
ICME | 4 |
| 2025 | A 20.98TOPS/W Energy-Efficient Binary BERT Model on Group Vector Systolic CIM AcceleratorabstractTransformer-based large language models (LLMs) impose significant bandwidth and compute challenges when deployed on edge devices. SRAM-based compute-in-memory (CIM) accelerators offer a promising solution to reduce data movement but are still limited by model size. This work develops a ternary weight splitting (TWS) binarization to obtain Brain-Floating-Point-16×INT1 (BF16×1-b) and INT8×INT1 (8-b×1-b) based transformers that exhibit competitive accuracy while significantly reducing model size compared to full precision counterparts. Then, a fully digital SRAM-based CIM accelerator is designed incorporating a bit-parallel SRAM macro within a highly efficient group vector systolic architecture, which can store one column of BERT-Tiny model with stationary systolic data reuse. The design in a 28nm technology only requires 2KB SRAM with an area of 2mm2. It achieves a throughput of 6.55TOPS and consumes a total power of 312.5mW and 221mW at 400MHz, resulting in a state-of-the-art area efficiency of 3.3TOPS/mm2and normalized energy efficiency of 20.98TOPS/W and 34.35TOPS/W for BF16×1-b and 8-b×1-b respectively on BERT-Tiny model, demonstrating a 10.25× improvement in area efficiency and a 2.23× improvement in energy efficiency compared to other state-of-the-art counterparts. Additionally, our proposed configuration compresses the model size by 32% with only a 0.5% accuracy loss on SST-2. Dingbang Liu, Qilong Chen, Jingyun Gu, Jiaqi Yang 0009, Kai Li 0024, Wei Mao 0002, Ngai Wong 0001, Chang Wen Chen, Hao Yu 0001 |
ISLPED | 9 |
| 2025 | Multimodal LLMs Can Reason about Aesthetics in Zero-ShotabstractThe rapid technical progress of generative art (GenArt) has democratized the creation of visually appealing imagery. However, achieving genuine artistic impact - the kind that resonates with viewers on a deeper, more meaningful level - remains formidable as it requires a sophisticated aesthetic sensibility. This sensibility involves a multifaceted cognitive process extending beyond mere visual appeal, which is often overlooked by current computational methods. This paper pioneers an approach to capture this complex process by investigating how the reasoning capabilities of Multimodal LLMs (MLLMs) can be effectively elicited to perform aesthetic judgment. Our analysis reveals a critical challenge: MLLMs exhibit a tendency towards hallucinations during aesthetic reasoning, characterized by subjective opinions and unsubstantiated artistic interpretations. We further demonstrate that these hallucinations can be suppressed by employing an evidence-based and objective reasoning process, as substantiated by our proposed baseline, ArtCoT. MLLMs prompted by this principle produce multifaceted, in-depth aesthetic reasoning that aligns significantly better with human judgment. These findings have direct applications in areas such as AI art tutoring and as reward models for image generation. Ultimately, we hope this work paves the way for AI systems that can truly understand, appreciate, and contribute to art that aligns with human aesthetic values. Project homepage: https://github.com/songrise/MLLM4Art. Ruixiang Jiang, Chang Wen Chen |
ACM Multimedia | 2 |
| 2025 | DiffArtist: Towards Structure and Appearance Controllable Image Stylization
Ruixiang Jiang, Chang Wen Chen |
ACM Multimedia | 2 |
| 2025 | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual ReasoningabstractRecent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conversely, less attention has been given to scaling fine-grained pixel-level understanding capabilities, where the models are expected to realize pixel-level alignment between visual signals and language semantics. Some previous studies have applied LMMs to related tasks such as region-level captioning and referring expression segmentation. However, these models are limited to performing either referring or segmentation tasks independently and fail to integrate these fine-grained perception capabilities into visual reasoning. To bridge this gap, we propose UniPixel, a large multi-modal model capable of flexibly comprehending visual prompt inputs and generating mask-grounded responses. Our model distinguishes itself by seamlessly integrating pixel-level perception with general visual understanding capabilities. Specifically, UniPixel processes visual prompts and generates relevant masks on demand, and performs subsequent reasoning conditioning on these intermediate pointers during inference, thereby enabling fine-grained pixel-level reasoning. The effectiveness of our approach has been verified on 10 benchmarks across a diverse set of tasks, including pixel-level referring/segmentation and object-centric understanding in images/videos. A novel PixelQA task that jointly requires referring, segmentation, and question answering is also designed to verify the flexibility of our method. Ye Liu 0002, Zongyang Ma, Junfu Pu, Zhongang Qi, Yang Wu 0001, Ying Shan, Chang Wen Chen |
NeurIPS | 7 |
| 2025 | Exploring Rich Subjective Quality Information for Image Quality Assessment in the WildabstractTraditional in the wild image quality assessment (IQA) models are generally trained with the quality labels of mean opinion score (MOS), while missing the rich subjective quality information contained in the quality ratings, for example, the standard deviation of opinion scores (SOS) or even distribution of opinion scores (DOS). In this paper, we propose a novel IQA method namedRichIQAto explore the rich subjective rating information beyond MOS to predict image quality in the wild. RichIQA is characterized by two key novel designs: 1) a three-stage image quality prediction network which exploits the powerful feature representation capability of the Convolutional vision Transformer (CvT) and mimics the short-term and long-term memory mechanisms of human brain; 2) a multi-label training strategy in which rich subjective quality information like MOS, SOS and DOS are concurrently used to train the quality prediction network. Powered by these two novel designs, RichIQA is able to predict the image quality in terms of a distribution, from which the mean image quality can be subsequently obtained. Extensive experimental results verify that the three-stage network is tailored to predict rich quality information, while the multi-label training strategy can fully exploit the potentials within subjective quality rating and enhance the prediction performance and generalizability of the network. RichIQA outperforms state-of-the-art competitors on multiple large-scale in the wild IQA databases with rich subjective rating labels. The code of RichIQA will be made publicly available on GitHub. Xiongkuo Min, Yuqin Cao, Guangtao Zhai, Wenjun Zhang 0001, Huifang Sun, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Counting Beyond Domains: Toward Alignment in Unsupervised Domain Adaptation in Remote Sensing Object Counting
Guanchen Ding, Daiqin Yang, Zhenzhong Chen 0001, Chang Wen Chen |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Motion Attention-Guided Relational Reasoning for Weakly Supervised Group Activity RecognitionabstractThe existing attention-based label-free weakly supervised group activity recognition methods can automatically learn tokens related to the actors. And they have difficulties generating sufficiently diverse token embeddings. To address these issues, we automatically obtain the grayscale motion mask of all the moving objects based on the motion direction not the motion amplitude. A Motion-Guided Mask Generator module (MGMG) is proposed to estimate the attention region mask under the supervision of the grayscale motion mask. MGMG involves four parts. A correlation layer measures the relative displacement between two adjacent feature maps. A cosine attention mechanism is designed to reduce the module's sensitivity to feature amplitude changes. A mask generator is built to generate the attention region mask. And a specifically designed activation function is used to refine the attention region mask and to enhance its focus on actor motion regions. We also customize a normalized relative error loss function for MGMG module. This loss can address the value range mismatch problem for the estimated attention mask as well as the grayscale motion mask. Furthermore, a Motion Attention-Guided Relational Reasoning (MAGRR) framework is presented for the weakly supervised condition. It uses the MGMG module to estimate the attention region automatically, and a Spatial-temporal Aggregation Stack (SAS) module to activate the attention regions of the features at the spatial level, then transform them into multiple tokens, which are further captured by the attention mechanism for their temporal dependencies and interrelationships. MAGRR is experimented on the Collective Activity dataset and the Collective Activity Extension dataset, achieving state-of-the-art performance and competitive performance on the Volleyball and the NBA datasets. Yihao Zheng 0002, Zhuming Wang, Lifang Wu, Liang Wang 0001, Chang Wen Chen |
IEEE Trans. Image Process. | 5 |
| 2025 | ZS-VAT: Learning Unbiased Attribute Knowledge for Zero-Shot Recognition Through Visual Attribute TransformerabstractIn zero-shot learning (ZSL), attribute knowledge plays a vital role in transferring knowledge from seen classes to unseen classes. However, most existing ZSL methods learn biased attribute knowledge, which usually results in biased attribute prediction and a decline in zero-shot recognition performance. To solve this problem and learn unbiased attribute knowledge, we propose a visual attribute Transformer for zero-shot recognition (ZS-VAT), which is an effective and interpretable Transformer designed specifically for ZSL. In ZS-VAT, we design an attribute-head self-attention (AHSA) that is capable of learning unbiased attribute knowledge. Specifically, each attribute head in AHSA first transforms the local features into attribute-reinforced features and then accumulates the attribute knowledge from all corresponding reinforced features, reducing the mutual influence between attributes and avoiding information loss. AHSA finally preserves unbiased attribute knowledge through attribute embeddings. We also propose an attribute fusion model (AFM) that learns to recover the correct category knowledge from the attribute knowledge. In particular, AFM takes all features from AHSA as input and generates global embeddings. We carried out experiments to demonstrate that the attribute knowledge from AHSA and the category knowledge from AFM are able to assist each other. During the final semantic prediction, we combine the attribute embedding prediction (AEP) and global embedding prediction (GEP). We evaluated the proposed scheme on three benchmark datasets. ZS-VAT outperformed the state-of-the-art generalized ZSL (GZSL) methods on two datasets and achieved competitive results on the other dataset. Zongyan Han, Zhenyong Fu, Shuo Chen 0003, Le Hui, Jian Yang 0003, Chang Wen Chen |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Learning to Aggregate Multi-Scale Context for Instance Segmentation in Remote Sensing ImagesabstractThe task of instance segmentation in remote sensing images, aiming at performing per-pixel labeling of objects at the instance level, is of great importance for various civil applications. Despite previous successes, most existing instance segmentation methods designed for natural images encounter sharp performance degradations when they are directly applied to top-view remote sensing images. Through careful analysis, we observe that the challenges mainly come from the lack of discriminative object features due to severe scale variations, low contrasts, and clustered distributions. In order to address these problems, a novel context aggregation network (CATNet) is proposed to improve the feature extraction process. The proposed model exploits three lightweight plug-and-play modules, namely, dense feature pyramid network (DenseFPN), spatial context pyramid (SCP), and hierarchical region of interest extractor (HRoIE), to aggregate global visual context at feature, spatial, and instance domains, respectively. DenseFPN is a multi-scale feature propagation module that establishes more flexible information flows by adopting interlevel residual connections, cross-level dense connections, and feature reweighting strategy. Leveraging the attention mechanism, SCP further augments the features by aggregating global spatial context into local regions. For each instance, HRoIE adaptively generates RoI features for different downstream tasks. Extensive evaluations of the proposed scheme on iSAID, DIOR, NWPU VHR-10, and HRSID datasets demonstrate that the proposed approach outperforms state-of-the-arts under similar computational costs. Source code and pretrained models are available at https://github.com/yeliudev/CATNet. Ye Liu 0002, Huifang Li 0001, Chang Wen Chen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Teaching Masked Autoencoder With Strong AugmentationsabstractMasked autoencoder (MAE) has been regarded as a capable self-supervised learner for various downstream tasks. Nevertheless, the model still lacks high-level discriminability, which results in poor linear probing performance. In view of the fact that strong augmentation plays an essential role in contrastive learning, can we capitalize on strong augmentation in MAE? The difficulty originates from the pixel uncertainty caused by strong augmentation that may affect the reconstruction, and thus, directly introducing strong augmentation into MAE often hurts the performance. In this article, we delve into the potential of strong augmented views to enhance MAE while maintaining MAE's advantages. To this end, we propose a simple yet effective masked Siamese autoencoder (MSA) model, which consists of a student branch and a teacher branch. The student branch derives MAE's advanced architecture, and the teacher branch treats the unmasked strong view as an exemplary teacher to impose high-level discrimination onto the student branch. We demonstrate that our MSA can improve the model's spatial perception capability and, therefore, globally favors interimage discrimination. Empirical evidence shows that the model pretrained by MSA provides superior performances across different downstream tasks. Notably, linear probing performance on frozen features extracted from MSA leads to 6.1% gains over MAE on ImageNet-1k. Fine-tuning (FT) the network on VQAv2 task finally achieves 67.4% accuracy, outperforming 1.6% of the supervised method DeiT and 1.2% of MAE. Codes and models are available at https://github.com/KimSoybean/MSA. Rui Zhu 0014, Yalong Bai, Ting Yao 0003, Jingen Liu, Zhenglong Sun 0001, Tao Mei 0001, Chang Wen Chen |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Streaming 360° VR Video With Statistical QoS Provisioning in mmWave Networks From Delay and Rate PerspectivesabstractMillimeter-wave$\!$(mmWave) technology has emerged as a pivotal catalyst for unleashing the full potential of 360° virtual reality (VR). Nonetheless, the explosive growth of VR services, combined with the quality-of-service (QoS) provisioning issues of mmWave, poses formidable challenges in wireless resource allocation for mmWave-enabled 360° VR. In this paper, we propose an innovative 360° VR streaming architecture that addresses three underexplored issues: overlapping fields-of-view (FoVs), statistical QoS provisioning (SQP), and loss-tolerant active data discarding. Specifically, we first design an overlapping FoV-based optimal joint unicast and multicast (JUM) task assignment scheme, which significantly conserves wireless resources by implementing non-redundant task allocation. Leveraging stochastic network calculus (SNC), we develop a comprehensive SNC-based SQP theoretical framework from delay and rate perspectives. Additionally, we propose two corresponding optimal adaptive joint resource allocation and active-discarding (ADAPT-JRAAD) transmission schemes to minimize resource consumption while guaranteeing SQP performance from delay and rate perspectives, respectively. Extensive simulations demonstrate the outstanding performance of the designed optimal JUM task assignment scheme in conserving wireless resources. Moreover, comprehensive comparisons against six benchmarks validate the superiority of the proposed two ADAPT-JRAAD schemes in resource utilization, flexible rate control, and robust queue management. Hancheng Lu, Langtian Qin, Chang Wu 0006, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 5 |
| 2024 | SD-DiT: Unleashing the Power of Self-Supervised Discrimination in Diffusion Transformer*abstractDiffusion Transformer (DiT) has emerged as the new trend of generative diffusion models on image generation. In view of extremely slow convergence in typical DiT, recent breakthroughs have been driven by mask strategy that significantly improves the training efficiency of DiT with additional intra-image contextual learning. Despite this progress, mask strategy still suffers from two inherent limitations: (a) training-inference discrepancy and (b) fuzzy relations between mask reconstruction & generative diffusion process, resulting in sub-optimal training of DiT. In this work, we address these limitations by novelly unleashing the self-supervised discrimination knowledge to boost DiT training. Technically, we frame our DiT in a teacher-student manner. The teacher-student discriminative pairs are built on the diffusion noises along the same Probability Flow Ordinary Differential Equation (PF-ODE). Instead of applying mask reconstruction loss over both DiT encoder and decoder, we decouple DiT encoder and decoder to separately tackle discriminative and generative objectives. In particular, by encoding discriminative pairs with student and teacher DiT encoders, a new discriminative loss is designed to encourage the inter-image alignment in the selfsupervised embedding space. After that, student samples are fed into student DiT decoder to perform the typical generative diffusion task. Extensive experiments are conducted on ImageNet dataset, and our method achieves a competitive balance between training cost and generative capacity. Rui Zhu 0014, Yingwei Pan, Yehao Li, Ting Yao 0003, Zhenglong Sun 0001, Tao Mei 0001, Chang Wen Chen |
CVPR | 7 |
| 2024 | Expanding Scene Graph Boundaries: Fully Open-Vocabulary Scene Graph Generation via Visual-Concept Alignment and Retention
Zuyao Chen, Jinlin Wu, Zhen Lei 0001, Zhaoxiang Zhang 0001, Chang Wen Chen |
ECCV (66) | 5 |
| 2024 | $\mathrm R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding
Ye Liu 0002, Jixuan He, Wanhua Li 0001, Junsik Kim 0001, Donglai Wei 0001, Hanspeter Pfister, Chang Wen Chen |
ECCV (41) | 7 |
| 2024 | Towards Omniscient Feature Alignment for Video RescalingabstractVideo super-resolution often reconstructs high-resolution (HR) video from low-resolution (LR) video that has been downsampled using predefined methods, which is an ill-posedness problem. Recent video rescaling algorithms alleviate this problem by jointly training the downsampling and upsampling processes. However, they primarily exploit the shallow temporal correlations among video frames, overlooking the intricate, long-term sequential depth dependencies within the video. In this paper, we propose an omniscient feature alignment to leverage the bidirectional deep temporal information for video rescaling, namely OFA-VRN. In the downsampling phase, the proposed method separates the input HR video into LR frames and high-frequency components using haar wavelet transform and explicitly embeds the high-frequency components into the LR frames. In this way, detailed information is stored in the frame and maintains visual perception quality in downsampled videos. During the upsampling phase, we use an advanced bidirectional propagation paradigm to enhance temporal information aggregation capabilities. By incorporating the proposed omniscient feature alignment, the network is capable of leveraging multi-frame feature information from the triplet dimension to further alleviate misalignment issues, thereby enhancing its capacity for deep temporal information utilization. The experiments on Vid4 and Vimeo90K-T demonstrate that our model achieves competitive performance compared to the state-of-the-art methods. Guanchen Ding, Chang Wen Chen |
ICASSP | 2 |
| 2024 | Removing Reflective Flare in Real-World ConditionsabstractThe increasing prevalence of mobile devices has led to significant advancements in mobile camera systems and improved image quality. Nonetheless, mobile photography still grapples with flare corruptions such as reflective flare. The absence of a comprehensive real image dataset tailored for mobile phones hinders the development of effective flare mitigation techniques. To address this issue, we present a novel real image dataset specifically designed for mobile camera systems, focusing on flare removal. Capitalizing on the distinct properties of real images, this dataset serves as a solid foundation for developing advanced flare removal algorithms. The dataset comprises over 1,100 pairs of high-quality, full-resolution images for reflective flare, which generate 2,200 paired patches, ensuring broad adaptability across various imaging conditions. Experimental results demonstrate that networks trained with synthesized data struggle to cope with the complex lighting settings present in this real image dataset. Our dataset is expected to enable an array of new research in flare removal and contribute to substantial improvements in mobile image quality, benefiting mobile photographers and end-users alike. Fengbo Lan, Chang Wen Chen |
ICIP | 2 |
| 2024 | Domain-Agnostic Crowd Counting via Uncertainty-Guided Style Diversity AugmentationabstractDomain shift significantly hinders crowd counting performance in unseen domains. Domain adaptation methods tackle this issue using target domain images but falter when acquiring these images is difficult. Moreover, they demand additional training time for fine-tuning. To address this issue, we propose an Uncertainty-Guided Style Diversity Augmentation (UGSDA) method, enabling the models to be trained solely on the source domain and directly generalized to various target domains. It is achieved by generating sufficiently diverse and realistic samples during the training process. Specifically, our UGSDA method incorporates three tailor-designed components: the Global Styling Elements Extraction (GSEE) module, the Local Uncertainty Perturbations (LUP) module, and the Density Distribution Consistency (DDC) loss. The GSEE extracts global style elements from the feature space of the whole source domain. The LUP aims to obtain uncertainty perturbations from the batch-level input to form style distributions beyond the source domain, which used to generate diversified stylized samples together with global style elements. To regulate the extent of perturbations, the DDC loss imposes constraints between the source samples and the stylized samples, ensuring the stylized samples maintain a higher degree of realism and reliability. Comprehensive experiments validate the superiority of our approach, demonstrating its strong generalization capabilities across various datasets and models. Code is available at https://github.com/gcding/UGSDA-pytorch. Guanchen Ding, Lingbo Liu, Zhenzhong Chen 0001, Chang Wen Chen |
ACM Multimedia | 4 |
| 2024 | Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment RetrievalabstractIn this paper, we explore the use of large language models (LLMs) to enhance video moment retrieval (VMR) by integrating general knowledge and pseudo-events as priors. We address the limitations of LLMs in generating continuous outputs, such as salience scores and inter-frame embeddings, which are critical for capturing inter-frame relations. To address these limitations, we propose using LLM encoders, which refine inter-concept relations in multimodal embeddings effectively, even without textual training. Our feasibility study shows that this capability extends to other embeddings like BLIP and T5 when they exhibit similar patterns to CLIP embeddings. We present a general framework for integrating LLM encoders into existing VMR architectures, specifically within the fusion module. The LLM encoder's ability to refine concept relation can help the model to achieve a balanced understanding of the foreground concepts (e.g., persons, faces) and background concepts (e.g., street, mountains) rather focusing only on the visually dominant foreground concepts. Additionally, we utilize pseudo-events, identified via event detection, to guide accurate moment prediction within event boundaries, reducing distractions from adjacent moments. Our plug-in approach for semantic refinement and pseudo-event regulation demonstrates state-of-the-art VMR performance through experimental validation. The source code can be accessed at https://github.com/fletcherjiang/LLMEPET. Wengyu Zhang, Xulu Zhang, Xiaoyong Wei, Chang Wen Chen, Qing Li 0001 |
ACM Multimedia | 5 |
| 2024 | Understanding and Tackling Scattering and Reflective Flare for Mobile Camera SystemsabstractThe rise of mobile devices has spurred advancements in camera technology and image quality. However, mobile photography still faces issues like scattering and reflective flares. While previous research has acknowledged the negative impact of the mobile devices' internal image signal processing pipeline (ISP) on image quality, the specific ISP operations that hinder flare removal have not been fully identified. In addition, current solutions only partially address ISP-related deterioration due to a lack of comprehensive raw image datasets for flare study. To bridge these research gaps, we introduce a new raw image dataset tailored for mobile camera systems, focusing on eliminating flare. This dataset encompasses over 2,000 high-quality, full-resolution raw image pairs for scattering flare, and 1,200 for reflective flare, captured across various real-world scenarios, mobile devices, and camera settings. It is designed to enhance the generalizability of flare removal algorithms across a wide spectrum of conditions. Through detailed experiments, we have identified that ISP operations, such as denoising, compression, and sharpening, may either improve or obstruct flare removal, offering critical insights into optimizing ISP configurations for better flare mitigation. Our dataset is poised to advance the understanding of flare-related challenges, enabling more precise incorporation of flare removal steps into the ISP. Ultimately, this work paves the way for significant improvements in mobile image quality, benefiting both enthusiasts and professional mobile photographers alike. Fengbo Lan, Chang Wen Chen |
ACM Multimedia | 2 |
| 2024 | Bridging the Gap: Sketch-Aware Interpolation Network for High-Quality Animation Sketch InbetweeningabstractHand-drawn 2D animation workflow is typically initiated with the creation of sketch keyframes. Subsequent manual inbetweens are crafted for smoothness, which is a labor-intensive process and the prospect of automatic animation sketch interpolation has become highly appealing. Yet, common frame interpolation methods are generally hindered by two key issues: 1) limited texture and colour details in sketches, and 2) exaggerated alterations between two sketch keyframes. To overcome these issues, we propose a novel deep learning method - Sketch-Aware Interpolation Network (SAIN). This approach incorporates multi-level guidance that formulates region-level correspondence, stroke-level correspondence and pixel-level dynamics. A multi-stream U-Transformer is then devised to characterize sketch inbetweening patterns using these multi-level guides through the integration of self / cross-attention mechanisms. Additionally, to facilitate future research on animation sketch inbetweening, we constructed a large-scale dataset - STD-12K, comprising 30 sketch animation series in diverse artistic styles. Comprehensive experiments on this dataset convincingly show that our proposed SAIN surpasses the state-of-the-art interpolation methods. Our code and dataset are avaliable in https://github.com/none-master/SAIN. Kun Hu 0008, Wei Bao 0001, Chang Wen Chen, Zhiyong Wang 0001 |
ACM Multimedia | 4 |
| 2024 | Semantic-aware Next-Best-View for Multi-DoFs Mobile System in Search-and-Acquisition based Visual PerceptionabstractEfficient visual perception using mobile systems is crucial, particularly in unknown environments such as search and rescue operations, where swift and comprehensive perception of objects of interest is essential. In such real-world applications, objects of interest are often situated in complex settings, making the selection of the 'Next Best' view based solely on maximizing visibility gain suboptimal. We argue that incorporating semantics-providing a higher-level interpretation of perception-can significantly contribute to the selection of viewpoints for various perception tasks. In this study, we formulate a novel information gain that integrates both visibility and semantic gain in a unified form to select the semantic-aware Next-Best-View. We also design an adaptive strategy with termination criterion to facilitate the two-stage search-and-acquisition manoeuvre on multiple objects of interest aided by a multi-degree-of-freedoms (Multi-DoFs) mobile system. To evaluate our approach, we introduce several semantically relevant reconstruction metrics, including perspective directivity and the region of interest (ROI)-to-full reconstruction volume ratio. Simulation experiments demonstrate that our approach outperforms the existing methods by up to 27.46% in the ROI-to-full reconstruction volume ratio and 0.88234 in average perspective directivity. Furthermore, the planned motion trajectory exhibits better perceiving coverage toward the target. Xiaotong Yu, Chang Wen Chen |
ACM Multimedia | 2 |
| 2024 | E.T. Bench: Towards Open-Ended Event-Level Video-Language UnderstandingabstractRecent advances in Video Large Language Models (Video-LLMs) have demonstrated their great potential in general-purpose video understanding. To verify the significance of these models, a number of benchmarks have been proposed to diagnose their capabilities in different scenarios. However, existing benchmarks merely evaluate models through video-level question-answering, lacking fine-grained event-level assessment and task diversity. To fill this gap, we introduce E.T. Bench (Event-Level & Time-Sensitive Video Understanding Benchmark), a large-scale and high-quality benchmark for open-ended event-level video understanding. Categorized within a 3-level task taxonomy, E.T. Bench encompasses 7.3K samples under 12 tasks with 7K videos (251.4h total length) under 8 domains, providing comprehensive evaluations. We extensively evaluated 8 Image-LLMs and 12 Video-LLMs on our benchmark, and the results reveal that state-of-the-art models for coarse-level (video-level) understanding struggle to solve our fine-grained tasks, e.g., grounding event-of-interests within videos, largely due to the short video context length, improper time representations, and lack of multi-event training data. Focusing on these issues, we further propose a strong baseline model, E.T. Chat, together with an instruction-tuning dataset E.T. Instruct 164K tailored for fine-grained event-level understanding. Our simple but effective solution demonstrates superior performance in multiple scenarios. Ye Liu 0002, Zongyang Ma, Zhongang Qi, Yang Wu 0001, Ying Shan, Chang Wen Chen |
NeurIPS | 6 |
| 2024 | Pin-CasNet: Detecting pin status in transmission lines based on cascade network
Fang Gao 0001, Rongwei Zhang, Jingfeng Tang, Shaomin Liu, Jun Yu 0001, Chang Wen Chen, Hanbo Zheng |
Eng. Appl. Artif. Intell. | 7 |
| 2024 | EGCN++: A New Fusion Strategy for Ensemble Learning in Skeleton-Based Rehabilitation Exercise AssessmentabstractSkeleton-based exercise assessment focuses on evaluating the correctness or quality of an exercise performed by a subject. Skeleton data provide two groups of features (i.e., position and orientation), which existing methods have not fully harnessed. We previously proposed an ensemble-based graph convolutional network (EGCN) that considers both position and orientation features to construct a model-based approach. Integrating these types of features achieved better performance than available methods. However, EGCN lacked a fusion strategy across the data, feature, decision, and model levels. In this paper, we present an advanced framework, EGCN++, for rehabilitation exercise assessment. Based on EGCN, a new fusion strategy called MLE-PO is proposed for EGCN++; this technique considers fusion at the data and model levels. We conduct extensive cross-validation experiments and investigate the consistency between machine and human evaluations on three datasets: UI-PRMD, KIMORE, and EHE. Results demonstrate that MLE-PO outperforms other EGCN ensemble strategies and representative baselines. Furthermore, the MLE-PO's model evaluation scores are more quantitatively consistent with clinical evaluations than other ensemble strategies. Bruce X. B. Yu, Yan Liu 0004, Keith C. C. Chan, Chang Wen Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | GLOCAL: A self-supervised learning framework for global and local motion estimation
Yihao Zheng 0002, Kunming Luo, Shuaicheng Liu, Zun Li 0001, Ye Xiang, Lifang Wu, Bing Zeng 0001, Chang Wen Chen |
Pattern Recognit. Lett. | 8 |
| 2024 | Un-Gaze: A Unified Transformer for Joint Gaze-Location and Gaze-Object DetectionabstractThis paper proposes an efficient and effective method for joint gaze location detection (GL-D) and gaze object detection (GO-D), i.e., gaze following detection. Current approaches frame GL-D and GO-D as two separate tasks, employing a multi-stage framework where human head crops must first be detected and then be fed into a subsequent GL-D sub-network, which is further followed by an additional object detector for GO-D. In contrast, we reframe the gaze following detection task as detecting human head locations and their gaze followings simultaneously, aiming at jointly detect human gaze location and gaze object in a unified and single-stage pipeline. To this end, we propose GTR, short for Gaze following detection TRansformer, streamlining the gaze following detection pipeline by eliminating all additional components, leading to the first unified paradigm that unites GL-D and GO-D in a fully end-to-end manner. GTR enables an iterative interaction between holistic semantics and human head features through a hierarchical structure, inferring the relations of salient objects and human gaze from the global image context and resulting in an impressive accuracy. Concretely, GTR achieves a 12.1 mAP gain ($\mathbf {25.1}\%$) on GazeFollowing and a 18.2 mAP gain ($\mathbf {43.3\%}$) on VideoAttentionTarget for GL-D, as well as a 19 mAP improvement ($\mathbf {45.2\%}$) on GOO-Real for GO-D. Meanwhile, unlike existing systems detecting gaze following sequentially due to the need for a human head as input, GTR has the flexibility to comprehend any number of people’s gaze followings simultaneously, resulting in high efficiency. Specifically, GTR introduces over a$\times 9$improvement in FPS and the relative gap becomes more pronounced as the human number grows. Danyang Tu, Wei Shen 0002, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Knowledge Augmented Relation Inference for Group Activity RecognitionabstractGroup activity recognition is a challenging task because it involves diverse individual actions and complex relations. Most existing methods enhance individual representation by introducing relation inference using appearance features. Some methods utilize extra knowledge, such as action labels, to enhance relation inference and refine the individual representation, but the knowledge they explored is simple and insufficient. In this paper, we propose a novel idea of knowledge concretization and further develop a Knowledge Augmented Relation Inference framework (KARI) for group activity recognition. Specifically, we first concretize knowledge from training data, and then represent them as Class-Class co-occurrence Map (C-C Map) and Class-Position distribution Map (C-P Map). On top of them, KARI explores concretized knowledge to integrate visual and semantic representation in a unified architecture for group activity recognition. Experimental results on two public datasets show that the proposed framework performs favorably compared with state-of-the-art approaches. Zhuming Wang, Zun Li 0001, Xianglong Lang, Yihao Zheng 0002, Lifang Wu, Liang Wang 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | Target-Aware Camera Placement for Large-Scale Video SurveillanceabstractIn large-scale surveillance of urban or rural areas, an effective placement of cameras is critical in maximizing surveillance coverage or minimizing economic cost of cameras. Existing Surveillance Camera Placement (SCP) methods generally focus on physical coverage of surveillance by implicitly assuming uniform distribution of interested targets or objects across all blocks, which is, however, uncommon in real-world scenarios. In this paper, we are the first to propose a target-aware SCP (tSCP) model, which prioritizes optimizing the task based on uneven target densities, allowing cameras to preferentially cover blocks with more interested targets. First, we define target density as the likelihood of interested targets occurring in a block, which is positively correlated with the importance of the block. Second, we combine aerial imagery with a lightweight object detection network to identify target density. Third, we formulate tSCP as an optimization problem to maximize target coverage in surveillance area, and solve this problem with a target-guided genetic algorithm. Our method optimizes the rational and economical utilization of cameras in large-scale video survillance. Compared with the state-of-the-art methods, our tSCP achieves the highest target coverage with a fixed number of cameras (8.31%-14.81% more than its peers), or utilizes the minimum number of cameras to achieve a preset target coverage. Codes are available athttps://github.com/wu-hongxin/tSCP_main. Hongxin Wu, Qinghou Zeng, Tiesong Zhao, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Self-Supervised Video Representation Learning by Serial Restoration With Elastic ComplexityabstractSelf-supervised video representation learning leaves out heavy manual annotation by automatically excavating supervisory signals. Although contrastive learning based approaches exhibit superior performances, pretext task based approaches still deserve further study. This is because the pretext tasks exploit the nature of data and encourage feature extractors to learn spatiotemporal logic by discovering dependencies among video clips or cubes, without manual engineering on data augmentations or manual construction of contrastive pairs. To utilize chronological property more effectively and efficiently, this work proposes a novel pretext task, named serial restoration of shuffled clips (SRSC), disentangled by an elaborately designed task network composed of an order-aware encoder and a serial restoration decoder. In contrast to other order based pretext tasks that formulate clip order recognition as a one-step classification problem, the proposed SRSC task restores shuffled clips into the right order in multiple steps. Owing to the excellent elasticity of SRSC, a novel taxonomy of curriculum learning is further proposed to equip SRSC with different pre-training strategies. According to the factors that affect the complexity of solving the SRSC task, the proposed curriculum learning strategies can be categorized into task based, model based and data based. Extensive experiments are conducted on the subdivided strategies to explore their effectiveness and noteworthy laws. Compared with existing approaches, this work demonstrates that the proposed approach achieves state-of-the-art performances in pretext task based self-supervised video representation learning and a majority of the proposed strategies further boost the performance of downstream tasks. For the first time, the features pre-trained by the pretext tasks are applied to video captioning by feature-level early fusion, and enhance the input of existing approaches as a lightweight plugin. Hanli Wang, Chang Wen Chen |
IEEE Trans. Multim. | 3 |
| 2024 | Hybrid Graph Reasoning With Dynamic Interaction for Visual DialogabstractAs a pivotal branch of intelligent human-computer interaction, visual dialog is a technically challenging task that requires artificial intelligence (AI) agents to answer consecutive questions based on image content and history dialog. Despite considerable progresses, visual dialog still suffers from two major problems: (1) how to design flexible cross-modal interaction patterns instead of over-reliance on expert experience and (2) how to infer underlying semantic dependencies between dialogues effectively. To address these issues, an end-to-end framework employing dynamic interaction and hybrid graph reasoning is proposed in this work. Specifically, three major components are designed and the practical benefits are demonstrated by extensive experiments. First, a dynamic interaction module is developed to automatically determine the optimal modality interaction route for multifarious questions, which consists of three elaborate functional interaction blocks endowed with dynamic routers. Second, a hybrid graph reasoning module is designed to explore adequate semantic associations between dialogues from multiple perspectives, where the hybrid graph is constructed by aggregating a structured coreference graph and a context-aware temporal graph. Third, a unified one-stage visual dialog model with an end-to-end structure is developed to train the dynamic interaction module and the hybrid graph reasoning module in a collaborative manner. Extensive experiments on the benchmark datasets of VisDial v0.9 and VisDial v1.0 demonstrate the effectiveness of the proposed method compared to other state-of-the-art approaches. The source code of this work can be found inhttps://mic.tongji.edu.cn. Shanshan Du, Hanli Wang, Tengpeng Li, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2024 | OARNet: Object-Attribute-Relation Network for Predicting Soccer EventsabstractEvent prediction involves analyzing and forecasting events that occur at a specific time and location to inform decision-making and take the next actions. Current event prediction approaches primarily employ deep learning methods to analyze regular patterns from large amounts of historical data. However, predicting adversarial soccer events remains a significant challenge due to strong antagonism and complex relationships between players. With this consideration, we propose an objectattribute-relation (OAR) network for predicting soccer events using multimodal data, including spatiotemporal trajectory data and video data. The proposed scheme aims to enhance prediction performance by transforming multimodal data into an OAR space that integrates global and local relationships (adversarial information and multi-objective information). In particular, the scheme consists mainly of a relation module, an object attribute module, and a graph prediction module. We first use ConvLSTM to extract the visual features of players from video data and use LSTM to extract the movement features of players from spatiotemporal data. Additionally, we apply a multihead GRU attention mechanism to calculate the relation weights. These three components are then combined into an OAR graph of a clip in a soccer game. Finally, an OAR GNN is designed to determine the influence of different objects and predict events. The entire process constitutes an end-to-end event prediction learning framework. Extensive experimental results on the two challenging datasets, namely, soccER and SkillCorner, verify the effectiveness of the proposed framework. Yiping Duan, Xiaoming Tao 0001, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2024 | Learned Video Compression via Heterogeneous Deformable Compensation NetworkabstractLearned video compression has recently emerged as an essential research topic in developing advanced video compression technologies, where motion compensation is considered one of the most challenging issues. In this article, we propose a learned video compression framework via heterogeneous deformable compensation strategy (HDCVC) to tackle the problems of unstable compression performance caused by single-size deformable kernels in downsampled feature domain. More specifically, instead of utilizing optical flow warping or single-size-kernel deformable alignment, the proposed algorithm extracts features from the two adjacent frames to estimate content-adaptive heterogeneous deformable (HetDeform) kernel offsets. Then we align the features extracted from the reference frames with the HetDeform convolution to accomplish motion compensation. Moreover, we design a Spatial-Neighborhood-Conditioned Divisive Normalization (SNCDN) to reduce spatial statistic dependencies and achieve more effective data Gaussianization combined with the Generalized Divisive Normalization. Furthermore, we propose a multi-frame enhanced reconstruction module for exploiting context and temporal information for final quality enhancement. Experimental results indicate that HDCVC achieves superior performance than the recent state-of-the-art learned video compression approaches. Huairui Wang, Zhenzhong Chen 0001, Chang Wen Chen |
IEEE Trans. Multim. | 3 |
| 2024 | End-to-End Video Scene Graph Generation With Temporal Propagation TransformerabstractVideo scene graph generation has been an emerging research topic, which aims to interpret a video as a temporally-evolving graph structure by representing video objects as nodes and their relations as edges. Existing approaches predominantly follow a multi-step scheme, including frame-level object detection, relation recognition and temporal association. Although effective, these approaches neglect the mutual interactions between independent steps, resulting in a sub-optimal solution. We present a novel end-to-end framework for video scene graph generation, which naturally unifies object detection, object tracking, and relation recognition via a new Transformer structure, namely Temporal Propagation Transformer (TPT). Particularly, TPT extends the existing Transformer-based object detector (e.g., DETR) along the temporal dimension by involving a query propagation module, which can additionally associate the detected instances by identities across frames. A temporal dynamics encoder is then leveraged to dynamically enrich the features of the detected instances for relation recognition by attending to their historic states in previous frames. Meanwhile, the relation propagation strategy is devised to emphasize the temporal consistency of relation recognition results among adjacent frames. Extensive experiments conducted on VidHOI and Action Genome benchmarks demonstrate the superior performance of the proposed TPT over the state-of-the-art methods. Yong Zhang 0056, Yingwei Pan, Ting Yao 0003, Rui Huang 0001, Tao Mei 0001, Chang Wen Chen |
IEEE Trans. Multim. | 6 |
| 2024 | Joint Identity-Aware Mixstyle and Graph-Enhanced Prototype for Clothes-Changing Person Re-IdentificationabstractIn recent years, considerable progress has been witnessed in the person re-identification (Re-ID). However, in a more realistic long-term scenario, the appearance shift arising from the clothes-changing inevitably deteriorates the conventional methods that heavily depend on the clothing color. Although the current clothes-changing person Re-ID methods introduce external human knowledge (i.e, contour, mask) and sophisticated feature decoupling strategy to alleviate the clothing shift, they still face the risk of overfitting to clothing due to the limited clothing diversity of training set. To more efficiently and effectively promote the clothes-irrelevant feature learning, we present a novel joint Identity-aware Mixstyle and Graph-enhanced Prototype method for clothes-changing person Re-ID. Specifically, by treating the cloth-changing as fine-grained domain/style shift, the identity-aware mixstyle (IMS) is proposed from the perspective of domain generalization, which mixes the instance-level feature statistics of samples within each identity to synthesize novel and diverse clothing styles, while retaining the correspondence between synthesized samples and latent label space. By incorporating the IMS module, the more diverse styles can be exploited to train a clothing-shift robust model. To further reduce the feature discrepancy caused by clothing variations, the graph-enhanced prototype constraint (GEP) module is proposed to explore the graph similarity structure of style-augmented samples across memory bank to build informative and robust prototypes, which serve as powerful exemplars for better clothing-irrelevant metric learning. The two modules are integrated into a joint learning framework and benefit each other. The extensive experiments conducted on clothes-changing person Re-ID datasets validate the superiority and effectiveness of our method. In addition, our method also shows good universality and corruption robustness on other Re-ID tasks. Bin Liu 0016, Yan Lu 0001, Qi Chu 0001, Nenghai Yu, Chang Wen Chen |
IEEE Trans. Multim. | 6 |
| 2024 | Prompt-Based Learning for Unpaired Image CaptioningabstractUnpaired Image Captioning (UIC) has been developed to learn image descriptions from unaligned vision-language sample pairs. Existing works usually tackle this task using adversarial learning and visual concept reward based on reinforcement learning. However, these existing works were only able to learn limited cross-domain information in vision and language domains, which restrains the captioning performance of UIC. Inspired by the success of Vision-Language Pre-Trained Models (VL-PTMs) in this research, we attempt to infer the cross-domain cue information about a given image from the large VL-PTMs for the UIC task. This research is also motivated by recent successes of prompt learning in many downstream multi-modal tasks, including image-text retrieval and vision question answering. In this work, a semantic prompt is introduced and aggregated with visual features for more accurate caption prediction under the adversarial learning framework. In addition, a metric prompt is designed to select high-quality pseudo image-caption samples obtained from the basic captioning model and refine the model in an iterative manner. Extensive experiments on the COCO and Flickr30 K datasets validate the promising captioning ability of the proposed model. We expect that the proposed prompt-based UIC model will stimulate a new line of research for the VL-PTMs based captioning. Peipei Zhu, Xiao Wang 0014, Lin Zhu 0012, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 7 |
| 2024 | Boosting Scene Graph Generation with Contextual InformationabstractScene graph generation (SGG) has been developed to detect objects and their relationships from the visual data and has attracted increasing attention in recent years. Existing works have focused on extracting object context for SGG. However, very few works have attempted to exploit implicit contextual correlations among relationships of the objects. Furthermore, most existing SGG schemes rely on high-level features to predict the predicates while overlooking the potential inherent association of low-level features with the object relationships. We present in this article a novel scheme to capture enhanced contextual information for both objects and relationships. We design a Dual-branch Context Analysis Transformer (DCAT) architecture to extract both object context and relationship context from the visual data with dual transformer branches and then effectively fuse both high-level and low-level features by an adaptive approach to facilitate relationship prediction. Specifically, we first conduct feature representation learning to enrich relation representations by the visual, spatial, and linguistic feature extractors. Next, two transformer branches are designed to leverage the modeling of global associative interaction and mine the hidden association among objects and relationships. Then, we devise a novel feature disentangling method to decouple contextualized high-level features with guidance from the visual semantics. Finally, we develop a refined attention module to perform low-level feature recalibration for the refinement of the final predicate prediction. Experiments on Visual Genome and Action Genome datasets demonstrate the effectiveness of DCAT for both image and video SGG settings. Moreover, we also test the quality of the generated image scene graphs to verify the generalizability on downstream tasks like sentence-to-graph retrieval and image retrieval. Shiqi Sun 0002, Danlan Huang, Xiaoming Tao 0001, Chengkang Pan, Guangyi Liu 0001, Chang Wen Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Statistical QoS Provisioning Analysis and Performance Optimization in xURLLC-Enabled Massive MU-MIMO Networks: A Stochastic Network Calculus PerspectiveabstractIn this paper, fundamentals and performance tradeoffs of next-generation ultra-reliable and low-latency communication (xURLLC) are investigated from the perspective of stochastic network calculus (SNC). An xURLLC-enabled massive MU-MIMO system model has been developed to accommodate xURLLC features. By leveraging and promoting SNC, we provide a quantitative statistical quality of service (QoS) provisioning analysis and derive the closed-form expression of upper-bounded statistical delay violation probability (UB-SDVP). Based on the proposed theoretical framework, we formulate the UB-SDVP minimization problem, which is first degenerated into a one-dimensional integer-search problem by deriving the minimum error probability (EP) detector, and then efficiently solved by the integer-form Golden-Section search algorithm. Moreover, two novel concepts, EP-based effective capacity (EP-EC) and EP-based energy efficiency (EP-EE), have been defined to characterize the tail distributions and performance tradeoffs for xURLLC. Subsequently, we formulate the EP-EC and EP-EE maximization problems, and the EP-EC maximization problem is proven to be equivalent to the UB-SDVP minimization problem, while the EP-EE maximization problem is solved with a low-complexity outer-descent inner-search collaborative algorithm. Extensive simulations demonstrate that the proposed framework can reduce computational complexity compared to reference schemes and provide various tradeoffs and optimization performance of xURLLC concerning UB-SDVP, EP, EP-EC, and EP-EE. Hancheng Lu, Langtian Qin, Chenwu Zhang, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 5 |
| 2023 | Interventional Bag Multi-Instance Learning On Whole-Slide Pathological ImagesabstractMulti-instance learning (MIL) is an effective paradigm for whole-slide pathological images (WSIs) classification to handle the gigapixel resolution and slide-level label. Prevailing MIL methods primarily focus on improving the feature extractor and aggregator. However, one deficiency of these methods is that the bag contextual prior may trick the model into capturing spurious correlations between bags and labels. This deficiency is a confounder that limits the performance of existing MIL methods. In this paper, we propose a novel scheme, Interventional Bag Multi-Instance Learning (IBMIL), to achieve deconfounded bag-level prediction. Unlike traditional likelihood-based strategies, the proposed scheme is based on the backdoor adjustment to achieve the interventional training, thus is capable of suppressing the bias caused by the bag contextual prior. Note that the principle of IBMIL is orthogonal to existing bag MIL methods. Therefore, IBMIL is able to bring consistent performance boosting to existing schemes, achieving new state-of-the-art performance. Code is available at https://github.com/HHHedo/IBMIL. Tiancheng Lin 0001, Zhimiao Yu, Hongyu Hu, Yi Xu 0001, Chang Wen Chen |
CVPR | 5 |
| 2023 | Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic SpaceabstractScene graph generation (SGG) aims to abstract an image into a graph structure, by representing objects as graph nodes and their relations as labeled edges. However, two knotty obstacles limit the practicability of current SGG methods in real-world scenarios: 1) training SGG models requires time-consuming ground-truth annotations, and 2) the closed-set object categories make the SGG models limited in their ability to recognize novel objects outside of training corpora. To address these issues, we novelly exploit a powerful pre-trained visual-semantic space (VSS) to trigger language-supervised and open-vocabulary SGG in a simple yet effective manner. Specifically, cheap scene graph supervision data can be easily obtained by parsing image language descriptions into semantic graphs. Next, the noun phrases on such semantic graphs are directly grounded over image regions through region-word alignment in the pre-trained VSS. In this way, we enable open-vocabulary object detection by performing object category name grounding with a text prompt in this VSS. On the basis of visually-grounded objects, the relation representations are naturally built for relation recognition, pursuing open-vocabulary SGG. We validate our proposed approach with extensive experiments on the Visual Genome benchmark across various SGG scenarios (i.e., supervised / language-supervised, closed-set / open-vocabulary). Consistent superior performances are achieved compared with existing methods, demonstrating the potential of exploiting pre-trained VSS for SGG in more practical scenarios. Yong Zhang 0056, Yingwei Pan, Ting Yao 0003, Rui Huang 0001, Tao Mei 0001, Chang Wen Chen |
CVPR | 6 |
| 2023 | Being Comes from Not-Being: Open-Vocabulary Text-to-Motion Generation with Wordless TrainingabstractText-to-motion generation is an emerging and challenging problem, which aims to synthesize motion with the same semantics as the input text. However, due to the lack of diverse labeled training data, most approaches either limit to specific types of text annotations or require online optimizations to cater to the texts during inference at the cost of efficiency and stability. In this paper, we investigate offline open-vocabulary text-to-motion generation in a zero-shot learning manner that neither requires paired training data nor extra online optimization to adapt for unseen texts. Inspired by the prompt learning in NLP, we pretrain a motion generator that learns to reconstruct the full motion from the masked motion. During inference, instead of changing the motion generator, our method reformulates the input text into a masked motion as the prompt for the motion generator to “reconstruct” the motion. In constructing the prompt, the unmasked poses of the prompt are synthesized by a text-to-pose generator. To supervise the optimization of the text-to-pose generator, we propose the first text-pose alignment model for measuring the alignment between texts and 3D poses. And to prevent the pose generator from over-fitting to limited training texts, we further propose a novel wordless training mechanism that optimizes the text-to-pose generator without any training texts. The comprehensive experimental results show that our method obtains a significant improvement against the baseline methods. The code is available at https://github.com/junfanlin/oohmg. Junfan Lin, Jianlong Chang, Lingbo Liu, Guanbin Li, Liang Lin 0004, Qi Tian 0001, Chang Wen Chen |
CVPR | 7 |
| 2023 | GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular Videoabstract3D human pose estimation has been researched for decades with promising fruits. 3D human pose lifting is one of the promising research directions toward the task where both estimated pose and ground truth pose data are used for training. Existing pose lifting works mainly focus on improving the performance of estimated pose, but they usually underperform when testing on the ground truth pose data. We observe that the performance of the estimated pose can be easily improved by preparing good quality 2D pose, such as fine-tuning the 2D pose or using advanced 2D pose detectors. As such, we concentrate on improving the 3D human pose lifting via ground truth data for the future improvement of more quality estimated pose data. Towards this goal, a simple yet effective model called Global-local Adaptive Graph Convolutional Network (GLA-GCN) is proposed in this work. Our GLA-GCN globally models the spatiotemporal structure via a graph representation and backtraces local joint features for 3D human pose estimation via individually connected layers. To validate our model design, we conduct extensive experiments on three benchmark datasets: Human3.6M, HumanEva-I, and MPI-INF-3DHP. Experimental results show that our GLA-GCN1implemented with ground truth 2D poses significantly outperforms state-of-the-art methods (e.g., up to 3%, 17%, and 14% error reductions on Human3.6M, HumanEva-I, and MPI-INF-3DHP, respectively). Bruce X. B. Yu, Zhi Zhang 0004, Yongxu Liu 0003, Shenghua Zhong, Yan Liu 0004, Chang Wen Chen |
ICCV | 6 |
| 2023 | Internet of Video Things: Technical Challenges and Emerging ApplicationsabstractThe worldwide flourishing of the Internet of Things (IoT) in the past decade has enabled numerous new applications through the internetworking of a wide variety of devices and sensors. In recent years, visual sensors have seen a considerable boom in IoT systems because they are capable of providing richer and more versatile information. Internetworking of large-scale visual sensors has been named the Internet of Video Things (IoVT). IoVT has a new array of unique characteristics in terms of sensing, transmission, storage, and analysis, all are fundamentally different from the conventional IoT. These new characteristics of IoVT are expected to impose significant challenges on existing technical infrastructures. In this keynote talk, an overview of recent advances in various fronts of IoVT will be introduced and a broad range of technological and systematic challenges will be addressed. Several emerging IoVT applications will be discussed to illustrate the great potential of IoVT in a broad range of practical scenarios. Chang Wen Chen |
ACM Multimedia | 1 |
| 2023 | CLIP-Count: Towards Text-Guided Zero-Shot Object CountingabstractRecent advances in visual-language models have shown remarkable zero-shot text-image matching ability that is transferable to downstream tasks such as object detection and segmentation. Adapting these models for object counting, however, remains a formidable challenge. In this study, we first investigate transferring vision-language models (VLMs) for class-agnostic object counting. Specifically, we propose CLIP-Count, the first end-to-end pipeline that estimates density maps for open-vocabulary objects with text guidance in a zero-shot manner. To align the text embedding with dense visual features, we introduce a patch-text contrastive loss that guides the model to learn informative patch-level visual representations for dense prediction. Moreover, we design a hierarchical patch-text interaction module to propagate semantic information across different resolution levels of visual features. Benefiting from the full exploitation of the rich image-text alignment knowledge of pretrained VLMs, our method effectively generates high-quality density maps for objects-of-interest. Extensive experiments on FSC-147, CARPK, and ShanghaiTech crowd counting datasets demonstrate state-of-the-art accuracy and generalizability of the proposed method. Code is available: https://github.com/songrise/CLIP-Count. https://github.com/songrise/CLIP-Count. Ruixiang Jiang, Lingbo Liu, Chang Wen Chen |
ACM Multimedia | 3 |
| 2023 | Toward Human Perception-Centric Video Thumbnail GenerationabstractVideo thumbnail plays an essential role in summarizing video content into a compact and concise image for users to browse efficiently. However, automatically generating attractive and informative video thumbnails remains an open problem due to the difficulty of formulating human aesthetic perception and the scarcity of paired training data. This work proposes a novel Human Perception-Centric Video Thumbnail Generation (HPCVTG) to address these challenges. Specifically, our framework first generates a set of thumbnails using a principle-based system, which conforms to established aesthetic and human perception principles, such as visual balance in the layout and avoiding overlapping elements. Then rather than designing from scratch, we ask human annotators to evaluate some of these thumbnails and select their preferred ones. A Transformer-based Variational Auto-Encoder (VAE) model is firstly pre-trained with Model-Agnostic Meta-Learning (MAML) and then fine-tuned on these human-selected thumbnails. The exploration of combining the MAML pre-training paradigm with human feedback in training can reduce human involvement and make the training process more efficient. Extensive experimental results show that our HPCVTG framework outperforms existing methods in objective and subjective evaluations, highlighting its potential to improve the user experience when browsing videos and inspire future research in human perception-centric content generation tasks. The code and dataset will be released via https://github.com/yangtao2019yt/HPCVTG. Junfan Lin, Zhongang Qi, Yang Wu 0001, Ying Shan, Chang Wen Chen |
ACM Multimedia | 8 |
| 2023 | SGCL: Spatial guided contrastive learning on whole-slide pathological images
Tiancheng Lin 0001, Zhimiao Yu, Zengchao Xu, Hongyu Hu, Yi Xu 0001, Chang Wen Chen |
Medical Image Anal. | 6 |
| 2023 | Knowledge-Enriched Attention Network With Group-Wise Semantic for Visual StorytellingabstractAs a technically challenging topic, visual storytelling aims at generating an imaginary and coherent story with narrative multi-sentences from a group of relevant images. Existing methods often generate direct and rigid descriptions of apparent image-based contents, because they are not capable of exploring implicit information beyond images. Hence, these schemes could not capture consistent dependencies from holistic representation, impairing the generation of reasonable and fluent stories. To address these problems, a novel knowledge-enriched attention network with group-wise semantic model is proposed. Three main novel components are designed and supported by substantial experiments to reveal practical advantages. First, a knowledge-enriched attention network is designed to extract implicit concepts from external knowledge system, and these concepts are followed by a cascade cross-modal attention mechanism to characterize imaginative and concrete representations. Second, a group-wise semantic module with second-order pooling is developed to explore the globally consistent guidance. Third, a unified one-stage story generation model with encoder-decoder structure is proposed to simultaneously train and infer the knowledge-enriched attention network, group-wise semantic module and multi-modal story generation decoder in an end-to-end fashion. Substantial experiments on the visual storytelling datasets with both objective and subjective evaluation metrics demonstrate the superior performance of the proposed scheme as compared with other state-of-the-art methods. The source code of this work can be found in https://mic.tongji.edu.cn. Tengpeng Li, Hanli Wang, Bin He 0003, Chang Wen Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Active Spatial Positions Based Hierarchical Relation Inference for Group Activity RecognitionabstractGroup activity recognition aims to recognize behaviors characterized by multiple individuals within a scene. Existing schemes rely on individual relation inference and usually take the individuals as tokens. Essentially they select the most relevant region of the group activity from the entire image while filtering out irrelevant background noises. However, these schemes require individual bounding box labeling in both training and testing stages. Since individuals have usually been presented at one scale, multi-scale individuals cannot be combined in an effective way. In this paper, we present a novel end-to-end hierarchical relation inference framework based on active spatial positions for group activity recognition. This framework is designed to locate active spatial positions and use them as visual tokens to infer the relations for token embeddings. It requires individual bounding box labeling only in the training stage while automatically eliminating the background after locating active spatial positions from the entire scene. The hierarchical relations can be naturally inferred based on the visual tokens at different scales, contributing to further performance improvement. Experimental results demonstrate that the proposed framework is competitive against existing schemes that require more laboring and computation to generate labels in both the training and testing stage. Lifang Wu, Xianglong Lang, Ye Xiang, Chang Wen Chen, Zun Li 0001, Zhuming Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | PhaseAnti: An Anti-Interference WiFi-Based Activity Recognition System Using Interference-Independent Phase ComponentabstractDriven by a wide range of essential applications, significant achievements have recently been made to explore WiFi-based Human Activity Recognition (HAR) techniques that utilize the information collected by commercial off-the-shelf (COTS) WiFi infrastructures to infer human activities without the need for the subject to carry any devices. Although existing WiFi-based HAR systems achieve satisfactory performance in some instances, they are faced with a severe challenge that the impacts of ubiquitous Co-channel Interference (CCI) on WiFi signals are inevitable. This downgrades the performance of these HAR systems significantly. To address this challenge, we propose PhaseAnti in this paper, a novel WiFi-based HAR system to exploit the CCI-independent phase component, Nonlinear Phase Error Variation (NLPEV), of WiFi Channel State Information (CSI) to cope with the negative effects of CCI. The stability of NLPEV data and the sensibility of this component to motions are rigorously analyzed. Furthermore, validated by extensive properly designed experiments, this phase component across subcarriers is invariant under various CCI scenarios while sufficiently distinct for different motions. Therefore, the NLPEV data can be used and processed effectively to perform HAR in CCI scenarios. Extensive experiments with various daily activities in different indoor rooms demonstrate the superior effectiveness and generalizability of the proposed PhaseAnti system under various CCI scenarios. Specifically, PhaseAnti achieves a$ 96.5\%$recognition accuracy rate (RAR) on average in different CCI scenarios, which can improve up to a$ 16.7\%$RAR compared with the amplitude component in the presence of CCI. Furthermore, the recognition speed is 10.3 × faster than the state-of-the-art solution. Jinyang Huang, Bin Liu 0016, Chenglin Miao, Yan Lu 0001, Qijia Zheng, Yu Wu 0020, Jiancun Liu, Lu Su 0001, Chang Wen Chen |
IEEE Trans. Mob. Comput. | 9 |
| 2023 | Unpaired Image Captioning by Image-Level Weakly-Supervised Visual Concept RecognitionabstractThe goal of unpaired image captioning (UIC) is to describe images without using image-caption pairs in the training phase. Although challenging, we expect the task can be accomplished by leveraging images aligned with visual concepts. Most existing studies use off-the-shelf algorithms to obtain the visual concepts because the Bounding Box (BBox) labels or relationship-triplet labels used for training are expensive to acquire. To avoid exhaustive annotations, we propose a novel approach to achieve cost-effective UIC. Specifically, we adopt image-level labels to optimize the UIC model in a weakly-supervised manner. For each image, we assume that only the image-level labels are available without specific locations and numbers. The image-level labels are utilized to train a weakly-supervised object recognition model to extract object information (e.g., instance), and the extracted instances are adopted to infer the relationships among different objects using an enhanced graph neural network (GNN). The proposed approach achieves comparable or even better performance compared with previous methods without expensive annotations. Furthermore, we design an unrecognized object (UnO) loss to improve the alignment of the inferred object and relationship information with the images. It can effectively alleviate the issue encountered by existing UIC models when generating sentences with nonexistent objects. To the best of our knowledge, this is the first attempt to address the problem of Weakly-Supervised visual concept recognition for UIC (WS-UIC) based only on image-level labels. Extensive experiments demonstrate that the proposed method achieves inspiring results on the COCO dataset while significantly reducing the labeling cost. Peipei Zhu, Xiao Wang 0014, Yong Luo 0002, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 7 |
| 2023 | Boosting Scene Graph Generation with Visual Relation SaliencyabstractThe scene graph is a symbolic data structure that comprehensively describes the objects and visual relations in a visual scene, while ignoring the inherent perceptual saliency of each visual relation (i.e., relation saliency). However, humans often quickly allocate attention to important/salient visual relations in a scene. To align with such human perception of a scene, we explicitly model the perceptual saliency of visual relation in scene graph by upgrading each graph edge (i.e., visual relation) with an attribute of relation saliency. We present a new design, named as Saliency-guided Message Passing (SMP), that boosts the generation of such scene graph structure with the guidance from the visual relation saliency. Technically, an object interaction encoder is first utilized to strengthen object relation representations by jointly exploiting the appearance, semantic, and spatial relations in between. A branch is further leveraged to estimate the relation saliency of each visual relation by ordinal regression. Next, conditioned on the object and relation features (coupled with the estimated relation saliency), our SMP enhances scene graph generation by performing message passing over the objects and the most salient relations. Extensive experiments on VG-KR and VG150 datasets demonstrate the superiority of SMP for the scene graph generation. Moreover, we empirically validate the compelling generalizability of the learned scene graphs via SMP on downstream tasks like cross-model retrieval and image captioning. Yong Zhang 0056, Yingwei Pan, Ting Yao 0003, Rui Huang 0001, Tao Mei 0001, Chang Wen Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Reconfigurable Intelligent Surfaces-Enhanced Uplink User-Centric Networks on Energy Efficiency OptimizationabstractUser-centric network (UCN) has been considered as a promising technology to improve the throughput of users, especially users at the cell edge, where each user is connected to and served by a group of access points. With the rapid growth of user’s uplink traffic demands, energy consumption has become an important issue in uplink UCN as user’s battery capacity is limited. To address this issue, in this research, we develop a reconfigurable intelligent surface (RIS)-enhanced uplink UCN, where RIS is exploited to improve the signal quality of uplink transmission for users with virtually no extra energy consumption. To approach an optimal energy efficiency, we jointly perform reflect beamforming at RISs and uplink power control at users. This is formulated into a non-convex and intractable energy efficiency maximization problem. To solve it effectively, we decompose it into two subproblems and carry out the optimization alternately. That is, we design the reflect beamforming matrices at RISs through fractional programming and an uplink power control at users through constructing a surrogate function via majorization-minimization algorithm. Numerical results demonstrate that the proposed algorithm could achieve significant gain both on energy efficiency and spectral efficiency compared to the benchmark algorithm. Chenwu Zhang, Hancheng Lu, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 3 |
| 2022 | UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionabstractFinding relevant moments and highlights in videos according to natural language queries is a natural and highly valuable common need in the current video content explosion era. Nevertheless, jointly conducting moment retrieval and highlight detection is an emerging research topic, even though its component problems and some related tasks have already been studied for a while. In this paper, we present the first unified framework, named Unified Multi-modal Transformers (UMT), capable of realizing such joint optimization while can also be easily degenerated for solving individual problems. As far as we are aware, this is the first scheme to integrate multi-modal (visual-audio) learning for either joint optimization or the individual moment retrieval task, and tackles moment retrieval as a keypoint detection problem using a novel query generator and query decoder. Extensive comparisons with existing methods and ablation studies on QVHighlights, Charades-STA, YouTube Highlights, and TVSum datasets demonstrate the effectiveness, superiority, and flexibility of the proposed method under various settings. Source code and pre-trained models are available at https://github.com/TencentARC/UMT. Ye Liu 0002, Siyuan Li 0026, Yang Wu 0001, Chang Wen Chen, Ying Shan, Xiaohu Qie |
CVPR | 4 |
| 2022 | Exploring Structure-aware Transformer over Interaction Proposals for Human-Object Interaction DetectionabstractRecent high-performing Human-Object Interaction (HOI) detection techniques have been highly influenced by Transformer-based object detector (i.e., DETR). Nevertheless, most of them directly map parametric interaction queries into a set of HOI predictions through vanilla Transformer in a one-stage manner. This leaves rich interor intra-interaction structure under-exploited. In this work, we design a novel Transformer-style HOI detector, i.e., Structure-aware Transformer over Interaction Proposals (STIP), for HOI detection. Such design decomposes the process of HOI set prediction into two subsequent phases, i.e., an interaction proposal generation is first performed, and then followed by transforming the non-parametric interaction proposals into HOI predictions via a structure-aware Transformer. The structure-aware Transformer upgrades vanilla Transformer by encoding additionally the holistically semantic structure among interaction proposals as well as the locally spatial structure of human/object within each interaction proposal, so as to strengthen HOI predictions. Extensive experiments conducted on V-COCO and HICO-DET benchmarks have demonstrated the effectiveness of STIP, and superior results are reported when comparing with the state-of-the-art HOI detectors. Source code is available at https://github.com/zyong812/STIP. Yong Zhang 0056, Yingwei Pan, Ting Yao 0003, Rui Huang 0001, Tao Mei 0001, Chang Wen Chen |
CVPR | 6 |
| 2022 | Joint Front-Edge-Cloud IoVT Analytics: Resource-Effective Design and SchedulingabstractA tremendous amount of visual data are bing collected by the Internet of Video Things (IoVT) systems in which ubiquitous cameras deployed in cities enable new applications in the domains of smart transportation and public security. However, the limited resources in terms of communication, computing, and caching (3C) in the conventional cellular network make it challenging to adopt centralized artificial intelligence (AI) to conduct real-time video-based data analytics. In this work, based on the 5G network architecture with edge servers, a three-phase resource-effective solution is proposed to perform surveillance operations in a large-scale wireless IoVT. The proposed strategy integrates front-end cameras with simple on-chip neural networks performing real-time object-of-interest segmentation, edge servers, and cloud servers with AI functionality carrying out image-based target recognition and video-based target analytics tasks. More importantly, we design the optimal 3C strategy to achieve the best video analytics performance constrained by computing offload ratio, network resource allocation and video-related parameters. Extensive simulations with deep neural networks implemented both at the front-end cameras and in the cloud server have validated the effectiveness of the proposed solution. Youjia Chen, Tiesong Zhao, Peng Cheng 0002, Ming Ding 0001, Chang Wen Chen |
IEEE Internet Things J. | 5 |
| 2022 | Beyond fine-tuning: Classifying high resolution mammograms using function-preserving transformationsabstractThe task of classifying mammograms is very challenging because the lesion is usually small in the high resolution image. The current state-of-the-art approaches for medical image classification rely on using the de-facto method for convolutional neural networks-fine-tuning. However, there are fundamental differences between natural images and medical images, which based on existing evidence from the literature, limits the overall performance gain when designed with algorithmic approaches. In this paper, we propose to go beyond fine-tuning by introducing a novel framework called MorphHR, in which we highlight a new transfer learning scheme. The idea behind the proposed framework is to integrate function-preserving transformations, for any continuous non-linear activation neurons, to internally regularise the network for improving mammograms classification. The proposed solution offers two major advantages over the existing techniques. Firstly and unlike fine-tuning, the proposed approach allows for modifying not only the last few layers but also several of the first ones on a deep ConvNet. By doing this, we can design the network front to be suitable for learning domain specific features. Secondly, the proposed scheme is scalable to hardware. Therefore, one can fit high resolution images on standard GPU memory. We show that by using high resolution images, one prevents losing relevant information. We demonstrate, through numerical and visual experiments, that the proposed approach yields to a significant improvement in the classification performance over state-of-the-art techniques, and is indeed on a par with radiology experts. Moreover and for generalisation purposes, we show the effectiveness of the proposed learning scheme on another large dataset, the ChestX-ray14, surpassing current state-of-the-art techniques. Angelica I. Avilés-Rivero, Shuo Wang 0011, Yuan Huang 0009, Fiona J. Gilbert, Carola-Bibiane Schönlieb, Chang Wen Chen |
Medical Image Anal. | 7 |
| 2022 | LensCast: Robust Wireless Video Transmission Over MmWave MIMO With Lens Antenna ArrayabstractIn this paper, we present LensCast, a novel cross-layer video transmission framework for wireless networks, which seamlessly integrates millimeter wave (mmWave) lens multiple-input multiple-output (MIMO) with robust video transmission. LensCast is designed to exploit the video content diversity at the application layer, together with the spatial path diversity of lens antenna array at the physical layer, to achieve graceful video transmission performance under varying channel conditions. In LensCast, a transmission distortion minimization problem is formulated with the consideration of video chunk scheduling, path matching and power allocation, which is an intractable mixed integer non-linear programming (MINLP) problem. The solution of this MINLP problem is converted into resource allocation (i.e., joint path matching and power allocation) plus chunk scheduling. First, resource allocation is investigated with given chunk scheduling results. By analyzing the optimality of the resource allocation problem, a winner-takes-all assignment is obtained to guide resource allocation. After that, a greedy water-filling algorithm is proposed as a near-optimal solution. Second, we propose a low-complexity chunk scheduling algorithm to schedule chunks for each transmission. Simulation results demonstrate that the proposed LensCast achieves an improved performance in terms of both peak signal-to-noise ratio and visual quality comparing with reference schemes. Yongqiang Gui, Hancheng Lu, Feng Wu 0001, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2022 | TWGAN: Twin Discriminator Generative Adversarial NetworksabstractGenerative Adversarial Networks (GAN) has become more and more popular these years. However, it is difficult to train and suffers from the training instability problem. To tackle this difficulty, this paper proposes a novel approach. Our idea is intuitive but proven to be very useful. In essence, it combines saturating loss and non-saturating loss into the loss function. Thus it will exploit the complementary statistical properties from two kinds of loss functions to effectively improve the training stability. We term our method twin discriminator Generative Adversarial Networks (TWGAN), which, unlike GAN, has a generator and a twin discriminator. The twin discriminator consists of two discriminators with identical architecture and both of them aim to distinguish whether the samples are from real data or fake data. We develop theoretical analysis to show that, given the optimal discriminators, optimizing the generator of TWGAN reduces to minimizing the Kullback-Leibler (KL) divergence between the distribution of generated data ($P_g$) and the distribution of real data ($P_data$), hence effectively addressing the training instability problem. Extensive experiments on MNIST, Fashion MNIST, CIFAR-10/100 and STL-10 datasets demonstrate that the competitive performance of our TWGAN in generating good quality and diverse samples over baselines. The obtained highest inception score (IS) and lowest Fr$\acute{e}$chet Inception Distance (FID), compared with other state-of-the-art GANs, show the superiority of our TWGAN. Zhaoyu Zhang 0001, Haonian Xie, Jun Yu 0001, Tongliang Liu, Chang Wen Chen |
IEEE Trans. Multim. | 6 |
| 2021 | Improving Contrastive Learning by Visualizing Feature TransformationabstractContrastive learning, which aims at minimizing the distance between positive pairs while maximizing that of negative ones, has been widely and successfully applied in unsupervised feature learning, where the design of positive and negative (pos/neg) pairs is one of its keys. In this paper, we attempt to devise a feature-level data manipulation, differing from data augmentation, to enhance the generic contrastive self-supervised learning. To this end, we first design a visualization scheme for pos/neg score1distribution, which enables us to analyze, interpret and understand the learning process. To our knowledge, this is the first attempt of its kind. More importantly, leveraging this tool, we gain some significant observations, which inspire our novel Feature Transformation proposals including the extrapolation of positives. This operation creates harder positives to boost the learning because hard positives enable the model to be more view-invariant. Besides, we propose the interpolation among negatives, which provides diversified negatives and makes the model more discriminative. It is the first attempt to deal with both challenges simultaneously. Experiment results show that our proposed Feature Transformation can improve at least 6.0% accuracy on ImageNet-100 over MoCo baseline, and about 2.0% accuracy on ImageNet-1K over the MoCoV2 baseline. Transferring to the downstream tasks successfully demonstrate our model is less task-bias. Visualization tools and codes: https://github.com/DTennant/CL-Visualizing-Feature-Transformation. Rui Zhu 0014, Bingchen Zhao, Jingen Liu, Zhenglong Sun 0001, Chang Wen Chen |
ICCV | 5 |
| 2021 | Modularized Morphing of Deep Convolutional Neural Networks: A Graph ApproachabstractNetwork morphism is an effective learning scheme to morph a well-trained neural network to a new one with the network function completely preserved. However, existing network morphism scheme addresses only basic morphing types on the layer level. In this research, we address the central problem of network morphism at a higher level, i.e., how a convolutional layer can be morphed into an arbitrary module of a neural network. To simplify the representation of a network, we abstract a module as a graph with blobs as vertices and convolutional layers as edges. Based on this graph, the morphing process can be formulated as a graph transformation problem. Two atomic morphing operations are introduced to construct the graphs, based on which modules are classified into two families, i.e., simple morphable modules and complex modules. We present practical morphing solutions for both families, and prove that any module can be morphed from a single convolutional layer. Extensive experiments have been conducted based on the state-of-the-art ResNet on benchmarks to verify the effectiveness of the proposed solution. Changhu Wang, Chang Wen Chen |
IEEE Trans. Computers | 3 |
| 2021 | Joint Learning of Latent Similarity and Local Embedding for Multi-View ClusteringabstractSpectral clustering has been an attractive topic in the field of computer vision due to the extensive growth of applications, such as image segmentation, clustering and representation. In this problem, the construction of the similarity matrix is a vital element affecting clustering performance. In this paper, we propose a multi-view joint learning (MVJL) framework to achieve both a reliable similarity matrix and a latent low-dimensional embedding. Specifically, the similarity matrix to be learned is represented as a convex hull of similarity matrices from different views, where the nuclear norm is imposed to capture the principal information of multiple views and improve robustness against noise/outliers. Moreover, an effective low-dimensional representation is obtained by applying local embedding on the similarity matrix, which preserves the local intrinsic structure of data through dimensionality reduction. With these techniques, we formulate the MVJL as a joint optimization problem and derive its mathematical solution with the alternating direction method of multipliers strategy and the proximal gradient descent method. The solution, which consists of a similarity matrix and a low-dimensional representation, is ultimately integrated with spectral clustering or K-means for multi-view clustering. Extensive experimental results on real-world datasets demonstrate that MVJL achieves superior clustering performance over other state-of-the-art methods. Aiping Huang, Tiesong Zhao, Chang Wen Chen |
IEEE Trans. Image Process. | 4 |
| 2021 | A Multilayer Pyramid Network Based on Learning for Vehicle Logo RecognitionabstractIn this paper, we present a novel learning-based scheme for vehicle logo recognition (VLR). This scheme is termed Multilayer Pyramid Network Based on Learning (MLPNL) and is based on the principle that considering multiple resolutions is helpful for extracting valuable features that benefit the final recognition performance. The innovations of this scheme include (1) a multilayer pyramid network, with pixel difference matrices (PDMs) as its input and output and feature parameters mapping one PDM to another; (2) an objective function and a corresponding optimization method designed to facilitate the learning of the feature parameters of the proposed multilayer pyramid network; and (3) a multi-codebook-based encoding method that makes best use of the features extracted from PDMs corresponding to different resolutions. Extensive experiments conducted with an open dataset, HFUT-VL, demonstrate that the proposed MLPNL scheme outperforms state-of-the-art handcrafted descriptors and non-deep-learning-based learning methods when fewer training samples exist. Experiments conducted with a benchmark dataset, XMU, demonstrate that MLPNL outperforms existing state-of-the-art VLR methods. Experiments conducted both on HFUT-VL and XMU demonstrate that MLPNL is faster than most deep-learning-based learning methods while maintaining nearly the same recognition rate. Code has been made available at:https://github.com/HFUT-CV/MLPNL. Jun Wang 0071, Hai Min, Wei Jia 0001, Jun Yu 0001, Chang Wen Chen |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2021 | Robust Video Broadcast for Users With Heterogeneous Resolution in Mobile NetworksabstractRecently, robust video transmission system that can eliminate the cliff effect in digital video transmission has attracted great interest from both academia and industry. By linearizing the whole system, robust video transmission is intrinsically scalable to channel conditions in mobile networks. However, the heterogeneity of user devices in terms of viewing resolution has not been well studied for robust video broadcast systems. In this paper, we propose a spatial scalability enabled robust video broadcast (SSRVB) system, aiming at accommodating diverse users with both heterogeneous resolutions and heterogeneous channel conditions. In SSRVB, a novel spatial decomposition method based on linear projection is first designed for robust video transmission. Then the transmission distortion minimization problem with joint subcarrier matching and power allocation is formulated. A near-optimal low-complexity subcarrier matching algorithm based on auction theory and an optimal power allocation strategy are also proposed. Furthermore, an iterative algorithm is designed to solve the problem of joint resource allocation. Simulation results demonstrate that SSRVB can achieve an average of 3 dB gain when compared with the reference schemes (i.e., ECast, MCast, discrete wavelet transform (DWT) based scheme, and scalable video coding (SVC) scheme) in terms of average peak signal-to-noise ratio under heterogeneous scenarios. Yongqiang Gui, Hancheng Lu, Feng Wu 0001, Chang Wen Chen |
IEEE Trans. Mob. Comput. | 4 |
| 2021 | NOMA-Based Scalable Video Multicast in Mobile Networks With Statistical ChannelsabstractTo cope with rapid growth of video services, we propose a non-orthogonal multiple access (NOMA) based scalable video multicast (NOMA-SVM) framework for mobile networks, by exploiting NOMA's specific potential in scalable video multicast transmission. We consider statistical channels, instead of channels with perfect estimation, in the proposed NOMA-SVM framework in order to capture the realistic channel behaviors. As quality of experience (QoE) is a better metric than throughput for video transmission, QoE-driven power allocation is performed among multiple video layers in the proposed NOMA-SVM framework, in which users can decode video with quality proportional to their channel conditions. Specifically, we formulate the power allocation problem with the goal to maximize the average QoE over all users while guaranteeing the basic services of these users. To solve such a non-convex discrete problem, an optimal algorithm is developed based on the hidden monotonicity of the problem. A suboptimal algorithm is also proposed with much lower complexity in order to meet the practical needs. Simulation results show that the proposed algorithms outperform existing orthogonal multiple access (OMA) and NOMA based algorithms under various multicast scenarios in terms of QoE. Ming Zhang 0029, Hancheng Lu, Feng Wu 0001, Chang Wen Chen |
IEEE Trans. Mob. Comput. | 4 |
| 2021 | Learning Face Image Super-Resolution Through Facial Semantic Attribute Transformation and Self-Attentive Structure EnhancementabstractFace super-resolution is a domain-specific super-resolution (SR) problem of generating high-resolution (HR) face images from low-resolution (LR) inputs. Even though existing face SR methods have achieved great performance on the global region evaluation, most of them cannot restore local attributes and structure reasonably, especially to ultra-resolve tiny LR face images (16 × 16 pixels) to its larger version (8 × upscaling factor). In this paper, we propose an open source face SR framework based on facial semantic attribute transformation and self-attentive structure enhancement. Specifically, the proposed framework introduces face semantic information (i.e., face attributes) and face structure information (i.e., face boundaries) in a successive two-stage fashion. In the first stage, an Attribute Transformation Network (AT-Net) is established. It upsamples LR face images to HR feature maps and then combines facial attributes with these features to generate the intermediate HR results with rational attributes. In the second stage, a Structure Enhancement Network (SE-Net) is built. It simultaneously extracts face features and estimates facial boundary heatmaps from the inputs, and then fuses them to output the final HR face images. Extensive experiments demonstrate that our method achieves superior super-resolved results and outperforms the state-of-the-art methods. Zhaoyu Zhang 0001, Jun Yu 0001, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2021 | Accurate and Efficient Image Super-Resolution via Global-Local Adjusting Dense NetworkabstractConvolutional neural network-based (CNN-based) method has shown its superior performance on the image super-resolution (SR) task. However, several researches have shown that obtaining a better reconstruction result often leads to the significant increase in parameters and computation. To alleviate the burden in computational needs, we propose a novel global-local adjusting dense super-resolution network (GLADSR) to build a powerful yet lightweight CNN-based SR model. To enhance the network capacity, we present a global-local adjusting module (GLAM) which can adaptively reallocate the processing resources with local selective block (LSB) and global guided block (GGB). The GLAMs are linked with nested dense connections to make better use of the global-local adjusted features. In addition, we also introduce a separable pyramid upsampling (SPU) module to replace the regular upsampling operation, which thus brings a substantial reduction of its parameters and obtains better results. Furthermore, we show that the proposed refinement structure is capable of reducing image artifacts in SR processing. Extensive experiments on benchmark datasets show that the proposed GLADSR outperforms the state-of-the-art methods with much fewer parameters and much less computational cost. Sunxiangyu Liu, Kongya Zhao, Guitao Li, Liuguo Yin, Chang Wen Chen |
IEEE Trans. Multim. | 7 |
| 2021 | Association and Caching in Relay-Assisted mmWave Networks: A Stochastic Geometry PerspectiveabstractLimited backhaul bandwidth and blockage effects are two main factors limiting the practical deployment of millimeter wave (mmWave) networks. To tackle these issues, we study the feasibility of relaying as well as caching in mmWave networks. A user association and relaying (UAR) criterion dependent on both caching status and maximum biased received power is proposed by considering the spatial correlation caused by the coexistence of base stations (BSs) and relay nodes (RNs). Using stochastic geometry tools, we decouple the UAR and caching placement issues by analyzing the relationship between UAR probabilities and caching placement probabilities. We then optimize the formulated caching placement problem based on polyblock outer approximation by exploiting the monotonic property in the general case and utilizing convex optimization in the noise-limited case. Accordingly, we propose a BS and RN selection algorithm where caching status at BSs and maximum biased received power are jointly considered. Experimental results demonstrate a significant enhancement of backhaul offloading using the proposed algorithms, and show that deploying more RNs and increasing cache size in mmWave networks is a more cost-effective alternative than increasing BS density to achieve similar backhaul offloading performance. Zhuojia Gu, Hancheng Lu, Ming Zhang 0029, Haizhou Sun, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 5 |
| 2021 | QoS-Aware User Grouping Strategy for Downlink Multi-Cell NOMA SystemsabstractIn multi-cell non-orthogonal multiple access (NOMA) systems, designing an appropriate user grouping strategy is an open problem due to diverse quality of service (QoS) requirements and inter-cell interference. In this paper, we exploit both game theory and graph theory to study QoS-aware user grouping strategies, aiming at minimizing power consumption in downlink multi-cell NOMA systems. Under different QoS requirements, we derive the optimal successive interference cancellation (SIC) decoding order with inter-cell interference, which is different from existing SIC decoding order of increasing channel gains, and obtain the corresponding power allocation strategy. Based on this, the exact potential game model of the user grouping strategies adopted by multiple cells is formulated. We prove that, in this game, the problem for each player to find a grouping strategy can be converted into the problem of searching for specific negative loops in the graph composed of users. Bellman-Ford algorithm is expanded to find these negative loops. Furthermore, we design a greedy based suboptimal strategy to approach the optimal solution with polynomial time. Extensive simulations confirm the effectiveness of grouping users with consideration of QoS and inter-cell interference, and show that the proposed strategies can considerably reduce total power consumption comparing with reference strategies. Fengqian Guo, Hancheng Lu, Xiaoda Jiang, Ming Zhang 0029, Jun Wu 0006, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 6 |
| 2021 | Distortion-Aware Cross-Layer Power Allocation for Video Transmission Over Multi-User NOMA SystemsabstractNon-orthogonal multiple access (NOMA) is promising to enable growing video services with requirements of massive traffic and low latency. In this paper, we investigate a novel multi-user NOMA system design for video delivery. Different from data delivery considered in existing NOMA systems, video distortion is taken into consideration for the design and optimization of the proposed scheme. We first propose a cross-layer scalable video delivery scheme over NOMA wireless networks based on given user grouping. By adopting a semi-analytical rate distortion model for encoded video sequences and investigating the physical layer model for the multi-carrier NOMA network, the distortion-aware power allocation problem is formulated with the goal to minimize the system end-to-end distortion with quality-of-service requirements. To solve such a non-convex problem, we design an efficient algorithm by exploiting its hidden monotonic property to approximate the optimal solution. Inspired by the decoding order in NOMA, a fast algorithm is developed by combining water-filling and greedy strategies to further tame the computational complexity. Simulations results have verified the advantages of the proposed cross-layer scheme with distortion-aware power allocation algorithms, against two existing NOMA schemes and an OMA scheme. Hancheng Lu, Xiaoda Jiang, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 3 |
| 2020 | ConsNet: Learning Consistency Graph for Zero-Shot Human-Object Interaction DetectionabstractWe consider the problem of Human-Object Interaction (HOI) Detection, which aims to locate and recognize HOI instances in the form of in images. Most existing works treat HOIs as individual interaction categories, thus can not handle the problem of long-tail distribution and polysemy of action labels. We argue that multi-level consistencies among objects, actions and interactions are strong cues for generating semantic representations of rare or previously unseen HOIs. Leveraging the compositional and relational peculiarities of HOI labels, we propose ConsNet, a knowledge-aware framework that explicitly encodes the relations among objects, actions and interactions into an undirected graph called consistency graph, and exploits Graph Attention Networks (GATs) to propagate knowledge among HOI categories as well as their constituents. Our model takes visual features of candidate human-object pairs and word embeddings of HOI labels as inputs, maps them into visual-semantic joint embedding space and obtains detection results by measuring their similarities. We extensively evaluate our model on the challenging V-COCO and HICO-DET datasets, and results validate that our approach outperforms state-of-the-arts under both fully-supervised and zero-shot settings. Ye Liu 0002, Junsong Yuan 0001, Chang Wen Chen |
ACM Multimedia | 3 |
| 2020 | Fusing motion patterns and key visual information for semantic event recognition in basketball videos
Lifang Wu, Qi Wang 0076, Meng Jian, Boxuan Zhao, Junchi Yan, Chang Wen Chen |
Neurocomputing | 7 |
| 2020 | Internet of Video Things: Next-Generation IoT With Visual SensorsabstractThe worldwide flourishing of the Internet of Things (IoT) in the past decade has enabled numerous new applications through the internetworking of a wide variety of devices and sensors. More recently, visual sensors have seen their considerable booming in IoT systems because they are capable of providing richer and more versatile information. Internetworking of large-scale visual sensors has been named Internet of Video Things (IoVT). IoVT has its own unique characteristics in terms of sensing, transmission, storage, and analysis, which are fundamentally different from the conventional IoT. These new characteristics of IoVT are expected to impose significant challenges to existing technical infrastructures. In this article, an overview of recent advances in various fronts of IoVT will be introduced and a broad range of technological and system challenges will be addressed. Several emerging IoVT applications will be discussed briefly to illustrate the potentials of IoVT in a broad range of practical scenarios. Chang Wen Chen |
IEEE Internet Things J. | 1 |
| 2020 | Integrating Mobile Display Energy Saving into Cloud-Based Video Streaming via Rate-Distortion-Display Energy ProfilingabstractMobile displays have been recognized as the major contributors to the energy consumption of contemporary mobile video services. Existing display energy reduction (DER) algorithms focus on local video/image processing by utilizing the computation resources on the mobile devices. As such, a per-device DER strategy is highly inefficient from the systematic perspective since the same computation is repeated among millions of individual mobile devices. In this new era of cloud-based video streaming, a natural question to ask is can ubiquitous cloud resources be exploited to overcome these drawbacks. It is based on this motivation that we design a paradigm-shifting strategy to integrate the display energy saving engine into cloud-based video streaming in order to simultaneously benefit massive mobile devices with a one-time video processing in the cloud. By taking full advantage of the computational and storage resources in the cloud, a family of Rate-Distortion-Display Energy (R-D-DE) profiles can be created for a given video source and for different types of mobile devices. The preliminary experimental results with OLED displays prove that we can achieve the desired DER performance through jointly managing the display energy and user experience by implementing the proposed R-D-DE integrated video encoding engine in the cloud. Qian Liu 0001, Zhisheng Yan, Chang Wen Chen |
IEEE Trans. Cloud Comput. | 3 |
| 2020 | Compressed Pseudo-Analog Transmission System for Remote Sensing Images Over Bandwidth-Constrained Wireless ChannelsabstractRecently, pseudo-analog transmission based on SoftCast has been proposed to improve the received quality of video/image by eliminating the cliff effect in traditional digital transmission. In this paper, we propose a Compressed Pseudo-analog Transmission System (ComPaTS) for remote sensing images over bandwidth-constrained wireless channels. This novel scheme is developed based on the observation that the inherent dropping strategy in pseudo-analog transmission is impractical for remote sensing images in which the transmission bandwidth is generally insufficient. In ComPaTS, to guarantee the content diversity gain under pseudo-analog transmission, block-based Compressive Sensing (CS) is applied to the wavelet domain of each remote sensing image, where the sampling ratio is proportional to the importance of different blocks. The main work of ComPaTS is to leverage the sampling ratio in block-based CS and the resource allocation in pseudo-analog transmission in order to minimize system distortion. Two components of system distortion, i.e., source distortion and channel distortion are analyzed respectively. To characterize the coupling relationship between these two different types of distortion, a joint bandwidth-power distortion optimization problem is formulated. Furthermore, we also propose a two-stage allocation algorithm to solve the problem efficiently. The simulation results demonstrate that the proposed ComPaTS scheme significantly outperforms reference schemes in terms of peak signal-to-noise ratio under different bandwidth-constrained scenarios. Yongqiang Gui, Hancheng Lu, Xiaoda Jiang, Feng Wu 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Ontology-Based Global and Collective Motion Patterns for Event Classification in Basketball VideosabstractIn multi-person videos, especially team sport videos, a semantic event is usually represented as a confrontation between two teams of players, which can be represented as collective motion. In broadcast basketball videos, specific camera motions are used to present specific events. Therefore, a semantic event in broadcast basketball videos is closely related to both the global motion (camera motion) and the collective motion. A semantic event in basketball videos can be generally divided into three stages: pre-event, event occurrence (event-occ), and post-event. By analyzing the influence of different stages of video segments to semantic events discrimination, it is observed that the pre-event and event-occ segments are effective for classification, while the post-events are effective for event success/failure classification. In this paper, we propose an ontology-based global and collective motion pattern (On_GCMP) algorithm for the basketball event classification. First, a two-stage GCMP-based event classification scheme is proposed. The GCMP is extracted using the optical flow. The two-stage scheme progressively combines a five-class event classification algorithm on event-occs and a two-class event classification algorithm on pre-events. Both algorithms utilize the sequential convolutional neural networks (CNNs) and the long short-term memory (LSTM) networks to extract the spatial and temporal features of GCMP for event classification. Second, we utilize the post-event segments to predict success/failure using deep features of images in the video frames (RGB_DF_VF)-based algorithms. Finally, the event classification results and success/failure classification results are integrated to obtain the final results. To evaluate the proposed scheme, we collected a new dataset called NCAA+, which is automatically obtained from the NCAA dataset by extending the fixed length of video clips forward and backward of the corresponding semantic events. The experimental results demonstrate that the proposed scheme achieves the mean average precision of 58.10% on NCAA+. It is higher by 6.50% than the state of the art on NCAA. Lifang Wu, Jiaoyu He, Meng Jian, Yaowen Xu, Dezhong Xu, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2020 | Deterministic Model Fitting by Local-Neighbor Preservation and Global-Residual OptimizationabstractGeometric model fitting has been widely used in many computer vision tasks. However, it remains as a challenging task when handing multiple-structural data contaminated by noises and outliers. Most previous work on model fitting cannot guarantee the consistency of their solutions due to their randomness, precluding them from many real-world applications. In this research, we propose a fast two-view approximately deterministic model fitting scheme (called LGF), to provide consistent solutions for multiple-structural data. The proposed LGF scheme starts from defining preference function by preserving local neighborhood relationship, and then adopts the min-hash technique to roughly sample subsets. By this way, it is able to cover all model instances in data in the parameter space with a high probability. After that, LGF refines the previous sampled subsets by globalresidual optimization. Furthermore, we propose a simple yet effective model selection framework to estimate the number and the parameters of model instances in data. Extensive experiments on real images show that the proposed LGF scheme is able to observe superior or very competitive performance on both accuracy and speed over several state-of-the-art model fitting methods. Guobao Xiao, Jiayi Ma 0001, Shiping Wang, Chang Wen Chen |
IEEE Trans. Image Process. | 4 |
| 2019 | AVT: Unsupervised Learning of Transformation Equivariant Representations by Autoencoding Variational TransformationsabstractThe learning of Transformation-Equivariant Representations (TERs), which is introduced by Hinton et al. \cite{hinton2011transforming}, has been considered as a principle to reveal visual structures under various transformations. It contains the celebrated Convolutional Neural Networks (CNNs) as a special case that only equivary to the translations. In contrast, we seek to train TERs for a generic class of transformations and train them in an {\em unsupervised} fashion. To this end, we present a novel principled method by Autoencoding Variational Transformations (AVT), compared with the conventional approach to autoencoding data. Formally, given transformed images, the AVT seeks to train the networks by maximizing the mutual information between the transformations and representations. This ensures the resultant TERs of individual images contain the {\em intrinsic} information about their visual structures that would equivary {\em extricably} under various transformations in a generalized {\em nonlinear} case. Technically, we show that the resultant optimization problem can be efficiently solved by maximizing a variational lower-bound of the mutual information. This variational approach introduces a transformation decoder to approximate the intractable posterior of transformations, resulting in an autoencoding architecture with a pair of the representation encoder and the transformation decoder. Experiments demonstrate the proposed AVT model sets a new record for the performances on unsupervised tasks, greatly closing the performance gap to the supervised models. Guo-Jun Qi, Liheng Zhang, Chang Wen Chen, Qi Tian 0001 |
ICCV | 3 |
| 2019 | Stable Network MorphismabstractDeep neural networks perform better when they are deeper. Network morphism is one of the paradigms to construct deeper neural networks. It makes developing deeper neural networks building on existing ones possible by morphing a well-trained neural network into a new one with the network function completely preserved. The morphed network also has the potential to continue growing into a more powerful one as it has more parameters. Existing network morphism schemes include Net2Net and NetMorph. However, both of them suffer from significant initial performance drop when the morphed network is continually trained. Such unstability is very much undesired for a continual learning system. In this research, we first identify the reason for the unstability, which is due to the large amount of zeros padded into the parameters. Based on this observation, we propose an algorithm based on modified gradient descent to decompose the network morphism equation. As a result, the morphed parameters are all non-zeros and the continual training process become stable. Experimental results on benchmark datasets demonstrate the effectiveness of the proposed stable network morphism scheme. Changhu Wang, Chang Wen Chen |
IJCNN | 3 |
| 2019 | Receiver-driven Video Multicast over NOMA Systems in Heterogeneous EnvironmentsabstractNon-orthogonal multiple access (NOMA) has shown potential for scalable multicast of video data. However, one key drawback for NOMA-based video multicast is the limited number of layers allowed by the embedded successive interference cancellation algorithm, failing to meet satisfaction of heterogeneous receivers. We propose a novel receiver-driven superposed video multicast (Supcast) scheme by integrating Softcast, an analog-like transmission scheme, into the NOMA-based system to achieve high bandwidth efficiency as well as gradual decoding quality proportional to channel conditions at receivers. Although Softcast allows gradual performance by directly transmitting power-scaled transformation coefficients of frames, it suffers performance degradation due to discarding coefficients under insufficient bandwidth and its power allocation strategy cannot be directly applied in NOMA due to interference. In Supcast, coefficients are grouped into chunks, which are basic units for power allocation and superposition scheduling. By bisecting chunks into base-layer chunks and enhanced-layer chunks, the joint power allocation and chunk scheduling is formulated as a distortion minimization problem. A two-stage power allocation strategy and a near-optimal low-complexity algorithm for chunk scheduling based on the matching theory are proposed. Simulation results have shown the advantage of Supcast against Softcast as well as the reference scheme in NOMA under various practical scenarios. Xiaoda Jiang, Hancheng Lu, Chang Wen Chen, Feng Wu 0001 |
INFOCOM | 3 |
| 2019 | 3D Singing Head for Music VR: Learning External and Internal Articulatory Synchronicity from Lyric, Audio and NotesabstractWe propose a real-time 3D singing head system to enhance the talking head on model integrity, keyframe generation and song synchronicity. The individual head appearance meshes are first obtained by matching multi-view visible images with face prior for accuracy, and then used to reconstruct entire head model by integrating with generic internal articulatory meshes for efficiency. After embedding physiology, the keyframes of each phoneme-music note correspondence are substantially synthesized from real articulation data. The song synchronicity of articulators is learned using a deep neural network to train visual co-articulation model (VCM) on parallel audio-visual data. Finally, the keyframes of adjacent phoneme-music note correspondences are blended by VCM to produce song synchronized animation. Compared to state-of-the-art baselines, our system can not only clearly distinguish phonemes and notes, but also significantly reduce the dependence on training data. Jun Yu 0001, Chang Wen Chen, Zengfu Wang |
ACM Multimedia | 2 |
| 2019 | Front-End Smart Visual Sensing and Back-End Intelligent Analysis: A Unified Infrastructure for Economizing the Visual System of City BrainabstractThe visual data, which are acquired from the ubiquitous visual sensors deployed in metropolitans, are of great value and paramount significance to enhance the effectiveness and pursue the future development of smart cities. In this paper, the essential building blocks of the unified visual data management and analysis infrastructure that serve as the foundation for the economical visual system in the city brain, are introduced to facilitate the utilization of the visual signal in the artificial intelligence era. In particular, we start by the discussion of the front-end smart visual sensing in the context of economical communication and service with the heterogeneous network, and the functionalities and necessities of compact visual feature and deep learning model representations are detailed. Subsequently, the utilities of the infrastructure are demonstrated through two intelligent applications at the back-end, including vehicle re-identification and person re-identification. The standardizations regarding compact feature and deep neural network representations, which are regarded as the key ingredients in this infrastructure and greatly facilitate the construction of the visual system in the city brain, are also discussed. Finally, we envision how the potential issues regarding the economical visual communications for future smart cities might be pragmatically approached within this unified infrastructure. Yihang Lou, Ling-Yu Duan, Shiqi Wang 0001, Ziqian Chen, Chang Wen Chen, Wen Gao 0001 |
IEEE J. Sel. Areas Commun. | 6 |
| 2019 | Replay attack detection based on distortion by loudspeaker for voice authentication
Yanzhen Ren, Zhong Fang, Dengkai Liu, Chang Wen Chen |
Multim. Tools Appl. | 4 |
| 2019 | Synthesizing 3D Trump: Predicting and Visualizing the Relationship Between Text, Speech, and Articulatory MovementsabstractThe movements of articulators, such as lips, tongue and teeth, play an important role in increasing the language expression capability by unmasking the information hid in text or speech. Hence, it is necessary to deeply mine and visualize the relationship between text, speech and articulatory movements for understanding language in multi-modality and multi-level. As a case study, given text and audio of President Donald John Trump, this paper synthesizes a high quality 3D animation of him speaking with accurate synchronicity between speech and articulators. First, visual co-articulation is modeled by predicting the mapping from text/speech to articulatory movements. Then, based on a reconstructed 3D head model, physiological characteristics and statistical learning are combined to visualize each phoneme. Finally, the visualization results of consecutive phonemes are fused by visual co-articulation model to generate synchronized articulatory animations. Experiments show that the system can not only produce photo-realistic results in front but also distinguish the visual differences among phonemes from unconstrained views. Jun Yu 0001, Qiang Ling 0001, Changwei Luo, Chang Wen Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Toward Guaranteed Video Experience: Service-Aware Downlink Resource Allocation in Mobile Edge NetworksabstractVideo delivery has been playing an essential role in video services over edge networks. Although HTTP segment-based streaming, e.g., Dynamic Adaptive Streaming over HTTP (DASH), has become the prevailing technique, it cannot provide guaranteed video playback in terms of bitrate to mobile users. In essence, HTTP streaming downloads the video segments in a best effort fashion, i.e., passively responding to the channel dynamics. This can cause unstable playback with frequent rebuffer and multi-client competition that degrades a network-wide performance. In this paper, we present a network-assisted streaming framework for Guaranteed Playback-Experience Streaming over HTTP (GESH) that leverages the proactive control of network resources and joint coordination among multiple clients for service-aware network resource allocation. Specifically, GESH is empowered by a new weighted proportional fair scheduling without modifying existing cellular infrastructure, a per-segment channel variation model, and a suite of algorithms to seek the optimal weights for the scheduling. Extensive evaluations show that GESH can maximally guarantee the video playback of multiple users, as well as significantly outperforming conventional HTTP streaming and current DASH systems. Zhisheng Yan, Miao Zhao, Cédric Westphal, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Noise Robust Multiobjective Evolutionary Clustering Image Segmentation Motivated by the Intuitionistic Fuzzy InformationabstractImages are always contaminated by noise, increasing uncertainty. Fuzzy set (FS) theory is a useful tool for dealing with uncertainty in images. When comparing with the FS, an intuitionistic fuzzy set (IFS) can better describe the blurred characteristic in images due to the membership, nonmembership, and hesitation degrees. However, when applied to an image segmentation, the IFS cannot completely overcome the influence of noise. With the aim of performing noisy image segmentation under several criteria, this paper defines a noise robust IFS (NR-IFS) for an image and then presents a novel noise robust multiobjective evolutionary intuitionistic fuzzy clustering algorithm (NR-MOEIFC). A majority dominated suppressed similarity measure using the neighborhood statistics and the competitive learning is proposed to obtain the NR-IFS representation for the image corrupted by noise. Then, the NR-IFS is fully used to motivate the whole process of multiobjective evolutionary clustering: first, computing a three-parameter intuitionistic fuzzy distance measure; second, constructing intuitionistic fuzzy fitness functions; third, designing a nonuniform intuitionistic fuzzy mutation operator; and forth, defining an intuitionistic fuzzy cluster validity index to select the optimal solution from the final nondominated solution set. The histogram statistics of NR-IFS are adopted in the NR-MOEIFC to greatly reduce the computational complexity. Experimental results on Berkeley and real magnetic resonance images reveal that the NR-MOEIFC behaves well in noise robustness and segmentation performance while requiring a low time cost. Feng Zhao 0005, JiuLun Fan 0001, Hanqiang Liu 0001, Rong Lan, Chang Wen Chen |
IEEE Trans. Fuzzy Syst. | 5 |
| 2019 | Real-Time Head Pose Estimation and Face Modeling From a Depth ImageabstractWe address the issues of 3-D head pose estimation and face modeling from a depth image. Given a depth image, random forests are effective for estimating the location and orientation of a person's head. However, the accuracy of the estimation is not high enough. We propose using corrected regression votes. The corrected votes are obtained by considering the cooperation of all trees, leading to significant improvement of head pose estimation accuracy. Based on the head pose estimator, we present a face modeling system. In our system, the face model is generated by aligning a deformable face model to a depth image using an iterative closest point (ICP) algorithm. The novelty of our approach is that an optimal weight for each vertex is incorporated into the ICP algorithm with point to plane constraints. Experiments show that our system can automatically estimate the head pose and generate a realistic face model from a single depth image. We also provide a detailed evaluation that shows the benefits of our approach. Changwei Luo, Juyong Zhang, Jun Yu 0001, Chang Wen Chen, Shengjin Wang |
IEEE Trans. Multim. | 4 |
| 2019 | Scalable Access Control For Privacy-Aware Media SharingabstractThe prevalence of social networks has made it easier than ever for users to share their photos, videos, and other media content with anybody from anywhere. However, the easy access of user-generated media content also brings about privacy concerns. Traditional access control mechanisms, where a single access policy is made for a specific piece of content, cannot satisfy the user privacy requirements in large-scale media sharing systems. Instead, configuring multiple levels of access privileges for the shared media content is desired. On one hand, it conforms to the principle of social networks in information propagation. On the other hand, it accords with the diverse and complex social relationship among social network users. In this paper, we propose a scalable media access control (SMAC) system to enable such a configuration in a secure and efficient manner. The proposed SMAC system is empowered by the scalable ciphertext policy attribute-based encryption algorithm as well as a comprehensive key-management scheme. We provide formal security proof to prove the security of the proposed SMAC system. In addition, we conduct extensive experiments on mobile devices to demonstrate its efficiency. Changsha Ma, Zhisheng Yan, Chang Wen Chen |
IEEE Trans. Multim. | 3 |
| 2019 | SSPA-LBS: Scalable and Social-Friendly Privacy-Aware Location-Based ServicesabstractPrivacy-aware location-based service (PA-LBS) preserves LBS users' privacy but undesirably sacrifices service quality. In order to balance the two factors with satisfactory user experience, existing frameworks are faced with two barriers, that is, scalability and social-friendliness. First, existing schemes do not enable LBS users to flexibly scale their privacy level on service provision. Such a lack of scalability easily results in either unacceptable service-quality degradation or insufficient privacy protection and fails to meet dynamic user requirements. Second, existing schemes handle privacy protection by merely considering the trust relationship between users and servers but ignore the complex trust relationships among users. As a result, users cannot preserve privacy in location-based social services that involve user-to-user interactions. In this paper, we present the first scalable and social-friendly PA-LBS system. In particular, we propose a novel camouflage algorithm with a formal privacy guarantee that enables LBS users to expose their location information by scaling two privacy related factors, that is, camouflage range and place type. Furthermore, we apply the scalable ciphertext policy attribute-based encryption algorithm to enable LBS users to effectively control the access from other users to their location information. Moreover, we also demonstrated the operational efficiency of the proposed system through successful implementations on Android devices. Changsha Ma, Zhisheng Yan, Chang Wen Chen |
IEEE Trans. Multim. | 3 |
| 2018 | DA-GAN: Instance-Level Image Translation by Deep Attention Generative Adversarial NetworksabstractUnsupervised image translation, which aims in translating two independent sets of images, is challenging in discovering the correct correspondences without paired data. Existing works build upon Generative Adversarial Networks (GANs) such that the distribution of the translated images are indistinguishable from the distribution of the target set. However, such set-level constraints cannot learn the instance-level correspondences (e.g. aligned semantic parts in object transfiguration task). This limitation often results in false positives (e.g. geometric or semantic artifacts), and further leads to mode collapse problem. To address the above issues, we propose a novel framework for instance-level image translation by Deep Attention GAN (DA-GAN). Such a design enables DA-GAN to decompose the task of translating samples from two sets into translating instances in a highly-structured latent space. Specifically, we jointly learn a deep attention encoder, and the instance-level correspondences could be consequently discovered through attending on the learned instances. Therefore, the constraints could be exploited on both set-level and instance-level. Comparisons against several state-of-the-arts demonstrate the superiority of our approach, and the broad application capability, e.g, pose morphing, data augmentation, etc., pushes the margin of domain translation problem.1 Jianlong Fu, Chang Wen Chen, Tao Mei 0001 |
CVPR | 3 |
| 2018 | FF-CMnet: A CNN-Based Model for Fine-Grained Classification of Car Models Based on Feature FusionabstractWe present in this paper a novel scheme for fine-grained car model classification based on convolutional neural network and feature fusion. This scheme is called FF-CMNET (Feature Fusion based Car Model Classification Net) and is based on the principle that the car frontal images can be partitioned into upper and lower parts that exhibit distinct feature distributions but are still structurally correlated to allow feature fusion. The characteristics of FF-CMNET include: (1) the design of two separate branches, named UpNet and DownNet, for extracting the features of upper parts and lower parts of the car frontal images separately; (2) a two-step fusion of features at the output of UpNet and DownNet and then again in FusionNet; and (3) the adoption of small convolution kernels and global average pooling. Extensive experiments conducted on a benchmark dataset, CompCars, show favorable results which demonstrate that the proposed FF-CMNET is able to outperform the state-of-the-art models in the classification of large datasets. Qiang Jin, Chang Wen Chen |
ICME | 3 |
| 2018 | Enabling Quality-Driven Scalable Video Transmission over Multi-User NOMA SystemabstractRecently, non-orthogonal multiple access (NOMA) has been proposed to achieve higher spectral efficiency over conventional orthogonal multiple access. Although it has the potential to meet increasing demands of video services, it is still challenging to provide high performance video streaming. In this research, we investigate, for the first time, a multi-user NOMA system design for video transmission. Various NOMA systems have been proposed for data transmission in terms of throughput or reliability. However, the perceived quality, or the quality-of-experience of users, is more critical for video transmission. Based on this observation, we design a quality-driven scalable video transmission framework with cross-layer support for multiuser NOMA. To enable low complexity multi-user NOMA operations, a novel user grouping strategy is proposed. The key features in the proposed framework include the integration of the quality model for encoded video with the physical layer model for NOMA transmission, and the formulation of multiuser NOMA-based video transmission as a quality-driven power allocation problem. As the problem is non-concave, a global optimal algorithm based on the hidden monotonic property and a suboptimal algorithm with polynomial time complexity are developed. Simulation results show that the proposed multi-user NOMA system outperforms existing schemes in various video delivery scenarios. Xiaoda Jiang, Hancheng Lu, Chang Wen Chen |
INFOCOM | 3 |
| 2018 | Fully Point-wise Convolutional Neural Network for Modeling Statistical Regularities in Natural ImagesabstractModeling statistical regularity plays an essential role in ill-posed image processing problems. Recently, deep learning based methods have been presented to implicitly learn statistical representation of pixel distributions in natural images and leverage it as a constraint to facilitate subsequent tasks, such as color constancy and image dehazing. However, the existing CNN architecture is prone to variability and diversity of pixel intensity within and between local regions, which may result in inaccurate statistical representation. To address this problem, this paper presents a novel fully point-wise CNN architecture for modeling statistical regularities in natural images. Specifically, we propose to randomly shuffle the pixels in the origin images and leverage the shuffled image as input to make CNN more concerned with the statistical properties. Moreover, since the pixels in the shuffled image are independent identically distributed, we can replace all the large convolution kernels in CNN with point-wise (1*1) convolution kernels while maintaining the representation ability. Experimental results on two applications: color constancy and image dehazing, demonstrate the superiority of our proposed network over the existing architectures, i.e., using 1/10~1/100 network parameters and computational cost while achieving comparable performance. Jing Zhang 0037, Yang Cao 0010, Yang Wang 0015, Chenglin Wen, Chang Wen Chen |
ACM Multimedia | 5 |
| 2018 | Generating hybrid interior structure for 3D printing
Yuxin Mao, Lifang Wu, Dong-Ming Yan 0001, Jianwei Guo 0003, Chang Wen Chen, Baoquan Chen |
Comput. Aided Geom. Des. | 5 |
| 2018 | A fast hybrid retargeting scheme with seam context and content aware strip partition
Lifang Wu, Chuncan Yan, Meng Jian, Weiming Dong, Chang Wen Chen |
Neurocomputing | 6 |
| 2018 | Intuitionistic fuzzy set approach to multi-objective evolutionary clustering with multiple spatial information for image segmentation
Feng Zhao 0005, Hanqiang Liu 0001, JiuLun Fan 0001, Chang Wen Chen, Rong Lan |
Neurocomputing | 4 |
| 2018 | λ-Domain Optimal Bit Allocation Algorithm for High Efficiency Video CodingabstractRate control typically involves two steps: bit allocation and bitrate control. The bit allocation step can be implemented in various fashions depending on how many levels of allocation are desired and whether or not an optimal rate- distortion (R-D) performance is pursued. The bitrate control step has a simple aim in achieving the target bitrate as precisely as possible. In our recent research, we have developed a λ-domain rate control algorithm that is capable of controlling the bitrate precisely for High Efficiency Video Coding (HEVC). The initial research showed that the bitrate control in the λ-domain can be more precise than the conventional schemes. However, the simple bit allocation scheme adopted in this initial research is unable to achieve an optimal R-D performance reflecting the inherent R-D characteristics governed by the video content. In order to achieve an optimal R-D performance, the bit allocation algorithms need to be developed taking into account the video content of a given sequence. The key issue in deriving the video-content-guided optimal bit allocation algorithm is to build a suitable R-D model to characterize the R-D behavior of the video content. In this paper, to complement the R-λ model developed in our initial work, a D-λ model is properly constructed to complete a comprehensive framework of λ-domain R-D analysis. Based on this comprehensive λ-domain R-D analysis framework, a suite of optimal bit allocation algorithms are developed. In particular, we design both picture-level and basic-unit-level bit allocation algorithms based on the fundamental R-D optimization theory to take full advantage of the content-guided principles. The proposed algorithms are implemented in HEVC reference software, and the experimental results demonstrate that they can achieve an obvious R-D performance improvement with a smaller bitrate control error. The proposed bit allocation algorithms have already been adopted by the Joint Collaborative Team on Video Coding and integrated into the HEVC reference software. Li Li 0040, Bin Li 0012, Houqiang Li, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Realizing Low-Cost Flash Memory Based Video Caching in Content Delivery SystemsabstractTo implement caching devices in content delivery systems, flash memory is preferable to hard disk drives from the performance perspective. Nevertheless, the higher bit cost of flash memory is one major obstacle for the wide real-life deployment of flash-based video caching. This paper presents a set of design solutions to address this cost issue. First, we present a flash memory error tolerance design strategy customized for video data storage, which can enable the use of lower-cost less-reliable flash memory chips for video storage. The cost challenge can also be addressed by reducing the video storage footprint through on-the-fly transcoding. However, direct transcoding suffers from a high implementation cost. We propose two design techniques that can largely reduce the transcoding complexity at minimal storage overhead in flash memory. All the developed design solutions share the common feature of cohesively exploring the characteristics of video coding and flash memory device physics. Their effectiveness has been well demonstrated through experiments with 20-nm MLC NAND flash memory chips and extensive simulations with representative video sequences. Danni Xiong, Kai Zhao 0005, Chang Wen Chen, Tong Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | IPAD: Intensity Potential for Adaptive De-QuantizationabstractDisplay devices at bit depth of 10 or higher have been mature but the mainstream media source is still at bit depth of eight. To accommodate the gap, the most economic solution is to render source at low bit depth for high bit-depth display, which is essentially the procedure of de-quantization. Traditional methods, such as zero-padding or bit replication, introduce annoying false contour artifacts. To better estimate the least-significant bits, later works use filtering or interpolation approaches, which exploit only limited neighbor information, cannot thoroughly remove the false contours. In this paper, we propose a novel intensity potential (IP) field to model the complicated relationships among pixels. The potential value decreases as the spatial distance to the field source increases and the potentials from different field sources are additive. Based on the proposed IP field, an adaptive de-quantization procedure is then proposed to convert low-bit-depth images to high-bit-depth ones. To the best of our knowledge, this is the first attempt to apply potential field for natural images. The proposed potential field preserves local consistency and models the complicated contexts well. Extensive experiments on natural, synthetic, and high-dynamic range image data sets validate the efficiency of the proposed IP field. Significant improvements have been achieved over the state-of-the-art methods on both the peak signal-to-noise ratio and the structural similarity. Jing Liu 0002, Guangtao Zhai, Anan Liu, Xiaokang Yang 0001, Xibin Zhao, Chang Wen Chen |
IEEE Trans. Image Process. | 6 |
| 2018 | CrowdDBS: A Crowdsourced Brightness Scaling Optimization for Display Energy Reduction in Mobile VideoabstractMobile display has become one of the most power-hungry components in mobile video viewing. Currently, mobile devices can reduce the display energy by performing dynamic brightness scaling (DBS) under the distortion constraint of video signals. We observe that there is a pitfall preventing current practice from systematic display energy reduction. In particular, existing objective DBS schemes lack direct connection to the subjective human perception on DBS-enabled videos, which is the key to achieving human-centered energy-experience optimization. To overcome this pitfall, we present CrowdDBS, a crowdsourced display energy reduction framework for mobile video viewing. CrowdDBS is empowered by a set of crowdsourcing studies that uncover the relationship between human perception and DBS frequency, magnitude, and temporal consistency, respectively. Motivated by the insights obtained from these studies, CrowdDBS employs a suit of designs and a DBS optimization framework to optimize the energy-experience tradeoff in mobile video viewing. Comprehensive experimental results and user evaluations under a variety of practical settings show that CrowdDBS can achieve 37 percent device energy reduction on average while guaranteeing satisfactory user experience in mobile video viewing. Zhisheng Yan, Qian Liu 0001, Tong Zhang 0002, Chang Wen Chen |
IEEE Trans. Mob. Comput. | 4 |
| 2018 | Blind Quality Assessment Based on Pseudo-Reference ImageabstractTraditional full-reference image quality assessment (IQA) metrics generally predict the quality of the distorted image by measuring its deviation from a perfect quality image called reference image. When the reference image is not fully available, the reduced-reference and no-reference IQA metrics may still be able to derive some characteristics of the perfect quality images, and then measure the distorted image's deviation from these characteristics. In this paper, contrary to the conventional IQA metrics, we utilize a new “reference” called pseudo-reference image (PRI) and a PRI-based blind IQA (BIQA) framework. Different from a traditional reference image, which is assumed to have a perfect quality, PRI is generated from the distorted image and is assumed to suffer from the severest distortion for a given application. Based on the PRI-based BIQA framework, we develop distortion-specific metrics to estimate blockiness, sharpness, and noisiness. The PRI-based metrics calculate the similarity between the distorted image's and the PRI's structures. An image suffering from severer distortion has a higher degree of similarity with the corresponding PRI. Through a two-stage quality regression after a distortion identification framework, we then integrate the PRI-based distortion-specific metrics into a general-purpose BIQA method named blind PRI-based (BPRI) metric. The BPRI metric is opinion-unaware (OU) and almost training-free except for the distortion identification process. Comparative studies on five large IQA databases show that the proposed BPRI model is comparable to the state-of-the-art opinion-aware- and OU-BIQA models. Furthermore, BPRI not only performs well on natural scene images, but also is applicable to screen content images. The MATLAB source code of BPRI and other PRI-based distortion-specific metrics will be publicly available. Xiongkuo Min, Ke Gu 0001, Guangtao Zhai, Jing Liu 0002, Xiaokang Yang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 6 |
| 2017 | Let Your Photos Talk: Generating Narrative Paragraph for Photo Stream via Bidirectional Attention Recurrent Neural NetworksabstractAutomatic generation of natural language description for individual images (a.k.a. image captioning) has attracted extensive research attention. In this paper, we take one step further to investigate the generation of a paragraph to describe a photo stream for the purpose of storytelling. This task is even more challenging than individual image description due to the difficulty in modeling the large visual variance in an ordered photo collection and in preserving the long-term language coherence among multiple sentences. To deal with these challenges, we formulate the task as a sequence-to-sequence learning problem and propose a novel joint learning model by leveraging the semantic coherence in a photo stream. Specifically, to reduce visual variance, we learn a semantic space by jointly embedding each photo with its corresponding contextual sentence, so that the semantically related photos and their correlations are discovered. Then, to preserve language coherence in the paragraph, we learn a novel Bidirectional Attention-based Recurrent Neural Network (BARNN) model, which can attend on the discovered semantic relation to produce a sentence sequence and maintain its consistence with the photo stream. We integrate the two-step learning components into one single optimization formulation and train the network in an end-to-end manner. Experiments on three widely-used datasets (NYC/Disney/SIND) show that the proposed approach outperforms state-of-the-art methods with large margins for both retrieval and paragraph generation tasks. We also show the subjective preference of the machine-generated stories by the proposed approach over the baselines through a user study with 40 human subjects. Yu Liu 0061, Jianlong Fu, Tao Mei 0001, Chang Wen Chen |
AAAI | 4 |
| 2017 | LARM: A Lifetime Aware Regression Model for Predicting YouTube Video PopularityabstractOnline content popularity prediction provides substantial value to a broad range of applications in the end-to-end social media systems, from network resource allocation to targeted advertising. While using historical popularity can predict the near-term popularity with a reasonable accuracy, the bursty nature of online content popularity evolution makes it difficult to capture the correlation between historical data and future data in the long term. Although various existing efforts have been made toward long-term prediction, they need to accumulate a long enough historical data before the prediction and their model assumptions cannot be applied to the complex YouTube networks with inherent unpredictability. Changsha Ma, Zhisheng Yan, Chang Wen Chen |
CIKM | 3 |
| 2017 | A-Lamp: Adaptive Layout-Aware Multi-patch Deep Convolutional Neural Network for Photo Aesthetic AssessmentabstractDeep convolutional neural networks (CNN) have recently been shown to generate promising results for aesthetics assessment. However, the performance of these deep CNN methods is often compromised by the constraint that the neural network only takes the fixed-size input. To accommodate this requirement, input images need to be transformed via cropping, warping, or padding, which often alter image composition, reduce image resolution, or cause image distortion. Thus the aesthetics of the original images is impaired because of potential loss of fine grained details and holistic image layout. However, such fine grained details and holistic image layout is critical for evaluating an images aesthetics. In this paper, we present an Adaptive Layout-Aware Multi-Patch Convolutional Neural Network (A-Lamp CNN) architecture for photo aesthetic assessment. This novel scheme is able to accept arbitrary sized images, and learn from both fined grained details and holistic image layout simultaneously. To enable training on these hybrid inputs, we extend the method by developing a dedicated double-subnet neural network structure, i.e. a Multi-Patch subnet and a Layout-Aware subnet. We further construct an aggregation layer to effectively combine the hybrid features from these two subnets. Extensive experiments on the large-scale aesthetics assessment benchmark (AVA) demonstrate significant performance improvement over the state-of-the-art in photo aesthetic assessment. Jing Liu 0002, Chang Wen Chen |
CVPR | 3 |
| 2017 | Fast Haze Removal for Nighttime Image Using Maximum Reflectance PriorabstractIn this paper, we address a haze removal problem from a single nighttime image, even in the presence of varicolored and non-uniform illumination. The core idea lies in a novel maximum reflectance prior. We first introduce the nighttime hazy imaging model, which includes a local ambient illumination item in both direct attenuation term and scattering term. Then, we propose a simple but effective image prior, maximum reflectance prior, to estimate the varying ambient illumination. The maximum reflectance prior is based on a key observation: for most daytime haze-free image patches, each color channel has very high intensity at some pixels. For the nighttime haze image, the local maximum intensities at each color channel are mainly contributed by the ambient illumination. Therefore, we can directly estimate the ambient illumination and transmission map, and consequently restore a high quality haze-free image. Experimental results on various nighttime hazy images demonstrate the effectiveness of the proposed approach. In particular, our approach has the advantage of computational efficiency, which is 10-100 times faster than state-of-the-art methods. Jing Zhang 0037, Yang Cao 0010, Shuai Fang, Yu Kang 0001, Chang Wen Chen |
CVPR | 5 |
| 2017 | Performance analysis for ZigBee under WiFi interference in smart homeabstractSmart home not only makes people's lives more convenient, but also saves energy and daily expenses for households. Both WiFi and ZigBee are widely deployed in the smart home and operated in the same 2.4GHz ISM band, which results in the coexistence interference. Since the transmitting power of WiFi devices is much higher than that of ZigBee devices, ZigBee is more susceptible to coexisting interference. In this paper, we propose an analytical model to evaluate the performance of ZigBee under WiFi interference in the practical smart home scenario. By considering both channel access behavior and path loss behavior of the WiFi and ZigBee devices, the proposed model is built by involving both the Markov chain model and the indoor path loss model, and the ZigBee performance under WiFi interference is theoretically derived. The simulation results demonstrate the effectiveness of the proposed model for evaluating the ZigBee performance under WiFi interference. Bin Liu 0016, Chang Wen Chen |
ICC | 3 |
| 2017 | On the effective capacities of distributed and co-located large-scale antenna systemsabstractEffective capacity analysis is a powerful tool to investigate the impact of physical layer designs on the link layer delay-sensitive QoS performance, which is important for real-time multimedia applications. In this paper, we rigorously analyze the effective capacities of downlink large-scale antenna systems. The main focus is to establish the fundamental effective capacity in a very-large MIMO system, and to characterize the performance difference between co-located and distributed antenna layouts. To that end, we first analytically derive the closed-form effective capacities for two widely used linear precoding schemes, conjugate beamforming and zero-forcing beamforming. We then analyze the asymptotic average effective capacities when the number of BS antennas and the number of users grow unboundedly with a fixed ratio. The effective capacity gain of the distributed antenna layout over the co-located layout is established via theoretical analysis. Cong Shen 0001, Chang Wen Chen, Feng Wu 0001 |
ICC | 3 |
| 2017 | IPAD: Intensity potential for adaptive de-quantizationabstractDisplay devices at bit-depth of 10 or higher have been mature but the mainstream media source is still at bit-depth as low as 8. To accommodate the gap, the most economic solution is to render source at low bit-depth for high bit-depth display, which is essentially the procedure of de-quantization. Traditional methods, like zero-padding or bit replication, introduce annoying false contour artifacts. To better estimate the least-significant bits, later works use filtering or interpolation approaches, which exploit only limited neighbor information, can not thoroughly remove the false contours. In this paper, we propose a novel intensity potential field to model the complicated relationships among pixels. Then, an adaptive de-quantization algorithm is proposed to convert low bit-depth images to high bit-depth ones. To the best of our knowledge, this is the first attempt to apply potential field for natural images. The proposed potential field preserves local consistency and models the complicated contexts very well. Extensive experiments on natural image datasets validate the efficiency of the proposed intensity potential field. Significant improvements have been achieved over the state-of-the-art methods on both PSNR and SSIM. Jing Liu 0002, Guangtao Zhai, Xiaokang Yang 0001, Menghan Hu, Chang Wen Chen |
ICME | 5 |
| 2017 | From talking head to singing head: A significant enhancement for more natural human computer interactionabstractThis paper proposes a 3D virtual animating head system, which can not only talk but also sing. With a reconstructed head mesh model, including external/internal articulators, from multi-source images, biology information are first used to visualize each phoneme with a musical note. The synchronicity between songs and articulatory movements is then modeled by a deep neural network trained on an audio/articulatory corpus. Finally, the visualization results of phonemes are blended by the synchronicity model to produce the song synchronized articulatory animations. Quantitative and qualitative improvements of singing ability on human computer interaction are demonstrated by comparing with other state-of-the-art talking head systems. Jun Yu 0001, Chang Wen Chen |
ICME | 2 |
| 2017 | Too Many Pixels to Perceive: Subpixel Shutoff for Display Energy Reduction on OLED SmartphonesabstractOrganic light-emitting diode (OLED) has been widely recognized as the next-generation mobile display. Recently, smartphone manufacturers have been pushing up the pixel density of OLED display. Unfortunately, such an effort does not necessarily improve the everyday viewing because of the limitation in human visual acuity. Instead, high pixel density OLED can drain the battery power even more quickly since the power dissipation of OLED is determined by the number of displayed pixels and their RGB values, or subpixels. This paper presents a new design dimension to remedy this prevailing issue by leveraging the intuition that shutting off redundant subpixels of the display content on OLED can reduce power consumption without impacting viewing perception. We introduce ShutPix, a power-saving display system for OLED smartphones that can optimally shut off the redundant subpixels before the content is displayed. Inspired by the motivational studies, ShutPix is empowered by a suite of designs based on visual acuity, human perception, and content redundancy. Experimental results show that ShutPix can, on average, reduce 21% of display power and 15% of system power without degrading user viewing experience. Zhisheng Yan, Chang Wen Chen |
ACM Multimedia | 2 |
| 2017 | A Structural Coupled-Layer Tracking Method Based on Correlation Filters
Bin Liu 0016, Chang Wen Chen |
MMM (1) | 3 |
| 2017 | Joint facial landmark detection and action estimation based on deep probabilistic random forestabstractRandom forest is effective and efficient for detecting facial landmark from visual images, and has achieved the state-of-the-art performance, both in accuracy and speed, by regressing local binary features (LBF). This paper aims to increase the detection accuracy of random forest for facial landmarks and extends it to facial action estimation. First, probabilistic features are designed to overcome the weaknesses of LBF, e.g., feature sparseness and tracking jitter. Second, a deep architecture is introduced to random forest for enhancing the capacity of representation learning. Third, the initial detected facial landmarks are refined and 3D facial actions are estimated jointly by registering a deformable facial model to images based on an optimized iterative closest point framework. Experiments show that the proposed methods significantly outperform the state-of-the-art ones in terms of accuracy, as well as achieve the excellent tracking stability and real-time ability at about 60 fps on an ordinary PC. Jun Yu 0001, Chang Wen Chen |
VCIP | 2 |
| 2017 | MLLDA: Multi-level LDA for modelling users on content curation social networks
Lifang Wu, Dan Wang 0004, Xiuzhen Zhang 0001, Chang Wen Chen |
Neurocomputing | 6 |
| 2017 | An Optimal Resource Allocation for Superposition Coding-Based Hybrid Digital-Analog SystemabstractHybrid digital-analog (HDA) video transmission is a new cross-layer design, which can be widely used in Internet of Things. The key problem of HDA video transmission is to find the optimal resource allocation between the digital and analog part. This paper presents a new general resource allocation algorithm for superposition coding-based HDA system. On one hand, in order to achieve successful decoding in digital part, the bitrate is controlled by the quantization parameter (QP), and the channel coding rate and modulation order are determined by the signal to interference noise power ratio, in which analog part is considered as the interference. On the other hand, the overall video quality is directly determined by the mean square error of analog part, which depends jointly on the data variance of the analog part, the power allocated to the analog part and the channel noise power. We propose a prediction model to describe how the data variance of the analog part changes with the QP in the digital part. Based on the proposed model, the power allocation of two parts can be quantitatively connected to form an optimization problem. We prove the convexity of the resource allocation problem and the gradient descent method is utilized in system implementation. With extensive simulations, the proposed algorithm is validated, achieving 1.4-5.3 dB gain over the conventional digital system, and 6.2-7.4 dB gain over pseudo-analog system in peak signal-to-noise ratio. Bin Tan 0001, Hao Cui 0001, Jun Wu 0006, Chang Wen Chen |
IEEE Internet Things J. | 4 |
| 2017 | Streaming Mobile Cloud Gaming Video Over TCP With Adaptive Source-FEC CodingabstractCloud gaming has emerged as a promising application to enable high-end game playing with thin clients. Transmission control protocol (TCP) is pervasively adopted as the transport-layer protocol in the mainstream cloud gaming systems for video communication. However, streaming mobile cloud gaming video using the TCP is challenged with several key technical barriers: (1) the performance limitations of wireless networks in bandwidth and reliability; (2) the high throughput demand and stringent delay constraint imposed by high-quality gaming video transmission; and (3) the deadline violations and throughput fluctuations caused by the packet retransmission and congestion control mechanisms in the TCP. To address these critical problems, this paper proposes an application-layer source-forward error correction (FEC) coding framework dubbed adaptive source-FEC coding over TCP (ESCOT). First, we analytically formulate the optimization problem of joint source-FEC coding to minimize the end-to-end distortion of real-time video communication over TCP. Second, we develop a heuristic solution for effective loss rate approximation, source rate control, and FEC coding adaptation. ESCOT is distinct from existing source-FEC coding schemes in proactively analyzing and leveraging the TCP characteristics. The proposed solution is able to effectively mitigate both consecutive and sporadic video frame drops caused by congestion and random packet losses. We conduct the performance evaluation through extensive emulations in the Exata platform using real-time gaming video encoded by the H.264 codec. Experimental results show that the ESCOT advances the state of the art with noticeable improvements in video peak signal-to-noise ratio, end-to-end delay, goodput, and frame success rate. Jiyan Wu, Chau Yuen, Ngai-Man Cheung, Junliang Chen 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | Prius: Hybrid Edge Cloud and Client Adaptation for HTTP Adaptive Streaming in Cellular NetworksabstractIn this paper, we present Prius, a hybrid edge cloud and client adaptation framework for HTTP adaptive streaming (HAS) by taking advantage of the new capabilities empowered by recent advances in edge cloud computing. In particular, emerging edge clouds are capable of accessing an application layer and radio access networks (RANs) information in real time. Coupled with powerful computation support, an edge cloud-assisted strategy is expected to significantly enrich mobile services. Meanwhile, although HAS has established itself as the dominant technology for video streaming, one key challenge for adapting HAS to mobile cellular networks is in overcoming the inaccurate bandwidth estimation and unfair bitrate adaptation under the highly dynamic cellular links. Edge cloud-assisted HAS presents a new opportunity to resolve these issues and achieve systematic enhancement of quality of experience (QoE) and QoE fairness in cellular networks. To explore this new opportunity, Prius overlays a layer of adaptation intelligence at the edge cloud to finalize the adaptation decisions while considering the initial bandwidth-irrelevant bitrate selection at the clients. Prius is able to exploit RAN channel status, client device characteristics, and application-layer information in order to jointly adapt the bitrate of multiple clients. Prius also adopts a QoE continuum model to track the cumulative viewing experience and an exponential smoothing estimation to accurately estimate a future channel under different moving patterns. Extensive trace-driven simulation results show that Prius with hybrid edge cloud and client adaptation is promising under both slow and fast-moving environments. Furthermore, the Prius adaptation algorithm achieves a near-optimal performance that outperforms the exiting strategies. Zhisheng Yan, Jingteng Xue, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | No-Reference Quality Metric of Contrast-Distorted Images Based on Information MaximizationabstractThe general purpose of seeing a picture is to attain information as much as possible. With it, we in this paper devise a new no-reference/blind metric for image quality assessment (IQA) of contrast distortion. For local details, we first roughly remove predicted regions in an image since unpredicted remains are of much information. We then compute entropy of particular unpredicted areas of maximum information via visual saliency. From global perspective, we compare the image histogram with the uniformly distributed histogram of maximum information via the symmetric Kullback-Leibler divergence. The proposed blind IQA method generates an overall quality estimation of a contrast-distorted image by properly combining local and global considerations. Thorough experiments on five databases/subsets demonstrate the superiority of our training-free blind technique over state-of-the-art full- and no-reference IQA methods. Furthermore, the proposed model is also applied to amend the performance of general-purpose blind quality metrics to a sizable margin. Ke Gu 0001, Weisi Lin, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Chang Wen Chen |
IEEE Trans. Cybern. | 6 |
| 2017 | Full Reference Quality Assessment for Image Retargeting Based on Natural Scene Statistics Modeling and Bi-Directional Saliency SimilarityabstractImage retargeting technology has been widely studied to adapt images for the devices with heterogeneous screen resolutions. Meanwhile effective objective retargeting quality assessment algorithms are also very important for optimizing and selecting favorable retargeting methods. Unlike previous assessment algorithms which rely on image local structure features and unidirectional prediction of information loss, we propose a bi-directional natural salient scene distortion model (BNSSD) including image natural scene statistics (NSS) measurement, salient global structure distortion measurement, and bi-directional salient information loss measurement. First, we propose a new NSS model in log-Gabor domain and verify its effectiveness in reflecting nature scene statistical distortions introduced during the retargeting process. Second, the concept of salient global structure distortion is proposed to measure the global structure uniformity in the corresponding salient regions between original and retargeted images. Finally, we propose a bidirectional salient information loss metric to measure the information loss between salient areas in original image and retargeted image. The effectiveness of the BNSSD model is verified on two widely recognized public databases, and the experimental results show that our method outperforms the state-of-the-art algorithms under different statistical assessment criteria. Zhibo Chen 0001, Ning Liao, Chang Wen Chen |
IEEE Trans. Image Process. | 4 |
| 2017 | Local Co-Occurrence Selection via Partial Least Squares for Pedestrian DetectionabstractChannel feature detectors are the most popular approaches for pedestrian detection recently. However, most of these approaches train the boosted decision trees by selecting a single feature at each node, which does not effectively exploit the multi-feature cues and spatial information. To address this issue, this paper proposes to construct the co-occurrence of multiple channel features in local image neighborhoods for pedestrian detection. In our approach, a binary pattern of feature co-occurrence is represented by combining the binary variables quantized from each channel feature, and the spatial information is incorporated by selecting the neighbors to jointly represent the feature co-occurrence in a local image block. However, feature co-occurrence selection leads to many possible feature combinations, which significantly increase the computational cost at the training stage. Therefore, in order to reduce the number of candidate features and obtain the most discriminative features effectively, a partial least squares-based feature selection approach called variable importance on projection is exploited. Comprehensive experiments are conducted on several challenging pedestrian data sets, and superior performances are achieved by the proposed approach in comparison with some state-of-the-art pedestrian detection approaches. Hanzi Wang, Yan Yan 0001, Bo Li 0006, Chang Wen Chen |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2017 | Medium Access Control for Wireless Body Area Networks with QoS Provisioning and Energy Efficient DesignabstractWith the promising applications in e-Health and entertainment services, wireless body area network (WBAN) has attracted significant interest. One critical challenge for WBAN is to track and maintain the quality of service (QoS), e.g., delivery probability and latency, under the dynamic environment dictated by human mobility. Another important issue is to ensure the energy efficiency within such a resource-constrained network. In this paper, a new medium access control (MAC) protocol is proposed to tackle these two important challenges. We adopt a TDMA-based protocol and dynamically adjust the transmission order and transmission duration of the nodes based on channel status and application context of WBAN. The slot allocation is optimized by minimizing energy consumption of the nodes, subject to the delivery probability and throughput constraints. Moreover, we design a new synchronization scheme to reduce the synchronization overhead. Through developing an analytical model, we analyze how the protocol can adapt to different latency requirements in the healthcare monitoring service. Simulations results show that the proposed protocol outperforms CA-MAC and IEEE 802.15.6 MAC in terms of QoS and energy efficiency under extensive conditions. It also demonstrates more effective performance in highly heterogeneous WBAN. Bin Liu 0016, Zhisheng Yan, Chang Wen Chen |
IEEE Trans. Mob. Comput. | 3 |
| 2017 | Message From the Outgoing Editor-in-ChiefabstractPresents the introductory editorial for this issue of the publication. Chang Wen Chen |
IEEE Trans. Multim. | 1 |
| 2017 | Analog Coded SoftCast: A Network Slice Design for Multimedia Broadcast/MulticastabstractThis paper presents a network slice design for ultra high definition (UHD) video broadcast/multicast to achieve higher network efficiency and improved quality of experience (QoE). The proposed network slice design consists of a rateless source compression scheme and an analog-coded SoftCast scheme. The rateless Spinal code is adopted to compress the video source at content server and the compressed source is transmitted from content server across wireless core network to the base station. Ana prioriinformation-assisted Spinal decoder is designed to utilize the sparsity of bit planes for compression. In the analog-coded SoftCast scheme, we design a new chaotic function-based analog code with negligible power penalty for the generalized Gaussian-distributed source in SoftCast because the existing chaotic functions designed for uniformly distributed sources suffer from serious power penalty in SoftCast. We also design a maximuma posterioriprobability decoding algorithm for the proposed analog code in order to exploit the statistics of video source asa prioriinformation to improve the performance. The experimental results show that the proposed rateless code-based compression scheme achieves efficient compression and approaches the bound of binary erasure channel. In particular, the 1/2 analog-coded SoftCast has almost 2 dB gain over conventional SoftCast with two repetitions, and the 1/3 analog-coded SoftCast has almost 3 dB gain over conventional SoftCast with three repetitions. The system simulations for the broadcast system show higher network capacity and improved QoE in the proposed UHD slice, because the reconstructed video quality of each user is commensurate with its channel condition. Bin Tan 0001, Jun Wu 0006, Ying Li 0020, Hao Cui 0001, Chang Wen Chen |
IEEE Trans. Multim. | 6 |
| 2016 | An energy-efficient and QoS-effective resource allocation scheme in WBANsabstractWireless Body Area Networks (WBANs) represent one of the most promising networks to provide health applications for improving the quality of life, such as ubiquitous e-Health services and real-time health monitoring. The resource allocation of an energy-constrained, heterogeneous WBAN is a critical issue that should consider both energy efficiency and Quality of Service (QoS) requirements with the dynamic link characteristics, especially when the limited resource cannot satisfy the expected QoS requirements. In this paper, we propose an Energy-efficient and QoS-effective resource allocation that considers a mix-cost parameter characterizing both energy cost and QoS cost between attainable QoS support and QoS requirements. Based on the mix-cost parameter, we first formulate the resource allocation problem as a mixed integer nonlinear programming (MINP) for optimizing the transmission power, the transmission rate and allocated time slots for each sensor to minimize total mix-cost of the system. Then we propose a sub-optimal greedy resource allocation algorithm, which has a much lower complexity compared to exhaustive search. Simulation results demonstrate the advantage of the mix-cost parameter to evaluate energy efficiency and attainable QoS support, as well as verifying the effectiveness of the proposed resource allocation algorithm. Bin Liu 0016, Chang Wen Chen |
BSN | 4 |
| 2016 | Buffer-aware and QoS-effective resource allocation scheme in WBANsabstractWireless Body Area Network (WBAN) represents one of the most promising networks to provide health applications for improving the quality of life, such as ubiquitous e-Health services and real-time health monitoring. The resource allocation of an energy-constrained, heterogeneous WBAN is a critical issue that should consider both energy efficiency and Quality of Service (QoS) requirements with the dynamic link characteristics, especially when the limited resource cannot satisfy the expected QoS requirements. In this paper, a buffer aware Energy-efficient and QoS-effective resource allocation scheme is proposed in which the sensor queue buffer states, constraints of QoS metrics and the characteristics of dynamic links are considered. Specifically, a buffer aware sensor evaluation method is designed to dynamically evaluate the sensor state with considering the sensor buffer states for improving the system performance. We then formulate the resource allocation problem for optimizing the transmission power, the transmission rate and the allocated time slots for each sensor to minimize the sum mix-cost, which is defined to characterize the energy cost and QoS cost between attainable QoS support and QoS requirements. Simulation results demonstrate the effectiveness of the buffer aware sensor evaluation method and the proposed energy-efficient and QoS-effective resource allocation scheme. Bin Liu 0016, Chang Wen Chen |
HealthCom | 3 |
| 2016 | A two-layer and multi-strategy framework for human activity recognition using smartphoneabstractHuman Activity Recognition (HAR) is widely used in many applications and HAR using smartphone only has been proved to be effective, flexible and unobtrusive for activity recognition. In this paper, a two-layer and multi-strategy HAR framework is proposed to overcome the major challenge of HAR using smartphone only, i.e., the variation in orientation and position of the device. In the first layer, the activities are classified into different groups with high accuracy and for each group in the second layer, the appropriate strategy is designed according to the characteristics of the group to improve the recognition performance. For static activity group, the transitional activities are introduced to help classifying the activities indirectly. For dynamic activity group sensitive to the position variation of the smartphone, a position-assisted strategy is proposed to alleviate the influence of position variation. The simulation results demonstrate the effectiveness of the proposed two-layer multi-strategy HAR framework. Bin Liu 0016, Chang Wen Chen |
ICC | 3 |
| 2016 | Bayesian based view synthesis for multi-planar structuresabstractWith the rapid development of mobile applications in recent years, there is a strong desire on light weight algorithm for view synthesis using uncalibrated images and limited geometry information. To address this challenge, we propose a Bayesian based view synthesis framework to support the rendering of complex scene with multiple planar structures. In this framework, we integrate image segmentation, reference plane selection and hole filling with a Bayesian formulation to perform view synthesis without using 3D geometry information. More specifically, we partition every reference image into multiple planes, estimate geometric and photometric parameters for each plane, synthesize the novel view by Bayesian modeling using selected reference planes, and refine the rendered image by a hole filling scheme. The entire view synthesis process is executed in an iterative manner to pursue high quality visual results. The experimental results show that the proposed method is able to achieve desired performance with less distortion and higher resolution. Jie Hu 0008, Chang Wen Chen |
ICIP | 2 |
| 2016 | Cloud-based video streaming with systematic mobile display energy saving: Rate-distortion-display energy profilingabstractMobile display has been considered as the major contributor to the energy consumption of the ever-increasing mobile video services. Current practices in display energy reduction (DER) utilize local computing resources to analyze the video content before DER strategies can be applied in a per-device fashion. For a given video, same analytical computations are repeated in millions of individual devices. In this paper, we demonstrate that a paradigm shifting framework can be designed to systematically move the common DER local processing to the streaming server with the emergence of cloud-based video services. This framework has the potential to replace the massive per-device DER computation by a one-time global video processing in the cloud. To accomplish this ultimate goal in DER, a new family of video rate-distortion (R-D) profile with embedded DER strategies shall be properly generated. This family of rate-distortion-display energy (R-D-DE) profiles contains a set of common DER parameters to be directly extracted and employed by individual mobile devices to achieve desired display energy saving without repeated local computation. Performance evaluations of the proposed design are carried out to show that this family of R-D-DE profiles is indeed able to command the systematic DER design based on the intrinsic relationships among bitrate, video quality and display energy saving. Qian Liu 0001, Zhisheng Yan, Chang Wen Chen |
ICIP | 3 |
| 2016 | Automatic creation of magazine-page-like social media visual summary for mobile browsingabstractToday, mobile users are struggling with accessing overloading and unstructured social media feeds on the severely constrained mobile display. To overcome the challenges associated with browsing social media feeds on mobile devices, we are developing an innovative scheme to automatically create and synthesize the mixed social media digest (pictures, texts and videos) into a magazine-page-like social media visual summary. Given a set of personalized social media digest, a multi-objective optimization is formulated to organize the digest into visual summary in a 9-block-partition fashion with consideration of informative delivery, aesthetic rules and visual perception principles. Each block will be optimized interactively in terms of size, position and color to best represent the overall social media digest. Extensive evaluation and analysis based on user studies demonstrate that the proposed approach is effective in presenting social media content in a visually appealing and compact way. It is expected that this visual summary will lead to much enhanced user experiences for browsing social media digest on mobile devices. Chang Wen Chen |
ICIP | 2 |
| 2016 | Forecasting initial popularity of just-uploaded user-generated videosabstractUser-generated videos (UGVs) have dominated contemporary social networking sites (SNSs). Forecasting their popularity is of great relevance to a broad range of online services. All existing studies forecast popularity of UGVs using their popularity statistics that are accumulated for a period of time after they are uploaded. Hence, there is always a substantial time lag (days to weeks) before popularity forecast can take effects. However, such a time lag is undesirable for timely popularity forecast as forecasting initial popularity during UGVs' lifetime is vitally important. In fact, we have found in our measurement that the most popular UGVs usually precede others starting from the beginning days and UGVs generally receive the highest attentions during the first few days. In this paper, we present the first exploration on forecasting initial popularity for UGVs at their uploading moment without accumulating their popularity statistics. Specifically, we first design an effective crawler framework to collect the publicly observable features of videos at their uploading moment. We then collect a representative and large YouTube video data set with 318,627 videos. Based on the data set, we select the most relevant features as predictors and design a neural network-based learning model to forecast initial popularity of just-uploaded UGVs. Experimental results validate the effectiveness of the proposed forecasting model and demonstrate the model's benefits for online services such as in-video advertising and video caching. Changsha Ma, Zhisheng Yan, Chang Wen Chen |
ICIP | 3 |
| 2016 | Automatic suggestion of presentation image for storytellingabstractDigital storytelling applications are playing an increasingly important role in people's daily life. In contemporary storytelling applications such as PowerPoint presentation and macro/micro blogs, good presentation images are always highly desired by content creators to boost their presentation in an intuitive and attractive way. Existing studies, however, have not yet addressed the challenging problem of how to select the most appropriate presentation images for storytelling. In this paper, we formulate this problem of presentation image suggestion (given a textual query) as selecting images by maximizing visual and semantic diversity from web image search results of suggested queries. The proposed framework consists of two novel components: 1) click-through-based query suggestion, which is designed to suggest textual queries by searching relevant queries in a constructed query graph that can reflect diverse aspects of a given query, and 2) query-based image selection, which selects the most appropriate presentation images by keeping semantic relevance while maximizing visual diversity and quality, using a novel model based on Conditional Random Field (CRF) by individual and correlation characters. We evaluate the proposed approach by comparing with several baselines and a thorough subjective survey. The evaluations show inspiring results using the proposed approach for automatic suggestion of images for storytelling. Yu Liu 0061, Tao Mei 0001, Chang Wen Chen |
ICME | 3 |
| 2016 | Attribute-based multi-dimension scalable access control for social media sharingabstractSocial media sharing is one of the most popular social interactions in online social networks (OSNs). Due to the diverse networking conditions and various privacy requirements of OSN users, scalable media sharing has become a promising paradigm. It allows a media data distributor to share a media content of different qualities with different data consumers. To guarantee user privacy in scalable media sharing, it is essential to design an effective scalable media access control (SMAC) mechanism. However, all existing schemes cannot support scalable media streams with more than two dimensions, which significantly limits the flexibility of OSN services. In this paper, we present the first multi-dimension SMAC (MD-SMAC) system for social media sharing, based on the proposed scalable ciphertext policy attribute-based encryption (SCP-ABE) algorithm. In the MD-SMAC system, secure and reliable access control can be performed on multidimension scalable media streams according to data consumers' attributes. Through the security analysis, we prove the security and reliability of the MD-SMAC system. We also conduct experiments on mobile devices to demonstrate the computation efficiency of the proposed system. Changsha Ma, Zhisheng Yan, Chang Wen Chen |
ICME | 3 |
| 2016 | Network MorphismabstractWe present a systematic study on how to morph a well-trained neural network to a new one so that its network function can be completely preserved. We define this as network morphism in this research. After morphing a parent network, the child network is expected to inherit the knowledge from its parent network and also has the potential to continue growing into a more powerful one with much shortened training time. The first requirement for this network morphism is its ability to handle diverse morphing types of networks, including changes of depth, width, kernel size, and even subnet. To meet this requirement, we first introduce the network morphism equations, and then develop novel morphing algorithms for all these morphing types for both classic and convolutional neural networks. The second requirement is its ability to deal with non-linearity in a network. We propose a family of parametric-activation functions to facilitate the morphing of any continuous non-linear activation neurons. Experimental results on benchmark datasets and typical neural networks demonstrate the effectiveness of the proposed network morphism scheme. Changhu Wang, Yong Rui, Chang Wen Chen |
ICML | 4 |
| 2016 | Visual saliency model based on minimum description lengthabstractIn this paper, a novel patch-wise visual saliency model based on Minimum Length Description (MDL) principle is presented. Visual saliency is measured as the unpredicted information of image patch through an order-adaptive predictor under MDL principle. Specifically, each image patch is estimated with a linear combination of several neighboring patches. The number and location of candidate patches are automatically tuned to local contexts based on MDL. Then the entropy of prediction residuals of center patch, which represents the surprise to the visual system, is used to measure the saliency. Furthermore, a structural redundancy operator is also involved to improve the saliency detection performance. Experimental results demonstrate that the predictor under MDL principle along with the structural redundancy operator can improve the accuracy of human fixations prediction. We show that the proposed model outperforms the mainstream algorithms in predicting human fixations. Jing Liu 0002, Xiaokang Yang 0001, Guangtao Zhai, Chang Wen Chen |
ISCAS | 4 |
| 2016 | User Profiling by Combining Topic Modeling and Pointwise Mutual Information (TM-PMI)
Lifang Wu, Dan Wang 0004, Chang Wen Chen |
MMM (2) | 5 |
| 2016 | RnB: rate and brightness adaptation for rate-distortion-energy tradeoff in HTTP adaptive streaming over mobile devicesabstractVideo streaming is a prevalent mobile service that drains a significant amount of battery power. While various efforts have been made toward saving both video transfer and display energy, they are independently designed in an ad-hoc way and thereby can cause some non-apparent yet critical performance issues. To fill in this gap, this paper presents a fundamentally new design by jointly considering the end-to-end pipeline from the initial video encoding to the final mobile display. In essence, we shift the classic R-D tradeoff that has governed streaming system designs for decades to a fresh rate-distortion-energy (R-D-E) tradeoff specifically tailored for mobile devices. We present RnB, a video bitrate and display brightness adaptation platform that is standard-compliant, backward compatible, and device-neutral in order to achieve the proposed R-D-E tradeoff. RnB is empowered by some new discovery about the inherent relationship among bitrate, display brightness, and video quality as well as by an control-theoretic formulation to dynamically adapt the bitrate and scale the display brightness. Experimental results based on real-time implementation show that RnB can achieve an average of 19% energy reduction with final video quality comparable to conventional R-D based schemes. Zhisheng Yan, Chang Wen Chen |
MobiCom | 2 |
| 2016 | A joint visual-inertial image registration for mobile HDR imagingabstractRobust image alignment is a necessary and challenging step for numerous computational photography applications. In particular, large camera motion poses significant challenge to Mobile High Dynamic Range (HDR) Imaging due to hand-held capture of input images and limited computational resources. Aligning images only by detecting and matching image features is computationally expensive and can also be erratic. We present a robust multi-sensory method for aligning exposure bracketed images on mobile cameras. We use inertial sensor based camera pose estimate to pre-warp images and iteratively align them by minimizing alignment error. We also simultaneously estimate local motion masks which can be used to eliminate ghosting artifacts in the final HDR image. We collected HDR image dataset with diverse scenes along with inertial sensor data, which is a novel contribution and have evaluated our performance with existing mobile HDR image alignment techniques in literature. Radhakrishna Dasari, Chang Wen Chen |
VCIP | 2 |
| 2016 | Are you what you look like? Exploring correlations in personality type and their wearingabstractEveryday, people choose their preferred clothing to wear before leaving home for various activities. It has been recognized that each person may consciously or unconsciously express individual personality through their wearing styles. It will be interesting to ask: Are you what you look like? This paper presents a novel scheme developed to infer personality type from their wearing. The proposed research is justifiably rooted in the psychological findings that reveal intrinsic correlations between clothing style and wearer's inner self in self-image, mood, and social aspirations. First, we build a relatively large dataset with more than 300 persons and over 10,000 portraits, each is labeled with personality type. Then, personality-related clothing features are explored through statistical analysis based on psychological theories. To extract the clothing features from the images, a suite of algorithms, including body detection, GrabCut algorithm and saliency detection have been developed. Binary logistic regression is then applied to verify the Significance Level of the extracted features for predicting personality types. Experimental results demonstrate that the proposed scheme is able to predict several types of personality combination or type pairs with relatively high precisions. Yan Yan 0003, Zhiqiang Wei 0002, Chang Wen Chen |
VCIP | 3 |
| 2016 | Irregular Repetition Slotted ALOHA with Priority (P-IRSA)abstractIn this paper, a random access with priority protocol named irregular repetition slotted ALOHA with priority (PIRSA) is introduced. By optimizing the packets repetition and using successive interference cancellation (SIC), irregular repetition slotted ALOHA (IRSA) has achieved a high throughput. Contention resolution diversity slotted ALOHA with access control (AC-CDRSA) enhances both throughput and efficiency under high traffic load. In P-IRSA, priority is considered in order to handle traffics from the users who demands enhanced access and throughput. Access control is added before the traffic, and users with high priority gain higher probability to access the channel under high traffic load. Further more, performance of the proposed P- IRSA in terms of the packet loss rate, throughput and channel efficiency are verified by simulations and are compared with IRSA and AC-CDRSA protocols. Simulation results show that the throughput of users with high priority can be enhanced under high traffic load with the proposed protocol. Jingyun Sun, Rongke Liu, Chang Wen Chen |
VTC Spring | 4 |
| 2016 | TCP-Oriented Raptor Coding for High-Frame-Rate Video Transmission Over Wireless NetworksabstractHigh-frame-rate (HFR) video technology is becoming widely implemented in popular multimedia applications (e.g., Youtube and cloud gaming) to provide a smooth viewing experience, while transmission control protocol (TCP) is pervasively adopted as the transport-layer solution for video communications to achieve firewall traversal and network friendliness. However, it is severely challenging to effectively deliver real-time HFR video over TCP: 1) HFR video streaming features high transmission rate, enhanced frame density, and stringent delay constraint and 2) the packet retransmission and congestion control mechanisms in TCP may cause frequent throughput fluctuations and deadline violations. Motivated by addressing these critical issues, this research presents an application-layer forward error correction (FEC) framework dubbed Raptor coded HFR video over TCP (ROCHET). First, we develop a mathematical model to analyze the frame-level distortion of systematic Raptor code-based HFR video communication over TCP in wireless networks. Second, we propose a joint approximate distortion estimation and Raptor coding adaption solution to minimize the sum of total distortion. The proposed ROCHET is able to effectively leverage unequal error protection and TCP state analysis to enhance streaming video quality. We conduct the performance evaluation through extensive emulations in Exata involving real-time HFR video encoded with H.264 codec. Compared with the existing FEC coding schemes, ROCHET achieves appreciable improvements in terms of video peak signal-to-noise ratio, goodput, and frame success rate. Thus, ROCHET is recommended for TCP-based HFR video transmission over wireless networks. Jiyan Wu, Chau Yuen, Ming Wang 0002, Junliang Chen 0001, Chang Wen Chen |
IEEE J. Sel. Areas Commun. | 5 |
| 2016 | Scalable Video Multicast for MU-MIMO Systems With Antenna HeterogeneityabstractIn contemporary multiuser multiple-input multipleoutput systems, it is common for the reception devices to have a varying number of antennas. When multicast is performed, the number of concurrent spatial streams is limited by the device with the least number of antennas, which prevents more capable devices from getting higher rates. In this paper, we address the antenna heterogeneity in wireless video multicast by the innovative design of multiple similar description (MSD) video coding and multiplexed space-time block coding (M-STBC). MSD coding generates multiple descriptions of a video and features that any linear combinations of the descriptions are decodable. The descriptions comprising of real numbers are further processed by transform and power allocation steps for efficient transmission in a power-constrained system. M-STBC puts symbols in similar descriptions to corresponding space-time positions and ensures decodability under any antenna settings and channel conditions. As a result, we build up a scalable video multicast system, named AirScale, which allows receivers with a various number of antennas to decode from a single transmission, and the reconstructed video quality improves with the number of equipped antennas. Evaluations on Sora shows that, in a {1, 2, 3, 4} × 4 system, AirScale provides baseline quality for one-antenna receiver and a much higher quality for multiantenna receivers. The gain over SoftCast is up to 3.5, 3.9, and 4.1 dB for two-, three-, and four-antenna receivers, respectively. Hao Cui 0001, Chong Luo 0001, Chang Wen Chen, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | A Unified Scheme for Super-Resolution and Depth Estimation From Asymmetric Stereoscopic VideoabstractReconstructing a full-resolution stereoscopic video from an asymmetric stereoscopic video is a challenging task. The existing approaches require depth information, which imposes an additional challenge in data acquisition. In this paper, we propose a novel scheme that is capable of obtaining super-resolution and depth estimation simultaneously from an asymmetric stereoscopic video. The proposed scheme models the video super-resolution and stereo matching with a unified energy function. Then, we apply an alternating optimization method to minimize this energy function, which can be implemented with a two-step algorithm. In the first step we calculate the initial depth map by using a region-based cooperative optimization technique while considering the temporal consistency in video. In the second step we resolve the super-resolution problem under the guidance of the depth information. It is effective because each step benefits from the additional improvement over the previous step. We iteratively update the two steps until stable depth and super-resolution results are obtained. We have conducted a series of experiments on public stereoscopic video sequences to evaluate the performance of the proposed method. Both objective indexes and subjective visual comparisons verify that the proposed scheme can achieve satisfactory super-resolution results and high-quality depth map simultaneously. In particular, the subjective evaluation experiments on a 3-D monitor show that this scheme outperforms others and achieves the best visual sharpness. Jing Zhang 0037, Yang Cao 0010, Zhengjun Zha, Zhigang Zheng, Chang Wen Chen, Zengfu Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Adaptive Hybrid Digital-Analog Video Transmission in Wireless Fading ChannelabstractWe propose an adaptive hybrid digital-analog video transmission (A-HDAVT) scheme for robust video streaming in mobile networks with realistic fading channels. This scheme is fundamentally different from recent research in hybrid digital-analog video transmission in which wireless channels are unrealistically assumed to be Gaussian. For fading channels, it is critical to take full advantage of diversity in both video contents and multiuser channels. Like all hybrid approaches, A-HDAVT is designed to exploit the benefits from both digital and analog systems. To achieve this goal, each group of pictures is first transformed into one low-pass frame and several high-pass frames with motion-compensated temporal filtering. The critical low-pass frame is reliably transmitted as base layer in a digital mode, while high-pass frames are transmitted as enhancement layers in an analog mode to achieve desired graceful degradation performance. In analog transmission, we introduce a channel prediction-based adaptive power-distortion optimization (P-APDO) scheme to combat channel fading in mobile networks. The basic idea behind P-APDO is to perform power allocation based on the video content as well as the predicted channel status. Furthermore, we also investigate the multiuser scenarios in which the content diversity and channel diversity among users are appropriately exploited. Extensive simulations have been carried out to evaluate the performance of A-HDAVT under various degrees of channel fading. The results show that A-HDAVT achieves significant performance gains over competing schemes Robust Uncoded Video Transmission, Parcast, and Softcast in both single-user and multiuser scenarios. Hancheng Lu, Chang Wen Chen, Jun Wu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | A Joint Source-Channel Adaptive Scheme for Wireless H.264/AVC Video AuthenticationabstractAuthentication has become an emerging issue for video streaming over lossy networks. Although the advanced video coding standards, such as H.264/AVC, efficiently reduce the amount of data to be transmitted, the coding dependency brings new challenges in designing efficient stream authentication scheme. In this paper, we propose a novel joint-designed-layered source-channel adaptive scheme that integrates authentication into source and channel coding components to sufficiently use the related information to efficiently address the coding dependency and to design the optimal rate allocation scheme for the sake of end-to-end video quality. The proposed layered framework is able to minimize end-to-end quality degradation incurred by both the wireless channel noise and the authentication failure. In particular, the competing requirements of high verification probability and low authentication overhead are concurrently satisfied by the elegant design of layered hash appending with efficient adaptation to the H.264 source coding and channel conditions. A joint source-channel-authentication rate allocation scheme is then developed to achieve optimal end-to-end video quality. The experimental results on H.264 video sequences confirm the efficacy of this joint adaptive scheme and demonstrate that it indeed outperforms the state-of-the-art graph-based authentication algorithms. Xinglei Zhu, Chang Wen Chen |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2016 | Sparse Representation With Spatio-Temporal Online Dictionary Learning for Promising Video CodingabstractClassical dictionary learning methods for video coding suffer from high computational complexity and interfered coding efficiency by disregarding its underlying distribution. This paper proposes a spatio-temporal online dictionary learning (STOL) algorithm to speed up the convergence rate of dictionary learning with a guarantee of approximation error. The proposed algorithm incorporates stochastic gradient descents to form a dictionary of pairs of 3D low-frequency and high-frequency spatio-temporal volumes. In each iteration of the learning process, it randomly selects one sample volume and updates the atoms of dictionary by minimizing the expected cost, rather than optimizes empirical cost over the complete training data, such as batch learning methods, e.g., K-SVD. Since the selected volumes are supposed to be independent identically distributed samples from the underlying distribution, decomposition coefficients attained from the trained dictionary are desirable for sparse representation. Theoretically, it is proved that the proposed STOL could achieve better approximation for sparse representation than K-SVD and maintain both structured sparsity and hierarchical sparsity. It is shown to outperform batch gradient descent methods (K-SVD) in the sense of convergence speed and computational complexity, and its upper bound for prediction error is asymptotically equal to the training error. With lower computational complexity, extensive experiments validate that the STOL-based coding scheme achieves performance improvements than H.264/AVC or High Efficiency Video Coding as well as existing super-resolution-based methods in rate-distortion performance and visual quality. Wenrui Dai, Yangmei Shen, Junni Zou, Hongkai Xiong, Chang Wen Chen |
IEEE Trans. Image Process. | 6 |
| 2016 | Light Field Multi-View Video Coding With Two-Directional Parallel Inter-View PredictionabstractLight field (LF) technology has been popularly adopted by a wide range of conventional industries. However, one problem when dealing with LFs is the sheer size of data volume. There have been many multi-view video coding (MVC)-based LF video coding methods reported in the literature, aiming at finding the best prediction structure for LF video coding. It is clear that the number of possible prediction structures is unlimited, and it is also observed that the coding bit-rate can be reduced by increasing the number of bi-directionally encoded views in the prediction structure. However, none work has been conducted to analyze the relationship of the prediction structure with its coding performance. In light of this observation, we first design a new LF-MVC prediction structure by extending the inter-view prediction into a two-directional parallel structure. Analytical models for source coding rate and encoding time are developed to analyze their relationships with the prediction structure, and are proven to be well-matched to our experimental results. Experimental evaluation of two LF video sequences demonstrates that the proposed LF-MVC prediction structure can achieve a factor of 26% bit-rate reduction against the conventional MVC prediction structure for an LF video with 5×5 views, and a further 34% bit-rate reduction for an LF video with a larger 10×10 views. Compared with the state-of-the-art MVC-based LF video coding prediction structures in the literature, LF-MVC can achieve the best coding performance, and with its high encoding efficiency, is well suited for deployment in practical LF-based 3D systems. Eric Wang 0001, Wei Xiang 0001, Mark R. Pickering, Chang Wen Chen |
IEEE Trans. Image Process. | 4 |
| 2016 | No-Reference Depth Assessment Based on Edge Misalignment Errors for T + D ImagesabstractThe quality of depth is crucial in all depth-based applications. Unfortunately, the error-free ground truth is often unattainable for depth. Therefore, no-reference quality assessment is very much desired. This paper presents a novel depth quality assessment scheme that is completely different from conventional approaches. In particular, this scheme focuses on depth edge misalignment errors in texture-plus-depth (T + D) images and develops a robust method to detect them. Based on the detected misalignments, a no-reference metric is calculated to evaluate the quality of depth maps. In the proposed scheme, misalignments are detected by matching texture and depth edges through three constraints: 1) spatial similarity; 2) edge orientation similarity; and 3) segment length similarity. Furthermore, the matching is performed on edge segments instead of individual pixels, which enables robust edge matching. Experimental results demonstrate that the proposed scheme can detect misalignment errors accurately. The proposed no-reference depth quality metric is highly consistent with the full-reference metric, and is also well-correlated with the quality of synthesized virtual views. Moreover, the proposed scheme can also use the detected edge misalignments to facilitate depth enhancement in various practical texture-plus-depth-based applications. Sen Xiang, Li Yu 0003, Chang Wen Chen |
IEEE Trans. Image Process. | 3 |
| 2016 | Adaptive Quantization Parameter Cascading in HEVC Hierarchical CodingabstractThe state-of-the-art High Efficiency Video Coding (HEVC) standard adopts a hierarchical coding structure to improve its coding efficiency. This allows for the quantization parameter cascading (QPC) scheme that assigns quantization parameters (Qps) to different hierarchical layers in order to further improve the rate-distortion (RD) performance. However, only static QPC schemes have been suggested in HEVC test model, which are unable to fully explore the potentials of QPC. In this paper, we propose an adaptive QPC scheme for an HEVC hierarchical structure to code natural video sequences characterized by diversified textures, motions, and encoder configurations. We formulate the adaptive QPC scheme as a non-linear programming problem and solve it in a scientifically sound way with a manageable low computational overhead. The proposed model addresses a generic Qp assignment problem of video coding. Therefore, it also applies to group-of-picture-level, frame-level and coding unit-level Qp assignments. Comprehensive experiments have demonstrated that the proposed QPC scheme is able to adapt quickly to different video contents and coding configurations while achieving noticeable RD performance enhancement over all static and adaptive QPC schemes under comparison as well as HEVC default frame-level rate control. We have also made valuable observations on the distributions of adaptive QPC sets in the videos of different types of contents, which provide useful insights on how to further improve static QPC schemes. Tiesong Zhao, Zhou Wang 0001, Chang Wen Chen |
IEEE Trans. Image Process. | 3 |
| 2016 | Editorial: On Building a Stronger Multimedia Community
Chang Wen Chen |
IEEE Trans. Multim. | 1 |
| 2016 | A High-Fidelity and Low-Interaction-Delay Screen Sharing SystemabstractThe pervasive computing environment and wide network bandwidth provide users more opportunities to share screen content among multiple devices. In this article, we introduce a remote display system to enable screen sharing among multiple devices with high fidelity and responsive interaction. In the developed system, the frame-level screen content is compressed and transmitted to the client side for screen sharing, and the instant control inputs are simultaneously transmitted to the server side for interaction. Even if the screen responds immediately to the control messages and updates at a high frame rate on the server side, it is difficult to update the screen content with low delay and high frame rate in the client side due to non-negligible time consumption on the whole screen frame compression, transmission, and display buffer updating. To address this critical problem, we propose a layered structure for screen coding and rendering to deliver diverse screen content to the client side with an adaptive frame rate. More specifically, the interaction content with small region screen update is compressed by a blockwise screen codec and rendered at a high frame rate to achieve smooth interaction, while the natural video screen content is compressed by standard video codec and rendered at a regular frame rate for a smooth video display. Experimental results with real applications demonstrate that the proposed system can successfully reduce transmission bandwidth cost and interaction delay during screen sharing. Especially for user interaction in small regions, the proposed system can achieve a higher frame rate than most previous counterparts. Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2016 | SecSIFT: Secure Image SIFT Feature Extraction in Cloud ComputingabstractThe image and multimedia data produced by individuals and enterprises is increasing every day. Motivated by the advances in cloud computing, there is a growing need to outsource such computational intensive image feature detection tasks to cloud for its economic computing resources and on-demand ubiquitous access. However, the concerns over the effective protection of private image and multimedia data when outsourcing it to cloud platform become the major barrier that impedes the further implementation of cloud computing techniques over massive amount of image and multimedia data. To address this fundamental challenge, we study the state-of-the-art image feature detection algorithms and focus on Scalar Invariant Feature Transform (SIFT), which is one of the most important local feature detection algorithms and has been broadly employed in different areas, including object recognition, image matching, robotic mapping, and so on. We analyze and model the privacy requirements in outsourcing SIFT computation and propose Secure Scalar Invariant Feature Transform (SecSIFT), a high-performance privacy-preserving SIFT feature detection system. In contrast to previous works, the proposed design is not restricted by the efficiency limitations of current homomorphic encryption scheme. In our design, we decompose and distribute the computation procedures of the original SIFT algorithm to a set of independent, co-operative cloud servers and keep the outsourced computation procedures as simple as possible to avoid utilizing a computationally expensive homomorphic encryption scheme. The proposed SecSIFT enables implementation with practical computation and communication complexity. Extensive experimental results demonstrate that SecSIFT performs comparably to original SIFT on image benchmarks while capable of preserving the privacy in an efficient way. Zhan Qin, Jingbo Yan, Kui Ren 0001, Chang Wen Chen, Cong Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2016 | Modeling and Optimization of High Frame Rate Video Transmission Over Wireless NetworksabstractHigh frame rate (HFR) video is emerging as a new paradigm in popular multimedia applications (e.g., cloud gaming) to achieve smooth viewing experience perceived by end-users. In the context of HFR streaming video, end-to-end distortion and sending frame rate are equally important to the perceptual quality. This study presents a modeling-based approach to optimizing the HFR video transmission over wireless networks. First, we develop an analytical model dubbed FRIED (Frame Rate versus vIdEo Distortion) to characterize the tradeoff between sending frame rate and end-to-end video distortion. Second, we propose a Joint frAme Selection and FEC (Forward Error Correction) cOding (JASCO) approach based on the FRIED model to optimize the transmission performance. The efficacy of the proposed JASCO is evaluated through extensive semi-physical emulations in Exata involving H.264 video streaming. Experimental results show that JASCO outperforms the reference approaches in improving video peak signal-to-noise ratio (PSNR) at the same frame rate. Or conversely, JASCO is able to achieve higher received frame rate while guaranteeing the same video PSNR. Jiyan Wu, Chau Yuen, Ngai-Man Cheung, Junliang Chen 0001, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 5 |
| 2015 | Energy-Efficient Resource Allocation with QoS Support in Wireless Body Area NetworksabstractWireless Body Area Network (WBAN) has become a promising type of networks to provide applications such as real-time health monitoring and ubiquitous e-Health services. One challenge in the design of WBAN is that energy efficiency needs to be ensured to increase the network lifetime in such a resourceconstrained network. Another critical challenge for WBAN is that quality of service (QoS) requirements, including packet loss rate (PLR), throughput and delay, should be guaranteed even under the highly dynamic environment due to changing of body postures. In this paper, we design a unified framework of energy efficient resource allocation scheme for WBAN, in which both constraints of QoS metrics and the characteristics of dynamic links are considered. A transmission rate allocation policy (TRAP) is proposed to carefully adjust the transmission rate at each sensor such that more strict PLR requirement could be achieved even when the link quality is very poor. A QoS optimization problem is then formulated to optimize the transmission power and allocated time slots for each sensor, which minimizes energy consumption subject to the QoS constraints. Numerical results demonstrate the effectiveness of the proposed transmission rate allocation policy and the resource allocation scheme. Bin Liu 0016, Chang Wen Chen |
GLOBECOM | 4 |
| 2015 | QoS-Driven Power Control for Inter-WBAN Interference MitigationabstractWireless Body Area Networks (WBANs) are usually designed for pervasive healthcare applications. Since the primary traffic in WBAN is vital physiological signals, guaranteeing the Quality of Service (QoS) is crucial while designing WBAN. However, QoS of WBAN will be degraded in strong inter-WBAN interference environment such as hospitals and senior communities, where WBANs are densely deployed. In this paper, by focusing on a more practical WBAN model, we propose a non- cooperative power control game to mitigate inter- WBAN interference, in which the cost function is well designed by considering both QoS requirement and energy constraint. The existence of at least one Nash equilibrium (NE) point for the game is proved and a sufficient condition for the uniqueness of the NE is derived. To guarantee non-cooperative among WBANs, an interference segmentation estimate (ISE) algorithm is proposed to obtain an approximation of the NE point. Simulation results demonstrate the effectiveness of the proposed ISE algorithm. Xiaosong Zhao, Bin Liu 0016, Chang Wen Chen |
GLOBECOM | 4 |
| 2015 | A hierarchical anti-occlusion tracking algorithm based on DMPF and ORBabstractAn important issue in video target tracking is to deal with occlusion problem. In this paper, a hierarchical anti-occlusion tracking algorithm based on Dual Mode Particle Filter (DMPF) and Oriented FAST and Rotated BRIEF (ORB) is proposed to improve the location accuracy under different occlusion conditions in video target tracking. In the first layer, DMPF is used to track target and preliminarily locate its position under various occlusion status. By using corner matching in the second layer, the target is precisely located according to ORB similarity and corner coordinate, thus a more accurate position of target is obtained. A modularized occlusion judgment scheme is also presented to switch tracking mode timely and accurately in DMPF and a Quantified Matrix based on Weighted RGB color space (WQM) is introduced for both color feature creation and corner detection to save operation time. Simulation results show that the proposed algorithm could provide high tracking accuracy in a real-time manner for different occlusion status. Kejia Liu, Bin Liu 0016, Chang Wen Chen |
ICIP | 4 |
| 2015 | Service provisioning and profit maximization in network-assisted adaptive HTTP streamingabstractMobile adaptive HTTP streaming with centralized consideration of multiple streams has gained increasing interest. It poses a special challenge that the interests of both content provider and network operator need to be deliberately balanced. More importantly, the adaptation is required to be flexible enough to be ported to various systems that work under different network environments, QoE levels, and economic objectives. To address these challenges, we propose a Markov Decision Process (MDP) based network-assisted adaptation framework, wherein cost of buffering, significant playback variation, bandwidth management and income of playback are jointly investigated. We then demonstrate its promising service provisioning and maximal profit for a mobile network in which fair or differentiated service is required. Zhisheng Yan, Cédric Westphal, Chang Wen Chen |
ICIP | 4 |
| 2015 | Multi-objective content preserving warping for image stitchingabstractImage stitching has been developed to enhance users' immersive experiences by constructing new images of higher spatial resolution and broader field of view from multiple input images. Most image stitching schemes assume 2D transformation of the whole frame between images. There usually exist local misalignments of contents that do not match with the transformation and may cause artifacts such as ghosting. Several schemes including seam cutting, image blending and other image composition algorithms, have been developed to relieve such artifacts. However, these schemes are essentially post-stitching processing that do not address the alignment problem and would fail when the parallax is large. In this paper, we propose a multi-objective content preserving warping (CPW) for local adjustment of image alignment which is optimized considering feature matching, line preserving, line matching, shape smoothness and other factors to significantly reduce the visual artifacts. In particular, we adjust the image warping from their pre-warping results using the homography estimated during the image registration step. Experimental results show that the proposed CPW scheme for image stitching is able to generate much smoother stitching results comparing with the state-of-the-art methods reported in the literature. Jie Hu 0008, Dong-Qing Zhang, Hong Heather Yu, Chang Wen Chen |
ICME | 4 |
| 2015 | Discontinuous seam cutting for enhanced video stitchingabstractVideo stitching requires proper seam cutting technique to decide the boundary of the sub video volume cropped from source videos. In theory, approaches such as 3D graph-cuts that search the entire spatiotemporal volume for a cutting surface should provide the best results. However, given the tremendous data size of the camera array video source, the 3D graph-cuts algorithm is extremely resource-demanding and impractical. In this paper, we propose a sequential seam cutting scheme, which is a dynamic programming algorithm that scans the source videos frame-by-frame, updates the pixels' spatiotemporal constraints, and gradually builds the cutting surface in low space complexity. The proposed scheme features flexible seam finding conditions based on temporal and spatial coherence as well as salience. Experimental results show that by relaxing the seam continuity constraint, the proposed video stitching can better handle abrupt motions or sharp edges in the source, reduce stitching artifacts, and render enhanced visual quality. Jie Hu 0008, Dong-Qing Zhang, Hong Heather Yu, Chang Wen Chen |
ICME | 4 |
| 2015 | Image inpainting with adaptive linear predictorabstractIn this paper, a novel examplar-based inpainting algorithm with adaptive linear predictor is proposed. The patches in the damaged region are sequentially estimated with a linear combination of several nearest neighboring patches. The number of candidate patch is automatically tuned to local contexts based on Bayesian Information Criterion (BIC). The flexibility of the order-adaptive predictor makes the proposed algorithm suitable for both structural regions and detailed textures. The multi-scale framework and a novel propagation order are also involved to further improve the inpainting performance. Compared to the state-of-the-art image inpainting algorithms, experimental results show that the proposed method gives comparative or better performance. Jing Liu 0002, Guangtao Zhai, Xiaokang Yang 0001, Chang Wen Chen |
ICME | 4 |
| 2015 | Image retargeting by combining fast seam carving with neighboring probability (FSc_Neip) and scalingabstractNo single retargeting approach performs well on all images and all target sizes; therefore, hybrid algorithms are often considered as promising alternatives. However, most hybrid schemes are time consuming. In this paper, we propose a fast hybrid framework in which the Fast Content-Aware Image Distance (FCAID) is used to connect fast seam carving with neighboring probability constraints (FSc_Neip) and scaling. FCAID is used to measure the image distance between the resized image given by FSc_Neip and the original image. This fast technique is embedded within the FSc_Neip framework. Our hybrid scheme is locally applied in strip regions. This makes the retargeting scheme globally non-homogeneous. Experimental results demonstrate that our approach comprehensively outperforms other state-of-the-art techniques in terms of image quality and computational complexity. Lifang Wu, Qingyang Zheng, Yuchen Jing, Chang Wen Chen |
ICME | 6 |
| 2015 | Color Photo Makeover via Crowd Sourcing and RecoloringabstractIt is not always easy for amateur photographers to capture photos with desired colors even on a classic hot spot as the appearance of color photo dependent on many factors. This paper proposes a novel approach to recolor given photos via a crowdsourcing based makeover scheme. When a user input a photo to be recolored, the proposed system will first conduct favorite exemplars suggestion from the images hosted by the social media sites, by jointly leveraging contextual and visual information associated with the images. The recommended exemplars shall reveal the scene and context dependent color compositions and provide users with diverse possible color styles. Then, a novel superpixel-based recoloring scheme, incorporating color statistics, texture characteristics and spatial constraints into soft matching, is applied to generate new photos of desired color. Experiments and a user study demonstrate that the proposed color photo makeover is able to achieve robust recoloring results for various outdoor photos. Wengang Cheng, Ruru Jiang, Chang Wen Chen |
ACM Multimedia | 3 |
| 2015 | Exploring QoE for Power Efficiency: A Field Study on Mobile Videos with LCD DisplaysabstractDisplay power consumption has become a major concern for both mobile users and design engineers, especially considering the prevalence of today's video-rich mobile services. The power consumption of liquid crystal display (LCD), a dominant mobile display technology, can be reduced by dynamic backlight scaling (DBS). However, such dynamic changes of screen brightness may degrade users' quality of experience (QoE) in viewing videos. How would QoE be impacted by different DBS strategies has not yet been understood clearly and thus obscures the way to achieve systematic power saving. In this paper, we take a first step to explore the QoE of DBS on smartphones and aim at maximally enhancing the display power performance without negatively impacting users' QoE. In particular, we conduct three motivational studies to uncover the inherent relationship between QoE and backlight scaling frequency, magnitude, and temporal consistency, respectively. Motivated by the findings of these studies, we design a suite of techniques to implement a comprehensive DBS strategy. We demonstrate an example application of the proposed DBS designs in a mobile video streaming system. Measurements and user evaluations show that more than 40% system power reduction, or equivalently, 20% more power savings than the non-QoE approaches, can be achieved without QoE impairment. Zhisheng Yan, Qian Liu 0001, Tong Zhang 0002, Chang Wen Chen |
ACM Multimedia | 4 |
| 2015 | Automatic Contrast Enhancement Technology With Saliency PreservationabstractIn this paper, we investigate the problem of image contrast enhancement. Most existing relevant technologies often suffer from the drawback of excessive enhancement, thereby introducing noise/artifacts and changing visual attention regions. One frequently used solution is manual parameter tuning, which is, however, impractical for most applications since it is labor intensive and time consuming. In this research, we find that saliency preservation can help produce appropriately enhanced images, i.e., improved contrast without annoying artifacts. We therefore design an automatic contrast enhancement technology with a complete histogram modification framework and an automatic parameter selector. This framework combines the original image, its histogram equalized product, and its visually pleasing version created by a sigmoid transfer function that was developed in our recent work. Then, a visual quality judging criterion is developed based on the concept of saliency preservation, which assists the automatic parameters selection, and finally properly enhanced image can be generated accordingly. We test the proposed scheme on Kodak and Video Quality Experts Group databases, and compare with the classical histogram equalization technique and its variations as well as state-of-the-art contrast enhancement approaches. The experimental results demonstrate that our technique has superior saliency preservation ability and outstanding enhancement effect. Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Smart Downlink Scheduling for Multimedia Streaming Over LTE Networks With Hard HandoffabstractThis paper presents a novel smart downlink scheduling scheme to enhance the performance of multimedia transmission over long-term evolution (LTE) networks. LTE represents a promising framework for the next-generation broadband multimedia services because of its significantly increased data rate over 3G cellular networks. However, the current LTE scheduling scheme has largely been designed for general data traffic without adequate consideration of multimedia characteristics. Moreover, the hard handoff (HO) procedure adopted in LTE will further degrade the multimedia services when the mobile user moves from one cell to another. Even with increased data rate, current LTE systems still cannot deliver the expected quality of service (QoS) to the mobile users under various mobility scenarios, especially when the hard HO is evoked. Aiming at overcoming these major challenges, we develop in this paper a QoS-driven smart downlink scheduling scheme for enhanced multimedia transmission over LTE networks. The proposed design shall take the following QoS metrics into consideration: 1) the delay constraint of voice-over-Internet protocol flow; 2) the packet deadline of video flow; and 3) the service degradation induced by the hard HO procedure. We achieve the design objectives by creating three QoS-driven operational control modules: 1) a transmission delay control module to ensure the on-time arrival of various types of multimedia data; 2) an HO control module to warrant continuous multimedia services when the user moves across a cell boundary; and 3) a resource allocation module to strategically map the requested flows to best fit radio resource blocks. The simulation results confirm the performance gain of the proposed scheme. Qian Liu 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Enabling Adaptive High-Frame-Rate Video Streaming in Mobile Cloud Gaming ApplicationsabstractHigh-frame-rate (HFR) video is emerging in popular gaming applications to enhance the smooth experience perceived by end users. However, it is challenging to guarantee the delivery quality of HFR video in mobile cloud gaming scenarios because of the high transmission rate and limited wireless resources. To address this critical problem, we develop a novel transmission scheduling framework dubbed AdaPtive HFR vIdeo Streaming (APHIS). The term adaptive indicates this scheme's capability in dynamically adjusting the video traffic load and forward error correction (FEC) coding. First, we propose an online video frame selection algorithm to minimize the total distortion based on the network status, input video data, and delay constraint. Second, we introduce an unequal FEC coding scheme to provide differentiated protection for Intra (I) and Predicted (P) frames with low-latency cost. The proposed APHIS framework is able to appropriately filter video frames and adjust data protection levels to optimize the quality of HFR video streaming. We conduct extensive emulations in Exata involving HFR video encoded with H.264 codec. Experimental results show that APHIS outperforms the reference transmission schemes in terms of video peak signal-to-noise ratio, end-to-end delay, and goodput. Therefore, we recommend APHIS for delivering HFR video streaming in mobile cloud gaming systems. Jiyan Wu, Chau Yuen, Ngai-Man Cheung, Junliang Chen 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2015 | Robust Transmission of Scalable Video Coding Bitstream Over Heterogeneous NetworksabstractVideo transmission using Scalable Video Coding (SVC) enables the functionalities of adaptation for both channel bandwidth and terminal device. However, link adaptation, which requires the video bitstream to provide error resilient scalability (ERS), is currently not supported in SVC. In this paper, we propose a robust SVC bitstream transmission framework, which integrates both error resilient (ER) coding and error concealment (EC) to enable the desired link adaptability. First, at the encoder, we design an ER coding scheme for SVC using redundant pictures, in which redundant picture information (RPI) is generated under rate-distortion criteria and transmitted to the media gateway together with the SVC bitstream. Then, ERS can be effectively achieved at the media gateway by selectively adding or removing coded pictures of different video coding layers according to the RPI and current network link status. Finally, at the decoder, we take advantage of the unique characteristics of SVC to develop an EC scheme based on both the proposed Virtual-BLSkip method and Wiener filtering to recover the lost or removed enhancement layer pictures. The proposed framework has negligible impact on the coding efficiency because it transmits RPIs rather than directly adding redundancy into the original bitstream. More importantly, it successfully strikes a balance between coding efficiency, error resiliency, and operation complexity. It is capable of providing the desired ERS with extremely low complexity, and maintaining virtually the same bit rate as input video bitstream. The experimental results show that the proposed framework achieves significant performance gains over existing ER transmission schemes. Houqiang Li, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Layered Compression for High-Precision Depth DataabstractWith the development of depth data acquisition technologies, access to high-precision depth with more than 8-b depths has become much easier and determining how to efficiently represent and compress high-precision depth is essential for practical depth storage and transmission systems. In this paper, we propose a layered high-precision depth compression framework based on an 8-b image/video encoder to achieve efficient compression with low complexity. Within this framework, considering the characteristics of the high-precision depth, a depth map is partitioned into two layers: 1) the most significant bits (MSBs) layer and 2) the least significant bits (LSBs) layer. The MSBs layer provides rough depth value distribution, while the LSBs layer records the details of the depth value variation. For the MSBs layer, an error-controllable pixel domain encoding scheme is proposed to exploit the data correlation of the general depth information with sharp edges and to guarantee the data format of LSBs layer is 8 b after taking the quantization error from MSBs layer. For the LSBs layer, standard 8-b image/video codec is leveraged to perform the compression. The experimental results demonstrate that the proposed coding scheme can achieve real-time depth compression with satisfactory reconstruction quality. Moreover, the compressed depth data generated from this scheme can achieve better performance in view synthesis and gesture recognition applications compared with the conventional coding schemes because of the error control algorithm. Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen |
IEEE Trans. Image Process. | 5 |
| 2015 | Message From the Editor-in-ChiefabstractPresents the Editor-in-Chief's message for this issue of the publication. Chang Wen Chen |
IEEE Trans. Multim. | 1 |
| 2014 | JSM-2 Based Joint ECG Compression Exploiting Temporal and Structural DependencyabstractIn Wireless Body Area Networks (WBAN), the electrocardiogram (ECG) signal is an important class of bio-signals which needs to be transmitted and stored for diseasesdiagnostics. Due to the resource limitation in WBAN, the large amount of ECG signals need to be compressed before transmission and reconstructed with high accuracy. In this paper, we propose a novel CS-based ECG compression scheme, which considers both structural dependency and temporal dependency among ECG signals. The received ECG heartbeats are first classified into different classes and the statistical support information (SSI) is then established for each class. By using the corresponding SSI, a more accurate partially known support (PKS) will be obtained and the joint reconstruction performance of ECG signals could be improved consequently. Simulation results show that the proposed ECG compression scheme outperforms existing schemes, especially when the dimension of sampled measurements is low. Jinguo Luo, Bin Liu 0016, Chang Wen Chen |
BSN | 3 |
| 2014 | Privacy-preserving outsourcing of image global feature detectionabstractThe amount and availability of user-contributed image data have been dramatically increased during the past ten years. Popular multimedia social networks, e.g. Flicker, commonly utilize user image data to construct user behavior models, social preferences, etc., for the purpose of effective advertisement, better user retention and attraction, and many others. Existing practices of data utilization, however, seriously deteriorate users' personal privacy and have led to increasing criticisms and legislation pressures. In this paper, we aim to construct a privacy-preserving feature detection scheme over encrypted image data. The proposed system enables an interested party to perform a variety of image feature detection tasks, including visual descriptors in MPEG-7 standard, while protecting user privacy relating to image contents. We implement a prototype system based on somewhat homomorphic encryption scheme and the benchmark Caltech256 database. The experimental results show that our system can guarantee effective image feature detection without sacrificing user privacy. Zhan Qin, Jingbo Yan, Kui Ren 0001, Chang Wen Chen, Cong Wang 0001, Xinwen Fu |
GLOBECOM | 4 |
| 2014 | Bayesian game based power control scheme for inter-WBAN interference mitigationabstractWireless body area network (WBAN) is an emerging technology that provides socialized health monitoring service. However, the quality of service can be severely degraded by concomitant inter-WBAN interference in some specific environments where multiple WBANs are densely deployed, e.g., hospitals and senior citizen communities. In this work, we propose a Bayesian game based power control scheme to mitigate the impact of inter-WBAN interference. By modeling WBANs as players and active links as types of players in the Bayesian game model, the proposed power control scheme tries to maximize each player's expected payoff involving both throughput and energy efficiency. We prove the existence of Bayesian equilibrium (BE) for the proposed power control game and also derive a practical sufficient condition for the uniqueness of BE. A harmonic mean based algorithm is then proposed to obtain an approximation of BE point without the need to pass message among WBANs, which satisfies the non-cooperative manner for inter-WBAN interference mitigation. Simulation results show that the proposed algorithm can converge to the BE point effectively. Bin Liu 0016, Chang Wen Chen |
GLOBECOM | 4 |
| 2014 | Long scene panorama generation for indoor environmentabstractLong scene panorama (LSP) can provide the immersive experience for an extended view in one single photograph. The state-of-the-art image mosaicing based LSP techniques are significantly constrained by their high computational complexity, limiting their applications to platforms with rich computing resource. Slit-scanning based schemes, although implementation friendly, heavily rely on the consistency of the camera's speed and suffer from pushbroom distortion. In this paper, we propose a novel scheme which can generate the LSP for the indoor environment via consumer-grade cameras. A novel adaptive resampling scheme is introduced to generate more natural looking panoramas from videos captured by cameras with non-constant moving speed. A multiple slits replacement scheme is also developed to reduce the zigzag distortion and to handle the depth change. Experimental results show that the proposed scheme can achieve satisfactory performance with handheld cameras. Jie Hu 0008, Dong-Qing Zhang, Hong Heather Yu, Chang Wen Chen |
ICIP | 4 |
| 2014 | Finding your spot: A photography suggestion system for placing human in the sceneabstractCapturing a professional like photo is always a challenging task, especially for novice users. This paper proposes a photography suggestion approach to assist users to take high visual quality photos with human in the scene. In this research, we first investigate a set of aesthetic composition rules and visual perception principles to construct an aesthetic score prediction model in order to measure visual quality in terms of photo composition. Then we conduct a study on professional photos to define a proper size for the enclosure of human into the picture of a given scene. The proposed approach is able to leverage saliency and geometric detection to represent composition features. Finally, we utilize an efficient hierarchical search to obtain the optimal enclosure for human in the scene. Extensive experiments have been performed for hot spot landmark locations. Through subjective evaluation, the results show that the proposed approach can effectively provide appealing composition recommendation to help users take high quality photos with human in the scene. Yangyu Fan, Chang Wen Chen |
ICIP | 3 |
| 2014 | High resolution free-view interpolation of planar structureabstractIn many applications, users desire to access a random viewing angle of a scene in high resolution while this specifically queried imagery is not among the acquired sample images. In theory, 3D reconstruction based rendering can generate such an image. However, accurate camera calibration over large scale photo collections is needed and is highly complex in nature. Image stitching based approachescan also be applied. However, such scheme is unable to provide free view interpolation or resolution enhancement. In this paper, we present a novel free view image super-resolution scheme to interpolate free views for planar structures. We construct a Bayesian model and marginalize it over photometric regulation and geometric registration parameters applying the Lie group theory. Experimental results show that the proposed scheme is able to achieve desired performance against the state-of-the-art image super-resolution approaches and successfully obtain registration in full 6 degree-of-freedom (DOF). Compared with existing image based rendering schemes, the proposed scheme achieves free view interpolation for planar structures with higher resolution and less distortion. Jie Hu 0008, Dong-Qing Zhang, Hong Heather Yu, Chang Wen Chen |
ICME | 4 |
| 2014 | Secure media sharing in the cloud: Two-dimensional-scalable access control and comprehensive key managementabstractMedia sharing in cloud environment, which supports sharing media content at any time and from anywhere, is a promising paradigm of social interaction. However, it also brings forth security issues in terms of data confidentiality and access control on media data consumers with different access privileges. One promising solution is scalable media access control, which is capable of providing data confidentiality by encrypting the media data and issuing the key to only authorized data consumers. More importantly, it can empower the data distributor to provide the same media content with various quality levels to the consumers with different privileges. Traditional schemes without utilizing the cloud resources achieve scalable media access control by generating access keys using hash chains. Despite of their computational efficiency, such schemes suffer from various problems including vulnerability to user collusion attack in two-dimensional case, inflexible key distribution, and ambiguous key revocation strategy. In this paper, we propose a novel two-dimensional-scalable access control by generating access keys based on Attribute-Based Encryption (ABE) algorithm. Moreover, the proposed scheme can efficiently achieve comprehensive key management including key distribution and key revocation by fully exploiting the cloud. Security analysis shows that the proposed scheme is able to provide collusion resistance, as well as forward and backward secrecy. We have also evaluated the efficiency of the scheme through numerical analysis and initial implementation. Changsha Ma, Chang Wen Chen |
ICME | 2 |
| 2014 | No-reference depth quality assessment for texture-plus-depth imagesabstractIn 3D video (3DV) and free-viewpoint video (FVV), it is vitally important to detect the errors and assess depth quality. However, since ground-truth depth maps are often unattainable, assessing depth quality without reference becomes an imperative task for many applications. This research considers the texture-plus-depth format in 3DV and FVV, and focuses on the misalignment error at depth discontinuities. A matching algorithm between depth and texture edges is proposed to determine corresponding matching pairs and to identify serious mismatches. The matching procedure is based on both spatial distances and direction similarities between texture and depth edges, and the algorithm is performed between edge segments, instead of edge pixels, in order to improve the robustness of matching. Furthermore, an adaptive algorithm is designed to divide depth edges into segments with different lengths based on the curvature of the edges. After the matching is completed, misalignments between matching pairs are used to generate a no-reference assessment metric. Experimental results demonstrate that the proposed matching scheme is able to achieve accurate matches between depth and texture edges. More importantly, it has been shown that the correlations between the proposed metric and the widely accepted full-reference metric are greater than 0.9, making this no-reference depth quality assessment scheme suitable for contemporary 3DV and FVV applications. Sen Xiang, Jingteng Xue, Li Yu 0003, Chang Wen Chen |
ICME | 4 |
| 2014 | QoE continuum driven HTTP adaptive streaming over multi-client wireless networksabstractDifferent from traditional HTTP adaptive streaming (HAS) in which only one client is considered, HAS over multi-client wireless networks faces new challenges. The Quality of Experience (QoE) of users becomes unstable due to users' competition for shared bandwidth. It is thus important to accurately estimate the perceived experience of users and then adapt the streaming process accordingly. Furthermore, the QoE fairness among multiple clients subscribing to the same services shall also be addressed. In this research, we propose a QoE continuum driven HAS adaptation algorithm to address these challenges. We model the QoE continuum as an integrated consideration of cumulative playback quality and playback smoothness. Based on this model, we jointly optimize the quality adaptation of multiple users by considering both QoE history and channel status. Moreover, we propose to use quantization parameter and segment size to represent the video files in a fine-grained fashion, in order to more effectively capture the bandwidth fluctuation. The results from extensive simulations show that the proposed scheme can provide balanced and satisfactory QoE among multiple clients. Zhisheng Yan, Jingteng Xue, Chang Wen Chen |
ICME | 3 |
| 2014 | Robust uncoded video transmission over wireless fast fading channelabstractThis research studies robust uncoded video transmission over wireless fast fading channel, where only statistical channel state information (CSI) is available at the transmitter. We observe that increasing channel diversity for high priority (HP) data is essential to improving the robustness of video transmission in fading channels. By utilizing the noise and loss resilient nature of video, we find it possible to design a more robust system by re-allocating the power and channel uses among HP and LP (low priority) data. With total power and channel use constraints, we derive an optimal resource allocation scheme under the squared error distortion criterion. In particular, we first propose a new power allocation algorithm at given channel allocation. Second, based on the proposed power allocation algorithm, we design a channel allocation algorithm to strike the tradeoff between the diversity increase of HP data and the information loss of LP data. Third, under known noise power distribution, we derive the optimal resource allocation for uncoded video multicast. Simulations show that the proposed system achieves 2dB and 5dB gain in average and outage PSNR over Softcast in video unicast, and around 1.4dB and 4dB gain in multicast. Hao Cui 0001, Chong Luo 0001, Chang Wen Chen, Feng Wu 0001 |
INFOCOM | 3 |
| 2014 | A privacy-aware cloud-assisted healthcare monitoring system via compressive sensingabstractWireless sensors are being increasingly used to monitor/collect information in healthcare medical systems. For resource-efficient data acquisition, one major trend today is to utilize compressive sensing, for it unifies traditional data sampling and compression. Despite the increasing popularity, how to effectively process the ever-growing healthcare data and simultaneously protect data privacy, while maintaining low overhead at sensors, remains challenging. To address the problem, we propose a privacy-aware cloud-assisted healthcare monitoring system via compressive sensing, which integrates different domain techniques with following benefits. By design, acquired sensitive data samples never leave sensors in unprotected form. Protected samples are later sent to cloud, for storage, processing, and disseminating reconstructed data to receivers. The system is privacy-assured where cloud sees neither the original samples nor underlying data. It handles well sparse and general data, and data tampered with noise. Theoretical and empirical evaluations demonstrate the system achieves privacy-assurance, efficiency, effectiveness, and resource-savings simultaneously. Cong Wang 0001, Bingsheng Zhang, Kui Ren 0001, Janet Roveda, Chang Wen Chen |
INFOCOM | 5 |
| 2014 | High frame rate screen video coding for screen sharing applicationsabstractIn this paper, we propose a high frame rate screen video compression scheme aiming at improving the interactive user experience on screen sharing applications. The proposed screen video compression is performed as two-layer coding: a base layer coding using the conventional video codec and an enhancement layer coding using the proposed open-loop coding scheme. For efficient frame level layer selection and compression, the content update of each frame is evaluated through global motion detection. The screen frame with significant content update is fed to the conventional video encoder in base layer. In contrast, the frame with little update is compressed in enhancement layer in which the duplicate content is indicated by global motion vector and skip flag while the updated content is encoded by distinct intra modes in terms of inherent local features. The experimental results demonstrate that for the screen video containing interaction, the proposed coding scheme can achieve 3.09ms/frame encoding rate and 2.33ms/frame decoding rate with efficient rate distortion performance. Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen |
ISCAS | 5 |
| 2014 | Pose Maker: A Pose Recommendation System for Person in the Landscape PhotographingabstractTo pose like a fashion model and to take professional grade photo are always two challenging tasks, especially for novice users. Pose Maker, an innovative system for pose and photo composition recommendation and synthesis, is developed in this work. Given a user-provided clothing color and gender, this system shall not only offer some suitable poses, but also assist users to take high visual quality photos by generating the visual effect of person in the landscape pictures. To recommend poses, we first propose a view specific professional learning model to help users select several compatible poses for a given image. Based on selection results, a hierarchical pose image synthesis module is designed to synthesize the selected poses along with the scene to be captured in the most suitable position and size taking into consideration several important factors in picture composition, including aesthetic photography principles, color harmony and focal length parameters. Extensive experimental evaluations and analysis on test images of various conditions demonstrate the effectiveness of the proposed system. Yangyu Fan, Chang Wen Chen |
ACM Multimedia | 3 |
| 2014 | Towards Efficient Privacy-preserving Image Feature Extraction in Cloud ComputingabstractAs the image data produced by individuals and enterprises is rapidly increasing, Scalar Invariant Feature Transform (SIFT), as a local feature detection algorithm, has been heavily employed in various areas, including object recognition, robotic mapping, etc. In this context, there is a growing need to outsource such image computation with high complexity to cloud for its economic computing resources and on-demand ubiquitous access. However, how to protect the private image data while enabling image computation becomes a major concern. To address this fundamental challenge, we study the privacy requirements in outsourcing SIFT computation and propose SecSIFT, a high performance privacy-preserving SIFT feature detection system. In previous private image computation works, one common approach is to encrypt the private image in a public key based homomorphic scheme that enables the original processing algorithms designed for plaintext domain to be performed over ciphertext domain. In contrast to these works, our system is not restricted by the efficiency limitations of homomorphic encryption scheme. The proposed system distributes the computation procedures of SIFT to a set of independent, co-operative cloud servers, and keeps the outsourced computation procedures as simple as possible to avoid utilizing homomorphic encryption scheme. Thus, it enables implementation with practical computation and communication complexity. Extensive experimental results demonstrate that SecSIFT performs comparably to original SIFT on image benchmarks while capable of preserving the privacy in an efficient way. Zhan Qin, Jingbo Yan, Kui Ren 0001, Chang Wen Chen, Cong Wang 0001 |
ACM Multimedia | 4 |
| 2014 | Admission Control for Wireless Adaptive HTTP Streaming: An Evidence Theory Based ApproachabstractIn this research, we propose an evidence theory based admission control scheme for wireless cellular adaptive HTTP streaming systems. This novel scheme allows us to effectively address the uncertainty and inaccuracy in QoE management and network estimation, and seamlessly grant or deny the access requests. Specifically, based on recent work of QoE continuum model and QoE continuum driven adaptation algorithm, we utilize Dempster-Shafer evidence theory to assign proper degree of belief to admission, rejection and an uncertainty decision for each user's evidence. We then can strategically combine the weighted evidence of multiple users and make the final decision. The evaluation results show that the proposed scheme can provide satisfactory QoE for both existing and new users while still achieving comparable bandwidth efficiency. Zhisheng Yan, Chang Wen Chen, Bin Liu 0016 |
ACM Multimedia | 2 |
| 2014 | Social TV analytics: a novel paradigm to transform TV watching experienceabstractThe blooming online social networks have revolutionized the way information is created, disseminated and consumed, positing significant challenges to the conventional information propagation carriers, especially for the television land-scape. In this paper, we design and develop a multi-screen cloud social TV integrated with social media via a second screen as a novel paradigm in response to this trend. Our system comprises three building blocks, including a cloud based social TV system, a social TV analytics system, and a multi-screen orchestration system. In particular, we leverage the cloud infrastructure to improve the system scalability, and design intelligent social media collection & analysis mechanisms to mine deeper social perception. Furthermore, we demonstrate two key features of our system based on a real user case. Han Hu 0003, Yonggang Wen 0001, Chang Wen Chen, Tat-Seng Chua |
MMSys | 5 |
| 2014 | A novel segmentation based video-denoising method with noise level estimation
Yang Cao 0010, Zhengjun Zha, Jing Zhang 0037, Chang Wen Chen |
Inf. Sci. | 5 |
| 2014 | Editorial: Special issue on QoE in 2D/3D video systems
Tasos Dagiuklas, Luigi Atzori, Periklis Chatzimisios, Chang Wen Chen, Weisi Lin |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | A new closed loop method of super-resolution for multi-view images
Jing Zhang 0037, Yang Cao 0010, Zhigang Zheng, Chang Wen Chen, Zengfu Wang |
Mach. Vis. Appl. | 4 |
| 2014 | Robust Linear Video Transmission Over Rayleigh Fading ChannelabstractThis research addresses the problem of robust linear video transmission over the Rayleigh fading channel, where only statistical channel state information (CSI) is available to the sender. We observe that discarding low-priority (LP) data and saving the channel uses for high-priority (HP) data can significantly improve the quality of the received video. We formulate an optimization problem that aims to minimize the total squared error of a multi-variant Gaussian random vector under the given bandwidth and power resources. To tame the complexity of this NP-hard problem, we analyze two sub-problems, namely power allocation and bandwidth allocation, and propose an iterative algorithm to approximate the solution. Subsequently, we propose a one-pass two-step fast algorithm that further reduces both algorithmic and computational complexity. A linear video transmission system is implemented based on the proposed algorithm. Simulations show that our system significantly outperforms Soft-Cast, and the PSNR gain at 5th percentile of 1000 test runs is between 4.0 dB and 7.5 dB under varying noise levels. Hao Cui 0001, Chong Luo 0001, Chang Wen Chen, Feng Wu 0001 |
IEEE Trans. Commun. | 3 |
| 2014 | Message from the Editor-in-Chief
Chang Wen Chen |
IEEE Trans. Multim. | 1 |
| 2014 | Cloud Mobile Media: Reflections and OutlookabstractThis paper surveys the emerging paradigm of cloud mobile media. We start with two alternative perspectives for cloud mobile media networks: an end-to-end view and a layered view. Summaries of existing research in this area are organized according to the layered service framework: i) cloud resource management and control in infrastructure-as-a-service (IaaS), ii) cloud-based media services in platform-as-a-service (PaaS), and iii) novel cloud-based systems and applications in software-as-a-service (SaaS). We further substantiate our proposed design principles for cloud-based mobile media using a concrete case study: a cloud-centric media platform (CCMP) developed at Nanyang Technological University. Finally, this paper concludes with an outlook of open research problems for realizing the vision of cloud-based mobile media. Yonggang Wen 0001, Joel J. P. C. Rodrigues, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2014 | Socialized Mobile Photography: Learning to Photograph With Social Context via Mobile DevicesabstractThe popularity of mobile devices equipped with various cameras has revolutionized modern photography. People are able to take photos and share their experiences anytime and anywhere. However, taking a high quality photograph via mobile device remains a challenge for mobile users. In this paper we investigate a photography model to assist mobile users in capturing high quality photos by using both the rich context available from mobile devices and crowdsourced social media on the Web. The photography model is learned from community-contributed images on the Web, and dependent on user's social context. The context includes user's current geo-location, time (i.e., time of the day), and weather (e.g., clear, cloudy, foggy, etc.). Given a wide view of scene, our socialized mobile photography system is able to suggest the optimal view enclosure (composition) and appropriate camera parameters (aperture, ISO, and exposure time). Extensive experiments have been performed for eight well-known hot spot landmark locations where sufficient context tagged photos can be obtained. Through both objective and subjective evaluations, we show that the proposed socialized mobile photography system can indeed effectively suggest proper composition and camera parameters to help the user capture high quality photos. Wenyuan Yin, Tao Mei 0001, Chang Wen Chen, Shipeng Li 0001 |
IEEE Trans. Multim. | 3 |
| 2014 | Joint Spectrum and Power Auction With Multiauctioneer and Multibidder in Coded Cooperative Cognitive Radio NetworksabstractIn the existing cooperative cognitive radio (CR), qualities of service (e.g., rate or outage probability) of primary users can be improved by cooperative transmission from secondary users. However, as owners of the spectrum, primary users' traffic demands are relatively easy to satisfy. The rationale is that they would be more interested in the benefits of other formats (e.g., revenue), rather than the enhanced rate. In this paper, we propose a new cooperative CR framework, where primary users assist in the transmissions of secondary users. In exchange for this concession, primary users receive payments from secondary users for the spectrum and cooperative transmit power being used in cooperation. An auction-theoretic model with multiple auctioneers, multiple bidders, and multiple commodities is developed for a joint spectrum and cooperative power allocation. Since the spectrum and power are two heterogeneous but correlated commodities, their properties are specifically considered in auction strategy design. Finally, we mathematically prove the convergence of the proposed auction game, and show with numerical results, that the proposed auction is beneficial to both primary and secondary users. Junni Zou, Hongkai Xiong, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 4 |
| 2013 | JSM-2 based ECG compression with statistical support predictionabstractThis paper addresses the problem of developing an efficient compression scheme with high quality and low computational complexity for ECG signal compression. Taking into account the joint sparsity existing in ECG data and the temporal dependencies in ECG signal sequence, a novel scheme for JSM-2 based ECG compression is developed to exploit these characteristics. We first predict support information in sparse domain from the previous ECG data for the current recovery process. Then a modified Simultaneous Orthogonal Matching Pursuit Algorithm (SOMP) algorithm is proposed to incorporate the idea of support information establishment for JSM-2 based ECG compression. Simulation results show that the proposed JSM-2 based ECG compression scheme with statistical support prediction outperforms existing schemes with enhanced performance and low computational complexity. Sucheng Yu, Bin Liu 0016, Chi Zhang 0001, Chang Wen Chen |
Healthcom | 5 |
| 2013 | MOS-Based Channel Allocation Schemes for Mixed Services over Cognitive Radio NetworksabstractIn cognitive radio (CR) networks, secondary users (SUs) may have various applications such as multimedia delivery and file download, resulting in different bandwidth requirements. In this paper, we propose a channel allocation scheme for mixed services, especially video streaming, based on mean opinion score (MOS) maximization. MOS is an effective metric of Quality of Experience (QoE) that directly measures the satisfaction of the end users. The cognitive radio network base station (CRNBS) collects all the SUs' application information and allocates available channel resource to the SUs with the overall user perceived MOS maximized and fairness among SUs ensured. The simulation results confirm that the proposed MOS-based channel allocation scheme outperforms the conventional good put-based scheme in terms of overall user satisfaction. Bin Liu 0016, Yuanzhi Yao, Nenghai Yu, Chang Wen Chen |
ICIG | 5 |
| 2013 | Enhancing multimedia QoS with device-to-device communication as an underlay in lte networksabstractWe present in this paper a new device-to-device (D2D) communication scheme to create a QoS enhanced multimedia services for LTE users who are physically close to each other. This scheme has a potential to provide higher bandwidth service with reduced delay and reduced power consumption. However, a significant challenge from such D2D opportunity is to establish the direct link between nearby devices without interference to other regular LTE users. In this research, we develop a novel user selective resource allocation scheme that allows D2D links to share the same radio resources with LTE regular users while proactively avoiding the interferences with intelligent user selection. The main contribution of the proposed scheme is in the innovative design of the sub-channel selection algorithms for both D2D links and regular LTE users. These algorithms serve dual purpose: (1) providing desired QoS enhancement to multimedia consumers in D2D links while minimizing the interference to regular users; (2) maximizing the sum rate of the LTE networks considering the interference from D2D links and the fairness issue. The simulation results show that the total achievable capacity of LTE networks is dramatically enhanced by D2D communication with this new user selective resource allocation. Qian Liu 0001, Hong Heather Yu, Chang Wen Chen |
ICME | 3 |
| 2013 | Layered screen video coding leveraging hardware video codecabstractIn this paper, we propose a layered screen video coding scheme based on existing video codecs to leverage hardware video codec for efficient screen video compression. In this scheme, the screen video compression is performed as two-layer coding: base layer coding and enhancement layer coding. The screen video is first analyzed in both frame and block levels for useful temporal and spatial information extraction to assist coding content selection in each layer. The non-skip screen frames are directly compressed by the conventional video codec in the base layer, while the screen contents sensitive to the video quality degradation are selected for improved coding in the enhancement layer. For contents to be enhanced, two intra coding modes are designed to improve the quality of the compressed text/graphics contents and suppress the artifacts introduced by chroma downsampling. The experimental results demonstrate that the screen video quality is improved objectively and subjectively by the proposed scheme with low cost on bitrate and computation complexity. Moreover, an average of 2.95dB coding gain is achieved in high bitrate. Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen |
ICME | 5 |
| 2013 | Mining visualnessabstractTo understand which concepts are visualizable and to what extent they can be visualized are worthwhile for multimedia and computer vision research. Unfortunately, few previous works have ever touched such topics. In this paper, we propose an unified model to automatically identify visual concepts and estimate their visual characteristics, or visualness, from a large-scale image dataset. To this end, an image heterogeneous graph is first built to integrate various visual features, and then a simultaneous ranking and clustering algorithm is introduced to generate visually and semantically compact image clusters, named visualsets. Based on the visualsets, visualizable concepts are discovered and their visualness scores are estimated. The experimental results demonstrate the effectiveness of the proposed schema. Zheng Xu 0002, Xin-Jing Wang, Chang Wen Chen |
ICME | 3 |
| 2013 | Fast and improved seam carving with strip partition and neighboring probability constraintsabstractSeam carving is an effective way of image retargeting. However, existing seam carving schemes often come with unacceptable artifacts and are quite time consuming. In this paper, we propose a fast seam carving scheme with strip partition and neighboring probability constraints to resolve these two problems simultaneously. Firstly, we split the original image into several strips of equal space and we estimate the importance of each strip by its average saliency values. This partition results that more seams are removed from the strips consisting of more unimportant regions while fewer seams are removed from that of more important regions. Then, we establish the adjacency relationship by maximum correlation [8]. The neighboring probability is obtained to describe the neighboring relationship between the seams. Finally, by combining the neighboring probability and their accumulated energy, least important seams are removed. The neighboring probability constraint ensures that the seam removal is distributed to avoid abrupt changes in the scene. This leads to an improved quality in the resized image. The experimental results show that the proposed approach performs better than the state-of-the-art seam carving schemes. Lifang Wu, Lianchao Cao, Chang Wen Chen |
ISCAS | 3 |
| 2013 | Automatic generation of social media snippets for mobile browsingabstractThe ongoing revolution in media consumption from traditional PCs to the pervasiveness of mobile devices is driving the adoption of social media in our daily lives. More and more people are using their mobile devices to enjoy social media content while on the move. However, mobile display constraints create challenges for presenting and authoring the rich media content on screens with limited display size. This paper presents an innovative system to automatically generate magazine-like social media visual summaries, which is called "snippet," for efficient mobile browsing. The system excerpts the most salient and dominant elements, i.e., a major picture element and a set of textual elements, from the original media content, and composes these elements into a text overlaid image by maximizing information perception. In particular, we investigate a set of aesthetic rules and visual perception principles to optimize the layout of the extracted elements by considering display constraints. As a result, browsing the snippet on mobile devices is just like quickly glancing at a magazine. To the best of our knowledge, this paper represents one of the first attempts at automatic social media snippet generation by studying aesthetic rules and visual perception principles. We have conducted experiments and user studies with social posts from news entities. We demonstrated that the generated snippets are effective at representing media content in a visually appealing and compact way, leading to a better user experience when consuming social media content on mobile devices. Wenyuan Yin, Tao Mei 0001, Chang Wen Chen |
ACM Multimedia | 3 |
| 2013 | Prediction-based dynamic relay transmission scheme for Wireless Body Area NetworksabstractTo support long-term pervasive healthcare services, communications in Wireless Body Area Networks (WBANs) need to be both reliable and energy-efficient. As a cooperative transmission method, relay transmission scheme works effectively in resisting shadowing effect and improving reliability in WBANs. However, the extra energy consumption introduced by relay transmission is very high, which can shorten the lifetime of the whole network. In this paper, temporal and spatial correlation models for on-body channels are first presented to better characterize the slow fading effect of on-body channels. Then a prediction-based dynamic relay transmission (PDRT) scheme that makes full use of the correlation characteristics of on-body channels is proposed. In the PDRT scheme, “when to relay” and “who to relay” are decided in an optimal way based on the last known channel states. Moreover, neither extra signaling procedure nor dedicated channel sensing period is needed. Simulation results show that the PDRT scheme achieves significant performance improvement in energy efficiency, as well as ensuring the transmission reliability. Bin Liu 0016, Zhisheng Yan, Chi Zhang 0001, Chang Wen Chen |
PIMRC | 5 |
| 2013 | Structured Set Intra Prediction With Discriminative Learning in a Max-Margin Markov Network for High Efficiency Video CodingabstractThis paper proposes a novel model on intra coding for High Efficiency Video Coding (HEVC), which simultaneously predicts blocks of pixels with optimal rate distortion. It utilizes the spatial statistical correlation for the optimal prediction based on 2-D contexts, in addition to formulating the data-driven structural interdependences to make the prediction error coherent with the probability distribution, which is desirable for successful transform and coding. The structured set prediction model incorporates a max-margin Markov network (M3N) to regulate and optimize multiple block predictions. The model parameters are learned by discriminating the actual pixel value from other possible estimates to maximize the margin (i.e., decision boundary bandwidth). Compared to existing methods that focus on minimizing prediction error, the M3N-based model adaptively maintains the coherence for a set of predictions. Specifically, the proposed model concurrently optimizes a set of predictions by associating the loss for individual blocks to the joint distribution of succeeding discrete cosine transform coefficients. When the sample size grows, the prediction error is asymptotically upper bounded by the training error under the decomposable loss function. As an internal step, we optimize the underlying Markov network structure to find states that achieve the maximal energy using expectation propagation. For validation, we integrate the proposed model into HEVC for optimal mode selection on rate-distortion optimization. The proposed prediction model obtains up to 2.85% bit rate reduction and achieves better visual quality in comparison to the HEVC intra coding. Wenrui Dai, Hongkai Xiong, Xiaoqian Jiang, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | Sparse Spatio-Temporal Representation With Adaptive Regularized Dictionary Learning for Low Bit-Rate Video CodingabstractFor promising vision-based video coding on low-quality data, this paper proposes a sparse spatio-temporal representation with adaptive regularized dictionary learning and develops a low bit-rate video coding scheme. In a reversed-complexity Wyner-Ziv coding manner, it selects a subset of key frames to code at original resolution, while the rest are down sampled and reconstructed by a sparse spatio-temporal approximation using key frames as a training dataset. Since primitive patches (geometry) are of low dimensionality and can be well learned from the primitive patches across frames in a scale space, a video frame is divided into three layers: a primitive layer, a nonprimitive coarse layer, and a nonprimitive smooth layer. The multiscale differential feature representations are invertible to facilitate reconstruction with dictionary learning, and the target is formulated as an optimization problem by constructing a sparse representation of 2-D patches and 3-D volumes over adaptive regularized dictionaries, a set of 2-D subdictionary pairs trained from primitive patches, and a 3-D dictionary trained from nonprimitive volumes. Specifically, the nonprimitive layer is constructed as volumes in to order keep it consistent along the motion trajectory, which enables sparse representations over a learned 3-D spatio-temporal dictionary. Through hierarchical bidirectional motion estimation and adaptive overlapped block motion compensation, the 3-D low-frequency and high-frequency dictionary pair is designed by the K-SVD algorithm to update the atoms for optimal sparse representation and convergence. In reconstruction, the lost high-frequency information of the down-sampled frames can be synthesized from the sparse spatio-temporal representation over the adaptive regularized dictionaries. Extensive experiments validate the compression efficiency of the proposed scheme versus H.264/AVC in terms of both objective and subjective comparisons. Hongkai Xiong, Zhiming Pan, Xinwei Ye, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | Multi-Level Video Frame Interpolation: Exploiting the Interaction Among Different LevelsabstractThis paper proposes a novel multi-level frame interpolation scheme by exploiting the interactions among different levels. The proposed scheme includes three major stages that work at block level, pixel level, and sequence level, respectively. Effective algorithms are designed for each stage, i.e., block-level motion estimation with dropping unreliable motion vectors, pixel-level motion vector-guided partial scale-invariant feature transform flow matching, and sequence-level 3-D total variation regularized completion. Compared to traditional methods that focus mostly at one single level, the proposed scheme manages to recognize and utilize the interactions among the three levels based on their distinct characteristics and intertwined relationships. With a proper exploitation of interactions, unique advantages for each level can be effectively preserved while inherent limitations of a given level can be overcome by utilizing information from other levels. Extensive experiments have confirmed its superior performance over several classical schemes, in both subjective visual quality and objective peak signal-to-noise ratio/structure similarity measurements, and typical artifacts can be significantly reduced. Zhefei Yu, Houqiang Li, Zhangyang Wang, Zeng Hu, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2013 | Guest Editorial - Special section on cloud-based mobile media: Infrastructure, services, and applicationsabstractIt is the aim of this special section to report on the latest research that explores various aspects of cloud-based mobile media. Through an open call for papers, we received 40 submissions. Fourteen (14) papers were accepted for final publication after two rounds of highly competitive reviews. The final papers were selected on the basis of originality and significance of the technical work, as well as relevance to the theme topic. The papers in this special section span a wide range of novel algorithms, techniques, applications, and system-level solutions. They naturally fall into two categories: addressing existing challenges or exploring emerging opportunities. Chang Wen Chen, Yonggang Wen 0001, Joel J. P. C. Rodrigues |
IEEE Trans. Multim. | 1 |
| 2013 | Compressive Coded Modulation for Seamless Rate AdaptationabstractThis paper presents a novel compressive coded modulation (CCM) which simultaneously achieves joint source-channel coding and seamless rate adaptation. The embedding of source compression into modulation brings significant throughput gain when the physical layer data contain non-negligible redundancy. The kernel of CCM is a new random projection (RP) code inspired by the compressive sensing (CS) theory. The RP code generates multilevel symbols from source binaries through weighted sum operations. Then, the generated RP symbols are mapped into a dense constellation for transmission. The receiver performs joint decoding based on received symbols. As the number of RP symbols can be adjusted in fine granularity, the rate adaptation becomes seamless. Two key design issues in the proposed CCM are addressed in this paper. First, we consider the RP code design for sources with different redundancies. Three principles are established and a concrete implementation is given. Second, we devise a linear-time decoding algorithm for the proposed RP code. In this belief propagation (BP) algorithm, we find that computing convolution in time domain is more efficient than that in frequency domain for binary variable nodes. Moreover, we invent a ZigZag deconvolution to further reduce the complexity. Analysis show that the proposed decoding algorithm is nearly 20 times faster than the state-of-the-art BP algorithm for CS called CS-BP. Emulations on traced data show that CCM achieves significant throughput gain, up to 33% and 70%, respectively, over the Hybrid ARQ with compression and BICM with compression, under practical time-varying wireless channels. Hao Cui 0001, Chong Luo 0001, Jun Wu 0006, Chang Wen Chen, Feng Wu 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2012 | JSM-2 based joint ECG compressed sensing with partially known support establishmentabstractCompressed sensing (CS) is a technique that enables sparse signal reconstruction from much fewer samples. In this paper, we propose ECG compressed sensing methods based on distributed compressed sensing to exploit the joint sparsity for both single- and multi-lead ECG signals. We apply JSM-2 (joint sparse model type 2) for jointly sparse ECG signals and formulate how to establish a partially known support based on this type of sparse model. Through careful analysis of joint partially known support, two-step ECG signal reconstruction schemes for single-lead and multi-lead ECG signals are developed. Simulation results show that the proposed schemes based on partially known support establishment outperforms existing schemes with enhanced performance measured by percentage root mean square difference (PRD). Bin Liu 0016, Chang Wen Chen |
Healthcom | 3 |
| 2012 | QoS-driven scheduling approach using optimal slot allocation for Wireless Body Area NetworksabstractWireless Body Area Network (WBAN) is a promising type of networks that mainly targets at applications in ubiquitous communication and e-Health services. Different from other types of networks, one important challenge for WBAN is that its quality of service (QoS) requirement, in terms of delivery probability and data rate, will be time varying since human body is a highly dynamic physical environment. Another significant challenge for WBAN is that energy efficiency needs to be guaranteed in such a resource-limited network. In this paper, a QoS-driven scheduling approach is proposed to address these challenges. We model the WBAN channel as a Markov model as suggested by the emerging IEEE 802.15.6 BAN standard and propose a threshold-based scheme to adjust the transmission order of nodes. The number of slots for each node is optimally assigned according to the QoS requirement while minimizing the energy consumption of nodes. The results from extensive simulations show that the proposed approach can provide high QoS and energy efficiency under different network conditions, especially in highly heterogeneous ones in WBAN. Zhisheng Yan, Bin Liu 0016, Chang Wen Chen |
Healthcom | 3 |
| 2012 | Receiver rate adaptation for MIMO systemabstractThis paper presents a novel design of linear random space time coding (LRSTC) scheme. This scheme is able to achieve the dual benefits for MIMO receiver rate adaptation: (1) rateless coding and modulation and (2) approximately universal tradeoff between diversity gain and multiplexing gain. In this scheme, an LRSTC symbol is the superposition of a block of randomly rotated QPSK symbols. Such superposition generates a dense constellation that ensures a high saturation rate. The random rotation guarantees the pseudo-orthogonality between QPSK symbols. Through utilizing the orthogonality of symbols, a power gain can be generated at receiver end. The power gain can be increased by transmitting more LRSTC symbols to the receiver in order to resist stronger noise of the MIMO system. Therefore, the rate can be seamlessly adapted through varying the number of transmitted symbols rather than the adopting conventional strategies in terms of modulation, coding and antenna selection. Moreover, the transmission of LRSTC symbols will span over all of the transmit antennas which allows the spatial diversity to be fully exploited. When the rank of MIMO matrix is greater than 1, more linear independent symbols can be received in each time slot. Hence, multiplexing gain can also be achieved. We have carried out both analysis and simulations to verify this LRSTC design. The simulation results show that LRSTC not only achieves the throughput comparable with the ideal conventional rate adaptation, but also accomplishes a seamless switching between diversity and multiplexing gain. Hao Cui 0001, Chang Wen Chen |
ICC | 2 |
| 2012 | Joint transceiver beamforming design and power allocation for multiuser MIMO systemsabstractMulti-user multiple input multiple output (MU-MIMO) system is an excellent choice for the next generation broadband communications because of its great potential in enhancing MIMO system capacity. One major challenge of such system is the jointly optimal transceiver beamforming design that maximizes the sum capacity under a total power constraint. Existing approaches have provided convex optimization solution by exploiting the duality between transmitter and receiver. However, with high computation complexity, these approaches are ineligible of practical implementation. In this paper, we aim to provide complexity efficient transceiver beamformer and power allocation design in MU-MIMO downlink (broadcast) channels. We first develop an iterative sub-optimal algorithm based on cyclic self-SINR-maximization (CSSM) beamforming design and water-filling power allocation. Comparing to convex-optimization-based (COB) approaches, the proposed solution (i.e. CSSM-WF algorithm) have insignificant sum-rate degradation but very low computational complexity and extremely fast converging speed. Trying to improve the performance of CSSM-WF algorithm, we then introduce another algorithm denoted as efficient COB (ECOB) algorithm in which CSSM-WF algorithm is used to generate a good start point for optimal COB approaches. Simulation results prove the efficiency of both proposed algorithms. Qian Liu 0001, Chang Wen Chen |
ICC | 2 |
| 2012 | A fast H.264 compressed domain watermarking scheme with minimized propagation error based on macro-block dependency analysisabstractEmbedding watermark into H.264 bit-streams directly can simplify encoding procedure and reduce computational cost comparing with conventional encoder based watermarking algorithms. However, possible error propagation may occur due to the prediction mechanism of the H.264 codec and cause video quality degradation. Expensive computation is often needed to compensate such drift. Furthermore, compensation in compressed domain usually leads to significant increase in the size of the watermarked video. To address these problems, we have developed a novel paradigm that can minimize the error propagation due to the watermarking process with very low complexity, high fidelity, robustness and security. This new paradigm is based on careful analysis of the dependency relation among the macro-blocks. We first extract the macro-block dependencies from the H.264 prediction mechanism and cache them as a graph. We then select these macro-blocks that have the least impact to the overall sequence to embed the desired watermark. This new way of embedding shall intrinsically minimize the error propagation without post processing of drift compensation. Extensive simulations have been carried out to confirm that the proposed watermark embedding is indeed able to achieve the desired performance with low complexity, high fidelity, robustness and security. Jingteng Xue, Chang Wen Chen, Hong Heather Yu |
ICIP | 3 |
| 2012 | Virtual View Reconstruction Using Temporal InformationabstractThe most significant problem in generating virtual views from a limited number of video camera views is handling areas that have become dis-occluded by shifting the virtual view away from the camera view. We propose using temporal information to address this problem, based on the notion that dis-occluded areas may have been seen by some camera in some previous frames. We formulate the problem as one of estimating the underlying state of the object in a stochastic dynamical system, given a sequence of observations. We apply the formulation to improving the visual quality of virtual views generated from a single “color plus depth” camera, and show that our algorithm achieves better results than depth image based rendering using standard inpainting. Shujie Liu 0001, Philip A. Chou, Cha Zhang, Zhengyou Zhang, Chang Wen Chen |
ICME | 5 |
| 2012 | QoS-driven and Fair Downlink Scheduling for Video Streaming over LTE Networks with Deadline and Hard Hand-offabstractLong-term evolution (LTE) represents a promising technique for ubiquitous multimedia communication because of its significant enhancement to the data transmission rate. However, the hard hand-off (HO) procedure standardized in LTE is a menace to multimedia quality of service (QoS), due to the service interruption introduced by the procedure. Thus, one major challenge of such a system is to design an effective scheduler that can guarantee quality-of-service (QoS) to users under various mobility scenarios, including the hard HO procedure. Existing downlink scheduling approaches do not consider the characteristics of the video sources together with the service degradation evoked by hard HO, and therefore are unable to meet the QoS requirements of multimedia consumers. In this paper, we develop a QoS-driven downlink scheduling scheme for video streaming that considers both QoS metrics of video packet deadline and hard HO service degradation, in order to guarantee the QoS requirements of multimedia consumers. It will be demonstrated that the proposed scheduler is not only able to achieve satisfactory QoS for users during HO period, but also offers fairness for both multimedia traffic and regular data traffic. Simulation results confirm the efficiency of the proposed scheme. Qian Liu 0001, Chang Wen Chen |
ICME | 3 |
| 2012 | Crowdsourced Learning to Photograph via Mobile DevicesabstractCapturing a professional photo with high visual quality is always a challenging task for mobile users. This paper presents a crowd sourced learning to photograph approach to assist mobile users for composing high quality photos via their mobile devices. The proposed approach is able to leverage the camera and scene context to search related images with similar context and content from social media communities, and then mine composition knowledge to guide photographing on mobile devices. We develop a patch-based feature generation and selection process to discover salient patches and positions that dominate photo composition aesthetics in the input scene. We then build a regression model to map the composition of salient patches to photo-aesthetic scores. Finally, we develop an efficient hierarchical approach to search for the optimal view enclosure for photograph suggestion. We conducted extensive simulations and subjective evaluations to verify the proposed approach. Wenyuan Yin, Tao Mei 0001, Chang Wen Chen |
ICME | 3 |
| 2012 | Texture-assisted Kinect depth inpaintingabstractThe emergence of Kinect facilitates the possibility of depth capture in real-time and with low cost by consumers. It also provides powerful tool and inspiration for researchers to engage in new array of technology development. However, the quality of the depth map captured from Kinect is still inadequate for many applications due to holes, noises and artifacts existing within the depth information. In this paper, we present a texture assisted Kinect depth inpainting framework, aiming at obtaining improved depth information. In this framework, the relationship between texture and depth is investigated, and the characteristics of depth are also exploited. More specifically, texture edge information is extracted to assist the depth inpainting. Furthermore, filtering and diffusion are designed for hole-filling and edge alignment. Experiment results demonstrate that the Kinect depth can be appropriately repaired in both smooth and edge region. Comparing with the original depth, the inpainted depth information enhances the quality of advanced processing such as 3D reconstruction. Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen |
ISCAS | 5 |
| 2012 | A novel Slepian-Wolf decoding algorithm exploiting geometric regularity constraints with anisotropic MRF modelingabstractThis paper proposes a novel Slepian-Wolf decoding algorithm for distributed video coding by exploiting not only the statistical correlation between the side-information and source but also the spatio-temporal consistency constraint of video sequences. The proposed algorithm models the log-likelihood-ratio (LLR) information for Slepian-Wolf decoding as an anisotropic MRF model and solving the inference by iteratively performing conventional probabilistic Slepian-Wolf decoding, which imposes the global bit-wise constraint from the Slepian-Wolf encoding process, and MRF optimization with belief propagation (BP) to enforce the local geometric regularity constraint of video frames. Experimental results demonstrate a considerable performance gain beyond existing Slepian-Wolf decoding algorithms in literature. Hongkai Xiong, Chang Wen Chen |
ISCAS | 3 |
| 2012 | Mobile JND: environment adapted perceptual model and mobile video quality enhancementabstractDesign of a quality-of-experience (QoE) optimized mobile video system should consider not only the video content and display specifications but also the fact that mobile devices are exposed to many different environments and viewing scenarios. For same device and same content, the viewer will perceive different visual qualities when the viewing environment changes. Current perceptual quality estimation approaches including the extensively adopted just noticeable distortion (JND) based models neglect significant influence of surroundings on perception. However, the environmental effects on perception have long been supported by psychophysical experiments. This paper proposes a novel viewing scenario adapted model that exploits the influence of various viewing conditions including display size, viewing distance, ambient luminance and body movement and apply the proposed model to the H.264 video encoding. With the help of multiple sensors widely equipped on handholds today, the mobile device is able to dynamically estimate the surrounding conditions. The estimated environment parameters are feedback to video encoder to generate encoded video source that best matches to the current scenario so as to improve the bandwidth efficiency and enhance visual quality for that particular environment. Our subjective experiments demonstrate a significant 30% saving on bit-rates without perceivable quality loss, or obvious improvements in visual qualities under same bandwidth constraint. Jingteng Xue, Chang Wen Chen |
MMSys | 2 |
| 2012 | Layered compression for high dynamic range depthabstractWith the rapid development of depth data acquisition technology, the high precision depth becomes much easier to access in real-time by depth sensors, and the generated high dynamic range (HDR) depth is widely adopted to benefit the depth assistant applications. Accordingly, the HDR depth compression becomes essential for the efficient depth storage and transmission. In this paper, we introduce a layered compression framework for HDR depth to achieve efficient and low-complexity depth compression. To leverage the state-of-art 8-bit image/video encoders, the HDR depth is partitioned into two layers: most significant bit (MSB) layer and least significant bit (LSB) layer. For MSB layer, an error controllable pixel domain encoding scheme is proposed to guarantee the compatibility for existing 8-bit codec by controlling quantization errors added back to LSB layer. Meanwhile, the efficient major color extraction and adaptive quantization enhance the coding performance of MSB layer. For LSB layer, the layer data with limited dynamic range is compressed by normal 8-bit image/video based encoding scheme. The experimental results demonstrate that our coding scheme can achieve real-time depth compression with the satisfactory reconstruction quality. The encoding time is less than 31ms/frame and the decoding time is around 20ms/frame in average. Our compression scheme can be easily integrated into the real-time depth transmission system. Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen |
VCIP | 5 |
| 2012 | Assessing photo quality with geo-context and crowdsourced photosabstractAutomatic photo quality assessment emerged as a hot topic in recent years for its potential in numerous applications. Most existing approaches to photo quality assessment have predominantly focused on image content itself, while ignoring various contexts such as the associated geo-location and timestamp. However, such a universal aesthetic assessment model may not work well with significantly different contexts, since the photography rules are always scene and context dependent. In real cases, professional photographers use different photography knowledge when shooting various scenes in different places. Motivated by this observation, we leverage the geo-context information associated with photos for visual quality assessment. Specifically, we propose in this paper a Scene-Dependent Aesthetic Model (SDAM) to assess photo quality, by jointly leveraging the geo-context and visual content. Geo-contextual leveraged searching is performed to obtain relevant images with similar content to discover the scene-dependent photography principles for accurate photo quality assessment. To overcome the problem that in many cases the number of the contextually searched images is insufficient for learning the SDAM, we adopt transfer learning to utilize auxiliary photos within the same scene category from other locations for learning photography rules. Extensive experiments shows that the proposed SDAM scheme indeed improves the photo quality assessment accuracy via leveraging photo geo-contexts, compared with traditional universal aesthetic models. Wenyuan Yin, Tao Mei 0001, Chang Wen Chen |
VCIP | 3 |
| 2012 | Guest Editorial QoE-Aware Wireless Multimedia SystemsabstractThe 11 papers in this special issue cover a range of topics and can be logically organized in three groups, focusing on QoE-aware media protection, QoE assessment and modelling, and multi-user-QoE management. Maria G. Martini, Chang Wen Chen, Zhibo Chen 0001, Tasos Dagiuklas, Lingfen Sun |
IEEE J. Sel. Areas Commun. | 2 |
| 2012 | Distributed Robust Optimization for Scalable Video Multirate Multicast Over Wireless NetworksabstractThis paper proposes a distributed robust optimization scheme to jointly optimize overall video quality and traffic performance for scalable video multirate multicast over practical wireless networks. In order to guarantee layered utility maximization, the initial nominal joint source and network optimization is defined, where each scalable layer is tailored in an incremental order and finds jointly optimal multicast paths and associated rates with network coding. To enhance the robustness of the nominal convex optimization formulation with nonlinear constraints, we reserve partial bandwidth for backup paths disjoint from the primal paths. It considers the path-overlapping allocation of backup paths for different receivers to take advantage of network coding, and takes into account the robust multipath rate-control and bandwidth reservation problem for scalable video multicast streaming when possible link failures of primary paths exist. Specifically, an uncertainty set of the wireless medium capacity is introduced to represent the uncertain and time-varying property of parameters related to the wireless channel. The targeted uncertainty in the robust optimization problem is studied in a form of protection functions with nonlinear constraints, to analyze the tradeoff between robustness and distributedness. Using the dual decomposition and primal-dual update approach, we develop a fully decentralized algorithm with regard to communication overhead. Through extensive experimental results under critical performance factors, the proposed algorithm could converge to the optimal steady-state more quickly, and adapt the dynamic network changes in an optimal tradeoff between optimization performance and robustness than existing optimization schemes. Hongkai Xiong, Junni Zou, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | A novel 3D video transcoding scheme for adaptive 3D video transmission to heterogeneous terminalsabstractThree-dimensional video (3DV) is attracting many interests with its enhanced viewing experience and more user driven features. 3DV has several unique characteristics different from 2D video: (1) It has a much larger amount of data captured and compressed, and corresponding video compression techniques can be much more complicated in order to explore data redundancy. This will lead to more constraints on users' network access and computational capability, (2) Most users only need part of the 3DV data at any given time, while the users' requirements exhibit large diversity, (3) Only a limited number of views are captured and transmitted for 3DV. View rendering is thus necessary to generate virtual views based on the received 3DV data. However, many terminal devices do not have the functionality to generate virtual views. To enable 3DV experience for the majority of users with limited capabilities, adaptive 3DV transmission is necessary to extract/generate the required data content and represent it with supported formats and bitrates for heterogeneous terminal devices. 3DV transcoding is an emerging and effective technique to achieve desired adaptive 3DV transmission. In this article, we propose the first efficient 3DV transcoding scheme that can obtain any desired view, either an encoded one or a virtual one, and compress it with more universal H.264/AVC. The key idea of the proposed scheme is to appropriately utilize motion information contained in the bitstream to generate candidate motion information. Original information of both the desired view and reference views are used to obtain this candidate information and a proper motion refinement process is carried out for certain blocks. Simulation results show that, compared to the straightforward cascade algorithm, the proposed scheme is able to output compressed bitstream of the required view with significantly reduced complexity while incurring negligible performance loss. Such a 3DV transcoding can be applied to most gateways that usually have constraints on computational complexity and time delay. Shujie Liu 0001, Chang Wen Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2012 | A joint layered scheme for reliable and secure mobile JPEG-2000 streamingabstractThis article presents a novel joint layered approach to simultaneously achieve both reliable and secure mobile JPEG-2000 image streaming. With a priori knowledge of JPEG-2000 source coding and channel coding, the proposed joint system integrates authentication into the media error protection components to ensure that every source-decodable media unit is authenticated. By such a dedicated design, the proposed scheme protects both compressed JPEG-2000 codestream and the authentication data from wireless channel impairments. It is fundamentally different from many existing systems that consider the problem of media authentication separately from the other operations in the media transmission system. By utilizing the contextual relationship, such as coding dependency and content importance between media slices for authentication hash appending, the proposed scheme generates an extremely low authentication overhead. Under this joint layered coding framework, an optimal rate allocation algorithm for source coding, channel coding, and media authentication is developed to guarantee end-to-end media quality. Experiment results on JPEG-2000 images validate the proposed scheme and demonstrate that the performance of the proposed scheme is approaching its upper bound, in which case no authentication is applied to the media stream. Xinglei Zhu, Chang Wen Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2011 | Wireless Video Streaming QoS Guarantees Based on Virtual Leaky BucketabstractWith recent rapid advances in wireless communication and networking, the demand for wireless video streaming is exploding. Unlike data streaming, video streaming is much more challenging due to unique characteristics such as bursty flow, stringent delay bound, and selective loss tolerance. Under the dynamic nature of wireless channels, wireless video streaming is virtually impossible to achieve deterministic quality-of-service (QoS) guarantee. Recent development in statistical QoS guarantees has shown great potential for wireless video streaming. However, in the case of statistical QoS guarantee, loss of data due to queue length bound violation or delay violation is often inevitable. Because of unequal importance of video data stream and their inherent dependency, different loss patterns will result in distinctive loss impact. It is critical to design an appropriate scheme to control the loss of video data in such a way that it will minimize the impact of such loss. We develop in this research a prioritized packet dropping scheme based on virtual leaky bucket to control the delivery of video streams. We have carried out analysis based on effective capacity, effective bandwidth, and virtual leaky bucket to demonstrate the potential of the proposed scheme. Preliminary results obtained from simulations show that an improved performance in QoS guarantees can be achieved for wireless video streaming. Zhengyong Feng, Guangjun Wen, Chang Wen Chen |
GLOBECOM | 4 |
| 2011 | Packet Loss Analysis for Media Streaming with Network Coding in Wireless Broadcast NetworksabstractFor media streaming, throughput are inadequate to fully determine media playback quality at receivers. Packet loss behavior is also crucial in designing a high performance media streaming system. In this paper, we carry out analysis on packet loss behavior when network coding is performed at the Base Station (BS) with limited buffer in a wireless broadcast network. Two types of packet loss at the BS are considered, namely packet loss at the input introduced by buffer overflow (Lin), and packet loss at the output due to channel errors and late arrivals (Lout). Our analysis is based on a bulk-service queue model denoted as M/G(a, b)/l/B. The major contribution of this research is the successful derivation of the probability distribution of queue states at an arbitrary time by applying the supplementary variable technique. Given particular channel conditions, this probability distribution enables us to obtain the relationship between two types of packet loss (Lin,Lout) and the supportable arrival rates at the BS. Then, the maximum allowable arrival rate when Linand Loutare constrained can be determined. This is critical for the design of appropriate congestion control for media streaming at the BS. Extensive simulations have been performed to demonstrate that the results from our analysis are consistent with that from simulations. Hancheng Lu, Chang Wen Chen |
GLOBECOM | 2 |
| 2011 | Dynamic Adaptive Streaming over HTTP from Multiple Content Distribution ServersabstractDynamic adaptive streaming over HTTP (DASH\footnote{In this paper, we use DASH as a general acronym, which does not represent 3GPP- DASH/MPEG-DASH standards.}) is attractive. It reuses popular web servers rather than relying on expensive streaming server. It takes the benefits of router and firewall optimizations for HTTP traffic. It promises to automatically adapt to bandwidth dynamics to provide smooth playback experience with best achievable visual quality. In this paper, we study a new architecture for DASH, i.e. DASH from multiple content distribution servers with scalable coded video. We take the advantage of content distribution network (CDN) architecture to offer even better DASH experience. In our scheme, users request one video program simultaneously from more than one CDN servers over HTTP. To optimize resource allocation among different servers, we propose a collaborative multi-scale scheduling algorithm (CMSS) which can effectively: 1) mitigate web server load and automatically balance loads among servers; 2) adapt to bandwidth dynamics of each server; 3) optimize aggregated streaming quality. We implemented a prototype of our algorithm and deployed it on PlanetLab open platform. Measurement results verify the effectiveness of CMSS. Chang Wen Chen |
GLOBECOM | 3 |
| 2011 | Joint source-channel rate control for pixel-domain distributed video codingabstractWe study the scenario of pixel-domain distributed video coding for noisy transmission environments and propose a method to allocate the available rate between source coding and channel coding to generate a robust video stream. Having observed in experiments the uncertainty of the source and the channel coding rate, we model them as random variables via offline training, estimate the decoding failure probability and calculate the mean end-to-end distortion. Adaptive quantization is performed for each slice to minimize its mean end-to-end distortion. With this joint source-channel rate allocation, we compare the robustness of two coding prototypes, namely distributed video coding and distributed video coding with forward error correction. According to our experimental results, under same total bit budget, the distributed video coding only scheme proves more robust than the latter one and the gain is up to 1 dB in PSNR. Eckehard G. Steinbach, Chang Wen Chen |
ICASSP | 3 |
| 2011 | A comparison of the error resiliency of bit-plane based and symbol based pixel-domain distributed video codingabstractThis work studies the error resilience of pixel-domain distributed video coding in noisy wireless transmission environment. Turbo codes are used to implement the DVC coder and the AWGN model is assumed for the transmission channel. The goal is to find out whether symbol based coding or bit-plane based coding is more robust against channel noise. We compare the two in the context of joint source-channel coding to ensure a high error resilience. First, we propose a framework to estimate the end-to-end distortion for these two schemes. Next, we allocate the rate between source coding and channel coding, aiming at a minimum end-to-end distortion. Then, we simulate the two schemes, setting the same bit budget for them, and compare their error resilience. Experimental results show that symbol based coding outperforms bit-plane based coding by up to 0.7 dB in PSNR of the decoded video. Eckehard G. Steinbach, Chang Wen Chen |
ICIP | 3 |
| 2011 | Fairness and QOS guaranteed user scheduling for multi-user MIMO broadcasting channelabstractThis paper investigates fairness and quality of service (QoS) guaranteed user scheduling scheme in multi-user MIMO broadcast channel. The main contribution of this paper is to consider characteristic of video source in user scheduling process. As a result, the proposed scheme can provide fair and satisfactory services to regular data users, while QoS of multimedia consumers is guaranteed. Since video source is delay-sensitive, in our framework, first we will estimate the transmitting time of each video packet. The estimation of transmission time serves dual purpose: one is to guarantee the transmission of any given video packet before it expires, and another one is to allow other unscheduled regular data users to access the channel when there is no pressing need to transmit any video packet for the video consumer. Simulation results show that the proposed user scheduling framework achieves high QoS performance for multimedia consumers. Qian Liu 0001, Chang Wen Chen |
ICIP | 2 |
| 2011 | New TCP video streaming proxy design for last-hop wireless networksabstractTCP based HTTP video streaming is becoming more and more popular in recent years. However, for last-hop wireless access, due to TCP's additive increase/multiplicative decrease (AIMD) congest control mechanism, when interference and fading are strong, TCP streaming behaves poorly. In this research, we propose a new design of TCP video streaming proxy at the wireless base station. The new proxy's mission is two folds. First, it transparently isolates the wireless network from the wired network. Specifically, for the last-hop wireless transmission, it adopts a new variant of TCP, we call it raw-TCP, which in combination with a fair scheduler component in the proxy, results in an increased average throughput with low dynamics. Second, through video adaptation, i.e. priority guided packet reordering and deadline driven adaptive truncation, the proxy eliminates playback jitter while maintaining small visual quality variation between successive frames. These two functions fundamentally differentiate our research from the vanilla split-TCP schemes. In our research, the TCP proxy is a multimedia aware element while in traditional split-TCP, the relay node is multimedia oblivious. Simulation results using ns-2 verify the effectiveness of the proposed solution. Chang Wen Chen |
ICIP | 3 |
| 2011 | A collusion resilient key management scheme for multi-dimensional scalable media access controlabstractThis paper proposes a novel key management scheme for multi-dimensional scalable multimedia access control. We build up a collusion-attack model and prove that the proposed scheme is indeed resilient to collusion attack. The lower bound of number of the segment keys is also given in the analysis. We consider the key management problem on partially order set and propose a general approach consists of a poset decomposition step and an element projection step. Under the proposed framework, a novel hierarchical 3D poset decomposition approach is further developed. The proposed scheme is the first provable collusion resilient scheme in 3-dimension (and higher) scenario. It can be easily adapted to other scalable dependence structure and higher dimension. Analysis on the proposed scheme proves its collusion-resilient property. Xinglei Zhu, Chang Wen Chen |
ICIP | 2 |
| 2011 | Towards viewing quality optimized video adaptationabstractFor video delivery over heterogeneous networks, the bandwidth mismatch at the receiver edge often leads to inevitable degradation of video quality during adaptation. While most studies in this field are based on maximizing PSNR (peak-signal-to-noise-ratio), many psychophysical experiments discover that viewing quality is not correlated well with PSNR but significantly influenced by the viewing conditions (e.g. display size and viewing distance), which are determined by the receiver device. In this research, we show that major picture quality evaluation schemes are not suitable for subjective quality driven video adaptation because they are unable to track the correspondence between viewing condition and viewing quality. In this paper we investigate a novel paradigm that performs adaptation with respect to the target display scenario in order to maximize the viewing quality. We demonstrate that viewing ratio (VR) plays a key role in linking viewing conditions and perceptual quality of given video signals. We convert the linkage between VR and perceptual quality into a computationally feasible algorithm for video adaptation. Experimental results show the VR based video adaptation scheme not only outperforms the benchmarks in terms of quality-of-experience for designated viewing scenarios, but also provides generally an improved visual quality for mobile devices with bandwidth constraint. Jingteng Xue, Chang Wen Chen |
ICME | 2 |
| 2011 | Towards maximal decodable rate for multi-rate multicast of digital media with network codingabstractMulti-rate multicasting has become more and more attractive in contemporary multimedia applications because of its efficiency in serving heterogeneous receivers with different rates commensurate with their capabilities. When network coding is applied to multi-rate media multicasting to achieve additional gain, we encounter several significant challenges in designing a scheme that attains maximum throughput for all heterogeneous users. In this paper we focus on the problem of designing optimal network coding based approach to maximize total received utility for multi-rate multicasting media encoded in a layered structure. Such a layered structure facilitates multi-rate media delivery that matches users' reception capabilities but creates certain undesired dependency between different layers. We propose a request-assign mechanism to enable sufficient information propagation between source node and receiver node before multicasting the media content. With the help of request messages, the proposed scheme can overcome challenges of the layered dependency and is able to determine maximal decodable layer for each individual receiver. Furthermore, the proposed scheme achieves maximal decodable rate with polynomial complexity. Experiment results on JPEG-2000 images verify the proposed scheme. Xinglei Zhu, Chang Wen Chen |
ICME | 2 |
| 2011 | A deadline-aware virtual contention free EDCA scheme for H.264 video over IEEE 802.11e wireless networksabstractContention-based access has been the key component for 802.11 wireless networks. As the video service is becoming more popular, contention-based access has been the key limiting factor in providing satisfactory quality-of-service (QoS) for video over wireless networks, especially when the data traffic includes both real-time video and other types of traffics. A contention free burst (CFB) scheme has been developed recently [1], would result in improved video QoS, but at the expenses of best effort and background traffics. In this paper, we proposed a deadline-aware virtual contention free (DVCF) scheme for H.264 video over 802.1 le wireless networks. The main contribution of this paper is to jointly consider the characteristic of video source, 802.lie protocol, and network specifications by adjusting the parameters of 802.lie protocol. As a result, the proposed scheme (DVCF EDCA) maintains the transmission of all types of traffics provided that the deadline of transported video packets can be strictly observed. In this framework, the deadline of each video packet is estimated first as in [2]. The estimated deadline serves dual purpose: one is to guarantee the transmission of any given video packet before it expires and another one is to allow other types of data traffic to transmit when there is no pressing needs to transmit any video packet. Simulation results show that the proposed framework achieves virtually no loss for video packets and noticeable improvement in combined throughput of video and non-video traffics. Qian Liu 0001, Chang Wen Chen |
ISCAS | 3 |
| 2011 | Resource allocation for cloud-based free viewpoint video rendering for mobile phonesabstractFree viewpoint video (FVV) / Free viewpoint TV (FTV) on mobile devices over cellular networks is very challenging due to the requirement for large bandwidth and limitations in computation and battery life on mobile phones. To address such challenges, in this paper we propose a cloud-based FVV / FTV rendering framework for mobile devices over cellular networks. In this framework, cloud performs rendering for mobile devices. In order to achieve maximum QoE (Quality of Experience) for mobile users, we propose a novel resource allocation scheme, which jointly considers rendering allocation between cloud and client based on user's QoE and rate allocation among texture, depth, and channel rate based on rate-distortion analysis. We formulate this resource allocation scheme as an optimization problem which can then be transformed into a convex optimization for the given rate ratio. Experimental results demonstrate that the proposed cloud-based FVV rendering solution can substantially improve video quality on mobile devices comparing with traditional approaches. Dan Miao, Wenwu Zhu 0001, Chong Luo 0001, Chang Wen Chen |
ACM Multimedia | 4 |
| 2011 | Seamless rate adaptation for wireless networkingabstractThis paper aims at designing a Seamless Rate Adaptation for wireless networking which achieves smooth rate adjustment in a broad dynamic range of channel conditions. Conventional rate adaptation can only achieve a stair-case rate adjustment. Even when combining with hybrid ARQ, it suffers from an irreconcilable conflict between throughput and dynamic range. We tackle this problem from a new perspective by relying on modulation, instead of channel coding, for rate adaptation. We propose rate compatible modulation (RCM), in which modulation signals are incrementally generated from information bits through weighted mapping. Rate adaptation is achieved through varying the number of modulated signals. As more signals are transmitted, information bits gradually accumulate energy. The weights in bit-to-symbol mapping are delicately designed to ensure fine-grained energy accumulation so that smoothness and efficiency can both be achieved. We design and implement a rate adaptation system, called SRA and evaluate its performance through a software radio testbed. Results show that, under highly dynamic channel conditions, SRA achieves over 80% throughput gain over 802.11a adaptive modulation and coding, and achieves 28.8% and 43.8% gain over HARQ systems implemented with Turbo code and Raptor code. We believe that the concept of rate compatible modulation opens up a fresh research avenue toward the wireless rate adaptation problem. Hao Cui 0001, Chong Luo 0001, Feng Wu 0001, Chang Wen Chen |
MSWiM | 5 |
| 2011 | Scalable video transmission: packet loss induced distortion modeling and estimationabstractTo provide enhanced multimedia services for heterogeneous networks and terminal devices, Scalable Video Coding (SVC) has been developed to embed different quality of video in a single bitstream. Similar to classical compressed video transmission, different packets of a video bitstream have different impacts on received video quality. Therefore, distortion modeling and estimation are necessary in designing a robust video transmission strategy under various network conditions. In the paper, we present the first scheme of packet loss induced distortion modeling and estimation in SVC transmission. The proposed scheme is applicable to numerous video communication and networking scenarios in which accurate distortion information can be utilized to enhance the performance of video transmission. One major challenge in scalable video distortion estimation is due to the adoption of more complicated prediction structure in SVC, which makes the tracking of error propagation much more difficult than the non-scalable encoded video. In this research, we tackle such challenge by systematically tracking the propagation of errors under various prediction trajectories. Supplemental information about the compressed video is embedded into data packets to substantially simplify the modeling and estimation. Moreover, with supplemental data of inter prediction information, distortion estimation can be processed without parsing video bitstream which results in much lower computation and memory cost. With negligible effects on the data size, experimental results show that the proposed scheme is able to track and estimate the distortion with very high accuracy. This first ever scalable video transmission distortion modeling and estimation scheme can be deployed at either gateways or receivers because of its low computation and memory cost. Shujie Liu 0001, Chang Wen Chen |
NOSSDAV | 2 |
| 2011 | MixCast modulation for layered video multicast over WLANsabstractThe major challenge in wireless multicast is the heterogeneous channel conditions of multiple users. In video multicast, the combination of a layered video coding scheme and a layered transmission scheme can gracefully accommodate user heterogeneity. This paper presents MixCast, a novel physical layer scheme for layered transmission. The key innovation in MixCast is the rateless Euclidean symbol mapping which mixes base layer and enhancement layer bits into arbitrary number of wireless symbols using arithmetic weighted sum operation. This design brings two benefits when compared with the state-of-the- art physical layer technique known as hierarchical modulation (HM). First, MixCast uses a fixed modulation constellation, avoiding the complexity in adaptive modulation and coding when channel condition varies. Second, the rateless symbol mapping allows MixCast to achieve much smoother rates than the stair-shaped rates in HM. We implemented MixCast for typical video multicast over OFDM physical layer, and evaluate its performance against HM through both simulations and software radio testbed. In simulations, MixCast shows consistent gain of 3dB to 5dB over HM under various rate combinations. In the testbed experiments, MixCast achieves significant gain up to 15dB in video PSNR over HM for football sequence. Hao Cui 0001, Chong Luo 0001, Chang Wen Chen, Feng Wu 0001 |
VCIP | 3 |
| 2011 | Wyner-Ziv video coding using progressive encoding and decodingabstractIn distributed single-view video coding (DSVC) or distributed multiview video coding (DMVC), the compression performance is greatly affected by the quality of side information (SI). Most existing DSVC or DMVC schemes generate SI by employing motion estimation (ME) at the decoder. Since the original frame is unavailable, the ME accuracy is low. In this paper we propose a novel distributed coding scheme for both single-view video and multi-view video. At the encoder, each Wyner-Ziv frame (WF) is sub-sampled using a block-based sub-sampling algorithm, and resulting sub-samples are successively encoded and transmitted. For the latter samples, adaptive intra prediction is employed with no intra mode bits transmitted. At the decoder, the current WF is progressively decoded, and thus the ME accuracy is iteratively improved. Experimental results demonstrate that the proposed DSVC and DMVC framework has achieved significant rate distortion (RD) gain. Houqiang Li, Chang Wen Chen |
VCIP | 4 |
| 2011 | Conditional random field based side-information fusion for distributed multi-view video codingabstractThis paper presents a new temporal and inter-view side-information fusion algorithm for distributed multi-view video coding (DMVC). Unlike existing fusion algorithms in DMVC schemes that produce the fusion mask by finding the motion vector outliers, it introduces conditional random fields (CRF) to exploit the intrinsic geometric regularity and temporal consistency constraint in multi-view video sequences. Specifically, Wyner-Ziv (WZ) frames are modeled by CRF with the temporal and the inter-view side-information as two observations. The observation distribution models the local accuracy of the temporal and the inter-view side-information. The transition distribution of the CRF model represents the local geometric regularity, e.g., the edge directions and the local smoothness of the WZ frame. Its parameters are trained from previously decoded WZ frames, and the inference is made on trained weights to generate fused side-information. The accurate modeling is validated to show a significant performance gain over the existing fusion algorithms by experiments. Hongkai Xiong, Hao Wang 0183, Chang Wen Chen |
VCIP | 4 |
| 2011 | Playback interruption probability analysis for Roadside-to-Vehicle media streamingabstractIn this paper, we propose a novel analytical framework for Roadside-to-Vehicle (R2V) media streaming to investigate the interruption probability when playback is performed at a vehicle. The results from such analysis can be used to design an R2V media streaming system with minimum interruption probability. We first introduce a general time-dependent queue model G(t)/G/1 to describe the playout buffer at the vehicle. This model characterizes the dynamic wireless channel conditions in a vehicular environment with time-varying packet arrival intervals. Based on the statistically monotonic behavior of the channel rates as the vehicle moves towards and away from a roadside Access Point (AP), we are able to derive the upper and lower bounds of the interruption probability through a time-dependent queuing analysis with diffusion approximation when the startup state is given. The parameters of startup state include startup delay and the number of cached packets in the playout buffer before playback. We divide the period during which the vehicle is guaranteed with low playback interruption probability into small time intervals. This way, the analysis of the G(t)/G/1 system can be decomposed into a series of transient analysis of stationary G/G/1 systems in a sequential fashion. For each stationary G/G/1 system, diffusion approximation with specified initial conditions is applied to obtain the transient queue length and the interruption probability of the playback. This scheme can be easily implemented since it only requires limited statistical information about fading channel and the playback statistics, such as mean and variance of the rate. The proposed analytical framework has been validated to show accurate characterization of the R2V media streaming system with extensive simulations. Hancheng Lu, Chang Wen Chen |
WOWMOM | 2 |
| 2011 | Joint Coding/Routing Optimization for Distributed Video Sources in Wireless Visual Sensor NetworksabstractThis paper studies a joint coding/routing optimization between network lifetime and video distortion by applying information theory to wireless visual sensor networks for correlated sources. Arbitrary coding [distributed video coding and network coding (NC)] from both combinatorial optimization and information theory could make significant progress toward the performance limit and tractable. Also, multipath routing can spread energy utilization across nodes within the entire network to keep a potentially longer lifetime, and solve the wireless contention issues by the splitting traffic. The objective function not only keeps the total energy consumption of encoding power, transmission power, and reception power minimized, but ensures the information received by sink nodes to approximately reconstruct the visual field. Also, a generalized power consumption model for distributed video sources is developed, in which the coding complexity of Key frames and Wyner-Ziv frames is measured by translating specific coding behavior into energy consumption. On the basis of the distributed multiview video coding and NC-based multipath routing, the balance problem between lifetime (costs) and distortion (capacity) is modeled as an optimization formulation with a fully distributed solution. Through a primal decomposition, a two-level optimization is relaxed with Lagrangian dualization and solved by the gradient algorithm. The low-level optimization problem is further decomposed into a secondary master dual problem with four cross-layer subproblems: a rate control problem, a channel contention problem, a distortion control problem, and an energy conservation problem. The implementation of the distributed algorithm is discussed with regard to the communication overhead and dynamic network change. Simulation results validate the convergence and performance of the proposed algorithm. Junni Zou, Hongkai Xiong, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2011 | Priority Belief Propagation-Based Inpainting Prediction With Tensor Voting Projected Structure in Video CompressionabstractThis paper presents a new video compression framework over H.264/AVC scheme that integrates our proposed structured priority belief propagation (BP)-based inpainting prediction (IP) to exploit the intrinsic nonlocal and geometric regularity in video samples. Unlike the existing edge-based inpainting adopted in lossy image coding, the optimal predictor could maintain the pixel-wise fidelity and the robust error resilience without any assistant information. Beyond the local prediction limitation of traditional intra and inter-modes, the priority BP with regularized structure priors of a spatio-temporal Markov random field is imposed on the predictor in an adaptive and more convergent sense. Specifically, the structured sparsity of the predicted macroblock region is inferred by tensor voting projected from the co-located decoded regions. In turn, the priority and visiting order of nodes are assigned according to the sets of updated beliefs as the propagation of messages. Through relatively few iterations of forward and backward process, the sparse inference of priority BP would ensure a stable marginal belief distribution on the structure and texture through updating local messages and beliefs. Within the optimal mode selection on rate-distortion optimization (RDO), the IP-mode with structured priority BP outperforms the existing vision-based approaches, and specially achieves a better objective rate-distortion performance besides visual quality. The IP-mode with structured priority BP can be applied to both I and P frames to generate low entropy residue, e.g., homogeneous visual patterns, and the computation complexity is also competitive with one iteration of sparse inference. Moreover, it behaves more resilient with an intrinsic probabilistic inference than the intra and inter-modes. Hongkai Xiong, Yang Xu 0048, Yuan F. Zheng, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2011 | An Error Resilient Video Coding Scheme Using Embedded Wyner-Ziv Description With Decoder Side Non-Stationary Distortion ModelingabstractIn this paper, we propose a generic error resilient video coding (ERVC) scheme using embedded Wyner-Ziv (WZ) description. At the encoder side, a joint source-channel R-D optimized mode selection (JSC-RDO-MS) algorithm with WZ-coded anchor frames is statistically studied and developed. Given a stationary first-order Markov Gaussian source, the proposed mode optimization is justified by an analysis of the RD impact on the WZ bit-rate. JSC-RDO-MS involves in the estimation of expected rate and distortion of WZ coding with the unavailable side information, and the WZ bit-rate of each coding mode is determined based on the error correction capability of the specific WZ codec. At the decoder side, an online correlation noise model between the source and the side-information is proposed with a mixture of Laplacians whose parameters are attained to reflect the coherence of the motion field of successive frames and the energy of prediction residual. Each mixture component represents the statistical distribution of prediction residuals, and the mixing coefficients represent the amount of errors in motion compensation. The proposed scheme achieves the so-called classification gain by exploiting the spatially non-stationary characteristics of the motion field and texture. Extensive experimental results show that the proposed WZ-ERVC scheme achieves a better overall RD performance than existing ERVC schemes, and the proposed modeling algorithm also significantly outperforms the conventional Laplacian model by up to 2 dB. Hongkai Xiong, Zhihai He, Songyu Yu, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2011 | Reconstruction for Distributed Video Coding: A Context-Adaptive Markov Random Field ApproachabstractWithin the existing reconstruction process of distributed video coding (DVC), there are two major approaches: the maximum probability reconstruction and the minimum mean square error (MMSE) reconstruction. Both of them assume that each node, a pixel in pixel domain DVC or a coefficient in transform domain DVC, is i.i.d., and reconstruct the value of each node independently by only exploiting statistical correlation between source and side-information. These kinds of models produce considerable amount of artifacts in decoded Wyner-Ziv (WZ) frames and degrade the objective performance. In this paper, we propose a context-adaptive Markov random field (MRF) reconstruction algorithm which exploits both the statistical correlation and the spatio-temporal consistency by modeling the corresponding MRF of a generic DVC architecture, and solve the inference by finding its MRF-based maximum a posteriori (MAP) estimate. The energy function of the MRF model consists of two terms: a data term measuring the statistical correlation, and a geometric regularity term enforcing local spatio-temporal structure consistency which is modeled by optical flow estimation with regard to the critical parameters under a wide variety of DVC scenarios. In case the unreliability of the derived local structure, a confidence parameter is introduced to prevent inappropriate penalizing. To find the reconstructed patch assignment with the largest expected probability in the context-adaptive MRF, the energy minimization for the MRF-based MAP estimate of the WZ frames is solved by global optimization and greedy strategies. Compared to the existing maximum probability and MMSE reconstruction with i.i.d. model, a better subjective and objective performance is validated by extensive experiments. Hongkai Xiong, Zhihai He, Songyu Yu, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2011 | Event-Based Semantic Image Adaptation for User-Centric Mobile Display DevicesabstractThis paper proposes a semantic image adaptation scheme for heterogeneous mobile display devices. This scheme aims to provide mobile users with the most desired image content by integrating the content semantic importance with user preferences under limited mobile display constraints. The main contributions of the proposed scheme are: 1) seamless integration of mobile user supplied query information with low level image features to identify semantically important image contents; 2) integration of semantic importance and user feedback to dynamically update mobile user preferences; and 3) perceptually optimized adaptation for image display on mobile devices. In order to bridge the semantic gap for adaptation, we design a Bayesian fusion approach to properly integrate low level features with high level semantics. To accommodate the variation of user preferences, the system involves mobile users in the adaptation process with only a few simple feedbacks so as to present to the users most interesting content on mobile devices. Eventually, perceptually optimized adaptation is performed to present the best image content for mobile users according to mobile display capacities. Extensive experiments have been carried out based on several common events defined in Kodak's consumer image database. These experiments show that by utilizing the proposed semantic adaptation scheme with integration of the semantics and mobile user preferences, perceptually relevant adaptation can be effectively carried out to tailor the image towards user intentions under the mobile environment constraints. Wenyuan Yin, Jiebo Luo 0001, Chang Wen Chen |
IEEE Trans. Multim. | 3 |
| 2010 | Blind Channel Equalization for Fast Moving Terminals in Prioritized Spatial Multiplexing MIMO SystemsabstractA prioritized spatial multiplexing scheme has recently been developed for the transmission of scalable video coded (SVC) video over MIMO systems. One unique feature of this scheme is its capability to channelize the sub-channels according to the importance of the individually transmitted data streams. We consider in this research a challenging task of providing high quality multimedia services over MIMO systems with fast moving terminals. With fast moving terminals many existing approaches for MIMO channel estimation and equalization cannot be applied as most of them rely on the help of periodic pilot signals. When the terminals are travelling in high speed, the undesired Doppler effects prevent the existing approaches from efficient use of pilot signals for channel estimation and equalization. Although existing blind channel estimation and equalization schemes are able to avoid the problems associated with pilot signals, these schemes cannot be directly adopted for the prioritized spatial multiplexing MIMO system for video streaming in which data streams under different modulations need to be transmitted through sub-channels. In this paper, we develop a blind constant modulus algorithm (CMA)-based channel equalization (CMACE) scheme to track the fast time-varying channel without the requirement of pilot signals. We show that the transmitted signal in each sub-channel can be modulated with different symbol constellation, and delivered at different rate according to their importance. Simulation results demonstrate that the proposed scheme is effective for channel equalization with fast moving terminal in the prioritized spatial multiplexing MIMO systems. Qian Liu 0001, Chang Wen Chen |
GLOBECOM | 2 |
| 2010 | Low-complexity rate control based on rho-domain model for Scalable Video CodingabstractIn this paper, we present a novel low complexity yet very efficient rate control algorithm for Scalable Video Coding (SVC) based on the ρ-domain model. Compared with the conventional ρ-domain model in determining quantization parameter, the proposed algorithm adopts a new linear model to obtain the quantization parameter at frame-level. This linear model is able to characterize the relationship between the percentage of zero among quantized coefficients and the quantization step. Since this model considers the Mean Absolute Difference (MAD) of each frame, individual frame only needs to be encoded once to control the bit-rate. In particular, the model parameter in the ρ-domain model can be adaptively estimated by utilizing either temporal or inter-layer information. This leads to significant improvement in estimation accuracy. Experimental results show that the proposed algorithm can achieve accurate bit-rate and maintain the maximum mismatch between actual bit-rate and target bit-rate within 0.5%. More importantly, the proposed low complexity algorithm is able to obtain PSNR improvement up to 0.6dB comparing with the algorithm adopted in the Joint Scalable Video Model (JSVM). Meng Liu 0006, Houqiang Li, Chang Wen Chen |
ICIP | 4 |
| 2010 | Sparse dyadic mode for depth map compressionabstractIn order to enable new video applications such as 3DTV and free-viewpoint video, new data formats including both 2D video sequences and corresponding depth map sequences have been proposed. One major characteristic making the depth maps different from video frames is that they typically consist of homogeneous areas separated by sharp edges representing depth discontinuities. Another characteristic of depth map sequences is that the edges exhibit quite similar boundary behaviors as the edges in the corresponding video frames. In this paper, we propose a novel sparse dyadic mode in the design of an efficient depth map compression algorithm through appropriately exploiting these characteristics. With sparse representations of depth blocks and effective reference of edge information from the corresponding video frames, sparse dyadic mode can achieve up to 1.5 dB gain on rendering quality as compared to depth sequences coded using MVC at the same bitrate. Shujie Liu 0001, PoLin Lai, Dong Tian, Cristina Gomila, Chang Wen Chen |
ICIP | 5 |
| 2010 | A novel prioritized spatial multiplexing for MIMO wireless system with application to H.264 SVC videoabstractSpatial multiplexing MIMO wireless system is an excellent choice for next generation broadband multimedia communications because of its parallel transmission capability. A major challenge for such MIMO system is that it usually requires appropriate decomposition of both wireless channels and multimedia data bitstreams. Existing approaches to MIMO channel decomposition have been passively based on simple channel feedback. To facilitate proper match between decomposed channels with compressed video that exhibit variable priority for different layer of bitstreams, it is necessary to proactively prioritize the decomposed MIMO channels. We present in this paper a novel pre-coding scheme capable of integrating both channel and source characteristics in order to achieve the desired prioritized spatial multiplexing. This pre-coding scheme is applied to the transmission of video data compressed with H.264 Scalable Video Coding (SVC) standard. Based on the desired bit-error-rate (BER) and signal-to-noise-ratio (SNR) for each SVC video layer, the base layer video is proactively assigned the highest priority sub-channel to guarantee minimum required quality while enhancement layers are mapped to other lower priority sub-channels. Comparing with existing schemes, the proposed pre-coder has low computational complexity and is suitable for rapid hardware implementation. The experimental results demonstrate that the proposed prioritized spatial multiplexing approach can efficiently enhance the quality of SVC video streaming over MIMO systems. Qian Liu 0001, Shujie Liu 0001, Chang Wen Chen |
ICME | 3 |
| 2010 | User guided semantic image adaptation for mobile display devicesabstractThis paper proposes a novel semantics-based consumer photo adaptation scheme for users of small-display mobile devices. The main contributions of the proposed scheme are: (1) seamless integration of mobile user supplied semantic information with low level image features to identify se-mantically important regions-of-interest (ROI), and (2) perceptually optimized adaptation for photo display on mobile devices. In order to bridge the semantic gap in photo search in a consumer photo collection and to perform appropriate adaptation of perceptually important region-of-interest from each photo to fit small displays on the mobile devices, we design a Bayesian fusion approach to properly integrate low level features with high level semantics. Low level features are extracted in a bottom-up fashion while the high level semantics is applied in a top-down style. Extensive experiments have been carried out based on several common events defined in the Kodak consumer photo database. These experiments show that by utilizing the semantics provided by the mobile device users, perceptually consistent adaptation can be effectively carried out. This new semantic adaptation scheme is able to outperform the conventional attention-based scheme for the common events and produce desired regions for the small displays on mobile devices. Wenyuan Yin, Jiebo Luo 0001, Chang Wen Chen |
ICME | 3 |
| 2010 | A joint source-channel adaptive scheme for wireless H.264 video authenticationabstractThis paper proposes a novel joint source-channel adaptive scheme that integrates the authentication into source and channel coding components to achieve 100% effective verification probability and an optimal end-to-end video quality. By jointly considering source coding and channel conditions with authentication, the proposed layered framework is able to minimize end-to-end quality degradation incurred by both wireless channel noise and authentication failure. In particular, the competing requirements of high verification probability and low authentication overhead are concurrently satisfied by elegant design of hash appending with efficient adaptation to H.264 source coding and channel conditions. A joint source-channel-authentication rate allocation scheme is then developed to achieve optimal end-to-end video quality. Experiment results on H.264 video sequences confirm the efficacy of this joint adaptive scheme and demonstrate that it indeed outperforms the state-of-the-art graph based authentication algorithms. Xinglei Zhu, Chang Wen Chen |
ICME | 2 |
| 2010 | Stable Maximum Throughput Broadcast in Wireless Fading ChannelsabstractThis research considers network coded broadcast system with multi-rate transmission and dual queue stability constraints. Existing network coded broadcast systems consider single rate transmission without receiver queue constraints. First, we shall illustrate that broadcast without network coding cannot support maximum throughput in wireless fading channels. However, the network coded broadcast poses new constraints for the receivers to manage stable queues while the fading channel characteristics suggest the broadcast to operate at multi-rate to achieve higher throughput. In this research, we propose a joint scheduling and network coding (JSNC) strategy for such network coded broadcast system to achieve maximum throughput under queue stability constraint. In a single cell broadcast networks with exogenous arrivals of packets at the base station, we prove that JSNC can stabilize the system as long as the rate of the exogenous arrival flow is within the capacity region. Sufficient control parameters are provided in JSNC for trading off between sender's buffer and the receivers' buffers. JSNC can be viewed as a generalization of the classical backpressure scheduling rule to coded information flow. Alternatively, JSNC can also be viewed as an extension of network coding theory to queuing system. Hao Cui 0001, Chong Luo 0001, Feng Wu 0001, Chang Wen Chen |
INFOCOM | 5 |
| 2010 | A cross-layer adaptation HCCA MAC for QoS-aware H.264 video communications over Wireless Mesh NetworksabstractWe present in this paper a novel scheme of QoS-guaranteed transmission of H.264 video over Wireless Mesh Networks (WMNs) based on a Cross-Layer Adaptation HCCA (CLA-HCCA) MAC protocol. This CLA-HCCA strategy utilizes the queue length information at the MAC layer and the MAC layer interference estimation for optimal routing selection at network layer. The scheme will exchange timely Link Capacity Estimation (LCE) information for an adaptive adjustment of spectrally capable transmission at the physical layer. This Cross-Layer Adaptation approach provides an optimized transmission to guarantee the minimum packets delay and drop rate needed for time bound applications, such as video. Since conventional IEEE 802.11e HCCA MAC does not offer cross-layer design framework, we need to resolve the problem associated with IEEE 802.11e HCCA MAC by designing a novel adaptive architecture based on LCE. This integrated scheme allows the physical layer to achieve the optimal transmission via Video-Adaptive FEC (VA-FEC) scheme implemented in the application layer. We evaluate the proposed scheme based on network-level metrics, including bit rate, packets delay and drop rate, in comparison with the state-of-the-art Physical Rate Based Admission Control scheme (PRBAC) HCCA MAC scheme based on IEEE 802.lie WMNs. Extensive simulation results based on H.264 video transmission for the proposed CLA-HCCA and the PRBAC-HCCA schemes have been obtained. We have confirmed that the proposed CLA-HCCA strategy outperforms the PRBAC-HCCA MAC by a significant margin as LCE and VA-FEC is able to adapt to the dynamic network conditions of IEEE 802.lie WMNs in a timely fashion. Byung Joon Oh, Chang Wen Chen |
ISCAS | 2 |
| 2010 | Semantic adaptation of consumer photo for mobile device accessabstractThis paper proposes a novel event-aware semantic image adaptation framework for accessing consumer photos via mobile devices. There are two major challenges to develop such an image adaptation framework: (1) how to bridge the semantic gap in the search for the desired photos from a consumer database; (2) how to perform appropriate adaptation for various small displays in order to maximize user's perceptual experience according to the semantic significance of objects in photos. To meet the first challenge, an event semantics-guided extraction is developed based on statistical fusion of bottom-up low level features and top-down semantic features. To meet the second challenge, a key object-based semantic adaptation approach is designed to obtain the perceptually optimized Region of Interest (ROI). We conducted experiments based on the events defined in the Kodak consumer photo database. These experiments show that by utilizing the semantics of the photos with event-based a priori knowledge, the adaptation result outperforms the conventional attention based scheme. More importantly, the proposed framework can be easily extended to other events by adopting the event semantics based a priori knowledge into the adaptation system. Wenyuan Yin, Jiebo Luo 0001, Chang Wen Chen |
ISCAS | 3 |
| 2010 | Error resilient scalability for video bit-stream over heterogeneous packet loss networksabstractRobust transmission of compressed video bit-streams over heterogeneous packet loss network is one of the key challenges in the contemporary video communication system. Recent researches have focused on enhancing error resilience of bit-stream through the deployment of transcoder between wired and wireless networks. However, the conventional error resilient algorithms (e.g., Intra refresh) through cascading decoding and encoding processes usually have high degree of complexity. We introduce in this paper a completely new concept of error resilient scalability. We develop an extremely low complexity scheme based on redundant pictures information. In this scheme, redundant picture information is generated at the encoder and can be applied at media gateway to determine the redundant quantity of bit-stream according to the packet loss rate of access network. By transmitting compressed bit-stream together with redundant picture information, we are able to achieve the desired error resilient scalability of video bit-stream at the media gateway for heterogeneous networks. Joint rate source-channel-distortion model is adopted to optimize the generation of redundant picture information under various packet loss rates. Experimental results demonstrate expected effectiveness of the proposed scheme in error resilience scalability. Houqiang Li, Chang Wen Chen |
ISCAS | 4 |
| 2010 | ACM workshop on mobile cloud media computingabstractSmart mobile devices such as camera phones typically will be carried by people all the time. These devices are true "multimedia" devices that acquire, process, transmit and present text, image, video and audio data. However, due to the limitations in hardware and networking, multimedia applications and systems have not been adequately supported on mobile devices. With the recent developments mobile hardware, wireless network, and cloud computing, it is now the prime time for us to realize intelligent mobile device centered multimedia applications with the support of a cloud computing platform. The focus of this workshop is on exploring challenges and opportunities of intelligent multimedia technologies, applications and systems on mobile devices, especially when a media cloud computing platform can be appropriately leveraged. Xian-Sheng Hua 0001, Gang Hua 0001, Chang Wen Chen |
ACM Multimedia | 3 |
| 2010 | 3D video transcoding for virtual viewsabstractRecent emerging development of three dimensional video (3DV) has been vigorously driving the Multiview Video Coding (MVC) standard developed by Joint Video Team as an amendment to H.264/AVC and the new 3DV standard developed by MPEG. It is expected that 3DV contents will soon be available to various media consumers. Because of the heterogeneous networks and terminal devices, a variety of end users shall not have (1) the 3DV decoder and view synthesize software installed on their devices, and (2) adequate network bandwidth for 3DV content delivery. 3DV transcoder is therefore necessary for users without 3DV decoder to enjoy 3DV services, as well as for service providers to better control virtual view quality of all 3DV customers. In this paper, we report the first 3DV transcoding scheme for virtual view, which is able to generate bitstream of one single view that can be decoded by H.264/AVC. The key idea of the proposed transcoding is to appropriately generate candidate motion and modes based on motion information from other views in the original bitstream making best use of inter-view correlation. These candidate motion and modes are then properly used to encode a virtual view with significantly reduced complexity for decoding by H.264/AVC. Simulation results show that, compared to the straightforward cascade algorithm, the proposed transcoding experiences minor performance loss with a significantly reduced computation cost. Such low complexity transcoding can be deployed at various media gateways to deliver 3DV content to non-3DV devices. Shujie Liu 0001, Chang Wen Chen |
ACM Multimedia | 2 |
| 2010 | Joint trilateral filtering for depth map compressionabstractNew data formats including 2D video and the corresponding depth maps enable new video applications in which virtual views can be rendered, such as 3DTV and free-viewpoint video (FVV). Different from video frames, depth maps typically consist of homogeneous areas (with no textures) separated by sharp edges representing depth value changes such as between foreground and background. Conventional video coding techniques with transforms followed by quantization typically result in large artifacts along such sharp edges. To suppress these coding artifacts while preserving edges, we propose in this paper a novel filtering method for depth coding, joint trilateral filter. The main contribution in the proposed filter design is the utilization of edge information in the collocated video frame as well as in the depth map. The filtering weights are determined by the following three factors: a domain (spatial) filter which measures the proximity of pixel positions, and two range filters. One range filter takes into account the similarity among depth samples and the other one considers the similarity among the collocated pixels in the video frame. By replacing the deblocking filter in H.264/AVC with the proposed trilateral filter, simulation results demonstrate up to 0.8 dB gain in rendering quality at given bitrate for depth signal. Shujie Liu 0001, PoLin Lai, Dong Tian, Cristina Gomila, Chang Wen Chen |
VCIP | 5 |
| 2010 | A deadline-aware transmission framework for H.264/AVC video over IEEE 802.11e EDCA wireless networksabstractOne of the most challenging issues in video transmission over wireless networks is to address the rigid time bounded constraint for video delivery. We propose in this paper a deadline-aware transmission framework (DATF) for video over IEEE 802.11e EDCA wireless networks. In this new framework, we estimate the deadline for time bounded video delivery for each packet in queues at MAC layer according to the sequence number of a frame a given video data packet belonging to as well as the current network delay. Then, the MAC layer determines whether a packet should be sent or should be dropped based on the estimated deadline information. To accomplish the scheme of DATF, we propose to modify the mapping scheme in IEEE 802.11e EDCA to facilitate unequal deadline requirement of video packets. Instead of mapping all video packets into class AC_VI, which is defined for video data in IEEE 802.11e EDCA, we differentiate video packets further based on the dependency characteristics of a given frame type. The proposed DATF scheme has been implemented with NS-2 simulation based on the scenario of wired-cum-wireless network architecture. We compare the proposed approach with several competing schemes and the simulation results show that the proposed scheme outperforms these competing schemes in terms of both wireless networking metrics and received video quality. Jianchao Du, Chang Wen Chen |
VCIP | 2 |
| 2010 | Hybrid bit-stream rewriting from scalable video coding to H.264/AVCabstractScalable Video Coding (SVC) is an extension of H.264/AVC standard. The base layer of SVC is compatible with H.264/AVC standard, while the enhancement layers provide desired temporal, quality and/or spatial scalabilities. Bit-stream rewriting in SVC standard allows an SVC bit-stream to be converted to an H.264/AVC bit-stream without quality loss and preferably with low computational complexity. However, current rewriting is only supported in quality scalability rather than spatial scalability, which limits the application in many practical scenarios. In this paper, a hybrid bit-stream rewriting approach to support both quality and spatial scalability is proposed based on the principle of residue upsampling in transform domain. The computational complexity of the proposed approach is much lower than the conventional scheme of cascading transcoding. Extensive experimental results demonstrate that the loss of the rate-distortion (RD) performance of the proposed rewritable SVC bit-stream is acceptable compared with the conventional SVC bit-stream, however, the RD performance is better than that of simulcast. Furthermore, the RD performance of the H.264/AVC bit-stream rewritten from the rewritable SVC bit-stream is even better than that of the input SVC bit-stream. Compared with the cascading transcoding scheme, the proposed hybrid rewriting can achieve 0.8 dB Y-PSNR gains while saving 80% processing time on average. Bin Li 0012, Houqiang Li, Chang Wen Chen |
VCIP | 4 |
| 2010 | An adaptive approach to human motion tracking from videoabstractVision based human motion tracking has drawn considerable interests recently because of its extensive applications. In this paper, we propose an approach to tracking the body motion of human balancing on each foot. The ability to balance properly is an important indication of neurological condition. Comparing with many other human motion tracking, there is much less occlusion in human balancing tracking. This less constrained problem allows us to combine a 2D model of human body with image analysis techniques to develop an efficient motion tracking algorithm. First we define a hierarchical 2D model consisting of six components including head, body and four limbs. Each of the four limbs involves primary component (upper arms and legs) and secondary component (lower arms and legs) respectively. In this model, we assume each of the components can be represented by quadrangles and every component is connected to one of others by a joint. By making use of inherent correlation between different components, we design a top-down updating framework and an adaptive algorithm with constraints of foreground regions for robust and efficient tracking. The approach has been tested using the balancing movement in HumanEva-I/II dataset. The average tracking time is under one second, which is much shorter than most of current schemes. Lifang Wu, Chang Wen Chen |
VCIP | 2 |
| 2010 | Special issue on multimedia networking and security in convergent networks
Chang Wen Chen, Stefanos Gritzalis, Pascal Lorenz, Shiguo Lian |
Comput. Commun. | 1 |
| 2010 | Feedback-free rate-allocation scheme for transform domain Wyner-Ziv video coding
Xinglei Zhu, Guogang Hua, Hongxing Guo, Jingli Zhou, Chang Wen Chen |
Multim. Syst. | 6 |
| 2010 | Color to Gray: Visual Cue PreservationabstractBoth commercial and scientific applications often need to transform color images into gray-scale images, e.g., to reduce the publication cost in printing color images or to help color blind people see visual cues of color images. However, conventional color to gray algorithms are not ready for practical applications because they encounter the following problems: 1) Visual cues are not well defined so it is unclear how to preserve important cues in the transformed gray-scale images; 2) some algorithms have extremely high time cost for computation; and 3) some require human-computer interactions to have a reasonable transformation. To solve or at least reduce these problems, we propose a new algorithm based on a probabilistic graphical model with the assumption that the image is defined over a Markov random field. Thus, color to gray procedure can be regarded as a labeling process to preserve the newly well--defined visual cues of a color image in the transformed gray-scale image. Visual cues are measurements that can be extracted from a color image by a perceiver. They indicate the state of some properties of the image that the perceiver is interested in perceiving. Different people may perceive different cues from the same color image and three cues are defined in this paper, namely, color spatial consistency, image structure information, and color channel perception priority. We cast color to gray as a visual cue preservation procedure based on a probabilistic graphical model and optimize the model based on an integral minimization problem. We apply the new algorithm to both natural color images and artificial pictures, and demonstrate that the proposed approach outperforms representative conventional algorithms in terms of effectiveness and efficiency. In addition, it requires no human-computer interactions. Mingli Song, Dacheng Tao, Chun Chen 0001, Xuelong Li 0001, Chang Wen Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2010 | A New Hybrid DCT-Wiener-Based Interpolation Scheme for Video Intra Frame Up-SamplingabstractVideo frame resizing has received more and more research attention as contemporary video distributions need to deliver video to various receiving devices with different display resolutions. It is now becoming necessary for video distribution systems to be able to generate higher resolution video from lower resolution one for some end users. Current schemes in video up-sampling either only focused on improving visual quality with little or no improvement in objective quality or increasing the up-sampling accuracy by adaptively optimizing interpolation methods. Many such algorithms are typically quite complex and require substantial extra implementation efforts. In this letter, we propose a new hybrid DCT-Wiener-based interpolation scheme for video intra frame up-sampling without referencing the original high resolution video frames. This scheme takes full advantage of interpolation in both DCT domain and spatial domain and seamlessly integrates these two approaches to design an improved up-sampling filter. Experiments have been carried out to demonstrate that noticeable improvement in both objective and visual quality are obtained. With similar complexity, the proposed algorithm can achieve up to 4 dB gain in PSNR over fixed parameter Wiener filter-based interpolation, 4 dB gain over popular bicubic interpolation, and 1 dB gain over up-sampling scheme in DCT domain. Chang Wen Chen |
IEEE Signal Process. Lett. | 3 |
| 2010 | Editorial Message From the Outgoing Editor-in-Chief
Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Efficient Measurement Generation and Pervasive Sparsity for Compressive Data GatheringabstractWe proposed compressive data gathering (CDG) that leverages compressive sampling (CS) principle to efficiently reduce communication cost and prolong network lifetime for large scale monitoring sensor networks. The network capacity has been proven to increase proportionally to the sparsity of sensor readings. In this paper, we further address two key problems in the CDG framework. First, we investigate how to generate RIP (restricted isometry property) preserving measurements of sensor readings by taking multi-hop communication cost into account. Excitingly, we discover that a simple form of measurement matrix [I R] has good RIP, and the data gathering scheme that realizes this measurement matrix can further reduce the communication cost of CDG for both chain-type and tree-type topology. Second, although the sparsity of sensor readings is pervasive, it might be rather complicated to fully exploit it. Owing to the inherent flexibility of CS principle, the proposed CDG framework is able to utilize various sparsity patterns despite of a simple and unified data gathering process. In particular, we present approaches for adapting CS decoder to utilize cross-domain sparsity (e.g. temporal-frequency and spatial-frequency). We carry out simulation experiments over both synthesized and real sensor data. The results confirm that CDG can preserve sensor data fidelity at a reduced communication cost. Chong Luo 0001, Feng Wu 0001, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 4 |
| 2009 | Stateful Scheduling with Network Coding for Roadside-to-Vehicle CommunicationabstractIn urban areas such as street blocks, media-rich services are delivering to moving vehicles through Access Points (APs) deployed on the roadside. To increase the capacity of wireless links for these services, a roadside-to-vehicle communication system based on a novel Stateful Scheduling with Network Coding (SSNC) strategy is proposed. A major innovation of SSNC is that it enables the scheduling on the roadside AP to fully utilize the states of received data from vehicles for an enhanced performance. Furthermore, a set of mechanisms are proposed to ensure reliable transmission in such highly dynamic and error-prone wireless channels. Extensive simulations show that SSNC achieves an increased throughput gain between 1.2 and 2.0 at various conditions. Hancheng Lu, Feng Wu 0001, Chang Wen Chen |
ICC | 3 |
| 2009 | Distributed multiview video coding using the fusion of triple side informationabstractTraditional multiview video coding schemes exploit all types of correlations at the encoder. The inherent high encoder complexity and the needs for large volume data communication between cameras make them impractical for many applications. In this paper, we propose a novel distributed coding scheme for multiview video. All the views are separately encoded but jointly decoded. At the decoder, we design a fusion algorithm that is able to make full use of three types of side information: temporal prediction, interview motion prediction and inter-view luma prediction. This is the key innovation in this research. This fusion scheme is able to exploit both intra-view and inter-view correlations with a low complexity encoder. Experimental results indicate significant R-D performance gain over single view distributed video coding and traditional intra coding. Houqiang Li, Chang Wen Chen |
ICME | 4 |
| 2009 | A cross-layer oriented multi-channel MAC protocol design for QoS-centric video streaming over wireless ad hoc networksabstractThis paper presents a cross-layer design for a QoS-centric video transmission over wireless ad hoc networks based on multi-channel MAC protocol with TDMA. First, we investigate a study of the multi-channel MAC protocol through Markov chain modeling. Based on this study, two innovative cross-layer modules are adopted for the design of multi-channel MAC protocol. In the first module, we adopt maximum latency rate (MLR) as the channel quality metric. Unlike the traditional MAC design based on network allocation vector (NAV), MLR is implemented to provide differentiated traffic so that the channel with smaller MLR time is initiated for higher priority traffic. In the second module we adopt two congestion aware metrics, namely MAC utilization and queue length of the MAC layer, to improve the congestion aware routing protocol with DSR. These two innovative modules allow the proposed MAC protocol design to achieve high performance video transmission over wireless ad hoc networks. Experimental results show that the proposed scheme outperforms the state-of-the-art schemes under multi-channel environments in wireless ad hoc networks for as much as 4.9 dB in PSNR. Such significant performance enhancement confirms that the cross-layer approach is very effective for multi-channel MAC protocol design. Byung Joon Oh, Chang Wen Chen |
ICME | 2 |
| 2009 | Distributed coding techniques for onboard lossless compression of multispectral imagesabstractIn this paper, we propose a low complexity lossless compression scheme for multispectral images based on distributed coding. Data decorrelation operations are moved to decoder on the ground in order to design a lightweight yet very efficient encoder suitable for onboard applications. The decoder with abundant resources will perform spectral-spatial adaptive fuzzy prediction to generate high quality side information for distributed coding by capturing the spatially varying spectral correlation. The LDPC decoding is integrated with Markov random field (MRF) modeling aiming at joint exploitation of spectral correlation and spatial correlation. Simulations have been carried out to demonstrate that the proposed scheme is able to provide competitive performance with respect to the state-of-the-art 3D DPCM technique but with significantly lower encoding complexity. Jinrong Zhang 0002, Houqiang Li, Chang Wen Chen |
ICME | 3 |
| 2009 | A joint layered coding scheme for unified reliable and secure media transmission with implementation on JPEG 2000 imagesabstractThis paper presents a novel stream-level joint layered coding scheme for unified reliable and secure media transmission over wireless networks. The proposed scheme simultaneously protects both compressed media content and the authentication data from wireless channel impairments. Therefore, the media quality degradation incurred by both channel noise and authentication constraints can be minimized. With a prior knowledge of source coding and channel coding, the proposed joint system integrates authentication into the media error protection components to ensure 100% effective verification probability, i.e. every source decodable media unit is authenticable. In particular, by utilizing the contextual relationship, such as coding dependency and content importance between media slices for authentication hash appending, the proposed scheme generates an extremely low authentication overhead. The proposed authentication scheme is fundamentally different from many existing systems that consider the problem of authenticating media content separately from the other operations in the media transmission system. Under this joint layered coding framework, an optimal rate allocation algorithm for source coding, channel coding and media authentication is developed to guarantee the end-to-end media quality. Experiment results on JPEG 2000 images validate the proposed scheme and demonstrate that the performance of the proposed approach is approaching its upper bound, in which case no authentication is applied to the media stream. Xinglei Zhu, Zhishou Zhang, Chang Wen Chen |
ICME | 3 |
| 2009 | An Opportunistic Multi Rate MAC for Reliable H.264/AVC Video Streaming over Wireless Mesh NetworksabstractWe present in this paper a scheme of reliable transmission of H.264/AVC video over wireless mesh networks (WMNs) based on an opportunistic multi rate (OMR) IEEE 802.11 MAC strategy. This OMR strategy employs the channel reservation control packets at the MAC layer to exchange timely channel quality estimation (CQE) information for an adaptive adjustment of spectrally efficient transmission rate at the physical layer. Such rate adaptation scheme offers an optimized transmission to guarantee the minimum packets delay and drop rate needed for time bound applications, such as video. Since IEEE 802.11 standard does not adopt a multi-rate MAC to support time critical applications, we need to resolve the problem associated with IEEE 802.11 standard by designing a novel adaptive rate architecture with minor modification on IEEE 802.11 MAC protocol based on CQE. This integrated scheme allows the physical layer to achieve the optimal transmission via smart FEC scheme implemented in the application layer. We evaluate the proposed scheme based on network-level metrics, including bit rate, packets delay and drop rate, in comparison with the state-of-the-art opportunistic auto-rate (OAR) MAC scheme based on IEEE 802.11 WMNs. Extensive simulation results based on H.264/AVC video transmission for both proposed OMR scheme and OAR MAC have confirmed that the proposed OMR strategy outperforms the OAR MAC by a significant margin as CQE and smart FEC is able to adapt to the dynamic network conditions of IEEE 802.11 WMNs in a timely fashion. Byung Joon Oh, Chang Wen Chen |
ISCAS | 2 |
| 2009 | Compressive data gathering for large-scale wireless sensor networksabstractThis paper presents the first complete design to apply compressive sampling theory to sensor data gathering for large-scale wireless sensor networks. The successful scheme developed in this research is expected to offer fresh frame of mind for research in both compressive sampling applications and large-scale wireless sensor networks. We consider the scenario in which a large number of sensor nodes are densely deployed and sensor readings are spatially correlated. The proposed compressive data gathering is able to reduce global scale communication cost without introducing intensive computation or complicated transmission control. The load balancing characteristic is capable of extending the lifetime of the entire sensor network as well as individual sensors. Furthermore, the proposed scheme can cope with abnormal sensor readings gracefully. We also carry out the analysis of the network capacity of the proposed compressive data gathering and validate the analysis through ns-2 simulations. More importantly, this novel compressive data gathering has been tested on real sensor data and the results show the efficiency and robustness of the proposed scheme. Chong Luo 0001, Feng Wu 0001, Chang Wen Chen |
MobiCom | 4 |
| 2009 | Multiview Video transcoding: From multiple views to single viewabstractAs multiview video is gaining more and more attentions, Multiview Video Coding (MVC) standard has been under development by the Joint Video Team as an extension to H.264/AVC. There will be increasingly more multiview video sources for both high end and low end consumers. Since a variety of end user devices shall not have multiview video decoders installed, multiview video transcoder is therefore needed in order for these users to enjoy the videos encoded by MVC standard. In this paper, we propose an algorithm for transcoding any viewpoint in bitstreams encoded by MVC to bitstreams that can be decoded by H.264/AVC decoders, regardless of the MVC prediction structure. The key component of this algorithm is the design of motion information reuse to implement simple motion refinement for inter-view coded macroblocks. Such reduced complexity implementation is especially applicable to the low power mobile devices. Experimental results show that, comparing with cascaded algorithm, the proposed scheme, with much reduced complexity, only suffers moderate performance loss. Shujie Liu 0001, Chang Wen Chen |
PCS | 2 |
| 2009 | Forepressure Transmission Control for Wireless Video Sensor NetworksabstractMulti-source data transmission in wireless video sensor networks is a challenging problem because of the high bandwidth demand of video streams and the many-to-one traffic pattern. This paper proposes forepressure transmission control for efficient and fair video transmission in this scenario. Contrary to traditional backpressure transmission control where downstream node creates backpressure and causes upstream node to hold off when its buffer is full, forepressure transmission control creates forepressure and signals upstream node to transmit when the downstream node's buffer is not full. Forepressure transmission control proactively avoids congestion at little communication overhead. In addition, it is able to provide fairness among concurrent flows. We evaluate forepressure transmission control via NS-2 simulations, and compare it with both backpressure congestion control and end-to-end rate regulation. Results in three typical topologies show that forepressure transmission control achieves much lower loss ratio and ensures better fairness than the other two schemes. Chong Luo 0001, Chang Wen Chen, Feng Wu 0001 |
SECON | 3 |
| 2009 | Test-pattern-reduced decoding for turbo product codes with multi-error-correcting eBCH codesabstractWe present a method to reduce the number of test patterns (TPs) decoded in the Chase-II algorithm for turbo product codes (TPCs) constructed with multi-error-correcting extended Bose-Chaudhuri-Hocquengem (eBCH) codes. We classify TPs into different conditions based on the relationship between syndromes and the number of errors so that TPs with the same codeword are not decoded except the one with the least number of errors. For eBCH with code length of 64, simulation results show that over 50% of TPs need not to be decoded without any performance degradation. Guo Tai Chen, Lei Cao 0001, Lun Yu, Chang Wen Chen |
IEEE Trans. Commun. | 4 |
| 2009 | Message From the Editor-in-Chief
Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Natural and Seamless Image Composition With Color ControlabstractWhile the state-of-the-art image composition algorithms subtly handle the object boundary to achieve seamless image copy-and-paste, it is observed that they are unable to preserve the color fidelity of the source object, often require quite an amount of user interactions, and often fail to achieve realism when there exists salient discrepancy between the background textures in the source and destination images. These observations motivate our research towards color controlled natural and seamless image composition with least user interactions. In particular, based on the Poisson image editing framework, we first propose a variational model that considers both the gradient constraint and the color fidelity. The proposed model allows users to control the coloring effect caused by gradient domain fusion. Second, to have less user interactions, we propose a distance-enhanced random walks algorithm, through which we avoid the necessity of accurate image segmentation while still able to highlight the foreground object. Third, we propose a multiresolution framework to perform image compositions at different subbands so as to separate the texture and color components to simultaneously achieve smooth texture transition and desired color control. The experimental results demonstrate that our proposed framework achieves better and more realistic results for images with salient background color or texture differences, while providing comparable results as the state-of-the-art algorithms for images without the need of preserving the object color fidelity and without significant background texture discrepancy. Jianmin Zheng, Jianfei Cai 0001, Susanto Rahardja, Chang Wen Chen |
IEEE Trans. Image Process. | 5 |
| 2009 | A Cross-Layer Approach to Multichannel MAC Protocol Design for Video Streaming Over Wireless Ad Hoc NetworksabstractThis paper presents a cross-layer design for a reliable video transmission over wireless ad hoc networks based on multichannel MAC protocol with TDMA. First, we conduct a study of the multichannel MAC protocol through Markov chain model. Based on this study, two novel cross-layer modules are adopted for the design of multichannel MAC protocol. First, we adopt maximum latency rate (MLR) as the channel quality metric. Unlike the traditional MAC design based on network allocation vector (NAV), MLR is implemented to provide differentiated traffic so that the channel with smaller MLR time is initiated for higher priority traffic. Second, we adopt two congestion-aware metrics, namely MAC utilization and queue length of MAC layer, to improve the congestion-aware routing protocols with AODV and DSR. These two novel modules allow the proposed MAC protocol design to achieve high performance video transmission over wireless ad hoc networks. Experimental results show that the proposed scheme outperforms the state-of-the-art schemes under multichannel environments in wireless ad hoc networks for as much as 3.6 dB in PSNR. Such significant performance enhancement confirms that the cross-layer approach is very effective for multichannel MAC protocol design. Byung Joon Oh, Chang Wen Chen |
IEEE Trans. Multim. | 2 |
| 2009 | QoS-driven network coded wireless multicastabstractEmerging wireless multicast applications simultaneously impose two requirements to the underlying communication networks: to provide sufficient bandwidth and to support a variety of quality-of-service (QoS) sensitivities. For the first requirement, network coding has been proposed recently as an effective way of improving bandwidth utilization. However, almost all previous works about network coding focus on the throughput gain without considering the QoS requirements. Optimal network code construction in wireless multicast under different QoS constraints remains as a significant challenge. In this research, we study the QoS-driven network coding problem. We use large deviation principle to establish the relationship among source rate, link condition, QoS requirement, and network code. Using this relationship, under given QoS requirements, we solve the optimal network code construction problem. The proposed network code supports maximal source rate without violating the QoS requirements. These results constitute the foundations for future designing and implementing network coding based wireless multicast protocols. Chong Luo 0001, Feng Wu 0001, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 4 |
| 2008 | Seamless Video Transmission over Wireless LANs based on an Effective QoS Model and Channel State EstimationabstractThis paper presents a scheme for video transmission over wireless LANs based on an effective QoS model and channel state estimation (CSE). The proposed QoS model is able to effectively characterize the performance of video transmission over error prone wireless channels. Unlike existing schemes that use complicated metrics based on the actual distortion computed from decoded video frames, the proposed model employs H.264 group of pictures (GOP) level estimator of decodable slice rate (GEDSR). This new GEDSR model can be obtained from simplified calculation of decodable I, P, and B slices without computing the actual distortions. With this model, it is possible for the estimation to include the dependence of error control such as forward error correction and redundant slices. In this research, we also develop an adaptive cross-layer feedback mechanism based on channel status, maximum retry limit value, and RTS/CTS function. Extensive simulations have been carried out to confirm the effectiveness of the proposed GEDSR model for the distortion analysis. The experimental results of the GEDSR model analysis have been applied to selecting forward error correction (FEC) coding parameters for video streams to respond to dynamic channel changes. Simulations on video transmission over wireless LAN show that the proposed scheme is able to achieve significant performance gains (7dB) over schemes without channel feedback. Byung Joon Oh, Guogang Hua, Chang Wen Chen |
ICCCN | 3 |
| 2008 | Robust Video Transmission Over Packet Erasure Wireless Channels Based on Wyner-Ziv Coding of Motion RegionsabstractThis paper presents a new scheme for robust video transmission over packet erasure wireless channels based on Wyner-Ziv coding of motion regions. The multipath fading and shading of the wireless channels usually lead to loss or erroneous video packets which on occasions result in some spontaneous drop in video quality. Existing approaches with forward error correction (FEC) and error concealment have not been able to provide the desired robustness in video transmission. We develop a new scheme with a motion-based Wyner-Ziv coding (MWZC) by leveraging distributed source coding (DSC) ideas for error robustness. This new scheme is based on the fact that motion regions of a given video frame are particularly important in both objective and perceptual video quality and hence should be given preferential Wyner-Ziv coding based embedded protection. To achieve high coding efficiency, we determine the underlining motion regions based on a rate-distortion model. Within the framework of H.264/AVC specification, motion region determination can be efficiently implemented using flexible macroblock ordering (FMO) and data partitioning (DP). The bit stream generated by the proposed scheme consists two parts: the systematic portion generated from conventional H.264/AVC bit stream and the supplementary bit stream for error robust video transmission generated by the Wyner-Ziv coding of motion regions. Experimental results demonstrate that the proposed scheme significantly outperforms both decoder-based error concealment (DBEC) and conventional FEC with DBEC approaches. Byung Joon Oh, Guogang Hua, Hongxing Guo, Jingli Zhou, Chang Wen Chen |
ICCCN | 6 |
| 2008 | Distributed video coding with zero motion skip and efficient DCT coefficient encodingabstractIn this paper, we propose a suite of efficient schemes for the representation of spatial and temporal correlations of video signals in distributed video coding (DVC). These schemes consist of zero motion skipping, Gray codes, and sign bits coding of DCT coefficients developed to improve the overall rate-distortion (R-D) performance of distributed video coding. Existing schemes have focused on exploiting the spatial and temporal correlation without sufficient investigation on the efficient representation of the correlation. We believe this is one of the reasons that there still exists substantial gap between current DVC and traditional hybrid video coding. In traditional hybrid coding, a suite of efficient representation schemes such as zig-zag scan, run length code, and skipped macroblock have contributed significantly to excellent R-D performance. The efficient representation schemes we developed for DVC in this research will also lead to improved rate-distortion performance comparing with existing DVC schemes. We present in this paper the overall framework for the proposed zero motion skip and the analysis for efficient representations of both DCT coefficients and their signs. The experiment results show that the distributed video coding based on these efficient representations is able to achieve considerably improved rate-distortion performance over existing schemes. Guogang Hua, Chang Wen Chen |
ICME | 2 |
| 2008 | Energy efficient h.264 video transmission over wireless ad hoc networks based on adaptive 802.11e EDCA MAC protocolabstractThis paper presents an energy efficient H.264 video transmission over wireless ad hoc networks based on an adaptive IEEE 802.11e EDCA (enhanced distributed channel access) mechanism. The standard IEEE 802.11e EDCA mechanism provides only service priority differentiation for time bound applications over ad hoc networks. However, this standard does not provide adaptation mechanism to respond to the time varying network link status. Furthermore, the critical issues in energy efficient transmission have not been addressed under EDCA mechanism. In this research we develop a new adaptive scheme to adjust the Contention Window (CW ) after each transmission attempt based on current network link status. This adaptation scheme is also integrated with energy efficient FHSS (BFSK) scheme in order to save the energy consumption during the transmission. An analytical model of Energy efficient estimator (E3metric) is developed to support the selection of modulation scheme. We also carry out NS-2 based simulation, in combination with E3model, to evaluate the proposed scheme in comparison with standard IEEE 802.11e EDCA. Based on extensive simulation with H.264/AVC video over the proposed AEDCA and the standard EDCA, we demonstrate that the proposed AEDCA with Energy efficient BFSK scheme outperforms the standard IEEE 802.11e EDCA with an improved video quality as well as reduced energy dissipation. Byung Joon Oh, Chang Wen Chen |
ICME | 2 |
| 2008 | Distributed image coding based on integrated Markov modeling and LDPC decodingabstractWe present in this paper a novel distributed image coding scheme by exploiting image spatial correlation via Markov modeling at the decoding end. The exploitation of image statistics at the decoding end allows us to design a simple yet efficient encoder suitable for various energy efficient imaging sensor network applications. Existing distributed coding schemes developed for imaging sensor networks mostly attempt to exploit inter-image correlation. We develop in this research an integration of LDPC decoding and Markov model estimation in order to jointly exploit both inter-image and intra-image correlation. Simulations have been carried out to demonstrate that this Markov model-based approach is able to achieve significant gains over the schemes without Markov model. The simulation results also show that the 2D Markov model is able to achieve additional gains over the 1D Markov model. Jinrong Zhang 0002, Houqiang Li, Chang Wen Chen |
ICME | 3 |
| 2008 | A joint ECC based media error and authentication protection schemeabstractThis paper presents a novel content-aware joint media error and authentication protection scheme based on error correcting coding (ECC). The innovation of the proposed scheme lies in the true joint design of error protection and authentication verification. This is fundamentally different from many existing schemes in which they consider media authentication separately from other media processing components such as error protection in media communication systems. By making use of the channel information and integrating the authentication with error protection necessary in contemporary media communication systems, we are able to achieve 100% complete verification with low authentication overhead. With such integration, the end-to-end media quality and media security guarantee can be obtained. Based on this joint error and authentication protection framework, an optimal rate allocation algorithm under certain source and channel models is also developed. Simulations based on JPEG 2000 images have been carried out to validate the proposed scheme and the simulation results show that the proposed approach is indeed able to achieve simultaneous error and authentication protection for JPEG 2000 images. Xinglei Zhu, Qibin Sun, Zhishou Zhang, Chang Wen Chen |
ICME | 4 |
| 2008 | Continuous Network Coding in Wireless Relay NetworksabstractNetwork coding has recently been applied to wireless networks and has achieved some initial success. Researches in wireless network coding have been mostly focusing on utilizing the broadcast nature of the wireless networks. In this paper, we propose a novel network coding framework for wireless relay networks that also takes into consideration the fading and error prone nature of the wireless networks. First, we extend the traditional network coding in lossless networks which operates on 0-1 bits, to a new framework which defines network coding on the posterior probability of each bit. This new framework allows an imperfect decode-recode process at a relay node and avoids possible error propagation when a hard decision is made at the relay node. It implicitly integrates decode-and-forward and estimate-and-forward strategies for wireless network coding to address the technical issues of channel fading and transmission errors. The proposed approach is validated through both theoretical analysis and extensive simulations. Both analysis and simulation confirm that this new framework is able to achieve significant gain over traditional network coding. This new framework also enables the introduction of adaptive scheme into network coding. We demonstrate a basic adaptation scheme and present some preliminary experimental results. The proposed adaptive scheme will lay down an essential foundation in this emerging field of wireless network coding in order to address issues related to link heterogeneity. Chong Luo 0001, Shipeng Li 0001, Chang Wen Chen |
INFOCOM | 4 |
| 2008 | Performance evaluation of H.264 video over ad hoc networks based on dual mode IEEE 802.11B/G and EDCA MAC architectureabstractWe present in this paper the performance evaluation for H.264/AVC video transmission over wireless ad hoc networks. In order to achieve robust video transmission under ad hoc mode, the proposed system not only develops a physical rate based strategy for dual band IEEE 802.11b/g, but also adopts EDCA MAC layer architecture of IEEE 802.11e for an improved QoS performance. The dual mode operation allows the ad hoc terminals to move within a larger coverage area while maintaining robust connection via 802.11b/g. The adopted EDCA MAC layer architecture enables the prioritized transmission of differentiated video traffic for various ad hoc terminals. We have also developed several network level metrics, including bit rate, packets delay and drop rate, to evaluate the proposed scheme in comparison with simple dual band IEEE 802.11b/g. Simulation results demonstrate that the proposed EDCA-based scheme is able to outperform the simple dual band IEEE 802.1lb/g with much improved received video quality. Byung Joon Oh, Chang Wen Chen |
ISCAS | 2 |
| 2008 | Rate allocation for transform domain Wyner-Ziv video coding without feedbackabstractIn this paper, we propose a new rate allocation algorithm for transform domain Wyner-Ziv video coding (WZVC) without feedback. In contrast to conventional video coding, Wyner-Ziv video coding aims to design simple intra-frame encoding and complex inter-frame decoding based on the Slepian-Wolf and Wyner-Ziv distributed source coding theorems. To allocate proper number of bits to each frame, most existing Wyner-Ziv video coding solutions need a feedback channel (FC) at the decoder. However, in many video coding applications, the FC is not allowed. Moreover, the FC will introduce latency and an increase of decoder complexity because several iterative decoding operations may be needed to decode the data to achieve target video quality. The proposed algorithm predicts the number of bits for each Wyner-Ziv frame at the encoder as a function of the coding mode and the quantization parameters. Such predictions will not significantly increase the complexity at the encoder. However, the prediction will be able to properly select the best mode and quantization parameter for encoding each Wyner-Ziv frame. Experimental results show that the proposed algorithms is able to achieve good encoder rate allocation while still maintains consistent coding efficiency. Comparing to the WZVC coder with FC, this new WZVC coder without FC induces only a small loss in Rate-Distortion performance. Guogang Hua, Hongxing Guo, Jingli Zhou, Chang Wen Chen |
ACM Multimedia | 5 |
| 2008 | Technical challenges in video coding and processing for future digital entertainmentabstractThis talk will present some technical challenges in video coding and processing to meet the paradigm shifting trends for future digital entertainment for consumers. Traditional consumer video services have been in the broadcasting mode, from terrestrial TV, to satellite and cable services, in which a single encoder is able to serve millions of decoders. The design principle has been the simple decoder of volume sets at the expense of very complicated encoder. The proliferation of mobile devices with video capture capabilities in the recent years has resulted in a paradigm shift trends that require simple encoder for the mobile devices. The burden of the performance has now shifted to decoder that resides at consumer’s home to manage volumetric video captures with desktop computers. This paradigm shift thus created an opportunity for new generations of video coding and processing algorithms and architectures to meet the challenges in more complicated video decoding. In this talk, we will present several examples of new video coding and processing schemes based on distributed source coding. Some detailed analysis and simulation results will be shown to demonstrate that distributed source coding based approach is indeed promising for video decoding and processing for future digital entertainment. Chang Wen Chen |
MMSP | 1 |
| 2008 | Error resilient transcoding of Scalable Video bitstreamsabstractWe propose in this paper a novel error resilient transcoding scheme that can be placed at the boundary between wired and wireless networks via heterogeneous network links. This error resilient transcoder shall seamlessly complement the standard Scalable Video Coding (SVC) bitstream to offer additional error resilient adaptation capability for receiving devices. The novel error resilient transcoding scheme consists of three different modules; each is designed to meet various levels of complexity need. The three modules are all based on the Loss-Aware Rate-Distortion Optimization (LA-RDO) mode decision algorithm we have previously developed for SVC. However, each individual module can be tailored to different complexity requirements depending on whether and how the LA-RDO mode decision is implemented. Another innovation of this approach is the design of a fast rate control algorithm in order to maintain consistent bitrates between input and output of the transcoder. This rate control algorithm only needs picture-level bit information for training target quantization parameters. Simulation results demonstrate that, comparing with standard SVC, the proposed approach is able to achieve up to 4 dB gain for the enhancement layer video and up to 1 dB gain for the base layer video. Houqiang Li, Ye-Kui Wang, Chang Wen Chen |
MMSP | 4 |
| 2008 | Distributed image coding based on integrated Markov random field modeling and LDPC decodingabstractWe present in this paper a novel distributed image coding scheme by exploiting image spatial correlation via Markov random field modeling at the decoding end. This allows us to design a simple yet efficient encoder suitable for various energy efficient imaging sensor network applications. The novelty is the integration of LDPC decoding and Markov random field modeling in order to jointly exploit both inter-image and intra-image correlation. The current research aims at improving our previous work in which the Markov model was defined by a state transition probability matrix. In this research, we model the image via a Markov random field described by Gibbs distribution. Both analysis and simulations have been carried out to demonstrate that this Markov model-based approach is able to achieve significant gains over the schemes without Markov modeling. Furthermore, this new Gibbs-based Markov model is less sensitive to correlated noise. Our approach also outperforms a JPEG codec by up to 4 dB even if the interimage correlation is not very high. Jinrong Zhang 0002, Houqiang Li, Chang Wen Chen |
MMSP | 3 |
| 2008 | Maximum-throughput delivery of SVC-based video over MIMO systems with time-varying channel capacity
Daewon Song, Chang Wen Chen |
J. Vis. Commun. Image Represent. | 2 |
| 2008 | Robust multiple description image coding over wireless networks based on wavelet tree coding, error resilient entropy coding, and error concealment
Daewon Song, Lei Cao 0001, Chang Wen Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2008 | Quality-Optimized and Secure End-to-End Authentication for Media DeliveryabstractThe need for security services, such as confidentiality and authentication, has become one of the major concerns in multimedia communication applications, such as video on demand and peer-to-peer content delivery. Conventional data authentication cannot be directly applied for streaming media when an unreliable channel is used and packet loss may occur. This paper begins by reviewing existing end-to-end media authentication schemes, which can be classified into stream-based and content-based techniques. We then motivate and describe how to design authentication schemes for multimedia delivery that exploit the unequal importance of different packets. By applying conventional cryptographic hashes and digital signatures to the media packets, the system security is similar to that achievable in conventional data security. However, instead of optimizing packet verification probability, we optimize the quality of the authenticated media, which is determined by the packets that are received and able to be decoded and authenticated. The quality of the authenticated media is optimized by allocating the authentication resources unequally across streamed packets based on their relative importance, thereby providing unequal authenticity protection. The effectiveness of this approach is demonstrated through experimental results on different media types (image and video), different compression standards (JPEG, JPEG2000, and H.264), and different channels (wired with packet erasures and wireless with bit errors). Qibin Sun, John G. Apostolopoulos, Chang Wen Chen, Shih-Fu Chang |
Proc. IEEE | 3 |
| 2008 | Video Error Concealment Using Spatio-Temporal Boundary Matching and Partial Differential EquationabstractError concealment techniques are very important for video communication since compressed video sequences may be corrupted or lost when transmitted over error-prone networks. In this paper, we propose a novel two-stage error concealment scheme for erroneously received video sequences. In the first stage, we propose a novel spatio-temporal boundary matching algorithm (STBMA) to reconstruct the lost motion vectors (MV). A well defined cost function is introduced which exploits both spatial and temporal smoothness properties of video signals. By minimizing the cost function, the MV of each lost macroblock (MB) is recovered and the corresponding reference MB in the reference frame is obtained using this MV. In the second stage, instead of directly copying the reference MB as the final recovered pixel values, we use a novel partial differential equation (PDE) based algorithm to refine the reconstruction. We minimize, in a weighted manner, the difference between the gradient field of the reconstructed MB in current frame and that of the reference MB in the reference frame under given boundary condition. A weighting factor is used to control the regulation level according to the local blockiness degree. With this algorithm, the annoying blocking artifacts are effectively reduced while the structures of the reference MB are well preserved. Compared with the error concealment feature implemented in the H.264 reference software, our algorithm is able to achieve significantly higher PSNR as well as better visual quality. Yan Chen 0007, Yang Hu 0006, Oscar C. Au, Houqiang Li, Chang Wen Chen |
IEEE Trans. Multim. | 5 |
| 2008 | Segmentation-Based View-Dependent 3-D Graphics Model TransmissionabstractFor wireless network based graphics applications, a key challenge is how to efficiently transmit complex 3-D models over bandwidth-limited wireless channels. Most existing 3-D mesh transmission systems do not consider such a view-dependent delivery issue, and thus transmit unnecessary portions of 3-D mesh models, which leads to the waste in precious wireless network bandwidth. In this paper, we propose a novel view-dependent 3-D model transmission scheme, where a 3-D model is partitioned into a number of segments, each segment is then independently coded using the MPEG-4 3DMC coding algorithm, and finally only the visible segments are selected and delivered to the client. Moreover, we also propose analytical models to find the optimal number of segments so as to minimize the average transmission size. Simulation results show that such a view-based 3-D model transmission is able to substantially save the transmission bandwidth and therefore has a significant impact on wireless graphics applications. Jianfei Cai 0001, Jianmin Zheng, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2007 | Combining Neural Network and Wavelet Transformn for Trigger Asynchrony Detection
Lan Chang, Pau-Choo Chung, Chang Wen Chen |
CIBCB | 3 |
| 2007 | Analysis of Retry Limit for Supporting VoIP in IEEE 802.11e EDCA WLANsabstractWe present in this paper an analysis of retry limit for supporting VoIP in IEEE 802.11e EDCA WLANs. The proposed research considers packet delay as an important factor of QoS performance measure for providing time-bounded service, such as VoIP and video in WLANs which usually implement default parameters. We also present an analytical method based on a Markov chain model that allows us to derive theoretical bounds for retry limit analysis. Based on the Markov chain model that considers retry limit, we can analyze the performance of VoIP over EDCA mechanism. In addition, in order to resolve the bottleneck transmission in VoIP in the downlink of AP, we maintain that the best service can be achieved when the throughputs of uplink and downlink are the same. Such a cross performance point is studied with respect to different retry limit parameters. Finally, extensive simulations have been carried out to validate the analytical model as well as to find out the corresponding cross point. We demonstrate that an enhanced VoIP capacity can be achieved by selecting appropriate retry limit and contention window size parameters. Byung Joon Oh, Chang Wen Chen |
ICCCN | 2 |
| 2007 | A New Network Layer for Mobile Ad Hoc Wireless Networks Based on Assignment Router Identity ProtocolabstractCustomers have complained that malicious nodes are mounting increasingly sophisticated attacking operations in mobile ad hoc networks (MANETs). It is obvious that the IP-based MANETs are vulnerable to attacks and therefore are insecure. In this paper, we design a novel assignment router identity protocol (ARIP) to establish a new layer of network architecture and to take full advantage of Dynamic Hybrid Multi Routing Protocol that we have recently developed for MANETs [2]. A new Identity in the flexible namespace of ARIP enhances the limit in forming prevention line in the security of MANET. A complete architecture is then derived as an instantiation of router identity routing protocol (RIRP) model whose architecture satisfies the condition of ARIP model in order to use this new Identify for routers/hosts. All applications in RIRP deal with this new Identity instead of the vulnerable IP addresses and therefore provide the security embedded seamlessly into the overall network architecture. The proposed RIRP have been implemented on the simulator Glomosim. Simulation results show that RIRP has some great impact on routing performance comparing with DHMRP. Chaorong Peng, Chang Wen Chen |
ICCCN | 2 |
| 2007 | View-Based 3D Model Transmission via Mesh SegmentationabstractFor network-based graphics applications, a key challenge is how to efficiently transmit complex three-dimensional (3D) models over bandwidth-limited communication channels such as wireless links. Most existing 3D mesh coding algorithms do not consider the view-dependent rendering issue, and therefore result in transmitting unnecessary portions of 3D mesh models which leads to the waste in precious network bandwidth. In this paper, we propose a novel view-dependent 3D model transmission scheme, where a 3D model is partitioned into a number of segments, each segment is then independently coded using the MPEG-4 3DMC coding algorithm, and finally only the visible segments are selected and delivered to the client. Such a view-based 3D model transmission is able to substantially save the transmission bandwidth and therefore has a significant impact on wireless network based graphics applications. Jianfei Cai 0001, Jianmin Zheng, Chang Wen Chen |
ICME | 4 |
| 2007 | Distributed Source Coding Under Noisy Transmission EnvironmentsabstractDistributed source coding (DSC) has been attracting more and more research interests recently, particularly for mobile wireless applications. However, most existing DSC schemes assume that the coded information is transmitted error free to the decoder end. In this paper, we report our latest investigations on distributed source coding under noisy transmission environments. It is expected that these investigations provide useful guidelines for practical DSC system design when the transmission is actually not error free. Especially, many current schemes in DSC adopted the channel coding principles to accomplish the source coding tasks. We shall illustrate that the channel coding-based DSC schemes possess some inherent noise combating capability to recover possible channel errors under noisy transmission environments. Specifically, we shall report the investigations of Turbo DSC under various environments: error free, noisy transmission with additional forward error correction (FEC) and noisy transmission without FEC. Our experimental results show that, with the same number of bits transmitted through noise channel, the unprotected Turbo DSC outperforms the FEC protected Turbo DSC. Guogang Hua, Lei Cao 0001, Chang Wen Chen |
ICME | 3 |
| 2007 | QoS Guaranteed Scalable Video Transmission Over MIMO Systems with Time-Varying Channel CapacityabstractIn this paper, we present a novel QoS guaranteed scalable video transmission scheme over multi-input multi-output (MIMO) wireless systems. The proposed scheme is able to not only guarantee the QoS (required BER for base and enhancement layer in compressed video) and but also maximize the system throughput over time-varying MIMO channel. We assume that the estimated channel state information (CSI) at the receiver is available at the transmitter through feedback. In this scenario, under a transmit power constraint, optimal power allocation can be designed to satisfy QoS-guaranteed maximum throughput by a combination of several adaptive operations. These operations include channel adaptive sub-channel selection, power allocation based on water-filling (WF), and data rate maximizing power reallocation. The proposed power reallocation scheme enables us to obtain surplus power after selecting proper adaptive mode and reallocate to other subchannels so as to maximize the data rate. We present in this paper some detailed analysis as well as simulation results to demonstrate that QoS guaranteed video transmission over MIMO wireless systems can indeed be achieved based on scalable video coding and a sequence of adaptive operations. Daewon Song, Chang Wen Chen |
ICME | 2 |
| 2007 | A Transform Domain Classification Based Wyner-Ziv Video CodecabstractIn theory, a Wyner-Ziv video codec should achieve the same efficiency as a joint encoding and decoding one. However, existing approaches still exhibit significant gaps. One main reason is the lack of the complete correlation exploitation between source and side information. In this paper, we propose a transform domain classification based Wyner-Ziv video codec, aiming at exploiting additional video statistics. In this proposed new scheme, the encoder exploits additional statistics by performing block classification to differentiate low motion blocks from high motion ones. In general, low motion blocks represent highly correlated regions. Such information is useful when the decoder performs motion-compensated interpolation to obtain better side information, thus improving the performance of Wyner-Ziv coding. Experimental results show that we are indeed able to achieve better rate-distortion performance compared to the existing Wyner-Ziv video codecs at the expense of some additional complexity in frame store and comparisons at the encoder. Jinrong Zhang 0002, Houqiang Li, Chang Wen Chen |
ICME | 4 |
| 2007 | Joint Source-Channel-Authentication Resource Allocation for Multimedia overWireless NetworksabstractIn our previous work (Li et al., 2006), we have presented unequal authenticity protection (UAP), the methodology of effective protecting multimedia stream transmitted over error-prone wireless networks. In this paper, we extend the previous analysis and consider integrating UAP into the joint source-channel coding (JSCC) framework to achieve optimization of end-to-end quality of media content. We further illustrate the effectiveness of this system using an implementation on progressive JPEG coder. Qibin Sun, Zhi Li 0001, Yong Lian 0001, Chang Wen Chen |
ISCAS | 4 |
| 2007 | Dynamic GOP structure for scalable video codingabstractMPEG is currently developing a new scalable video coding (SVC) standard which should provide compression efficiency similar to current non-scalable video coding standard. The extension of H.264/AVC hybrid video coding using motion-compensated temporal filtering (MCTF) is a current solution. This paper presents a dynamic group of picture (GOP) structure to improve the quality of the SVC coded video by using variable GOP sizes along the video sequences without the restriction of the fixed GOP size or the limited adaptive GOP size. The dynamic GOP structure keeps temporal importance information of frames in the encoded video, which helps in coding efficiency and provides better user perception. An effective GOP size determination method is also proposed in this paper. The proposed scheme has been validated by experimental results. Haksoo Kim, Chang Wen Chen |
VCIP | 2 |
| 2007 | Distributed video coding based on constrained rate adaptive low density parity check codesabstractIn this paper, we present a distributed video coding scheme based on zero motion identification at the decoder and constrained rate adaptive low density parity check (LDPC) codes. Zero-motion-block identification mechanism is introduced at the decoder, which takes the characters of video sequence into account. The constrained error control decoder can use the bits in the zero motion blocks as a constraint to achieve a better decoding performance and further improve the overall video compression efficiency. It is only at the decoder side that the proposed scheme exploits temporal and spatial redundancy without introducing any additional processing at the encoder side, which keeps the complexity of the encoding as low as possible with certain compression efficiency. As a powerful alternative to Turbo codes, LDPC codes have been applied to our scheme. Since video data are highly non-ergodic, we use rate-adaptive LDPC codes to fit this variation of the achievable compression rate in our scheme. We propose a constrained LDPC decoder not only to improve the decoder efficiency but also to speed the convergence of the iterative decoding. Simulation demonstrates that the scheme has significant improvement in the performances. In addition, the proposed constrained LDPC decoder may benefit other application. Rongke Liu, Guogang Hua, Chang Wen Chen |
VCIP | 3 |
| 2007 | Progressive image transmission with RCPT protectionabstractIn this paper, a joint source-channel coding scheme is proposed for progressive image transmission over channels with both random bit errors and packet loss by using rate-compatible punctured Turbo codes (RCPT) protection only. Two technical components which are different from existing methods are presented. First, a data frame is divided into multiple CRC blocks before being coded by a turbo code. This is to secure a high turbo coding gain which is proportional to the data frame size. In the mean time, the beginning blocks in a frame may still be usable although the decoding of the entire frame fails. Second, instead of employing product codes, we only use RCPT, along with an interleaver, to protect images over channels with combined distortion including random errors and packet loss. With this setting, the effect of packet loss is equivalent to randomly puncturing turbo codes. As a result, the optimal allocation of channel code rates is required for the random errors only, which largely reduces the complexity of the optimization process. The effectiveness of the proposed schemes is demonstrated with extensive simulation results. Lei Cao 0001, Chang Wen Chen |
VCIP | 3 |
| 2007 | An Efficient Decoder for Turbo Product Codes with Multi-Error Correcting CodesabstractIn this paper, Chase decoding algorithm for turbo product codes (TPCs) built with multi-error-correcting codes is investigated. A method is proposed to reduce the decoding of test patterns (TPs) as well as the calculation of syndromes and metrics. In the beginning, syndromes and metrics of the first half of TPs are calculated recursively. After decoding of one TP, according to the error bit positions, other TPs that will have the same decoded codeword are identified and rejected from decoding. In the mean time, if these TPs fall into the range of second half of TPs, their syndromes and metrics need not to be calculated either. Simulations results for TPCs built with eBCH(64,51,6,2) or eBCH(64,45,8,3) are presented to demonstrate the efficiency of the proposed method. Moreover, this method that reduces the inherent redundancy in Chase decoding does not cause any degradation in coding gain. Guo Tai Chen, Lei Cao 0001, Lun Yu, Chang Wen Chen |
WCNC | 4 |
| 2007 | UEP Video Transmission Based on Dynamic Resource Allocation in MIMO OFDM SystemabstractThis paper presents a novel unequal error protection (UEP) video transmission scheme based on dynamic resource allocation in a MIMO-OFDM system in order to improve the video transmission performance over wideband wireless channels. In this new scheme, a low density parity code (LDPC) is combined with a complexity reduced power reallocation algorithm to utilize the excess powers which margin from the powers necessary to meet the modulation level selected under a certain BER constraint. Furthermore, by taking full advantage of the residual excess powers, unequal error protection on important bits of video streams is provided by the LDPC code. Accordingly, the transmission efficiency and quality are both increased. Simulation results have shown that the proposed scheme may achieve great improvement in both systems with or without perfect channel estimation. Congchong Ru, Liuguo Yin, Jianhua Lu, Chang Wen Chen |
WCNC | 4 |
| 2007 | Guest Editorial Cross-layer Optimized Wireless Multimedia CommunicationsabstractThe 19 papers in this special issue focus on cross-layer optimized wireless multimedia communications. The papers are organized into four sections: quality of service support for wireless networks; system architecture for multimedia over wireless networks; resource allocation in wireless multimedia communications, and multimedia coding and scheduling issues in wireless networks. Pascal Frossard, Chang Wen Chen, Cormac J. Sreenan, K. P. Subbalakshmi, Dapeng Oliver Wu, Qian Zhang 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2007 | Rate-distortion analysis of leaky prediction based FGS video for constant quality constrained rate adaptation
Jianhua Wu 0003, Jianfei Cai 0001, Chang Wen Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2007 | Single-Pass Rate-Smoothed Video Encoding With Quality ConstraintabstractIn this letter, we study the rate smoothing problem in single-pass video encoding, i.e., given a certain video quality constraint, how to smooth out the traffic rate so that the complexity of the video delivery can be reduced. We apply the low-pass filtering idea, originally proposed for single-pass quality-smoothed video encoding, into the problem of rate-smoothed video encoding. In particular, we use the arithmetic averaging filter to smooth out the rate during single-pass video encoding. Both theoretical analysis and experimental results show that our proposed scheme can not only smooth out the bit rate but also automatically achieve the targeted average video quality. Jianhua Wu 0003, Jianfei Cai 0001, Chang Wen Chen |
IEEE Signal Process. Lett. | 3 |
| 2007 | Message from the Editor-in-Chief
Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Message From the Editor-in-Chief
Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Scalable H.264/AVC Video Transmission Over MIMO Wireless Systems With Adaptive Channel Selection Based on Partial Channel InformationabstractIn this paper, we present a novel joint application physical-layer design (JAPLD) strategy to cost-effectively transmit scalable H.264/AVC video over multi-input multi-output (MIMO) wireless systems. With this approach, the application layer cooperates with the physical layer to maximize the system performance. First, in physical layer, we propose a new layered video transmission scheme over MIMO: adaptive channel selection (ACS). ACS-MIMO is fundamentally different from parallel transmission MIMO (PT-MIMO). While each bit stream is continuously transmitted through a fixed antenna in PT-MIMO, ACS-MIMO is able to periodically switch each bit stream among multiple antennas. In application layer, Scalable Video Coding (SVC) generates layered bit streams that need prioritized delivery. Then, we obtain the ordering of each subchannel's SNR strength as partial channel information (CI) at the receiver. The partial CI is acquired via the estimated channel state information based on training sequences. The JAPLD strategy we developed in this research shall switch the bit stream automatically to match the ordering of SNR strength for the subchannels. Essentially, we will launch higher priority layer bit stream into higher SNR strength subchannel by the proposed JAPLD algorithm. In this fashion, we can implicitly achieve automatic unequal error protection (UEP) for layered SVC transmission over MIMO system without power control at the transmitter. Experimental results show that the proposed ACS-MIMO system is able to achieve UEP with the obtained partial CI and the reconstructed video peak signal-to-noise ratio demonstrate the performance improvement of the proposed system as compared with open loop PT-MIMO system. Daewon Song, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Joint Source-Channel-Authentication Resource Allocation and Unequal Authenticity Protection for Multimedia Over Wireless NetworksabstractThere have been increasing concerns about the security issues of wireless transmission of multimedia in recent years. Wireless networks, by their nature, are more vulnerable to external intrusions than wired ones. Many applications demand authenticating the integrity of multimedia content delivered wirelessly. In this work, we describe a framework for jointly coding and authenticating multimedia to be delivered over heterogeneous wireless networks. We firstly introduce a novel concept called unequal authenticity protection (UAP), which unequally allocate resources to achieve an optimal authentication result. We then consider integrating UAP with specific source and channel-coding models, to obtain optimal end-to-end quality by the means of joint source-channel-authentication analysis. Lastly, we present an implementation of the proposed joint coding and authentication system on a progressive JPEG coder. Experimental results demonstrate that the proposed approach is indeed able to achieve the desired authentication of multimedia over wireless networks Zhi Li 0001, Qibin Sun, Yong Lian 0001, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2006 | Admission Control with Traffic Shaping for Variable Bit Rate Traffic in IEEE 802.11e WLANsabstractWith the increasing popularity of using WLANs for Internet access, the controlled channel access mechanism in IEEE 802.11e WLANs, HCCA, has received much more attentions since its inherent centralized mechanism is more efficient in handling time-bounded multimedia traffics. So far, only a few research works address the admission control problem of variable bit rate (VBR) traffic over HCCA. These existing works consider each traffic flow individually and thus cannot exploit the statistical multiplexing gain among multiple VBR traffic flows. In this paper, we apply the existing statistical multiplexing works to the studied admission control problem with all the features of 802.11e HCCA being taken into consideration. Experimental results show that our proposed admission control achieves significant improvement in network utilization while still satisfying all the QoS requirements. Deyun Gao, Jianfei Cai 0001, Chang Wen Chen |
GLOBECOM | 3 |
| 2006 | Capacity Analysis of Supporting VoIP in IEEE 802.11e EDCA WLANsabstractDirectly implementing voice over Internet protocol (VoIP) over infrastructure wireless local area networks (WLANs) will have the bottleneck problem in the access point (AP). In this paper, we propose to use the service differentiation provided by the new IEEE 802.11e standard to solve the bottleneck problem and improve the voice capacity. In particular, we propose to allocate higher priority access category (AC) to the AP while allocating lower priority AC to mobile stations. We develop a simple Markov chain model, which considers not only the important enhanced distributed channel access (EDCA) parameters but also the channel errors. Based on the developed analytical model, we analyze the performance of VoIP over EDCA. Through appropriately selecting the EDCA parameters, we are able to differentiate the services for the downlink and the uplink. The experimental results are very promising: with the adjustment of only one EDCA parameter, we improve the VoIP capacity by 20~30%. Deyun Gao, Jianfei Cai 0001, Chang Wen Chen |
GLOBECOM | 3 |
| 2006 | Low Punctured Turbo Codes and Zero Motion Skip Encoding Strategy for Distributed Video CodingabstractIn this paper, we present a distributed video coding scheme based on low punctured turbo codes and zero motion encoding strategy. Most recent distributed video coding schemes have a common structure: the encoded bit streams consist of both side information derived from previous frames and syndrome bits derived from channel coding such as Turbo codes. Turbo codes are often punctured to generate scalable compressed video bit streams. We demonstrate in this paper that the generation of side information and the design of the punctured Turbo codes can be integrated to improve the overall coding performance. In the encoding of the video frames, zero motion blocks can be identified and made known to decoder so that the reference frames at the decoder shall have some blocks remain unchanged. While constructing the frame at the decoder, the Turbo decoder may be able to use this information as a constraint to achieve an improved performance. However, we discover a better way of making use of zero motion blocks by skipping the encoding of these blocks. The skip at the encoder will result in the direct reduction of the video coding bit rate as well as lower punctured Turbo codes. Since low punctured Turbo codes perform better, the combined effects on the reduced rate side information and an improved Turbo coding performance leads to an overall performance improvement. Simulation results based on this approach show that integrated low punctured Turbo codes and zero motion skip encoding strategy indeed achieves better performance than the constrained punctured Turbo coding. Furthermore, this scheme can be extended to other channel coding schemes, such as Product Accumulate code, that have been adopted in distributed video coding. Guogang Hua, Chang Wen Chen |
GLOBECOM | 2 |
| 2006 | QoS Guaranteed SVC-based Video Transmission over MIMO Wireless Systems with Channel State InformationabstractThis paper presents a novel scheme for QoS guaranteed video transmission based on scalable video coding (SVC) over multi-input multi-output (MIMO) wireless systems. The proposed scheme is able to not only guarantee the QoS (required BER for base and enhancement layers in compressed video) and but also maximize the system throughput. We assume that the estimated channel state information (CSI) at the receiver is available at the transmitter through feedback. In this scenario, under a transmit power constraint, optimal power allocation can be achieved to satisfy QoS-guaranteed maximum throughput by a combination of several operations. These operations include singular value decomposition (SVD) for decomposing the MIMO channel into equivalent parallel SISO channels, adaptive QAM-modulation, and channel adaptive sub-channel selection for the transmission of base and enhancement layers. We present in this paper with detailed analysis and simulation results to demonstrate that we can achieve QoS optimized video transmission over MIMO wireless systems based on scalable video coding. Daewon Song, Chang Wen Chen |
ICIP | 2 |
| 2006 | Syndrome-Based Light-Weight Video Coding for Mobile Wireless ApplicationabstractIn conventional video coding, the complexity of an encoder is generally much higher than that of a decoder because of operations such as motion estimation consume significant computational resources. Such codec architecture is suitable for downlink transmission model of broadcast. However, in the contemporary applications of mobile wireless video uplink transmission, it is desirable to have low complexity video encoder to meet the resource limitations on the mobile devices. Recent advances in distributed video source coding provide potential reverse in computational complexity for encoder and decoder. In the same spirit, we proposed in this paper a syndrome-based light-weight video encoding scheme for mobile wireless applications. This scheme is based on two innovations: (1) adoption of low resolution low quality reference frames for motion estimation at the decoder; (2) introduction of more powerful product accumulate code. Extensive experimental results have confirmed that this syndrome based encoding can reduce computational complexity at the encoder while maintaining good reconstruction quality at the decoder. Therefore, this light weight video coding scheme is suitable for mobile wireless applications Min Wu 0007, Guogang Hua, Chang Wen Chen |
ICME | 3 |
| 2006 | Dynamic Hybrid Multi Routing Protocol For Ad Hoc Wireless NetworkabstractDynamic hybrid multi routing protocol (DHMRP) is proposed to overcome re-discovered route path to be reply path in traditional routing protocol. The protocol utilizes the reply path based on the hybrid clustering hierarchical establishment of multi routing path to gain an automatic monitoring and repairing broken links in ad hoc networks. And due to reply path and multi routing path shall exist separately in network to gain mitigation traffic "bottlenecks" of ClusterHeads so that improving clusters stability. Performance comparison of DHMRP with AOMDV using Glomosim simulation shows that DHMRP is able to achieve a lower data delay and route discovery ratio and higher packets deliver ratio Chaorong Peng, Chang Wen Chen |
SECON | 2 |
| 2006 | An Attention Based Spatial Adaptation Scheme for H.264 Videos on MobilesabstractWith the growing popularity of personal digital assistant devices and smart phones, consumers have become increasingly enthusiastic to watching videos from these mobile devices. However, when browsing videos in mobiles, users often feel that the display resolution greatly affects their perceptual experience with the limited screen size. In this paper, an attention based spatial video adaptation scheme is proposed to overcome the display constraints by producing and displaying the region of interest. According to the size of the target display, we automatically detect and crop the informative region in each frame to generate a smooth sequence. To avoid costly full encoding operations, we develop a set of transcoding techniques based on the H.264 standard. Experimental results show that this approach not only improves the perceptual quality but also saves the bandwidth and computation, especially for the videos which have not been well edited. Yi Wang 0037, Houqiang Li, Xin Fan 0001, Chang Wen Chen |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2006 | A novel frame-level bit allocation based on two-pass video encoding for low bit rate video streaming applications
Jianfei Cai 0001, Zhihai He, Chang Wen Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2005 | Optimal packetization of fine granularity scalability codestreams for error-prone channelsabstractAn optimal source-channel packetization scheme for MPEG-4 fine granularity scalability (FGS) codestreams is proposed in this paper. The channel is modeled with a uniform error distribution to the enhancement layer transmission. A cost function that models the error expansion for a MPEG-4 FGS stream is derived, and then used in the optimal packetization problem subject to the same overhead as the conventional packetization scheme. An efficient scheme to find the optimal solution is described, which takes time similar to encoding an MPEG-4 FGS codestream. Experiments show that our scheme has up to 1.96 dB gain over the conventional packetization scheme. Bin B. Zhu, Yang Yang 0059, Chang Wen Chen, Shipeng Li 0001 |
ICIP (2) | 3 |
| 2005 | Error Resilient Multiple Description Coding Based on Wavelet Tree Coding and EREC for Wireless NetworksabstractIn this paper, we propose an integrated error resilient MDC scheme for wireless networks with both packet loss and random bit errors. Two descriptions are first generated independently by using index assignment MDSQ. For each description, multiple bitstreams are then generated based on wavelet trees along the spatial orientations. The spatial-orientation trees in the wavelet domain are individually encoded using SPIHT. Error propagation is thus limited within each bitstreams. However, synchronization words are usually needed to avoid error propagation across multiple independent bitstreams. In order to maintain high compression efficiency robust synchronization, we adopt EREC to re-organize these variable-length bitstreams into fixed-length data slots before multiplexing and transmission. Therefore, the synchronization of the start of each bitstream can be automatically obtained at the receiver. Finally, to alleviate the devastating image degradation resulted from errors in the beginning of the bitstreams, we propose an error concealment technique to both constrain the EREC/MDC decoding and post-process the decoded wavelet coefficients Daewon Song, Lei Cao 0001, Chang Wen Chen |
ICME | 3 |
| 2005 | Fine Granularity Scalability Encryption of MPEG-4 FGS BitstreamsabstractIn this paper, we present an encryption scheme for MPEG-4 FGS which provides the same or a little coarser granularity of scalability after encryption. The scheme encrypts compressed data of each video packet or block independently. Initialization vectors are generated with a method to minimize the overhead. The scalability provided in an encrypted codestream using this scheme enables intermediate nodes to truncate an encrypted bitstream at near R-D optimality directly without decryption, which enhances system security. The scheme has virtually negligible overhead, and produces encrypted codestream with virtually the same error resilience performance as the unencrypted case. These features are very desirable in many applications Bin B. Zhu, Yang Yang 0059, Chang Wen Chen, Shipeng Li 0001 |
MMSP | 3 |
| 2005 | Distributed Source Coding in Wireless Sensor NetworksabstractThis paper offers a practical distributed data compression in wireless sensor networks based on convolutional code and turbo code. An improved Viterbi algorithm for distributed source coding (VA-DSC) is proposed to take the advantage of the known parity bits at the decoder. When the algorithm is applied to recursive systematic convolutional (RSC) and turbo codes, it can decrease both the decoding error probability and the computation complexity. Also a scheme of applying distributed source coding to wireless sensor networks is proposed to ensure receiving the data correctly as well as reducing the energy consumption in the networks. Guogang Hua, Chang Wen Chen |
QSHINE | 2 |
| 2005 | Constrained decoding for turbo-CRC code with high spectral efficient modulationabstractIn this paper, we propose a constrained decoding algorithm for turbo-CRC codes with high spectral efficient modulation. We follow a data structure adopted in 3GPP in which the large data frame used for turbo coding consists of multiple CRC packets. First, we modify the turbo encoder by introducing an additional interleaver to permute correctly decoded bits at early stages of decoding over the entire frame. Then, we propose a constrained decoding algorithm at the receiver, making use of the information of correctly decoded bits to help the decoding of other bits and reduce the complexity. Experimental results demonstrate the excellent performance of the entire system. In addition, the error localization property inherent in the system may benefit other applications such as ARQ. Huijun Chen, Lei Cao 0001, Chang Wen Chen |
WCNC | 3 |
| 2005 | Joint source-channel coding of GGD sources with allpass filtering source reshaping
Jianfei Cai 0001, Chang Wen Chen |
J. Vis. Commun. Image Represent. | 2 |
| 2005 | Special issue on visual communication in the ubiquitous era
Chang Wen Chen, Mohammed Ghanbari 0001, King Ngi Ngan |
J. Vis. Commun. Image Represent. | 1 |
| 2005 | Low-pass filtering of rate-distortion functions for quality smoothing in real-time video communicationabstractIn variable-bit-rate video coding, the video is preprocessed to collect sequence-level statistics, which are used for global bit allocation in the actual encoding stage to obtain a smoothed video presentation quality. However, in real-time video recording and network streaming, this type of two-pass encoding scheme is not allowed because the access to future frames and global statistics is not available. To address this issue, we introduce the concept of low-pass filtering of rate-distortion functions and develop a smoothed rate control (SRC) framework for real-time video recording and streaming. Theoretically, we prove that, using a geometric averaging filter, the SRC algorithm is able to maintain a smoothed video presentation quality while achieving the target bit rate automatically. We also analyze the buffer requirement of the SRC algorithm in real-time video streaming, and propose a scheme to seamlessly integrate robust buffer control into the SRC framework. The proposed SRC algorithm has very low computational complexity and implementation cost. Our extensive experimental results demonstrate that the SRC algorithm significantly reduces the picture quality variation in the encoded video clips. Zhihai He, Wenjun Zeng 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2005 | Layered unequal loss protection with pre-interleaving for fast progressive image transmission over packet-loss channelsabstractMost existing unequal loss protection (ULP) schemes do not consider the minimum quality requirement and usually have high computation complexity. In this research, we propose a layered ULP (L-ULP) scheme to solve these problems. In particular, we use the rate-based optimal solution with a local search to find the average forward error correction (FEC) allocation and use the gradient search to find the FEC solution for each layer. Experimental results show that the executing time of L-ULP is much faster than the traditional ULP scheme but the average distortion is worse. Therefore, we further propose to combine the L-ULP with the pre-interleaving to have an improved L-ULP (IL-ULP) system. By using the pre-interleaving, we are able to delay the occurrence of the first unrecoverable loss in the source bitstream and thus improve the loss resilience performance. With the better loss resilience performance in the source bitstream, our proposed IL-ULP scheme is allowed to have a weaker FEC protection and allocate more bits to the source coding which leads to the improvement of overall performance. Experimental results show that our proposed IL-ULP scheme even outperforms the global optimal result obtained by any traditional ULP scheme while the complexity of IL-ULP is almost the same as L-ULP. Jianfei Cai 0001, Chang Wen Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2004 | Layered unequal loss protection with pre-interleaving for progressive image transmission over packet loss channelsabstractMost existing ULP (unequal loss protection) schemes do not consider the minimum quality requirement and usually have high computation complexity. Previously, we proposed a layered ULP (L-ULP) scheme to solve the mentioned problems at the cost of performance degradation. In this paper, we propose to combine the L-ULP with the preinterleaving, which is able to delay the occurrence of the first unrecoverable loss in the source data bitstream while still keeping the original priorities among different layers. Experimental results show that the proposed joint L-ULP and pre-interleaving scheme is able to achieve as good performance as that of the ULP while the complexity is much lower. Jianfei Cai 0001, Chang Wen Chen |
ICIP | 3 |
| 2004 | Lowpass filtering of rate-distortion functions for quality smoothing for real-time video recording and streamingabstractIn this work, we introduce the concept of low-pass filtering of rate-distortion (R-D) functions and develop a smoothed rate control (SRC) framework for real-time video recording and streaming. Theoretically, we prove that using a geometric averaging filter the SRC algorithm is able to maintain a very smooth video presentation quality while achieving the target bit rate automatically. The proposed SRC algorithm has very low computational complexity and implementation cost. Our extensive experimental results demonstrate that the proposed SRC algorithm significantly reduces the picture quality variation in the encoded video clips while matching the encoding bit rate target very accurately. Zhihai He, Chang Wen Chen, Jianfei Cai 0001 |
ICIP | 2 |
| 2004 | A novel video coding scheme for mobile devicesabstractIn this paper, we propose a novel video coding profile for the multimedia applications oriented for mobile wireless communication. Because mobile devices generally have limited computational capability and constrained power consumption, we have considered jointly both video coding efficiency and implementation feasibility when developing the proposed video coding scheme. We achieve a good trade-off between coding efficiency and complexity by minimizing the video reconstruction distortion caused by adopting reduced complexity algorithms. A number of experiments have been conducted and the results have shown the efficiency of the proposed profile. A demo implemented on the Nokia 6600 cell phone demonstrates the feasibility of video coding scheme for mobile devices in wireless communication applications. Yi Wang 0037, Houqiang Li, Chang Wen Chen |
MUM | 3 |
| 2004 | Layered unequal loss protection for progressive image transmission over packet loss channelsabstractIn the past, many schemes have been proposed for progressive image transmission using unequal error protection (UEP) or unequal loss protection (ULP). However, most existing UEP/ULP schemes do not consider the minimum image quality requirement and usually have high computation complexity. In this paper, we propose a layered ULP (L-ULP) scheme for progressive image transmission over packet loss channels, which is able to solve the mentioned problems of existing ULP schemes by smartly choosing the layers. The numerical results show that the proposed L-ULP scheme is quite promising for fast image transmission over packet loss networks. Jianfei Cai 0001, Chang Wen Chen |
VCIP | 3 |
| 2004 | A half D1 MPEG-4 encoder on the BSP-15 DSPabstractIn this paper, we present the work on implementation of a half-D1interlaced MPEG-4 encoder with Equator Technology DSP chip, BSP-15. The BSP-15 DSP consists mainly of a VLIW core, Co-processors, and media I/O interfaces. The encoder utilizes several BSP-15 functional blocks in parallel. In general, the VLIW performs pixel procesing that is computationally intensive. The VLx coprocessor completes variable length coding. Further parallelism is obtained by pre-loading data cache and doubling data buffers. Given the DSP processing power and real time requirements, a complexity control scheme is implemented. A frame-level quantization scheme with quality and rate control is employed. The current implementation for video at 30 fps consumes about 90% of the chip performance at a bit rate ~2Mbps. Lulin Chen, Zhihai He, Chang Wen Chen, Michael A. Isnardi |
VCIP | 3 |
| 2004 | Effective quality analysis for video streaming over wireless ad hoc networkabstractVideo encoding and streaming over wireless ad hoc network operates under severe conditions, such as time-varying channel characteristics with bursty errors, limit power for data transmission, and dynamic topology of the self-organized network, stringent time delay for packet delivery, etc. Due to the dynamic topology, complex mechanism, and time-varying nature of the wireless ad hoc network, the network system and the streaming service often exhibit unpredictable behaviors. The ultimate goal in the wireless video streaming service is to provide the end user with the best possible video presentation quality. The video streaming quality, often measured by the end-to-end picture distortion, is affected by the scene coding characteristics S of the input video, and the configuration of the network parameters N which includes bandwidth, transmission power, bit error ratio, delay, etc. This brings up the following important and challenging issue: given an input video with scene characteristics S and a network configuration N, what is the average video streaming quality the receiver could expect the system to provide. To address this issue, in this work, we examine the behaviors and constraints of the major components in the streaming system, and propose an across-layer framework to model, control, and optimize the end-to-end video streaming quality. Zhihai He, Chang Wen Chen |
VCIP | 2 |
| 2004 | Collaborative image transmission over wireless sensor networksabstractThe imaging sensors are able to provide intuitive visual information for quick recognition and decision. However, imaging sensors usually generate vast amount of data. Thus, processing of image data collected in the sensor network for the purpose of energy efficient transmission poses a significant technical challenge. In particular, when a cluster of imaging sensors is activated to track certain moving target, multiple sensors may be collecting similar visual information simultaneously. With correlated image data, we need to intelligently reduce the redundancy among the neighboring sensors so as to minimize the energy for transmission, the primary source of sensor energy consumption. We propose in this paper a novel collaborative image transmission scheme for wireless sensor networks. First, we apply a shape matching method to coarsely register images to find out maximal overlap in order to exploiting the spatial correlation between images acquired from neighboring sensors. A transformation is generated according to the matching results. We encode the original image and the difference between the transformed image and reference image. Then, we transmit the coded bit stream together with the transformation parameters. This will significantly reduce the transmission energy comparing with transmitting two individual images independently. To exploiting the temporal correlation among images in the same sensor, we assume that the imaging sensors and the background scenes remain stationary over the data acquisition period. For a given image sequence, we transmit background image only once. A simple background subtraction method is employed to detect targets. Whenever targets are detected, only the regions of target and their spatial locations are transmitted to the monitoring center. At the monitoring center, the whole image can be reconstructed by fusing the background and the target image as well as its spatial location to further reduce energy consumption. Experimental results show that the transmission energy can be greatly reduced. Min Wu 0007, Chang Wen Chen |
VCIP | 2 |
| 2004 | Constrained SOVA decoding in concatenated codesabstractIn practical communication systems, error detection and correction codes such as cyclic redundancy check (CRC) codes and turbo codes (TCs) are almost always concatenated. While TCs are used to remove errors, CRC codes may be applied to applications such as ARQ, early stop of turbo decoding. Considering the scenario that a large data frame for TC consists of multiple short packets, each with CRC parity bits, we propose a constrained-SOVA (C-SOVA) algorithm to fully take advantage of the interim CRC detection results to redesign the TC decoding algorithm. One more interleaver is added at the encoder so that bits in a correctly decoded packet are permuted to the entire range of frame. At the decoder, the C-SOVA has been designed that includes constraining trellis, updating the extrinsic information, and reducing the number of tracing back. As a result, not only a the decoding complexity is significantly reduced, but a high coding gain has been preserved due to elimination of certain error patterns as well. Simulations have been presented to demonstrate the excellent features of the proposed scheme. Although SOVA algorithm and concatenated codes are considered in this paper, the proposed idea using constraints is applicable to any trellis-based decoding algorithms including maximum a posterior probability (MAP) algorithm, and to any scenarios that partial bits are known correct in the course of turbo decoding. Lei Cao 0001, Chang Wen Chen |
WCNC | 2 |
| 2004 | A novel algorithm for computer-assisted measurement of cervical length from transvaginal ultrasound imagesabstractThe cervical length measured by transvaginal ultrasound is a proven clinical tool for predicting premature birth. The standard manual measurement of the cervix is limited by variability in the technique. In this research, we develop the first computer algorithm that is able to identify the anatomic landmarks of the cervix on a transvaginal ultrasound image and determine the standard cervical length. The system is composed of four stages: The first stage is adaptive speckle suppression using variable length sticks algorithm. The second stage is the location of the internal cervical opening or "os" using a region-based segmentation. The third stage is delineation of the cervical canal. The fourth stage uses gray level summation patterns and prior knowledge to first localize the tissue boundary of the external cervix, and then use a template to determine the specific location of the external os. The cervical length is determined and calculated to image scale. To validate the proposed algorithm, 101 cervical ultrasound images were selected from a series of 37 examinations performed on 17 patients over an eight-month period. Repeated measurements of cervical length using the computer-assisted method were compared with those carried out by two experienced sonographers. The median intraobserver variability for the 101 images using the computer-assisted method was significantly smaller than that of the manual method by either sonographer. In a pairwise comparison, the mean cervical length for the computer method matches with the mean manual cervical length. Min Wu 0007, Robert F. Fraser II, Chang Wen Chen |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2003 | Modeling of subband coefficients for clustering-based adaptive quantization with spatial constraints
Jiebo Luo 0001, Chang Wen Chen |
J. Vis. Commun. Image Represent. | 2 |
| 2003 | Guest Editorial
Yücel Altunbasak, Chang Wen Chen, M. Reha Civanlar, King Ngi Ngan |
Signal Process. Image Commun. | 2 |
| 2003 | A novel product coding and recurrent alternate decoding scheme for image transmission over noisy channelsabstractIn this letter, we present a novel product channel coding and decoding scheme for image transmission over noisy channels. Two convolutional codes with at least one recursive systematic convolutional code are employed to construct the product code. Received data are decoded alternately in two directions. A constrained Viterbi algorithm is proposed to exploit the detection results of cyclic redundancy check codes so that both reduction in error patterns and fast decoding speed are achieved. Experiments with image data coded by the algorithm of set partitioning in hierarchical trees exhibit results better than those currently reported in the literature. Lei Cao 0001, Chang Wen Chen |
IEEE Trans. Commun. | 2 |
| 2002 | Optimal bit allocation for low bit rate video streaming applicationsabstractCurrent rate control schemes in video coding standards do not have efficient frame-level bit allocation because of the inherent constraints in real-time encoding. In this paper, we assume an offline video encoding environment and proposed a rate control scheme based on optimal bit allocation for low bit rate streaming applications. Specifically, we apply a /spl rho/-domain rate-distortion (R-D) model, originally applied at macroblock (MB) level, to frame-level. Based on this frame-level R-D model and a two-pass encoding method, we are able to allocate bits among video frames in an optimal way so that video sequences can be coded at low bit rate with an improved quality. Experimental results demonstrate the proposed scheme is able to achieve not only noticeable reduction in average distortion but also a more consistent and smoother visual quality. Jianfei Cai 0001, Zhihai He, Chang Wen Chen |
ICIP (1) | 3 |
| 2002 | Rate-reduction transcoding design for wireless video streamingabstractThis paper presents two types of techniques suitable for rate-reduction transcoding for wireless video streaming applications. We begin this paper by reviewing existing approaches and addressing several issues related to transcoding. Next, we describe the first type of transcoding based on intra refresh architecture for spatial resolution reduction. Then, we present the second type of transcoding scheme based on rate-distortion (R-D) characteristics of the pre-encoded video. This R-D based approach can be applied to architecture simplification, rate control, frame dropping control, and channel adaptive transcoding. We conclude this paper by pointing out that transcoding is an integral part of wireless video streaming because it provide a flexible interface between the wired network and the wireless network. Anthony Vetro, Chang Wen Chen |
ICIP (1) | 2 |
| 2002 | Video coding using joint temporal-spatial compensationabstractIn motion compensation (MC) based video coding system, macroblock data is always predicted by its temporal neighborhood in its previous frame. In this work, we observe that the macroblock data can often be better predicted by its spatial neighborhood than its temporal neighborhood, especially at low coding bit rates. Based on this observation, we propose a joint temporal-spatial compensation (JTSC) scheme for video coding, with the conventional MC being its subset. Our experimental results demonstrate that the coded picture quality is significantly improved by this novel JTSC scheme. Zhihai He, Chang Wen Chen |
ICME (1) | 2 |
| 2002 | End-to-end video quality analysis and modeling for video streaming over IP networkabstractWe derive an analytical end-to-end video quality prediction and control model for video streaming over an IP network. This model maps the quality of service (QoS) parameters defined at the connection level to actual video presentation quality at the receiver end. Specifically, we provide an analysis on the effects of packet loss on the decoded video quality and relate the packet loss ratio, one of the most important QoS parameters at the connection level, to the actual video quality. The model also characterizes the rate-distortion (R-D) behavior of the encoder at the video sequence level, predicting the average encoding quality for a given channel bandwidth. Our experimental results demonstrate the accuracy of the proposed video quality model. Zhihai He, Chang Wen Chen |
ICME (1) | 2 |
| 2002 | Two-pass video encoding for low-bit-rate streaming applications
Jianfei Cai 0001, Chang Wen Chen |
VCIP | 2 |
| 2002 | Encoder-based rate smoothing and quality control for low-delay video coding and communication
Zhihai He, Chang Wen Chen |
VCIP | 2 |
| 2002 | ?-domain optimum bit allocation and accurate rate control for DCT video coding
Zhihai He, Chang Wen Chen |
VCIP | 2 |
| 2002 | Blind channel estimation and equalization using Viterbi algorithmsabstractIn this paper, a trellis based blind channel estimation and equalization technique is presented. First, blind channel estimation is accomplished by incorporating a list parallel Viterbi algorithm with the least mean square (LMS) updating approach. In this operation, multiple trellis mappings are preserved and ranked in terms of path metrics. Equivalently, multiple channel estimates are maintained and updated once a single symbol is received. Second, at a certain point, which is determined by the evolution of path metrics and the linear constraint possessed in the trellis mapping, only the best channel estimate is maintained. Third, this channel estimate is adopted to construct the whole trellis that is used by a conventional adaptive Viterbi algorithm. Signal detection and further channel updating are then conducted alternatively. To alleviate the noise impact, a small delay is introduced before the feedback of detected symbols to further update the channel estimate. Simulation has shown the overall good performance of the proposed scheme in terms of mean square error (MSE) convergence of the channel estimation, stability to the initial channel guess, and computational complexity etc. Lei Cao 0001, Chang Wen Chen, Philip V. Orlik, Jinyun Zhang, Daqing Gu |
VTC Spring | 2 |
| 2002 | A high-performance and low-complexity video transcoding scheme for video streaming over wireless linksabstractVideo streaming over wireless links involves two basic needs: rate reduction transcoding and error control channel coding. For traditional transcoding systems, the performance is usually proportional to the system complexity. In this paper, we propose a novel transcoding scheme, which can achieve better performance with reduced system complexity. Coupled with a simple joint source-channel bit allocation approach, the proposed transcoding can accurately throttle the source coding rate so that there is sufficient bandwidth left for channel coding to correct channel errors in wireless links. Simulation results demonstrate that the proposed video streaming system can achieve a good performance without much increasing in system complexity. Jianfei Cai 0001, Chang Wen Chen |
WCNC | 2 |
| 2002 | Multiple Hierarchical Image Transmission over Wireless Channels
Lei Cao 0001, Chang Wen Chen |
J. Vis. Commun. Image Represent. | 2 |
| 2002 | Special Issue on Recent Advances in Wireless Video SIGNAL PROCESSING: IMAGE COMMUNICATION
Yücel Altunbasak, Chang Wen Chen, M. Reha Civanlar, King Ngi Ngan |
Signal Process. Image Commun. | 2 |
| 2002 | Introduction to the special issue on wireless communication
Chang Wen Chen, Reginald L. Lagendijk, Amy R. Reibman, Wenwu Zhu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Joint source channel rate-distortion analysis for adaptive mode selection and rate control in wireless video codingabstractWe first develop a rate-distortion (R-D) model for DCT-based video coding incorporating the macroblock (MB) intra refreshing rate. For any given bit rate and intra refreshing rate, this model is capable of estimating the corresponding coding distortion even before a video frame is coded. We then present a theoretical analysis of the picture distortion caused by channel errors and the subsequent inter-frame propagation. Based on this analysis, we develop a statistical model to estimate such channel errors induced distortion for different channel conditions and encoder settings. The proposed analytic model mathematically describes the complex behavior of channel errors in a video coding and transmission system. Unlike other experimental approaches for distortion estimation reported in the literature, this analytic model has very low computational complexity and implementation cost, which are highly desirable in wireless video applications. Simulation results show that this model is able to accurately estimate the channel errors induced distortion with a minimum delay in processing. Based on the proposed source coding R-D model and the analytic channel-distortion estimation, we derive an analytic solution for adaptive intra mode selection and joint source-channel rate control under time-varying wireless channel conditions. Extensive experimental results demonstrate that this scheme significantly improves the end-to-end video quality in wireless video coding and transmission. Zhihai He, Jianfei Cai 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2002 | Content-based multiple bitstream image transmission over noisy channelsabstractIn this paper, we propose a novel combined source and channel coding scheme for image transmission over noisy channels. The main feature of the proposed scheme is a systematic decomposition of image sources so that unequal error protection can be applied according to not only bit error sensitivity but also visual content importance. The wavelet transform is adopted to hierarchically decompose the image. The association between the wavelet coefficients and what they represent spatially in the original image is fully exploited so that wavelet blocks are classified based on their corresponding image content. The classification produces wavelet blocks in each class with similar content and statistics, therefore enables high performance source compression using the set partitioning in hierarchical trees (SPIHT) algorithm. To combat the channel noise, an unequal error protection strategy with rate-compatible punctured convolutional/cyclic redundancy check (RCPC/CRC) codes is implemented based on the bit contribution to both peak signal-to-noise ratio (PSNR) and visual quality. At the receiving end, a postprocessing method making use of the SPIHT decoding structure and the classification map is developed to restore the degradation due to the residual error after channel decoding. Experimental results show that the proposed scheme is indeed able to provide protection both for the bits that are more sensitive to errors and for the more important visual content under a noisy transmission environment. In particular, the reconstructed images illustrate consistently better visual quality than using the single-bitstream-based schemes. Lei Cao 0001, Chang Wen Chen |
IEEE Trans. Image Process. | 2 |
| 2002 | Special issue: Multimedia over mobile IPabstractThe rapid advance in wireless communication and Internet has ushered in a new era of mobile/wireless multimedia applications and services. Tremendous growth opportunities lie ahead of us as academia, industry and government agencies make giant strides in the development of Multimedia Over Mobile IP. Apart from the marketing and economic impacts, the convergence of wireless communication, Internet, and multimedia is creating a new paradigm of research and development in order to deliver multimedia content over Internet and mobile wireless networks. With the benefit of increased bandwidth in such networks, newer and more powerful mobile devices, and continued effort towards standardization, wireless communication is quickly moving beyond voice and text-based emails into the reality of multimedia content delivery and access at anytime and from anywhere. In this Special Issue, we are extremely pleased to bring together a team of the leading experts from academia, industry and government agencies to provide an in-depth, comprehensive overview of the rapidly evolving field of Multimedia Over Mobile IP. Mobile person to person speech communication has turned out to be extremely popular with more than one billion users at present. Today, mobile networks are evolving quickly to support multimedia communications. This Special Issue opens with a paper on the latest products and services for carrying Multimedia Over Mobile IP. The paper ‘Wireless meets multimedia’ by Jukka Yrjänäinen and Yrjö Neuvo provides an excellent overview of the wireless multimedia products and services in coming years. The access to various sources of multimedia will further enrich mobile communication; new mobile phones with graphical color displays and integrated cameras will provide a natural platform for large sets of multimedia applications; and open software architectures in these terminals are giving new possibilities to the developer community. Clearly, evolution of the mobile multimedia will create new challenges and exciting opportunities for related research and development activities. 3G has been a buzzword in the past couple of years. With the development of the 3G network infrastructure, emerging wireless communication standards and products, wireless multimedia applications and services are imminent and are poised to significantly change the way people live around the world. In their paper ‘3G wireless multimedia: technologies and practical issues’, Wenjun Zeng and Jiangtao Wen present an overview of the emerging wireless communication standards, end-to-end wireless streaming systems, and relevant wireless multimedia technologies. The paper highlights some of the challenges in the deployment of 3G wireless multimedia services, using an commercially available solution as an example. Video streaming is a highly demanding multimedia application for wireless channels. Bernd Girod, Mark Kalman, Yi Liang and Rui Zhang review recent advances in channel-adaptive video streaming in their paper ‘Advances in channel-adaptive video streaming’. Adaptive media playout at the client can be used to reduce receiver buffering and therefore average latency, and provide limited rate scalability. Rate-distortion optimized packet scheduling determines the best packet to send given the distortion reduction associated with sending that packet, interpacket dependencies, and the success of past transmissions. Channel-adaptive packet dependency control can greatly improve the error-robustness of streaming video and reduce or eliminate the need for packet retransmissions. Three architectures are considered for wireless video streaming, along with discussions on the utility of the related techniques for each architecture. A very important aspect of multimedia experience is perception and interaction. Interaction is desired for playing games on wireless devices. At present, most multimedia experience is delivered and enjoyed in a passive mode. Between audio and video, we argue that the more important perception is visual perception. As consumers move beyond the initial “wow” stage, the quality of the multimedia experience becomes increasingly important. This often over-looked aspect is addressed by a unique paper in this Special Issue on ‘Displaying images on mobile devices: capabilities, issues, and solutions’ by Jiebo Luo, Amit Singhal, Gustav Braun, Robert Gray, Nicolas Touchard and Olivier Seignol. Indeed, wireless imaging is enabling visual communication at any time and from anywhere to become a reality. However, a key technical challenge is how to achieve best-perceived image quality given the limited screen size and display bit depth of the mobile devices. In this paper, the authors present an overview of the current capabilities of various mobile devices, highlight some of the technical issues, and present potential solutions. In addition, to help sort through the myriad of commercial solutions and anticipate what is to come, a review of some of the major software products on the market and an outlook of the trend towards more capable devices are provided. In many applications such as construction, manufacturing, ground robotic vehicles, and rescue operations, there are many issues that necessitate the capability of transmitting digital video and that such transmissions should be performed wirelessly and in an ad-hoc manner. In ‘IEEE 802.11 FHSS receiver design for cluster-based multihop video communications’, Koichiro Ban and Hamid Gharavi proposed an ad-hoc, cluster-based, multi-hop network architecture for video communications. For implementation, the IEEE 802.11 FHSS wireless LAN system using 2GFSK modulation has been deployed. To help analyze ways to enhance the overall throughput rate for higher quality video communications, a performance evaluation of the IEEE 802.11 FHSS was conducted when 4GFSK modulation option is selected. It was found that the 2 Mb/s system utilizing 4GFSK modulation is not very efficient in terms of RF range. Therefore, to improve its performance for multihop applications, a combination of diversity and non-coherent Viterbi based receiver is considered. For the video transmission part, the team has considered a bitstream splitting technique together with a packet-based error protection strategy to combat packet drops under multipath fading conditions. Finally, the paper presents the simulation results, including the effects of the receiver design and diversity on the quality of the received video signals. Transport of multimedia content over mobile networks is challenged by the error-prone nature of wireless channels. This may result in loss or erroneous decoding of the video. In their paper, ‘Second-generation error concealment for video transport over error-prone channels’, Trista Pei-chun Chen and Tsuhan Chen review different error concealment methods and introduce a new framework, which can be considered as second-generation error concealment. All the error concealment methods reconstruct the lost video content by making use of some a priori knowledge about the video content. First-generation error concealment builds such a priori in a heuristic manner. The proposed second-generation error concealment builds the a priori by modeling the statistics of the video content. Context-based models are trained with the correctly decoded video content, and then used to replenish the lost video content. Trained models capture the statistics of the video content and thus reconstruct the lost video content better than reconstruction by heuristics. Bandwidth and the diversity in bandwidth have been major limiting factors for transmitting multimedia content across mobile wireless networks. Transcoding becomes a necessity in order to deal with channels with different bandwidth capabilities. Two types of techniques suitable for rate reduction transcoding for wireless video streaming applications are presented by Anthony Vetro, Jianfei Cai and Chang Wen Chen in ‘Rate-reduction transcoding design for wireless video streaming’. They begin by reviewing existing approaches and addressing several issues related to transcoding. Next, they describe the first type of transcoding based on intra refresh architecture for spatial resolution reduction and the second type of transcoding scheme based on rate-distortion (R-D) characteristics of the pre-encoded video. This R-D based approach can be applied to architecture simplification, rate control, frame dropping control, and channel adaptive transcoding. This paper concludes by pointing out that transcoding is an integral part of wireless video streaming because it provides a flexible interface between the wired network and the wireless network. Joint source-channel coding schemes have been proven to be very effective for reliable multimedia communications. Numerous joint source-channel coding approaches have been proposed to address the important issue on how to allocate limited bit budget between source and channel coding. In their paper entitled “Combined hidden Markov source estimation and low density parity-check coding: a novel joint source–channel coding scheme for multimedia communications,” Liuguo Yin, Jianhua Lu, and Youshou Wu describe a novel joint source-channel coding scheme for multimedia communications. This approach combines the hidden Markov source estimation and the low-density parity-check (LDPC) codes with an iterative estimation/decoding scheme. With this innovative combination, multimedia source redundancy could be accurately extracted by the hidden Markov estimation without a priori information about the source. Moreover, the interleaver that is usually used to separate the source coding and channel coding can be avoided by exploiting the randomizing property of the LDPC codes. Furthermore, the channel decoding procedure may be implemented in parallel, resulting in good performance with a fairly low decoding complexity and delay. Simulation results have shown that the proposed scheme can achieve much better performance than that of the standard coding scheme over the binary input additive white Gaussian noise channels. The Guest Editors of this Special Issue would like to thank all the authors for submitting their excellent work using their best efforts in making sure the successful and timely publication of this Special Issue. We would also like to thank Professor Mohsen Guizani, Editor-In-Chief of Wireless Communications and Mobile Computing for his vision and encouragement, and Laura Kempster and Claire Bailey of John Wiley & Sons Ltd. editorial office for their support throughout the entire publication process of this Special Issue. Chang Wen Chen, Jiebo Luo 0001 |
Wirel. Commun. Mob. Comput. | 1 |
| 2002 | Rate-reduction transcoding design for wireless video streamingabstractAbstract Owing to the heterogeneity existing in wireless video streaming, such as terminal capabilities, network conditions, user preferences, and the natural environment in which a user is located, rate‐reduction transcoding is usually necessary to adjust the coding rate to the available channel bandwidth. We begin this paper by reviewing existing approaches and addressing several issues related to transcoding. Next, two types of rate‐reduction coding techniques are studied. The first type is based on standard coding schemes and specifically considers transcoding from MPEG‐2 to MPEG‐4 with a reduced spatial resolution. A method for drift compensation based on an intrarefresh technique is also presented. The second type of scheme is a rate‐reduction transcoder that makes use of frame‐level R‐D information that has already been extracted at the server before transmission. This R‐D‐based approach can be applied to architecture simplification, rate control, frame dropping control, and channel adaptive transcoding. We conclude this paper by pointing out that transcoding is an integral part of wireless video streaming because it provides a flexible interface between the wired network and the wireless network. Copyright © 2002 John Wiley & Sons, Ltd. Anthony Vetro, Jianfei Cai 0001, Chang Wen Chen |
Wirel. Commun. Mob. Comput. | 3 |
| 2001 | FEC-Based Wireless Video Streaming with Pre-Interleaving
Jianfei Cai 0001, Chang Wen Chen |
Data Compression Conference | 2 |
| 2001 | Use of pre-interleaving for video streaming over wireless access networksabstractWe present an FEC-based end-to-end error control scheme for video streaming over wireless access networks, considering both bit errors and packet-loss. We propose a novel robust video streaming system in which an interleaving is applied to the compressed bitstream before channel coding. The application of such pre-interleaving is able to significantly improve the error-combating performance of video streaming because the adoption of this pre-interleaving can simultaneously satisfy different requirements arising from both channel coding and source coding. Experimental results demonstrate the improved performance of the proposed pre-interleaving scheme, especially in the case of highly bursty and regular channel errors. Jianfei Cai 0001, Chang Wen Chen |
ICIP (1) | 2 |
| 2001 | Robust image transmission based on wavelet tree coding and ERECabstractWe propose a robust image transmission scheme based on wavelet tree coding and error resilient entropy coding (EREC). After the wavelet decomposition, the coefficients are re-arranged into wavelet trees according to their spatial representation. Each wavelet tree is coded independently so that the scheme is essentially a block-based coding in the spatial domain. Since the self-similarity across subbands is preserved, a high source coding efficiency can be achieved. EREC is then adopted to enhance the error resilience capability of the compressed bitstream. At the receiving end, the automatic re-synchronization of each block is obtained. Furthermore, the bits impacted by the error propagation are more likely located in the low bit-layers so that the resultant degradation is less severe in terms of PSNR. Experimental results show that this proposed scheme can achieve a good error resilient performance. Lei Cao 0001, Chang Wen Chen |
ICIP (3) | 2 |
| 2001 | Synthesis of directional texture based on multiresolution block sampling and constrained block movementabstractThe synthesis of directional texture is particularly challenging. We present a novel directional texture method based on our previously proposed multiresolution block sampling (MBS) and constrained block movement. First, we estimate the dominant direction of a given directional texture. Using a modified random movement-control set, we incorporate an additional directional constraint into our basic method to produce a new directional texture. A number of directional textures are used to show the effectiveness of this extended method. Jiebo Luo 0001, Chang Wen Chen |
ICIP (2) | 3 |
| 2001 | VBR transcoding architecture for video streaming
Chang Wen Chen |
VCIP | 2 |
| 2001 | A Novel Fast Block Motion Estimation Algorithm Based on Combined Subsamplings on Pixels and Search Candidates
Chang Wen Chen |
J. Vis. Commun. Image Represent. | 3 |
| 2001 | Uniform threshold TCQ with block classification for image transmission over noisy channelsabstractA combined source-channel coding scheme without explicit error protection is proposed to transmit images over noisy channels. Major components of the proposed coding scheme include 2-D DCT with block classification, fixed-length uniform threshold trellis coded quantization, optimal bit-allocation algorithm, and noise reduction filters. The integration of these components allows us to organize the compressed bitstream in such a way that it is less sensitive to channel noise, and hence achieves data compression and error resilience at the same time. This paper reports our previous study by incorporating the block classification into the integrated scheme. Experimental results show that, in the case of noise-free channels and at the bit rate of 0.5 bpp, an improvement of 2.33 dB can be achieved with the classification. In the case of noisy channels, the gain decreases as the bit error rate increases. However, we can still achieve an average improvement of 0.46 dB, even for highly noisy channels with BER=0.1. Our proposed system uses no error protection, no synchronization codewords and no entropy coding. However, it shows a decent compression ratio and graceful degradation with respect to increasing channel errors. Jianfei Cai 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | Robust image transmission with bidirectional synchronization and hierarchical error correctionabstractWe present a novel joint-source and channel image-coding scheme for noisy channel transmission. The proposed scheme consists of two innovative components: (1) intelligent bidirectional synchronization and, (2) layered bit-plane error protection. The bidirectional synchronization is able to recover the coding synchronization when any single or even when two consecutive synchronization codes are corrupted by the channel noise. With synchronized partition, unequal error protection for each bit plane can be designed based on the analysis of bit-plane error sensitivity for compressed image data transmission over noisy channels. Experimental results over extensive channel simulations show that the proposed scheme outperforms the well-known approach proposed by Sherwood and Zeger (see IEEE Signal Processing Lett., vol.4, p.189-91, 1997). Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | A novel two-pass VBR coding algorithm for fixed-size storage applicationabstractWe propose a novel variable bit rate (VBR) algorithm that adopts two-pass encoding to implement high-performance coding for fixed-size storage applications. In the first-pass encoding, we use the original frame as the reference for motion estimation. This way, we are able to obtain the exact relationship between the amount of bits generated for each frame (rate) and each possible quantization factor (Q) so that an R-Q function can be built. The corresponding computational expense is virtually the same as a typical VBR encoder. With this R-Q function, we can optimize the quantization factor for each frame throughout the entire sequence based on the constraints in total storage size and decoder buffer size before the second-pass encoding is executed. This procedure not only guarantees full use of the fixed storage size for the compressed bitstream, but also prevents the decoder buffer from underflowing. Unlike a typical constant bit rate (CBR) decoder, this VBR decoder is capable of preventing the buffer from overflowing by introducing the pausing operation at some designated time. Because the quantization factor optimization is a simple table lookup, the computational requirement is minimized. In the second-pass encoding, the estimated quantization factors for all the frames through the above procedures is used to encode the entire video sequence as a free VBR encoder. Experimental results on real video sequences show that the proposed VBR encoding can provide more consistent visual quality and an improved coding efficiency. The application of the proposed VBR coding scheme includes video streaming over the Internet, digital versatile disk, digital library, and video on demand. Yiliang Wang, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2000 | Product Code and Recurrent Alternative Decoding for Wireless Image TransmissionabstractSummary form only given. We have developed a novel channel coding scheme for image transmission over wireless channels. A key component of this scheme is the construction of the product code using a convolutional code and a recursive systematic code (RSC) in the horizontal and vertical directions, respectively. High performance transmission has been achieved through an innovative recurrent alternative decoding scheme based on constrained Viterbi decoding. We used RSC because of its systematic structure so that we did not need to transmit the original data portion of the RSC coding outcome since this portion can be obtained from the other direction of channel decoding. At the receiving end, the data will at first be decoded along the horizontal direction. After decoding, the position of a correctly decoded row will be assigned a positive sign and be preserved. When we decode along the vertical direction, the positive signs after horizontal decoding will be used as constraints in the Viterbi decoding algorithm to force the survival path's passing through these known positions of correctly decoded bits. The constrained Viterbi algorithm using known correct bits enables us to derive likely correct values for other bits. This reduces greatly possible paths along the trellis diagram even though some or the paths dropped may have smaller path cost due to channel noise corruption and would otherwise be preserved by the unconstrained Viterbi algorithm. Moreover, this also reduces significantly the comparisons between the branches and the received symbols. Similarly the use of the correctly localized columns are also preserved and used again to refine the next round of horizontal decoding. Such recurrent decoding will alternate between horizontal and vertical directions until all data is correctly decoded or until there is no more improvement in both directions. We investigated the performance of the proposed scheme for both BSC and GEC channels and report the results here. Lei Cao 0001, Chang Wen Chen |
Data Compression Conference | 2 |
| 2000 | A Novel Product Coding and Decoding Scheme for Wireless Image TransmissionabstractThis paper presents a novel generic channel coding scheme for image transmission over wireless channels. The product code is constructed using one classical convolutional code in one direction and one recursive systematic code (RSC) in the other direction. At the receiving end, the data are alternatively decoded along the two directions. In each iteration, any correctly decoded data in one direction will be used to constrain the decoding along the other direction. The constrained Viterbi decoding algorithm is designed in such a way that the survival paths are forced to pass through the positions where the bits have already been correctly decoded. With such a constraint, we achieve a reduction of candidate paths for the maximum likelihood (ML) decoding as well as a speed up of convergence. Experimental results show that the proposed scheme is able to outperform all existing schemes under some random and bursty error channel conditions. Lei Cao 0001, Chang Wen Chen |
ICIP | 2 |
| 2000 | A High Performance VBR Coding Algorithm for Fixed Size Storage ApplicationsabstractWe propose a novel variable bit rate (VBR) algorithm that adopts two-pass VBR encoding to implement high performance coding for fixed size storage applications. In the first pass encoding, we use the original frame as the reference frame to compute motion estimation so that the exact relationship between the amount of bits generated for each frame (Rate) and each possible quantization factor (Q) can be obtained. With this R-Q function, we are able to optimize the quantization factor throughout an entire sequence according to the constraints of the storage size and the decoder buffer size before the second pass encoding is executed. In addition, the decoder is able to prevent the buffer from overflowing by pausing the operation of reading the bit stream into the buffer at some specific time. In the second pass encoding, the selected quantization factors for all the frames will be used to encode the entire video sequence as a free VBR encoder. Experimental results show that the proposed VBR encoding can provide more consistent visual quality as well as significantly improved coding efficiency. Yiliang Wang, Chang Wen Chen |
ICIP | 4 |
| 2000 | An FEC-based error control scheme for wireless MPEG-4 video transmissionabstractIn this paper, we propose an FEC-based error control scheme for the transmission of low bit rate MPEG-4 video over wireless channels. The proposed error control scheme, through integrating with MPEG-4 inherent error resilient techniques, divides the MPEG-4 bitstream into several classes. Such division of the bitstream not only facilitates unequal error protection for different classes of data but also allows local reorganization of the bitstream into a fixed-length structure. This reorganization enables robust decoding at the receiving end, however, results in little side information and causes negligible delay. The proposed scheme can also be combined with motion-compensation based error concealment for effective post processing. Experimental results demonstrate that the performance of the proposed scheme is much better than that of the simple equal error protection scheme, not only in PSNR but also in visual quality, in terms of the reconstructed video frames at the receiver. Jianfei Cai 0001, Qian Zhang 0001, Wenwu Zhu 0001, Chang Wen Chen |
WCNC | 4 |
| 2000 | Scene adaptive multiple coding scheme for robust image transmissionabstractWe propose a combined source and channel coding scheme for image transmission over noisy channels. The key component is extracting and preserving the scene information and incorporating it with unequal error protection to combat the channel errors. After hierarchical wavelet decomposition of the image, wavelet coefficients with a parent-child relationship are grouped into wavelet blocks and classified according to corresponding scenes in the original image. For each class, spatial neighborhood coefficients in the high frequency subbands are constrained so that the spatially isolated coefficients are removed and clustered coefficients are retained at the same time. All the wavelet blocks in the same class are grouped together and coded using the set partitioning in hierarchical trees (SPIHT) algorithm. High source coding efficiency is preserved even though multiple source coded bitstreams are generated since the wavelet tree structure is intact. In order to combat the channel errors, an unequal error protection strategy implemented by RCPC/CRC channel coding is designed based on the bit contribution to both PSNR and human visual sensitivity. Finally, a postprocessing method is developed at the receiving end to restore the degradation due to the residual error after channel decoding. Experimental results show that the proposed scheme is indeed able to provide protection for more important bits and more important visual content under a noisy transmission environment. In particular, the reconstructed images illustrate consistently better visual quality. Lei Cao 0001, Chang Wen Chen |
WCNC | 2 |
| 2000 | SNR scalable transcoding for video over wireless channelsabstractWe propose a novel scheme integrating transcoding, SNR scalability and error protection channel coding for video transmission over wireless channels. The target application for such error-resilient transcoding is to provide robust access to the pre-encoded high quality video server from mobile wireless terminals. The high bit rate of pre-encoded video sequence is reduced through transcoding to fit greatly the limited bandwidth of wireless links. The proposed transcoding is able to reduce both size of the video frames and the bit rate of the pre-encoded video. We also design the transcoder such that the transcoding output bitstream has the desired SNR scalability in order to guarantee a basic quality of video transmission under the hostile wireless transmission environment. With SNR scalable bitstream output, we are able to incorporate unequal error protection through channel coding for the base layer and the enhancement layers, respectively. Different rate of CRC/RCPC is adopted so as to simplify the implementation of a single channel decoder at the receiver. Experimental results show that the proposed integrated scheme of SNR scalable transcoding can guarantee the basic decoded visual quality under random and bursty error channel conditions and outperform the conventional single layer transcoding without channel coding. Chang Wen Chen |
WCNC | 2 |
| 2000 | Robust joint source-channel coding for image transmission over wireless channelsabstractWe present a fixed-length robust joint source-channel coding (RJSCC) scheme for transmitting images over wireless channels. The system integrates a joint source-channel coding (JSCC) scheme with all-pass filtering source shaping to enable robust image transmission. In particular, we are able to incorporate both transition probability and bit error rate of a bursty channel model into an end-to-end rate-distortion (R-D) function to achieve an optimum tradeoff between source coding accuracy and channel error protection under a fixed transmission rate. Experimental results show that the proposed scheme can achieve not only high peak signal-to-noise ratio performance, but also excellent perceptual quality, especially when the channel mismatch occurs. Jianfei Cai 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | Operational rate-distortion design for joint source-channel coding over noisy channelsabstractAn optimal joint source-channel coding (OJSCC) scheme is developed for memoryless generalized Gaussian distribution (GGD) sources encoding and transmission over noisy channels. Two channel models are studied, binary symmetric channels (BSC) for memoryless channels and Gilbert-Elliott channels (GEC) for bursty channels. The operational rate-distortion (R-D) function we adopted represents an end-to-end error measurement that includes errors due to both quantization and channel noise. In particular, we are able to incorporate both the channel transition probability and channel bit error rate in the case of bursty channels. With the operational R-D function, we can achieve an optimum tradeoff between source coding accuracy and channel error protection under a fixed transmission rate. Experiments show that for BSC, OJSCC outperforms the best channel optimized scalar quantization (COSQ) system at high bit rate constraint; while for GEC, we show that the optimal design achieves better performance than the popular designs based on either the average BER or the worst BER. Moreover, based on the results of OJSCC, we propose a robust joint source-channel coding (RJSCC) scheme based on a combination of OJSCC with allpass filtering, RJSCC can be applied to a broad class of GGD sources with shape factor /spl nu/<2.0 to achieve an improved transmission performance. Jianfei Cai 0001, Chang Wen Chen |
WCNC | 2 |
| 1999 | Multiple hierarchical image transmission over wireless channelsabstractIn this paper, we present a multiple hierarchical coding and transmission scheme in which both the fragile structure of set partitioning in hierarchical trees (SPIHT) source coding and the unequal importance of bits in the coded bitstream have been taken into account. Multiple subsampling is applied to split the wavelet coeffcients of the original image source into multiple subsources so that each subsource is a crude representation of the original image. The sequential dependence of coded bitstream is broken. As a result, the error propagation is limited to a single subsource which is also coded with SPIHT to achieve desired coding performance. Cyclic redundancy coder/rate compatible punctured convolutional coder (CRC/RCPC) channel coding has been used to offer unequal error protection (UEP) to the multiple coded bitstreams. In each substream all the bits in the same bit layer are protected with the same channel code. The higher the bit-layer, the more the channel coding protects. Experimented results show that this new scheme has a good error resilient performance over wireless channels with time varying and bursty errors characteristics. In particular the reconstructed images often demonstrate good visual quality. Lei Cao 0001, Chang Wen Chen |
WCNC | 2 |
| 1999 | Uniform trellis coded quantization for image transmission over noisy channels
Chang Wen Chen, Zhaohui Sun |
Signal Process. Image Commun. | 1 |
| 1999 | Image transmission over noisy channels with variable-coefficient fixed-length coding schemeabstractA variable-coefficient fixed-length coding scheme is proposed for wavelet-based image transmission over noisy channels. Multiple subband coefficients are grouped into an extended source symbol and coded as a fixed-length code word such that the extended source symbols have almost the same probabilities. Part of the codebook is fixed based on the observation of the coefficient spatial distribution patterns in each subband to alleviate the transmission of the codebook. The remaining code positions within the fixed-length codebook can be utilized to combat channel errors by carefully arranging the code positions or be filled with other frequently appearing coefficient sequences to achieve higher compression ratio. Chang Wen Chen, Zhaohui Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | A Real-Time Monitoring System for the Cardiac and Respiratory Parameters Using Bioimpedance TechniqueabstractWe have implemented a bioimpedance system based on a digital signal processor (DSP) chip (ADSP-2101, from Analog Device Corp.) and a personal computer (PC) with digital and modular design approaches for monitoring respiratory and cardiac parameters. The overall system contains a digital EPROM (eraseable programmable read-only-memory) oscillator, eight independent current generators, a voltmeter with digital demodulation, an amplifier, an I/O interface, and a friendly man-machine interface through the PC. Information such as respiratory flow, respiratory rate, tidal volume, heart rate, cardiac output, stroke volume and electrocardiograph (ECG) are presented in real time on the PC's monitor. Ji-Jer Huang, Yueh-Nong Hong, Kuo-Sheng Cheng, Chang Wen Chen |
CBMS | 4 |
| 1998 | An Improved Approach for Cardiac Dynamics Analysis based on Constrained Local Force ModelabstractProposes a constrained local force model to estimate the non-rigid motion of left ventricle over a cardiac cycle. The local force model is derived from the dynamics of independent point masses driven by local constant forces over a short period of time. The trajectory that minimizes the kinetic energy required to move a point mass from one surface to another is considered as the local displacement vector. However, the original model suffers from problems of potential one-to-many point correspondence and inconsistent motion estimation. The proposed constrained local force model attempts to achieve better and more accurate motion estimation. First, Gaussian curvature is incorporated as constraint on point correspondence to solve the one-to-many correspondence problem. Second, motion estimations of neighboring points are taken into account to obtain more consistent estimation. Experimental results show that the constrained model achieves an improved estimation in cardiac motion analysis. Chang Wen Chen |
ICIP (1) | 2 |
| 1998 | Joint Source and Channel Optimized Block TCQ with Layered Transmission and RCPCabstractWe present a fixed length block trellis coded quantization (TCQ) image coding method with layered transmission and rate compatible punctured convolutional codes (RCPC) for noisy channel image communication. Employing optimal bit allocation and TCQ, this fixed length coding scheme achieves a good compression ratio while avoiding the loss of synchronization between encoder and decoder in a noisy environment. The loss of synchronization, caused often by entropy coding in the presence of channel noise leads to catastrophic decoding errors and degrades the coding performance substantially. With layered transmission and RCPC protection, the experiments show that the PSNR performance degrades gradually compared with the schemes layered operation as the channel bit error rate increases. Chang Wen Chen |
ICIP (1) | 2 |
| 1998 | Image segmentation via adaptive K-mean clustering and knowledge-based morphological operations with biomedical applicationsabstractImage segmentation remains one of the major challenges in image analysis. In medical applications, skilled operators are usually employed to extract the desired regions that may be anatomically separate but statistically indistinguishable. Such manual processing is subject to operator errors and biases, is extremely time consuming, and has poor reproducibility. We propose a robust algorithm for the segmentation of three-dimensional (3-D) image data based on a novel combination of adaptive K-mean clustering and knowledge-based morphological operations. The proposed adaptive K-mean clustering algorithm is capable of segmenting the regions of smoothly varying intensity distributions. Spatial constraints are incorporated in the clustering algorithm through the modeling of the regions by Gibbs random fields. Knowledge-based morphological operations are then applied to the segmented regions to identify the desired regions according to the a priori anatomical knowledge of the region-of-interest. This proposed technique has been successfully applied to a sequence of cardiac CT volumetric images to generate the volumes of left ventricle chambers at 16 consecutive temporal frames. Our final segmentation results compare favorably with the results obtained using manual outlining. Extensions of this approach to other applications can be readily made when a priori knowledge of a given object is available. Chang Wen Chen, Jiebo Luo 0001, Kevin J. Parker |
IEEE Trans. Image Process. | 1 |
| 1997 | Ultrasound Image Compression Based on Subband Decomposition and Speckle SynthesisabstractWe present a study of ultrasound image compression based on subband decomposition and speckle synthesis. It is well-known that one signature characteristic of ultrasound images is the presence of speckle noise. The high frequency nature of speckle noise poses great challenges to standard coding techniques. In this study, ultrasound images are first decomposed into subbands with good spatial-frequency separation. To better separate the noise spectrum from high frequency signal content, selected high frequency subbands are further decomposed. The separation results in a structure image and a speckle image. Ultrasound images are coded such that the structure image is coded using a wavelet-based method and the speckle noise is reproduced using synthesized high frequency subbands at the decoder end. While the impact on diagnostics is yet to be investigated, this coding scheme enables efficient compression while maintaining the signature appearance of ultrasound images. Jiebo Luo 0001, Chang Wen Chen |
ICIP (3) | 3 |
| 1997 | A scene adaptive and signal adaptive quantization for subband image and video compression using waveletsabstractThe discrete wavelet transform (DWT) provides an advantageous framework of multiresolution space-frequency representation with promising applications in image processing. The challenge as well as the opportunity in wavelet-based compression is to exploit the characteristics of the subband coefficients with respect to both spectral and spatial localities. A common problem with many existing quantization methods is that the inherent image structures are severely distorted with coarse quantization. Observation shows that subband coefficients with the same magnitude generally do not have the same perceptual importance. We propose in this paper a scene adaptive and signal adaptive quantization scheme capable of exploiting the spectral and spatial localization properties resulting from the wavelet transform. The quantization is implemented as maximum a posteriori probability estimation-based clustering in which subband coefficients are quantized to their cluster means, subject to local spatial constraints. The intensity distribution of each cluster within a subband is modeled by an optimal Laplacian source to achieve signal adaptivity, while spatial constraints are enforced by appropriate Gibbs random fields (GRF) to achieve scene adaptivity. With spatially isolated coefficients removed and clustered coefficients retained at the same time, the available bits are allocated to visually important scene structures so that the information loss is least perceptible. Furthermore, the reconstruction noise in the decompressed image can be suppressed using another GRF-based enhancement algorithm. Jiebo Luo 0001, Chang Wen Chen, Kevin J. Parker, Thomas S. Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1997 | A Novel Volumetric Feature Extraction Technique, with Applications to MR ImagesabstractA semiautomated feature extraction algorithm is presented for the extraction and measurement of the hippocampus from volumetric magnetic resonance imaging (MRI) head scans. This algorithm makes use of elements of both deformable model and region growing techniques and allows incorporation of a priori operator knowledge of hippocampal location and shape. Experimental results indicate that the algorithm is able to estimate hippocampal volume and asymmetry with an accuracy which approaches that of laborious manual outlining techniques. Edward A. Ashton, Kevin J. Parker, Michel J. Berg, Chang Wen Chen |
IEEE Trans. Medical Imaging | 4 |
| 1996 | A cellular neural network for clustering-based adaptive quantization in subband video compressionabstractThis paper presents a novel cellular connectionist model for the implementation of a clustering-based adaptive quantization in video coding applications. The adaptive quantization has been designed for a wavelet-based video coding system with a desired scene adaptive and signal adaptive quantization. Since the adaptive quantization is accomplished through a maximum a posteriori probability (MAP) estimation-based clustering process, its massive computation of neighborhood constraints makes it difficult for a software-based real-time implementation of video coding applications. The proposed cellular connectionist model aims at designing an architecture for the real-time implementation of the clustering-based adaptive quantization. With a cellular neural network architecture mapping onto the image domain, the powerful Gibbs spatial constraints are realized through interactions among neurons connected with their neighbors. In addition, the computation of coefficient distribution is designed as an external input to each component of a neuron or processing element (PE). We prove that the proposed cellular neural network does converge to the desired steady state with the proposed, update scheme. This model also provides a general architecture for image processing tasks with Gibbs spatial constraint-based computations. Chang Wen Chen, Lulin Chen, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1996 | Face location in wavelet-based video compression for high perceptual quality videoconferencingabstractWe present a human face location technique based on contour extraction within the framework of a wavelet-based video compression scheme for videoconferencing applications. In addition to an adaptive quantization in which spatial constraints are enforced to preserve perceptually important information at low bit rates, semantic information of the human face is incorporated to design a hybrid compression scheme for videoconferencing since the human face is often the most important portion within a frame and should be coded with high fidelity. The human face is detected based on contour extraction and feature point analysis. An approximated face mask is then used in the quantization of the decomposed subbands. At the same total bit rate, coarser quantization of the background enables the face region to be quantized finer and coded with higher quality. Simulation results have shown that the perceptual image quality can be greatly improved using the proposed scheme. Jiebo Luo 0001, Chang Wen Chen, Kevin J. Parker |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1996 | Artifact reduction in low bit rate DCT-based image compressionabstractThis correspondence presents a scheme for artifact reduction of low bit rate discrete-cosine-transform-compressed (DCT-compressed) images. First, the DC coefficients are calibrated using gradient continuity constraints. Then, an improved Huber-Markov-random-field-based (HMRF-based) smoothing is applied. The constrained optimization is implemented by the iterative conditional mode (ICM). Final reconstructions of typical images with improvements in both visual quality and peak signal-to-noise ratio (PSNR) are also shown. Jiebo Luo 0001, Chang Wen Chen, Kevin J. Parker, Thomas S. Huang |
IEEE Trans. Image Process. | 2 |