VLDB 2026 Research / reviewers in the wild / expert
Ya-Qin Zhang
dblp:09/2187
· DBLP profile ↗
158ranked-venue papers
11as first author
37since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 92 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 35 · 2 first-author · 34 since 2021Computer networks · 19 · 2 first-author · 1 since 2021Systems, architecture and hardware · 13 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S²Drug: Bridging Protein Sequence and 3D Structure in Contrastive Representation Learning for Virtual ScreeningabstractVirtual screening (VS) is an essential task in drug discovery, focusing on the identification of small-molecule ligands that bind to specific protein pockets. Existing deep learning methods, from early regression models to recent contrastive learning approaches, primarily rely on structural data while overlooking protein sequences, which are more accessible and can enhance generalizability. However, directly integrating protein sequences poses challenges due to the redundancy and noise in large-scale protein-ligand datasets. To address these limitations, we propose S²Drug, a two-stage framework that explicitly incorporates protein Sequence information and 3D Structure context in protein-ligand contrastive representation learning. In the first stage, we perform protein sequence pretraining on ChemBL using an ESM2-based backbone, combined with a tailored data sampling strategy to reduce redundancy and noise on both protein and ligand sides. In the second stage, we fine-tune on PDBBind by fusing sequence and structure information through a residue-level gating module, while introducing an auxiliary binding site prediction task. This auxiliary task guides the model to accurately localize binding residues within the protein sequence and capture their 3D spatial arrangement, thereby refining protein-ligand matching. Across multiple benchmarks, S²Drug consistently improves virtual screening performance and achieves strong results on binding site prediction, demonstrating the value of bridging sequence and structure in contrastive learning. Bowei He, Yankai Chen 0001, Yanyan Lan, Chen Ma 0001, Philip S. Yu, Ya-Qin Zhang, Wei-Ying Ma |
AAAI | 7 |
| 2026 | Learning Protein-Ligand Binding in Hyperbolic SpaceabstractProtein-ligand binding prediction is central to virtual screening and affinity ranking, two fundamental tasks in drug discovery. While recent retrieval-based methods embed ligands and protein pockets into Euclidean space for similarity-based search, the geometry of Euclidean embeddings often fails to capture the hierarchical structure and fine-grained affinity variations intrinsic to molecular interactions. In this work, we propose HypSeek, a hyperbolic representation learning framework that embeds ligands, protein pockets, and sequences into Lorentz-model hyperbolic space. By leveraging the exponential geometry and negative curvature of hyperbolic space, HypSeek enables expressive, affinity-sensitive embeddings that can effectively model both global activity and subtle functional differences–particularly in challenging cases such as activity cliffs, where structurally similar ligands exhibit large affinity gaps. Our model unifies virtual screening and affinity ranking in a single framework, introducing a protein-guided three-tower architecture to enhance representational structure. HypSeek improves early enrichment in virtual screening on DUD-E from 42.63 to 51.44 (+20.7%) and affinity ranking correlation on JACS from 0.5774 to 0.7239 (+25.4%), demonstrating the benefits of hyperbolic geometry across both tasks and highlighting its potential as a powerful inductive bias for protein-ligand modeling. Wenyu Zhu, Ya-Qin Zhang, Wei-Ying Ma, Yanyan Lan |
AAAI | 5 |
| 2026 | Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement LearningabstractXuanyu Lei, Chenliang Li, Yuning Wu, Kaiming Liu, Weizhou Shen, Peng Li, Ming Yan, Fei Huang, Ya-Qin Zhang, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xuanyu Lei, Chenliang Li 0003, Yuning Wu 0001, Kaiming Liu, Weizhou Shen, Peng Li 0030, Ming Yan 0008, Fei Huang 0002, Ya-Qin Zhang, Yang Liu 0005 |
ACL (1) | 9 |
| 2025 | PICD: Versatile Perceptual Image Compression with Diffusion RenderingabstractRecently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when compressing text. To tackle this challenge, we propose versatile perceptual screen image compression with diffusion rendering (PICD), a codec that works well for both screen and natural images. More specifically, we propose a compression framework that encodes the text and image separately, and renders them into one image using diffusion model. For this diffusion rendering, we integrate conditional information into diffusion models at three distinct levels: 1). Domain level: We fine-tune the base diffusion model using text content prompts with screen content. 2). Adaptor level: We develop an efficient adaptor to control the diffusion model using compressed image and text as input. 3). Instance level: We apply instance-wise guidance to further enhance the decoding process. Empirically, our PICD surpasses existing perceptual codecs in terms of both text accuracy and perceptual quality. Additionally, without text conditions, our approach serves effectively as a perceptual codec for natural images. Tongda Xu, Jiahao Li 0001, Bin Li 0012, Yan Wang 0105, Ya-Qin Zhang, Yan Lu 0001 |
CVPR | 5 |
| 2025 | Universal Actions for Enhanced Embodied Foundation ModelsabstractTraining on diverse, internet-scale data is a key factor in the success of recent large foundation models. Yet, using the same recipe for building embodied agents has faced noticeable difficulties. Despite the availability of many crowd-sourced embodied datasets, their action spaces often exhibit significant heterogeneity due to distinct physical embodiment and control interfaces for different robots, causing substantial challenges in developing embodied foundation models using cross-domain data. In this paper, we introduce UniAct, a new embodied foundation modeling framework operating in a Universal Act ion Space. Our learned universal actions capture the generic atomic behaviors across diverse robots by exploiting their shared structural features, and enable enhanced cross-domain data utilization and cross-embodiment generalizations by eliminating the notorious heterogeneity. The universal actions can be efficiently translated back to heterogeneous actionable commands by simply adding embodiment-specific details, from which fast adaptation to new robots becomes simple and straightforward. Our 0.5B instantiation of Uni-Act reaches 14X larger SOTA embodied foundation models in extensive evaluations on various real-world and simulation robots, showcasing exceptional cross-embodiment control and adaptation capability, highlighting the crucial benefit of adopting universal actions. Project page: https://2toinf.github.io/UniAct/ Jinliang Zheng, Dongxiu Liu, Yinan Zheng, Zhonghong Ou, Yu Liu 0015, Ya-Qin Zhang, Xianyuan Zhan |
CVPR | 9 |
| 2025 | InvRGB+L: Inverse Rendering of Complex Scenes with Unified Color and LiDAR Reflectance ModelingabstractWe present InvRGB+L, a novel inverse rendering model that reconstructs large, relightable, and dynamic scenes from a single RGB+LiDAR sequence. Conventional inverse graphics methods rely primarily on RGB observations and use LiDAR mainly for geometric information, often resulting in suboptimal material estimates due to visible light interference. We find that LiDAR's intensity values-captured with active illumination in a different spectral range-offer complementary cues for robust material estimation under variable lighting. Inspired by this, InvRGB+L leverages LiDAR intensity cues to overcome challenges inherent in RGB-centric inverse graphics through two key innovations: (1) a novel physics-based LiDAR shading model and (2) RGB-LiDAR material consistency losses. The model produces novel-view RGB and LiDAR renderings of urban and indoor scenes and supports relighting, night simulations, and dynamic object insertions, achieving results that surpass current state-of-the-art methods in both scene-level urban inverse rendering and LiDAR simulation. Xiaoxue Chen, Bhargav Chandaka, Chih-Hao Lin, Ya-Qin Zhang, David A. Forsyth, Shenlong Wang |
ICCV | 4 |
| 2025 | Reframing Structure-Based Drug Design Model Evaluation via Metrics Correlated to Practical NeedsabstractRecent advances in structure-based drug design (SBDD) have produced surprising results, with models often generating molecules that achieve better Vina docking scores than actual ligands. However, these results are frequently overly optimistic due to the limitations of docking score accuracy and the challenges of wet-lab validation. While generated molecules may demonstrate high QED (drug-likeness) and SA (synthetic accessibility) scores, they often lack true drug-like properties or synthesizability. To address these limitations, we propose a model-level evaluation framework that emphasizes practical metrics aligned with real-world applications. Inspired by recent findings on the utility of generated molecules in ligand-based virtual screening, our framework evaluates SBDD models by their ability to produce molecules that effectively retrieve active compounds from chemical libraries via similarity-based searches. This approach provides a direct indication of therapeutic potential, bridging the gap between theoretical performance and real-world utility. Our experiments reveal that while SBDD models may excel in theoretical metrics like Vina scores, they often fall short in these practical metrics. By introducing this new evaluation strategy, we aim to enhance the relevance and impact of SBDD models for pharmaceutical research and development. Haichuan Tan, Yanwen Huang, Minsi Ren, Wei-Ying Ma, Ya-Qin Zhang, Yanyan Lan |
ICLR | 7 |
| 2025 | Redefining the task of Bioactivity PredictionabstractSmall molecules are vital to modern medicine, and accurately predicting their bioactivity against protein targets is crucial for therapeutic discovery and development. However, current machine learning models often rely on spurious features, leading to biased outcomes. Notably, a simple pocket-only baseline can achieve results comparable to, and sometimes better than, more complex models that incorporate both the protein pockets and the small molecules. Our analysis reveals that this phenomenon arises from insufficient training data and an improper evaluation process, which is typically conducted at the pocket level rather than the small molecule level. To address these issues, we redefine the bioactivity prediction task by introducing the SIU dataset-a million-scale Structural small molecule-protein Interaction dataset for Unbiased bioactivity prediction task, which is 50 times larger than the widely used PDBbind. The bioactivity labels in SIU are derived from wet experiments and organized by label types, ensuring greater accuracy and comparability. The complexes in SIU are constructed using a majority vote from three commonly used docking software programs, enhancing their reliability. Additionally, the structure of SIU allows for multiple small molecules to be associated with each protein pocket, enabling the redefinition of evaluation metrics like Pearson and Spearman correlations across different small molecules targeting the same protein pocket. Experimental results demonstrate that this new task provides a more challenging and meaningful benchmark for training and evaluating bioactivity prediction models, ultimately offering a more robust assessment of model performance. Yanwen Huang, Yinjun Jia, Hongbo Ma, Wei-Ying Ma, Ya-Qin Zhang, Yanyan Lan |
ICLR | 6 |
| 2025 | Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a PosteriorabstractRecent advancements in diffusion models have been leveraged to address inverse problems without additional training, and Diffusion Posterior Sampling (DPS) (Chung et al., 2022a) is among the most popular approaches. Previous analyses suggest that DPS accomplishes posterior sampling by approximating the conditional score. While in this paper, we demonstrate that the conditional score approximation employed by DPS is not as effective as previously assumed, but rather aligns more closely with the principle of maximizing a posterior (MAP). This assertion is substantiated through an examination of DPS on 512$\times$512 ImageNet images, revealing that: 1) DPS’s conditional score estimation significantly diverges from the score of a well-trained conditional diffusion model and is even inferior to the unconditional score; 2) The mean of DPS’s conditional score estimation deviates significantly from zero, rendering it an invalid score estimation; 3) DPS generates high-quality samples with significantly lower diversity. In light of the above findings, we posit that DPS more closely resembles MAP than a conditional score estimator, and accordingly propose the following enhancements to DPS: 1) we explicitly maximize the posterior through multi-step gradient ascent and projection; 2) we utilize a light-weighted conditional score estimator trained with only 100 images and 8 GPU hours. Extensive experimental results indicate that these proposed improvements significantly enhance DPS's performance. The source code for these improvements is provided in https://github.com/tongdaxu/Rethinking-Diffusion-Posterior-Sampling-From-Conditional-Score-Estimator-to-Maximizing-a-Posterior. Tongda Xu, Xiyan Cai, Xingtong Ge, Dailan He, Ya-Qin Zhang, Yan Wang 0105 |
ICLR | 8 |
| 2025 | Contrastive Private Data Synthesis via Weighted Multi-PLM FusionabstractSubstantial quantity and high quality are the golden rules of making a good training dataset with sample privacy protection equally important. Generating synthetic samples that resemble high-quality private data while ensuring Differential Privacy (DP), a formal privacy guarantee, promises scalability and practicality. However, existing methods relying on pre-trained models for data synthesis often struggle in data-deficient scenarios, suffering from limited sample size, inevitable generation noise and existing pre-trained model bias. To address these challenges, we propose a novel contr**A**stive private data **S**ynthesis via **W**eighted multiple **P**re-trained generative models framework, named as **WASP**. WASP utilizes limited private samples for more accurate private data distribution estimation via a Top-*Q* voting mechanism, and leverages low-quality synthetic samples for contrastive generation via collaboration among dynamically weighted multiple pre-trained models. Extensive experiments on 6 well-developed datasets with 6 open-source and 3 closed-source PLMs demonstrate the superiority of WASP in improving model performance over diverse downstream tasks. Code is available at https://github.com/LindaLydia/WASP. Tianyuan Zou, Yang Liu 0165, Peng Li 0030, Yufei Xiong, Jianqing Zhang, Xiaozhou Ye, Ye Ouyang, Ya-Qin Zhang |
ICML | 9 |
| 2025 | Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal LearningabstractMultimodal task specification is essential for enhanced robotic performance, where Cross-modality Alignment enables the robot to holistically understand complex task instructions. Directly annotating multimodal instructions for model training proves impractical, due to the sparsity of paired multimodal data. In this study, we demonstrate that by leveraging unimodal instructions abundant in real data, we can effectively teach robots to learn multimodal task specifications. First, we endow the robot with strong Crossmodality Alignment capabilities, by pretraining a robotic multimodal encoder using extensive out-of-domain data. Then, we employ two Collapse and Corrupt operations to further bridge the remaining modality gap in the learned multimodal representation. This approach projects different modalities of identical task goal as interchangeable representations, thus enabling accurate robotic operations within a well-aligned multimodal latent space. Evaluation across more than 130 tasks and 4000 evaluations on both simulated LIBERO benchmark and real robot platforms showcases the superior capabilities of our proposed framework, demonstrating significant potential in overcoming data constraints in robotic learning. Website: zh1hao.wang/Robo_MUTUAL Jinliang Zheng, Xiaoai Zhou, Guanming Wang, Guanglu Song, Yu Liu 0015, Ya-Qin Zhang, Junzhi Yu 0001, Xianyuan Zhan |
ICRA | 9 |
| 2025 | IROAM: Improving Roadside Monocular 3D Object Detection Learning from Autonomous Vehicle Data DomainabstractIn autonomous driving, The perception capabilities of the ego-vehicle can be improved with roadside sensors, which can provide a holistic view of the environment. However, existing monocular detection methods designed for vehicle cameras are not suitable for roadside cameras due to viewpoint domain gaps. To bridge this gap and Improve ROAdside Monocular 3D object detection, we propose IROAM, a semantic-geometry decoupled contrastive learning framework, which takes vehicle-side and roadside data as input simultaneously. IROAM has two significant modules. In-Domain Query Interaction module utilizes a transformer to learn content and depth information for each domain and outputs object queries. Cross-Domain Query Enhancement To learn better feature representations from two domains, Cross-Domain Query Enhancement decouples queries into semantic and geometry parts and only the former is used for contrastive learning. Experiments demonstrate the effectiveness of IROAM in improving roadside detector's performance. The results validate that IROAM has the capabilities to learn cross-domain information. Zhe Wang 0070, Xiaoliang Huo, Siqi Fan 0002, Ya-Qin Zhang, Yan Wang 0105 |
ICRA | 5 |
| 2025 | CoopDETR: A Unified Cooperative Perception Framework for 3D Detection via Object QueryabstractCooperative perception enhances the individual perception capabilities of autonomous vehicles (AVs) by providing a comprehensive view of the environment. However, balancing perception performance and transmission costs remains a significant challenge. Current approaches that transmit regionlevel features across agents are limited in interpretability and demand substantial bandwidth, making them unsuitable for practical applications. In this work, we propose CoopDETR, a novel cooperative perception framework that introduces objectlevel feature cooperation via object query. Our framework consists of two key modules: single-agent query generation, which efficiently encodes raw sensor data into object queries, reducing transmission cost while preserving essential information for detection; and cross-agent query fusion, which includes Spatial Query Matching (SQM) and Object Query Aggregation (OQA) to enable effective interaction between queries. Our experiments on the OPV2V and V2XSet datasets demonstrate that CoopDETR achieves state-of-the-art performance and significantly reduces transmission costs to 1/782 of previous methods. Zhe Wang 0070, Shaocong Xu, Xucai Zhuang, Tongda Xu, Yan Wang 0105, Ya-Qin Zhang |
ICRA | 8 |
| 2025 | AutoDroid-V2: Boosting SLM-based GUI Agents via Code GenerationabstractLarge language models (LLMs) have brought exciting new advances to mobile UI agents, a long-standing research field that aims to complete arbitrary natural language tasks through mobile UI interactions. However, existing UI agents usually demand powerful large language models that are difficult to be deployed locally on end-users' devices, raising huge concerns about user privacy and centralized serving cost. Inspired by the remarkable coding abilities of recent small language models (SLMs), we propose to convert the UI task automation problem to a code generation problem, which can be effectively solved by an on-device SLM and efficiently executed with an on-device code interpreter. Unlike normal coding tasks that can be extensively pre-trained with public datasets, generating UI automation code is challenging due to the diversity, complexity, and variability of target apps. Therefore, we adopt a document-centered approach that automatically builds fine-grained API documentation for each app and generates diverse task samples based on this documentation. By guiding the agent with the synthetic documents and task samples, it learns to generate precise and efficient scripts to complete unseen tasks. Based on detailed comparisons with state-of-the-art mobile UI agents, our approach effectively improves the mobile task automation with significantly higher success rates and lower latency/token consumption. Code is open-sourced at https://github.com/MobileLLM/AutoDroid-V2. Hao Wen 0004, Shizuo Tian, Borislav Pavlov, Wenjie Du 0004, Ge Chang 0002, Shanhui Zhao, Yunxin Liu 0001, Ya-Qin Zhang, Yuanchun Li 0003 |
MobiSys | 10 |
| 2025 | CIDD: Collaborative Intelligence for Structure-Based Drug Design Empowered by LLMsabstractStructure-guided molecular generation is pivotal in early-stage drug discovery, enabling the design of compounds tailored to specific protein targets. However, despite recent advances in 3D generative modeling, particularly in improving docking scores, these methods often produce rare and intrinsically irrational molecular structures that deviate from drug-like chemical space. To quantify this issue, we propose a novel metric, the Molecule Reasonable Ratio (MRR), which measures structural rationality and reveals a critical gap between existing models and real-world approved drugs. To address this, we introduce the Collaborative Intelligence Drug Design (CIDD) framework, the first approach to unify the 3D interaction modeling capabilities of generative models with the general knowledge and reasoning power of large language models (LLMs). By leveraging LLM-based Chain-of-Thought reasoning, CIDD generates molecules that not only bind effectively to protein pockets but also exhibit strong structural drug-likeness, rationality, and synthetic accessibility. On the CrossDocked2020 benchmark, CIDD consistently improves drug-likeness metrics, including QED, SA, and MRR, across different base generative models, while maintaining competitive binding affinity. Notably, it raises the combined success rate (balancing drug-likeness and binding) from 15.72% to 34.59%, more than doubling previous results. These findings demonstrate the value of integrating knowledge reasoning with geometric generation to advance AI-driven drug design. Yanwen Huang, Yiqiao Liu, Wenxuan Xie, Bowei He, Haichuan Tan, Wei-Ying Ma, Ya-Qin Zhang, Yanyan Lan |
NeurIPS | 8 |
| 2025 | DAPO: An Open-Source LLM Reinforcement Learning System at ScaleabstractInference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the **D**ecoupled Clip and **D**ynamic s**A**mpling **P**olicy **O**ptimization (**DAPO**) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL. Qiying Yu, Zheng Zhang 0001, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, Lingjun Liu, Xin Liu 0039, Haibin Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang 0022, Mofan Zhang, Ru Zhang 0006, Wang Zhang 0017, Jiaze Chen, Jiangjie Chen, Hongli Yu, Yuxuan Song 0002, Xiangpeng Wei, Hao Zhou 0012, Wei-Ying Ma, Ya-Qin Zhang, Mingxuan Wang |
NeurIPS | 33 |
| 2025 | AANet: Virtual Screening under Structural Uncertainty via Alignment and AggregationabstractVirtual screening (VS) is a critical component of modern drug discovery, yet most existing methods—whether physics-based or deep learning-based—are developed around *holo* protein structures with known ligand-bound pockets. Consequently, their performance degrades significantly on *apo* or predicted structures such as those from AlphaFold2, which are more representative of real-world early-stage drug discovery, where pocket information is often missing. In this paper, we introduce an alignment-and-aggregation framework to enable accurate virtual screening under structural uncertainty. Our method comprises two core components: (1) a tri-modal contrastive learning module that aligns representations of the ligand, the *holo* pocket, and cavities detected from structures, thereby enhancing robustness to pocket localization error; and (2) a cross-attention based adapter for dynamically aggregating candidate binding sites, enabling the model to learn from activity data even without precise pocket annotations. We evaluated our method on a newly curated benchmark of *apo* structures, where it significantly outperforms state-of-the-art methods in blind apo setting, improving the early enrichment factor (EF1\%) from 11.75 to 37.19. Notably, it also maintains strong performance on *holo* structures. These results demonstrate the promise of our approach in advancing first-in-class drug discovery, particularly in scenarios lacking experimentally resolved protein-ligand complexes. Our implementation is publicly available at [https://github.com/Wiley-Z/AANet](https://github.com/Wiley-Z/AANet). Wenyu Zhu, Yinjun Jia, Haichuan Tan, Ya-Qin Zhang, Wei-Ying Ma, Yanyan Lan |
NeurIPS | 6 |
| 2025 | Distributional Soft Actor-Critic With Three RefinementsabstractReinforcement learning (RL) has shown remarkable success in solving complex decision-making and control tasks. However, many model-free RL algorithms experience performance degradation due to inaccurate value estimation, particularly the overestimation of Q-values, which can lead to suboptimal policies. To address this issue, we previously proposed the Distributional Soft Actor-Critic (DSAC or DSACv1), an off-policy RL algorithm that enhances value estimation accuracy by learning a continuous Gaussian value distribution. Despite its effectiveness, DSACv1 faces challenges such as training instability and sensitivity to reward scaling, caused by high variance in critic gradients due to return randomness. In this paper, we introduce three key refinements to DSACv1 to overcome these limitations and further improve Q-value estimation accuracy: expected value substitution, twin value distribution learning, and variance-based critic gradient adjustment. The enhanced algorithm, termed DSAC with Three refinements (DSAC-T or DSACv2), is systematically evaluated across a diverse set of benchmark tasks. Without the need for task-specific hyperparameter tuning, DSAC-T consistently matches or outperforms leading model-free RL algorithms, including SAC, TD3, DDPG, TRPO, and PPO, in all tested environments. Additionally, DSAC-T ensures a stable learning process and maintains robust performance across varying reward scales. Its effectiveness is further demonstrated through real-world application in controlling a wheeled robot, highlighting its potential for deployment in practical robotic tasks. Jingliang Duan, Wenxuan Wang 0004, Liming Xiao, Jiaxin Gao 0002, Shengbo Eben Li, Chang Liu 0002, Ya-Qin Zhang, Bo Cheng 0003, Keqiang Li 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Defending Against Inference and Backdoor Attacks in Vertical Federated Learning via Mutual Information RegularizationabstractVertical Federated Learning (VFL) is widely utilized in real-world applications to enable collaborative learning while protecting local data and models. However, previous works show that parties without labels (passive parties) in VFL can infer the sensitive label or feature information owned by the party with labels (active party), or execute backdoor attacks. Meanwhile, active party can also infer sensitive feature or attribute information from passive party. All these pose great challenges to VFL systems. Former defense methods tend to experience either a loss in overall effectiveness or are too specialized for specific tasks. In this work, we propose a novel method and a unified framework for defending various attacks in VFL altogether, namely Mutual Information Regularization Defense (MID), which limits the mutual information between private raw data and intermediate outputs to achieve a consistently better trade-off between model utility and privacy. We provide both theoretical and experimental evidence to confirm the effectiveness of our MID framework in defending against a wide range of label and feature inference attacks, along with backdoor attacks in VFL. These showcase its promising potential as a versatile and effective defense mechanism, not tied to any specific task. Tianyuan Zou, Yang Liu 0165, Xiaozhou Ye, Ye Ouyang, Ya-Qin Zhang |
IEEE Big Data | 5 |
| 2024 | Time-conditioned Illumination for Inverse Rendering of Outdoor Scenes
Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang |
BMVC | 4 |
| 2024 | FuseGen: PLM Fusion for Data-generation based Zero-shot LearningabstractData-generation based zero-shot learning, although effective in training Small Task-specific Models (STMs) via synthetic datasets generated by Pre-trained Language Models (PLMs), is often limited by the low quality of such synthetic datasets.Previous solutions have primarily focused on single PLM settings, where synthetic datasets are typically restricted to specific sub-spaces and often deviate from real-world distributions, leading to severe distribution bias.To mitigate such bias, we propose FuseGen, a novel data-generation based zero-shot learning framework that introduces a new criteria for subset selection from synthetic datasets via utilizing multiple PLMs and trained STMs.The chosen subset provides in-context feedback to each PLM, enhancing dataset quality through iterative data generation.Trained STMs are then used for sample re-weighting as well, further improving data quality.Extensive experiments across diverse tasks demonstrate that FuseGen substantially outperforms existing methods, highly effective in boosting STM performance in a PLM-agnostic way. 1 Tianyuan Zou, Yang Liu 0005, Peng Li 0030, Jianqing Zhang, Ya-Qin Zhang |
EMNLP | 6 |
| 2024 | Query-Policy Misalignment in Preference-Based Reinforcement LearningabstractPreference-based reinforcement learning (PbRL) provides a natural way to align RL agents’ behavior with human desired outcomes, but is often restrained by costly human feedback. To improve feedback efficiency, most existing PbRL methods focus on selecting queries to maximally improve the overall quality of the reward model, but counter-intuitively, we find that this may not necessarily lead to improved performance. To unravel this mystery, we identify a long-neglected issue in the query selection schemes of existing PbRL studies: Query-Policy Misalignment. We show that the seemingly informative queries selected to improve the overall quality of reward model actually may not align with RL agents’ interests, thus offering little help on policy learning and eventually resulting in poor feedback efficiency. We show that this issue can be effectively addressed via policy-aligned query and a specially designed hybrid experience replay, which together enforce the bidirectional query-policy alignment. Simple yet elegant, our method can be easily incorporated into existing approaches by changing only a few lines of code. We showcase in comprehensive experiments that our method achieves substantial gains in both human feedback and RL sample efficiency, demonstrating the importance of addressing query-policy misalignment in PbRL tasks. Xianyuan Zhan, Qing-Shan Jia, Ya-Qin Zhang |
ICLR | 5 |
| 2024 | Idempotence and Perceptual Image CompressionabstractIdempotence is the stability of image codec to re-compression. At the first glance, it is unrelated to perceptual image compression. However, we find that theoretically: 1) Conditional generative model-based perceptual codec satisfies idempotence; 2) Unconditional generative model with idempotence constraint is equivalent to conditional generative codec. Based on this newfound equivalence, we propose a new paradigm of perceptual image codec by inverting unconditional generative model with idempotence constraints. Our codec is theoretically equivalent to conditional generative codec, and it does not require training new models. Instead, it only requires a pre-trained mean-square-error codec and unconditional generative model. Empirically, we show that our proposed approach outperforms state-of-the-art methods such as HiFiC and ILLM, in terms of Fréchet Inception Distance (FID). The source code is provided in https://github.com/tongdaxu/Idempotence-and-Perceptual-Image-Compression. Tongda Xu, Ziran Zhu, Dailan He, Yanghao Li, Zhe Wang 0070, Hongwei Qin, Yan Wang 0105, Ya-Qin Zhang |
ICLR | 11 |
| 2024 | VFLAIR: A Research Library and Benchmark for Vertical Federated LearningabstractVertical Federated Learning (VFL) has emerged as a collaborative training paradigm that allows participants with different features of the same group of users to accomplish cooperative training without exposing their raw data or model parameters. VFL has gained significant attention for its research potential and real-world applications in recent years, but still faces substantial challenges, such as in defending various kinds of data inference and backdoor attacks. Moreover, most of existing VFL projects are industry-facing and not easily used for keeping track of the current research progress. To address this need, we present an extensible and lightweight VFL framework VFLAIR (available at https://github.com/FLAIR-THU/VFLAIR), which supports VFL training with a variety of models, datasets and protocols, along with standardized modules for comprehensive evaluations of attacks and defense strategies. We also benchmark $11$ attacks and $8$ defenses performance under different communication and model partition settings and draw concrete insights and recommendations on the choice of defense strategies for different practical VFL deployment scenarios. Tianyuan Zou, Zixuan Gu, Hideaki Takahashi, Yang Liu 0165, Ya-Qin Zhang |
ICLR | 6 |
| 2024 | DecisionNCE: Embodied Multimodal Representations via Implicit Preference LearningabstractMultimodal pretraining is an effective strategy for the trinity of goals of representation learning in autonomous robots: $1)$ extracting both local and global task progressions; $2)$ enforcing temporal consistency of visual representation; $3)$ capturing trajectory-level language grounding. Most existing methods approach these via separate objectives, which often reach sub-optimal solutions. In this paper, we propose a universal unified objective that can simultaneously extract meaningful task progression information from image sequences and seamlessly align them with language instructions. We discover that via implicit preferences, where a visual trajectory inherently aligns better with its corresponding language instruction than mismatched pairs, the popular Bradley-Terry model can transform into representation learning through proper reward reparameterizations. The resulted framework, DecisionNCE, mirrors an InfoNCE-style objective but is distinctively tailored for decision-making tasks, providing an embodied representation learning framework that elegantly extracts both local and global task progression features, with temporal consistency enforced through implicit time contrastive learning, while ensuring trajectory-level instruction grounding via multimodal joint encoding. Evaluation on both simulated and real robots demonstrates that DecisionNCE effectively facilitates diverse downstream policy learning tasks, offering a versatile solution for unified representation and reward learning. Project Page: https://2toinf.github.io/DecisionNCE/ Jinliang Zheng, Yinan Zheng, Liyuan Mao, Sijie Cheng, Jihao Liu, Yu Liu 0015, Ya-Qin Zhang, Xianyuan Zhan |
ICML | 11 |
| 2024 | EMIFF: Enhanced Multi-scale Image Feature Fusion for Vehicle-Infrastructure Cooperative 3D Object DetectionabstractIn autonomous driving, cooperative perception makes use of multi-view cameras from both vehicles and infrastructure, providing a global vantage point with rich semantic context of road conditions beyond a single vehicle viewpoint. Currently, two major challenges persist in vehicle-infrastructure cooperative 3D (VIC3D) object detection: 1) inherent pose errors when fusing multi-view images, caused by time asynchrony across cameras; 2) information loss in transmission process resulted from limited communication bandwidth. To address these issues, we propose a novel camera-based 3D detection framework for VIC3D task, Enhanced Multi-scale Image Feature Fusion (EMIFF). To fully exploit holistic perspectives from both vehicles and infrastructure, we propose Multi-scale Cross Attention (MCA) and Camera-aware Channel Masking (CCM) modules to enhance infrastructure and vehicle features at scale, spatial, and channel levels to correct the pose error introduced by camera asynchrony. We also introduce a Feature Compression (FC) module with channel and spatial compression blocks for transmission efficiency. Experiments show that EMIFF achieves SOTA on DAIR-V2X-C datasets, significantly outperforming previous early-fusion and late-fusion methods with comparable transmission costs. Zhe Wang 0070, Siqi Fan 0002, Xiaoliang Huo, Tongda Xu, Yan Wang 0105, Ya-Qin Zhang |
ICRA | 8 |
| 2024 | Defending Batch-Level Label Inference and Replacement Attacks in Vertical Federated LearningabstractIn a vertical federated learning (VFL) scenario where features and models are split into different parties, it has been shown that sample-level gradient information can be exploited to deduce crucial label information that should be kept secret. An immediate defense strategy is to protect sample-level messages communicated with Homomorphic Encryption (HE), exposing only batch-averaged local gradients to each party. In this paper, we show that even with HE-protected communication, private labels can still be reconstructed with high accuracy by gradient inversion attack, contrary to the common belief that batch-averaged information is safe to share under encryption. We then show that backdoor attack can also be conducted by directly replacing encrypted communicated messages without decryption. To tackle these attacks, we propose a novel defense method, Confusional AutoEncoder (termedCAE), which is based on autoencoder and entropy regularization to disguise true labels. To further defend attackers with sufficient prior label knowledge, we introduce DiscreteSGD-enhanced CAE (termedDCAE), and show that DCAE significantly boosts the main task accuracy than other known methods when defending various label inference attacks. Tianyuan Zou, Yang Liu 0165, Yan Kang 0001, Wenhan Liu, Yuanqin He, Zhihao Yi, Qiang Yang 0001, Ya-Qin Zhang |
IEEE Trans. Big Data | 8 |
| 2024 | Vertical Federated Learning: Concepts, Advances, and ChallengesabstractVertical Federated Learning (VFL) is a federated learning setting where multiple parties with different features about the same set of users jointly train machine learning models without exposing their raw data or model parameters. Motivated by the rapid growth in VFL research and real-world applications, we provide a comprehensive review of the concept and algorithms of VFL, as well as current advances and challenges in various aspects, including effectiveness, efficiency, and privacy. We provide an exhaustive categorization for VFL settings and privacy-preserving protocols and comprehensively analyze the privacy attacks and defense strategies for each protocol. In the end, we propose a unified framework, termed VFLow, which considers the VFL problem under communication, computation, privacy, as well as effectiveness and fairness constraints. Finally, we review the most recent advances in industrial applications, highlighting open challenges and future directions for VFL. Yang Liu 0165, Yan Kang 0001, Tianyuan Zou, Yanhong Pu, Yuanqin He, Xiaozhou Ye, Ye Ouyang, Ya-Qin Zhang, Qiang Yang 0001 |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2023 | DPF: Learning Dense Prediction Fields with Weak SupervisionabstractNowadays, many visual scene understanding problems are addressed by dense prediction networks. But pixel-wise dense annotations are very expensive (e.g., for scene parsing) or impossible (e.g., for intrinsic image decomposition), motivating us to leverage cheap point-level weak supervision. However, existing pointly-supervised methods still use the same architecture designed for full supervision. In stark contrast to them, we propose a new paradigm that makes predictions for point coordinate queries, as inspired by the recent success of implicit representations, like distance or radiance fields. As such, the method is named as dense prediction fields (DPFs). DPFs generate expressive intermediate features for continuous sub-pixel locations, thus allowing outputs of an arbitrary resolution. DPFs are naturally compatible with point-level supervision. We showcase the effectiveness of DPFs using two substantially different tasks: high-level semantic parsing and low-level intrinsic image decomposition. In these two cases, supervision comes in the form of single-point semantic category and two-point relative reflectance, respectively. As benchmarked by three large-scale public datasets PASCALContext, ADE20K and IIW, DPFs set new state-of-the-art performance on all of them with significant margins. Code can be accessed at https://github.com/cxx226/DPF. Xiaoxue Chen, Yuhang Zheng 0004, Yupeng Zheng, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang |
CVPR | 7 |
| 2023 | Mind the Gap: Offline Policy Optimization for Imperfect Rewards
Haoran Xu 0003, Xianyuan Zhan, Qing-Shan Jia, Ya-Qin Zhang |
ICLR | 7 |
| 2023 | When Data Geometry Meets Deep Function: Generalizing Offline Reinforcement Learning
Xianyuan Zhan, Haoran Xu 0003, Ya-Qin Zhang |
ICLR | 6 |
| 2023 | Bit Allocation using OptimizationabstractIn this paper, we consider the problem of bit allocation in Neural Video Compression (NVC). First, we reveal a fundamental relationship between bit allocation in NVC and Semi-Amortized Variational Inference (SAVI). Specifically, we show that SAVI with GoP (Group-of-Picture)-level likelihood is equivalent to pixel-level bit allocation with precise rate & quality dependency model. Based on this equivalence, we establish a new paradigm of bit allocation using SAVI. Different from previous bit allocation methods, our approach requires no empirical model and is thus optimal. Moreover, as the original SAVI using gradient ascent only applies to single-level latent, we extend the SAVI to multi-level such as NVC by recursively applying back-propagating through gradient ascent. Finally, we propose a tractable approximation for practical implementation. Our method can be applied to scenarios where performance outweights encoding speed, and serves as an empirical bound on the R-D performance of bit allocation. Experimental results show that current state-of-the-art bit allocation algorithms still have a room of $\approx 0.5$ dB PSNR to improve compared with ours. Code is available at https://github.com/tongdaxu/Bit-Allocation-Using-Optimization. Tongda Xu, Han Gao 0012, Chenjian Gao, Dailan He, Jinyong Pi, Jixiang Luo, Mao Ye 0001, Hongwei Qin, Yan Wang 0080, Ya-Qin Zhang |
ICML | 13 |
| 2023 | LODE: Locally Conditioned Eikonal Implicit Scene Completion from Sparse LiDARabstractScene completion refers to obtaining dense scene representation from an incomplete perception of complex 3D scenes. This helps robots detect multi-scale obstacles and analyse object occlusions in scenarios such as autonomous driving. Recent advances show that implicit representation learning can be leveraged for continuous scene completion and achieved through physical constraints like Eikonal equations. However, former Eikonal completion methods only demonstrate results on watertight meshes at a scale of tens of meshes. None of them are successfully done for non-watertight LiDAR point clouds of open large scenes at a scale of thousands of scenes. In this paper, we propose a novel Eikonal formulation that conditions the implicit representation on localized shape priors which function as dense boundary value constraints, and demonstrate it works on SemanticKITTI and SemanticPOSS. It can also be extended to semantic Eikonal scene completion with only small modifications to the network architecture. With extensive quantitative and qualitative results, we demonstrate the benefits and drawbacks of existing Eikonal methods, which naturally leads to the new locally conditioned formulation. Notably, we improve IoU from 31.7% to 51.2% on SemanticKITTI and from 40.5% to 48.7% on SemanticPOSS. We extensively ablate our methods and demonstrate that the proposed formulation is robust to a wide spectrum of implementation hyper-parameters. Codes and models are publicly available at https://github.com/AIR-DISCOVER/LODE Pengfei Li 0007, Ruowen Zhao, Yongliang Shi, Hao Zhao 0002, Jirui Yuan, Guyue Zhou, Ya-Qin Zhang |
ICRA | 7 |
| 2023 | Idempotent Learned Image Compression with Right-InverseabstractWe consider the problem of idempotent learned image compression (LIC).
The idempotence of codec refers to the stability of codec to re-compression.
To achieve idempotence, previous codecs adopt invertible transforms such as DCT and normalizing flow.
In this paper, we first identify that invertibility of transform is sufficient but not necessary for idempotence. Instead, it can be relaxed into right-invertibility. And such relaxation allows wider family of transforms.
Based on this identification, we implement an idempotent codec using our proposed blocked convolution and null-space enhancement.
Empirical results show that we achieve state-of-the-art rate-distortion performance among idempotent codecs. Furthermore, our codec can be extended into near-idempotent codec by relaxing the right-invertibility. And this near-idempotent codec has significantly less quality decay after $50$ rounds of re-compression compared with other near-idempotent codecs. Yanghao Li, Tongda Xu, Yan Wang 0105, Ya-Qin Zhang |
NeurIPS | 5 |
| 2023 | Towards Autonomous DrivingabstractThe automotive and transportation industry is going through a tectonic shift in the next decade with the advent of Connectivity, Automation, Sharing, and Electrification (CASE). Autonomous driving presents a historical opportunity to transform the academic, technological, and industrial landscape with advanced sensing and actuation, high definition mapping, new machine learning algorithms, smart planning and control, increasing computing powers, and new infrastructure with 5G, cloud and edge computing. Indeed, we have witnessed unprecedented innovation and activities in the past five years in R&D, investment, joint ventures, road tests and commercial trials, from auto makers, tier-ones, and new forces from the internet and high-tech industries. Ya-Qin Zhang |
WSDM | 1 |
| 2022 | Cerberus Transformer: Joint Semantic, Affordance and Attribute ParsingabstractMulti-task indoor scene understanding is widely considered as an intriguing formulation, as the affinity of different tasks may lead to improved performance. In this paper, we tackle the new problem of Joint semantic, affordance and attribute parsing. However, successfully resolving it requires a model to capture long-range dependency, learn from weakly aligned data and properly balance sub-tasks during training. To this end, we propose an attention-based architecture named Cerberus and a tailored training framework. Our method effectively addresses aforementioned challenges and achieves state-of-the-art performance on all three tasks. Moreover, an in-depth analysis shows concept affinity consistent with human cognition, which inspires us to explore the possibility of weakly supervised learning. Surprisingly, Cerberus achieves strong results using only 0.1%–1% annotation. Visualizations further confirm that this success is credited to common attention maps across tasks. Code and models can be accessed at https://github.com/OPEN-AIR-SUN/Cerberus. Xiaoxue Chen, Tianyu Liu 0008, Hao Zhao 0001, Guyue Zhou, Ya-Qin Zhang |
CVPR | 5 |
| 2022 | TOIST: Task Oriented Instance Segmentation Transformer with Noun-Pronoun DistillationabstractCurrent referring expression comprehension algorithms can effectively detect or segment objects indicated by nouns, but how to understand verb reference is still under-explored. As such, we study the challenging problem of task oriented detection, which aims to find objects that best afford an action indicated by verbs like sit comfortably on. Towards a finer localization that better serves downstream applications like robot interaction, we extend the problem into task oriented instance segmentation. A unique requirement of this task is to select preferred candidates among possible alternatives. Thus we resort to the transformer architecture which naturally models pair-wise query relationships with attention, leading to the TOIST method. In order to leverage pre-trained noun referring expression comprehension models and the fact that we can access privileged noun ground truth during training, a novel noun-pronoun distillation framework is proposed. Noun prototypes are generated in an unsupervised manner and contextual pronoun features are trained to select prototypes. As such, the network remains noun-agnostic during inference. We evaluate TOIST on the large-scale task oriented dataset COCO-Tasks and achieve +10.7% higher $\rm{mAP^{box}}$ than the best-reported results. The proposed noun-pronoun distillation can boost $\rm{mAP^{box}}$ and $\rm{mAP^{mask}}$ by +2.6% and +3.6%. Codes and models are publicly available. Pengfei Li 0007, Beiwen Tian, Yongliang Shi, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang |
NeurIPS | 7 |
| 2008 | Cross-Layer Design for QoS Support in Multihop Wireless NetworksabstractDue to such features as low cost, ease of deployment, increased coverage, and enhanced capacity, multihop wireless networks such as ad hoc networks, mesh networks, and sensor networks that form the network in a self-organized manner without relying on fixed infrastructure is touted as the new frontier of wireless networking. Providing efficient quality of service (QoS) support is essential for such networks, as they need to deliver real-time services like video, audio, and voice over IP besides the traditional data service. Various solutions have been proposed to provide soft QoS over multihop wireless networks from different layers in the network protocol stack. However, the layered concept was primarily created for wired networks, and multihop wireless networks oppose strict layered design because of their dynamic nature, infrastructureless architecture, and time-varying unstable links and topology. The concept of cross-layer design is based on architecture where different layers can exchange information in order to improve the overall network performance. Promising results achieved by cross-layer optimizations initiated significant research activity in this area. This paper aims to review the present study on the cross-layer paradigm for QoS support in multihop wireless networks. Several examples of evolutionary and revolutionary cross-layer approaches are presented in detail. Realizing the new trends for wireless networking, such as cooperative communication and networking, opportunistic transmission, real system performance evaluation, etc., several open issues related to cross-layer design for QoS support over multihop wireless networks are also discussed in the paper. Qian Zhang 0001, Ya-Qin Zhang |
Proc. IEEE | 2 |
| 2008 | Edge-Oriented Uniform Intra PredictionabstractWe propose an intra prediction solution to block-based image compression. In order to adapt to local image features during intra prediction, we consider the distinct image singularities within the model of piece-wise smooth functions. With such singularities, i.e., edges in this paper, intra prediction can be performed by solving Laplace equations. Moreover, since edges exhibit spatial correlations, we design a rate-distortion optimized method for edge extraction and edge coding. Our edge-oriented intra prediction thus consists of the prediction of smooth regions as well as the prediction of edges. We compare our intra prediction with that in H.264 and achieve superior performance. Our intra prediction can also be integrated into a block-based image coding scheme, which is comparable to JPEG2000 in terms of objective quality. An important advantage of our intra prediction is the improvement in visual quality at low bit-rate due to the preservation of edges. Dong Liu 0002, Xiaoyan Sun 0001, Feng Wu 0001, Ya-Qin Zhang |
IEEE Trans. Image Process. | 4 |
| 2007 | Image Compression With Edge-Based InpaintingabstractIn this paper, image compression utilizing visual redundancy is investigated. Inspired by recent advancements in image inpainting techniques, we propose an image compression framework towards visual quality rather than pixel-wise fidelity. In this framework, an original image is analyzed at the encoder side so that portions of the image are intentionally and automatically skipped. Instead, some information is extracted from these skipped regions and delivered to the decoder as assistant information in the compressed fashion. The delivered assistant information plays a key role in the proposed framework because it guides image inpainting to accurately restore these regions at the decoder side. Moreover, to fully take advantage of the assistant information, a compression-oriented edge-based inpainting algorithm is proposed for image restoration, integrating pixel-wise structure propagation and patch-wise texture synthesis. We also construct a practical system to verify the effectiveness of the compression approach in which edge map serves as assistant information and the edge extraction and region removal approaches are developed accordingly. Evaluations have been made in comparison with baseline JPEG and standard MPEG-4 AVC/H.264 intra-picture coding. Experimental results show that our system achieves up to 44% and 33% bits-savings, respectively, at similar visual quality levels. Our proposed framework is a promising exploration towards future image and video compression. Dong Liu 0002, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2007 | Subband Coupling Aware Rate Allocation for Spatial Scalability in 3-D Wavelet Video CodingabstractThe motion compensated temporal filtering (MCTF) technique, which is extensively used in 3-D wavelet video coding schemes nowadays, leads to signal coupling among various spatial subbands because motion alignment is introduced in the temporal filtering. Using all spatial subbands as a reference enables MCTF to fully take advantage of temporal correlation across frames but inevitably brings drifting problem in supporting spatial scalability. This paper first analyzes the signal coupling phenomenon and then proposes a quantitative model to describe signal propagation across spatial subbands during the MCTF process. The signal propagation is modeled for a single MC step based on the shifting effect of wavelet synthesis filters and then it is extended to multilevel MCTF. This model is called subband coupling aware signal propagation (SCASP) model in this paper. Based on the model, we further propose a subband coupling aware rate allocation scheme as one possible solution to the above dilemma in supporting spatial scalability. To find the optimal rate allocation among all subbands for a specified reconstruction resolution, the SCASP model is used to approximate the reconstruction process and derive the synthesis gain of each subband with regard to that reconstruction. Experimental results have fully demonstrated the advantages of our proposed rate allocation scheme in improving both objective and subjective qualities of reconstructed low-resolution video, especially at middle bit rates and high bit rates. Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2006 | Video loss recovery with FEC and stream replicationabstractPacket loss is inevitable in video multicast. In this paper, we propose and study an effective feedback-free loss recovery scheme for layered video which combines forward error correction (FEC) and stream replication. In our scheme, the server multicasts the video in parallel with FEC packets and a number of replicated delayed (ReD) version of the stream. Receivers autonomously and dynamically join the FEC and ReD streams to repair their losses. On the server side, we analyze and optimize the number of replicated streams and FEC packets to meet a certain residual loss requirement (i.e., error after correction). On the receiver side, we analyze the optimal combination of FEC and ReD packets to minimize its loss. We also present a fast yet accurate approximation algorithm for receiver to make such decision. We show that FEC combined with merely one or two replicated streams can effectively reduce the residual error rate (by as much as 50%) as compared with pure FEC or replication alone. Both subjective and objective video measures confirm that our recovery scheme achieves much better visual quality. Shueng-Han Gary Chan, Xing Zheng, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
IEEE Trans. Multim. | 5 |
| 2006 | Optimal stream replication for video simulcastingabstractVideo simulcasting enables a sender to generate replicated streams of different rates, serving receivers of diverse access bandwidths. As replication introduces noticeable redundancy, balancing bandwidth consumption with user satisfaction becomes a critical concern in simulcasting. This paper investigates the above issue; more explicitly, we seek answers to the following two questions: what is the number of streams that should be generated, and what is the bandwidth that should be allocated to each stream? We derive optimal and efficient solutions, and evaluate their performance under a variety of configurations. The results demonstrate that an optimal and adaptive bandwidth allocation significantly improves user satisfaction under stringent resource constraints, and an optimal choice of the stream number yields further improvements. Jiangchuan Liu, Bo Li 0001, Ya-Qin Zhang |
IEEE Trans. Multim. | 3 |
| 2006 | On Routing for Multiple Description Video Over Wireless Ad Hoc NetworksabstractWe study the problem of multipath routing for double description (DD) video in wireless ad hoc networks. We follow an application-centric cross-layer approach and formulate an optimal routing problem that minimizes the application layer video distortion. We show that the optimization problem has a highly complex objective function and an exact analytic solution is not obtainable. However, we find that a meta-heuristic approach such as genetic algorithms (GAs) is eminently effective in addressing this type of complex cross-layer optimization problems. We provide a detailed solution procedure for the GA-based approach. Simulation results demonstrate the superior performance of the GA-based approach versus several other approaches. Our efforts in this work provide an important methodology for addressing complex cross-layer optimization problems, particularly those involved in the application and network layers Shiwen Mao, Y. Thomas Hou 0001, Xiaolin Cheng, Hanif D. Sherali, Scott F. Midkiff, Ya-Qin Zhang |
IEEE Trans. Multim. | 6 |
| 2006 | Region-based rate control and bit allocation for wireless video transmissionabstractIn this paper, we propose a joint source-channel region-based rate control algorithm for real-time video transmissions over wireless systems. During the video transmission, the channel throughput available to the video encoder in the wireless systems is inherently variable, due to the retransmission of the error packets using the automatic repeat request (ARQ) error control. The variable data rate of the wireless system is characterized by the packet-level Gilbert two-state Markov Model, the parameters of which are extracted from the statistical properties of the channel information obtained from the wireless channel simulator. The proposed algorithm adopts a fast but effective block-based segmentation method to extract the regions of interest. Unlike traditional bit allocation methods used in the region/content-based rate control, the algorithm exploits the most effective criteria "coding qualities" as quantitative factors to directly control bit allocation among different regions so as to achieve better visual quality in the regions of interest. The computational complexity of the algorithm is low making it suitable for real-time applications. Compared with the MPEG-4 rate control algorithm, our algorithm can effectively enhance the perceptual quality for the regions of interest and significantly reduce the number of frame skipping; thereby, improve the smoothness of the video. Yu Sun 0003, Ishfaq Ahmad 0001, Dongdong Li 0009, Ya-Qin Zhang |
IEEE Trans. Multim. | 4 |
| 2006 | Lateral error recovery for media streaming in application-level multicastabstractWe consider media streaming using application-level multicast (ALM) where packet loss has to be recovered via retransmission in a timely manner. Since packets may be lost due to congestion, node failures, and join and leave dynamics, traditional "vertical" recovery approach where upstream nodes retransmit the lost packets is no longer effective. We therefore propose lateral error recovery (LER). In LER, hosts are divided into a number of planes, each of which forms an independent ALM tree. Since error correlation across planes is low, a node effectively recovers its error by "laterally" requesting retransmission from nearby nodes in other planes. We present analysis on the complexity and recovery delay on LER. Using Internet-like topologies, we show via simulations that LER is an effective error recovery mechanism. It achieves low overhead in terms of delivery delay (i.e., relative delay penalty) and physical link stress. As compared with traditional recovery schemes, LER attains much lower residual loss rate (i.e., loss rate after retransmission) under a certain deadline constraint. The performance can be substantially improved in the presence of some reliable proxies. Wai-Pun Ken Yiu, Kin Fung Simon Wong, Shueng-Han Gary Chan, Wan-Ching Wong, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
IEEE Trans. Multim. | 7 |
| 2005 | End-to-End QoS for Video Delivery Over Wireless InternetabstractProviding end-to-end quality of service (QoS) support is essential for video delivery over the next-generation wireless Internet. We address several key elements in the end-to-end QoS support, including scalable video representation, network-aware end system, and network QoS provisioning. There are generally two approaches in QoS support: the network-centric and the end-system centric solutions. The fundamental problem in a network-centric solution is how to map QoS criterion at different layers respectively, and optimize total quality across these layers. We first present the general framework of a cross-layer network-centric solution, and then describe the recent advances in network modeling, QoS mapping, and QoS adaptation. The key targets in end-system centric approach are network adaptation and media adaptation. We present a general framework of the end-system centric solution and investigate the recent developments. Specifically, for network adaptation, we review the available bandwidth estimation and efficient video transport protocol; for media adaptation , we describe the advances in error control, power control, and corresponding bit allocation. Finally, we highlight several advanced research directions. Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
Proc. IEEE | 3 |
| 2005 | A proxy-assisted adaptation framework for object video multicastingabstractVideo multicast is a challenging problem due to the heterogeneous and best-effort nature of the Internet. In this paper, we present a novel video multicast framework that exploits the potential of object scalability offered by MPEG-4. Specifically, we introduce the concept of object transmission proxy (OTP), which filters incoming streams using object-based bandwidth adaptation to meet dynamic network conditions. Multiple OPTs can form an overlay network that interconnects diverse multicast islands with semi-uniform demands within each single island. We concur with the wisdom that an application best knows the utility of its data. Hence, the bandwidth-adaptation algorithm for the OTPs adaptively allocates bandwidth among video objects according to their respective utilities and then performs application-level filtering based on an effective stream classification and packetization scheme. Extensive simulation results demonstrate that our framework has substantial performance improvement over conventional bandwidth-adaptation schemes. It is particularly suitable for object-based video multicasting where the objects are of different importance. Jiangchuan Liu, Bo Li 0001, Huai-Rong Shao, Wenwu Zhu 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2005 | Accelerate Video Decoding With Generic GPUabstractMost modern computers or game consoles are equipped with powerful yet cost-effective graphics processing units (GPUs) to accelerate graphics operations. Though the graphics engines in these GPUs are specially designed for graphics operations, can we harness their computing power for more general nongraphics operations? The answer is positive. In this paper, we present our study on leveraging the GPUs graphics engine to accelerate the video decoding. Specifically, a video decoding framework that involves both the central processing unit (CPU) and the GPU is proposed. By moving the whole motion compensation feedback loop of the decoder to the GPU, the CPU and GPU have been made to work in parallel in a pipelining fashion. Several techniques are also proposed to overcome the GPUs constraints or to optimize the GPU computation. Initial experimental results show that significant speed-up can be achieved by utilizing the GPU power. We have achieved real-time playback of high definition video on a PC with an Intel Pentium III 667-MHz CPU and an nVidia GeForce3 GPU. Guobin Shen, Guang-ping Gao, Shipeng Li 0001, Harry Shum, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2005 | Sender-Adaptive and Receiver-Driven Layered Multicast for Scalable Video Over the InternetabstractIn this paper, we propose and analyze a new system architecture for video multicast over Internet, namely, the sender-adaptive and receiver-driven layered multicast (SARLM). In SARLM, the sender of a video source splits the video data coded by a scalable codec and a channel codec into multiple data streams, each of which corresponds to a separate multicast group. The sender can adjust the way in which the video sequence is split dynamically based on the receivers' network parameters collected through feedback. Meanwhile, a receiver can estimate available bandwidth based on a modified packet-pair technique and choose to reassemble and playback the video sequence for a given quality level by dynamically subscribing a given part or all of the data streams according to its network conditions. To optimize the sender's adaptation strategy, we introduce a quality-space (Q-Space) model to describe and analyze the mathematical relationship between the sending rate of different SARLM layers and the video quality received by a given receiver identified by its network characteristics including available bandwidth and packet loss ratio. Our simulation results demonstrate that, under the same network topology and condition, the SARLM architecture can achieve higher network throughput and better video qualities on the receiver side than the existing approaches. Qian Zhang 0001, Quji Guo, Qiang Ni, Wenwu Zhu 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2005 | Video transcoding: an overview of various techniques and research issuesabstractOne of the fundamental challenges in deploying multimedia systems, such as telemedicine, education, space endeavors, marketing, crisis management, transportation, and military, is to deliver smooth and uninterruptible flow of audio-visual information, anytime and anywhere. A multimedia system may consist of various devices (PCs, laptops, PDAs, smart phones, etc.) interconnected via heterogeneous wireline and wireless networks. In such systems, multimedia content originally authored and compressed with a certain format may need bit rate adjustment and format conversion in order to allow access by receiving devices with diverse capabilities (display, memory, processing, decoder). Thus, a transcoding mechanism is required to make the content adaptive to the capabilities of diverse networks and client devices. A video transcoder can perform several additional functions. For example, if the bandwidth required for a particular video is fluctuating due to congestion or other causes, a transcoder can provide fine and dynamic adjustments in the bit rate of the video bitstream in the compressed domain without imposing additional functional requirements in the decoder. In addition, a video transcoder can change the coding parameters of the compressed video, adjust spatial and temporal resolution, and modify the video content and/or the coding standard used. This paper provides an overview of several video transcoding techniques and some of the related research issues. We introduce some of the basic concepts of video transcoding, and then review and contrast various approaches while highlighting critical research issues. We propose solutions to some of these research issues, and identify possible research directions. Ishfaq Ahmad 0001, Xiaohui Wei 0003, Yu Sun 0003, Ya-Qin Zhang |
IEEE Trans. Multim. | 4 |
| 2004 | Layered motion estimation and coding for fully scalable 3d wavelet video codingabstractThis paper proposes a framework of scalable motion estimation and coding with the structure of multilayers for 3D wavelet video coding. The motion representation consists of multiple layers. The encoder uses motion of all layers to perform analysis, while the decoder may receive only part of motion for synthesis. Different from other schemes, each layer of motion is a point optimized at a certain range of bit-rate. We observe that the distortion introduced by motion mismatch is highly independent with the rate for texture in a wide range. Therefore, to make the best trade-off between motion and texture under the constraint of a given bit rate, a motion layer decision algorithm is used to find the appropriate number of motion layers to be included into the bit-stream. The proposed framework also supports the spatial and temporal scalabilities of motion. Experimental results show significant improvement at low bit-rates and nearly no loss at high bit-rates with layered motion coding and optimal motion decision. The performance is approaching to the convex hull of those with multiple sets of nonscalable motion. Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang |
ICIP | 5 |
| 2004 | Lateral Error Recovery for Application-Level MulticastabstractWe consider the delivery of reliable and streaming services using application-level multicast (ALM) by means of UDP, where packet loss has to be recovered via retransmission in a timely manner in order to offer high level of service. Since packets may be lost due to congestion, tree-reconfiguration or node failure, the traditional "vertical" recovery, whereby upstream nodes retransmit the lost packet is no longer effective. We therefore propose and investigate lateral error recovery (LER). In LER, hosts are divided into a number of planes, each of which forms an independent ALM tree. Since the correlation of error among the planes is likely to be low, a node can effectively recover its error "laterally" from nearby nodes in other planes. We employ the technique of global network positioning (GNP) to map the hosts into a coordinate space and identify a set of close neighbors for error recovery by constructing a Voronoi diagram for each plane. We present centralized and distributed algorithm on how to construct the Voronoi diagrams. Using Internet-like topologies, we show via simulations that our system achieves low overheads in terms of relative delay penalty and physical link stress. For reliable service, lateral recovery greatly reduces the average recovery time as compared with vertical recovery schemes. For streaming applications, LER achieves much lower residual loss rate under a certain deadline constraint. Kin Fung Simon Wong, Shueng-Han Gary Chan, Wan-Ching Wong, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
INFOCOM | 6 |
| 2004 | Streaming and Bit Allocation for Scalable Video over Mobile Wireless InternetabstractWith the convergence of wired line Internet and mobile wireless networks as well as tremendous demand on video application in mobile wireless Internet, it's essential to design an effective video streaming protocol and an efficient resource allocation scheme for video delivery over wireless Internet. In this paper, we employ WMSTFP, an end-to-end TCP-friendly multimedia streaming protocol, to detect the status of the wired and wireless part of the wireless Internet, where only the last hop is wireless link. By accurately distinguishing the packet losses due to transmission errors from the congestive losses and smoothing out the pathologic round-trip-time values caused by the highly dynamic wireless environment, in WMSTFP higher throughput in wireless Internet can be achieved and rate can be adjusted in a smooth and TCP-friendly manner. Based upon WMSTFP, we propose a novel loss pattern differentiated bit allocation scheme while applying unequal loss protection (ULP) for scalable video streaming over wireless Internet. Specifically, a rate-distortion (R-D) based bit allocation scheme which considers both wired and wireless network status is proposed to minimize the expected end-to-end distortion. The optimal solution of global optimization for the bit allocation scheme is obtained by a local search algorithm taking the characteristics of progressive fine granularity scalable (PFGS) video into account. Analytical and simulation results demonstrate the effectiveness of our proposed schemes Fan Yang 0024, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
INFOCOM | 4 |
| 2004 | Exploiting temporal correlation with adaptive block-size motion alignment for 3D wavelet codingabstractThis paper proposes an adaptive block-size motion alignment technique in 3D wavelet coding to further exploit temporal correlations across pictures. Similar to B picture in traditional video coding, each macroblock can motion align from forward and/or backward for temporal wavelet de-composition. In each direction, a macroblock may select its partition from one of seven modes - 16x16, 8x16, 16x8, 8x8, 8x4, 4x8 and 4x4 - to allow accurate motion alignment. Furthermore, the rate-distortion optimization criterions are proposed to select motion mode, motion vectors and partition mode. Although the proposed technique greatly improves the accuracy of motion alignment, it does not directly bring the coding efficiency gain because of smaller block size and more block boundaries. Therefore, an overlapped block motion alignment is further proposed to cope with block boundaries and to suppress spatial high-frequency components. The experimental results show the proposed adaptive block-size motion alignment with the overlapped block motion alignment can achieve up to 1.0 dB gain in 3D wavelet video coding. Our 3D wavelet coder outperforms the MC-EZBC for most sequences by 1~2dB and we are doing up to 1.5 dB better than H.264. Ruiqin Xiong, Feng Wu 0001, Shipeng Li 0001, Zixiang Xiong, Ya-Qin Zhang |
VCIP | 5 |
| 2004 | End-to-end TCP-friendly streaming protocol and bit allocation for scalable video over wireless InternetabstractWith the convergence of wired-line Internet and mobile wireless networks, as well as the tremendous demand on video applications in mobile wireless Internet, it is essential to an design effective video streaming protocol and resource allocation scheme for video delivery over wireless Internet. Taking both network conditions in the Internet and wireless networks into account, in this paper, we first propose an end-to-end transmission control protocol (TCP)-friendly multimedia streaming protocol for wireless Internet, namely WMSTFP, where only the last hop is wireless. WMSTFP can effectively differentiate erroneous packet losses from congestive losses and filter out the abnormal round-trip time values caused by the highly varying wireless environment. As a result, WMSTFP can achieve higher throughput in wireless Internet and can perform rate adjustment in a smooth and TCP-friendly manner. Based upon WMSTFP, we then propose a novel loss pattern differentiated bit allocation scheme, while applying unequal loss protection for scalable video streaming over wireless Internet. Specifically, a rate-distortion-based bit allocation scheme which considers both the wired and the wireless network status is proposed to minimize the expected end-to-end distortion. The global optimal solution for the bit allocation scheme is obtained by a local search algorithm taking the characteristics of the progressive fine granularity scalable video into account. Analytical and simulation results demonstrate the effectiveness of our proposed schemes. Fan Yang 0024, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
IEEE J. Sel. Areas Commun. | 4 |
| 2004 | Channel-adaptive resource allocation for scalable video transmission over 3G wireless networkabstractThe paper addresses the important issues of resource allocation for scalable video transmission over third generation (3G) wireless networks. By taking the time-varying wireless channel/network condition and scalable video codec characteristic into account, we allocate resources between source and channel coders based on the minimum-distortion or minimum-power consumption criterion. Specifically, we first present how to estimate the time-varying wireless channel/network condition through measurements of throughput and error rate in a 3G wireless network. Then, we propose a new distortion-minimized bit allocation scheme with hybrid unequal error protection (UEP) and delay-constrained automatic repeat request (ARQ), which dynamically adapts to the estimated time-varying network conditions. Furthermore, a novel power-minimized bit allocation scheme with channel-adaptive hybrid UEP and delay-constrained ARQ is proposed for mobile devices. In our proposed distortion/power-minimized bit-allocation scheme, bits are optimally distributed among source coding, forward error correction, and ARQ according to the varying channel/network condition. Simulation and analysis are performed using a progressive fine granularity scalability video codec. The simulation results show that our proposed schemes can significantly improve the reconstructed video quality under the same network conditions. Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2004 | An end-to-end adaptation protocol for layered video multicast using optimal rate allocationabstractLayered transmission is a promising solution to video multicast over the heterogeneous Internet. However, since the number of layers is practically limited, noticeable mismatches would occur between the coarse-grained layer subscription levels and the heterogeneous and dynamic rate requirements from the receivers. In this paper, we show that such mismatch can be effectively reduced using a dynamic and fine-grained layer rate allocation on the sender's side. Specifically, we study the optimization criteria for rate allocation, and propose a metric called application-aware fairness index. This metric takes into consideration 1) the nonlinear relation between the perceived video quality and the delivered rate and 2) the degree of satisfaction for receivers with heterogeneous bandwidth requirements. We formulate the rate allocation into an optimization problem with the objective of maximizing the expected fairness index for all receivers in a multicast session. We then derive an efficient and scalable solution, and demonstrate that it can be seamlessly integrated into an end-to-end adaptation protocol, called hybrid adaptation layered multicast (HALM). This protocol takes advantage of the emerging fine-grained layered coding, and is fully compatible with the best-effort Internet infrastructure. Simulation and numerical results show that HALM noticeably improves the degree of fairness, and interacts with TCP traffic better than static allocation based protocols. More important, increasing the number of layers in HALM generally improves the degree of fairness; it is sufficient to obtain satisfactory performance with a small number of layers (three to five layers). Jiangchuan Liu, Bo Li 0001, Ya-Qin Zhang |
IEEE Trans. Multim. | 3 |
| 2004 | Seamless switching of scalable video bitstreams for efficient streamingabstractEfficient adaptation to channel bandwidth is broadly required for effective streaming video over the Internet. To address this requirement, a novel seamless switching scheme among scalable video bitstreams is proposed in this paper. It can significantly improve the performance of video streaming over a broad range of bit rates by fully taking advantage of both the high coding efficiency of nonscalable bitstreams and the flexibility of scalable bitstreams, where small channel bandwidth fluctuations are accommodated by the scalability of a single scalable bitstream, whereas large channel bandwidth fluctuations are tolerated by flexible switching between different scalable bitstreams. Two main techniques for switching between video bitstreams are proposed. Firstly, a novel coding scheme is proposed to enable drift-free switching at any frame from the current scalable bitstream to one operated at lower rates without sending any overhead bits. Secondly, a switching-frame coding scheme is proposed to greatly reduce the number of extra bits needed for switching from the current scalable bitstream to one operated at higher rates. Compared with existing approaches, such as switching between nonscalable bitstreams and streaming with a single scalable bitstream, our experimental results clearly show that the proposed scheme brings higher efficiency and more flexibility in video streaming. Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001, Ya-Qin Zhang |
IEEE Trans. Multim. | 5 |
| 2004 | Peer-to-peer based multimedia distribution serviceabstractRecently, there are many research interests in providing efficient and scalable multimedia distribution service. However, stringent quality-of-service (QoS) requirements for media distribution, as well as dynamically changing and heterogeneous network capacity in today's best effort Internet, bring many challenges. In this paper, we introduce a novel framework for multimedia distribution service based on peer-to-peer (P2P) networks. A topology-aware overlay is proposed in which hosts self-organize into groups. End hosts within the same group have similar network conditions and can easily collaborate with each other to achieve QoS awareness. In order to improve media delivery quality and provide high service availability, we further propose two distributed heuristic replication strategies, intergroup replication and intragroup replication, based on this topology-aware overlay. Specifically, intergroup replication is aimed to improve the efficiency of media content delivery between the group where a request is issued and the group where the content is stored. Also, intragroup replication is targeted at improving the availability of the content. Extensive simulation results show that the latency in our proposed architecture is 20% less than that of the FreeNet and 50% less than that of the randomly replication system. Simulation results also show that the video quality in our system is much better than that in the other two systems. Our P2P-based approach is also distributed, scalable, cost effective, and aware of the performance. Zhe Xiang, Qian Zhang 0001, Wenwu Zhu 0001, Zhensheng Zhang, Ya-Qin Zhang |
IEEE Trans. Multim. | 5 |
| 2003 | Accelerating video decoding using GPUabstractMost modern computers or game consoles are equipped with powerful graphics processing units (GPU) to accelerate graphics operations. There is a trend that the power of GPU outgrows that of the CPU (central processing unit). However, the GPU engines are specially designed for graphics operations. Can we take advantage of the powerful GPU engines for more general operations other than pure graphics operations? The answer is positive. In this study, we present schemes that map other non-graphics operations into graphics engines with an example application of accelerating video decoding with the assistance of GPU. Our results show that significant speed-up can be achieved by leveraging the GPU power. Specifically, we have achieved real-time playback of high definition video on a PC with an Intel Pentium III 667 MHz CPU and an nVidia GeForce3 GPU. Guobin Shen, Lihua Zhu, Shipeng Li 0001, Harry Shum, Ya-Qin Zhang |
ICASSP (4) | 5 |
| 2003 | Advances in networked media - theory and practiceabstractSummary form only given. We have witnessed the increasing convergence of digital media, wireless and networking technologies in the past decade. This has profoundly transformed the way media is being represented, processed, delivered, and presented. For over half a century, Shannon's rate-distortion (R-D) theory has been the theoretical foundation for information representation. With the emergence of new media, devices and applications in networked environment, the conventional R-D theory needs to be extended to enable more effective representation and processing of connected media. This article summarizes our attempt to develop a 'networked R-D theory' as well as some initial applications. In particular, it addresses the following research initiatives currently undertaken by Microsoft Research Asia: (a) new sampling and rendering structure for computer graphics and digital ink; (b) media delivery over a network with errors, congestion, retransmission, and multi-user interaction; (c) and media summarization where maximum information could be extracted for a given time boundary. Some of these technologies have already been transferred into Microsoft's mainstream products, which will help enable a plethora of applications such as high-quality media streaming, intelligent note-taking, networked games, and home audio/photo/video editing. Ya-Qin Zhang |
ICASSP (1) | 1 |
| 2003 | An efficient algorithm for adaptive cell sectoring in CDMA systemsabstractCell sectorization is a promising method for improving the capacity of code division multiple access (CDMA) systems. It has been shown that the use of adaptive antenna arrays with dynamic cell sectoring is particularly suitable for non-uniformly distributed users. In this paper, we present a novel cluster-based sectoring algorithm for adaptive cell sectorization. The complexity of our algorithms, and more important, in a high-density case, it does not depend on the number of users in a cell. Extensive simulations show that the performance of our solution is comparable to the optimal solutions. Jihui Zhang 0002, Jiangchuan Liu, Qian Zhang 0001, Wenwu Zhu 0001, Bo Li 0001, Ya-Qin Zhang |
ICC | 6 |
| 2003 | Feedback-free packet loss recovery for video multicastabstractIn video streaming over multicast networks, error recovery is essential to alleviate the effect of packet loss. In this paper, we study a feedback-free recovery scheme which combines the strength of FEC (namely, a parity packet can repair any lost packet) and pseudo-ARQ (namely, incremental recovery). To account for the receiver heterogeneity, the server multicasts layered video streams for the receivers to join. For each layer, the receivers may join dynamically additional multicast channels of FEC and pseudo-ARQ packets to recover local losses. In order to offer quality video, we address the following issues: 1) "menu creation": given a certain maximum error rate after correction (i.e., residual error rate), what is the combination of FEC and pseudo-ARQ packets for the server to send so as to minimize a target receiver's bandwidth; and 2) "menu selection": given the server's menu, what's the combination of FEC and pseudo-ARQ packets for the receiver to join so as to minimize its residual error rate given its local probability, bandwidth and loss pattern. We present the analysis of the scheme and show that our scheme can substantially reduce a receiver's residual error rate as compared with pure FEC or pure pseudo-ARQ alone (by cutting it as much as half). Xing Zheng, Shueng-Han Gary Chan, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
ICC | 5 |
| 2003 | On the rate constraint of transmitting multiple priority classes with QoSabstractThe rate constraint of transmitting multiple priority classes over a time-varying service-rate channel is studied in this work. This constraint specifies the maximum data rate that can be transmitted reliably with QoS (quality of service) guarantee. In our framework, the time-varying service channel is modeled by an N-state discrete Markov process, where each Markov state is associated with a channel service rate, and the absolute priority scheduling is used to transport packets of different classes. The transmission rate constraint is derived based on effective bandwidth and capacity. To be more specific, given channel parameters and the maximum buffer size for each priority class, statistical QoS guarantees in terms of packet loss probabilities can be determined and translated to the transmission rate constraint. The derived result is verified by simulation in a time-varying wireless environment. Wuttipong Kumwilaisak, Qian Zhang 0001, Wenwu Zhu 0001, C.-C. Jay Kuo, Ya-Qin Zhang |
ICME | 5 |
| 2003 | An end-to-end TCP-friendly streaming protocol for multimedia over wireless InternetabstractWith the convergence of wired line Internet and mobile wireless networks, it's important to study its impacts on continuous media delivery and media streaming protocols. In this paper, we propose an end-to-end (wireless) multimedia streaming TCP-friendly protocol for media delivery over wireless Internet (WMSTFP). WMSTFP can effectively differentiate erroneous packet losses from congestive losses and filter out the abnormal round-trip-time values due to the highly varying wireless environment. Analytical and simulation results show that WMSTFP can achieve higher throughput in wireless Internet and can perform rate adjustment in a smooth and TCP-friendly manner. Fan Yang 0024, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
ICME | 4 |
| 2003 | A cross-Layer quality-of-service mapping architecture for video delivery in wireless networksabstractProviding quality-of-service (QoS) to video delivery in wireless networks has attracted intensive research over the years. A fundamental problem in this area is how to map QoS criterion at different layers and optimize QoS across the layers. In this paper, we investigate this problem and present a cross-layer mapping architecture for video transmission in wireless networks. There are several important building blocks in this architecture, among others, QoS interaction between video coding and transmission modules, QoS mapping mechanism, video quality adaptation, and source rate constraint derivation. We describe the design and algorithms for each building block, which either builds upon or extend the state-of-the-art algorithms that were developed without much considerations of other layers. Finally, we use simulation results to demonstrate the performance of the proposed architecture for progressive fine granularity scalability video transmission over time-varying and nonstationary wireless channel. Wuttipong Kumwilaisak, Y. Thomas Hou 0001, Qian Zhang 0001, Wenwu Zhu 0001, C.-C. Jay Kuo, Ya-Qin Zhang |
IEEE J. Sel. Areas Commun. | 6 |
| 2003 | Scalable portrait video for mobile video communicationabstractWireless networks have been rapidly developing in recent years. General Packet Radio Service (GPRS) and Code Division Multiple Access (CDMA 1X) for wide areas, and 802.11 and Bluetooth for local areas have already emerged. Broadband wireless networks urgently call for rich contents for consumers. Among various possible applications, video communication is one of the most promising for mobile devices on wireless networks. This paper describes the generation, coding, and transmission of an effective video form, scalable portrait video for mobile video communication. As an expansion to bilevel video, portrait video is composed of more gray levels, and therefore possesses higher visual quality while it maintains a low bit rate and low computational costs. Portrait video is a scalable video in that each video with a higher level always contains all the information of the video with a lower level. The bandwidths of 2-4-level portrait videos fit into the bandwidth range of 20-40 kbps that GPRS and CDMA 1X can stably provide; therefore, portrait video is very promising for video broadcast and communication on 2.5-G wireless networks. With portrait video technology, we are the first to enable two-way video communication on pocket PCs and handheld PCs. Jiang Li 0008, Keman Yu, Tielin He, Yunfeng Lin, Shipeng Li 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2003 | QoS-adaptive proxy caching for multimedia streaming over the InternetabstractThis paper proposes a quality-of-service (QoS)-adaptive proxy-caching scheme for multimedia streaming over the Internet. Considering the heterogeneous network conditions and media characteristics, we present an end-to-end caching architecture for multimedia streaming. First, a media-characteristic-weighted replacement policy is proposed to improve the cache hit ratio of mixed media including continuous and noncontinuous media. Secondly, a network-condition- and media-quality-adaptive resource-management mechanism is introduced to dynamically re-allocate cache resource for different types of media according to their request patterns. Thirdly, a pre-fetching scheme is described based on the estimated network bandwidth, and a miss strategy to decide what to request from the server in case of cache miss based on real-time network conditions is presented. Lastly, request and send-back scheduling algorithms, integrating with unequal loss protection (ULP), are proposed to dynamically allocate network resource among different types of media. Simulation results demonstrate effectiveness of our proposed schemes. Fang Yu 0002, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2002 | Efficient and universal scalable video codingabstractThis paper proposes a unified efficient and universal scalable video coding framework that supports different scalabilities, such as fine granularity quality, temporal, spatial and complexity scalabilities. The proposed framework is established upon the recent studies in fine granularity scalable (FGS) video coding. It contains two key points. Firstly, in order to improve the coding efficiency of the proposed framework, more than one motion compensation loop is used. Since high quality references are introduced into the enhancement layer coding, the proposed framework can efficiently compress different-resolution video at different layers for the purpose of the complexity and spatial scalability. Secondly, the drifting reduction techniques are studied in this paper. This helps the proposed framework to maintain good performance at lower enhancement bit rates. By defining coding modes, a macroblock level control mechanism is developed to achieve a better trade-off between low drifting errors and high coding efficiency. Feng Wu 0001, Shipeng Li 0001, Xiaoyan Sun 0001, Ya-Qin Zhang |
ICIP (2) | 5 |
| 2002 | A Hybrid Adaptation Protocol for TCP-Friendly Layered Multicast and Its Optimal Rate AllocationabstractLayered transmission has been proposed as a solution to video multicast over the Internet. Existing protocols usually perform adaptation at the receiver and use static rate allocation techniques at the sender. As a result, significant mismatches between the fixed transmission rates and the heterogeneous and dynamic rate requirements for the receivers can occur. We show that such mismatches can be minimized by employing dynamic layer rate allocation at the sender by taking advantage of the recent development in layered video coding. Specifically, we study the optimization criteria for layer rate allocation, and propose a metric, called fairness index, which fairly reflects the degree of a receiver's satisfaction. We then formulate this into an optimization problem with the objective of maximizing the expected fairness index, and derive an efficient and scalable algorithm to solve it. We further demonstrate that such sender rate adaptation can be seamlessly integrated into an end-to-end adaptation protocol called HALM (hybrid adaptation layered multicast). This protocol is designed for the current best-effort Internet and is TCP-friendly. Its control overhead is also kept at a low level. Simulation results show that HALM improves the degree of fairness for receivers with heterogeneous bandwidth requirements, and interacts with TCP flows substantially better than static allocation based protocols. In addition, increasing the number of layers in HALM always leads to a higher degree of fairness and 3 to 5 layers are usually sufficient. However, this is not true for the static allocation based protocols. Jiangchuan Liu, Bo Li 0001, Ya-Qin Zhang |
INFOCOM | 3 |
| 2002 | Allocation of layer bandwidths and FECs for video multicast over wired and wireless networksabstractLayered multicast is an efficient technique to deliver video to heterogeneous receivers over wired and wireless networks. We consider such a multicast system in which the server adapts the bandwidth and forward-error correction code (FEC) of each layer so as to maximize the overall video quality, given the heterogeneous client characteristics in terms of their end-to-end bandwidth, packet drop rate over the wired network, and bit-error rate in the wireless hop. In terms of FECs, we also study the value of a gateway which "transcodes" packet-level FECs to byte-level FECs before forwarding packets from the wired network to the wireless clients. We present an analysis of the system, propose an efficient algorithm on FEC allocation for the base layer, and formulate a dynamic program with a fast and accurate approximation for the joint bandwidth and FEC allocation of the enhancement layers. Our results show that a transcoding gateway performs only slightly better than the nontranscoding one in terms of end-to-end loss rate, and our allocation is effective in terms of FEC parity and bandwidth served to each user. T. W. Angus Lee, Shueng-Han Gary Chan, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2002 | Memory-constrained 3D wavelet transform for video coding without boundary effectsabstractThree-dimensional (3D) wavelet-based scalable video coding provides a viable alternative to standard MC-DCT coding. However, many current 3D wavelet coders experience severe boundary effects across group of pictures (GOP) boundaries. This paper proposes a memory-efficient transform technique via lifting that effectively computes wavelet transforms of a video sequence continuously on the fly, thus eliminating the boundary effects due to limited length of individual GOPs. Coding results show that the proposed scheme completely eliminates the boundary effects and gives superb video playback quality. Jizheng Xu, Zixiang Xiong, Shipeng Li 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2002 | Power-minimized bit allocation for video communication over wireless channelsabstractVideo communication over wireless links using handheld devices is a challenging task due to the time-varying characteristics of the wireless channels and limited battery resources. Rate-distortion (RD) analysis plays a key role in video coding and communication systems, and usually the RD relation does not assume any power constraint. We investigate the relations of rate, distortion and power consumption. Based on those relations, we propose a power-minimized bit-allocation scheme considering the processing power, for source coding and channel coding, jointly with the transmission power. The total bits are allocated between source and channel coders, according to wireless channel conditions and video quality requirements, to minimize the total power consumption for a single user and a group of users in a cell, respectively. Simulation results show that our proposed joint power-control and bit-allocation scheme achieves high power savings compared to the conventional scheme. Qian Zhang 0001, Zhu Ji, Wenwu Zhu 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2002 | 3-D wavelet compression and progressive inverse wavelet synthesis rendering of concentric mosaicabstractUsing an array of photo shots, the concentric mosaic offers a quick way to capture and model a realistic three-dimensional (3-D) environment. We compress the concentric mosaic image array with a 3-D wavelet transform and coding scheme. Our compression algorithm and bitstream syntax are designed to ensure that a local view rendering of the environment requires only a partial bitstream, thereby eliminating the need to decompress the entire compressed bitstream before rendering. By exploiting the ladder-like structure of the wavelet lifting scheme, the progressive inverse wavelet synthesis (PIWS) algorithm is proposed to maximally reduce the computational cost of selective data accesses on such wavelet compressed datasets. Experimental results show that the 3-D wavelet coder achieves high-compression performance. With the PIWS algorithm, a 3-D environment can be rendered in real time from a compressed dataset. Lin Luo 0004, Yunnan Wu, Jin Li 0001, Ya-Qin Zhang |
IEEE Trans. Image Process. | 4 |
| 2001 | Optimal allocation of packet-level and byte-level FEC in video multicasting over wired and wireless networksabstractMulticast is an efficient technique to deliver video content over a network. We consider such a multicast system to serve both wireless and wireline users when there are errors over the wired network and the wireless hop. Since packets are likely to be dropped in the wired networks while bit errors are more likely over the wireless hop, a combination of both packet-level and byte-level FEC is required to recover these errors. Given the estimated error and bandwidth characteristics reported by end users, the server needs to allocate optimally the packet-level and byte-level FEC to achieve maximum video quality. We study two schemes pertaining to whether or not the wireless gateway is able to transcode the video packets from the wired network before forwarding them to the wireless users. We first develop a model to analyze the system; and then propose an efficient algorithm for the FEC computation. We finally compare the schemes in terms of the optimal parameters used in the FEC, and the video quality achieved. T. W. Angus Lee, Shueng-Han Gary Chan, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
GLOBECOM | 5 |
| 2001 | A scalable bit-stream packetization and adaptive rate control framework for object-based video multicastabstractLayered and agent-based transmission schemes have been recommended as the solutions to the inherent heterogeneity problem of video multicast over the Internet. Such schemes either require the integration of a scalable codec and a layered transport protocol, or involve computationally intensive transcoding at the application level. In addition, how to support user interactions for object-based video applications (e.g. MPEG-4) has seldom been considered. To address these problems, we present a novel framework specifically designed for object-based video multicast. First, we propose a novel bit-stream re-organization and packetization scheme by taking: advantage of the error resilience and concealment features in the MPEG-4 object-based codec. Second, we introduce the concept of a Video Transmission Agent (VTA) that eliminates the need for application level transcoding and works well with network-level rate control algorithm. We demonstrate that the combination of such approaches not only enables adaptive rate control more efficiently but also significantly reduces system complexity. Jiangchuan Liu, Huai-Rong Shao, Bo Li 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
GLOBECOM | 5 |
| 2001 | Fine-granularity spatially scalable video codingabstractWe propose a novel architecture for spatially scalable video coding, namely, fine-granularity spatially scalable (FGSS) coding. The traditional layered spatially scalable coding provides only coarse scalability in which the bit-stream can be decoded only at a few fixed resolutions, but not something in between. The proposed FGSS scheme provides a fine-granularity property to the spatial scalability. In this scheme, the bit plane technique is combined with spatial scalability, thus a fine granularity increase in the image quality from low-resolution to high-resolution can be obtained. In addition, the proposed scheme provides a flexible embedded bitstream that can be decoded up to any point in the enhancement layer bitstream from low-resolution to high-resolution. This feature further enables efficient video streaming over the Internet where the scalable bitstream, can adapt to the widely fluctuating bandwidth. The FGSS coding scheme extends new functionalities such as multi-resolution, fine granularity, channel adaptation and error-recovery properties to scalable video coding, thus it can satisfy different user clients with a wide range of channel bandwidth and screen resolution. Feng Wu 0001, Shipeng Li 0001, Yuzhuo Zhong, Ya-Qin Zhang |
ICASSP | 5 |
| 2001 | On the optimal rate allocation for layered video multicastabstractLayered transmission has been recommended as one of the solutions to the inherent heterogeneity problem of video multicast over the Internet. Unlike traditional sender-driven adaptation schemes, current layered multicast protocols usually use distributed control algorithms which perform adaptation on the receiver side. With the assumption that the layer rates are fixed, the control granularity of such protocols is considerably coarse and there is usually a mismatch between the fixed rates and the requirements of the receivers. We argue that dynamically allocating layer rates by the sender can minimize this mismatch. We study the optimization criteria for layer rate allocation, and propose a novel metric, application-aware fairness index, to fairly reflect the degree of the receiver's satisfaction. We then formulate the problem of optimal allocation for heterogeneous receivers, and derive an efficient algorithm for solving it. The algorithm is highly scalable and outperforms the traditional static allocation schemes by 10% or more. Jiangchuan Liu, Kin-Man Cheung, Bo Li 0001, Ya-Qin Zhang |
ICCCN | 4 |
| 2001 | Macroblock-based progressive fine granularity scalable (PFGS) video coding with flexible temporal-SNR scalablilitiesabstractWe proposed a flexible and efficient architecture for scalable video coding, namely, the macroblock (MB)-based progressive fine granularity scalable video coding with temporal-SNR scalabilities (PFGST). The proposed architecture can provide not only much improved coding efficiency but also simultaneous SNR scalability and temporal scalability. Building upon the original frame-based progressive fine granularity scalable (PFGS) coding approach, the MB-based PFGS scheme is first proposed. Three INTER modes and the corresponding mode selection mechanism are presented for coding the SNR enhancement MBs in order to make a good trade-off between low drifting errors and high compression efficiency. Furthermore, temporal scalability is introduced into the MB-based PFGS, which forms the MB-based PFGST scheme. Two coding modes are proposed for coding the temporal enhancement MBs. Since it would not cause any error propagation if using the high quality reference in the temporal enhancement MB coding, the coding efficiency of the PFGST is highly improved by always choosing the most suitable reference for the temporal scalable coding. Experimental results show that the MB-based PFGST video coding scheme can significantly improve the coding efficiency up to 2.8 dB compared with the FGST scheme adopted in MPEG-4, while supporting full SNR, full temporal, and hybrid SNR-temporal scalabilities according to the different requirements from the channels, the clients or the servers. Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001, Ya-Qin Zhang |
ICIP (2) | 5 |
| 2001 | Network-adaptive scalable video streaming over 3G wireless networkabstractScalable video streaming over a wireless link with quality of service (QoS) is a very challenging task due to the time-varying characteristics of the wireless channel and limited battery resource in handheld devices. This paper proposes an end-to-end network-adaptive architecture for video streaming over 3G wireless networks. The proposed architecture not only dynamically estimates the varying network status on the fly, but also simultaneously performs application-level error control and transmission-level power control to protect against the random and fading error occurring across the 3G network. Distortion/power-minimized rate allocation is presented to achieve the minimum end-to-end distortion or minimum total power consumption based on the user's requirement. Simulation results demonstrate the effectiveness of our proposed scheme. Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
ICIP (3) | 3 |
| 2001 | Motion Compensated Lifting Wavelet And Its Application In Video CodingabstractA motion compensated lifting (MCLIFT) framework is proposed for the 3D wavelet video coder. By using bi-directional motion compensation in each lifting step of the temporal direction, the video frames are effectively de-correlated. With proper entropy coding and bitstream packaging schemes, the MCLIFT wavelet video coder can be scalable in frame rate and quality level. Experimental results show that the MCLIFT video coder outperforms the 3D wavelet video coder with the same entropy coding scheme by an average of 1.1-1.6dB, and outperforms MPEG-4 coder by an average of 0.9-1.4dB. Lin Luo 0004, Jin Li 0001, Shipeng Li 0001, Zhenquan Zhuang, Ya-Qin Zhang |
ICME | 5 |
| 2001 | An End-to-End Probing-Based Admission Control Scheme for Multimedia ApplicationsabstractThis paper proposes a new probing-based admission control scheme in which delay is taken as important admission criterion in conjunction with packet loss rate. The delay feature obtained by probing allows end systems to detect the network state fluctuation more sensitively and quickly without overloading the network. Our proposed approach is able to not only keep very low loss rate, but also provide end-to-end delay upper bound more accurately. Simulation results demonstrate the effectiveness of our approach. Jianming Qiu, Huai-Rong Shao, Wenwu Zhu 0001, Ya-Qin Zhang |
ICME | 4 |
| 2001 | Macroblock-Based Progressive Fine Granularity Scalable Video CodingabstractIn this paper, we proposed a flexible and efficient architecture for scalable video coding, namely, the macroblock (MB)-based progressive fine granularity scalable video coding with temporal-SNR scalabilities (PFGST in short). The proposed architecture can provide not only much improved coding efficiency but also simultaneous SNR scalability and temporal scalability. Building upon the original frame-based progressive fine granularity scalable (PFGS) coding approach, the MB-based PFGS scheme is first proposed. Three INTER modes and the corresponding mode selection mechanism are presented for coding the SNR enhancement MBs in order to make a good trade-off between low drifting errors and high compression efficiency. Furthermore, temporal scalability is introduced into the MB-based PFGS, which forms the MB-based PFGST scheme. Two coding modes are proposed for coding the temporal enhancement MBs. Since it would not cause any error propagation if using the high quality reference in the temporal enhancement MB coding, the coding efficiency of the PFGST is highly improved by always choosing the most suitable reference for the temporal scalable coding. Experimental results show that the MB-based PFGST video coding scheme can significantly improve the coding efficiency up to 2.8dB compared with the FGST scheme adopted in MPEG-4, while supporting full SNR, full temporal, and hybrid SNR-temporal scalabilities according to the different requirements from the channels, the clients or the servers. 1. Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001, Ya-Qin Zhang |
ICME | 5 |
| 2001 | End-to-end power-optimized video communication over wireless channelsabstractVideo communication over a wireless link is a very challenging task due to the time-varying characteristics of the wireless channel and limited battery resources in the handheld devices. This paper proposes an end-to-end power-optimized approach for video communication. The power consumed in complexity-reconfigurable source coding, channel-adaptive unequal error protection (UEP), and video transmission is analyzed. Power-minimized rate allocation is presented to achieve the minimal power consumption in the mobile host while maintaining the desired video quality. Simulation results show that our power-optimized scheme achieves significant power saving compared to the existing scheme. Zhu Ji, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
MMSP | 4 |
| 2001 | Channel-adaptive unequal error protection for scalable video transmission over wireless channel
Guijin Wang, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
VCIP | 4 |
| 2001 | Joint power control and source-channel coding for video communication over wireless networksabstractVideo communication over wireless link is a challenging task due to the time-varying characteristics of a wireless channel and limited battery resource in the handheld devices. This paper proposes a power-optimized approach for video communication, which simultaneously controls the transmission power, source rate and error protection level to minimize the total power consumption for all users. The performance of the cellular CDMA system using our proposed scheme is compared with the one using a fixed power-control scheme. The simulation results show that our proposed joint power control and source-channel coding scheme achieves significant power saving compared to the fixed scheme. Zhu Ji, Qian Zhang 0001, Wenwu Zhu 0001, Jianhua Lu, Ya-Qin Zhang |
VTC Fall | 5 |
| 2001 | Scalable video coding and transport over broadband wireless networksabstractWith the emergence of broadband wireless networks and increasing demand of multimedia information on the Internet, wireless multimedia services are foreseen to become widely deployed in the next decade. Real-time video transmission typically has requirements on quality of service (QoS). However, wireless channels are unreliable and the channel bandwidth varies with time, which may cause severe degradation in video quality. In addition, for video multicast, the heterogeneity of receivers makes it difficult to achieve efficiency and flexibility. To address these issues, three techniques, namely, scalable video coding, network-aware adaptation of end systems, and adaptive QoS support from networks, have been developed. This paper unifies the three techniques and presents an adaptive framework, which specifically addresses video transport over wireless networks. The adaptive framework consists of three basic components: (1) scalable video representations; (2) network-aware end systems; and (3) adaptive services. Under this framework, as wireless channel conditions change, mobile terminals and network elements can scale the video streams and transport the scaled video streams to receivers with a smooth change of perceptual quality. The key advantages of the adaptive framework are: (1) perceptual quality is changed gracefully during periods of QoS fluctuations and hand-offs; and (2) the resources are shared in a fair manner. Dapeng Oliver Wu, Y. Thomas Hou 0001, Ya-Qin Zhang |
Proc. IEEE | 3 |
| 2001 | User-aware object-based video transmission over the next generation Internet
Huai-Rong Shao, Wenwu Zhu 0001, Ya-Qin Zhang |
Signal Process. Image Commun. | 3 |
| 2001 | Streaming video over the Internet: approaches and directionsabstractDue to the explosive growth of the Internet and increasing demand for multimedia information on the Web, streaming video over the Internet has received tremendous attention from academia and industry. Transmission of real-time video typically has bandwidth, delay, and loss requirements. However, the current best-effort Internet does not offer any quality of service (QoS) guarantees to streaming video. Furthermore, for video multicast, it is difficult to achieve both efficiency and flexibility. Thus, Internet streaming video poses many challenges. In this article we cover six key areas of streaming video. Specifically, we cover video compression, application-layer QoS control, continuous media distribution services, streaming servers, media synchronization mechanisms, and protocols for streaming media. For each area, we address the particular issues and review major approaches and mechanisms. We also discuss the tradeoffs of the approaches and point out future research directions. Dapeng Oliver Wu, Y. Thomas Hou 0001, Wenwu Zhu 0001, Ya-Qin Zhang, Jon M. Peha |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2001 | A framework for efficient progressive fine granularity scalable video codingabstractA basic framework for efficient scalable video coding, namely progressive fine granularity scalable (PFGS) video coding is proposed. Similar to the fine granularity scalable (PGS) video coding in MPEG-4, the PFGS framework has all the features of FGS, such as fine granularity bit-rate scalability, channel adaptation, and error recovery. On the other hand, different from the PGS coding, the PFGS framework uses multiple layers of references with increasing quality to make motion prediction more accurate for improved video-coding efficiency. However, using multiple layers of references with different quality also introduces several issues. First, extra frame buffers are needed for storing the multiple reconstructed reference layers. This would increase the memory cost and computational complexity of the PFGS scheme. Based on the basic framework, a simplified and efficient PFGS framework is further proposed. The simplified PPGS framework needs only one extra frame buffer with almost the same coding efficiency as in the original framework. Second, there might be undesirable increase and fluctuation of the coefficients to be coded when switching from a low-quality reference to a high-quality one, which could partially offset the advantage of using a high-quality reference. A further improved PFGS scheme can eliminate the fluctuation of enhancement-layer coefficients when switching references by always using only one high-quality prediction reference for all enhancement layers. Experimental results show that the PFGS framework can improve the coding efficiency up to more than 1 dB over the FGS scheme in terms of average PSNR, yet still keeps all the original properties, such as fine granularity, bandwidth adaptation, and error recovery. A simple simulation of transmitting the PFGS video over a wireless channel further confirms the error robustness of the PFGS scheme, although the advantages of PFGS have not been fully exploited. Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2001 | Arbitrarily shaped video-object coding by waveletabstractVideo-object coding is one of the most important functionalities proposed by MPEG-4. We propose a new wavelet method to encode the texture of an arbitrarily shaped object, both for the still and for the video-object. The method uses the shape adaptive wavelet transform (SA-DWT) in MPEG-4 still object coding, but with a computationally more efficient lifting implementation. The transformed object coefficients are then quantized and entropy encoded with a partial bit-plane embedded coder, which greatly improves the coding efficiency. We denote the coding algorithm as the video-object wavelet (VOW) coder. Experimental results show that VOW significantly outperforms MPEG-4 in still-object coding, and achieves a comparable performance in video-object coding in terms of PSNR. Moreover, the VOW decoded object looks better subjectively, with less annoying blocking artifacts than that of MPEG-4. Guiwei Xing, Jin Li 0001, Shipeng Li 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2001 | Resource allocation for multimedia streaming over the InternetabstractThis paper addresses the resource allocation problem for multiple media streaming over the Internet. First, we present an end-to-end transport architecture for multimedia streaming over the Internet. Second, we propose a new multimedia streaming TCP-friendly protocol (MSTFP), which combines forward estimation of network conditions with information feedback control to optimally track the network conditions. Third, we propose a novel resource allocation scheme to adapt media rate to the estimated network bandwidth using each media's rate-distortion function under various network conditions. By dynamically allocating resources according to network status and media characteristics, we improve the end-to-end quality of services (QoS). Simulation results demonstrate the effectiveness of our proposed schemes. Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
IEEE Trans. Multim. | 3 |
| 2000 | Optimal Mode Selection in Internet Video Communication: An End-To-End ApproachabstractWe present an end-to-end approach to generalize the classical theory of rate distortion (R-D) optimized mode selection for point-to-point video communication. We introduce a notion of global distortion by taking into consideration of both the path characteristics and the receiver behavior, in addition to the source behavior. We derive, for the first time, a set of accurate global distortion metrics for any packetization scheme. Equipped with the global distortion metrics, we design an R-D optimized mode selection algorithm to provide the best trade-off between compression efficiency and error resilience. As an application, we integrate our theory with point-to-point MPEG-4 video conferencing over the Internet. Simulation results conclusively demonstrate that our end-to-end approach offers superior performance over the classical approach for Internet video conferencing. Dapeng Oliver Wu, Y. Thomas Hou 0001, Ya-Qin Zhang, H. Jonathan Chao |
ICC (1) | 3 |
| 2000 | On the Compression of Image Based Rendering SceneabstractIn image based rendering (IBR), a 3D scene is recorded through a set of photos, and a novel view is rendered by assembling data from the photo set. Compression is essential to reduce the huge data amount of IBR. We examine three categories of IBR compression algorithms: the block coder, the reference coder and the high dimensional transform (wavelet) coder. It is observed that the block coder consumes the least computation resource, however, its compression ratio is low. The reference coder achieves a good compression ratio with reasonable computation complexity. The high dimensional wavelet coder achieves the best compression ratio, however, it is also the most complex. Jin Li 0001, Harry Shum, Ya-Qin Zhang |
ICIP | 3 |
| 2000 | Scalable Object-Based Video Multicasting over the InternetabstractThis paper presents a bitstream classification, prioritization, and transmission scheme specifically designed for MPEG-4 video consisting of multiple video objects (MVO). Different types of data in the compressed bitstream such as shape, motion, and texture are re-assembled and assigned to different priority classes. We then propose a heterogeneous multicasting scheme for efficient packet delivery in a multi-party environment. Our proposed multicasting method is a single-layer solution with the advantage of multi-layer multicasting without the need of separate network sessions. Performance evaluation and test results indicate the advantages of the proposed scalable video transmission and multicasting scheme. Huai-Rong Shao, Wenwu Zhu 0001, Ya-Qin Zhang |
ICIP | 3 |
| 2000 | DCT-Prediction Based Progressive Fine Granularity Scalability CodingabstractWe propose a novel architecture for scalable video coding, namely, progressive fine granularity scalable (PEGS) coding, which can provide a high coding efficiency along with good bandwidth adaptation and error recovery properties. Unlike the fine granularity scalable (FGS) coding in the MPEG-4 proposal, some of the enhancement layers in a current frame are predicted from a high quality enhancement layer in a reference frame, rather than always from the base layer. Using a high quality enhancement layer as the reference makes the motion prediction more accurate to improve the coding efficiency. On the other hand, the use of multiple layers of different quality references may also result in increases and fluctuations of the prediction residues to be coded when switching the references, which may limit the coding efficiency improvement. A multiple-layer conditional replenishment approach is used to eliminate this kind of fluctuation. Experimental results show that our coding scheme can improve the coding efficiency up to 0.5 dB compared with fine granularity scalability coding. Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang |
ICIP | 3 |
| 2000 | Automatic extraction of moving objects using multiple features and multiple framesabstractThis paper introduces a novel automatic video object extraction algorithm based on combination of color and motion segmentation results. The algorithm includes five parts: preprocessing, color segmentation, motion segmentation, combination of color and motion segmentation of multiple frames, post-processing. The performance of this algorithm is very promising, resulting in pixel-wise accuracy of extracted objects. Since it is an automatic extraction algorithm, it can be very useful in some real time video processing system based on video objects. Jinhui Pan, Shipeng Li 0001, Ya-Qin Zhang |
ISCAS | 3 |
| 2000 | A confidence measure based moving object extraction system built for compressed domainabstractAs the proliferation of compressed video sequences in MPEG formats continues, the ability to perform video analysis directly in the compressed domain becomes increasingly attractive. The availability of motion vectors and pixel values in coded forms can indirectly provide motion and intensity information for object analysis, avoiding the need to re-perform motion estimation. Albeit that the embedded motion field is contaminated with matching modeling errors and measurement errors, we will illustrate several motion field filtering and correction techniques to combat with noisy motion fields. We strive to reconstruct smooth true motion fields with a minimal amount of decoding, reducing computational resource and time requirement. In this paper, we describe the whole moving object extraction system with the general framework and component designs and show their effectiveness with two test sequences. Roy Wang, HongJiang Zhang, Ya-Qin Zhang |
ISCAS | 3 |
| 2000 | Object-based multiresolution watermarking of images and videoabstractThis paper proposes a new approach to digital watermarking of image and video objects with arbitrary regions of support (AROS) based on the 2D and 3D shape adaptive wavelet transforms. The hierarchical nature of the wavelet representation of objects allows detection of the digital watermark at various resolutions. We show that, when subjected to image/video compression, the corresponding watermark of objects can still be correctly identified at each resolution (excluding the lowest one) in the wavelet domain. Such a multiresolution watermarking scheme for objects has computational advantage, especially for the video case. Potential applications of our proposed scheme include watermark-based object searching and indexing. Xiaoyun Wu, Wenwu Zhu 0001, Zixiang Xiong, Ya-Qin Zhang |
ISCAS | 4 |
| 2000 | Adaptive QOS control for MPEG-4 video communication over wireless channelsabstractThis paper proposes an adaptive quality-of-service (QoS) control to increase the robustness of MPEG-4 video communication over wireless channels. More specifically, the proposed adaptive QoS control consists of optimal mode selection and delay-constrained hybrid automatic repeat request (ARQ). The optimal mode selection is employed to provide QoS support on the compression layer while delay-constrained hybrid ARQ is used to provide QoS support on the link layer. Simulation results show that the proposed adaptive QoS control achieves satisfactory quality for MPEG-4 video under dynamically changing wireless channel conditions and utilizes network resources efficiently. Dapeng Oliver Wu, Y. Thomas Hou 0001, Ya-Qin Zhang, Wenwu Zhu 0001, H. Jonathan Chao |
ISCAS | 3 |
| 2000 | Arbitrarily shaped video object coding by waveletabstractVideo object coding is one of the most important functionalities proposed by MPEG4. In this paper, we propose a new wavelet method to encode the texture of an arbitrarily shaped object, both for the still and video object. The method uses the shape adaptive wavelet transform (SA-DWT) in MPEG4 still object coding, but with a computationally more efficient lifting implementation. The transformed object coefficients are then quantized and entropy encoded with a partial bitplane embedded coder, which greatly improves the coding efficiency. We denote the coding algorithm as a video object wavelet (VOW) coder. Experimental results show that VOW significantly outperforms MPEG4 in still object coding, and achieves a comparable performance in video object coding in terms of PSNR. Moreover, the VOW decoded object looks better subjectively, with less annoying blocking artifacts than that of MPEG4. Guiwei Xing, Jin Li 0001, Shipeng Li 0001, Ya-Qin Zhang |
ISCAS | 4 |
| 2000 | Resource allocation for audio and video streaming over the InternetabstractStreaming of compressed audio and video (AV) over the Internet is a very challenging task since the current Internet only provides best-effort services and lacks Quality of Service (QoS) guarantee. This paper addresses resource allocation of multiple compressed AV streams delivered over the Internet to achieve end-to-end optimal quality. A multimedia streaming TCP-friendly transport protocol (MSTFP) is proposed to adaptively estimate the network bandwidth and smooth the sending rate. The resources are then allocated dynamically according to the media encoding distortion and network degradation. MPEG-4 streams with multiple video objects are used in the simulation to demonstrate the effectiveness of our proposed scheme. Qian Zhang 0001, Ya-Qin Zhang, Wenwu Zhu 0001 |
ISCAS | 2 |
| 2000 | Scalable video transport over wireless IP networksabstractThere has been great interest in transporting real-time video over wireless IP networks from both industry and academia. Real-time video applications have quality-of-service (QoS) requirements. However, the fluctuations of wireless channel conditions pose many challenges to providing QoS for video transmission over wireless IP networks. It has been shown that scalable video coding and adaptive services are viable solutions under a time-varying wireless environment. We propose an adaptive framework to support quality video communication over wireless IP networks. The adaptive framework includes: (1) scalable video representations, (2) network-aware video applications, and (3) adaptive services. Under this framework, as wireless channel conditions change, the mobile terminal and network elements can scale the video streams and transport the scaled video streams to receivers with acceptable perceptual quality. The key advantages of the adaptive framework are: (1) perceptual quality is degraded gracefully under severe channel conditions; (2) network resources are efficiently utilized; and (3) the resources are shared in a fair manner. Dapeng Oliver Wu, Y. Thomas Hou 0001, Ya-Qin Zhang |
PIMRC | 3 |
| 2000 | User- and content-aware object-based video streaming over the Internet
Huai-Rong Shao, Wenwu Zhu 0001, Ya-Qin Zhang |
VCIP | 3 |
| 2000 | Rendering of 3D-wavelet-compressed concentric mosaic scenery with progressive inverse wavelet synthesis (PIWS)
Yunnan Wu, Lin Luo 0004, Jin Li 0001, Ya-Qin Zhang |
VCIP | 4 |
| 2000 | Three-dimensional shape-adaptive discrete wavelet transforms for efficient object-based video coding
Jizheng Xu, Shipeng Li 0001, Ya-Qin Zhang |
VCIP | 3 |
| 2000 | Video transcoding for multiple clients
Jeongnam Youn, Jun Xin, Ming-Ting Sun, Ya-Qin Zhang |
VCIP | 4 |
| 2000 | Resource allocation with adaptive QoS for multimedia transmission over W-CDMA channelsabstractThis paper addresses the important issues of resource allocation and rate adaptation for multiple media, such as audio, video, email, and Web traffic, transmitted over W-CDMA (wideband code division multiple access) channel with adaptive QoS (quality of service) support. In order to have QoS support for different types of media, we develop an architecture combining the link layer with application layer controls. It consists of the following contributions: (1) an appropriate model to estimate the varying fading channel is proposed; (2) a hybrid delay-constrained ARQ (automatic repeat request) and UEP (unequal error protection) mechanism that dynamically adapt to the time-varying channel is presented to meet the QoS requirements for different applications; (3) a new resource allocation scheme that considers varying media characteristics is described to be adapted to changing bit error rate (BER) conditions. Simulation results demonstrate the effectiveness of our proposed scheme. Qian Zhang 0001, Wenwu Zhu 0001, Guijin Wang, Ya-Qin Zhang |
WCNC | 4 |
| 2000 | An end-to-end approach for optimal mode selection in Internet video communication: theory and applicationabstractRate-distortion (R-D) optimized mode selection is a fundamental problem for video communication over packet-switched networks. The classical R-D optimized mode selection only considers quantization distortion at the source. Such an approach is unable to achieve global optimality under the error-prone environment since it does not consider the packetization behavior at the source, the transport path characteristics, and receiver behavior. This paper presents an end-to-end approach to generalize the classical theory of R-D optimized mode selection for point-to-point video communication. We introduce a notion of global distortion by taking into consideration both the path characteristics (i.e., packet loss) and the receiver behavior (i.e., the error concealment scheme), in addition to the source behavior (i.e., quantization distortion and packetization). We derive, for the first time, a set of accurate global distortion metrics for any packetization scheme. Equipped with the global distortion metrics, we design an R-D optimized mode selection algorithm to provide the best tradeoff between compression efficiency and error resilience. The theory developed in this paper is general and is applicable to many video coding standards, including H.261/263 and MPEG-1/2/4. As an application, we integrate our theory with point-to-point MPEG-4 video conferencing over the Internet, where a feedback mechanism is employed to convey the path characteristics (estimated at the receiver) and receiver behavior (error concealment scheme) to the source. Simulation results are discussed. Dapeng Oliver Wu, Y. Thomas Hou 0001, Bo Li 0001, Wenwu Zhu 0001, Ya-Qin Zhang, H. Jonathan Chao |
IEEE J. Sel. Areas Commun. | 5 |
| 2000 | Transporting real-time video over the Internet: challenges and approachesabstractDelivering real-time video over the Internet is an important component of many Internet multimedia applications. Transmission of real-time video has bandwidth, delay, and loss requirements. However the current Internet does not offer any quality of service (QoS) guarantees to video transmission over the Internet. In addition, the heterogeneity of the networks and end systems makes it difficult to multicast Internet video in an efficient and flexible way. Thus, designing protocols and mechanisms for Internet video transmission poses many challenges. In this paper, we take a holistic approach to these challenges and present solutions from both transport and compression perspectives. With the holistic approach, we design a framework for transporting real-time Internet video, which includes two components, namely, congestion control and error control. Specifically congestion control consists of rate control, rate-adaptive encoding, and rate shaping; error control consists of forward error correction (FEC), retransmission error resilience, and error concealment. For the design of each component in the framework, we classify approaches and summarize representative research work. We point out there exists a design space which can be explored by video application designers and suggest that the synergy of both transport and compression could provide good solutions. Dapeng Oliver Wu, Y. Thomas Hou 0001, Ya-Qin Zhang |
Proc. IEEE | 3 |
| 2000 | Scalable rate control for MPEG-4 videoabstractThis paper presents a scalable rate control (SRC) scheme based on a more accurate second-order rate-distortion model. A sliding-window method for data selection is used to mitigate the impact of a scene change. The data points for updating a model are adaptively selected such that the statistical behavior is improved. For video object (VO) shape coding, we use an adaptive threshold method to remove shape-coding artifacts for MPEG-4 applications. A dynamic bit allocation among VOs is implemented according to the coding complexities for each VO. SRC achieves more accurate bit allocation with low latency and limited buffer size. In a single framework, SRC offers multiple layers of controls for objects, frames, and macroblocks (MBs). At MB level, SRC provides finer bit rate and buffer control. At multiple VO level, SRC offers superior VO presentation for multimedia applications. The proposed SRC scheme has been adopted as part of the International Standard of the emerging ISO MPEG-4 standard. Hung-Ju Lee, Tihao Chiang, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2000 | New fast binary pyramid motion estimation for MPEG2 and HDTV encodingabstractA novel fast binary pyramid motion estimation (FBPME) algorithm is presented in this paper. The proposed FBPME scheme is based on binary multiresolution layers, exclusive-or (XOR) Boolean block matching, and a N-scale tiling search scheme. Each video frame is converted into a pyramid structure of K-1 binary layers with resolution decimation, plus one integer layer at the lowest resolution. At the lowest resolution layer, the N-scale tiling search is performed to select initial motion vector candidates. Motion vector fields are gradually refined with the XOR Boolean block-matching criterion and the N-scale tiling search schemes in higher binary layers. FBPME performs several thousands times faster than the conventional full-search block-matching scheme at the same PSNR performance and visual quality. It also dramatically reduces the bus bandwidth and on-chip memory requirement. Moreover, hardware complexity is low due to its binary nature. Fully functional software MPEG-2 MP@ML encoders and Advanced Television Standard Committee high definition television encoders based on the FBPME algorithm have been implemented. FBPME hardware architecture has been developed and is being incorporated into single-chip MPEG encoders. A wide range of video sequences at various resolutions has been tested. The proposed algorithm is also applicable to other digital video compression standards such as H.261, H.263, and MPEG4. Xudong Song, Tihao Chiang, Xiaobing Lee, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2000 | On end-to-end architecture for transporting MPEG-4 video over the InternetabstractWith the success of the Internet and flexibility of MPEG-4, transporting MPEG-4 video over the Internet is expected to be an important component of many multimedia applications in the near future. Video applications typically have delay and loss requirements, which cannot be adequately supported by the current Internet. Thus, it is a challenging problem to design an efficient MPEG-4 video delivery system that can maximize the perceptual quality while achieving high resource utilization. This paper addresses this problem by presenting an end-to-end architecture for transporting MPEG-4 video over the Internet. We present a framework for transporting MPEG-4 video, which includes source rate adaptation, packetization, feedback control, and error control. The main contributions of this paper are: (1) a feedback control algorithm based on the Real Time Protocol (RTP) and the Real Time Control Protocol (RTCP); (2) an adaptive source-encoding algorithm for MPEG-4 video which is able to adjust the output rate of MPEG-4 video to the desired rate; and (3) an efficient and robust packetization algorithm for MPEG video bit-streams at the sync layer for Internet transport. Simulation results show that our end-to-end transport architecture achieves good perceptual picture quality for MPEG-4 video under low bit-rate and varying network conditions and efficiently utilizes network resources. Dapeng Oliver Wu, Y. Thomas Hou 0001, Wenwu Zhu 0001, Hung-Ju Lee, Tihao Chiang, Ya-Qin Zhang, H. Jonathan Chao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 1999 | An End-To-End Architecture for Mpeg-4 Video Streaming over the InternetabstractIt is a challenging problem to design an efficient MPEG-4 video delivery system that can machine the perceptual quality while achieving high resource utilization. This paper addresses this problem by presenting an architecture of transporting MPEG-4 video over the Internet, which includes an end-to-end feedback control algorithm and a source encoding rate control algorithm. Our feedback control algorithm is capable of estimating the available bandwidth in the network based on the feedback information from the receiver, while our source encoding rate control algorithm is able to adjust the encoding rate of MPEG-4 video to the desired rate. Simulation results demonstrate that our architecture achieves good perceptual picture quality under low bit-rate and varying network conditions while efficiently utilizing network resources. Y. Thomas Hou 0001, Dapeng Oliver Wu, Wenwu Zhu 0001, Hung-Ju Lee, Tihao Chiang, Ya-Qin Zhang |
ICIP (1) | 6 |
| 1999 | Dynamic frame rate control for video streamsabstractA mechanism for dynamically varying the frame rate of pre-encoded video clips is described. An off-line encoder creates a high quality bitstream encoded at 30 fps, as well as separate files containing motion vectors for the same clip at lower frame rates. An on-line encoder decodes the bitstream (if necessary) and re-encodes it at lower frame-rates in real-time using the pre-computed, stored motion information. Dynamic Frame Rate Control, used in conjunction with dynamic bit-rate control, allows clients to solve the rate mismatch between the bandwidth available to them and the bit-rate of the pre-encoded bitsream. It also provides a means for implementing Fast Forward control for video streaming without increasing bandwidth consumption. Sassan Pejhan, Tihao Chiang, Ya-Qin Zhang |
ACM Multimedia (1) | 3 |
| 1999 | Robust video coding algorithms and systemsabstractWireless video communication is particularly challenging because it combines the already difficult problem of efficient compression with the additional and usually contradictory need to make the compressed bit stream robust to channel errors. We describe design and implementation strategies for error-robust video communications with an emphasis on techniques compatible with the coding approaches used in the ISO (MPEG-4) and ITU standards organizations. These techniques include modifications to the video coding algorithms as well as to the system layers that perform packetization and multiplexing. John D. Villasenor, Ya-Qin Zhang, Jiangtao Wen |
Proc. IEEE | 2 |
| 1999 | Scalable wavelet coding for synthetic/natural hybrid imagesabstractThis paper describes the texture representation scheme adopted for MPEG-4 synthetic/natural hybrid coding (SNHC) of texture maps and images. The scheme is based on the concept of multiscale zerotree wavelet entropy (MZTE) coding technique, which provides many levels of scalability layers in terms of either spatial resolutions or picture quality, MZTE, with three different modes (single-Q, multi-Q, and bilevel), provides much improved compression efficiency and fine-gradual scalabilities, which are ideal for hybrid coding of texture maps and natural images. The MZTE scheme is adopted as the baseline technique for the visual texture coding profile in both the MPEG-4 video group and SNHC group. The test results are presented in comparison with those coded by the baseline JPEG scheme for different types of input images, MZTE was also rated as one of the top five schemes in terms of compression efficiency in the JPEG2000 November 1997 evaluation, among 27 submitted proposals. Iraj Sodagar, Hung-Ju Lee, Paul Hatrack, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 1999 | A comparative study of DCT- and wavelet-based image codingabstractWe undertake a study of the performance difference of the discrete cosine transform (DCT) and the wavelet transform for both image and video coding, while comparing other aspects of the coding system on an equal footing based on the state-of-the-art coding techniques. The studies reveal that, for still images, the wavelet transform outperforms the DCT typically by the order of about 1 dB in peak signal-to-noise ratio. For video coding, the advantage of wavelet schemes is less obvious. We believe that the image and video compression algorithm should be addressed from the overall system viewpoint: quantization, entropy coding, and the complex interplay among elements of the coding system are more important than spending all the efforts on optimizing the transform. Zixiang Xiong, Kannan Ramchandran, Michael T. Orchard, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 1999 | Multiresolution watermarking for images and videoabstractThis paper proposes a unified approach to digital watermarking of images and video based on the two- and three-dimensional discrete wavelet transforms. The hierarchical nature of the wavelet representation allows multiresolutional detection of the digital watermark, which is a Gaussian distributed random vector added to all the high-pass bands in the wavelet domain. We show that when subjected to distortion from compression or image halftoning, the corresponding watermark can still be correctly identified at each resolution (excluding the lowest one) in the wavelet domain. Computational savings from such a multiresolution watermarking framework is obvious, especially for the video case. Wenwu Zhu 0001, Zixiang Xiong, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1998 | Multiresolution Watermarking for Images and Video: A Unified ApproachabstractThis paper proposes a unified approach to digital watermarking of images and video based on the 2D and 3D discrete wavelet transforms. The hierarchical nature of the wavelet representation allows multiresolutional detection of the digital watermark, which is a Gaussian distributed random vector added to all the high pass bands in the wavelet domain. We show that, when subjected to distortion from compression or image halftoning, the corresponding watermark can still be correctly identified at each resolution (excluding the lowest one) in the wavelet domain. Computational saving from such a multiresolution watermarking framework is obvious, especially for the video case. Wenwu Zhu 0001, Zixiang Xiong, Ya-Qin Zhang |
ICIP (1) | 3 |
| 1998 | End-to-end modeling and simulation of MPEG-2 transport streams over ATM networks with jitterabstractThe operation of MPEG-2 systems is modeled and simulated when an MPEG-2 transport stream is delivered through an ATM network with jitter. End-to-end packet-based analysis is performed for delivery of MPEG-2 transport streams over ATM networks. A novel approach to analyzing the decoder buffer behaviour in the presence of network jitter is presented. The probability density function of the interarrival time of the ATM adaptation layer 5 (AAL5) protocol data unit (PDU) is derived from an MPEG-2 video source model and an ATM network jitter model. Based on a real-time decoding requirement of the MPEG-2 transport stream (TS) system target decoder (T-STD), the decoder buffer behaviour is simulated. In this simulation, the packets' arrivals follow the derived probability density function of the AAL5 PDU interarrival time. The modeling and simulation results show the interactions among packet loss ratio, decoder buffer size, and network jitter level. We found that jitter affects decoder buffer size and packet loss ratio in a significant way. Wenwu Zhu 0001, Y. Thomas Hou 0001, Yao Wang 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 1997 | An optical flow based motion compensation algorithm for very low bit-rate video codingabstractWe propose an efficient compression algorithm for very low bit-rate video applications. The algorithm is based on (1) optical-flow motion estimation to achieve more accurate motion prediction fields; (2) DCT-coding of the motion vectors from the optical-flow estimation to further reduce the motion overheads; and (3) a region adaptive threshold technique to match optical flow motion prediction and minimize the residual errors. Unlike the classic block-matching based discrete cosine transformation (DCT) video coding schemes in MPEG 1/2 and H.261/3, the proposed algorithm uses optical flow for motion compensation and the DCT is applied to the optical flow field instead of predictive errors. Thresholding techniques are used to treat different regions to complement the optical flow technique and to efficiently code residual data. While maintaining comparable peak signal to noise ratio (PSNR) and computational complexity with that of ITU-T H.263/TMN5, the reconstructed video frames of the proposed coder are free of annoying blocking artifacts, and hence visually much more pleasant. Yun Q. Shi 0001, Ya-Qin Zhang |
ICASSP | 3 |
| 1997 | Scalable Rate Control for Very Low Bit Rate (VLBR) VideoabstractThis paper presents a scalable rate control scheme (SRC) based on three new concepts: (1) a more accurate rate distortion model, (2) a sliding window method, and (3) an adaptive selection criterion of data points. The SRC scheme achieves a more accurate target bit allocation under the constraints of low latency and limited buffer size. In addition to the picture-level rate control, the SRC scheme is also applicable to the macroblock-level rate control for finer bit allocation and buffer control, and multiple video objects (VOs) rate control for better video object presentation in different applications. The proposed SRC scheme has been adopted in verification model (VM) 5.0 and VM 8.0 of the emerging ISO MPEG-Q standard. Hung-Ju Lee, Tihao Chiang, Ya-Qin Zhang |
ICIP (2) | 3 |
| 1997 | Demonstration of the MPEG-2, MPEG-4 and H.263 video coding standardsabstractIn this demonstration session, we will demonstrate several examples of the state-of-the-art video compression standards MPEG-2, MPEG-4 and H.263. These demonstrations captured some of the recent work developed in the Digital Video Communications group at Sarnoff. For the MPEG-2 demonstration, the compression process uses standards compliant syntax with Sarnoff's proprietary algorithms to select the encoding parameters. For the MPEG-4 demonstration, we show the encoding results of the Verification Model including Sarnoff's contribution in scalable rate control and texture coding mode. In addition, we show the encoding results of a video streaming technology using H.263 and Sarnoff's proprietary technologies. Tihao Chiang, Hung-Ju Lee, Sassan Pejhan, Iraj Sodagar, Ya-Qin Zhang |
MMSP | 5 |
| 1997 | Global motion compensation for low bitrate video codingabstractA global motion compensation (GMC) scheme is described and implemented for low bitrate coding. Experiments show that GMC gives excellent coding results. When coupled with the H.263 low bitrate video coding standard, GMC is capable of achieving up to 27% bitrate saving (or 1.8 dB gain in PSNR) for a high motion sequence such as Stefan. For a moderate motion sequence such as Cost-guard, we obtained an 8% bitrate saving over H.263 at the same quality using GMC. Zixiang Xiong, Tihao Chiang, Ya-Qin Zhang |
MMSP | 3 |
| 1997 | Modeling and simulation of MPEG-2 video transport over ATM networks considering the jitter effectabstractIn this paper, the operation of MPEG-2 systems is modeled and simulated when an MPEG-2 transport stream is delivered through a ATM network with jitter. A novel approach to analyzing the decoder buffer behavior in the presence of network jitter is presented. The probability density function of the interarrival time of the ATM adaptation layer 5 (AAL5) Protocol Data Unit (PDU) is derived from a MPEG-2 video source model and an ATM network jitter model. Based on a real-time decoding requirement of the MPEG-2 transport stream (TS) system target decoder (T-STD), the decoder buffer behavior is simulated. The modeling; and simulation results show that jitter affects decoder buffer size and packet loss ratio in a significant way. Wenwu Zhu 0001, Y. Thomas Hou 0001, Yao Wang 0001, Ya-Qin Zhang |
MMSP | 4 |
| 1997 | A new rate control scheme using quadratic rate distortion modelabstractA new rate control scheme is used to calculate the target bit rate for each frame based on a quadratic formulation of the rate distortion function. The distortion measure is assumed to be the average quantization scale of a frame. The rate distortion function is modeled as a second-order function of the inverse of the distortion measure. We present a closed form solution for the target bit allocation which includes the MPEG-2 TM5 rate control scheme as a special case. The model parameters are estimated using statistical linear regression analysis. Since the estimation uses the past encoded frames of the same picture prediction type (I, P, B pictures), the proposed approach is a single pass rate control technique. Because of the improved accuracy of the rate distortion function, the fluctuations of the bit counts are significantly reduced by 20-65% in the standard deviation of the bit count while the picture quality remains the same. Thus, the buffer requirement is reduced at a small increase in complexity. This technique has been adopted by the MPEG committee as part of VM5.0 in November 1996. Tihao Chiang, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1997 | A video coding algorithm using vector-based techniquesabstractThis paper presents an algorithm proposal submitted to MPEG-4 for video coding. The proposed algorithm addresses the functionality of improved coding efficiency for compression. It uses vector-based techniques for coding intraframes (the first frame and subsequent refreshing key frames) and motion-compensated difference frames. It uses the same motion estimation and motion compensation techniques as H.263. A video frame (I or P frame) is first decomposed into a set of vector bands using a vector wavelet transform. This stage of vector-based signal processing makes subsequent vector quantization in the vector bands very efficient. Lattice vector quantization is then used in the vector bands. A 100% labeling efficiency is achieved for lattice vector quantization by using a set of generalized labeling algorithms for various important lattices with pyramid and sphere boundaries. Finally, entropy coding is used to code the indexes generated from lattice vector quantization. Our coding results have shown that a gain in peak signal-to-noise ratio (PSNR) up to 8 dB for intraframe coding and up to 6 dB for interframe coding can be achieved over H.263. Subjective quality improvement of the proposed algorithm over H.263 can be easily observed. Hugh Q. Cao, Shipeng Li 0001, Fan Ling, Scott A. Segan, Hongqiao Sun, John Wus, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 1997 | A zerotree wavelet video coderabstractThis paper describes a hybrid motion-compensated wavelet transform coder designed for encoding video at very low bit rates. The coder and its components have been submitted to MPEG-4 to support the functionalities of compression efficiency and scalability. Novel features of this coder are the use of overlapping block motion compensation in combination with a discrete wavelet transform followed by adaptive quantization and zerotree entropy coding, plus rate control. The coder outperforms the VM of MPEG-4 for coding of I-frames and matches the performance of the VM for P-frames while providing a path to spatial scalability, object scalability, and bitstream scalability. Stephen A. Martucci, Iraj Sodagar, Tihao Chiang, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 1997 | A deblocking algorithm for JPEG compressed images using overcomplete wavelet representationsabstractThis paper introduces a new approach to deblocking of JPEG compressed images using overcomplete wavelet representations. By exploiting cross-scale correlations among wavelet coefficients, edge information in the JPEG compressed images is extracted and protected, while blocky noise in the smooth background regions is smoothed out in the wavelet domain. Compared with the iterative methods reported in the literature, our simple wavelet-based method has much lower computational complexity, yet it is capable of achieving the same peak signal-to-noise ratio (PSNR) improvement as the best iterative method and giving visually very pleasing images as well. Zixiang Xiong, Michael T. Orchard, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1996 | A new rate control scheme using quadratic rate distortion modelabstractA new rate control scheme is used to calculate the target bit rate for each frame based on a quadratic formulation of the rate distortion function. The distortion measure are assumed to be either the average quantization scale of a frame or mean square error. The rate distortion function is modeled as a second order function of the inverse of the distortion measure. We presented a closed form solution for the target bit allocation which applies to both the digital storage media (DSM) applications (e.g., MPEG-2) and very low bit rate (VLBR) applications (e.g., MPEG-4). We perform experiments on a wavelet-based coder to demonstrate the versatility of our scheme. Model parameters are estimated using statistical linear regression analysis. Since the estimation uses the past encoded frames of the same picture prediction type (I, P and B pictures), the proposed approach is a single pass rate control technique. For the MPEG-2 coder the fluctuations of the bit counts are significantly reduced by 20 percent to 65 percent in the standard deviation of the bit count while the picture quality remains the same. For the MPEG-4 VM coder, the PSNR is improved by almost one dB as compared to the best MPEG-4 VM rate control using exhaustive combinations of quantization scale parameters. Tihao Chiang, Ya-Qin Zhang |
ICIP (2) | 2 |
| 1996 | Very low bit rate video coding using vector-based techniquesabstractThis paper reports the advances of using vector-based techniques for very low bit rate video coding. High efficiency has been achieved for coding intraframes (the first frame and subsequent refreshing key frames) and motion compensated difference frames of a video sequence. A video frame (I or P frame) is first decomposed into a set of vector bands using a vector wavelet transform. Adaptive lattice vector quantization is then used in the vector bands. A 100% labeling efficiency is achieved for lattice vector quantization by using a set of generalized labeling algorithms for various important lattices with pyramid and sphere boundaries. An adaptive algorithm is used to determine the type of lattice vector quantisation according to the statistical distribution of the vectors in the vector wavelet domain. Finally, entropy coding is used to code the indexes generated from lattice vector quantization. Our coding results show that a gain of 3 to 8 dBs in PSNR for intraframe coding can be achieved at a bitrate level of 16 Kbpf over H.263. Subjective quality improvement of the proposed algorithm over H.263 can be observed. Hugh Q. Cao, Shipeng Li 0001, Fan Ling, Scott A. Segan, H. Q. Sun, John Wus, Ya-Qin Zhang |
ICIP (1) | 8 |
| 1996 | A fast hierarchical motion-compensation scheme for video coding using block feature matchingabstractThis paper presents a fast hierarchical feature matching-motion estimation scheme (HFM-ME) that can be used in H.263, H.261, MPEG 1, MPEG 2, and HDTV applications. In the HFM-ME scheme, the sign truncated feature (STF) is defined and used for block template matching, as opposed to the pixel intensity values used in conventional block matching methods. The STF extraction process can be considered as a zero-crossing phase detection with the mean as the bias and binary sign pattern as the phase deviation. Using the STF definition, a data block can be represented by a mean and a set of binary features with a much reduced data set. The block matching motion estimation is then divided into mean matching and binary phase matching. The proposed technique enables a significant reduction in computational complexity compared with the conventional full-search block matching ME because binary phase matching only involves Boolean logic operations. This feature also significantly reduces the data transfer time between the frame buffer and motion estimator. The proposed HFM-ME algorithm is implemented and compared with the conventional full-search block matching schemes. Our test results using three full-motion MPEG sequences indicate that the performance of the HFM-ME is comparable with the full-search block matching under the same search ranges, however, HFM-ME can be implemented about 64 times faster than the conventional full-search schemes. The proposed scheme can be combined with other fast algorithms to further reduce the computational complexity, at the expense of picture quality. Xiaobing Lee, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1995 | Vector-based signal processing and quantization for image and video compressionabstractImage and video compression has become an increasingly important and active area. Many techniques have been developed in this area. Any compression technique can be modeled as a three-stage process. The first stage can be generally called a signal processing stage where an image or video signal is converted into a different domain. Usually, there is no or little loss of information in this stage. The second stage is quantization where loss of information occurs. The third stage is lossless coding that generates the compressed bit stream. The purpose of the signal processing stage is to convert an image or video signal into such a form that quantization can achieve better performance than without the signal processing stage. Because the quantization stage is the place where most of compression is achieved and loss of information occurs, it is naturally the central stage of any compression technique. Since scalar quantization or vector quantization may be used in the second stage, the operation in the first stage should be scalar-based or vector-based respectively in order to match the second stage so that the compression performance can be optimized. In this paper, we summarize the most recent research results on vector-based signal processing and quantization techniques that have shown high compression performance.> Ya-Qin Zhang |
Proc. IEEE | 2 |
| 1995 | Information loss recovery for block-based image coding techniques-a fuzzy logic approachabstractA new technique to recover the information loss in a block-based image coding system is developed in this paper. The proposed scheme is based on fuzzy logic reasoning and can be divided into three main steps: (1) hierarchical compass interpolation/extrapolation (HCIE) in the spatial domain for initial recovery of lost blocks that mainly contain low-frequency information such as smooth background (2) coarse spectra interpretation by fuzzy logic reasoning for recovery of lost blocks that contain high-frequency information such as complex textures and fine features (3) sliding window iteration (SWI), which is performed in both spatial and spectral domains to efficiently integrate the results obtained in steps (1) and (2) such that the optimal result can be achieved in terms of surface continuity on block boundaries and a set of fuzzy inference rules. The proposed method, which is suitable for recovering both isolated and contiguous block losses, provides a new approach for error concealment of block-based image coding systems such as the JPEG coding standard and vector quantization-based coding algorithms. The principle of the proposed scheme can also be applied to block-based video compression schemes such as the H.261, MPEG, and HDTV standards. Simulation results are presented to illustrate the effectiveness of the proposed method. Xiaobing Lee, Ya-Qin Zhang, Alberto Leon-Garcia |
IEEE Trans. Image Process. | 2 |
| 1994 | VQ-Based Image Coding and Vector Filter BankabstractIt is well known that vector quantisation (VQ) is a good technique for image coding. For VQ-based image coding, vector-based signal processing techniques should be used to better match with VQ. The paper reports some preliminary results of extending the concept of the filter bank from a scalar-based operation to a vector-based operation and the performance of a vector filter bank for image coding.> John Wus, Ya-Qin Zhang |
ICIP (1) | 3 |
| 1994 | Image Coding Using Vector Filter Bank and Vector QuantizationabstractSubband coding has proven to be a very effective technique when used for image compression purposes. In this paper, the concept of Vector Filter Bank (VFB), which extends the conventional Scalar Filter Bank (SFB) to the vector case, is introduced. A new image coding algorithm, called Vector Subband Coding (VSC), is proposed. Detailed vector formation, filtering processes, vector codebook generation, vector encoding and decoding are examined. Finally, the test results are presented.> John Wus, Ya-Qin Zhang |
ISCAS | 3 |
| 1994 | Performance of MPEG Codecs in the Presence of Errors
Ya-Qin Zhang, Xiaobing Lee |
J. Vis. Commun. Image Represent. | 1 |
| 1994 | A study of vector transform coding of subband-decomposed imagesabstractStudies vector transform coding (VTC), a new image coding scheme, on subband-decomposed images. It is shown that vector transformation (VT) reduces the inter-vector correlation, although not as much as the discrete cosine transform (DCT). However, it is also shown that VT preserves the intra-vector correlation much better than the DCT so that vector quantization (VQ) in the VT domain can be made more efficient. VTC of subband-decomposed images introduces another dimension of adaptivity, in which coding parameters, bit allocation, and VQ codebooks can be adapted to each level of the subband pyramid as well as to each vector in the VT domain. The new subband/VTC scheme is compared with VQ of original images, VQ of subband-decomposed images, DCT-based transform coding, and subband/DCT/VQ schemes. Simulation results indicate that the new scheme achieves 1 to 3dB improvement over the other schemes in terms of peak signal-to-noise ratio. This improvement is also supported by subjective evaluations.> Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1994 | Modeling and queueing analysis of variable-bit-rate coded video sources in ATM networksabstractTraffic models and queueing performance of variable-bit-rate (VBR) video sources in an asynchronous transfer mode (ATM) network are studied. A discrete-time discrete-state Markov chain is used to model the aggregate video traffic with each VBR-coded source being modeled by a renewal process, which has been successfully applied in the analysis of packet voice traffic. Three different methods including the stationary-interval (SI) method, the asymptotic method (ASM), and the hybrid method for queueing network analyzer (QNA) are used to approximate the average queue size. Results for different traffic conditions and different number of VBR sources are compared with the simulation results. It can be observed that as the number of sources increases the aggregate traffic becomes more predictable and the congestion at the common queue becomes smaller. This result verifies the fact that multiplexing a large number of identical video sources on a single high speed link statistically yields significant bandwidth saving. It is also interesting to note that the SI method provides an upper bound and the ASM method yields a lower bound for the average queue size for the type of traffic used in the study. When the number of VBR sources increases, the result deviates from the SI method and approaches the ASM method. In general, the QNA method provides a close match to the simulation result.> Nasser M. Marafih, Ya-Qin Zhang, Raymond L. Pickholtz |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1993 | New insights and results on transform domain VQ of images
Ya-Qin Zhang |
ICASSP (5) | 2 |
| 1993 | On the optimal transform for vector quantization of images
M. Atif Cay, Ya-Qin Zhang |
ISCAS | 3 |
| 1993 | Information loss recovery for block-based image coding techniques: a fuzzy logic approachabstractA new technique to recover the information loss in a block-based image coding system is developed in this paper. The proposed scheme is based on the fuzzy logic reasoning and can be divided into three main steps: (1) hierarchical compass interpolation/extrapolation in the spatial domain for initial recovery of lost blocks that mainly contain low-frequency information such as smooth background; (2) coarse spectra interpretation by fuzzy logic reasoning for recovery of lost blocks that contain high-frequency information such as complex textures and fine features; (3) sliding window iteration in both spatial and spectral domains to efficiently integrate the results obtained in step (1) and (2) such that optimal results can be achieved in terms of surface continuities on block boundaries and the established inference rules. The proposed method, suitable for recovering both isolated and contiguous block losses, provides a new approach for error concealment of block-based image coding systems such as the JPEG coding standard and vector quantization based coding algorithms. The principle of the proposed scheme can also be applied to block-based video compression schemes such as the H.261, MPEG, and HDTV standards. Simulation results are presented to illustrate the effectiveness of the proposed method. Xiaobing Lee, Ya-Qin Zhang, Alberto Leon-Garcia |
VCIP | 2 |
| 1993 | Performance of MPEG codecs in the presence of errorsabstractMPEG is emerging as a major international standard for applications in compressed video storage, transmission, and interactive communications. MPEG will be used in applications which involve imperfect channel conditions or storage defects. This work studies the effects of different types of errors on the compressed MPEG bit streams and possible approaches to minimize such effects. The use of forward error correction, block interleaving, error concealment, and their inter-relationship and applicability to MPEG bit streams, were investigated. Ya-Qin Zhang, Xiaobing Lee |
VCIP | 1 |
| 1993 | Security Analysis of the INTELSAT VI and VII Command NetworkabstractSome results of a study of the command security issues associated with the INTELSAT VI and VII satellites are reported. The configuration and protocols of the INTELSAT command system are briefly described. Three possible configurations for connecting the INTELSAT headquarters and the telemetry, tracking, and command (TTC) stations are distinguished, and a layered architecture is introduced to illustrate the command protocols corresponding to the ISO layers. The impact of introducing command security is then studied for the INTELSAT command network. In order to analyze the effect of errors on the operation and performance of the command network, a finite-state Markov chain is proposed for modeling the command and telemetry channels with their different error and delay characteristics. Possible risks that can threaten the INTELSAT secure command network operations, including single-event upsets, Vcc corruption, and manipulation of transmitted messages, are analyzed. The end-to-end performance of the INTELSAT command network is examined for three possible configurations-terrestrial link, single-hop satellite, and two-hop satellite.> Raymond L. Pickholtz, David B. Newman Jr., Ya-Qin Zhang, Makoto Tatebayashi |
IEEE J. Sel. Areas Commun. | 3 |
| 1993 | Multiscale Video Representation Using Multiresolution Motion Compensation and Wavelet DecompositionabstractA multiscale video representation using wavelet decomposition and variable-block-size multiresolution motion estimation (MRME) is presented. The multiresolution/multifrequency nature of the discrete wavelet transform makes it an ideal tool for representing video sources with different resolutions and scan formats. The proposed variable-block-size MRME scheme utilizes motion correlation among different scaled subbands and adapts to their importance at different layers. The algorithm is well suited for interframe HDTV coding applications and facilitates conversions and interactions between different video coding standards. Four scenarios for the proposed motion-compensated coding schemes are compared. A pel-recursive motion estimation scheme is implemented in a multiresolution form. The proposed approach appears suitable for the broadcast environment where various standards may coexist simultaneously.> Sohail Zafar, Ya-Qin Zhang, Bijan Jabbari |
IEEE J. Sel. Areas Commun. | 2 |
| 1993 | A new approach to reduce the 'blocking effect' of transform coding [image coding]abstractA combined-transform coding (CTC) scheme to reduce the blocking effect of conventional block transform coding and hence to improve the subjective performance is presented. The scheme is described, and its information-theoretic properties are discussed. Computer simulation results for a chest X-ray image are presented. The CTC scheme, the JPEG baseline scheme, and the conventional discrete Walsh-Hadamard transform (DWHT) are compared to demonstrate the performance improvement for the CTC scheme. The advantages of the CTC scheme include no ringing effect as there is no error propagation across the boundary, no additional computation, and distortion always held within a certain level.> Ya-Qin Zhang, Raymond L. Pickholtz, Murray H. Loew |
IEEE Trans. Commun. | 1 |
| 1993 | Statistical characterization and block-based modeling of motion-adaptive coded videoabstractThe statistical characteristics of full-motion video sources using motion-adaptive variable-bit-rate coding techniques is studied. Analytical models are developed to describe the behavior of the coded video signals based on the encoder structure. The video-compression algorithm used is in compliance with the general MPEG syntax and bit-stream definition. Statistical characteristics associated with each block type and their aggregate are presented. A composite model to represent the number of bits per field for the encoded video traffic that comprises multiple autoregressive models for the number of blocks per field and the number of bits in each coded block is derived. The statistics measured from a sample video sequence are compared to those obtained by the model and it is observed that the model captures the coded video behavior for each block type and their combination reasonably well. This model can be used to further study the cell-generation process of full-motion video codecs and the aggregation of such video sources at the statistical multiplexers.> Bijan Jabbari, Ferit Yegenoglu, Yu Kou, Sohail Zafar, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 1993 | A modified short-kernel filter pair for perfect reconstruction of HDTV signalsabstractA modified short-kernel filter pair is proposed for perfect reconstruction of HDTV signals. An interband prediction scheme based on the proposed filter pair is suggested to further reduce the average entropy of subband luminance signals. Simulations are conducted, and a modest reduction of the average entropy is obtained.> Cheng-Chang Lu, Norhanim Omar, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1993 | Motion-classified autoregressive modeling of variable bit rate videoabstractA motion-adaptive variable-bit-rate (VBR) video codec is considered, and a motion-classified model is developed to represent the characteristics of various classes of motion activities, including scene changes. The codec switches between interframe, motion-compensated, and intraframe coding corresponding to low, medium, and high amounts of motion and scene changes, respectively. The model captures the motion of various video scenes by providing the statistics of VBR-coded video traffic through a first-order autoregressive process with time-varying parameters. The parameters of this model are obtained from a VBR-coded sample video sequence with the objective of matching the bit-rate distribution and the autocorrelation among the bit rates. The validity and accuracy of the model are evaluated, and the characteristics of aggregated traffic sources obtained with the model are discussed.> Ferit Yegenoglu, Bijan Jabbari, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1993 | Rate-distortion bound for a class of non-Gaussian sources with memoryabstractThe rate-distortion performance of a class of non-Gaussian source with memory is studied. The source model is generated from a correlated Gaussian source through a memoryless transform, which actually constitutes a special class of the frequently-used Gibbs model in speech and image processing applications. Its entropy and a rate-distortion bound are evaluated to obtain further insights into this class of source models. Several examples are also given to illustrate the usefulness and tightness of the rate-distortion bound.> Ya-Qin Zhang, Raymond L. Pickholtz, Murray H. Loew |
IEEE Trans. Inf. Theory | 1 |
| 1992 | Modeling of Motion Classified VBR Video CodecsabstractThe authors use a motion adaptive variable-bit-rate (VBR) video codec and propose a motion classified model to represent the characteristics of various classes of motion activities. The codec switches between interframe, motion compensated, and intraframe coding corresponding to low, medium, and high motions and scene changes, respectively. The model captures the motion of various video scenes and the codec structure by providing the statistics of VBR-coded video' traffic through a first-order composite autoregressive process with three motion classes. The parameters of this model are derived from a VBR-coded sample video sequence such that the bit rate distribution and the autocorrelation in bit rates of two successive frames are matched. The validity and accuracy of the model are verified. Using this model, the characteristics of aggregated traffic sources are discussed.> Ferit Yegenoglu, Bijan Jabbari, Ya-Qin Zhang |
INFOCOM | 3 |
| 1992 | Motion-compensated wavelet transform coding for color video compressionabstractA variable-block-size multiresolution motion compensation (MRMC) scheme in which the size of a block is adapted to its level in the wavelet pyramid is proposed. This scheme not only considerably reduces the searching and matching time but also provides a meaningful characterization of the intrinsic motion structure. The variable-block-size approach also avoids the drawback of the constant-size MRMC in describing small object motion activities. After wavelet decomposition, each scaled subframe tends to have different statistical properties. An adaptive truncation process is implemented, and a bit allocation scheme similar to that in transform coding is examined. Four variations of the proposed motion-compensated wavelet video compression system are identified, and it is shown that the coding approach has a superior performance in terms of the peak-to-peak signal-to-noise ratio as well as the subjective quality.> Ya-Qin Zhang, Sohail Zafar |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1992 | A combined-transform coding (CTC) scheme for medical imagesabstractA combined-transform coding (CTC) scheme is proposed to reduce the blocking artifact of conventional block transform coding and hence to improve the subjective performance. The proposed CTC scheme is described and its information-theoretic properties are investigated. Computer simulation results for a class of chest X-ray images are presented. A comparison between the CTC scheme and the conventional discrete cosine transform (DCT) and discrete Walsh-Hadamard transform (DWHT) demonstrates the performance improvement of the proposed scheme. In addition, combined coding can also be used in noiseless coding, yielding a slight improvement in the compression performance if it is used properly. Ya-Qin Zhang, Murray H. Loew, Raymond L. Pickholtz |
IEEE Trans. Medical Imaging | 1 |
| 1991 | Variable bit-rate video transmission in the broadband ISDN environmentabstractMany compensative measures have been proposed recently the make the cell loss in ATM (asynchronous transfer mode) networks subjectively imperceptible. These schemes include the simple automatic repeat quest (ARQ) scheme, error concealment by command refreshment, appropriate queueing disciple and priority switching design and layered source coding schemes. The authors summarize these efforts and, in particular, elaborate on different layered source coding schemes. Four types of signal priority classification schemes are identified: bit-plane separation (FDS) combined with bit-plane-frequency separation (CBFS), and feature plane separation (FPS). Different layered coding techniques are discussed and compared. Some open questions and recommendations are also presented for further research and study.> Ya-Qin Zhang, Wiliam W. Wu, Kap S. Kim, Raymond L. Pickholtz, Jay Ramasastry |
Proc. IEEE | 1 |
| 1990 | On modeling the distribution of chest X-ray images and their stochastic propertiesabstractThe probabilistic distribution properties of a set of medical images are studied. It is shown that the generalized Gaussian function provides a good approximation to the distribution of antero-posterior chest radiographs. Based on this result and a goodness-of-fit test. a generalized Gaussian autoregressive model (GGAR) is proposed. Its properties and limitations are discussed. It is expected that the GGAR model will be useful in describing the stochastic characteristics of some classes of medical images and can be used in image data compression and other applications.> Ya-Qin Zhang, Murray H. Loew, Raymond L. Pickholtz |
ICPR (2) | 1 |
| 1990 | Variable-bit-rate video transmission in the broadband ISDN environmentabstractThe simple automatic repeat request (ARQ) scheme, error concealment by command refreshment, appropriate queuing discipline and priority switching design, layered source coding schemes, and other signal processing techniques are summarized with emphasis on different layered source coding schemes. Four types of signal priority classification schemes are identified: bit-plane separation (BPS), frequency-domain separation (FDS), combined bit-plane-frequency separation (CBFS) and feature plane separation (FPS). CBFS is shown to offer a good alternative to the FDS for possible cell dropping compensations in the ATM network. Different layered source coding schemes are described and compared, and a CBFS-based combined-transform coding (CTC) scheme is highlighted.> Ya-Qin Zhang, William W. Wu, Kap S. Kim, Raymond L. Pickholtz, Jay Ramasastry |
LCN | 1 |