VLDB 2026 Research / reviewers in the wild / expert
Jiangtao Wen
dblp:67/5635
· DBLP profile ↗
108ranked-venue papers
16as first author
15since 2021 · last 2026
0000-0002-0711-6132ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 74 · 13 first-author · 11 since 2021Databases, data management, data science and information retrieval · 18 · 6 first-author · 1 since 2021Computer networks · 12 · 1 first-authorArtificial intelligence and machine learning · 11 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 9Systems, architecture and hardware · 5 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-scale semantics meet SE(3) geometry: A sequence-aware hierarchical VLA for long-horizon manipulation
Zhitao Wang, Jiangtao Wen, Roberto Horowitz, Yanke Wang, Yuxing Han 0001 |
Knowl. Based Syst. | 2 |
| 2025 | EP-SAM: An Edge-Detection Prompt SAM Based Efficient Framework for Ultra-Low Light Video SegmentationabstractThe Segment Anything Model (SAM) excels at generating high-quality object masks with various prompts but struggles in ultra-low light. We developed EP-SAM (Edge-Detection Prompt SAM) with a Low Light Edge-Detection Network (LLEN), offering strong robustness and lightweight performance in ultra-low light. Using edge information as prompts, EP-SAM achieves high-precision segmentation in low-light videos.By combining motion estimation with reference frame optimization, the initial frame can predict the next 29 frames, reducing inference time by over 80% and computational complexity by 86%. Tests show LLEN accurately extracts edges even with about 80 photons per pixel, enabling EP-SAM to produce precise masks and significantly outperform SAM and SAM-2. EP-SAM improves mean Intersection over Union (mIoU) by 5.85% over SAM on the CamVid dataset. Video demos: https://github.com/wzt22thu/EP-SAM/releases/tag/DEMO. Zhitao Wang, Jiangtao Wen, Yuxing Han 0001 |
ICASSP | 2 |
| 2025 | A Zero Decoding Approach to Video ClassificationabstractClassifying videos into distinct categories, such as Sport and Music Video, is crucial for multimedia understanding and retrieval, especially with growing content volume. Traditional methods require video decompression to extract pixel-level features like color, texture, and motion, thereby increasing computational and storage demands. We present a novel approach that examines only the compressed bitstream of a video to perform classification, eliminating the need for bitstream decoding. To validate our approach, we built a comprehensive data set comprising over 29,000 YouTube video clips, totaling 6,000 hours and spanning 11 distinct categories. Our evaluations indicate precision, accuracy, and recall rates consistently above 80%, many exceeding 90%, and some reaching 99%. The algorithm operates approximately 15,000 times faster than real-time for 30fps videos, outperforming traditional Dynamic Time Warping (DTW) algorithm by seven orders of magnitude and state-of-the-art video classification model by three orders of magnitude. Chen Ye Gan, Jiangtao Wen, Yuxing Han 0001 |
ICME | 2 |
| 2025 | BackSlash: Rate Constrained Optimized Training of Large Language ModelsabstractThe rapid advancement of large-language models (LLMs) has driven extensive research into parameter compression after training has been completed, yet compression during the training phase remains largely unexplored. In this work, we introduce Rate-Constrained Training (BackSlash), a novel training-time compression approach based on rate-distortion optimization (RDO). BackSlash enables a flexible trade-off between model accuracy and complexity, significantly reducing parameter redundancy while preserving performance. Experiments in various architectures and tasks demonstrate that BackSlash can reduce memory usage by 60\% - 90\% without accuracy loss and provides significant compression gain compared to compression after training. Moreover, BackSlash proves to be highly versatile: it enhances generalization with small Lagrange multipliers, improves model robustness to pruning (maintaining accuracy even at 80\% pruning rates), and enables network simplification for accelerated inference on edge devices. Jiangtao Wen, Yuxing Han 0001 |
ICML | 2 |
| 2024 | Judging a video by its bitstream coverabstractClassifying videos into distinct categories, such as Sport and Music Video, is crucial for multimedia understanding and retrieval. Traditional methods require video decompression to extract pixel-level features like color, texture, and motion, thereby increasing computational and storage demands. We introduce a novel direction for video classification that does not rely on pixel domain information. Instead, we use the sequence of video frame sizes extracted from compressed bitstreams as input for a ResNet-based deep neural network, without the need for bitstream decoding or parsing. This approach leverages information captured by modern video compression algorithms, particularly advanced spatial and temporal prediction methods found in modern video coding standards such as H.264/AVC, H.265.HEVC and H.266/VVC. Yuxing Han 0001, Yunan Ding, Chen Ye Gan, Jiangtao Wen |
DCC | 4 |
| 2023 | DarkFeat: Noise-Robust Feature Detector and Descriptor for Extremely Low-Light RAW ImagesabstractLow-light visual perception, such as SLAM or SfM at night, has received increasing attention, in which keypoint detection and local feature description play an important role. Both handcraft designs and machine learning methods have been widely studied for local feature detection and description, however, the performance of existing methods degrades in the extreme low-light scenarios in a certain degree, due to the low signal-to-noise ratio in images. To address this challenge, images in RAW format that retain more raw sensing information have been considered in recent works with a denoise-then-detect scheme. However, existing denoising methods are still insufficient for RAW images and heavily time-consuming, which limits the practical applications of such scheme. In this paper, we propose DarkFeat, a deep learning model which directly detects and describes local features from extreme low-light RAW images in an end-to-end manner. A novel noise robustness map and selective suppression constraints are proposed to effectively mitigate the influence of noise and extract more reliable keypoints. Furthermore, a customized pipeline of synthesizing dataset containing low-light RAW image matching pairs is proposed to extend end-to-end training. Experimental results show that DarkFeat achieves state-of-the-art performance on both indoor and outdoor parts of the challenging MID benchmark, outperforms the denoise-then-detect methods and significantly reduces computational costs up to 70%. Code is available at https://github.com/THU-LYJ-Lab/DarkFeat. Yubin Hu 0001, Wang Zhao 0001, Jisheng Li, Yong-Jin Liu 0001, Yuxing Han 0001, Jiangtao Wen |
AAAI | 7 |
| 2023 | Efficient Semantic Segmentation by Altering Resolutions for Compressed VideosabstractVideo semantic segmentation (VSS) is a computationally expensive task due to the per-frame prediction for videos of high frame rates. In recent work, compact models or adaptive network strategies have been proposed for efficient VSS. However, they did not consider a crucial factor that affects the computational cost from the input side: the input resolution. In this paper, we propose an altering resolution framework called AR-Seg for compressed videos to achieve efficient VSS. AR-Seg aims to reduce the computational cost by using low resolution for non-keyframes. To prevent the performance degradation caused by downsampling, we design a Cross Resolution Feature Fusion (CR-eFF) module, and supervise it with a novel Feature Similarity Training (FST) strategy. Specifically, CReFF first makes use of motion vectors stored in a compressed video to warp features from high-resolution keyframes to low-resolution non-keyframes for better spatial alignment, and then selectively aggregates the warped features with local attention mechanism. Furthermore, the proposed FST supervises the aggregated features with high-resolution features through an explicit similarity loss and an implicit constraint from the shared decoding layer. Extensive experiments on CamVid and Cityscapes show that AR-Seg achieves state-of-the-art performance and is compatible with different segmentation backbones. On CamVid, AR-Seg saves 67% computational cost (measured in GFLOPs) with the PSPNet18 back-bone while maintaining high segmentation accuracy. Code: https://github.com/THU-LYJ-Lab/AR-Seg. Yubin Hu 0001, Yanghao Li, Jisheng Li, Yuxing Han 0001, Jiangtao Wen, Yong-Jin Liu 0001 |
CVPR | 6 |
| 2022 | Vision Perception Unit: Next-Generation Smart CMOS Image SensorabstractAs we reach the end of Moore’s Law and Dennard Scaling, it has become highly desirable to design a highly integrated and optimized pipeline specifically for computer vision. A new generation of integrated "smart" visual processors that streamline an end-to-end optimized visual information acquisition and processing pipeline (VIAPP) becomes necessary to lower the cost, power consumption, and latency.We describe a new paradigm for VIAPP as Vision Perception Unit (VPU), wherein electric signals generated by photons are amplified before converting to the digital signals to emulate an initial layer of a convolutional neural network (CNN). The outputs from these layers are then converted to digital signals and processed by following layers of a deep CNN. Wenqi Ji, Yuxing Han 0001, Jiangtao Wen, Yubin Hu 0001, Futang Wang, Jun Zhang 0006 |
HCS | 3 |
| 2022 | Rate Control for Learned Video CompressionabstractRate control is a critical part for video compression, especially in bandwidth-limited tasks such as live and broadcast. The newly-rising learned video compression has shown advantageous rate-distortion (RD) performance in previous research, but lack of rate control heavily limits its usage in real coding scenarios. In this work, we present the first rate control scheme tailored for learned video compression. Specifically, we explore the inter-frame dependency of learned video compression and propose a novel R-D-λ model accordingly for efficient rate allocation. Additionally, a staged update algorithm is developed for robust parameter estimation. Experiments on public datasets show that, the proposed rate control scheme achieves low rate error while maintaining equal or even higher RD performance, without introducing coding time overhead. Yanghao Li, Jisheng Li, Jiangtao Wen, Yuxing Han 0001, Shan Liu 0001, Xiaozhong Xu |
ICASSP | 4 |
| 2022 | Multiple hierarchical compression for deep neural network toward intelligent bearing fault diagnosis
Jiedi Sun, Jiangtao Wen |
Eng. Appl. Artif. Intell. | 3 |
| 2021 | Decision Tree Based Inter Partition Termination For Av1 EncodingabstractAs a next-generation video coding standard, AV1 introduces numerous new coding tools, leading to high computational complexity and high time cost. To deal with this problem, in this paper, we propose a decision tree based algorithm to early terminate the inter prediction process by predicting splitting decisions at each depth. Motion compensated block is introduced to provide temporal neighborhood information. Nine attributes are selected and analyzed in this paper, and a set of decision trees are generated for different block sizes. According to experimental results, our algorithm can save 23.6% of encoding time on average, with a negligible BD-rate loss of 0.73% under low-delay encoding mode. Yiwei Zhang 0009, Yanghao Li, Jiangtao Wen |
ICASSP | 4 |
| 2021 | Learning to Estimate Kernel Scale and Orientation of Defocus Blur with Asymmetric Coded ApertureabstractConsistent in-focus input imagery is an essential precondition for machine vision systems to perceive the dynamic environment. A de-focus blur severely degrades the performance of vision systems. To tackle this problem, we propose a deep-learning-based framework estimating the kernel scale and orientation of the defocus blur to ad-just lens focus rapidly. Our pipeline utilizes 3D ConvNet for a variable number of input hypotheses to select the optimal slice from the input stack. We use random shuffle and Gumbel-softmax to improve network performance. We also propose to generate synthetic defocused images with various asymmetric coded apertures to facilitate training. Experiments are conducted to demonstrate the effectiveness of our framework. Jisheng Li, Jiangtao Wen |
ICASSP | 3 |
| 2021 | Learning Model-Blind Temporal Denoisers without Ground TruthsabstractDenoisers trained with synthetic noises often fail to cope with the diversity of real noises, giving way to methods that can adapt to unknown noise without noise modeling or ground truth. Previous image-based method leads to noise overfitting if directly applied to temporal denoising, and has inadequate temporal information management especially in terms of occlusion and lighting variation. In this paper, we propose a general framework for temporal denoising that successfully addresses these challenges. A novel twin sampler assembles training data by decoupling inputs from targets without altering semantics, which not only solves the noise overfitting problem, but also generates better occlusion masks by checking optical flow consistency. Lighting variation is quantified based on the local similarity of aligned frames. Our method consistently outperforms the prior art by 0.6-3.2dB PSNR on multiple noises, datasets and network architectures. State-of-the-art results on reducing model-blind video noises are achieved. Yanghao Li, Bichuan Guo, Jiangtao Wen, Zhen Xia, Shan Liu 0001, Yuxing Han 0001 |
ICASSP | 3 |
| 2021 | Learning To Compose 6-DOF Omnidirectional Videos Using Multi-Sphere ImagesabstractOmnidirectional video is an essential component of Virtual Reality. Although various methods have been proposed to generate content that can be viewed with six degrees of freedom (6-DoF), existing systems usually involve complex depth estimation, image inpainting or stitching pre-processing. In this paper, we propose a system that uses a 3D ConvNet to generate a multi-sphere images (MSI) representation that can be experienced in 6-DoF VR. The system utilizes conventional omnidirectional VR camera footage directly without the need for a depth map or segmentation mask, thereby significantly simplifying the overall complexity of the 6-DoF omnidirectional video composition. By using a newly designed weighted sphere sweep volume (WSSV) fusing technique, our approach is compatible with most panoramic VR camera setups. A ground truth generation approach for high-quality artifact-free 6-DoF contents is proposed and can be used by the research and development community for 6-DoF content generation. Jisheng Li, Yubin Hu 0001, Yuxing Han 0001, Jiangtao Wen |
ICIP | 5 |
| 2021 | Extending 6-DoF VR Experience Via Multi-Sphere Images InterpolationabstractThree-degrees-of-freedom (3-DoF) omnidirectional imaging has been widely used in various applications ranging from street maps to 3-DoF VR live broadcasting. Although allowing for navigating viewpoints rotationally inside a virtual world, it does not provide motion parallax key for human 3D perception. Recent research mitigates this problem by introducing 3 transitional degrees of freedom (6-DoF) using multi-sphere images (MSI) which is beginning to show promises in handling occlusions and reflective objects. However, the design of MSI naturally limits the range of authentic 6-DoF experiences, as existing mechanisms for MSI rendering cannot fully utilize multi-layer information when synthesizing novel views between multiple MSIs. To tackle this problem and extend the 6-DoF range, we propose an MSI interpolation pipeline that utilizes adjacent MSIs' 3D information embedded inside their layers. In this work, we describe an MSI projection scheme along with an MSI interpolation network to predict intermediate MSIs in order to facilitate the need for extended range. We demonstrate that our system significantly improves the range of 6-DoF experience compared with other MSI-based methods. With extensive experiments, we show our algorithm outperforms state-of-the-art methods both qualitatively and quantitatively in synthesizing novel view panoramas. Jisheng Li, Jinghui Jiao, Yubin Hu 0001, Yuxing Han 0001, Jiangtao Wen |
ACM Multimedia | 6 |
| 2020 | Deep Material Recognition in Light-Fields via Disentanglement of Spatial and Angular Information
Bichuan Guo, Jiangtao Wen, Yuxing Han 0001 |
ECCV (24) | 2 |
| 2020 | Online Multi-modal Person Search in Videos
Jiangyue Xia, Anyi Rao, Qingqiu Huang, Linning Xu, Jiangtao Wen, Dahua Lin |
ECCV (12) | 5 |
| 2020 | Social Data Assisted Multi-Modal Video Analysis For Saliency DetectionabstractVideo saliency should be taken into consideration to facilitate optimization of the end-to-end video production, delivery and consumption ecosystem to improve user experience at lowered cost. Although recent studies have significantly increased the accuracy of saliency prediction, the approaches are mostly video-centric, without considering any prior "bias" that viewers may have with regard to the video contents. In this paper, we propose a novel learning-based multi-modal method for optimizing user-oriented video analysis. In particular, we generate a face-popularity mask using face recognition results and popularity information obtained from social media, and combine it with conventional content-only saliency analysis to produce multi-modal popularity-motion features. A convolutional long short-term memory (ConvL- STM) network discovers temporal correlation of human attention across frames. Experiments show that our method outperforms the state-of-the-art video saliency prediction approaches in representing human viewing preferences in real world applications, and demonstrate the necessity as well as the potential for integrating user bias information into attention detection. Jiangyue Xia, Jingqi Tian, Jiankai Xing, Jiawen Cheng, Jiangtao Wen, Zhengguang Li, Jian Lou 0003 |
ICASSP | 6 |
| 2020 | Asymmetric Convolutional Residual Network for AV1 Intra in-Loop FilteringabstractIn video compression standards, in-loop filtering plays an important role in alleviating blocking, blurring and ringing artifacts caused by lossy compression, which enhances visual quality and benefits coding efficiency. The boom of neural network applications in super-resolution and image restoration brings insights into solutions of in-loop filtering in video codecs. In this paper, we design an asymmetric convolutional residual network (ACRN) for in-loop filtering in the state-of-the-art AV1 codec. With the asymmetric convolutional blocks, directional features can be extracted to restore textures and improve quality. The cascading structure of wide-activated residual blocks with pruned dense connections enables reflecting hierarchical coding unit (CU) partition characteristics of video coding without losing overall details. Experiments show that the proposed lightweight ACRN can bring up to 12.78% coding efficiency improvement in intra coding of AV1. Jiangyue Xia, Jiangtao Wen |
ICIP | 2 |
| 2020 | Multimodal Video Saliency Analysis With User-Biased InformationabstractVideo saliency is widely used in various video understanding and processing related applications. Despite the fact that studies have indicated the influence of user preferences on visual attention when watching videos, current researches on saliency are based on visual contents and have not taken viewer-related information into account. In this paper, we propose a learning-based multimodal framework to predict video saliency aided by social data analysis. We introduce a popularity assisted attention mechanism into a content-specific neural network to extract spatio-motion features, and utilize a convolutional long short-term memory (ConvLSTM) network to discover temporal characteristics. Experiments demonstrate that our approach outperforms the state-of-the-art video saliency analysis methods, which validates the effectiveness of incorporating external user-biased information into saliency prediction. Jiangyue Xia, Jingqi Tian, Jiangtao Wen, Yuxing Han 0001 |
ICME | 5 |
| 2020 | Genetic Algorithm Based Rate Control for AV1abstractRate control persists as a core problem in video coding area. This paper proposes a robust rate control framework for the recently released AV1 standard, which is a royalty-free video codec specification generated by AOM. Firstly, an exponential rate control model as well as its initialization process and update strategy is proposed to characterize the correspondences between rate and quantization parameter. Next, this paper proposes RC-EMD to measure the similarity between two video frames. Then based on the RC-EMD, a dynamic bit allocation method using genetic algorithm is proposed for AV1 encoder, which can simultaneously detect scene changes. Unlike the existing work, the proposed method explores a larger solution space and automatically adapts to actual input videos, which results in improved performance. Meiyuan Fang, Yuxing Han 0001, Jiangtao Wen |
IEEE Signal Process. Lett. | 3 |
| 2019 | Long and Diverse Text Generation with Planning-based Hierarchical Variational ModelabstractZhihong Shao, Minlie Huang, Jiangtao Wen, Wenfei Xu, Xiaoyan Zhu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhihong Shao, Minlie Huang, Jiangtao Wen, Wenfei Xu, Xiaoyan Zhu 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Mid-depth Based Block Structure Determination for AV1abstractAV1 is an emerging open-source and royalty-free video compression format as a successor to VP9.The increase in coding efficiency and complexity over VP9 is due to the time required to find the optimal partition structure among the more flexible encoding modes for the coding units (CUs) and prediction units (PUs). Due to differences in the frame structure, existing fast block structure determination algorithm cannot be directly applied to AV1. To tackle this problem, we proposed a novel mid-depth based fast block structure determination algorithm for AV1. It checks the partition from mid-depth to provide information for estimating the posterior probabilistic distribution of the partition decisions as well as fast pruning in the PU prediction. Experimental results show that the proposed method can save up to 29.06% time saving with only 0.95% BD-Rate increase. Jiawen Gu, Jiangtao Wen |
ICASSP | 2 |
| 2019 | A Conditional Bayesian Block Structure Inference Model for Optimized AV1 EncodingabstractAV1, a next-generation open-source and royalty-free video coding standard, achieves high compression performance at high computational cost. To meet the requirements of HD and UHD video applications, extensive optimizations in both the algorithm and implementation of AV1 are required. In this paper, we analyze the similarities between the block structure decisions after rate-distortion (RD) optimized AV1 and HEVC encodings of the same input. Taking advantage of such similarities, we propose a conditional Bayesian inference model to perform early termination in block partition determination of AV1 based on HEVC encoding outputs. An estimation algorithm is designed to iteratively calculate the prior probability for Bayesian inference. Experiment results show that our proposed algorithm could realize an average time saving of 35.7% and negligible BD-rate loss (0.61%), with the pre-encoding time taken into consideration. Bichuan Guo, Minhao Tang, Yuxing Han 0001, Jiangtao Wen |
ICME | 5 |
| 2019 | High Efficiency Light Field Compression via Virtual Reference and Hierarchical MV-HEVCabstractEfficient storage and delivery of the light field (LF) information rely on high performance compression. In this paper, we propose a high efficiency light field compression algorithm that utilizes a hierarchical coding structure with synthetic virtual references. Specifically, a LF image are interpreted as a multi-view sequence that is efficiently compressed using the multi-view extension of high efficiency video coding (MV-HEVC). Using deep neural networks, we synthesize virtual references from reconstructed neighbor frames, they serve as extra reference candidates in our novel hierarchical coding structure. Compared with previous work, the proposed algorithm further exploits the intrinsic similarities in LF images. Experimental results show that the proposed algorithms demonstrate a superior performance that achieves up to 55.2% BD-rate reduction and 2.55dB BD-PSNR improvement compared with the HEVC benchmark and outperforms the state-of-the-art. Jiawen Gu, Bichuan Guo, Jiangtao Wen |
ICME | 3 |
| 2019 | AGEM: Solving Linear Inverse Problems via Deep Priors and SamplingabstractIn this paper we propose to use a denoising autoencoder (DAE) prior to simultaneously solve a linear inverse problem and estimate its noise parameter. Existing DAE-based methods estimate the noise parameter empirically or treat it as a tunable hyper-parameter. We instead propose autoencoder guided EM, a probabilistically sound framework that performs Bayesian inference with intractable deep priors. We show that efficient posterior sampling from the DAE can be achieved via Metropolis-Hastings, which allows the Monte Carlo EM algorithm to be used. We demonstrate competitive results for signal denoising, image deblurring and image devignetting. Our method is an example of combining the representation power of deep learning with uncertainty quantification from Bayesian statistics. Bichuan Guo, Yuxing Han 0001, Jiangtao Wen |
NeurIPS | 3 |
| 2019 | SMER: a secure method of exchanging resources in heterogeneous internet of things
Yuxing Han 0001, Jiangtao Wen |
Frontiers Comput. Sci. | 3 |
| 2019 | Hadamard Transform-Based Optimized HEVC Video CodingabstractThe High Efficiency Video Coding (HEVC/H.265) standard achieves great improvement in compression efficiency over the widely used H.264/AVC standard at a cost of much higher complexity. When encoding videos using HEVC, the selection of the quantization parameter (QP) can significantly affect the coding efficiency. Typical algorithms for adaptive quantization employ fixed bitrate budgeting or fixed QP offsets for different frames and different blocks, without considering detailed input video characteristics, some of which might be captured using computer vision methods. The problem of determining adaptive settings of HEVC coding parameters has not been satisfactorily solved. In this paper, we proposed a Hadamard (“HAD”) energy-based optimized HEVC video encoder, in which HAD energy is used to measure the amount of residual information to be encoded in a block so as to determine the QP value for each block for a better coding efficiency. HAD energy is also used to expedite the time-consuming mode decision process and for a precise scene change detection to avoid flicker artifacts. Experiment using the widely used open source HEVC encoder x265-v1.8 showed that the proposed algorithm was able to achieve an average of 14% saving in Bjøntegaard-delta-rate and an average of 20% saving in encoding time as compared the “medium” preset of x265, while the proposed algorithm also produced an improvement of 3.3% in coding efficiency for the HEVC reference software HM-16.6. Minhao Tang, Jiangtao Wen, Yuxing Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | A Universal Optical Flow Based Real-Time Low-Latency Omnidirectional Stereo Video SystemabstractOmnidirectional stereoscopic video (ODSV) is a key element of creating an immersive experience for virtual reality that has attracted extensive interest while presenting many technical challenges. Two such key challenges are real-time, low-latency high-quality seamless video stitching from multiple cameras, and faithful reconstruction of 3-D information. Even though various attempts have been made to achieve different combinations of real-time, low-latency, automation, and high output resolution in stereoscopic panoramic video communication, achieving these characteristics simultaneously remains a challenge to be tackled. In this paper, we present a universally applicable and practical end-to-end system based on a novel real-time optical flow algorithm to produce high-quality real-time ODSV with reconstructed 3-D depth information at low latency. Through a configurable process, various camera systems can be calibrated and seamlessly stitched together using the proposed system. The stitched 3-D panoramic video is encoded with a standard compliant video encoder that is optimized for panoramic video. Thanks to various optimizations introduced in this paper, the proposed system is capable of producing real-time ODSV of ultra High definition resolution with a glass to glass latency of 2.2 s using a desktop computer with a single Nvidia graphic card. Experiments show that the proposed system achieves an encoding performance superior to existing open-source HEVC implementations and an optical flow estimation performance better than the Facebook algorithm while running two orders of magnitudes faster. Minhao Tang, Jiangtao Wen, Jiawen Gu, Philip Junker, Bichuan Guo, Guansyun Jhao, Yuxing Han 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | A Bayesian Approach to Block Structure Inference in AV1-Based Multi-Rate Video EncodingabstractDue to differences in frame structure, existing multi-rate video encoding algorithms cannot be directly adapted to encoders utilizing special reference frames such as AV1 without introducing substantial rate-distortion loss. To tackle this problem, we propose a novel bayesian block structure inference model inspired by a modification to an HEVC-based algorithm. It estimates the posterior probabilistic distributions of block partitioning, and adapts early terminations in the RDO procedure accordingly. Experimental results show that the proposed method provides flexibility for controlling the tradeoff between speed and coding efficiency, and can achieve an average time saving of 36.1% (up to 50.6%) with negligible bitrate cost. Bichuan Guo, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen |
DCC | 5 |
| 2018 | Convex Optimization Based Bit Allocation for Light Field Compression Under Weighting and Consistency ConstraintsabstractCompared with conventional image and video, light field images introduce the weight channel, as well as the visual consistency of rendered view, information that has to be taken into account when compressing the pseudo-temporal-sequence (PTS) created from light field images. In this paper, we propose a novel frame level bit allocation framework for PTS coding. A joint model that measures weighted distortion and visual consistency, combined with an iterative encoding system, yields the optimal bit allocation for each frame by solving a convex optimization problem. Experimental results show that the proposed framework is effective in producing desired distortion distribution based on weights, and achieves up to 24.7% BD-rate reduction comparing to the default rate control algorithm. Bichuan Guo, Yuxing Han 0001, Jiangtao Wen |
DCC | 3 |
| 2018 | Multi-Representations Encoding Framework for Adaptive Http StreamingabstractAdaptive HTTP streaming requires a video to be encoded at multiple representations of different target bitrates. To achieve both smooth streaming and good quality, the representations need to be encoded with accurate rate control and the best quality possible for the target bitrates. However, in practical applications, accurate rate control and good video quality are very difficult to achieve at the same time with one-pass and real-time encoding required by low latency applications. In this paper, we proposed a multi-representation encoding framework that reused the encoding information from low-quality representations to accelerate and optimize higher bitrate encodings. The proposed framework is implemented on two different rate control models, namely R-λ model in HM-16.3 and complexity model in x264, to demonstrate the universality. To be best of our knowledge, the optimization in compression performance of multi - representation is first proposed in this paper. The proposed algorithm can achieve better video quality with only small latency in video coding. Results show that up to 49.6% BDRate savings and 4.43dB BDPSNR improvement are achieved in HM as compared with independent one-pass encodings. Meanwhile, 12.5% time and 14.65% BDRate savings were observed in x264. Jiawen Gu, Jiangtao Wen, Bichuan Guo, Yuxing Han 0001 |
ICIP | 2 |
| 2018 | Fast Block Structure Determination in Av1-Based Multiple Resolutions Video EncodingabstractThe widely used adaptive HTTP streaming requires an efficient algorithm to encode the same video to different resolutions. In this paper, we propose a fast block structure determination algorithm based on the AV1 codec that accelerates high resolution encoding, which is the bottle-neck of multiple resolutions encoding. The block structure similarity across resolutions is modeled by the fineness of frame detail and scale of object motions, this enables us to accelerate high resolution encoding based on low resolution encoding results. The average depth of a block's co-located neighborhood is used to decide early termination in the RDO process. Encoding results show that our proposed algorithm reduces encoding time by 30.1%-36.8%, while keeping BD-rate low at 0.71%-1.04%. Comparing to the state-of-the-art, our method halves performance loss without sacrificing time savings. Bichuan Guo, Yuxing Han 0001, Jiangtao Wen |
ICME | 3 |
| 2018 | A Deep Convolutional Network Based Supervised Coarse-to-Fine Algorithm for Optical Flow MeasurementabstractThe measurement of optical flow is an important problem in image processing. There are a number of methods available for optical flow estimation, including traditional variational methods, deep learning based supervised/unsupervised methods. In this work, we propose a deep convolutional network (CNN) based supervised coarse-to-fine approach, which is trained in end-to-end fashion. The proposed method is tested on standard optical flow benchmark datasets including Flying Chairs, MPI Sintel Clean and Final, KITTI. Experimental results show that the proposed framework is able to achieve comparable results to previous approaches with much smaller network architecture. Meiyuan Fang, Yanghao Li, Yuxing Han 0001, Jiangtao Wen |
MMSP | 4 |
| 2018 | Wavefront Parallel Processing for AV1 EncoderabstractThe emerging AV1 coding standard brings even higher computational complexity than current coding standards, but does not support traditional Wavefront Parallel Processing (WPP) approach due to the lacking of syntax support. In this paper we introduced a novel framework to implement WPP for AV1 encoder that is compatible with current decoder without additional bitstream syntax support, where mode selection is processed in wavefront parallel before entropy encoding and entropy contexts for rate-distortion optimization are predicted. Based on this framework, context prediction algorithms that use same data dependency model as previous works in H.264 and HEVC are implemented. Furthermore, we proposed an optimal context prediction algorithm specifically for AV1. Experimental results showed that our framework with proposed optimal algorithm yields good parallelism and scalability (over 10x speed-up with 16 threads for 4k sequences) with little coding performance loss (less than 0.2% bitrate increasing). Jiangtao Wen |
PCS | 2 |
| 2018 | Adaptive Intra Candidate Selection With Early Depth Decision for Fast Intra Prediction in HEVCabstractTo better exploit spatial correlations in a video frame, the High Efficiency Video Coding (HEVC) standard has adopted a great many more intra prediction modes than H.264/AVC. As a result, the complexity of rate-distortion-optimized (RDO) HEVC intra mode selection is very high. Many techniques have been proposed to expedite the intra mode selection process to achieve a better overall tradeoff between complexity and RD performance. In this paper, two novel techniques for adaptive intra mode candidate selection and bidirectional depth search algorithms are utilized to accelerate intra prediction. The proposed techniques show an average of 63% (up to 67%) time saving with only 1% BD-Rate increase, outperforming most of existing intra prediction algorithms. Jiawen Gu, Minhao Tang, Jiangtao Wen, Yuxing Han 0001 |
IEEE Signal Process. Lett. | 3 |
| 2018 | Accelerating HEVC Encoding Using Early-SplitabstractThe increase in coding efficiency and complexity of high efficiency video coding (HEVC) over H.264 is due to, among other factors, the time needed to find the optimal partition structure among the more flexible encoding modes for the coding units (CUs) and prediction units (PUs). Although many classification-based algorithms have been proposed to expedite the partition decision, the features that can be acquired from current HEVC encoding order are not sufficient to minimize the loss in coding efficiency. In this letter, we proposed an early-split (ES) order for HEVC CU-level encoding, where the encoder checks the split mode before the nonsquare PU partition modes and utilizes the encoding output of the subCUs to expedite subsequent encoding. Experiments show that the proposed algorithm can save 48% of encoding time on average with only about 0.8% loss in coding performance. Minhao Tang, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen, Shiqiang Yang |
IEEE Signal Process. Lett. | 5 |
| 2017 | Probabilistic Graphical Model Based Fast HEVC Inter PredictionabstractIn this article, we propose a probabilistic graphic model based fast HEVC encoding framework. A Bayesian network is characterized by the structure of the network (nodes and edges in the graph) and the probabilistic distributions. It can be constructed in three main steps: 1) data collection and pre-processing, 2) learning network structure, 3) learning parameters of the probabilistic distributions. In the first step, we select a subset of possible HEVC encoding parameters to be modeled by the Bayesian network. Data were collected using the first 150 frames of the HEVC common test condition Class D sequences. Then the model structure is trained so that it is consistent with the conditional dependencies from the observation. Finally, we use a Gaussian Bayesian network to properly model both discrete and continuous valued variables. Based on the Bayesian network and parameters, we can calculate the conditional probabilities for different status in the encoding process and neglect the events with probabilities smaller than a threshold. In this way, a fast HEVC encoding is achieved. Meiyuan Fang, Jiangtao Wen |
DCC | 2 |
| 2017 | SATD Based Fast Intra Prediction for HEVCabstractSummary form only given. To better exploit spatial correlations in a video frame, the HEVC video coding standard has introduced many intra prediction modes and a recursive quadtree-based coding unit (CU) structure. As a result, the complexity of Rate-distortion optimized (RDO) HEVC intra mode selection is significantly higher. Many techniques have been proposed to expedite the intra mode selection process to achieve a good overall trade-off between complexity and RD performance. In this paper, we proposed a fast intra decision algorithm based on Hadamard Transform. The algorithm consist of three parts: SATD calculation reduction, adaptive intra candidate selection, and SATD based early termination. Experiments conducted using the HEVC common test conditions show an average of 56.4% (up to 64.1%) time saving with only 1.2% increase in Bjontegaard delta rate (BD-rate) using the proposed algorithm. Jiawen Gu, Minhao Tang, Jiangtao Wen |
DCC | 3 |
| 2017 | Early-Split Based Fast HEVC EncodingabstractThe High Efficiency Video Coding (HEVC) standard achieves 50% improvement incompression efficiency over the widely used H.264/AVC standard at a cost of much higher complexity. The increase in complexity is due to, among other factors, the time needed to findthe optimal partition structure among the more flexible possibilities for the coding units (CUs) and prediction units (PUs). Many classification based algorithms have been proposed to reduce this partition decision time, but the features that can be acquired from current HEVC encoding order may not be sufficient to control the loss in coding efficiency. In this paper, we proposed an Early-Split (ES) order for HEVC encoding, where the encoder checks the split mode before the non-square PU partition modes and utilizes the encoding output of the subCUs to expedite subsequent encoding. Experiments show that the proposed algorithm achieved an average of 48% saving in encoding time with only 0.92% loss in the coding performance. Minhao Tang, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen |
DCC | 5 |
| 2017 | HEVC-based motion compensated joint temporal-spatial video denoisingabstractA novel HEVC-based efficient video denoising algorithm is proposed in this paper. It uses a spatial Gaussian filter for the chrominance components and then utilizes the HEVC motion estimation process to find the best temporal correspondence for low-pass filtering. Other HEVC tools such as quantization, the interpolation and the in-loop filters are also used. Experiments implementing the proposed algorithm in the open-source HEVC encoder ×265 showed a good denoising performance with a much lower computing complexity than the competitors. The performance was comparable to those highly sophisticated algorithms such as the VBM4D, which is 200 times slower. The proposed algorithm can be easily integrated into the real-world video processing systems due to its compatibility with the HEVC standard. Minhao Tang, Yuxing Han 0001, Jiangtao Wen, Shiqiang Yang |
ICASSP | 3 |
| 2017 | Probabilistic graphical model based fast HEVC inter predictionabstractThe High Efficiency Video Coding (HEVC) standard achieves 50% improvement in coding efficiency compared with H.264/ AVC by introducing many more video encoding tools achieving different coding performance and complexity tradeoffs [1]. Various techniques have been proposed to reduce the complexity of HEVC encoding. In this paper, an effective Bayesian network based complexity reduction framework for HEVC encoding is proposed. The proposed framework is able to calculate the conditional probabilities of different status in the encoding process, which can be used to speedup the encoders by neglecting small probability events. Experimental results show that the proposed algorithm can save 52.5% of computational complexity with a loss of 1.41% in compression performance on average. Meiyuan Fang, Jiangtao Wen, Yuxing Han 0001 |
ICIP | 2 |
| 2017 | A novel satd based fast intra prediction for HEVCabstractTo better exploit spatial correlations in a video frame, the HEVC video coding standard has introduced many intra prediction modes and a recursive quadtree-based coding unit (CU) structure. As a result, the complexity of Rate-distortion optimized (RDO) HEVC intra mode selection is significantly higher. Many techniques have been proposed to expedite the intra mode selection process to achieve a good overall trade-off between complexity and RD performance. In this paper, we proposed a fast intra decision algorithm consisting of three parts: calculation reduction, adaptive intra candidate selection, and fast depth decision used early termination. Experiments conducted using the HEVC common test conditions show an average of 61.1% (up to 67.7%) time saving with only 1.03% increase in Bjontegaard delta rate (BD-rate) using the proposed algorithm, out-performing existing state-of-the-art algorithms. Jiawen Gu, Minhao Tang, Jiangtao Wen |
ICIP | 3 |
| 2017 | Optimized video coding for omnidirectional videosabstractThe ever widening application of virtual reality requires the ultra high resolution omnidirectional videos (OVs) to be transmitted over the wired and wireless Internet at low cost (i.e. bitrate). Various solutions have been proposed to intelligently reduce the bitrate, e.g. adapting the spatial resolution of the video for different directions of the panorama with regard to current direction that the viewer is looking at and the distribution of the probability of each direction to be watched according to the video content. Due to various reasons, spatial resolution adaptation may often cause perceivable quality degradation to user experience. In this paper, we proposed two adaptive encoding techniques to reduce the bitrate of OVs after compression. The first is a content adaptive temporal resolution adaptation scheme for OVs using cube map projection. The second is a quantization and rate-distortion optimization scheme for equirectangular projection. Experiments implementing the proposed algorithms in the open source HEVC encoder x265 show that the proposed algorithms can save 14.69% and 13.01% in BD-rate on average for the cube map and equirectangular projected OVs respectively, while no degradation to the visual experience was reported in subjective tests. The proposed algorithms are compatible with and therefore can collaborate with the currently used adaptive spatial resolution scheme for further bitrate reduction. Minhao Tang, Jiangtao Wen, Shiqiang Yang |
ICME | 3 |
| 2017 | TCP-ACC: performance and analysis of an active congestion control algorithm for heterogeneous networks
Jun Zhang 0006, Jiangtao Wen, Yuxing Han 0001 |
Frontiers Comput. Sci. | 2 |
| 2017 | The IoT electric business model: Using blockchain technology for the internet of things
Jiangtao Wen |
Peer-to-Peer Netw. Appl. | 2 |
| 2016 | Intra Frame Flicker Reduction for Parallelized HEVC EncodingabstractThe existing intra flicker artifact reduction approaches, targeting at one of the major artifacts in current video encoding techniques, are not compatible with the distributed encoding structure, which is increasingly important in modern computing systems. To settle this problem, we propose a flicker reduction approach, which is effective, standard compliant, and especially suitable for parallel and distributed systems. Experimental results show that the proposed approach can reduce the flicker artifact by up to 60% on x265 and 14% on HM. Ziyu Wen, Jisheng Li, Jiangtao Wen |
DCC | 5 |
| 2016 | Novel 3D-WPP algorithms for parallel HEVC encodingabstractAlthough wavefront parallel processing (WPP) proposed in the HEVC standard and various inter frame WPP algorithms can achieve comparatively high parallelism, their scalability for its parallelism is still very limited due to various dependencies introduced in spatial and temporal prediction in HEVC. In this paper, we propose three types of 3 Dimensional WPP (3D-WPP) algorithms that can significantly improve the parallelism, while achieving good tradeoffs between implementation complexity, determinism, and rate-distortion (RD) performance. Experimental results show that the proposed algorithms can lead to up to 2.8 × speed up compared with existing inter frame WPP methods. While the Simple 3D-WPP and Static 3D-, WPP algorithm may introduce an BD rate loss between 0 to 4.9% as compared with existing algorithms, the more complex Dynamic 3D-WPP algorithm achieves better parallelism with virtually no coding performance loss. Ziyu Wen, Bichuan Guo, Jisheng Li, Yao Lu 0006, Jiangtao Wen |
ICASSP | 6 |
| 2016 | Novel tile segmentation scheme for omnidirectional videoabstractRegular omnidirectional video encoding technics use map projection to flatten a scene from a spherical shape into one or several 2D shapes. Common projection methods including equirectangular and cubic projection have varying levels of interpolation that create a large number of non-information-carrying pixels that lead to wasted bitrate. In this paper, we propose a tile based omnidirectional video segmentation scheme which can save up to 28% of pixel area and 20% of BD-rate averagely compared to the traditional equirectangular projection based approach. Jisheng Li, Ziyu Wen, Bichuan Guo, Jiangtao Wen |
ICIP | 6 |
| 2016 | Optimized HEVC encoding with complexity constraintsabstractThe High Efficiency Video Coding (HEVC) standard provides a number of encoding tools with different encoding quality and complexity tradeoffs [1]. Although numerous studies have been conducted on the complexity-compression trade-off of HEVC encoding tools [2], [4]-[11], and based on such studies, many algorithms have been proposed to reduce the complexity of various HEVC encoding modules and tasks, in most cases, these algorithms work in a statistical sense, i.e. best efforts are made to select a sensible subset of HEVC encoding tools that are expected to achieve good rate distortion (RD) results for the current input video sequence. Even though such algorithms have received wide applications in HEVC encoders for real time applications [3], it remains a challenge to achieve the optimal RD HEVC encoding while fully utilizing the available computational resources. In this paper, we propose a framework that addresses the issue of optimal HEVC encoding under an overall complexity constraint. Experimental results show that the proposed system can accurately achieve pre-defined target computational complexity while achieving RD performance superior to other encoders with similar complexity. Meiyuan Fang, Jiangtao Wen |
PCS | 2 |
| 2016 | A novel low delay in-loop filtering WPP process for parallel HEVC encodingabstractWavefront parallel processing (WPP) is a parallelization technique that enables processing of several rows of Largest Coding Units (LCUs) in parallel. It achieves a relatively high level of parallelism with reasonable loss in compression performance. At the same time, the HEVC standard specifies two in-loop filters, namely the deblocking filter and the sample adaptive offset (SAO), to improve subjective quality as well as coding efficiency. Because SAO parameters cannot be precisely determined until the lower right deblocked samples are available, a delay of several rows is introduced to implement the filtering process. Since the reconstruction of the LCUs will not start until finish filtering, the total delay of WPP and filtering is at least two rows, which obviously influences the performance of parallel processing. In this paper, we propose a novel In-Loop Filtering WPP method that reduces the row delay into four LCUs and significantly improves the parallelism with little rate-distortion (RD) performance loss. Experimental results show that the proposed algorithms can achieve up to the 2.89× speedup compared with the existing WPP method with 24-core server, where the speedup improves with the increase of core number. Jiawen Gu, Yuxing Han 0001, Jiangtao Wen |
VCIP | 3 |
| 2016 | Low lighting image enhancement using local maximum color value prior
Xuan Dong 0001, Jiangtao Wen |
Frontiers Comput. Sci. | 2 |
| 2016 | Novel mixing matrix estimation approach in underdetermined blind source separation
Jiedi Sun, Yuxia Li, Jiangtao Wen, Shengnan Yan |
Neurocomputing | 3 |
| 2016 | Metadata Feedback and Utilization for Data Deduplication Across WAN
Jiangtao Wen |
J. Comput. Sci. Technol. | 2 |
| 2016 | Improving Metadata Caching Efficiency for Data Deduplication via In-RAM Metadata Utilization
Jiangtao Wen |
J. Comput. Sci. Technol. | 2 |
| 2016 | A Data Deduplication Framework of Disk Images with Adaptive Block Skipping
Jiangtao Wen |
J. Comput. Sci. Technol. | 2 |
| 2016 | TCP-FIT: An improved TCP algorithm for heterogeneous networks
Jingyuan Wang 0001, Jiangtao Wen, Jun Zhang 0006, Zhang Xiong 0001, Yuxing Han 0001 |
J. Netw. Comput. Appl. | 2 |
| 2015 | R-(lambda) Model Based Improved Rate Control for HEVC with Pre-EncodingabstractIn this paper, we proposed a new rate control algorithm for High Efficiency Video Coding (HEVC), the latest video coding standard from the ITU/ISO. We use the information of pre-encoded 16x16 coding units (CUs) to estimate the characteristics of the largest coding unit (LCU). Based on the estimates, the proposed R -- λ model can be refined before the real encoding process. This is in contrast to rate control algorithms such as that in the HEVC reference software, where the model is updated based on a previously encoded picture. Experimental results show that the proposed rate control scheme can achieve accurate rate control with a BD-PSNR gain up to 5.37dB, compared to the state-of-the-art rate control algorithm in the HEVC test model (HM) 16.0. The largest PSNR improvement was over 6dB. Jiangtao Wen, Meiyuan Fang, Minhao Tang, Kuang Wu |
DCC | 1 |
| 2015 | RETCP: A ratio estimation approach for data center incast applicationsabstractTCP incast has deep impairments to the performance of today's data center applications, especially in many-to-one communication patterns. The impairments can be recognized as gross under-utilization of link capacity, high latency and packets loss rate, which due to the RTO mechanism of TCP proposed last century and the complex applications which have strict demand in performance in data center environment. We propose a delay-based novel and adaptive mechanism for TCP incast prevention in data center applications without modifying the switch or router. Theoretical analysis and simulation results show significant performance improvement as compared with existing methods. Jiangtao Wen |
ICC | 2 |
| 2015 | Improvement of re-sample template matching for lossless screen content videoabstractScreen Content (SC) video coding becomes more important for screen sharing and screen broadcasting applications. There are many easy to see different characters between screen content video and camera-captured video. We proposed a template matching prediction method for lossless SC intra picture coding with higher compression ratio. The pixels are re-sampled to form the Virtual Largest Coding Unit (VLCU) firstly. About 80% pixels in VLCU can be predicted exactly by template matching with zero error. Then, pixels with non-zero prediction error should be coded with three information, index, position and value. Among these three, position will consume the most bits than the other two. In order to handle this challenge, we propose to apply the similarity of non-zero prediction error pixel positions of neighbor VLCU which can greatly help to improve the compression performance. The VLCU can be divided into sub CU as the same as the division in the standard High Efficiency Video Coding(HEVC) intra coding, and RDO is applied to find the best coding efficiency. Pin Tao, Lixin Feng, Sichao Song 0002, Jiangtao Wen, Shiqiang Yang |
ICME | 4 |
| 2015 | An efficient HEVC to H.264/AVC transcoding systemabstractThe latest High Efficiency Video Coding (HEVC) achieves significant compression performance improvement over Advanced Video Coding (AVC). The high compression efficiency of HEVC together with the wide availabilities of H.264/AVC decoders necessitates transcoding between these two technologies. In this paper, we propose a novel algorithm for software-based HEVC to H.264/AVC transcoding. By utilizing the information extracted from the input HEVC stream, the transcoding process can be accelerated with relatively minor compression efficiency loss. Experiment results show that the proposed transcoding algorithm can save around 60% of the time cost of re-encoding process compared with x264, one of the most widely used H.264 encoders, with very small loss of compression performance. Minhao Tang, Jiangtao Wen |
ISCAS | 2 |
| 2015 | A pixel-based outlier-free motion estimation algorithm for scalable video quality enhancement
Xuan Dong 0001, Jiangtao Wen |
Frontiers Comput. Sci. | 2 |
| 2015 | DC-Vegas: A delay-based TCP congestion control algorithm for datacenter applications
Jingyuan Wang 0001, Jiangtao Wen, Chao Li 0001, Zhang Xiong 0001, Yuxing Han 0001 |
J. Netw. Comput. Appl. | 2 |
| 2015 | Efficient Software H.264/AVC to HEVC Transcoding on Distributed Multicore ProcessorsabstractThe latest High Efficiency Video Coding (HEVC) standard achieves a significant compression efficiency improvement over the H.264/Advanced Video Coding (AVC) standard, but with a much higher computational complexity. In this paper, we propose a novel framework for software-based H.264/AVC to HEVC transcoding, integrated with tools such as wavefront parallel processing that are useful for achieving higher levels of parallelism on multicore processors and distributed systems. By utilizing information extracted from the input H.264/AVC bitstream, the transcoding process can be greatly accelerated with a visual quality loss that is modest for many applications. Based on the HEVC HM 14.0 reference software and using standard HEVC test bitstreams, the proposed transcoder can achieve up to 60× speedup on a Quad Core 8-thread server over decoding-re-encoding based on FFMPEG and the HM software with a BD-rate loss of 15%-20%. By implementing a group of picture-level task distribution on a distributed system with nine processing units, the proposed software transcoder can achieve a speed for transcoding 720 p at 30 Hz in real time. Yucong Chen, Ziyu Wen, Jiangtao Wen, Minhao Tang, Pin Tao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Accelerating HEVC using heterogeneous platforms
Gabriel Cebrián-Márquez, José Luis Hernández-Losada, José Luis Martínez 0001, Pedro Cuenca 0001, Minhao Tang, Jiangtao Wen |
J. Supercomput. | 6 |
| 2014 | TCP-ACC: An active congestion compensation TCP for wireless networksabstractTCP is a widely used protocol in modern communication networks. However, the performances of existing TCP congestion control algorithms degrade severely in wireless networks due to wireless-related packet losses and packet reorderings, in addition to congestion. In this paper, a novel TCP algorithm, named TCP-ACC, is proposed. The algorithm detects the level of packet reordering as well as packet losses by combining a packet reordering measurement and congestion control so as to avoid unnecessary slowing down of the data transmission rate while preventing congestion and maintaining good fairness. Theoretical analysis and experiment results show that the algorithm achieves significant throughput improvement in wireless networks as compared with other state-of-the-art algorithms. Jiangtao Wen |
ISCC | 2 |
| 2014 | Achieving high throughput and TCP Reno fairness in delay-based TCP over large networks
Jingyuan Wang 0001, Jiangtao Wen, Yuxing Han 0001, Jun Zhang 0006, Chao Li 0001, Zhang Xiong 0001 |
Frontiers Comput. Sci. | 2 |
| 2013 | Ultra Fast H.264/AVC to HEVC TranscoderabstractThe emerging High Efficiency Video Coding (HEVC) standard achieves significant performance improvement over H.264/AVC standard at a cost of much higher complexity. In this paper, we propose a ultra fast H.264/AVC to HEVC transcoder for multi-core processors implementing Wave front Parallel Processing (WPP) and SIMD acceleration, along with expedited motion estimation (ME) and mode decision (MD) by utilizing information extracted from the input H.264/AVC stream. Experiments using standard HEVC test bit streams show that the proposed transcoder achieves 70x speed up over the HEVC HM 8.1 reference software (including H.264 encoding) at very small rate distortion (RD) performance loss. Yao Lu 0006, Ziyu Wen, Linxi Zou, Yucong Chen, Jiangtao Wen |
DCC | 6 |
| 2013 | Cross Segment Decoding for Improved Quality of Experience for Video ApplicationsabstractIn this paper, we present an improved algorithm for decoding live streamed or pre-encoded video bit streams with time-varying qualities. The algorithm extracts information available to the decoder from a high visual quality segment of the clip that has already been received and decoded, but was encoded independently from the current segment. The proposed decoder is capable of significantly improve the Quality of Experience of the user without incurring significant overhead to the storage and computational complexities of both the encoder and the decoder. We present simulation results using the HEVC reference encoder and standard test clips, and discuss areas of improvements to the algorithm and potential ways of incorporating the technique to a video streaming system or standards. Jiangtao Wen, Shunyao Li, Yao Lu 0006, Meiyuan Fang, Xuan Dong 0001, Huiwen Chang, Pin Tao |
DCC | 1 |
| 2013 | Hysteresis Re-chunking Based Metadata Harnessing Deduplication of Disk ImagesabstractMetadata-related overhead can significantly impact the performance of data deduplication systems, including the real duplication elimination ratio and the deduplication throughput. The amount of metadata produced is mainly determined by the chunking mechanism for the input data stream. In this paper, we propose a metadata harnessing deduplication (MHD) algorithm utilizing a duplication-distribution-based hysteresis re-chunking strategy. MHD harnesses the metadata by dynamically merging multiple non-duplicate chunks into one big chunk represented by one hash value while dividing big chunks straddling duplicate and non-duplicate data regions into small chunks represented with multiple hashes. Experimental results show that the proposed algorithm achieves a lower metadata overhead and a higher deduplication throughput for a given duplication elimination ratio, as compared with other state-of-the-art algorithms such as the Bimodal, Sub Chunk and Sparse Indexing algorithms. Jiangtao Wen |
ICPP | 2 |
| 2013 | A new binarization method for non-uniform illuminated document images
Jiangtao Wen, Shumo Li, Jiedi Sun |
Pattern Recognit. | 1 |
| 2013 | Image Super-Resolution Via Analysis Sparse PriorabstractIn this letter, we present a new algorithm for a single image super-resolution using the analysis sparse prior in thelαβ color space. Experimental results show that our algorithm outperforms other existing state-of-the-art methods. In addition, due to the high scalability of our algorithm, key modules of the proposed algorithm can be integrated with other super resolution algorithms. Qiang Ning, Li Yi 0001, Chuchu Fan, Yao Lu 0006, Jiangtao Wen |
IEEE Signal Process. Lett. | 6 |
| 2012 | Compressive video sensing using non-linear mappingabstractCompressive sensing provides a formalized mathematical framework to acquire and reconstruct sparse signals using sub-Nyquist sampling rate, and has great potential in the application of image and video acquisition and compression. In this paper, by incorporating improved OMP algorithm via non-linear mapping, our proposed compressive video sensing framework has the advantages of lower complexity than that of other convex optimization based framework, and improved reconstruction performance compared with traditional OMP algorithm. Experimental results have demonstrated the effectiveness of our framework. Jiangtao Wen |
ICIP | 2 |
| 2012 | Highly Scalable Parallel Arithmetic Coding on Multi-Core Processors Using LDPC CodesabstractWe describe a highly scalable parallel arithmetic coder for Markov inputs suitable for implementation on modern multi-core processors. The algorithm divides the input into interleaved sub-sequences which can be then processed independently on different processing units using LDPC-based Slepian-Wolf coding. Experimental simulations show good scalability of the proposed algorithm while also maintaining good compression performance. Notably, when compared with traditional parallel arithmetic coding, the proposed method maintains a much higher efficiency both respect to the entropy limit as well as in terms of the ability to distribute computations across multiple cores without performance loss. WeiDong Hu, Jiangtao Wen, Weiyi Wu, Yuxing Han 0001, Shiqiang Yang, John D. Villasenor |
IEEE Trans. Commun. | 2 |
| 2012 | Efficient Video Coding Using Legacy Algorithmic ApproachesabstractWe show that for high bit rates, a video coding algorithm using a suitable combination of the QM coder and on other methods first published over 20 years ago can deliver video quality rivaling that of H.264 at lower complexity. This has implications both technically, since encoders built using these methods can be more power efficient, and commercially, given the complex licensing and intellectual property issues that accompany newer coding methods such as H.264 and MPEG-4. The methods described in this paper are the basis for the recent decision of the MPEG standards group to begin work on what is referred to as the “Type-1 Video Coding” standard, which, in addition to aiming for high coding efficiency, is intended to minimize royalty issues. John D. Villasenor, Yuxing Han 0001, Yaocheng Rong, Cliff Reader, Jiangtao Wen |
IEEE Trans. Multim. | 9 |
| 2011 | A Compressive Sensing Reconstruction Algorithm for Trinary and Binary Sparse Signals Using Pre-mappingabstractIn this paper, we first analyze impact of the distribution of sparse signals on reconstruction quality in compressive sensing through experimental results and heuristic analysis. We suggest that trinary/binary sparse signals are one of the most difficult signals to reconstruct in terms of error bounds. We then show that by incorporating linear or non-linear mapping prior to sensing, significant improvement in the recovery performance can be achieved. Zhuoyuan Chen, Jiangtao Wen, Jianwei Ma 0006, Yuxing Han 0001, John D. Villasenor |
DCC | 3 |
| 2011 | Fast efficient algorithm for enhancement of low lighting videoabstractWe describe a novel and effective video enhancement algorithm for low lighting video. The algorithm works by first inverting an input low-lighting video and then applying an optimized image de-haze algorithm on the inverted video. To facilitate faster computation, temporal correlations between subsequent frames are utilized to expedite the calculation of key algorithm parameters. Simulation results show excellent enhancement results and 4× speed up as compared with the frame-wise enhancement algorithms. Xuan Dong 0001, Yi Pang, Weixin Li 0001, Jiangtao Wen, Wei Meng 0001, Yao Lu 0006 |
ICME | 5 |
| 2011 | TCP-FIT: An improved TCP congestion control algorithm and its performanceabstractThe Transport Control Protocol (TCP) has been widely used by wired and wireless Internet applications such as FTP, email and HTTP. Numerous congestion algorithms have been proposed to improve the performance of TCP in various scenarios, especially for high bandwidth-delay product (BDP) and wireless networks. Although different algorithms may achieve different performance improvements under different network conditions, designing a congestion algorithm that performs well across a wide spectrum of network conditions remains a great challenge. In this paper, we propose a novel congestion control algorithm, named TCP-FIT, which could perform gracefully in both wireless and high BDP networks. The algorithm was inspired by parallel TCP, but with the important distinctions that only one TCP connection with one congestion window is established for each TCP session, and that no modifications to other layers (e.g. the application layer) of the end-to-end system need to be made. Extensive experimental results obtained using both network simulators as well as over “live” wired line, WiFi and 3G networks at different geographical locations and at different times of the day are presented. The performance of the algorithm shown in the experiment results is significantly improved as compared to other state-of-the-art algorithms, while maintaining good fairness. Jingyuan Wang 0001, Jiangtao Wen, Jun Zhang 0006, Yuxing Han 0001 |
INFOCOM | 2 |
| 2011 | Probabilistic Estimation of the Number of Frequency-Hopping TransmittersabstractWe present two probabilistic estimation techniques for identifying the most likely number of frequency-hopping transmitters for a range of different scenarios and compare their performances. In the first technique, cumulative estimation, a Gaussian approximation methodology is developed based on a single integrated measurement over the observation time window. In the second technique, time-distributed estimation, a maximum-likelihood formulation is adopted, and time-specific observation data is used. We give specific analytical consideration to the potential that not all transmitters that are present will be always detected, and explore the effects of the probability of misdetection on the overall estimation process. Simulation results confirm that the approaches presented here can lead with high probability to a correct decision regarding the number of transmitters. Yuxing Han 0001, Jiangtao Wen, Danijela Cabric, John D. Villasenor |
IEEE Trans. Wirel. Commun. | 2 |
| 2010 | Image Compression Using the DCT and Noiselets: A New Algorithm and Its Rate Distortion PerformanceabstractWe describe an image coding algorithm combining the DCT and noiselet information. The algorithm first transmits DCT information sufficient to reproduce a "low-quality" version of the image at the decoder. This image is then used both at the decoder and encoder to create a mutually known list of locations of likely significant noiselet coefficients. The coefficient values themselves are then transmitted to the decoder differentially, by subtracting, at the encoder, the low-quality image from the original image, obtaining the noiselet values and subjecting them to quantization and entropy coding. There remain significant opportunities for further work combining CS-inspired information theoretic techniques with the rate-distortion considerations that are critical in practical image communications. Zhuoyuan Chen, Jiangtao Wen, Shiqiang Yang, Yuxing Han 0001, John D. Villasenor |
DCC | 2 |
| 2010 | Horizontal Spatial Prediction for High Dimension Intra CodingabstractMacroblock level Horizontal Spatial Prediction(HSP) based intra frame coding scheme for High Dimension(HD) video sequences was proposed in this paper. According to the correlation experiment on HD sequences, most HD pictures have the stronger horizontal spatial correlation than the vertical spatial correlation, about 2dB stronger. This phenomena drop a valuable hint to us that the horizontal spatial prediction can be used in HD video intra coding without considering the vertical spatial prediction. An adaptive divide and predict intra frame coding scheme has been proposed by Piao which has the similar idea. But this method divided the whole picture into several parts which is not conform to the conventional macroblock based video coding framework and it has the high computation complexity in motion estimation procedure. Pin Tao, Wenting Wu, Chao Wang 0063, Mou Xiao, Jiangtao Wen |
DCC | 5 |
| 2010 | Reconstruction of Sparse Binary Signals Using Compressive SensingabstractSummary form only given. This paper has described an improved algorithm for reconstructing sparse binary signals using compressive sensing. The algorithm is based on the reweighted lqnorm optimization algorithm, but with the important additional operation of bounding in each round of the interior-point method iteration, and progressive reduction of q. Experimental results confirm that the algorithm performs well both in terms of the ability to recover an input signal as well as in terms of speed. We also found that both the progressive reduction and the bounding are integral to the improvement in performance. Future work includes extending this approach to Gaussian distributed, as opposed to binary inputs. Jiangtao Wen, Zhuoyuan Chen, Shiqiang Yang, Yuxing Han 0001, John D. Villasenor |
DCC | 1 |
| 2010 | Fast Rate Distortion Optimized Quantization for H.264/AVCabstractIn this paper, a fast RDO (rate-distortion optimization) quantization algorithm for H.264/AVC is proposed. In this algorithm, the searching space of level adjustments is reduced by filtering the input quantized coefficients in a hierarchical way. The well quantized coefficients is first filtered out, and then the RD tradeoff of each level adjustment to each of the rest coefficients is examined to select some good candidates with their associated level adjustments. Finally these good candidates are combined to find the best combination of level adjustments which gives the minimal rate-distortion cost. Furthermore, a fast rate estimation technique is adopted to save the rate-distortion estimation time. Experimental results show that about 44% quantization time on average can be saved at the cost of negligible PSNR loss compared with RDO quantization algorithm implemented in JM. Jiangtao Wen, Mou Xiao, Pin Tao, Chao Wang 0063 |
DCC | 1 |
| 2010 | A compressive sensing image compression algorithm using quantized DCT and noiselet informationabstractInspired by recent theoretical advances in compressive sensing (CS), we propose a new framework that combines the classical local discrete cosine transform used in image compression algorithms such as JPEG with a global noiselet measure which is solved using second order cone programming (SOCP). Jiangtao Wen, Zhuoyuan Chen, Yuxing Han 0001, John D. Villasenor, Shiqiang Yang |
ICASSP | 1 |
| 2010 | A Probabilistic Approach to Identifying the Number of Transmitters in the Presence of MisdetectionabstractWe present an analytical framework for identifying the most likely number of frequency-hopping transmitters in the presence of potential misdetection. The problem is formulated for the case where there is a single, global misdetection probability characterizing all transmitter/sensor links, and for the case where the misdetection probabilities are permitted to be different across the different pairwise transmitter/sensor links. Simulation results confirm that the approach can lead with high probability to a correct decision regarding the number of interferers. Yuxing Han 0001, Jiangtao Wen, Danijela Cabric, Sateesh Addepalli, John D. Villasenor |
ICC | 2 |
| 2010 | Improved intra prediction for high definition video using localized horizontal spatial predictionabstractIn this paper, a localized horizontal spatial prediction (HSP) based algorithm was proposed for Intra coding of high definition (HD) inputs. In the algorithm, a block of size 32×16 is divided into two 16×16 MBs, consisting of the pixels from the even-numbered columns and the odd-numbered columns (termed the even MB and the odd MB) respectively. The even MB is encoded using conventional Intra coding techniques and then its reconstruction is used for the prediction of the odd MB. Experimental results show that up to 0.79 dB and on average 0.39 dB gain can be achieved for HD sequences with the proposed framework at lower complexity than H.264 Intra coding. Wenting Wu, Pin Tao, Mou Xiao, Jiangtao Wen, Ruiping Li |
ICIP | 4 |
| 2010 | AN efficient algorithm for joint QP and quantization optimization for H.264/AVCabstractWe describe an efficient algorithm for jointly optimizing the quantization parameter (QP) and the quantization decisions in H.264/AVC. Based on bitrate estimation for quantized and entropy coded coefficients in H.264/AVC, a technique was introduced to select only a small subset of transform coefficients which have the most significant impact on the overall quantization performance. Only the quantization decisions for the coefficients in this subset need to be explicitly examined along with the QP. Simulation results showed that the proposed algorithm achieves similar quantization performances as existing state-of-the-art optimized algorithms at roughly half the complexity. Mou Xiao, Jiangtao Wen, Pin Tao |
ICIP | 2 |
| 2010 | Macroblock level hybrid temporal-spatial prediction for H.264/AVCabstractIn this paper, a novel macroblock level hybrid temporal-spatial video coding framework is proposed. In this framework, a new Hybrid Temporal Spatial Prediction (HTSP) coding mode is adopted for each macroblock. Macroblock is divided into two partitions. The first partition is temporally predicted using motion compensation and encoded, while the second partition is spatially predicted using the reconstruction of the first partition. Experimental results show that up to 0.4 dB coding gain and 0.2 dB on average can be achieved for HD sequences. Mou Xiao, Pin Tao, Wenting Wu, Jiangtao Wen |
ISCAS | 5 |
| 2009 | A Probabilistic Approach to Identifying the Number of Frequency Hoppers for Spectrum SensingabstractCharacterizing the number and type of transmitters occupying a given frequency band is a critical aspect of spectrum sensing specifically and cognitive radio generally. We present an analytical framework based on probability to identify the number of frequency hopping transmitters of one specific type in a band of interest, and show that the probability mass functions associated with the different potential number of transmitters quickly becomes Gaussian as the number of channel observations increases. Simulation results confirm that the approach can lead with high probability to a correct decision regarding the number of interferers. Thus, the methods here can serve as a valuable complement to other spectrum sensing approaches. Yuxing Han 0001, Shaunak Joshi, Lillian L. Dai, Danijela Cabric, Sateesh Addepalli, Jiangtao Wen, John D. Villasenor |
GLOBECOM | 6 |
| 2009 | Frame-level heuristic scheduling Multi-view Video Coding on symmetric multi-core architectureabstractIn this paper, we propose a frame-level heuristic scheduling parallel emerging Multi-view Video Coding (MVC) using Directed Acyclic Graph (DAG) on Intel multi-core processor. We illustrate the reason to choose heuristic scheduling and formulate the problem. Through defining dependent degree and concurrent degree, we demonstrate why to choose frame as parallel granularity. Experimental results demonstrate the effectiveness and scalability of our parallel MVC. Yi Pang, Jiangtao Wen, Lifeng Sun, WeiDong Hu, Shiqiang Yang |
ICIP | 2 |
| 2009 | Key issues in secure, error resilient compressive sensing of multimedia contentabstractMultimedia applications have played an increasingly important role in entertainment, security, remote sensing, monitoring and other civil and military applications. The traditional approach to multimedia acquisition, communication and consumption chain starts with high resolution content acquisition, followed by content processing, image and video compression and then delivery over wireless and wired line networks. After decades of theory, technology and product developments, the traditional paradigm has approached the limits established by a number of physical barriers. In this talk, we review a few critical issues when applying compressive sensing theory to multimedia applications and propose, on a high level, some preliminary potential solutions. Jiangtao Wen |
ICME | 1 |
| 2009 | A Framework for Heuristic Scheduling for Parallel Processing on Multicore Architecture: A Case Study With Multiview Video CodingabstractIn this paper, using the Intel multicore architectures and the emerging multiview video coding standard, we introduce a framework for performing analysis, simulation, and evaluation of heuristics scheduling algorithms for implementing computationally intensive algorithms on multicore processors. The framework allows for accurate and quantitative characterization of the performance of dynamic scheduling algorithms for multimedia applications on different multicore processors without actual implementation of the scheduling algorithm and application on the actual platform. Experimental results demonstrate the effectiveness and scalability of our framework. Yi Pang, Lifeng Sun, Jiangtao Wen, Fengyan Zhang, WeiDong Hu, Shiqiang Yang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | Binary arithmetic coding with key-based interval splittingabstractBinary arithmetic coding involves recursive partitioning the range [0,1) in accordance with the relative probabilities of occurrence of the two input symbols. We describe a modification of this approach in which the overall length within the range [0,1) allocated to each symbol is preserved, but the traditional assumption that a single contiguous interval is used for each symbol is removed. A key known to both the encoder and decoder is used to describe where the intervals are "split" prior to encoding each new symbol. The repeated splitting has the effect of both scrambling the intervals and altering their lengths, thereby allowing both encryption and compression to be obtained simultaneously. Jiangtao Wen, Hyungjin Kim 0003, John D. Villasenor |
IEEE Signal Process. Lett. | 1 |
| 2004 | Heuristic Search Based Soft-Input Soft-Output Decoding of Arithmetic CodesabstractThe paper described an optimal search based and a heuristic search based decoding algorithm for arithmetic codes and compared their performance and complexity with the traditional "hard" bit based AC decoder. To approach a different error resilient AC decoding, a heuristic search algorithms (HSAs) in artificial intelligence for finding the minimal-weight path through directed and nonnegatively-weighted graphs is utilized when the bitstream is transmitted over error prone channels. Simulation results showed that both algorithms easily outperformed traditional "hard" bits based arithmetic decoder, while the heuristic search based algorithm achieved a very good tradeoff between performance and complexity. Jiangtao Wen |
Data Compression Conference | 2 |
| 2002 | 3G wireless multimedia: technologies and practical issuesabstractThis paper provides an overview of the emerging wireless communication standards, end-to-end wireless streaming systems, and relevant wireless multimedia technologies. It highlights some of the challenges in the deployment of 3G wireless multimedia services, using PacketVideo's solutions as an example. Wenjun Zeng 0001, Jiangtao Wen |
ICIP (1) | 2 |
| 2002 | Fast self-synchronous content scrambling by spatially shuffling codewords of compressed bitstreamsabstractThis paper presents a content access control method by spatially shuffling codewords of the compressed bitstream. This approach is lightweight and incurs no bit overhead. One important advantage of the approach is that the resulting scrambled bitstream can be made compliant to the compression format, thus providing some level of scalability, error resiliency, network friendliness and capability of performing signal processing directly on the encrypted bitstream. In addition, we propose a method for generating or updating the shuffling tables on the fly, based on encrypting some local-content-specific bits using a standard cipher. This local-content-specific bits based table generation process is self-synchronous, which is critical in the presence of packet loss. It also enhances the resistance of this spatial shuffling approach to plain-text attack. Wenjun Zeng 0001, Jiangtao Wen, Mike Severa |
ICIP (3) | 2 |
| 2002 | Soft-input soft-output decoding of variable length codesabstractWe present a method for utilizing soft information in decoding of variable length codes (VLCs). When compared with traditional VLC decoding, which is performed using "hard" input bits and a state machine, the soft-input VLC decoding offers improved performance in terms of packet and symbol error rates. Soft-input VLC decoding is free from the risk, encountered in hard decision VLC decoders in noisy environments, of terminating the decoding in an unsynchronized state, and it offers the possibility to exploit a priori knowledge, if available, of the number of symbols contained in the packet. Jiangtao Wen, John D. Villasenor |
IEEE Trans. Commun. | 1 |
| 2002 | A format-compliant configurable encryption framework for access control of videoabstractWe introduce new methods of performing selective encryption and spatial/frequency shuffling of compressed digital content that maintain syntax compliance after content has been secured. The tools described have been proposed to the MPEG-4 Intellectual Property Management and Protection (IPMP) standardization group and have been adopted into the MPEG-4 IPMP Final Proposed Draft Amendment (FPDAM). We describe the application of the new methods to the protection of MPEG-4 video content in the wireless environment, and illustrate how they are used to leverage established encryption algorithms for the protection of only the information fields in the bitstream that are critical to the reconstructed video quality, while maintaining compliance to the syntax of MPEG-4 video, and thereby reduces the amount of data to be encrypted and guarantees the inheritance of many of the good properties of the unprotected bitstreams that have been carefully studied and built, such as error resiliency and network friendliness. The encrypted content bitstream works with many existing random access, network bandwidth adaptation, and error control techniques that have been developed for standard-compliant compressed video, thus making it especially suitable for wireless multimedia applications. Standard compliance also allows subsequent signal processing techniques to be applied to the encrypted bitstream. Jiangtao Wen, Mike Severa, Wenjun Zeng 0001, Max H. Luttrell, Weiyin Jin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | 3G wireless multimedia: technologies and practical issuesabstractAbstract This paper provides overviews of the emerging wireless communication standards, end‐to‐end wireless streaming systems, and relevant wireless multimedia technologies. It highlights some of the challenges in the deployment of 3G wireless multimedia services, using PacketVideo's solutions as an example. Copyright © 2002 John Wiley & Sons, Ltd. Wenjun Zeng 0001, Jiangtao Wen |
Wirel. Commun. Mob. Comput. | 2 |
| 2001 | A format-compliant configurable encryption framework for access control of multimediaabstractWe present a framework for access control of standard-compliant video bitstreams for entertainment purposes. The approach leverages well-known encryption algorithms and maintains standard-compliance of the encrypted bitstream. The standard compliance feature guarantees the inheritance of the error resiliency properties of the video compression standards, and works with many existing network bandwidth adaptation and error control techniques that have been developed for standards-compliant compressed video, thus making it especially suitable for wireless multimedia applications. Standards compliance also allows subsequent signal processing techniques to be applied to the encrypted bitstream. The approach can be applied to any of the common video coding standards. It is also capable of providing layered compromises between security, complexity, delay, bit overhead, etc. Jiangtao Wen, Mike Severa, Wenjun Zeng 0001, Max H. Luttrell, Weiyin Jin |
MMSP | 1 |
| 2000 | Trellis-Based R-D Optimal Quantization in H.263+abstractWe describe a trellis-based algorithm which enables R-D optimum quantization decisions in the H.263+ video coding standard. The use of the trellis allows the quantization decisions for all coefficients in a block to be made jointly, and contrasts with more commonly used R-D optimizations which operate on one coefficient at a time. Experiments conducted using H.263+ for video coding rates of 40-50 kbps show an average improvement of 3.5% in bit rate, or equivalently 0.17 dB in PSNR, relative to implementations which follow the International Telecommunications Union (ITU) Test Model specifications in making quantization decisions. Max H. Luttrell, Jiangtao Wen, John D. Villasenor |
ICIP | 2 |
| 2000 | Priority Dropping in Network Transmission of Scalable VideoabstractBy constructing a model which takes into account the frame dependence of multimedia streams, we analyze the performance of different packet dropping mechanisms and find that scalable video combined with the priority dropping mechanism can bring higher throughput, lower delay and lower delay jitter. We also find the optimum system parameters for multimedia transmission. This study is important for achieving better performance for video transmission over IP networks. Tao Tian, Adam H. Li, Jiangtao Wen, John D. Villasenor |
ICIP | 3 |
| 2000 | Trellis-based R-D optimal quantization in H.263+abstractWe describe a trellis-based algorithm which enables R-D optimum quantization decisions in the H.263+ video coding standard. The algorithm allows the quantization decisions for all coefficients in a block to be made jointly, and experiments conducted using H.263+ for video coding rates of 40-50 kbps show an average improvement of 3.5% in bitrate relative to implementations which follow the International Telecommunications Union (ITU) Test Model specifications in making quantization decisions. Jiangtao Wen, Max H. Luttrell, John D. Villasenor |
IEEE Trans. Image Process. | 1 |
| 1999 | Utilizing Soft Information in Decoding of Variable Length CodesabstractWe present a method for utilizing soft information in decoding of variable length codes (VLCs). When compared with traditional VLC decoding, which is performed using "hard" input bits and a state machine, soft-input VLC decoding offers improved performance in terms of packet and symbol error rates. Soft-input VLC decoding is free from the risk, encountered in hard decision VLC decoders in noisy environments, of terminating the decoding in an unsynchronized state, and it offers the possibility to exploit a priori knowledge, if available, of the number of symbols contained in the packet. Jiangtao Wen, John D. Villasenor |
Data Compression Conference | 1 |
| 1999 | Robust video coding algorithms and systemsabstractWireless video communication is particularly challenging because it combines the already difficult problem of efficient compression with the additional and usually contradictory need to make the compressed bit stream robust to channel errors. We describe design and implementation strategies for error-robust video communications with an emphasis on techniques compatible with the coding approaches used in the ISO (MPEG-4) and ITU standards organizations. These techniques include modifications to the video coding algorithms as well as to the system layers that perform packetization and multiplexing. John D. Villasenor, Ya-Qin Zhang, Jiangtao Wen |
Proc. IEEE | 3 |
| 1999 | Structured Prefix Codes for Quantized Low-Shape-Parameter Generalized Gaussian SourcesabstractThe highly peaked wide-tailed pdf's that are encountered in many image coding algorithms are often modeled using the family of generalized Gaussian (GG) pdf's. We study entropy coding of quantized GG sources using prefix codes that are highly structured, and which therefore involve low computational complexity. We provide bounds for the redundancy associated with applying these codes to quantized GG sources. We also explore code efficiency and code choice for a wide range of GG source and quantizer parameters. Jiangtao Wen, John D. Villasenor |
IEEE Trans. Inf. Theory | 1 |
| 1998 | Reversible Variable Length Codes for Efficient and Robust Image and Video CodingabstractThe International Telecommunications Union (ITU) has adopted reversible variable length codes (RVLCs) for use in the emerging H.263+ video compression standard. As the name suggests, these codes can be decoded in two directions and can therefore be used by a decoder to enhance robustness in the presence of transmission bit errors. In addition, these RVLCs involve little or no efficiency loss relative to the corresponding non-reversible variable length codes. We present the ideas behind two general classes of RVLCs and discuss the results of applying these codes in the framework of the H.263+ and MPEG-4 video coding standards. Jiangtao Wen, John D. Villasenor |
Data Compression Conference | 1 |
| 1997 | A Class of Reversible Variable Length Codes for Robust Image and Video CodingabstractWe describe a class of parameterized reversible variable length codes that have length distributions identical to Golomb-Rice codes and exp-Golomb codes. The pdfs to which these codes correspond are well matched to statistics of image and video data, thus enabling an increase in robustness to channel errors with no penalty in coding efficiency. These codes are applicable to MPEG-4 and other algorithms that aim to use variable length codes in error-prone environments. Jiangtao Wen, John D. Villasenor |
ICIP (2) | 1 |