VLDB 2026 Research / reviewers in the wild / expert
Huifang Sun
dblp:18/4852
· DBLP profile ↗
83ranked-venue papers
7as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 66 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 5Applied, interdisciplinary, general and emerging computing · 5Computer networks · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Federated Learning via Clients-to-Server Knowledge Distillation (Student Abstract)abstractTo diminish the substantial communication costs incurred by federated learning during the training of the global model and enhance the model update efficiency across both clients and server domains, we have integrated knowledge distillation into the federated learning framework. This integration has led to the development of a novel approach termed ClientsToServerKDFL, which streamlines the distillation process by directly transferring model insights from clients to the server for computational learning without the need for extensive computations across numerous clients. This iterative process ensures model accuracy and curtails communication expenses. Experimental data analysis has validated the efficacy of this algorithm. Huifang Sun, Jiaming Pei, Lukun Wang |
AAAI | 1 |
| 2025 | Emerging Advances in Learned Video Compression: Models, Systems and BeyondabstractVideo compression is a fundamental topic in the visual intelligence, bridging visual signal sensing/capturing and high-level visual analytics. The broad success of artificial intelligence (AI) technology has enriched the horizon of video compression into novel paradigms by leveraging end-to-end optimized neural models. In this survey, we first provide a comprehensive and systematic overview of recent literature on end-to-end optimized learned video coding, covering the spectrum of pioneering efforts in both uni-directional and bi-directional prediction based compression model designation. We further delve into the optimization techniques employed in learned video compression (LVC), emphasizing their technical innovations, advantages. Some standardization progress is also reported. Furthermore, we investigate the system design and hardware implementation challenges of the LVC inclusively. Finally, we present the extensive simulation results to demonstrate the superior compression performance of LVC models, addressing the question that why learned codecs and AI-based video technology would have with broad impact on future visual intelligence research. Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001, Huifang Sun, Leonardo Chiariglione |
IJCAI | 5 |
| 2025 | Exploring Rich Subjective Quality Information for Image Quality Assessment in the WildabstractTraditional in the wild image quality assessment (IQA) models are generally trained with the quality labels of mean opinion score (MOS), while missing the rich subjective quality information contained in the quality ratings, for example, the standard deviation of opinion scores (SOS) or even distribution of opinion scores (DOS). In this paper, we propose a novel IQA method namedRichIQAto explore the rich subjective rating information beyond MOS to predict image quality in the wild. RichIQA is characterized by two key novel designs: 1) a three-stage image quality prediction network which exploits the powerful feature representation capability of the Convolutional vision Transformer (CvT) and mimics the short-term and long-term memory mechanisms of human brain; 2) a multi-label training strategy in which rich subjective quality information like MOS, SOS and DOS are concurrently used to train the quality prediction network. Powered by these two novel designs, RichIQA is able to predict the image quality in terms of a distribution, from which the mean image quality can be subsequently obtained. Extensive experimental results verify that the three-stage network is tailored to predict rich quality information, while the multi-label training strategy can fully exploit the potentials within subjective quality rating and enhance the prediction performance and generalizability of the network. RichIQA outperforms state-of-the-art competitors on multiple large-scale in the wild IQA databases with rich subjective rating labels. The code of RichIQA will be made publicly available on GitHub. Xiongkuo Min, Yuqin Cao, Guangtao Zhai, Wenjun Zhang 0001, Huifang Sun, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | MPAI-EEV: Standardization Efforts of Artificial Intelligence Based End-to-End Video CodingabstractThe rapid advancement of artificial intelligence (AI) technology has led to the prioritization of standardizing the processing, coding, and transmission of video using neural networks. To address this priority area, the Moving Picture, Audio, and Data Coding by Artificial Intelligence (MPAI) group is developing a suite of standards called MPAI-EEV for "end-to-end optimized neural video coding." The aim of this AI-based video standard project is to compress the number of bits required to represent high-fidelity video data by utilizing data-trained neural coding technologies. This approach is not constrained by how data coding has traditionally been applied in the context of a hybrid framework. This paper presents an overview of recent and ongoing standardization efforts in this area and highlights the key technologies and design philosophy of EEV. It also provides a comparison and report on some primary efforts such as the coding efficiency of the reference model. Additionally, it discusses emerging activities such as learned Unmanned-Aerial-Vehicles (UAVs) video coding which are currently planned, under development, or in the exploration phase. With a focus on UAV video signals, this paper addresses the current status of these preliminary efforts. It also indicates development timelines, summarizes the main technical details, and provides pointers to further points of reference. The exploration experiment shows that the EEV model performs better than the state-of-the-art video coding standard H.266/VVC in terms of perceptual evaluation metric. Chuanmin Jia, Fanke Dong, Leonardo Chiariglione, Siwei Ma 0001, Huifang Sun, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Learning to Compress Unmanned Aerial Vehicle (UAV) Captured Video: Benchmark and AnalysisabstractIn this paper, we propose to build a novel benchmark and neural video coding task named learning based Unmanned Aerial Vehicle (UAV) video coding. We collect the UAV videos with different content variations, including in-door and out-door scenes, object-scale variations and viewpoint distance, different climate condition etc. Then we encode those properly-selected videos using popular end-to-end optimized video codecs and conventional hybrid codecs, to form a comprehensive benchmark for learned drone video compression. We also provide a detailed analysis and envision the challenge of such task for future research. The main contributions of this paper are three folds. First, we construct a comprehensive benchmark for the task of drone video compression which consists of the rate-distortion (R-D) behavior of both hybrid and learned video codecs. To our knowledge, it is the first attempt in end-to-end optimized solution to compress drone videos. Second, we provide the review and analysis of the learned drone video compression schemes and further discuss the challenges of encoding UAV videos. Third, this benchmark and related research is accomplished as a milestone MPAI End-to-end Video (EEV) coding project. The proposed benchmark has constructed a solid baseline for compressing UAV videos and facilitates the future research works for related task. Chuanmin Jia, Huifang Sun, Siwei Ma 0001, Wen Gao 0001 |
DCC | 3 |
| 2022 | Tensor Product and Tensor-Singular Value Decomposition Based Multi-Exposure Fusion of ImagesabstractConsidering multidimensional structure of the multi-exposure images, a new Tensor product and Tensor-singular value decomposition based Multi-Exposure image Fusion (TT-MEF) method is proposed. The main innovation of this work is to explore a new feature representation of multi-exposure images in the new tensor domain and design the fusion strategy on this basis. Specifically, the luminance and the chrominance channels are fused separately to maintain color consistency. For the luminance fusion, the luminance channel of multi-exposure images is divided into two parts, that is, de-mean term and mean term. The de-mean term is represented as a tensor to extract the feature. Then, the tensor product and tensor-singular value decomposition (T-SVD) are used to design a tensor feature extractor. Furthermore, a fusion strategy of the de-mean term is presented according to the visual saliency model, and a fusion strategy of the mean term is defined by the local and the global visual weights to control counterpoise between the local and global luminance. For the chrominance fusion, a new fusion strategy is also designed by the tensor product and T-SVD, similar to the luminance fusion. Finally, the fused image is obtained by combining the luminance and chrominance fusion. Experimental results show that the proposed TT-MEF method generally outperforms the existing state-of-the-art in terms of subjective visual quality and objective evaluation. Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Zhongjie Zhu, Yongqiang Bai, Yang Song 0015, Huifang Sun |
IEEE Trans. Multim. | 7 |
| 2021 | Divisively Normalized Sparse Coding: Toward Perceptual Visual Signal RepresentationabstractSparse representation has been shown to be highly correlated with the visual perception of natural images, which can be characterized by a linear combination of neuronal responses in the visual cortex. Divisive normalization transform (DNT) has been proven to be an effective method in reducing statistical and perceptual dependencies for nonlinear properties in primary visual cortex. In this paper, we develop a divisively normalized sparse coding scheme, aiming to further bridge the gap between sparse representation and human visual perception. We show that such a scheme is perceptually meaningful for representing visual signals, with which the pixel-domain image representation and processing tasks can be feasibly and efficiently achieved in the divisively normalized sparse-domain. Specifically, we develop a sparse-domain similarity (SDS) index for perceptual quality evaluation, where the DNT is employed for transforming image signals into a perceptually uniform space. Furthermore, the proposed SDS index is employed to optimize the sparse coding process when representing natural images. The experimental results indicate that the SDS can provide accurate and consistent predictions of perceived image quality, and the performance of sparse coding can be significantly improved in terms of both objective and subjective quality evaluations. Xiang Zhang 0004, Siwei Ma 0001, Shiqi Wang 0001, Jian Zhang 0018, Huifang Sun, Wen Gao 0001 |
IEEE Trans. Cybern. | 5 |
| 2021 | Blind Quality Assessment of Screen Content Images Via Macro-Micro Modeling of Tensor Domain DictionaryabstractScreen content images (SCIs) have been rapidly and widely applied in interactive multimedia applications. The problem of quality assessment for SCIs is an interesting research topic. Most of the existing methods use subjective and independent features in gray domain to predict the image quality, which cannot comprehensively characterize the image properties or lack unified mathematical explanation for SCIs. To address these problems, we propose a novel blind quality assessment method based on macro-micro modeling of tensor domain dictionary for SCIs in this article. In the proposed method, the tensor decomposition is explored first to avoid the loss of color information, and then a target dictionary is learned more effectively with the principal components. Furthermore, a macro-micro model is established to characterize the micro and macro features in the target dictionary space, which can provide a systematic mathematical interpretation for feature extraction. For the micro features, a log-normal pooling scheme is designed to enhance the effectiveness of feature aggregation by analyzing the particularity of the statistical distribution of sparse codes. Additionally, the statistical properties are mainly discussed and studied based on the Bernoulli law of large numbers, and then a reliable macro feature is generated to describe the relationship between the statistical distribution and quality degradation of SCIs. Experimental results determined by using three public SCI databases show that the proposed method can perform better than relevant existing methods in the prediction of the visual quality of SCIs, especially in terms of the generalization for distortion type and interpretability for feature generation. Yongqiang Bai, Zhongjie Zhu, Gangyi Jiang, Huifang Sun |
IEEE Trans. Multim. | 4 |
| 2020 | Collaborative Localization Based on Traffic Landmarks for Autonomous DrivingabstractLocalizing an autonomous vehicle in real-time is critical for robust autonomous driving. As a standard approach, the map-based localization is robust and fast; however, it is expensive to create and maintain a large-scale high-definition map. In this paper, we propose an online localization technique based on the vehicle-to-vehicle communication and traffic landmark detection; called collaborative localization. This can potentially serve as a new complement to the standard localization solutions. We theoretically show that multiple vehicles with multiple traffic landmarks would significantly improve the localization performance. We then propose a practical algorithm, which leverages graph matching to handle practical issues, such as traffic landmark association. The experimental results validate the potential of the proposed methods. Siheng Chen, Ningxiao Zhang, Huifang Sun |
ISCAS | 3 |
| 2020 | Extracting influence relationships in China's industrial ecological transformation using a rough set based machine learning methodabstractChina's industry urgently needs to be transformed from the development patterns driven by traditional production factor to achieving industrial ecological transformation (IET). The IET is influenced by diversified factors including resource input, allocation and flow, environmental regulations and technological innovations in different situation. Revealing the complex influence mechanisms between IET and its influence factors is necessary for effectively analyzing, evaluating and improving the performance of IET. A three stages machine learning method including learning, verification and generalization based on dominance-based rough set approach is presented to extract the influential relationships between the IET and its contextual influence factors. The proposed method excavates and learns the historical panel data of China's 30 provinces, and the cross-validation is conducted to produce a set of highly credible "If-Then" decision rules to generalize the synergistic influential relationships and intensities in IET. The results show that China's investment strength, resource allocation efficiency, command controlled and economic incentive environmental regulations are determinants to enhance the performance of IET, which helps to select the optimal transformation patterns by taking the historical development characteristics as lessons. Wenxin Mao, Huifang Sun |
SMC | 3 |
| 2020 | Bounded consensus control for stochastic multi-agent systems with additive noises
Zhongmei Wang, Huifang Sun, Huanshui Zhang, Xiyu Liu 0001 |
Neurocomputing | 2 |
| 2018 | Compact Analytical Description of Digital Radio-Frequency Pulse-Width Modulated SignalsabstractRadio frequency pulse-width modulation (RF-PWM) has been used as a power coding method in all-digital transmitters, which employ highly efficient switched-mode power amplifiers (SMPA). The main drawback of RF-PWM is the high level of in-band harmonic distortion when digitally implemented. In order to reduce spectral aliasing effects and produce acceptable levels of harmonic noise, ultra-fast clock speeds are required, making it commercially infeasible. In this paper, we derive a novel compact analytical model of a multilevel digital RF-PWM, driven by an arbitrary bounded baseband signal. We show that the spectral aliasing effects are equivalent to a particular amplitude quantization of the input baseband signal. This result implies that highly linear digital RF-PWM can be realized with modest clock speeds if and only if the input baseband signal is pre-quantized according to the inherent quantization process. We provide full description of this quantization process and describe its dependence on RF-PWM design parameters. Presented results enable a complete understanding of the nonlinear behavior of digitally implemented RF-PWM, and therefore can aid in optimal transceiver design. Numerical simulations in MATLAB were used to verify the derived analytical expressions. Omer Tanovic, Rui Ma 0022, Huifang Sun |
ISCAS | 3 |
| 2018 | Optimizing Multistage Discriminative Dictionaries for Blind Image Quality AssessmentabstractState-of-the-art algorithms for blind image quality assessment (BIQA) typically have two categories. The first category approaches extract natural scene statistics (NSS) as features based on the statistical regularity of natural images. The second category approaches extract features by feature encoding with respect to a learned codebook. However, several problems need to be addressed in existing codebook-based BIQA methods. First, the high-dimensional codebook-based features are memory-consuming and have the risk of over-fitting. Second, there is a semantic gap between the constructed codebook by unsupervised learning and image quality. To address these problems, we propose a novel codebook-based BIQA method by optimizing multistage discriminative dictionaries (MSDDs). To be specific, MSDDs are learned by performing the label consistent K-SVD (LC-KSVD) algorithm in a stage-by-stage manner. For each stage, a new quality consistency constraint called “quality-discriminative regularization” term is introduced and incorporated into the reconstruction error term to form a unified objective function, which can be effectively solved by LC-KSVD for discriminative dictionary learning. Then, the latter stage takes the reconstruction residual data in the former stage as input based on which LC-KSVD is repeatedly performed until the final stage is reached. Once the MSDDs are learned, multistage feature encoding is performed to extract feature codes. Finally, the feature codes are concatenated across all stages and aggregated over the entire image for quality prediction via regression. The proposed method has been evaluated on five databases and experimental results well confirm its superiority over existing relevant BIQA methods. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Ke Gu 0001, Gangyi Jiang, Huifang Sun |
IEEE Trans. Multim. | 6 |
| 2017 | uAVS2 - Fast encoder for the 2nd generation IEEE 1857 video coding standard
Zhenyu Wang 0002, Ronggang Wang, Kui Fan, Huifang Sun, Wen Gao 0001 |
Signal Process. Image Commun. | 4 |
| 2017 | Entropy of Primitive: From Sparse Representation to Visual Information EvaluationabstractIn this paper, we propose a novel concept in evaluating the visual information when perceiving natural images-the entropy of primitive (EoP). Sparse representation has been successfully applied in a wide variety of signal processing and analysis applications due to its high efficiency in dealing with rich varied and directional information contained in natural scenes. Inspired by this observation, in this paper, the visual signal can be decomposed into structural and nonstructural layers according to the visual importance of sparse primitives. Accordingly, the EoP is developed in measuring the visual information. It has been found that the EoP changing tendency in image sparse representation is highly relevant with the hierarchical perceptual cognitive process of human eyes. Extensive mathematical explanations as well as experimental verifications have been presented in order to support the hypothesis. The robustness of the EoP is evaluated in terms of varied block sizes. The dictionary universality is also studied by employing both universal and adaptive dictionaries. With the convergence characteristics of the EoP, a novel top-down just-noticeable difference (JND) profile is proposed. The simulation results have shown that the EoP-based JND outperforms the state-of-the-art JND models according to the subjective evaluation. Siwei Ma 0001, Xiang Zhang 0004, Shiqi Wang 0001, Jian Zhang 0018, Huifang Sun, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | A Joint Compression Scheme of Video Feature Descriptors and Visual ContentabstractHigh-efficiency compression of visual feature descriptors has recently emerged as an active topic due to the rapidly increasing demand in mobile visual retrieval over bandwidth-limited networks. However, transmitting only those feature descriptors may largely restrict its application scale due to the lack of necessary visual content. To facilitate the wide spread of feature descriptors, a hybrid framework of jointly compressing the feature descriptors and visual content is highly desirable. In this paper, such a content-plus-feature coding scheme is investigated, aiming to shape the next generation of video compression system toward visual retrieval, where the high-efficiency coding of both feature descriptors and visual content can be achieved by exploiting the interactions between each other. On the one hand, visual feature descriptors can achieve compact and efficient representation by taking advantages of the structure and motion information in the compressed video stream. To optimize the retrieval performance, a novel rate-accuracy optimization technique is proposed to accurately estimate the retrieval performance degradation in feature coding. On the other hand, the already compressed feature data can be utilized to further improve the video coding efficiency by applying feature matching-based affine motion compensation. Extensive simulations have shown that the proposed joint compression framework can offer significant bitrate reduction in representing both feature descriptors and video frames, while simultaneously maintaining the state-of-the-art visual retrieval performance. Xiang Zhang 0004, Siwei Ma 0001, Shiqi Wang 0001, Xinfeng Zhang 0001, Huifang Sun, Wen Gao 0001 |
IEEE Trans. Image Process. | 5 |
| 2016 | Keypoint trajectory coding on compact descriptor for video analysisabstractIn contrast to still image analysis, motion information offers a powerful means to analyze video. In particular, motion trajectories determined from keypoints have become very popular in recent years for a variety of video analysis tasks, including search, retrieval and classification. Additionally, cloud-based analysis of media content has been gaining momentum, so efficient communication of salient video information to perform the necessary analysis of video at the cloud server is needed. This paper describes a novel framework to efficiently represent the keypoint trajectories. In particular, an interframe prediction is designed with the option to operate in a low-delay mode. Additionally, a scalable coding method is proposed that allows for a subset of the coded trajectories in a video segment to be easily accessed. Experimental results on several popular datasets including Stanford MAR and Hopkin155 demonstrate a significant rate saving of up to 25% with our proposed trajectory coding approaches relative to a state-of-the-art reference approach. Dong Tian, Huifang Sun, Anthony Vetro |
ICIP | 2 |
| 2016 | Grey dominance-based rough set approach to decision system with three-parameter interval grey numberabstractA method of knowledge acquisition for the decision information system whose attribute value of alternatives is three-parameter interval grey number is proposed in this paper. First, in classic rough set, the decision table must be given in advance, but we can only establish information system from the collected data. So, we establish the decision table from information system with the grey relational clustering decision method. Then, we construct the grey dominance relation based on the dominance extent between two three-parameter interval grey numbers, and put forward a method of extracting decision rules and attribute reduction. The last case about the comprehensive evaluation of icebreaking car is given to illustrate the effectiveness of the proposed method. Dang Luo, Wenxin Mao, Huifang Sun |
SMC | 3 |
| 2014 | View Synthesis Prediction in the 3-D Video Coding Extensions of AVC and HEVCabstractAdvanced multiview video systems are able to generate intermediate viewpoints of a 3-D scene. To enable low-complexity free view generation, texture and its associated depth are used as input data for each viewpoint. To improve the coding efficiency of such content, view synthesis prediction (VSP) is proposed to further reduce interview redundancy in addition to traditional disparity compensated prediction. This paper describes and analyzes rate-distortion optimized VSP designs, which were adopted in the 3-D extensions of both Advanced Video Coding (AVC) and High Efficiency Video Coding (HEVC). In particular, we propose a novel backward-VSP scheme using a derived disparity vector, as well as efficient signalling methods in the context of AVC and HEVC. In addition, we put forward a novel depth-assisted motion vector prediction method to optimize the coding efficiency. A thorough analysis of coding performance is provided using different VSP schemes and configurations. Experimental results demonstrate average bit rate reductions of 2.5% and 1.2% in AVC and HEVC coding frameworks, respectively, with up to 23.1% bit rate reduction for dependent views. Feng Zou 0006, Dong Tian, Anthony Vetro, Huifang Sun, Oscar C. Au, Shinya Shimizu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | An Analytical Model for Synthesis Distortion Estimation in 3D VideoabstractWe propose an analytical model to estimate the synthesized view quality in 3D video. The model relates errors in the depth images to the synthesis quality, taking into account texture image characteristics, texture image quality, and the rendering process. Especially, we decompose the synthesis distortion into texture-error induced distortion and depth-error induced distortion. We analyze the depth-error induced distortion using an approach combining frequency and spatial domain techniques. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that our analytical model can accurately estimate the synthesis noise power. Thus, the model can be used to estimate the rendering quality for different system designs. Lu Fang 0001, Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun, Oscar C. Au |
IEEE Trans. Image Process. | 5 |
| 2014 | Image Interpolation via Graph-Based Bayesian Label PropagationabstractIn this paper, we propose a novel image interpolation algorithm via graph-based Bayesian label propagation. The basic idea is to first create a graph with known and unknown pixels as vertices and with edge weights encoding the similarity between vertices, then the problem of interpolation converts to how to effectively propagate the label information from known points to unknown ones. This process can be posed as a Bayesian inference, in which we try to combine the principles of local adaptation and global consistency to obtain accurate and robust estimation. Specially, our algorithm first constructs a set of local interpolation models, which predict the intensity labels of all image samples, and a loss term will be minimized to keep the predicted labels of the available low-resolution (LR) samples sufficiently close to the original ones. Then, all of the losses evaluated in local neighborhoods are accumulated together to measure the global consistency on all samples. Moreover, a graph-Laplacian-based manifold regularization term is incorporated to penalize the global smoothness of intensity labels, such smoothing can alleviate the insufficient training of the local models and make them more robust. Finally, we construct a unified objective function to combine together the global loss of the locally linear regression, square error of prediction bias on the available LR samples, and the manifold regularization term. It can be solved with a closed-form solution as a convex optimization problem. Experimental results demonstrate that the proposed method achieves competitive performance with the state-of-the-art image interpolation algorithms. Xianming Liu 0005, Debin Zhao, Jiantao Zhou 0001, Wen Gao 0001, Huifang Sun |
IEEE Trans. Image Process. | 5 |
| 2013 | Synthesis distortion estimation in 3D video using frequency and spatial analysisabstractWe propose an analytical model to estimate the synthesized view quality in 3D video. Specifically, we estimate the depth-error induced distortion using an approach that combines frequency and spatial domain analysis. We also propose to decompose the spatial-variant video signals into gradient-based representations to capture the interaction between image gradients, depth errors and synthesis distortion. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that our analytical model can accurately estimate the synthesis noise power. Lu Fang 0001, Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun, Lu Yu 0003 |
ICIP | 5 |
| 2013 | Reconfigurable media coding: An overview
Euee S. Jang, Marco Mattavelli, Marius Preda, Mickaël Raulet, Huifang Sun |
Signal Process. Image Commun. | 5 |
| 2013 | Special issue on MPEG CCF
Marco Mattavelli, Euee S. Jang, Marius Preda, Mickaël Raulet, Huifang Sun |
Signal Process. Image Commun. | 5 |
| 2012 | On modeling the rendering error in 3D videoabstractWe propose an analytical model to estimate the rendering quality in 3D video. The model relates errors in the depth images to the rendering quality, taking into account texture image characteristics, texture image quality, the camera configuration and the rendering process. Specifically, we derive position (disparity) errors from the depth errors, and the probability distribution of the position errors is used to calculate the power spectral density of the rendering errors. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that the model can accurately estimate the synthesis noise up to a constant offset. Thus, the model can be used to estimate the change in rendering quality for different system designs. Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun |
ICIP | 4 |
| 2012 | Predictive coding of intra prediction modes for high efficiency video codingabstractThe High Efficiency Video Coding (HEVC) standardization process currently underway includes many tools for the coding of intra pictures. HEVC allows for many more intra prediction modes or directions as compared to previous standards. Efficient coding of these modes is therefore important because the modes consume a non-negligible portion of the total bit-stream used for coding intra pictures. In this paper, a predictive coding method is proposed to reduce the number of bits needed for signaling the intra prediction modes, where the spatial angular correlation between the intra prediction mode of the current Prediction Unit (PU) and the neighboring PUs is computed using a few modulo-N arithmetic operations that do not impact encoder or decoder run-times. The proposed method provides similar or greater improvements in compression efficiency as compared to competing tools, while requiring no changes to the existing bit-stream syntax. Xiaozhong Xu, Robert A. Cohen, Anthony Vetro, Huifang Sun |
PCS | 4 |
| 2012 | Regulation of cell proliferation and apoptosis by growth hormone during zebrafish auditory hair cell regenerationabstractBackground In order to develop treatments or preventive measures for auditory hair cell loss, an understanding of both the process of auditory hair cell regeneration and factors that influence this process, is needed. Our previous microarray analysis showed that growth hormone (GH) was significantly upregulated during zebrafish auditory hair cell regeneration, coupled with cell proliferation [1,2]. We further tested the effects of GH on zebrafish auditory hair cell regeneration by injecting GH after sound exposure and found that GH can efficiently promote post-trauma auditory hair cell regeneration, which may be achieved through stimulating proliferation and suppressing apoptosis [3]. In the current study, we used Next Generation Sequencing (NGS) to examine the possible GH pathways involved in zebrafish auditory hair cell regeneration. Gopinath Rajadinakaran, Huifang Sun, Claire Rinehart, Eric C. Rouchka, Michael E. Smith 0001 |
BMC Bioinform. | 2 |
| 2011 | Transductive Regression with Local and Global Consistency for Image Super-ResolutionabstractIn this paper, we propose a novel image super-resolution algorithm, referred to as interpolation based on transductive regression with local and global consistency (TRLGC). Our algorithm first constructs a set of local interpolation models which can predict the intensity labels of all image samples, and a loss term will be minimized to keep the predicted labels of available low-resolution (LR) samples sufficiently close to the original ones. Then, all of the losses evaluated in local neighborhoods are accumulated together to measure the global consistency on all samples. Furthermore, a graph-Laplacian based manifold regularization term is incorporated to penalize the global smoothness of intensity labels, such smoothing can alleviate the insufficient training of the local models and make them more robust. Finally, we construct a unified objective function to combine together the accumulated loss of the locally linear regression, square error of prediction bias on the available LR samples and the manifold regularization term, which could be solved with a closed-form solution as a convex optimization problem. In this way, a transductive regression algorithm with local and global consistency is developed. Experimental results on benchmark test images demonstrate that the proposed image super-resolution method achieves very competitive performance with the state-of-the-art algorithms. Xianming Liu 0005, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001, Huifang Sun |
DCC | 6 |
| 2011 | Image Interpolation Via Regularized Local Linear RegressionabstractThe linear regression model is a very attractive tool to design effective image interpolation schemes. Some regression-based image interpolation algorithms have been proposed in the literature, in which the objective functions are optimized by ordinary least squares (OLS). However, it is shown that interpolation with OLS may have some undesirable properties from a robustness point of view: even small amounts of outliers can dramatically affect the estimates. To address these issues, in this paper we propose a novel image interpolation algorithm based on regularized local linear regression (RLLR). Starting with the linear regression model where we replace the OLS error norm with the moving least squares (MLS) error norm leads to a robust estimator of local image structure. To keep the solution stable and avoid overfitting, we incorporate the l(2)-norm as the estimator complexity penalty. Moreover, motivated by recent progress on manifold-based semi-supervised learning, we explicitly consider the intrinsic manifold structure by making use of both measured and unmeasured data points. Specifically, our framework incorporates the geometric structure of the marginal probability distribution induced by unmeasured samples as an additional local smoothness preserving constraint. The optimal model parameters can be obtained with a closed-form solution by solving a convex optimization problem. Experimental results on benchmark test images demonstrate that the proposed method achieves very competitive performance with the state-of-the-art interpolation algorithms, especially in image edge structure preservation. Xianming Liu 0005, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001, Huifang Sun |
IEEE Trans. Image Process. | 6 |
| 2010 | Direction-adaptive transforms for coding prediction residualsabstractIn this paper, we present 2-D direction-adaptive transforms for coding prediction residuals of video. These Direction-Adaptive Residual Transforms (DART) are shown to be more effective than the traditional 2-D DCT when coding residual blocks that contain directional features. After presenting the directional transform structures and improvements to their efficiency, we outline how they are used to code both Inter and Intra prediction residuals. For Intra coding, we also demonstrate the relation between the prediction mode and the optimal DART orientation. Experimental results exhibit up to 7% and 9.3% improvements in compression efficiency in JM 16.0 and JM-KTA 2.6r1 respectively, as compared to using only the conventional H.264/AVC transform. Robert A. Cohen, Sven Klomp, Anthony Vetro, Huifang Sun |
ICIP | 4 |
| 2010 | Growth hormone induces proliferation in the zebrafish inner earabstractFigure 1 The effect of growth hormone injection on mean (±SE) number of BrdU-labeled cells (A and B) and hair cell bundle density (C) in zebrafish ear sensory tissues.Fish in (A) were not exposed to a sound stimulus, while fish in (B and C) were dissected 48 and 60 h post-sound exposure, respectively.N=6-12.* P<0.05. Michael E. Smith 0001, Huifang Sun, Julie B. Schuck, Shunsuke Moriyama |
BMC Bioinform. | 2 |
| 2010 | Deinterlacing Using Hierarchical Motion AnalysisabstractA motion-compensated deinterlacing scheme based on hierarchical motion analysis is presented. According to deinterlacing steps, our contribution can be divided into four parts: motion estimation, motion state analysis, motion consistency analysis, and finer-grained interpolation. In motion estimation, we introduce a Gaussian noise model for choosing the best motion vector for each block, and make a tradeoff between utilizing previous de-interlaced frames and avoiding error propagation. A directional interpolation method is also introduced in this part for backward fields. In motion state analysis, we define two motion states for each pixel, thus achieve a compromise between traditional block-based strategies and the extreme pixel-based case. In motion consistency analysis, we propose to measure both the motion vector consistency and the motion state consistency in order to determine whether the previous two parts should be performed again with a different block size. In finer-grained interpolation, we utilize a combination of recursive median filters to generate the final results. Experimental results show that all of the proposed techniques are effective, either objectively or subjectively. As a result, we can achieve much higher image quality, with an average gain of about 1.83 dB in terms of peak signal-to-noise ratio. Moreover, the increased computation complexity is marginal. Qian Huang 0008, Debin Zhao, Siwei Ma 0001, Wen Gao 0001, Huifang Sun |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2010 | Compressed Domain Video Object SegmentationabstractWe present a compressed domain video object segmentation method for the MPEG encoded video sequences. For a fraction of the raw domain analysis, compressed domain segmentation provides the essentiala prioriinformation to many vision tasks from surveillance to transcoding that require fast processing of large volumes of data where pixel-resolution boundary extraction is not required. Our method generates accurate segmentation maps in block resolution at hierarchically varying object levels, which empowers application to determine the most pertinent partition of images. It exploits the block structure of the compressed video to minimize the amount of data to be processed. All the available motion flow within a group of pictures is projected onto a single layer, which also consists of the frequency decomposition of color pattern. Then, by starting from the blocks where the spatial energy is small, it expands homogeneous regions while automatically adapting local similarity criteria. We also formulate an alternative solution that applies a kernel-based clustering where separate spatial, transform, and motion kernels are used to establish the affinity. We show that both region expansion and mean shift produce similar results as the computationally expensive raw domain segmentation. Finally, a binary clustering iteratively merges the most similar regions to generate a hierarchical partition tree. Fatih Porikli, Faisal I. Bashir, Huifang Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | New Standardized Extensions of MPEG4-AVC/H.264 for Professional-Quality Video ApplicationsabstractTo support high quality video applications, the Joint Video Team (JVT) has recently added five new profiles, two new supplemental enhancement information (SEI) messages, and two new extended gamut color space indicators to the MPEG4-AVC/H.264 video coding standard. The new profiles include substantial feature enhancements for high-quality video applications, including improved-efficiency 4:4:4 video format coding, improved-efficiency lossless macroblock coding, coding 4:4:4 video pictures using three separately-coded color planes, and support of bit depths up to 14 bits per sample. The new features were developed to support a wide range of applications where high quality video compression is demanded, including professional and semi-professional scenarios in particular. They also anticipate the introduction of higher fidelity displays. In this paper, the new extensions are presented along with quantitaive estimates of the benefits of the new features and a discussion of the target application environments. Gary J. Sullivan, Haoping Yu, Shun-ichi Sekiguchi, Huifang Sun, Thomas Wedi, Steffen Wittmann, Yung Lyul Lee, C. Andrew Segall, Teruhiko Suzuki |
ICIP (1) | 4 |
| 2007 | Motion Mapping for MPEG-2 to H.264/AVC TranscodingabstractThis paper describes novel motion mapping algorithms aimed for low-complexity MPEG-2 to AVC transcoding. The proposed algorithms efficiently map incoming MPEG-2 motion vectors to outgoing AVC motion vectors regardless of the block sizes that the motion vectors correspond to. Extensive simulation results show that our proposed transcoder incorporating the proposed algorithms achieves very good rate-distortion performance with low complexity. Compared with the cascaded decoder-encoder solution, the proposed approach could achieve similar coding efficiency while significantly reduce the complexity. Jun Xin, Anthony Vetro, Huifang Sun, Shun-ichi Sekiguchi |
ISCAS | 4 |
| 2007 | An overview of scalable video streamingabstractAbstract Significant progress in digital video processing and communication technology has enabled the dream of high‐quality and real‐time video streaming over various networks. Compression technologies have bridged the gap between huge amount of visual data required for video streaming and limited bandwidth of communication channels. Communication networks, including wired and wireless networks, provide the platform for video streaming applications. However, video streaming is still a challenging problem due to limited network bandwidth, the presence of channel errors, and the variability in consumer terminals. In this article, we present an overview of scalable video streaming techniques that have been developed in recent years. These techniques include scalable video coding (SVC), video transcoding, and scalable video streaming methods. Copyright © 2007 John Wiley & Sons, Ltd. Huifang Sun, Anthony Vetro, Jun Xin |
Wirel. Commun. Mob. Comput. | 1 |
| 2006 | Extensions of H.264/AVC for Multiview Video CompressionabstractWe consider multiview video compression: the problem of jointly compressing multiple views of a scene recorded by different cameras. To take advantage of the correlation between views, we propose using disparity compensated view prediction and view synthesis and describe how these features can be implemented by extending the H.264/AVC compression standard. Finally, we discuss experimental results on the test sequences from the MPEG call for proposals on multiview video. Emin Martinian, Alexander Behrens, Jun Xin, Anthony Vetro, Huifang Sun |
ICIP | 5 |
| 2006 | Preface
Feng Wu 0001, Huifang Sun |
J. Comput. Sci. Technol. | 2 |
| 2006 | Constant quality rate allocation for FGS coding using composite R-D analysisabstractIn this correspondence, we propose a constant quality rate allocation algorithm for fine granularity scalability (FGS) coded videos. The rate allocation problem is formulated as a constrained minimization of quality fluctuation. The minimization is solved using a composite rate distortion (R-D) analysis. For a set of video frames, a composite R-D curve is first computed and then used for computing the optimal rate allocation. This algorithm is efficient because it is neither iterative nor recursive. After the composite R-D curve is computed, it can be used for optimal rate allocation of any rate budget. Moreover, the composite R-D curve can be updated efficiently over sliding windows. Experimental results have shown both the effectiveness and the efficiency of the proposed algorithm. Xi Min Zhang, Yun Q. Shi 0001, Anthony Vetro, Huifang Sun |
IEEE Trans. Multim. | 5 |
| 2005 | Fast adaptive fuzzy post-filtering for coding artifacts removal in interlaced videoabstractThe paper presents a new method using fuzzy filtering to remove the coding artifacts in compressed video. The method takes the interlaced video format into consideration and processes each field separately. For deblocking, a 1D fuzzy filter with different window sizes is used to remove the horizontal and vertical blocking artifacts respectively. For deringing, each 8/spl times/8 block in a field is first classified into one of the four categories, i.e., strong edge, weak edge, texture and smooth blocks. According to each block's type and the neighboring block's type, the spread parameter of a 2D fuzzy filter is adaptively decided and the filter is applied To speed up the process, the fuzzy filter weights are generated using a piecewise linear membership function instead of the conventional Gaussian function. Experimental results show that the proposed method has better detail preservation and lower computational costs than our previous method. It achieves comparable deblocking and superior deringing performance to the MPEG-4 standard method at similar computation costs. Yao Nie, Hao-Song Kong, Anthony Vetro, Huifang Sun, Kenneth E. Barner |
ICASSP (2) | 4 |
| 2005 | Layered dynamic mixture model for pattern discovery in asynchronous multi-modal streams [video applications]abstractWe propose a layered dynamic mixture model for asynchronous multi-modal fusion for unsupervised pattern discovery in video. The lower layer of the model uses generative temporal structures such as a hierarchical hidden Markov model to convert the audiovisual streams into mid-level labels, it also models the correlations in text with probabilistic latent semantic analysis. The upper layer fuses the statistical evidence across diverse modalities with a flexible meta-mixture model that assumes loose temporal correspondence. Evaluation on a large news database shows that multi-modal clusters have better correspondence to news topics than audio-visual clusters alone; novel analysis techniques suggest that meaningful clusters occur when the prediction of salient features by the model concurs with those shown in the story clusters. Lexing Xie, Lyndon S. Kennedy, Shih-Fu Chang, Ajay Divakaran, Huifang Sun, Ching-Yung Lin |
ICASSP (2) | 5 |
| 2004 | Adaptive fuzzy post-filtering for highly compressed video
Hao-Song Kong, Yao Nie, Anthony Vetro, Huifang Sun, Kenneth E. Barner |
ICIP | 4 |
| 2004 | Discovering meaningful multimedia patterns with audio-visual concepts and associated textabstractThe work presents the first effort to automatically annotate the semantic meanings of temporal video patterns obtained through unsupervised discovery processes. This problem is interesting in domains where neither perceptual patterns nor semantic concepts have simple structures. The patterns in video are modeled with hierarchical hidden Markov models (HHMM), with efficient algorithms to learn the parameters, the model complexity and the relevant features; the meanings are contained in words of the speech transcript of the video. The pattern-word association is obtained via cooccurrence analysis and statistical machine translation models. Promising results are obtained through extensive experiments on 20+ hours of TRECVID news videos: video patterns that associate with distinct topics such as el-nino and politics are identified; the HHMM temporal structure model compares favorably to a nontemporal clustering algorithm. Lexing Xie, Lyndon S. Kennedy, Shih-Fu Chang, Ajay Divakaran, Huifang Sun, Ching-Yung Lin |
ICIP | 5 |
| 2004 | Error resilience video coding in H.264 encoder with potential distortion trackingabstractIn this paper, an efficient rate-distortion (RD) model for an H.264 video encoder in a packet loss environment is presented. The encoder keeps tracking the potential error propagation on a block basis by taking into account the source characteristics, network conditions as well as the error concealment method. The end-to-end distortion invoked in this RD model is estimated according to the potential error-propagated distortion stored in a distortion map. The distortion map, in terms of each frame, is derived after the frame is encoded, which can be used for the RD-based encoding of the subsequent frames. Since the channel distortion has been considered in the proposed RD model, the new Lagarangian parameter is derived accordingly. The proposed method outperforms the error robust rate-distortion optimization method in the H.264 test model better in terms of both transmission efficiency and computational complexity. Yuan Zhang 0014, Wen Gao 0001, Huifang Sun, Qingming Huang, Yan Lu 0001 |
ICIP | 3 |
| 2004 | Coding artifacts reduction using edge map guided adaptive and fuzzy filteringabstractThe work presents a new adaptive approach for blocking and ringing artifact reduction. In order to avoid smearing of the image details, the proposed method first performs visual artifact detection and then applies adaptive filtering to the corrupted blocks. Both visual artifact detection and filtering are guided by an edge map which is constructed based on local features. A fuzzy identity filter is used for image de-ringing. Since it possesses a good edge preserving property and the filtering operation is applied to the edge blocks only (smooth and textured blocks are unaltered), the proposed method shows great effectiveness of both artifact reduction and detail preservation. Experiments demonstrate better results compared with other methods. Hao-Song Kong, Yao Nie, Anthony Vetro, Huifang Sun, Kenneth E. Barner |
ICME | 4 |
| 2004 | Structure analysis of soccer video with domain knowledge and hidden Markov models
Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun |
Pattern Recognit. Lett. | 5 |
| 2003 | Combined rate control and mode decision optimization for MPEG-2 transcoding with spatial resolution reductionabstractThis paper presents a new algorithm for MPEG-2 transcoding with spatial resolution reduction. The proposed method combines rate control and mode decision to achieve optimal transcoding performance using Lagrange multiplier algorithm. Since the proposed method incorporates motion vector mapping and mode decision into a single Lagrange multiplier formula, by minimizing the Lagrangian cost function, the optimal solution for mode decision can be obtained. The proposed transcoding scheme has demonstrated better subjective and objective results compared with other methods. Hao-Song Kong, Anthony Vetro, Huifang Sun |
ICIP (1) | 3 |
| 2003 | Feature selection for unsupervised discovery of statistical temporal structures in videoabstractIn this paper, we present algorithms for automatic feature selection for of structure discovery from video sequences. Feature selection in this scenario is hard because of the absence of class labels to evaluate against, and the temporal correlation among samples that prevents the direct estimation of posterior probabilities of the cluster given the sequence. The overall problem of structure discovery is formulated as simultaneously finding the statistical descriptions of structure and locating segments that matches the descriptions. Under Markov assumptions among events, structures in the video are modelled with hierarchical hidden Markov models, with efficient algorithms to jointly learn the model parameters and the optimal model complexity. Feature selection iterates between a wrapper step that partitions the large feature pool into consistent subsets, and a filter step that eliminate redundancy within these subsets, respectively. The feature subsets are then ranked according to the normalized Bayesian Information criteria, and the learning results from these ranked subsets can be evaluated and interpreted by a human observer. Results on soccer and baseball videos show that the automatically selected feature set coincides with those selected with domain knowledge and intuition, while achieving a correspondence comparable to that of supervised learning against manually labelled ground truth. Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun |
ICIP (1) | 4 |
| 2003 | Rate allocation for FGS coded video using composite R-D analysisabstractIn this paper, we propose a constant quality rate allocation algorithm for MPEG-4 FGS (fine granularity scalability) coded video sequences. The rate allocation problem is formulated as a constrained minimization of quality fluctuation. The minimization is solved using a novel composite rate distortion analysis. For a set of video frames, a composite rate distortion curve is first computed and then used for computing the optimal rate allocation. The proposed algorithm is very efficient because it is neither iterative nor recursive. In addition, after the composite rate distortion curve is computed, it can be used to calculate optimal rate allocation for any rate budget. Therefore, it is suitable for FGS coded bitstreams, which need to be transmitted and decoded many times at many different rates. Moreover, the composite rate distortion curve can be updated efficiently over sliding windows. This further reduces the computational complexity. Experiments using both synthetic and real FGS coded videos have shown the effectiveness and the efficiency of the proposed algorithm. Xi Min Zhang, Yun Q. Shi 0001, Anthony Vetro, Huifang Sun |
ICME | 5 |
| 2003 | Object-based coding for long-term archive of surveillance videoabstractThis paper describes video coding and segmentation techniques that can be used to achieve significant increase in storage capacity. Specifically, we examine the possibility to use object- based coding for efficient long-term archiving of surveillance video. We consider surveillance systems with many camera sources in which we are required to store several months of video data for each source, thus storage capacity is a major concern. The paper considers several automatic segmentation algorithms. With each algorithm, we analyze the shape coding overhead and implication on overall storage requirements, as well as the effect each algorithm has on the reconstructed quality of frames. Additionally, this paper reviews techniques to dynamically control the temporal rate of objects in the scene and perform bit allocation. Experimental results show that up to 90% savings in storage can be achieved with the proposed method compared to frame-based video coding techniques. The cost for this savings is that the accuracy of the background is compromised; however, we feel that this is satisfactory for the application under consideration. Anthony Vetro, Tetsuji Haga, Kazuhiko Sumi, Huifang Sun |
ICME | 4 |
| 2003 | Unsupervised discovery of multilevel statistical video structures using hierarchical hidden Markov modelsabstractStructure elements in a time sequence (e.g. video) are repetitive segments with consistent deterministic or stochastic characteristics. While most existing work in detecting structures follows a supervised paradigm, we propose a fully unsupervised statistical solution in this paper. We present a unified approach to structure discovery from long video sequences as simultaneously finding the statistical descriptions of structure and locating segments that matches the descriptions. We model the multilevel statistical structure as hierarchical hidden Markov models, and present efficient algorithms for learning both the parameters and the model structure. When tested on a specific domain, soccer video, the unsupervised learning scheme achieves very promising results: it automatically discovers the statistical descriptions of high-level structures, and at the same time achieves even slightly better accuracy in detecting discovered structures in unlabelled videos than a supervised approach designed with domain knowledge and trained with comparable hidden Markov models. Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun |
ICME | 4 |
| 2003 | Coding mode optimization for MPEG-2 transcoding with spatial resolution reduction
Hao-Song Kong, Anthony Vetro, Huifang Sun |
VCIP | 3 |
| 2003 | Survey of compressed-domain features used in audio-visual indexing and analysis
Hualu Wang, Ajay Divakaran, Anthony Vetro, Shih-Fu Chang, Huifang Sun |
J. Vis. Commun. Image Represent. | 5 |
| 2003 | Constant quality constrained rate allocation for FGS-coded videoabstractThis paper proposes an optimal rate-allocation scheme for fine-granular scalability (FGS) coded bitstreams that can achieve constant quality reconstruction of frames under a dynamic rate budget constraint. In doing so, we also aim to minimize the overall distortion at the same time. To achieve this, we propose a novel rate-distortion (R-D) labeling scheme to characterize the R-D relationship of the source coding process. Specifically, sets of R-D points are extracted during the encoding process and linear interpolation is used to estimate the actual R-D curve of the enhancement-layer signal. The extracted R-D information is then used by an enhancement-layer transcoder to determine the bits that should be allocated per frame. A sliding-window-based rate-allocation method is proposed to realize constant quality among frames. This scheme is first considered for a single FGS-coded source, then extended to operate on multiple sources. With the proposed scheme, the rate allocation can be performed in a single pass; hence, the complexity is quite low. Experimental results confirm the effectiveness of the proposed scheme under static and dynamic bandwidth conditions. Xi Min Zhang, Anthony Vetro, Yun Q. Shi 0001, Huifang Sun |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2003 | Rate-distortion modeling for multiscale binary shape coding based on Markov random fieldsabstractThe purpose of this paper it to explore the relationship between the rate-distortion characteristics of multiscale binary shape and Markov random field (MRF) parameters. For coding, it is important that the input parameters that will be used to define this relationship be able to distinguish between the same shape at different scales, as well as different shapes at the same scale. We consider an MRF model, referred to as the Chien model, which accounts for high-order spatial interactions among pixels. We propose to use the statistical moments of the Chien model as input to a neural network to accurately predict the rate and distortion of the binary shape when coded at various scales. Anthony Vetro, Yao Wang 0001, Huifang Sun |
IEEE Trans. Image Process. | 3 |
| 2002 | A probabilistic approach for rate-distortion modeling of multiscale binary shapeabstractThe purpose of this paper it to explore the relationship between the rate-distortion (R-D) characteristics of multi scale binary shape and Markov Random Field (MRF) parameters. In our experiments, we consider two prior models. The first MRF model takes into account pair-wise interaction between pels, and for the binary case, is typically referred to as the auto-logistic model; the second MRF model accounts for higher order spatial interactions and is referred to as the Chien model. Experimental results indicate that the auto-logistic model is not sufficient to characterize the R-D characteristics of multi scale binary shape data. However, higher order models, such as the Chien model, do seem feasible. We propose to use the statistical moments of the Chien model as input to a neural network to accurately predict the rate and distortion of the binary shape when coded at various scales. Anthony Vetro, Yao Wang 0001, Huifang Sun |
ICASSP | 3 |
| 2002 | Structure analysis of soccer video with hidden Markov modelsabstractIn this paper, we present algorithms for parsing the structure of produced soccer programs. The problem is important in the context of a personalized video streaming and browsing system. While prior work focuses on the detection of special events such as goals or corner kicks, this paper is concerned with generic structural elements of the game. We begin by defining two mutually exclusive states of the game, play and break based on the rules of soccer. We select a domain-tuned feature set, dominant color ratio and motion intensity, based on the special syntax and content characteristics of soccer videos. Each state of the game has a stochastic structure that is modeled with a set of hidden Markov models. Finally, standard dynamic programming techniques are used to obtain the maximum likelihood segmentation of the game into the two states. The system works well, with 83.5% classification accuracy and good boundary timing from extensive tests over diverse data sets. Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Huifang Sun |
ICASSP | 4 |
| 2002 | Drift compensation architectures and techniques for reduced-resolution transcoding
Peng Yin 0002, Anthony Vetro, Huifang Sun, Bede Liu |
VCIP | 3 |
| 2002 | Constant-quality constrained-rate allocation for FGS video coded bitstreams
Xi Min Zhang, Anthony Vetro, Yun Q. Shi 0001, Huifang Sun |
VCIP | 4 |
| 2002 | Drift compensation for reduced spatial resolution transcodingabstractThis paper discusses the problem of reduced-resolution transcoding of compressed video bitstreams. An analysis of drift errors is provided to identify the sources of quality degradation when transcoding to a lower spatial resolution. Two types of drift error are considered: a reference picture error, which has been identified in previous works, and error due to the noncommutative property of motion compensation and down-sampling, which is unique to this work. To overcome these sources of error, four novel architectures are presented. One architecture attempts to compensate for the reference picture error in the reduced resolution, while another architecture attempts to do the same in the original resolution. We present a third architecture that attempts to eliminate the second type of drift error and a final architecture that relies on an intrablock refresh method to compensate for all types of errors. In all of these architectures, a variety of macroblock level conversions are required, such as motion vector mapping and texture down-sampling. These conversions are discussed in detail. Another important issue for the transcoder is rate control. This is especially important for the intra-refresh architecture since it must find a balance between number of intrablocks used to compensate for errors and the associated rate-distortion characteristics of the low-resolution signal. The complexity and quality of the architectures are compared. Based on the results, we find that the intra-refresh architecture offers the best tradeoff between quality and complexity and is also the most flexible. Peng Yin 0002, Anthony Vetro, Bede Liu, Huifang Sun |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2001 | Constant pace skimming and temporal sub-sampling of video using motion activityabstractWe describe a "constant pace" framework for video summarization via fast playback or temporal sub-sampling. The pace of the summary serves as a parameter that enables production of a video summary of any desired length. Earlier, we showed that the intensity of motion activity (or pace) of a video sequence is a good indication of its "summarizability". Here we build on this notion by adapting the playback frame-rate or the temporal subsampling rate to the pace. Either the less active parts of the sequence are played back at a faster frame rate or the less active parts of the sequence are sub-sampled more heavily than are the more active parts, so as to produce a summary with constant pace. The basic idea is to skip over the less interesting parts of the video. Our technique is computationally simple and gives satisfactory results with surveillance and entertainment video. Ajay Divakaran, Kadir A. Peker, Huifang Sun |
ICIP (3) | 3 |
| 2001 | Rate-distortion optimized video coding considering frameskipabstractThe general problem of optimized video encoding has received a great deal of attention in recent years. This paper focuses on the optimization of video coding with frameskip. We propose models that estimate the distortion for coded frames as well as non-coded frames. Using these models in conjunction with well-know models that estimate the rate allows us to formulate a rate control problem that trades-off spatial and temporal quality. Simulation results indicate moderate improvements for low motion test sequences. Anthony Vetro, Huifang Sun, Yao Wang 0001 |
ICIP (3) | 2 |
| 2001 | Estimating Distortion Of Coded And Non-Coded Frames For Frameskip-Optimized Video CodingabstractThis paper focuses on the problem of estimating the distortion for coded and non-coded frames in a video coder that employs variable frameskip. The distortion for coded frames is given by classic rate-distortion models, however the distortion for non-coded frames has not been considered. Based on the optical flow equation, we formulate a method for estimating the distortion of the non-coded frames. Anthony Vetro, Yao Wang 0001, Huifang Sun |
ICME | 3 |
| 2001 | Algorithms And System For Segmentation And Structure Analysis In Soccer VideoabstractIn this paper, we present a novel system and effective algorithms for soccer video segmentation. The output, about whether the ball is in play, reveals high-level structure of the content. The first step is to classify each sample frame into 3 kinds of view using a unique domain-specific feature, grass-area-ratio. Here the grass value and classification rules are learned and automatically adjusted to each new clip. Then heuristic rules are used in processing the view label sequence, and obtain play/break status of the game. The results provide good basis for detailed content analysis in next step. We also show that lowlevel features and mid-level view classes can be combined to extract more information about the game, via the example of detecting grass orientation in the field. The results are evaluated under different metrics intended for different applications; the best result in segmentation is 86.5%. 1. Lexing Xie, Shih-Fu Chang, Ajay Divakaran, Anthony Vetro, Huifang Sun |
ICME | 6 |
| 2001 | Object-based transcoding for adaptable video content deliveryabstractThis paper introduces a new framework for video content delivery that is based on the transcoding of multiple video objects. Generally speaking, transcoding can be defined as the manipulation or conversion of data into another more desirable format. We consider manipulations of object-based video content, and more specifically, from one set of bit streams to another. Given the object-based framework, we present a set of new algorithms that are responsible for manipulating the original set of video bit streams. Depending on the particular strategy that is adopted, the transcoder attempts to satisfy network conditions or user requirements in various ways. One of the main contributions of this paper is to discuss the degrees of freedom within an object-based transcoder and demonstrate the flexibility that it has in adapting the content. Two approaches are considered: a dynamic programming approach and an approach that is based on available meta-data. Simulations with these two approaches provide insight regarding the bit allocation among objects and illustrates the tradeoffs that can be made in adapting the content. When certain meta-data about the content is available, we show that bit allocation can be significantly improved, key objects can be identified, and varying the temporal resolution of objects can be considered. Anthony Vetro, Huifang Sun, Yao Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | A Region Based Descriptor for Spatial Distribution of Motion Activity for Compressed VideoabstractIn this paper we present a new descriptor for the spatial distribution of motion activity in video sequences. We construct a histogram of areas of distinct regions (or "blobs") of "motion active" regions over the entire video shot. We carry out another thresholding process on the histogram to get our descriptor, which is a histogram normalized with respect to the average size of the blobs, and thus normalized with respect to frame size. We get similar precision-recall performance to the spatial activity descriptor in the current MPEG-7 experimental model. We are also able to successfully capture the effects of camera motion as well as the effects of non-camera motion in distinct uncorrelated parts of our descriptor. Since the feature extraction is in the compressed domain and simple, it is extremely fast. We find that our descriptor enables fast and accurate indexing of video. Ajay Divakaran, Kadir A. Peker, Huifang Sun |
ICIP | 3 |
| 2000 | An overview of the encoding tools in the MPEG-4 reference softwareabstractIn this paper, we summarize several encoding tools that were introduced by the video group of the MPEG-4 committee. The review the key features of these algorithms. The original proposers in the MPEG-4 committee have implemented these algorithms in the reference software. We used the most recent reference software to generate a sample set of results. A profiling was performed on a Pentium 233 MHz PC to understand the complexity aspect of an MPEG-4 encoder. Tihao Chiang, Hung-Ju Lee, Huifang Sun |
ISCAS | 3 |
| 2000 | Object-based transcoding for scalable quality of serviceabstractIn this paper, we focus on the methods for delivering object-based video data. More specifically, me exploit the fact that a finer level of scalability can be achieved when the video frame has been decomposed into objects and coded using MPEG-4. A new framework is proposed. Anthony Vetro, Huifang Sun, Yao Wang 0001 |
ISCAS | 2 |
| 1999 | Rate-Distortion Modeling of Binary Shape Using State PartitioningabstractIn this paper, the rate-distortion (R-D) characteristics of binary shapes are modeled. Specifically we are interested in predicting the rate and distortion that is produced by the shape coding techniques that have been adopted into the MPEG-4 standard. The shape coding algorithm is a context-based arithmetic encoder and operates on a per block basis. Currently, there is no efficient way of estimating the rate and distortion at various levels of resolution. Consequently, we propose a model that is based on a set of parameters that can be easily extracted from the binary blocks. The parameters represent states that arise from the possible binary patterns that can occur in a small neighborhood around the current pixel. Symmetry is exploited to keep the number of states to a minimum. It is shown that the proposed model is computationally efficient and provides accurate estimates of the R-D characteristics of a binary shape. Anthony Vetro, Huifang Sun, Yao Wang 0001, Onur G. Guleryuz |
ICIP (2) | 2 |
| 1999 | MPEG-4 rate control for multiple video objectsabstractThis paper describes an algorithm which can achieve a constant bit rate when coding multiple video objects. The implementation is a nontrivial extension of the MPEG-4 rate control algorithm for single video objects which employs a quadratic rate quantizer model. The algorithm is organized into two stages: a pre- and a post-encoding stage. In the pre-encoding stage, an initial target estimate is made for each object. Based on the buffer fullness, the total target is adjusted and then distributed proportional to the relative size, motion, and variance of each object. Based on the new individual targets and rate-quantizer relation for texture, appropriate quantization parameters are calculated. After each object is encoded, the model parameters for each object are updated, and if necessary, frames are skipped to ensure that the buffer does not overflow. A preframeskip control is exercised to avoid buffer overflow when the motion and shape information occupies a significant portion of the bit budget. The rate control algorithm switches between two operation modes so that the coder can reduce the spatial coding accuracy for an improved temporal resolution. A shape-coding control mechanism is also proposed, which provides a tradeoff between texture and shape coding accuracy. Overall, the algorithm is able to successfully achieve the target bit rate, effectively code arbitrarily shaped objects, and maintain a stable buffer level. These techniques have been adopted by the MPEG committee in July 1997 as part of the video verification model (VM8). Anthony Vetro, Huifang Sun, Yao Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1998 | Frame-rate up-conversion using transmitted true motion vectorsabstractIn this paper, we present a video frame-rate up-conversion scheme that uses transmitted true motion vectors for motion-compensated interpolation. In a past work, we demonstrated that a neighborhood-relaxation motion tracker can provide more accurate true motion information than a conventional minimal-residue block-matching algorithm. Although the technique to estimate the true motion vectors is a novelty in its own right, the strength of this technique can be further demonstrated through various spatio-temporal interpolation applications. In this work, we focus on the particular problem of frame-rate up-conversion. In the proposed scheme, the true motion field is derived by the encoder and transmitted by normal means (e.g., MPEG or H.263 encoding). Then, it is recovered by the decoder and is used not only for motion compensated predictions but also used to reconstruct missing data. It is shown that the use of our neighborhood-relaxation motion estimation provides a method of constructing high quality image sequences in a practical manner. Yen-Kuang Chen, Anthony Vetro, Huifang Sun, Sun-Yuan Kung |
MMSP | 3 |
| 1997 | Frequency Domain Down Conversion of HDTV Using Adaptive Motion CompensationabstractThe paper investigates the possible means of down converting an incoming HDTV signal into one which is suitable for display onto an NTSC monitor. We present a novel frequency synthesis algorithm which constructs a set of 8/spl times/8 DCT coefficients from four original 8/spl times/8 DCT blocks. Also, a down-conversion decoder employing a unique adaptive motion compensation algorithm is presented. Using this fixed motion compensation algorithm, the proposed down-conversion technique is compared to the conventional method. Both yield acceptable image sequences with different trade-offs in complexity and quality. Anthony Vetro, Huifang Sun, Jay Bao, Tommy Poon |
ICIP (1) | 2 |
| 1997 | Joint rate control for coding multiple video objectsabstractThis paper describes an algorithm which achieves a constant bit rate when coding multiple video objects. This implementation is based on the current MPEG-4 video verification model. Each object maintains its own set of parameters. With these parameters an initial target estimate is made for each object. Based on the buffer fullness, the total target is adjusted and then distributed proportional to the size, motion, and variance of each object. Appropriate quantization parameters can be calculated for each video object based on the new individual targets and second order model parameters. This algorithm assures that the target bit rate is achieved for low latency video coding. Anthony Vetro, Huifang Sun |
MMSP | 2 |
| 1997 | Error concealment algorithms for robust decoding of MPEG compressed video
Huifang Sun, Joel W. Zdepski, Wilson Kwok, Dipankar Raychaudhuri |
Signal Process. Image Commun. | 1 |
| 1997 | MPEG coding performance improvement by jointly optimizing coding mode decisions and rate controlabstractThis paper presents a new algorithm for determining the optimal MPEG coding strategy in terms of the selection of macroblock coding modes and quantizer scales. In the algorithm proposed in the test model the rate control operates independently from the coding mode selection for each macroblock. The coding mode is decided based only upon the energy of predictive residues. Actually, the two processes of coding mode decision and rate control are intimately related to each other and should be determined jointly in order to achieve optimal coding performance. We formulate the constrained optimization problem and present solutions based upon rate-distortion characteristics, or R(D) curves, for all the macroblocks that compose the picture being coded. Distortion for the entire picture is assumed to be decomposable and expressible as a function of individual macroblock distortions, with this being the objective function to minimize. The determination of the optimal solution is complicated by the MPEG differential encoding of motion vectors and DC coefficients, which introduce dependencies that carry over from macroblock to macroblock for a duration equal to the slice length. As an approximation, a near optimum greedy algorithm is proposed. Once the upper bound in performance is calculated, it can be used to assess how well practical suboptimum methods perform. Finally, such a practical suboptimum algorithm is proposed and evaluated. Huifang Sun, Wilson Kwok, Max Chien, C. H. John Ju |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1996 | Architectures for MPEG compressed bitstream scalingabstractThe idea of moving picture expert group (MPEG) bitstream scaling relates to altering or scaling the amount of data in a previously compressed MPEG bitstream. The new scaled bitstream conforms to constraints that are not known nor considered when the original preceded bitstream was constructed. Numerous applications for video transmission and storage are being developed based on the MPEG video coding standard. Applications such as video on demand, trick-play track on digital video tape recorders (VTR's) and extended-play recording on VTR's motivate the idea of bitstream scaling. In this paper, we present several bitstream scaling methods for the purpose of reducing the rate of constant bit rate (CBR) encoded bitstreams. The different methods have varying hardware implementation complexity and associated trade-offs in resulting image quality. Simulation results on MPEG test sequences demonstrate the typical performance trade-offs of the methods. Huifang Sun, Wilson Kwok, Joel W. Zdepski |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1995 | Spatial scalable HDTV codingabstractIn this paper, we discuss the coding strategy when a two-layer digital HDTV service is considered. We propose to use-three different strategies in performing compression to yield the best quality based on the respective requirements on bandwidth, picture quality and efficiency. These three modes are as follows. Strategy A: if the bandwidth for the enhancement layer is not sufficient to achieve high quality, the base layer encoder should perform the compression at a reduced spatial resolution using the combined bandwidth. The top layer video signal will be obtained from upsampling the base layer. Strategy B: if the bandwidth is sufficient to support quality requirements at both layers, the typical two-layer spatial scalable coding scheme is used. Strategy C: if the quality requirement at the top layer is critical, a single layer compression at a full spatial resolution is used. To receive the low resolution signal, certain bit stream scaling can be used to get a usable quality signal. A theoretical qualitative analysis gives an insight of our experimental results. Such a strategy is useful in designing future HDTV services for television sets with different capabilities, resolutions and complexities. The results are also applicable to the problem of HDTV/standard TV compatibility. Tihao Chiang, Huifang Sun, Joel W. Zdepski |
ICIP | 2 |
| 1995 | Architectures for MPEG compressed bitstream scalingabstractThe Moving Picture Expert Group (MPEG) video coding standard has been proposed for a variety of applications for video transmission and storage. Several applications such as video on demand, trick-play track on digital VTRs and extended-play recording VTRs motivate the idea of bitstream scaling which intends to reduce the compressed bitstream size according to the outgoing channel capacity. We propose several methods for implementation of bitstream scaling. The different methods have varying hardware implementation complexity, with each having its own degree of tradeoff between required hardware and resulting image quality. Huifang Sun, Wilson Kwok, Joel W. Zdepski |
ICIP | 1 |
| 1995 | Concealment of damaged block transform coded images using projections onto convex setsabstractAn algorithm for lost signal restoration in block-based still image and video sequence coding is presented. Problems arising from imperfect transmission of block-coded images result in lost blocks. The resulting image is flawed by the absence of square pixel regions that are notably perceived by human vision, even in real-time video sequences. Error concealment is aimed at masking the effect of missing blocks by use of temporal or spatial interpolation to create a subjectively acceptable approximation to the true error-free image. This paper presents a spatial interpolation algorithm that addresses concealment of lost image blocks using only intra-frame information. It attempts to utilize spatially correlated edge information from a large local neighborhood of surrounding pixels to restore missing blocks. The algorithm is a Gerchberg (1974) type spatial domain/spectral domain constraint-satisfying iterative process, and may be viewed as an alternating projections onto convex sets method. Huifang Sun, Wilson Kwok |
IEEE Trans. Image Process. | 1 |
| 1993 | ATM transport and cell-loss concealment techniques for MPEG video
Dipankar Raychaudhuri, Huifang Sun, Régis Saint-Girons |
ICASSP (1) | 2 |
| 1988 | Frame-adaptive vector quantization for image sequence codingabstractAn adaptive technique for image sequence coding that is based on vector quantization is described. Each frame in the sequence is first decomposed into a set of vectors. A codebook is generated using the vectors of the first frame as the training sequence, and a label map is created by quantizing the vectors. The vectors of the second frame are then used to generate a new codebook, starting with the first codebook as seeds. The updated codebook is then transmitted. At the same time, the label map is replenished by coding the position and the new values of the labels that have changed from one frame to the other. The process is repeated for subsequent frames. Experimental results for a test sequences demonstrate that the technique can track the changes and maintain a nearly constant distortion over the entire sequence.> Morris Goldberg, Huifang Sun |
IEEE Trans. Commun. | 2 |
| 1986 | Frame Adaptive Vector Quantization
Morris Goldberg, Huifang Sun |
ICC | 2 |
| 1986 | Image Sequence Coding Using Vector QuantizationabstractIn this paper, a new interframe coding technique based upon vector quantization is presented. This algorithm has as its basis twodimensional block vector quantization, at the frame level, onto which is grafted the concept of adaptive codebook replenishment and frame replenishment. This algorithm includes three strategies: label replenishment, label replenishment with mean shift, and label replenishment with mean shift and selective codeword replacement. The last two strategies efficiently update the codebook to track the changes in local statistics on a frame basis. Compared with three-dimensional block vector quantization [16], for bit rates between 0.5 and 0.65 the normalized mean Squared error (NMSE) is reduced by a factor ranging between 2 and 2.5 by the use of replenishment. Morris Goldberg, Huifang Sun |
IEEE Trans. Commun. | 2 |