Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ming-Ting Sun

dblp:82/1447 · DBLP profile ↗
← Back
151ranked-venue papers
4as first author
3since 2021 · last 2021
0009-0002-7082-9044ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 115 · 3 first-author · 2 since 2021Systems, architecture and hardware · 23 · 1 first-authorArtificial intelligence and machine learning · 9 · 1 since 2021Databases, data management, data science and information retrieval · 4Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
11 papers
Image and video coding · 45% Image and video processing · 34% Rendering · 12%
Artificial intelligence
5 papers
3D vision · 27% Representation and self-supervised learning · 25% Language models and text generation · 22%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 85% Web and social media mining · 13% Recommender systems · 2%

Topics — the 30 heaviest of 47, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video coding
video compression
0.742017
Signal Dependent Transform Based on SVD for HEVC Intracoding · IEEE Trans. Multim. 2017
Improving Intra Prediction in High-Efficiency Video Coding · IEEE Trans. Image Process. 2016
Mode-Dependent Templates and Scan Order for H.264/AVC-Based Intra Lossless Coding · IEEE Trans. Image Process. 2012
Image and video coding › video compression
intra prediction
0.522017
Signal Dependent Transform Based on SVD for HEVC Intracoding · IEEE Trans. Multim. 2017
Improving Intra Prediction in High-Efficiency Video Coding · IEEE Trans. Image Process. 2016
Image and video processing › super-resolution › image super-resolution
depth super-resolution
0.522016
Edge-Guided Single Depth Image Super Resolution · IEEE Trans. Image Process. 2016
Joint Super Resolution and Denoising From a Single Depth Image · IEEE Trans. Multim. 2015
Machine learning › Representation and self-supervised learning › visual representation
compact binary descriptor
0.412019
Unsupervised Deep Learning of Compact Binary Descriptors · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Natural language and speech › Language models and text generation › controllable text generation
text style transfer
0.412019
Domain Adaptive Text Style Transfer · EMNLP/IJCNLP (1) 2019
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.412019
Unsupervised Deep Learning of Compact Binary Descriptors · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Information retrieval › hashing
binary descriptor learning
0.412019
Unsupervised Deep Learning of Compact Binary Descriptors · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Information retrieval
hashing
0.412019
Unsupervised Deep Learning of Compact Binary Descriptors · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Rendering › image-based rendering
depth-image-based rendering
0.312018
Hole Filling With Multiple Reference Views in DIBR View Synthesis · IEEE Trans. Multim. 2018
Geometric modeling and processing › mesh processing › mesh repair
hole filling
0.312018
Hole Filling With Multiple Reference Views in DIBR View Synthesis · IEEE Trans. Multim. 2018
Rendering
novel view synthesis
0.312018
Hole Filling With Multiple Reference Views in DIBR View Synthesis · IEEE Trans. Multim. 2018
Image and video processing › video restoration
video compression artifact reduction
0.312018
Deep Kalman Filtering Network for Video Compression Artifact Reduction · ECCV (14) 2018
Computer vision › 3D vision
point cloud processing
0.312017
A Data-Driven Point Cloud Simplification Framework for City-Scale Image-Based Localization · IEEE Trans. Image Process. 2017
Computer vision › 3D vision › point cloud processing
point cloud simplification
0.312017
A Data-Driven Point Cloud Simplification Framework for City-Scale Image-Based Localization · IEEE Trans. Image Process. 2017
Natural language and speech › Language models and text generation
text generation
0.312017
Adversarial Ranking for Language Generation · NIPS 2017
Robotics › Robot navigation and mapping › localization
vision-based localization
0.312017
A Data-Driven Point Cloud Simplification Framework for City-Scale Image-Based Localization · IEEE Trans. Image Process. 2017
Image and video coding › video compression
intra coding
0.312017
Signal Dependent Transform Based on SVD for HEVC Intracoding · IEEE Trans. Multim. 2017
Image and video coding
transform coding
0.312017
Signal Dependent Transform Based on SVD for HEVC Intracoding · IEEE Trans. Multim. 2017
Machine learning › Transfer learning and domain adaptation › knowledge transfer
label transfer
0.212016
Semantic Instance Annotation of Street Scenes by 3D to 2D Label Transfer · CVPR 2016
Computer vision › Segmentation and scene understanding › instance segmentation
semantic instance segmentation
0.212016
Semantic Instance Annotation of Street Scenes by 3D to 2D Label Transfer · CVPR 2016
Image and video coding › video compression › video codec
HEVC
0.212016
Improving Intra Prediction in High-Efficiency Video Coding · IEEE Trans. Image Process. 2016
Image and video processing
image enhancement
0.212016
Edge-Guided Single Depth Image Super Resolution · IEEE Trans. Image Process. 2016
Image and video processing › super-resolution
image super-resolution
0.212016
Edge-Guided Single Depth Image Super Resolution · IEEE Trans. Image Process. 2016
Image and video processing › image restoration
image denoising
0.212015
Joint Super Resolution and Denoising From a Single Depth Image · IEEE Trans. Multim. 2015
Image and video processing
image restoration
0.212015
Joint Super Resolution and Denoising From a Single Depth Image · IEEE Trans. Multim. 2015
Image and video coding
entropy coding
0.112012
Mode-Dependent Templates and Scan Order for H.264/AVC-Based Intra Lossless Coding · IEEE Trans. Image Process. 2012
Image and video coding › video compression
lossless video compression
0.112012
Mode-Dependent Templates and Scan Order for H.264/AVC-Based Intra Lossless Coding · IEEE Trans. Image Process. 2012
Computational photography and imaging
depth sensing
0.122016
Edge-Guided Single Depth Image Super Resolution · IEEE Trans. Image Process. 2016
Joint Super Resolution and Denoising From a Single Depth Image · IEEE Trans. Multim. 2015
Wireless sensing and localization › radar sensing
mmwave radar sensing
0.112011
An Algorithm for Power Line Detection and Warning Based on a Millimeter-Wave Radar Video · IEEE Trans. Image Process. 2011
Wireless sensing and localization
radar sensing
0.112011
An Algorithm for Power Line Detection and Warning Based on a Millimeter-Wave Radar Video · IEEE Trans. Image Process. 2011

Methods — techniques the papers use, named apart from their topics

quantization loss · 0.8deep neural network · 0.8backpropagation · 0.8domain adaptation · 0.4adversarial training · 0.4view interpolation · 0.3view extrapolation · 0.3selective warping · 0.3recurrent neural network · 0.3deep kalman filtering · 0.3weighted k-cover · 0.3structure from motion · 0.3singular value decomposition · 0.3policy gradient · 0.3generative adversarial network · 0.3discrete sine transform · 0.3discrete cosine transform · 0.3gaussian markov model · 0.2
YearPublicationVenuePosition
2021 Learning Nonparametric Human Mesh Reconstruction From A Single Image Without Ground Truth Meshes
abstract
We present a novel approach to learn human mesh reconstruction without ground truth mesh labels. This is made possible by introducing two new terms into the loss function of a graph convolutional neural network (Graph CNN). The first term is the Laplacian prior that acts as a regularizer on the mesh reconstruction. The second term is the part segmentation loss that forces the projected region of the reconstructed mesh to match the part segmentation. Extensive experiments validate the effectiveness of the proposed approach.
Zicheng Liu 0001, Ming-Ting Sun
ICIP5
2021 Contextualized Perturbation for Textual Adversarial Attack
abstract
Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, Bill Dolan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Dianqi Li, Yizhe Zhang 0002, Hao Peng 0009, Liqun Chen 0001, Chris Brockett, Ming-Ting Sun, William B. Dolan
NAACL-HLT6
2021 Cross-Domain Complementary Learning Using Pose for Multi-Person Part Segmentation
abstract
Supervised deep learning with pixel-wise training labels has great successes on multi-person part segmentation. However, data labeling at pixel-level is very expensive. To solve the problem, people have been exploring to use synthetic data to avoid the data labeling. Although it is easy to generate labels for synthetic data, the results are much worse compared to those using real data and manual labeling. The degradation of the performance is mainly due to the domain gap, i.e., the discrepancy of the pixel value statistics between real and synthetic data. In this paper, we observe that real and synthetic humans both have a skeleton (pose) representation. We found that the skeletons can effectively bridge the synthetic and real domains during the training. Our proposed approach takes advantage of the rich and realistic variations of the real data and the easily obtainable labels of the synthetic data to learn multi-person part segmentation on real images without any human-annotated labels. Through experiments, we show that without any human labeling, our method performs comparably to several state-of-the-art approaches which require human labeling on Pascal-Person-Parts and COCO-DensePose datasets. On the other hand, if part labels are also available in the real-images during training, our method outperforms the supervised state-of-the-art methods by a large margin. We further demonstrate the generalizability of our method on predicting novel keypoints in real images where no real data labels are available for the novel keypoints detection. Code and pre-trained models are available at https://github.com/kevinlin311tw/CDCL-human-part-segmentation.
Yinpeng Chen, Zicheng Liu 0001, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.6
2019 Domain Adaptive Text Style Transfer
abstract
Dianqi Li, Yizhe Zhang, Zhe Gan, Yu Cheng, Chris Brockett, Bill Dolan, Ming-Ting Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Dianqi Li, Yizhe Zhang 0002, Zhe Gan, Yu Cheng 0001, Chris Brockett, William B. Dolan, Ming-Ting Sun
EMNLP/IJCNLP (1)7
2019 Unsupervised Deep Learning of Compact Binary Descriptors
abstract
Binary descriptors have been widely used for efficient image matching and retrieval. However, most existing binary descriptors are designed with hand-craft sampling patterns or learned with label annotation provided by datasets. In this paper, we propose a new unsupervised deep learning approach, called DeepBit, to learn compact binary descriptor for efficient visual object matching. We enforce three criteria on binary descriptors which are learned at the top layer of the deep neural network: 1) minimal quantization loss, 2) evenly distributed codes and 3) transformation invariant bit. Then, we estimate the parameters of the network through the optimization of the proposed objectives with a back-propagation technique. Extensive experimental results on various visual recognition tasks demonstrate the effectiveness of the proposed approach. We further demonstrate our proposed approach can be realized on the simplified deep neural network, and enables efficient image matching and retrieval speed with very competitive accuracies.
Jiwen Lu, Chu-Song Chen, Jie Zhou 0001, Ming-Ting Sun
IEEE Trans. Pattern Anal. Mach. Intell.5
2018 Deep Kalman Filtering Network for Video Compression Artifact Reduction
Guo Lu, Wanli Ouyang, Dong Xu 0001, Xiaoyun Zhang 0001, Ming-Ting Sun
ECCV (14)6
2018 A Framework for Surface Light Field Compression
abstract
Surface Light Fields (SLF) have previously been proposed for representing 3D scenes under complex lighting conditions' enabling immersive viewing experiences from arbitrary observation directions. In this work, we present a new approach for SLF representation and a framework for SLF compression. Specifically, the SLF is compactly represented in a B-Spline wavelet basis. This representation is capable of modeling diverse surface materials and complex lighting conditions. The coefficients of the B-Spline wavelet are then compressed by removing the spatial redundancy over surface points. Compared with image based light field compression, the proposed scheme is functionally advanced because it enables rendering objects from arbitrary viewpoints with both good quality and high efficiency. In terms of bitrate and distortion, experimental results have shown that the proposed method can achieve competitive performance but with much lower decoder computational complexity, indicating its potential in practical virtual and augmented reality applications.
Xiang Zhang 0004, Philip A. Chou, Ming-Ting Sun, Maolong Tang, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
ICIP3
2018 Hole Filling With Multiple Reference Views in DIBR View Synthesis
abstract
Depth-image-based rendering (DIBR) oriented view synthesis has been widely employed in the current depth-based 3-D video systems by synthesizing a virtual view from an arbitrary viewpoint. However, holes may appear in the synthesized view due to disocclusion, thus significantly degrading the quality. Consequently, efforts have been made on developing effective and efficient hole-filling algorithms. Current hole-filling techniques generally extrapolate/interpolate the hole regions with the neighboring information based on an assumption that the texture pattern in the holes is similar to that of the neighboring background information. However, in many scenarios, especially of complex texture, the assumption may not hold. In other words, hole-filling techniques can only provide an estimation for a hole which may not be good enough or may even be erroneous considering a wide variety of complex scene of images. In this paper, we first examine the view interpolation with multiple reference views, demonstrating that the problem of emerging holes in a target virtual view can be greatly alleviated by making good use of other neighboring complementary views in addition to its two (commonly used) most neighboring primary views. The effects of using multiple views for view extrapolation in reducing holes are also investigated in this paper. In view of the 3D Video and ongoing free-viewpoint TV standardization, we propose a new view synthesis framework, which employs multiple views to synthesize output virtual views. Furthermore, a scheme of selective warping of complementary views is developed by efficiently locating a small number of useful pixels in the complementary views for hole reduction, to avoid full warping of additional complementary views thus lowering greatly the warping complexity. Experimental results show that the hole size based on two primary reference views may be reduced by up to about 70% with the help of two complementary reference views in the case of view interpolation, while the hole size based on one primary reference view may be reduced by about 27% with the help of one more complementary reference view in view extrapolation. Moreover, it is shown that by using one more pair of views in view interpolation and one more view in view extrapolation, 10% hole pixels may be reduced additionally.
Shuai Li 0005, Ce Zhu, Ming-Ting Sun
IEEE Trans. Multim.3
2017 Learning Spatiotemporal and Geometric Features with ISA for Video-Based Facial Expression Recognition
ChenHan Lin, Junfeng Yao, Ming-Ting Sun, Jinsong Su
ICONIP (3)4
2017 Adversarial Ranking for Language Generation
abstract
Generative adversarial networks (GANs) have great successes on synthesizing data. However, the existing GANs restrict the discriminator to be a binary classifier, and thus limit their learning capacity for tasks that need to synthesize output with rich structures such as natural language descriptions. In this paper, we propose a novel generative adversarial network, RankGAN, for generating high-quality language descriptions. Rather than training the discriminator to learn and assign absolute binary predicate for individual data sample, the proposed RankGAN is able to analyze and rank a collection of human-written and machine-written sentences by giving a reference group. By viewing a set of data samples collectively and evaluating their quality through relative ranking scores, the discriminator is able to make better assessment which in turn helps to learn a better generator. The proposed RankGAN is optimized through the policy gradient technique. Experimental results on multiple public datasets clearly demonstrate the effectiveness of the proposed approach.
Dianqi Li, Xiaodong He 0001, Ming-Ting Sun, Zhengyou Zhang
NIPS4
2017 Fast Intra-Mode and CU Size Decision for HEVC
abstract
The latest video coding standard High Efficiency Video Coding (HEVC) achieves about a 50% bit-rate reduction compared with H.264/AVC under the same perceptual video quality. For intra coding, a coding unit (CU) is recursively divided into a quadtree-based structure from the largest CU 64 × 64 to the smallest CU 8 × 8. Also, up to 35 intra-prediction modes are allowed. These two techniques improve the intra-coding performance significantly. However, the encoding complexity increases several times compared with H.264/AVC intra coding. In this paper, fast intra-mode decision and CU size decision are proposed to reduce the complexity of HEVC intra coding while maintaining the rate-distortion (RD) performance. For fast intra-mode decision, a gradient-based method is proposed to reduce the candidate modes for rough mode decision and RD optimization. For fast CU size decision, the homogenous CUs are early terminated first. Then two linear support vector machines that employ the depth difference and HAD cost ratio (and RD cost ratio) as features are proposed to perform the decisions of early CU split and early CU termination for the rest of the CUs. Experimental results show that the proposed fast intra-coding algorithm achieves about a 54% encoding time reduction on average with only a 0.7% BD-rate increase for the HEVC reference software HM 14.0 under all-intra configuration.
Tao Zhang 0013, Ming-Ting Sun, Debin Zhao, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2017 A Data-Driven Point Cloud Simplification Framework for City-Scale Image-Based Localization
abstract
City-scale 3D point clouds reconstructed via structure-from-motion from a large collection of Internet images are widely used in the image-based localization task to estimate a 6-DOF camera pose of a query image. Due to prohibitive memory footprint of city-scale point clouds, image-based localization is difficult to be implemented on devices with limited memory resources. Point cloud simplification aims to select a subset of points to achieve a comparable localization performance using the original point cloud. In this paper, we propose a data-driven point cloud simplification framework by taking it as a weighted K-Cover problem, which mainly includes two complementary parts. First, a utility-based parameter determination method is proposed to select a reasonable parameter K for K-Cover-based approaches by evaluating the potential of a point cloud for establishing sufficient 2D-3D feature correspondences. Second, we formulate the 3D point cloud simplification problem as a weighted K-Cover problem, and propose an adaptive exponential weight function based on the visibility probability of 3D points. The experimental results on three popular datasets demonstrate that the proposed point cloud simplification framework outperforms the state-of-the-art methods for the image-based localization application with a well predicted parameter in the K-Cover problem.
Weisi Lin, Xinfeng Zhang 0001, Michael Goesele, Ming-Ting Sun
IEEE Trans. Image Process.5
2017 Signal Dependent Transform Based on SVD for HEVC Intracoding
abstract
Transform is used to compact the energy of the blocks into a small number of coefficients and is widely used in recent image/video coding standards. In the latest video coding standard high efficiency video coding (HEVC), a combination of discrete cosine transform (DCT) and discrete sine transform (DST) is adopted to transform the residuals from intra prediction. Since the DCT and DST are the fixed transforms that are derived from the Gauss-Markov model, some of residual blocks may not be compacted well by the DCT/DST. In this paper, we propose a signal dependent transform based on singular value decomposition (SVD) for HEVC intracoding. The proposed transform (SDT-SVD) is derived by performing SVD on the synthetic block and applied to the residual block considering the structural similarity between them. Furthermore, we extend SDT-SVD to template matching prediction (TMP) to further improve the intracoding performance. Experimental results show that the proposed transform on angular intra prediction (AIP) outperforms the latest HEVC reference software with a bit rate reduction of 1.0% on average and it can be up to 2.1%. When the proposed transform is extended to TMP-based intracoding, the overall bit rate reduction is 2.7% on average and can be up to 5.8%.
Tao Zhang 0013, Haoming Chen, Ming-Ting Sun, Debin Zhao, Wen Gao 0001
IEEE Trans. Multim.3
2016 Semantic Instance Annotation of Street Scenes by 3D to 2D Label Transfer
abstract
Semantic annotations are vital for training models for object recognition, semantic segmentation or scene understanding. Unfortunately, pixelwise annotation of images at very large scale is labor-intensive and only little labeled data is available, particularly at instance level and for street scenes. In this paper, we propose to tackle this problem by lifting the semantic instance labeling task from 2D into 3D. Given reconstructions from stereo or laser data, we annotate static 3D scene elements with rough bounding primitives and develop a model which transfers this information into the image domain. We leverage our method to obtain 2D labels for a novel suburban video dataset which we have collected, resulting in 400k semantic and instance image annotations. A comparison of our method to state-of the-art label transfer baselines reveals that 3D information enables more efficient annotation while at the same time resulting in improved accuracy and time-coherent labels.
Martin Kiefel, Ming-Ting Sun, Andreas Geiger 0001
CVPR3
2016 Lagrangian Multiplier Adaptation for Rate-Distortion Optimization With Inter-Frame Dependency
abstract
Rate-distortion optimization (RDO) is widely used in video coding, which plays a critical role in enhancing the coding efficiency substantially. Currently, the RDO process is performed in a way that coding efficiency of each coding unit (CU) is maximized independently without considering the dependency among CUs. As we know, in the current hybrid video coding structure, spatial/temporal prediction techniques are extensively used, which introduce strong dependency among CUs. In this paper, we investigate RDO with inter-frame dependency, where the impact of coding performance of the current CU on that of the following frames is considered. Accordingly, an RDO scheme taking the inter-frame dependency into account is proposed by adapting the Lagrangian multiplier. The experimental results show that the proposed scheme can achieve about 3.22% and 3.19% BD-rate saving in average over the state-of-the-art High Efficiency Video Coding (HEVC) reference software HM15.0 in the low-delay $P$ (LDP) and low-delay $B$ (LDB) coding structures, respectively, with no extra encoding time. The proposed scheme can obtain a significantly higher coding gain than the multiple quantization parameter (MQP) (±3) optimization technique that would greatly increase the encoding time by a factor of about six. Coupled with MQP optimization, the proposed scheme can further achieve about 5.96% and 5.57% BD-rate savings in average over the HEVC and about 4.03% and 4.07% over the HEVC with MQP optimization, under the specified common test conditions for LDP and LDB coding structures, respectively.
Shuai Li 0005, Ce Zhu, Yanbo Gao, Yimin Zhou 0002, Frédéric Dufaux, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.6
2016 Improving Intra Prediction in High-Efficiency Video Coding
abstract
Intra prediction is an important tool in intra-frame video coding to reduce the spatial redundancy. In current coding standard H.265/high-efficiency video coding (HEVC), a copying-based method based on the boundary (or interpolated boundary) reference pixels is used to predict each pixel in the coding block to remove the spatial redundancy. We find that the conventional copying-based method can be further improved in two cases: 1) the boundary has an inhomogeneous region and 2) the predicted pixel is far away from the boundary that the correlation between the predicted pixel and the reference pixels is relatively weak. This paper performs a theoretical analysis of the optimal weights based on a first-order Gaussian Markov model and the effects when the pixel values deviate from the model and the predicted pixel is far away from the reference pixels. It also proposes a novel intra prediction scheme based on the analysis that smoothing the copying-based prediction can derive a better prediction block. Both the theoretical analysis and the experimental results show the effectiveness of the proposed intra prediction method. An average gain of 2.3% on all intra coding can be achieved with the HEVC reference software.
Haoming Chen, Tao Zhang 0013, Ming-Ting Sun, Ankur Saxena, Madhukar Budagavi
IEEE Trans. Image Process.3
2016 Edge-Guided Single Depth Image Super Resolution
abstract
Recently, consumer depth cameras have gained significant popularity due to their affordable cost. However, the limited resolution and the quality of the depth map generated by these cameras are still problematic for several applications. In this paper, a novel framework for the single depth image superresolution is proposed. In our framework, the upscaling of a single depth image is guided by a high-resolution edge map, which is constructed from the edges of the low-resolution depth image through a Markov random field optimization in a patch synthesis based manner. We also explore the self-similarity of patches during the edge construction stage, when limited training data are available. With the guidance of the high-resolution edge map, we propose upsampling the high-resolution depth image through a modified joint bilateral filter. The edge-based guidance not only helps avoiding artifacts introduced by direct texture prediction, but also reduces jagged artifacts and preserves the sharp edges. Experimental results demonstrate the effectiveness of our method both qualitatively and quantitatively compared with the state-of-the-art methods.
Rogério Feris, Ming-Ting Sun
IEEE Trans. Image Process.3
2015 Efficient 24/7 object detection in surveillance videos
abstract
We address the problem of 24/7 object detection in urban surveillance videos, which presents unique challenges due to significant object appearance variations caused by lighting effects such as shadows and specular reflections, object pose variation, multiple weather conditions, and different times of the day. Rather than training a generic detector and adapting its parameters over time to handle all these variations, we rely on a large set of complementary and extremely efficient detector models, covering multiple overlapping appearance subspaces. At run time, our method continuously selects the most suitable detectors for a given scene and condition, using a novel approach inspired by parametric background modeling algorithms. We provide a comprehensive experimental analysis to show the effectiveness of our approach, considering traffic monitoring as our application domain. Our system runs at 100 frames per second on a standard laptop computer.
Rogério Feris, Russell Bobbitt, Sharath Pankanti, Ming-Ting Sun
AVSS4
2015 Anomaly detection by using random projection forest
abstract
In this paper, we present a novel method for detecting anomalies from surveillance videos, which utilizes the random projection forest for evaluating the rarity of visual clues in a video frame. Given the hierarchical clustering of the data in a random projection tree and the aggregation process in the random forest, we achieve both efficient estimation of incoming samples and improved robustness against under-fitting and over-fitting under improperly selected models. Random forest is also online updatable, which is meaningful for future online anomaly detection. We designed the splitting rule for anomaly detection, the system framework and the criterion of anomaly determination. The efficiency of the proposed methods has been validated by experiments on public UCSD datasets and compared with previously reported results.
Zicheng Liu 0001, Ming-Ting Sun
ICIP3
2015 Inter-frame dependent rate-distortion optimization using lagrangian multiplier adaption
abstract
It is known that, in the current hybrid video coding structure, spatial and temporal prediction techniques are extensively used which introduce strong dependency among coding units. Such dependency poses a great challenge to perform a global rate-distortion optimization (RDO) when encoding a video sequence. RDO is usually performed in a way that coding efficiency of each coding unit is optimized independently without considering dependeny among coding units, leading to a suboptimal coding result for the whole sequence. In this paper, we investigate the inter-frame dependent RDO, where the impact of coding performance of the current coding unit on that of the following frames is considered. Accordingly, an inter-frame dependent rate-distortion optimization scheme is proposed and implemented on the newest video coding standard High Efficiency Video Coding (HEVC) platform. Experimental results show that the proposed scheme can achieve about 3.19% BD-rate saving in average over the state-of-the-art HEVC codec (HM15.0) in the low-delay B coding structure, with no extra encoding time. It obtains a significantly higher coding gain than the multiple QP (±3) optimization technique which would greatly increase the encoding time by a factor of about 6. Coupled with the multiple QP optimization, the proposed scheme can further achieve a higher BD-rate saving of 5.57% and 4.07% in average than the HEVC codec and the multiple QP optimization enabled HEVC codec, respectively.
Shuai Li 0005, Ce Zhu, Yanbo Gao, Yimin Zhou 0001, Frédéric Dufaux, Ming-Ting Sun
ICME6
2015 Improvements on Intra Block Copy in natural content video coding
abstract
The Intra Block Copy (IntraBC) is a newly adopted tool in the HEVC extension for the screen content video coding. The IntraBC tool efficiently encodes repeating patterns in a picture. The current IntraBC scheme achieves about 1.0% bit-rate reduction on average and up to 4.3% bitrate reduction on natural content video for a database consisting of 2K, 4K, and 8K sequences. In this paper, we propose to improve the IntraBC with a template matching block vector and a fractional search IntraBC. With these two tools, the gain on natural content video coding can be further improved by 0.5% on average and up to 2.0%.
Haoming Chen, Yu-Sheng Chen, Ming-Ting Sun, Ankur Saxena, Madhukar Budagavi
ISCAS3
2015 Hybrid angular intra/template matching prediction for HEVC intra coding
abstract
In the latest HEVC video coding standard, angular intra prediction (AIP) applies 35 modes including 33 angular modes which can handle blocks with direction information well, and other 2 modes (DC and planar) which are used to predict smooth blocks. However, for blocks with complex texture, these modes may not give good predictions. Some complex blocks can be predicted well by template matching prediction (TMP) which was proposed to predict blocks having similar patterns in the coded regions of the same frame without the cost of large overheads. In this paper, a novel hybrid AIP/TMP is proposed to improve the prediction efficiency. Experimental results show that the proposed method can achieve a coding gain of about 2.5% for high resolution sequences on average compared to the AIP in HEVC. The gain can be up to 4.2%.
Tao Zhang 0013, Haoming Chen, Ming-Ting Sun, Debin Zhao, Wen Gao 0001
VCIP3
2015 Real time gaze estimation with a consumer depth camera
Zicheng Liu 0001, Ming-Ting Sun
Inf. Sci.3
2015 Adaptive intra-refresh for low-delay error-resilient video coding
Haoming Chen, Chen Zhao 0002, Ming-Ting Sun, Aaron Drake
J. Vis. Commun. Image Represent.3
2015 Fine registration of 3D point clouds fusing structural and photometric information using an RGB-D camera
Yu-Feng Hsu, Rogério Feris, Ming-Ting Sun
J. Vis. Commun. Image Represent.4
2015 Depth Coding Based on Depth-Texture Motion and Structure Similarities
abstract
This paper addresses high performance depth coding in 3D video by making good use of its coded texture video counterpart. The relationship between the depth and its associated texture video in terms of coding mode and motion vector is carefully examined. Our statistical study suggests that the skip-coding mode and its associated motion vectors in the coded texture can be shared for depth coding by saving bit rate at the cost of little increase of distortion, which subsequently results in a nonsequential coding of the depth map. In this sense, coding/prediction of a block can be performed using the skip-coded blocks below and right, which are not available in the conventional sequential coding, thus producing the so-called omnidirectional blocks predicted in the intra-coding by making the best use of (at most) four neighboring blocks. Moreover, in view of the depth-texture structure similarity, a depth-texture cooperative clustering-based prediction method is proposed for cluster-based depth prediction in the intra-coding, which exploits the structure similarity for the current coding block and its neighboring pixels around the block. On the other hand, some large prediction errors may be present for the depth-texture misaligned pixels, which may greatly compromise the coding performance. To deal with these large residuals induced by the depth-texture misalignment, a simple yet effective detection and rectification approach is incorporated in the proposed depth coding scheme. Experimental results show that our proposed depth coding scheme achieves superior rate-distortion performance compared with other relevant coding methods.
Jianjun Lei 0001, Shuai Li 0005, Ce Zhu, Ming-Ting Sun, Chunping Hou
IEEE Trans. Circuits Syst. Video Technol.4
2015 Joint Super Resolution and Denoising From a Single Depth Image
abstract
This paper describes a new algorithm for depth image super resolution and denoising using a single depth image as input. A robust coupled dictionary learning method with locality coordinate constraints is introduced to reconstruct the corresponding high resolution depth map. The local constraints effectively reduce the prediction uncertainty and prevent the dictionary from over-fitting. We also incorporate an adaptively regularized shock filter to simultaneously reduce the jagged noise and sharpen the edges. Furthermore, a joint reconstruction and smoothing framework is proposed with an L0 gradient smooth constraint, making the reconstruction more robust to noise. Experimental results demonstrate the effectiveness of our proposed algorithm compared to previously reported methods.
Rogério Feris, Shiaw-Shian Yu, Ming-Ting Sun
IEEE Trans. Multim.4
2014 Joint Denoising and demosaicking of noisy CFA images based on inter-color correlation
abstract
Most digital cameras use a single sensor coupled with a Color Filter Array (CFA) to capture images, and apply demosaicking to interpolate the full color images. In reality, the CFA image is noisy, which causes problems in the demosaicking process. This paper proposes a Joint Denoising and Demo-saicking based on inter-Color correlation (JDDC) scheme. We propose a new framework that linearly combines an extracted luminance image and a low-passed RGB images to get a full color image. Given the noise in the extracted luminance image and the low-passed RGB images are non-stationary and partially correlated, we modify the classical Non-Local Means (NLM) filter to denoise the extracted luminance image and the low-passed RGB images before the combination. Experimental results verify the effectiveness of the proposed scheme both objectively and subjectively.
Ming-Ting Sun, Lu Fang 0001, Oscar C. Au
ICASSP2
2014 Edge guided single depth image super resolution
abstract
Recently, consumer depth cameras have gained significant popularity due to their affordable cost. However, the limited resolution and quality of the depth map generated by these cameras are still problems for several applications. In this paper, we propose a novel framework for single depth image super resolution guided by a high resolution edge map constructed from the edges in the low resolution depth image via a Markov Random Field (MRF) optimization. With the guidance of the high resolution edge map, the high resolution depth image is up-sampled via a joint bilateral filter. The edge guidance not only helps avoid artifacts introduced by direct texture prediction, but also reduces the jagged artifacts and preserves the sharp edges. Experimental results demonstrate the effectiveness of our proposed algorithm compared to previously reported methods.
Rogério Feris, Ming-Ting Sun
ICIP3
2014 Video Summarization based on Nonnegative Linear Reconstruction
abstract
With the development of imaging techniques and the Internet, it is hard to effectively and efficiently manage, index and store large amounts of videos. Video summarization appears to this need by obtaining useful information in videos. Recently, the reconstruction concept has been introduced into video summarization. However, most of these existing reconstruction methods where a dictionary of key frames is selected to reconstruct the original video best might come up with redundant information. In this paper, we propose a novel framework named Video Summarization based on Nonnegative Linear Reconstruction (VSNLR) which allows only additive, not subtractive, linear combinations. Our approach consists of two core steps: (1) we detect the shot boundary and choose a representative frame for every shot, then all the representative frames form the candidate set; (2) for every frame in the candidate set, VSNLR selects related frames to reconstruct the given frame by using the nonnegative linear reconstruction function. The key frames are selected by minimizing the sum of reconstruction errors. Experiments on a dataset and comparison to the state-of-art method demonstrate our advantage.
Qiao Luan, Mingli Song, Chu Yee Liau, Jiajun Bu, Zicheng Liu 0001, Ming-Ting Sun
ICME6
2014 Realtime gaze estimation with online calibration
abstract
For an eye gaze estimation system, calibration is an unavoidable procedure to determine certain person-specific parameters, either explicitly or implicitly. Although several offline implicit calibration methods have been proposed to ease the calibration burden, the calibration procedure is still cumbersome and the gaze estimation accuracy needs further improvement. In this paper, we propose a novel 3D gaze estimation system with online calibration. The proposed system uses a new 3D model-based gaze estimation method with a single consumer camera (Kinect). Unlike previous gaze estimation methods using explicit offline calibration with fixed number of calibration points or implicit calibration, our approach constantly improves person-specific eye parameters through online calibration, which enables the system to adapt gradually to a new user. The experimental results and the human-computer interaction (HCI) application show that the proposed system can work in realtime with superior gaze estimation accuracy (<; 2°) and minimal calibration burden.
Mingli Song, Zicheng Liu 0001, Ming-Ting Sun
ICME4
2014 Single depth image super resolution and denoising via coupled dictionary learning with local constraints and shock filtering
abstract
Recently, consumer depth cameras have gained significant popularity due to their affordable cost. However, the limited resolution and quality of the depth map generated by these cameras are still problems for several applications. In this paper, we propose a new algorithm for depth image super resolution using a single depth image as input. We reconstruct the corresponding high resolution depth map through a robust coupled dictionary learning algorithm with local coordinate constraints. The local constraints remove the prediction uncertainty and prevent the dictionary from over-fitting. We also incorporate an adaptively regularized Shock filter to simultaneously reduce the noise and sharpen the edges. Experimental results demonstrate the effectiveness of our proposed algorithm compared to previously reported methods.
Cheng-Chuan Chou, Rogério Feris, Ming-Ting Sun
ICME4
2014 Deblur a blurred RGB image with a sharp NIR image through local linear mapping
abstract
Image acquisition in a low light environment requires long exposure to achieve acceptable signal-to-noise ratio, which however causes blurry effect. This paper addresses this problem by using a sharp near-infrared (NIR) image when the environment has sufficient NIR light. We assume that an RGB and NIR image pair has a linear mapping in a local area and that the mapping function is valid for both the blur and sharp image pairs. Using this property, we solve the sharp RGB images from a blurred RGB image and the corres ponding s harp NIR image. The effectiveness of the proposed algorithm is verified with both synthetic and real captured datasets.
Tao Yue 0003, Ming-Ting Sun, Zhengyou Zhang, Jin-Li Suo, Qionghai Dai
ICME2
2014 Appearance-Based Object Detection Under Varying Environmental Conditions
abstract
Practical surveillance systems deployed in urban scenarios need to operate 24/7 under a wide range of environmental conditions. As modern video analytics shift from blob-based to object-centered architectures, appearance-based object detection under different weather conditions and lighting effects emerges as a critical yet largely unaddressed problem. This paper investigates this research topic, using as a case study the problem of vehicle detection in urban surveillance environments. In particular, we show that a simple and efficient Winsorized lighting correction technique improves performance significantly when outliers due to shadows, specularities, headlights, and occluders are present. Moreover, we demonstrate that a self-training mechanism utilizing a balanced training set automatically acquired from the target domain yields superior performance. Our experimental results are carried out on a novel dataset of vehicle images collected from a public traffic camera and categorized according to multiple environmental conditions.
Rogério Feris, Lisa M. Brown, Sharath Pankanti, Ming-Ting Sun
ICPR4
2014 Recognizing object manipulation activities using depth and visual cues
Matthai Philipose, Martin Pettersson, Ming-Ting Sun
J. Vis. Commun. Image Represent.4
2014 Automatic objects segmentation with RGB-D cameras
Matthai Philipose, Ming-Ting Sun
J. Vis. Commun. Image Represent.3
2013 Fine registration of 3D point clouds with iterative closest point using an RGB-D camera
abstract
We address the problem of accurate and efficient alignment of 3D point clouds captured by an RGB-D (Kinect-style) camera from different viewpoints. Our approach introduces a new cost function for the iterative closest point (ICP) algorithm that balances the significance of structural and photometric features with dynamically adjusted weights to improve the error minimization process. We also enhance the algorithm with a novel outlier rejection method, which relies on adaptive thresholding at each ICP iteration, using both the structural information of the object and the spatial distances of sparse SIFT feature pairs. The effectiveness of our proposed approach is demonstrated in challenging scenarios, involving objects lacking structural features, and significant camera view and lighting changes. We obtained superior registration accuracy than existing related methods while requiring low computational processing.
Yu-Feng Hsu, Rogério Feris, Ming-Ting Sun
ISCAS4
2013 Boosting object detection performance in crowded surveillance videos
abstract
We present a novel approach to automatically create efficient and accurate object detectors tailored to work well on specific video surveillance cameras (specific-domain detectors), using samples acquired with the help of a more expensive, general-domain detector (trained using images from multiple cameras). Our method requires no manual labels from the target domain. We automatically collect training data using tracking over short periods of time from high-confidence samples selected by the general-domain detector. In this context, a novel confidence measure is proposed for detectors based on a cascade of classifiers, which are frequently adopted for computer vision applications that require real-time processing. We demonstrate our proposed approach on the problem of vehicle detection in crowded surveillance videos, showing that an automatically generated detector significantly outperforms the original general-domain detector with much less feature computations.
Rogério Feris, Ankur Datta, Sharath Pankanti, Ming-Ting Sun
WACV4
2013 Complexity Reduction and Performance Improvement for Geometry Partitioning in Video Coding
abstract
Geometry partitioning for video coding involves establishing a partition line boundary within each block-shaped region and applying motion-compensated prediction to the two sub-regions created by the partition line. This paper presents techniques for enhancing the effectiveness and reducing the complexity of geometry partitioning schemes. A texture-difference-based approach is described to simplify the process of selecting the partition lines. Applying this approach together with a described skipping strategy for blocks with uniform texture can achieve a 94% reduction of encoding time while retaining a similar rate-distortion (R-D) performance to the full-search partitioning approach, when implemented for wedge-based geometry partitioning (WGP) in the context of H.264/MPEG-4 AVC JM 16.2. A bit rate improvement of approximately 6% is shown relative to not using geometry partitioning. For further R-D improvement, we describe a background-compensated prediction scheme to reduce the number of overhead bits used for motion vectors. Additionally, for systems in which high-quality depth maps are available, we incorporate depth map usage into the described approaches to generate a more accurate partitioning. Using these approaches with object-boundary-based geometry partitioning can achieve about 9% bit rate savings relative to using WGP, while keeping a similar computational complexity to the described complexity-reduced WGP.
Qifei Wang, Xiangyang Ji, Ming-Ting Sun, Gary J. Sullivan, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.3
2012 Automatic object segmentation with 3-D cameras
abstract
Recently, active 3-D cameras, which provide streams of depth and color images, have become widespread and popular. The depth data provides useful information for identifying object boundaries, making automatic object segmentation possible. However, the depth images are extremely noisy, and due to different response time of the color and depth sensors, the depth and color information often lose synchronization when the object is moving fast. In this work, we show how to combine depth and color information to clean up the depth maps and produce an accurate segmentation of the object. On a large dataset, we show that our proposed techniques are effective.
Matthai Philipose, Ming-Ting Sun
ICIP3
2012 Sparsity-based online missing sensor data recovery
abstract
In sensor networks, due to power outage at a sensor node, hardware dysfunction, or bad environmental conditions, not all sensor samples can be successfully gathered at the sink. Additionally, in the data stream scenario, some nodes may continually miss samples for a period of time. In this paper, a sparsity-based online data recovery approach is proposed. We construct an over complete dictionary composed of past data frames and traditional fixed transform bases. Assuming the current frame can be sparsely represented using only a few elements of the dictionary, missing samples in each frame can be estimated by Basis Pursuit. Our method was tested on data from a real sensor network application: monitoring the temperatures of the disk drive racks at a data center. Simulations show that in terms of estimation accuracy and stability, the proposed approach outperforms existing average-based interpolation methods, and is more robust to burst missing along the time dimension.
Di Guo 0003, Xiaobo Qu 0001, Lianfen Huang, Zicheng Liu 0001, Ming-Ting Sun
ISCAS6
2012 Complexity-reduced geometry partition search and high efficiency prediction for video coding
abstract
To reduce the complexity of searching for wedge-based geometry partitions in video coding, we propose a texture-difference based partition line selection approach with a skipping strategy. Applying this approach can reduce the encoding time by 90% while retaining the similar rate-distortion performance to that of exhaustive searching. We also propose a background-compensated prediction scheme to improve the rate-distortion performance of the geometry partitioning prediction by reducing its motion vector overhead. Incorporating our proposed approach into object-boundary-based geometry partitioning can achieve about 10% bit-rate savings relative to the full-search approach while keeping the complexity at about the same level as our proposed complexity-reduced wedge-based geometry partitioning.
Qifei Wang, Ming-Ting Sun, Gary J. Sullivan
ISCAS2
2012 Region-Based Rate Control for H.264/AVC for Low Bit-Rate Applications
abstract
Rate control plays an important role in video coding. However, in the conventional rate control algorithms, the number and position of macroblocks (MBs) inside one basic unit for rate control is inflexible and predetermined. The different characteristics of the MBs are not fully considered. Also, there is no overall optimization of the coding of basic units. This paper proposes a new region-based rate control scheme for H.264/advanced video coding to improve the coding efficiency. The inter-frame information is explored to objectively divide one frame into multiple regions based on their rate-distortion (R-D) behaviors. The MBs with similar characteristics are classified into the same region, and the entire region, instead of a single MB or a group of contiguous MBs, is treated as a basic unit for rate control. A linear rate-quantization stepsize model and a linear distortion-quantization stepsize model are proposed to accurately describe the R-D characteristics for the region-based basic units. Moreover, based on the above linear models, an overall optimization model is proposed to obtain suitable quantization parameters for the region-based basic units. Experimental results demonstrate that the proposed region-based rate control approach can achieve both better subjective and objective quality by performing the rate control adaptively with the content, compared to the conventional rate control approaches.
Hai-Miao Hu, Bo Li 0006, Weiyao Lin, Wei Li 0209, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.5
2012 Mode-Dependent Templates and Scan Order for H.264/AVC-Based Intra Lossless Coding
abstract
In H.264/advanced video coding (AVC), lossless coding and lossy coding share the same entropy coding module. However, the entropy coders in the H.264/AVC standard were original designed for lossy video coding and do not yield adequate performance for lossless video coding. In this paper, we analyze the problem with the current lossless coding scheme and propose a mode-dependent template (MD-template) based method for intra lossless coding. By exploring the statistical redundancy of the prediction residual in the H.264/AVC intra prediction modes, more zero coefficients are generated. By designing a new scan order for each MD-template, the scanned coefficients sequence fits the H.264/AVC entropy coders better. A fast implementation algorithm is also designed. With little computation increase, experimental results confirm that the proposed fast algorithm achieves about 7.2% bit saving compared with the current H.264/AVC fidelity range extensions high profile.
Zhouye Gu, Weisi Lin, Bu-Sung Lee, Chiew Tong Lau, Ming-Ting Sun
IEEE Trans. Image Process.5
2012 Resource-Efficient FPGA Architecture and Implementation of Hough Transform
abstract
Hough transform is widely used for detecting straight lines in an image, but it involves huge computations. For embedded application, field-programmable gate arrays are one of the most used hardware accelerators to achieve real-time implementation of Hough transform. In this paper, we present a resource-efficient architecture and implementation of Hough transform on an FPGA. The incrementing property of Hough transform is described and used to reduce the resource requirement. In order to facilitate parallelism, we divide the image into blocks and apply the incrementing property to pixels within a block and between blocks. Moreover, the locality of Hough transform is analyzed to reduce the memory access. The proposed architecture is implement on an Altera EP2S180F1508C3 device and can operate at a maximum frequency of 200 MHz. It could compute the Hough transform of 512 × 512 test images with 180 orientations in 2.07-3.16 ms without using many FPGA resources (i.e., one could achieve the performance by adopting a low-cost low-end FPGA).
Zhong-Ho Chen, Alvin Wen-Yu Su, Ming-Ting Sun
IEEE Trans. Very Large Scale Integr. Syst.3
2011 Accelerating Statistical LOR Estimation for a High-Resolution PET Scanner Using FPGA Devices and a High Level Synthesis Tool
abstract
In this paper, we use an FPGA platform and a high level synthesis tool, called Impulse C, to speedup a statistical Line Of Reaction (LOR) estimation for a high-resolution Positron Emission Tomography (PET) scanner. The estimation algorithm provides a significant improvement over conventional methods, but the execution time is too long to be practical for clinic applications. Impulse C allows us to rapidly map a C program into a platform with a host processor coupled to an FPGA device. However, the generated HDLs from the original codes are very inefficient, and the execution time is even worse than the software code. We describe some optimization methods for the algorithm using Impulse C. These methods could also be applied to other applications or used to improve the high level synthesis tools. The results show that the FPGA implementation can obtain a 82× speedup over the optimized software.
Zhong-Ho Chen, Alvin Wen-Yu Su, Ming-Ting Sun, Scott Hauck
FCCM3
2011 An effecive night video enhancement algorithm
abstract
Night video enhancement is important for video surveillance since many objects or activities of interest occur in a dark environment which cannot be seen easily without enhancement. In this paper, we discuss several problems of existing techniques for illumination-fusion based night video enhancement, which fuses video frames from day-time backgrounds and night-time video. We then present a simple enhancement algorithm without these problems. The algorithm uses an additive enhancement term with foreground object extraction and constrained low-passed object illumination to avoid light-inversion and sensitivity problems and to reduce ghost patterns. Experimental results show the effectiveness and robustness of the proposed algorithm.
Yunbo Rao, Zhong-Ho Chen, Ming-Ting Sun, Yu-Feng Hsu, Zhengyou Zhang
VCIP3
2011 Reduced-complexity search for video coding geometry partitions using texture and depth data
abstract
In this paper, a texture-space geometry partitioning approach is proposed to reduce the computational complexity of searching for geometry partitions for video coding. Additionally, for systems that capture both video and depth data, an enhanced geometry partitioning approach using both texture and depth information is proposed to further improve the partitioning accuracy and reduce the search complexity. Compared with a full-search approach, the proposed geometry partition search approaches achieve about 94% reduction of the encoding time while retaining similar rate-distortion performance.
Qifei Wang, Gary J. Sullivan, Ming-Ting Sun
VCIP4
2011 Performance analysis, parameter selection and extensions to H.264/AVC FRExt for high resolution video coding
Chenwei Deng, Weisi Lin, Bu-Sung Lee, Chiew Tong Lau, Ming-Ting Sun
J. Vis. Commun. Image Represent.5
2011 A rate-control algorithm using inter-layer information for H.264/SVC for low-delay applications
Hai-Miao Hu, Bo Li 0006, Weiyao Lin, Ming-Ting Sun
J. Vis. Commun. Image Represent.4
2011 A region-based rate-control scheme using inter-layer information for H.264/SVC
Hai-Miao Hu, Weiyao Lin, Bo Li 0006, Ming-Ting Sun
J. Vis. Commun. Image Represent.4
2011 Automatic video activity detection using compressed domain motion trajectories for H.264 videos
Ming-Ting Sun, Ruei-Cheng Wu, Shiaw-Shian Yu
J. Vis. Commun. Image Represent.2
2011 Rate-distortion optimized rate-allocation for motion-compensated predictive video codecs using PixelRank
Jian Lou 0006, Ming-Ting Sun
J. Vis. Commun. Image Represent.2
2011 A Fast Sub-Pixel Motion Estimation Algorithm for H.264/AVC Video Coding
abstract
Motion estimation (ME) is one of the most time-consuming parts in video coding. The use of multiple partition sizes in H.264/AVC makes it even more complicated when compared to ME in conventional video coding standards. It is important to develop fast and effective sub-pixel ME algorithms since: 1) the computation overhead by sub-pixel ME has become relatively significant while the complexity of integer-pixel search has been greatly reduced by fast algorithms, and 2) reducing sub-pixel search points can greatly save the computation for sub-pixel interpolation. In this letter, a novel fast sub-pixel ME algorithm is proposed which performs a “rough” sub-pixel search before the partition selection, and performs a “precise” sub-pixel search for the best partition. By reducing the searching load for the large number of non-best partitions, the computation complexity for sub-pixel search can be greatly decreased. Experimental results show that our method can reduce the sub-pixel search points by more than 50% compared to existing fast sub-pixel ME methods with negligible quality degradation.
Weiyao Lin, Krit Panusopone, David M. Baylon, Ming-Ting Sun, Zhenzhong Chen 0001, Hongxiang Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2011 An Algorithm for Power Line Detection and Warning Based on a Millimeter-Wave Radar Video
abstract
Power-line-strike accident is a major safety threat for low-flying aircrafts such as helicopters, thus an automatic warning system to power lines is highly desirable. In this paper we propose an algorithm for detecting power lines from radar videos from an active millimeter-wave sensor. Hough Transform is employed to detect candidate lines. The major challenge is that the radar videos are very noisy due to ground return. The noise points could fall on the same line which results in signal peaks after Hough Transform similar to the actual cable lines. To differentiate the cable lines from the noise lines, we train a Support Vector Machine to perform the classification. We exploit the Bragg pattern, which is due to the diffraction of electromagnetic wave on the periodic surface of power lines. We propose a set of features to represent the Bragg pattern for the classifier. We also propose a slice-processing algorithm which supports parallel processing, and improves the detection of cables in a cluttered background. Lastly, an adaptive algorithm is proposed to integrate the detection results from individual frames into a reliable video detection decision, in which temporal correlation of the cable pattern across frames is used to make the detection more robust. Extensive experiments with real-world data validated the effectiveness of our cable detection algorithm.
Qirong Ma, Darren S. Goshi, Yi-Chi Shih, Ming-Ting Sun
IEEE Trans. Image Process.4
2011 Introduction to the ICME2010 Special Issue
abstract
The 15 papers in this special issue are extended versions of papers presented at the 2010 IEEE International Conference on Multimedia and Expo (ICME), held in Singapore on July 19-23, 2010. These papers cover a wide range of topics in multimedia including user interface, content understanding, mobility, 3-D processing, storage, and forensics.
Zicheng Liu 0001, Ming-Ting Sun, Chia-Wen Lin, Zhengyou Zhang, Zhu Liu 0001, Homer H. Chen, Yap-Peng Tan, Oscar C. Au
IEEE Trans. Multim.2
2011 Robust Detection of Abandoned and Removed Objects in Complex Surveillance Videos
abstract
Tracking-based approaches for abandoned object detection often become unreliable in complex surveillance videos due to occlusions, lighting changes, and other factors. We present a new framework to robustly and efficiently detect abandoned and removed objects based on background subtraction (BGS) and foreground analysis with complement of tracking to reduce false positives. In our system, the background is modeled by three Gaussian mixtures. In order to handle complex situations, several improvements are implemented for shadow removal, quick-lighting change adaptation, fragment reduction, and keeping a stable update rate for video streams with different frame rates. Then, the same Gaussian mixture models used for BGS are employed to detect static foreground regions without extra computation cost. Furthermore, the types of the static regions (abandoned or removed) are determined by using a method that exploits context information about the foreground masks, which significantly outperforms previous edge-based techniques. Based on the type of the static regions and user-defined parameters (e.g., object size and abandoned time), a matching method is proposed to detect abandoned and removed objects. A person-detection process is also integrated to distinguish static objects from stationary people. The robustness and efficiency of the proposed method is tested on IBM Smart Surveillance Solutions for public safety applications in big cities and evaluated by several public databases, such as The Image library for intelligent detection systems (i-LIDS) and IEEE Performance Evaluation of Tracking and Surveillance Workshop (PETS) 2006 datasets. The test and evaluation demonstrate our method is efficient to run in real-time, while being robust to quick-lighting changes and occlusions in complex environments.
Yingli Tian, Rogério Feris, Arun Hampapur, Ming-Ting Sun
IEEE Trans. Syst. Man Cybern. Part C5
2010 Unsupervised action classification using space-time link analysis
abstract
In this paper we address the problem of unsupervised discovery of action classes in video data. Different from all existing methods thus far proposed for this task, we present a space-time link analysis approach which matches the performance of traditional unsupervised action categorization methods in a standard dataset. Our method is inspired by the recent success of link analysis techniques in the image domain. By applying these techniques in the space-time domain, we are able to naturally take into account the spatio-temporal relationships between the video features, while leveraging the power of graph matching for action classification. We present an experiment to demonstrate that our approach is capable of handling cluttered backgrounds, activities with subtle movements, and video data from moving cameras.
Rogério Feris, Volker Krüger, Ming-Ting Sun
ISCAS4
2010 Video activity detection using compressed domain motion trajectories for H.264 videos
abstract
Surveillance videos are often compressed for transmission or storage. It is desirable to be able to perform automatic event detection in the compressed domain directly. In this paper, we investigate the use of motion trajectories for video activity detection in the compressed domain. We show that it is possible to extract reliable motion trajectories directly from compressed H.264 video streams. To overcome the problems caused by unreliable motion vectors, we propose to include the information from the compressed domain prediction residuals to make the tracking more robust. We also show a real world application based on the classification of the motion trajectories to detect vacant or occupied parking spaces.
Ming-Ting Sun, Ruei-Cheng Wu, Shiaw-Shian Yu
ISCAS2
2010 Rate-Distortion Cost Estimation for H.264/AVC
abstract
Rate-distortion cost estimation is useful for many H.264/advanced video coding (AVC) applications including rate-distortion optimization (RDO) for mode-decision and rate-control. In this paper, we propose a new rate-prediction model and an adaptive algorithm to provide more accurate estimation of the number of coding bits for encoding intra and inter-blocks compared to previously proposed methods. The rate estimation is modeled by a linear combination of existing coding parameters, which are more accurately related to entropy coding and transform coefficients. Based on the proposed model, a cost estimation function is also proposed to give a more accurate rate-distortion cost estimation. Furthermore, we propose a block classification approach to further improve the effectiveness of the proposed scheme. The proposed schemes can achieve better results in estimating the bit-rate and rate-distortion cost compared to previously proposed approaches.
Ke-Ying Liao, Jar-Ferr Yang, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.3
2010 A Computation Control Motion Estimation Method for Complexity-Scalable Video Coding
abstract
In this paper, a new computation-control motion estimation (CCME) method is proposed which can perform motion estimation (ME) adaptively under different computation or power budgets while keeping high coding performance. We first propose a new class-based method to measure the macroblock (MB) importance where MBs are classified into different classes and their importance is measured by combining their class information as well as their initial matching cost information. Based on the new MB importance measure, a complete CCME framework is then proposed to allocate computation for ME. The proposed method performs ME in a one-pass flow. Experimental results demonstrate that the proposed method can allocate computation more accurately than previous methods and, thus, has better performance under the same computation budget.
Weiyao Lin, Krit Panusopone, David M. Baylon, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.4
2010 Group Event Detection With a Varying Number of Group Members for Video Surveillance
abstract
This paper presents a novel approach for automatic recognition of group activities for video surveillance applications. We propose to use a group representative to handle the recognition with a varying number of group members, and use an asynchronous hidden Markov model (AHMM) to model the relationship between people. Furthermore, we propose a group activity detection algorithm which can handle both symmetric and asymmetric group activities, and demonstrate that this approach enables the detection of hierarchical interactions between people. Experimental results show the effectiveness of our approach.
Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang
IEEE Trans. Circuits Syst. Video Technol.2
2010 Edge-Directed Error Concealment
abstract
In this paper we propose an edge-directed error concealment (EDEC) algorithm, to recover lost slices in video sequences encoded by flexible macroblock ordering. First, the strong edges in a corrupted frame are estimated based on the edges in the neighboring frames and the received area of the current frame. Next, the lost regions along these estimated edges are recovered using both spatial and temporal neighboring pixels. Finally, the remaining parts of the lost regions are estimated. Simulation results show that compared to the existing boundary matching algorithm [1] and the exemplar-based inpainting approach [2] , the proposed EDEC algorithm can reconstruct the corrupted frame with both a better visual quality and a higher decoder peak signal-to-noise ratio.
Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.4
2009 Healthcare audio event classification using Hidden Markov Models and Hierarchical Hidden Markov Models
abstract
Audio is a useful modality complement to video for healthcare monitoring. In this paper, we investigate the use of Hierarchical Hidden Markov Models (HHMMs) for healthcare audio event classification. We show that HHMM can handle audio events with recursive patterns to improve the classification performance. We also propose a model fusion method to cover large variations often existing in healthcare audio events. Experimental results from classifying key eldercare audio events show the effectiveness of the model fusion method for healthcare audio event classification.
Ya-Ti Peng, Ching-Yung Lin, Ming-Ting Sun, Kun-Cheng Tsai
ICME3
2009 Multimodal signal fusion
abstract
Our daily life involves multimodal signals (e.g., visual, audio, text, and signals from various sensors). Multimodal signals are highly correlated. For example, researchers have been using audio activities for video summarization of ball games and violence detection in movies. Multimodal signals are complementary to each other. For example, in multimodal surveillance, audio may carry important information not available in video. Usually, people don't hire a blind or deaf person for surveillance duties, since a human naturally fuses multimodal signals for the best results.
Ming-Ting Sun
ICME1
2009 A New Class-based Early Termination Method for Fast Motion Estimation in Video Coding
abstract
Motion Estimation (ME) is one of the most time-consuming parts in video coding. It is always desirable to develop fast ME algorithms to reduce the ME complexity. In this paper, a new early termination method is proposed for fast motion estimation. The proposed method first classifies each macroblock into one of three classes based on the estimation of the possible matching cost improvement from future search points. Different early termination strategies are then applied to different classes. Experimental results show that the proposed method can significantly reduce the search points with little quality degradation.
Weiyao Lin, Krit Panusopone, David M. Baylon, Ming-Ting Sun
ISCAS4
2009 A New One-pass Complexity-Scalable Computation-control Method for Video Coding
abstract
In this paper, a new complexity-scalable computation control method is proposed which can perform motion estimation (ME) adaptively under different computation or power budgets while keeping high coding performance. We first propose a new class-based method to measure the macroblock (MB) importance where MBs are classified into different classes and their importance is measured by combining their class information as well as their initial matching cost information in the ME. Based on the new MB importance measure, a computation control framework is then proposed to allocate computation for ME. Experimental results demonstrate that the proposed method can allocate computation more accurately than previous methods and thus has better performance under the same computation budget.
Weiyao Lin, Krit Panusopone, David M. Baylon, Ming-Ting Sun
ISCAS4
2009 Group Event Detection for Video Surveillance
abstract
This paper presents a novel approach for automatic recognition of group activities for video surveillance applications. We propose to use a group representative to handle the recognition with flexible or varying number of group members, and use an asynchronous hidden Markov model (AHMM) to model the relationship between two people. Furthermore, we propose a group activity detection algorithm which can handle symmetric and asymmetric group activities, and demonstrate that this approach enables the detection of hierarchical interactions between people. Experimental results show the effectiveness of our approach.
Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang
ISCAS2
2009 Error Concealment for Spatially Scalable Video Coding using Hallucination
abstract
We propose a new error concealment method based on hallucination for Scalable Video Coding with spatial scalability. In this method, parts of the frames which lose the enhancement layer are up-sampled from base layer and “hallucinated” as concealment frames. The database for hallucination is generated from the high-resolution and low-resolution frame-pairs near the lost frames in the video sequence. The effectiveness of hallucination here lies in the similarity between the nearby frames in the video sequence. Experiments show that the proposed method has superior results over the state-of-the-art error concealment method for spatially Scalable Video Coding.
Qirong Ma, Ming-Ting Sun
ISCAS3
2009 H.264 Deblocking Speedup
abstract
This letter tackles the problem of reducing the complexity of H.264 decoding. Since deblocking accounts for a significant percentage of H.264 decoding time, our focus is on the H.264 in-loop deblocking filter. Observing that branch operations are costly and that in the deblocking process there are events with significantly high probability of occurrence, we regroup and simplify the branch operations. We apply the idea of Huffman tree optimization to speed up the boundary strength derivation and the true-edge detection by taking advantage of the biased statistical distribution. Our analyses and experiments show that the proposed techniques can reduce the deblocking computation time typically by a factor of more than seven times, while maintaining the bit-exact output.
Jian Lou 0006, Ashish Jagmohan, Dake He, Ligang Lu, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.5
2008 Adaptive rate estimation for H.264/AVC intra mode decision
abstract
In this paper, a modified bit-rate estimation method is proposed to reduce the computation for 4×4 intra mode decision of H.264/AVC video encoder. The number of coded bits is modeled by a linear combination of existing coding parameters, which are highly related to the entropy coding of H.264/AVC. Furthermore, to improve the accuracy of the estimation, the proposed scheme is made adaptive to the information obtained from previously coded blocks. Comparing to the original rate distortion optimized (RDO) encoding process, which needs to calculate the actual encoded bits of H.264/AVC for each coding mode, the proposed adaptive rate estimation can save about 28% and 21% of the total encoding time for QCIF and VGA sequences, respectively. For the coding performance, the proposed method achieves nearly no loss in visual quality with only slight bit-rate increases.
Ke-Ying Liao, Jar-Ferr Yang, Ming-Ting Sun
ICASSP3
2008 Data scaling classification in stream analysis systems
abstract
We formalize data scaling classification (DSC) as a technique to trade the accuracy of classification with the network transmission load in stream analysis frameworks. We apply the proposed data scaling approaches to ECG classification in remote health monitoring systems. Experimental results show satisfactory resource savings for small amounts of utility degradation (e.g., 33% of bandwidth saving for a 3.3% of accuracy degradation).
Ya-Ti Peng, Ching-Yung Lin, Ming-Ting Sun
ICME3
2008 Fast sub-pixel motion estimation and mode decision for H.264
abstract
Motion Estimation (ME) is one of the most time-consuming parts in video coding. The use of multiple partition sizes in H.264 makes the ME even more complicated. It is important to develop fast sub-pixel ME algorithms due to (1) The computation overhead by sub-pixel ME has become relatively significant while the complexity of integer-pel search has been greatly reduced by fast algorithms and (2) Reducing sub-pel searching points can save the computation for interpolating sub-pixel values. In this paper, a new fast sub-pixel ME algorithm is proposed which performs a ‘rough’ sub-pel search before the partition selection and only performs the ‘precise’ sub-pel search for the best partition. Experimental results show that our method can reduce the sub-pel search points by more than 50% compared to existing fast sub-pel ME methods with little quality degradation.
Weiyao Lin, David M. Baylon, Krit Panusopone, Ming-Ting Sun
ISCAS4
2008 Human activity recognition for video surveillance
abstract
This paper presents a novel approach for automatic recognition of human activities from video sequences. We first group features with high correlations into Category Feature Vectors (CFVs). Each activity is then described by a combination of GMMs (Gaussian Mixture Models) with each GMM representing the distribution of a CFV. We show that this approach offers flexibility to add new events and to deal with the problem of lacking training data for building models for unusual events. For improving the recognition accuracy, a Confident-Frame-based Recognizing algorithm (CFR) is proposed to recognize the human activity, where the video frames which have high confidence for recognition an activity (Confident-Frames) are used as a specialized model for classifying the rest of the video frames. Experimental results show the effectiveness of the proposed approach.
Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang
ISCAS2
2008 Complexity and memory efficient GOP structures supporting VCR functionalities in H.264/AVC
abstract
Supporting digital video cassette recording (VCR) trick-play functionalities (e.g. random access, fast-forward play, fast-reverse play) is desirable for compressed video streams. However, due to strong inter-frame dependencies introduced by motion compensated prediction (MCP), the computational complexity and memory requirement is drastically increased. Tradeoffs between coding efficiency and decoding complexity can be made with different group of pictures (GOP) structures. In this paper, we investigate two flexible GOP structures, named G-group and binary reference GOP structure (BRGS), which can achieve trick-play functionalities while keeping low decoder complexity and memory requirement for the state-of-the-art H.264/AVC video coding standard. The schemes are drift-free since they utilize the compression and memory management tools adopted in H.264/AVC. Our analysis and experimental results show that they can greatly reduce the decoder complexity and buffer size while introducing only about 4.0%-7.6% bitrate increase on average. The computational complexity saving for the worst case and average case can be up to 77.8% and 62.5%, respectively. The memory buffer for the fast-reverse play mode can be reduced to 33.3% compared to the conventional scheme. Moreover, the schemes are flexible and can be easily adapted to achieve a good tradeoff between compression performance and complexity saving for the trick-play modes.
Jian Lou 0006, Anthony Vetro, Ming-Ting Sun
ISCAS4
2008 Audio event classification using binary hierarchical classifiers with feature selection for healthcare applications
abstract
In this paper, a binary hierarchical classifier with feature selection is proposed for multi-class audio event classification for healthcare applications. We consider the hierarchical clustering and the feature selection problems jointly when building a binary hierarchical classifier. The proposed method results in the classifier structure as well as a compact feature subset for each component classifier for constructing the overall binary hierarchical classifier. With Support Vector Machine (SVM) for the component classifiers in our experiment, results from classifying several key audio events for the eldercare application show competitive performance to the traditional one-against-one method while the number of training and testing SVM is less in our proposed scheme. Moreover, feature selection facilitates the training of the component classifier by filtering out possible redundant and irrelevant feature components.
Ya-Ti Peng, Ching-Yung Lin, Ming-Ting Sun
ISCAS3
2008 Special Issue on Video Surveillance
abstract
The 14 regular papers and two brief papers in this special issue capture some of the state-of-the-art research on video surveillance issues, provide comprehensive overview of existing techniques, and propose novel solutions for important research problems.
Ishfaq Ahmad 0001, Zhihai He, Hong-Yuan Mark Liao, Fernando Pereira 0001, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.5
2008 Activity Recognition Using a Combination of Category Components and Local Models for Video Surveillance
abstract
This paper presents a novel approach for automatic recognition of human activities for video surveillance applications. We propose to represent an activity by a combination of category components and demonstrate that this approach offers flexibility to add new activities to the system and an ability to deal with the problem of building models for activities lacking training data. For improving the recognition accuracy, a confident-frame-based recognition algorithm is also proposed, where the video frames with high confidence for recognizing an activity are used as a specialized local model to help classify the remainder of the video frames. Experimental results show the effectiveness of the proposed approach.
Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang
IEEE Trans. Circuits Syst. Video Technol.2
2007 High Speed H.264 High Profile Deblocking using Statistical Analysis and Logic Optimization
abstract
In-loop deblocking filter is identified as the most time consuming part for H.264 high profile decoders. This paper proposes an improved platform and encoder independent deblocking scheme for H.264 high profile codec speedup. Two key techniques are introduced in the proposed algorithm: a statistical analysis based hybrid boundary strength derivation scheme and a more efficient logic expression for the B-slice boundary strength derivation. Compared to previously proposed algorithms, significant computation can be saved, while maintaining the bit-exact output. The proposed techniques can be used in both standard conforming encoders and decoders.
Jian Lou 0006, Ashish Jagmohan, Dake He, Ligang Lu, Ming-Ting Sun
ICME5
2007 Statistical Analysis Based H.264 High Profile Deblocking Speedup
abstract
This paper proposes a novel scheme to achieve deblocking speedup for H.264 high profile decoders. The proposed approach is to use statistics dependent decoding which takes advantage of the biased statistical distribution in video streams. Specifically, in the proposed scheme, Huffman tree structures are introduced for boundary strength derivation, and hierarchical true edge detection is applied in the boundary filtering process to reduce the computation. As a result, significant computation can be saved in the deblocking process, while bit-exact output is maintained. This platform and encoder independent scheme can be incorporated into both standard conforming encoders and decoders. Since deblocking accounts for a significant percentage of decoding time, the scheme is especially important for decoder implementations. The analyses and experiments show that the proposed scheme could reduce the deblocking computational load by a factor of more than three times.
Jian Lou 0006, Ashish Jagmohan, Dake He, Ligang Lu, Ming-Ting Sun
ISCAS5
2007 Rate-Distortion Modeling for Efficient H.264/AVC Encoding
abstract
The rate-distortion (R-D) optimization technique plays an important role in optimizing video encoders. Modeling the rate and distortion functions accurately in acceptable complexity helps to make the optimization more practical. In this paper, we propose a bit-rate estimation function and a distortion measure by modeling the transform coefficients with spatial-domain variance. Furthermore, with quantization-based thresholding to determine the number, the absolute sum, and the squared sum of transform coefficients, the simplified transform-domain R-D measurement is introduced. The proposed algorithms can reduce the computation complexity for the R-D optimized mode-decision by using the new cost function evolving from the simplified transform-domain R-D model. Based on the proposed estimations, a rate-control scheme in the macroblock layer is also proposed to improve the coding efficiency
Yu-Kuang Tu, Jar-Ferr Yang, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.3
2006 An Efficient Criterion for Mode Decision in H.264/AVC
abstract
In this paper, an efficient cost function for mode decision in H.264/AVC is proposed. The proposed cost function is based on integer transform coefficients, where the rate and the distortion are jointly modeled by the number of nonzero quantized coefficients, the sum of absolute integer transformed differences (SAITD) and sum of squared integer transformed differences (SSITD). Comparing to the high-complexity cost function, which should be calculated from real bit-consumption and true reconstructed distortion for each coding mode, the proposed efficient cost function can achieve 79.93% and 22.61% time savings of computing rate-distortion cost and overall encoding, respectively, while introducing only slight degradation with 1.05% bit-rate increment and 0.049 dB PSNR drop
Yu-Kuang Tu, Jar-Ferr Yang, Ming-Ting Sun
ICME3
2006 Network condition detection for video transport over wireless Internet
abstract
Real-time video transport over wireless Internet faces many challenges due to the heterogeneous network environment. A robust network condition classification algorithm using multiple end-to-end metrics and support vector machine (SVM) is proposed to classify different network events and model the transition pattern of network conditions. End-to-end quality-of-service (QoS) mechanisms like congestion control, error control, and power control can benefit from the network condition information and react to different network situations appropriately. The proposed network condition classification algorithm uses SVM as a classifier to cluster different end-to-end metrics such as end-to-end delay, delay jitter, throughput and packet loss-rate for the HDP traffic with TCP-friendly rate control (TFRC), which is used for video transport. The algorithm is also flexible to classify different numbers of states representing different levels of network events such as wireline congestion and wireless channel loss. Simulation results using ns2 show the effectiveness of the proposed scheme.
Siu-Ping Chan, Ming-Ting Sun
ISCAS2
2006 Sleep condition inferencing using simple multimodality sensors
abstract
In this paper, we investigate the possibility of using simple multimodality sensors to automatically detect a person's sleep condition. Sleep latency and sleep efficiency are critical to both sleep-related diseases and sleep quality measurements. We propose a system which consists of heart-rate, video, and audio sensors, and apply machine learning methods to infer the sleep-awake condition during the time a user spends on the bed. The sleep-awake conditions will be useful information for inferring the sleep quality. Our experimental results are promising and show the potential use of the proposed novel economical alternative to the traditional medical measurement equipment, with competitive performance on the sleep-related activity monitoring and the sleep quality measurements
Ya-Ti Peng, Ching-Yung Lin, Ming-Ting Sun, Ming-Whei Feng
ISCAS3
2006 Statistical rate-distortion estimation for H.264/AVC coders
abstract
Based on statistical analyses of transform coefficients, we suggest estimations of bit-rate and distortion functions directly from spatial-domain statistics. Simulations show that the proposed rate and distortion estimators achieve similar results as the true rate-quantization (R-Q) and distortion-quantization (D-Q) curves. By exploiting the proposed statistical estimations, two major applications are suggested. One is the fast rate-distortion optimized (RDO) mode decision and the other is the improved rate control (RC) in video encoding. Simulation results show that the rate-distortion estimators have sufficient accuracy such that we could achieve considerably better performance in both applications. We achieved 51.69% R-D cost computation time reduction in average for the rate-distortion optimized mode decision, and a 0.20 dB average peak signal-to-noise ratio (PSNR) improvement for rate-control in H.264/AVC, respectively.
Yu-Kuang Tu, Jar-Ferr Yang, Ming-Ting Sun
ISCAS3
2006 Modeling Evolutionary Behaviors for Community-based Dynamic Recommendation
abstract
We exploit dynamic patterns from both documents' and users' aspects to build models for recommendation. We propose a Community-Based Dynamic Recommendation (CBDR) scheme to make recommendations by taking content semantics, evolutionary patterns, and user communities into consideration. A Time-Sensitive Adaboost algorithm is proposed to build adaptive user models for ranking document candidates based on leveraging dynamic factors such as freshness, popularity, and other attributes. Our experimental results on a large online application system demonstrate the recommendation usefulness of the CBDR scheme is 259% better than the collaborative filtering, 126% better than the community-based static recommendation algorithm, and 106% better than the optimal global recommendation bound.
Xiaodan Song, Ching-Yung Lin, Belle L. Tseng, Ming-Ting Sun
SDM4
2006 Personalized recommendation driven by information flow
abstract
We propose that the information access behavior of a group of people can be modeled as an information flow issue, in which people intentionally or unintentionally influence and inspire each other, thus creating an interest in retrieving or getting a specific kind of information or product. Information flow models how information is propagated in a social network. It can be a real social network where interactions between people reside; it can be, moreover, a virtual social network in that people only influence each other unintentionally, for instance, through collaborative filtering. We leverage users' access patterns to model information flow and generate effective personalized recommendations. First, an early adoption based information flow (EABIF) network describes the influential relationships between people. Second, based on the fact that adoption is typically category specific, we propose a topic-sensitive EABIF (TEABIF) network, in which access patterns are clustered with respect to the categories. Once an item has been accessed by early adopters, personalized recommendations are achieved by estimating whom the information will be propagated to with high probabilities. In our experiments with an online document recommendation system, the results demonstrate that the EABIF and the TEABIF can respectively achieve an improved (precision, recall) of (91.0%, 87.1%) and (108.5%, 112.8%) compared to traditional collaborative filtering, given an early adopter exists.
Xiaodan Song, Belle L. Tseng, Ching-Yung Lin, Ming-Ting Sun
SIGIR4
2006 FGS enhancement layer truncation with reduced intra-frame quality variation
Huai-Rong Shao, Ming-Ting Sun
J. Vis. Commun. Image Represent.3
2006 Fast motion estimation and Inter-mode decision for H.264/MPEG-4 AVC encoding
Zhi Zhou 0010, Jun Xin, Ming-Ting Sun
J. Vis. Commun. Image Represent.3
2006 Fast multiple reference frame motion estimation for H.264/AVC
abstract
Multiple reference frame motion compensation is a new feature introduced in H.264/MPEG-4 AVC to improve video coding performance. However, the computational cost of multiple reference frame motion estimation (MRF-ME) is very high. In this paper, we propose an algorithm that takes into account the correlation/continuity of motion vectors among different reference frames. We show that the algorithm effectively reduces the computations of MRF-ME, and achieves similar coding gain compared to the motion search approaches in the reference software.
Yeping Su, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.2
2006 Efficient rate-distortion estimation for H.264/AVC coders
abstract
In video coders, the optimal coding mode decision for each coding block can be achieved by exhaustively calculating the rate-distortion cost, which simultaneously considers the distortion performance and the coding bit consumption of all possible modes. The best mode is chosen from the one with the minimum Lagrange cost. To avoid the expensive computation of Lagrange costs, in this paper, we propose transform-domain bit-rate estimation and distortion measures, based on quantized and inverse quantized integer transform coefficients, for the inter-mode decision in H.264/AVC coders. With the proposed scheme, entropy coding, inverse transform, and pixel-reconstructions are not required in the process. With ignorable degradation in coding performance, simulations demonstrate that the proposed estimation method achieves about 40% reduced computation time of rate-distortion cost for the inter-mode decision and saves about 17% total encoding time while combining with fast motion estimation and fast mode decision algorithms.
Yu-Kuang Tu, Jar-Ferr Yang, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.3
2005 A key management scheme in distributed sensor networks using attack probabilities
abstract
Clustering approaches have been found useful in providing scalable data aggregation, security and coding for large scale distributed sensor networks (DSNs). Clustering (also known as subgrouping) has also been effective in containing and compartmentalizing node compromise in large scale networks. We consider the problem of designing a clustered DSN when the probability of node compromise in different deployment regions is known a priori. We make use of the a priori probability to design a variant of random key predistribution method that improves the resilience and hence the fraction of compromised communications compared to seminal works. We further relate the key ring size of the subgroup node to the probability of node compromise, and design an effective scalable security mechanism that increases the resilience to the attacks for the sensor subgroups. Simulation results show that by using our scheme, the performance can be substantially improved in the sensor network (including the resilience and the fraction of compromised communications) that only sacrifices a small extent in the probability of a shared key exists between two nodes, compared to those of the prior results.
Siu-Ping Chan, Radha Poovendran, Ming-Ting Sun
GLOBECOM3
2005 Speech-Based Visual Concept Learning Using Wordnet
abstract
Modeling visual concepts using supervised or unsupervised machine learning approaches are becoming increasing important for video semantic indexing, retrieval, and filtering applications. Naturally, videos include multimodality data such as audio, speech, visual and text, which are combined to infer therein the overall semantic concepts. However, in the literature, most researches were conducted within only one single domain. In this paper we propose an unsupervised technique that builds context-independent keyword lists for desired visual concept modeling using WordNet. Furthermore, we propose an Extended Speech-based Visual Concept (ESVC) model to reorder and extend the above keyword lists by supervised learning based on multimodality annotation. Experimental results show that the context-independent models can achieve comparable performance compared to conventional supervised learning algorithms, and the ESVC model achieves about 53% and 28.4% improvement in two testing subsets of the TRECVID 2003 corpus over a state-of-the-art speech-based video concept detection algorithm.
Xiaodan Song, Ching-Yung Lin, Ming-Ting Sun
ICME3
2005 Rate-distortion estimation for H.264/AVC coders
abstract
In a video coder, the optimal coding mode decision for each coding block could be achieved by exhaustively calculating the Lagrange cost (which includes the coding distortion plus the Lagrange parameter times the coding bit consumption) of all possible modes. The best mode can then be chosen as the one with the minimum Lagrange cost. To speed up the computationally intensive Lagrange cost computation, in this paper, we propose transform-domain bit-rate estimation and distortion measures for the inter-mode decision in H.264/AVC coders. With the proposed scheme, entropy coding, inverse DCT, and pixel-reconstructions are not required in the process. Simulation results show that the proposed estimation method is accurate for the inter-mode decision and about 46.42% time reduction can be achieved.
Yu-Kuang Tu, Jar-Ferr Yang, Ming-Ting Sun
ICME3
2005 Modeling and predicting personal information dissemination behavior
abstract
In this paper, we propose a new way to automatically model and predict human behavior of receiving and disseminating information by analyzing the contact and content of personal communications. A personal profile, called CommunityNet, is established for each individual based on a novel algorithm incorporating contact, content, and time information simultaneously. It can be used for personal social capital management. Clusters of CommunityNets provide a view of informal networks for organization management. Our new algorithm is developed based on the combination of dynamic algorithms in the social network field and the semantic content classification methods in the natural language processing and machine learning literatures. We tested CommunityNets on the Enron Email corpus and report experimental results including filtering, prediction, and recommendation capabilities. We show that the personal behavior and intention are somewhat predictable based on these models. For instance, "to whom a person is going to send a specific email" can be predicted by one's personal social network and content analysis. Experimental results show the prediction accuracy of the proposed adaptive algorithm is 58% better than the social network-based predictions, and is 75% better than an aggregated model based on Latent Dirichlet Allocation with social network enhancement. Two online demo systems we developed that allow interactive exploration of CommunityNet are also discussed.
Xiaodan Song, Ching-Yung Lin, Belle L. Tseng, Ming-Ting Sun
KDD4
2005 Special Issue on Advances in Video Coding and Delivery
Wenwu Zhu 0001, Ming-Ting Sun, Liang-Gee Chen, Thomas Sikora
Proc. IEEE2
2005 Digital Video Transcoding
abstract
Video transcoding, due to its high practical values for a wide range of networked video applications, has become an active research topic. We outline the technical issues and research results related to video transcoding. We also discuss techniques for reducing the complexity, and techniques for improving the video quality, by exploiting the information extracted from the input video bit stream.
Jun Xin, Chia-Wen Lin, Ming-Ting Sun
Proc. IEEE3
2005 Fast variable-size block motion estimation for efficient H.264/AVC encoding
Yu-Kuang Tu, Jar-Ferr Yang, Ming-Ting Sun, Yuesheng T. Tsai
Signal Process. Image Commun.3
2005 Global motion estimation from coarsely sampled motion vector field and the applications
abstract
Global motion estimation is a powerful tool widely used in video processing and compression as well as in computer vision areas. We propose a new approach for estimating global motions from coarsely sampled motion vector fields. The proposed method minimizes the fitting error between the input motion vectors and the motion vectors generated from the estimated motion model using the Newton-Raphson method with outlier rejections. Applications of the proposed method in video coding include fast global motion estimation for MPEG-4 Advanced Simple Profile coding, MPEG-2 to MPEG-4 ASP transcoding, and error concealments. Simulation results and analyses are provided for the proposed method and the applications, which show the effectiveness of the method in terms of accuracy, robustness, and speed.
Yeping Su, Ming-Ting Sun, Vincent Hsu
IEEE Trans. Circuits Syst. Video Technol.2
2004 Fast macroblock inter mode decision and motion estimation for H.264/MPEG-4 AVC
abstract
In H.264/MPEG-4 AVC, MacroBlock (MB) mode decision and motion estimation (ME) is one of the most computationally expensive processes. This paper proposes a fast intermode decision and ME algorithm. Experimental results show that the proposed algorithm could achieve similar Rate-Distortion (R-D) performance with about 50% computation saving, compared to the JM low-complexity mode with a Fast Full-Search (FFS) motion estimation algorithm.
Zhi Zhou 0010, Ming-Ting Sun
ICIP2
2004 Fast multiple reference frame motion estimation for H.264
abstract
Multiple reference frame motion compensation is a new feature introduced in H.264/MPEG-4 AVC to improve video coding performance. However, the computational cost of multiple reference frame motion estimation (MRF-ME) is very high. We propose an algorithm that takes into account the correlation/continuity of motion vectors among different reference frames. We also show that the algorithm effectively reduces the computations of MRF-ME, and achieves similar coding gain compared to the full-search approach.
Yeping Su, Ming-Ting Sun
ICME2
2004 A non-iterative motion vector based global motion estimation algorithm
abstract
Global motion compensation (GMC) is a coding tool introduced in MPEG-4 advanced simple profile (ASP); it can improve the coding efficiency on video sequences with global motion. Conventional pixel-domain iterative global motion estimation (GME) algorithms for estimating the perspective transform model parameters are very computationally demanding, and thus its applicability is restricted. A fast non-iterative GME algorithm is proposed for estimating the perspective transform global motion parameters from the motion vectors (MV) obtained from the block matching process. The proposed algorithm first fits multiple global motion models from groups of input MVs, then, a robust estimate is formed by a histogram-based postprocessing. Simulation results show the proposed algorithm can achieve superior coding performance while using much fewer computations compared to the iterative approaches.
Yeping Su, Ming-Ting Sun
ICME2
2004 Early-stop and motion vector reuse for MPEG-2 to H.264 transcoding
abstract
In this paper, early-stop and Motion Vector (MV) re-use approaches are proposed for the MPEG-2 to H.264 transcoding to reduce the computation of the variable block-size motion estimation. By combining the two approaches, the number of MV search points is reduced by more than 80% without significantly affecting the video quality. The proposed approaches can also be used in fast variable block-size motion estimation for the H.264 video encoding.
Mehmet Kucukgoz, Ming-Ting Sun
VCIP2
2003 Fast variable-size block motion estimation using merging procedure with an adaptive threshold
abstract
A merging procedure for variable-size block motion estimation is proposed to reduce the computation for block-size decision. A smaller block-size is initially used for motion estimation. An adaptive threshold, which depends on the information obtained from the motion estimation, quantization parameter, and rate distortion cost function is used to determine if the motion vectors of the neighboring blocks should be merged or not. We simulate the proposed merging procedure using an H.264 video encoder to determine if 16x16, 16x8, 8x16, or 8x8 blocksize should be used for each macroblock. Based on the simulations, performance of the proposed method is close to that of H.264 JM2.1 when 16x16, 16x8, 8x16 and 8x8 block modes are enabled and the exhaustive search method is used to determine the block-sizes. Computational complexity can be significantly reduced by the proposed algorithm since it only performs the motion search for one block-size.
Yu-Kuang Tu, Jar-Ferr Yang, Yi-Nung Shen, Ming-Ting Sun
ICME4
2003 Diversity-based fast block motion estimation
abstract
Many fast search strategies reduce the complexity of motion estimation by limiting search locations, including the center-biased diamond search (DS), which performances better for small motion, and the nonbiased three step search (TSS), which works better for large motion. To achieve improved performance, we propose a diversity search strategy (DSS) by combining DS and TSS and applying them to decimated block matching error functions. The diversity in search strategies and self-similarity of the decimated error functions can overcome the disadvantages of individual search without increasing the computational complexity. Extensive simulations show that DSS approaches the performance of full-search and outperforms individual strategies. Based on DSS, and efficient adaptive DSS (ADSS) is proposed. It is shown to outperform DS and TSS, both in quality and complexity. Finally, a fast motion estimation using diversity in matching criteria (DMC) is presented, with potential applications for variable-block-size motion estimation.
Jun Xin, Ming-Ting Sun, Vincent Hsu
ICME2
2003 FGS enhancement layer truncation with minimized intra-frame quality variation
abstract
This paper proposes an enhancement layer truncation scheme for the fine-granularity-scalability (FGS) video. Our target is to minimize the quality variation of different parts within each frame when the last transmitted enhancement layer is truncated according to the available network bandwidth. We propose to redistribute the bits in the enhancement layer that can only be partially kept with the available bit-budget so that it is able to cover the whole frame area to raise the quality of different parts uniformly. Simulation results confirm the effectiveness of the proposed method in improving the decoded visual quality and reducing the intra-frame quality variation.
Huai-Rong Shao, Chia Shen, Ming-Ting Sun
ICME4
2003 Finding structure in home videos by probabilistic hierarchical clustering
abstract
Accessing, organizing, and manipulating home videos present technical challenges due to their unrestricted content and lack of storyline. We present a methodology to discover cluster structure in home videos, which uses video shots as the unit of organization, and is based on two concepts: (1) the development of statistical models of visual similarity, duration, and temporal adjacency of consumer video segments and (2) the reformulation of hierarchical clustering as a sequential binary Bayesian classification process. A Bayesian formulation allows for the incorporation of prior knowledge of the structure of home video and offers the advantages of a principled methodology. Gaussian mixture models are used to represent the class-conditional distributions of intra- and inter-segment visual and temporal features. The models are then used in the probabilistic clustering algorithm, where the merging order is a variation of highest confidence first, and the merging criterion is maximum a posteriori. The algorithm does not need any ad-hoc parameter determination. We present extensive results on a 10-h home-video database with ground truth which thoroughly validate the performance of our methodology with respect to cluster detection, individual shot-cluster labeling, and the effect of prior selection.
Daniel Gatica-Perez, Alexander C. Loui, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.3
2003 Dynamic region of interest transcoding for multipoint video conferencing
abstract
This paper presents a region of interest transcoding scheme for multipoint video conferencing to enhance visual quality. In a multipoint video conference, usually there are only one or two active conferees at one time, which are the regions of interest to the other conferees involved. We propose a dynamic sub-window skipping scheme to firstly identify the active participants from the multiple incoming encoded video streams by calculating the motion activity of each sub-window and then dynamically reduce the frame rates of the motion inactive participants by skipping these less-important sub-windows. The bits saved from the skipping operation are reallocated to the active sub-windows to enhance the regions of interest. We also propose a low-complexity scheme to compose, as well as trace, the unavailable motion vectors with a good accuracy in the dropped inactive sub-windows after performing sub-window skipping. Simulation results show that the proposed methods not only significantly improve the visual quality of the active sub-windows without introducing serious visual quality degradation in the inactive ones, but also reduce the computational complexity and avoid whole-frame skipping. Moreover, the proposed algorithm is fully compatible with the H.263 video coding standard.
Chia-Wen Lin, Yung-Chang Chen, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.3
2002 Probabilistic home video structuring: feature selection and performance evaluation
abstract
We previously proposed a method to find the cluster structure in home videos based on statistical models of visual and temporal features of video segments and sequential binary Bayesian classification. In this paper, we present analysis and improved results on two key issues: feature selection and performance evaluation, using a ten-hour database (30 video clips, 1,075,000 frames). From multiple features and similarity measures, visual features are selected in order to minimize the empirical probability of misclassification. Temporal features are chosen to reflect the patterns existing in both shot and cluster duration and adjacency. Finally, we describe a detailed performance evaluation procedure that includes cluster detection, individual shot-cluster labeling, and prior selection.
Daniel Gatica-Perez, Alexander C. Loui, Ming-Ting Sun
ICIP (1)3
2002 Linking objects in videos by importance sampling
abstract
We present an approach to create hyper-links between video segments that contain objects of interest, based on video structuring, object definition, and stochastic object localization in the video structure. Localization is formulated in the metric mixture model framework, which allows for the joint probabilistic modeling of a (user-defined) set of color appearance exemplars and their geometric transformations. Candidate object configurations are drawn from a prior distribution using importance sampling - which guides the search towards regions of the configuration space likely to contain the correct object configuration, thus avoiding exhaustive processing - and evaluated using Bayes' rule. Results of linking real objects (with changes of size and pose) in several home videos illustrate the performance of the method.
Daniel Gatica-Perez, Ming-Ting Sun
ICME (2)2
2002 Bit-allocation for transcoding pre-encoded video streams
Jun Xin, Ming-Ting Sun, Kang Wook Chun
VCIP2
2002 Wireless video transport using conditional retransmission and low-delay interleaving
abstract
We consider the scenario of using Automatic Repeat reQuest (ARQ) retransmission for two-way low-bit-rate video communications over wireless Rayleigh fading channels. Low-delay constraint may require that a corrupted retransmitted packet not be retransmitted again, and thus there will be packet errors at the decoder which results in video quality degradation. We propose a scheme to improve the video quality. First, we propose a low-delay interleaving scheme that uses the video encoder buffer as a part of interleaving memory. Second, we propose a conditional retransmission strategy that reduces the number of retransmissions. Simulation results show that our proposed scheme can effectively reduce the number of packet errors and improve the channel utilization. As a result, we reduce the number of skipped frames and obtain a peak signal-to-noise ratio improvement up to about 4 dB compared to H.263 TMN-8.
Supavadee Aramvith, Chia-Wen Lin, Sumit Roy 0001, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.4
2002 An HDTV-to-SDTV spatial transcoder
abstract
Both high-definition television (HDTV) and, standard-definition television (SDTV) use the MPEG-2 video coding standard, but they have different spatial resolutions. In order to support the interlaced video coding, MPEG-2 incorporates various macroblock prediction modes. Thus, the HDTV-to-SDTV transcoding needs to handle spatial resolution downscaling and various MEPG-2 macroblock prediction modes. We investigate schemes to exploit the correlations between the input and output video in the design of an HDTV-to-SDTV transcoder so that the computation can be greatly saved while the quality of video is preserved as much as possible. First, by utilizing the motion vectors and macroblock coding modes of the input video, efficient motion reestimation and macroblock mode decision algorithms are proposed. Then a novel picture target bit allocation algorithm taking advantage of the coding statistics of the input video is presented. Simulation results showing the effectiveness of the proposed approaches are also presented.
Jun Xin, Ming-Ting Sun, Byung Sun Choi, Kang Wook Chun
IEEE Trans. Circuits Syst. Video Technol.2
2001 Consumer Video Structuring by Probabilistic Merging of Video Segments
abstract
Accessing, organizing, and manipulating home videos constitutes a technical challenge due to their unrestricted content and the lack of storyline. In this paper, we present a methodology for structuring consumer video, based on the development of statistical models of similarity and adjacency between video segments in a probabilistic formulation. Learned Gaussian mixture models of inter-segment visual similarity, temporal adjacency, and segment duration are used to represent the classconditional densities of observed features. Such models are then used in a sequential merging algorithm consisting of a binary Bayes classifier, where the merging order is determined by a variation of Highest Confidence First (HCF), and the merging criterion is Maximum a Posteriori (MAP). The merging algorithm can be efficiently implemented and does not need any empirical parameter determination. Finally, the representation of the merging sequence by a tree provides for hierarchical, nonlinear access to the video content. Results on an eight-hour home video database illustrate the validity of our approach.
Daniel Gatica-Perez, Ming-Ting Sun, Alexander C. Loui
ICME2
2001 Implementation Of A Realtime Object-Based Virtual Meeting System
abstract
This paper presents an H.323 standard compliant video conferencing system implementation. The proposed system not only serves as an MCU (Multipoint Control Unit) for multipoint connection but also provides a gateway function between the H.323 LAN (Local Area Network) and the H.324 WAN (Wide Area Network) users. The proposed video conferencing system provides user-friendly object compositing and manipulation features including 2-D video object scaling, re-positioning, rotating, and dynamic bit-allocation in a 3-D virtual environment. A segmentation scheme based on pre- stored background information is proposed for real-time segmentation of the foreground video objects at the client side. Chroma-key insertion is used to facilitate video objects extraction and manipulation. We have implemented the virtual conference system prototype with an integrated graphic user interface to demonstrate the feasibility of the proposed methods.
Chia-Wen Lin, Yao-Jen Chang, Yung-Chang Chen, Ming-Ting Sun
ICME4
2001 Bit Allocation For Joint Transcoding Of Multiple Mpeg Coded Video Streams
abstract
For the transcoding of pre-encoded video streams, certain useful picture statistics can be extracted from the input bit-streams to help bit-allocation. In this paper, a novel approach to estimate the picture-complexity of the transcoded output video streams from the input streams is proposed. Based on the picturecomplexity estimation, a bit-allocation strategy is presented for joint transcoding of multiple pre-encoded MPEG video streams. Simulation results show that our proposed algorithm can achieve better video quality compared to previously reported algorithms.
Jun Xin, Ming-Ting Sun, KouSou Kan
ICME2
2001 Encoding DCT Coefficients Based on Rate-Distortion Measurement
I-Ming Pao, Ming-Ting Sun, Shawmin Lei
J. Vis. Commun. Image Represent.2
2001 A rate-control scheme for video transport over wireless channels
abstract
We investigate the scenario of using the automatic repeat request (ARQ) retransmission scheme for two-way video communications over wireless Rayleigh fading channels. Video quality is the major concern of these applications. We show that, during the retransmissions of error packets, due to the reduced channel throughput, the video encoder buffer may fill-up quickly and cause the TMN8 rate-control algorithm to significantly reduce the bits allocated to each video frame. This results in PSNR degradation and many skipped frames. To minimize the number of frames skipped, we propose improved rate-control schemes that take into consideration the effects of the video buffer fill-up, an a priori channel model, and the channel feedback information. We also show that a discrete cosine transform (DCT) coefficient soft-thresholding scheme can be applied to further improve video quality. As a result, our proposed rate-control schemes encode the video sequences with less frame skipping and with higher PSNR compared to TMN8.
Supavadee Aramvith, I-Ming Pao, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.3
2001 Semantic video object extraction using four-band watershed and partition lattice operators
abstract
We conceive the problem of multiple semantic video object (SVO) extraction as an issue of designing extensive operators on a complete lattice of partitions. As a result, we propose a framework based on spatial partition generation and application of optimal operators on the generated partitions. Based on a statistical analysis of the watershed algorithm, we develop a multivalued morphological spatial segmentation method that incorporates an edge-driven marker extraction algorithm and a growing method which integrates both color and edge information. Having embedded the problem in the partition lattice framework, we propose a spatio-temporal regional maximum likelihood operator for extraction purposes. Some theoretical properties of the operator are established. Experimental results on several MPEG-4 test video sequences show that our scheme improves the precision of the extracted SVO boundaries compared to traditional watershed algorithms and provides accurate tracking of multiple SVOs in both static and moving camera scenarios. Furthermore, this scheme can be extended to deal with more general interactive video authoring systems.
Daniel Gatica-Perez, Chuang Gu, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.3
2001 Multiview extensive partition operators for semantic video object extraction
abstract
Occlusion/disocclusion is one of the fundamental problems for semantic video object (SVO) extraction, where pixel-wise accuracy is required. This issue is critical because the degradation in tracking due to object occlusion/disocclusion significantly increases the amount of user interaction required in off-line video editing applications. We present an approach based on the application of an extensive operator on a lattice of partitions, which exploits information from various views of the scene, based on a probabilistic formulation. Our multiview operator builds on the regional application of the maximum a posteriori principle, by integrating a single-view region classification stage with a multiview stage that improves classification for those disoccluded regions labeled as uncertain. Results on several real sequences show that our approach improves the SVO tracking compared to the single-view case and that, as a result, increases the quality of the extracted SVOs and reduces the total amount of user interaction.
Daniel Gatica-Perez, Ming-Ting Sun, Chuang Gu
IEEE Trans. Circuits Syst. Video Technol.2
2001 MPEG video streaming with VCR-functionality
abstract
With the proliferation of online multimedia content, the popularity of multimedia streaming technology, and the establishment of MPEG video coding standards, it is important to investigate how to efficiently implement an MPEG video streaming system. Digital video cassette recording (VCR) functionality (such as random access, fast forward, fast reverse, etc.) enables quick and user-friendly browsing of multimedia content, and thus is highly desirable in streaming video applications. The implementation of full VCR functionality, however, presents some technical challenges that have not yet been well resolved. We investigate the impacts of the VCR functionality on the network traffic and the video decoder complexity. We propose a least-cost scheme for the efficient implementation of MPEG streaming video system to provide full VCR functionality over a network with minimum requirements on the network bandwidth and the decoder complexity. We also discuss our implementation of an IP-based MPEG-4 video streaming platform which provides full VCR-functionality.
Chia-Wen Lin, Jeongnam Youn, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.4
2001 Encoding stored video for streaming applications
abstract
In streaming video applications, video sequences are encoded off-line and stored in a server. Users may access the server over a constant bit rate channel. Examples of the streaming video applications are video-on-demand, archived video news, and noninteractive distance learning. Before the playback, part of the video bitstream is pre-loaded in the decoder buffer to ensure that every frame can be decoded at the scheduled time. For these streaming video applications, since the video is encoded off-line and the future video frames are available to the encoder, a more sophisticated bit-allocation scheme can be used to achieve better video quality. During the encoding process for streaming video, two requirements need to be considered: the pre-loading time that the video viewers have to wait and the physical buffer-size at the receiver (decoder) side. In this paper, we propose a sliding-window rate-control scheme that uses statistical information of the future video frames as a guidance to generate better video quality for video streaming involving constant bit rate channels. A quantized discrete cosine transform coefficient selection scheme based on the rate-distortion measurement is also used to improve the video quality. Simulation results show video quality improvements over the regular H.263 TMN8 encoder.
I-Ming Pao, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.2
2001 Extensive partition operators, gray-level connected operators, and region merging/classification segmentation algorithms: theoretical links
abstract
The relation between morphological gray-level connected operators and segmentation algorithms based on region merging/classification strategies has been pointed out several times in the literature. However, to the best of our knowledge, the formal relation between them has not been established. This paper presents the link between the two domains based on the observation that both connected operators and segmentation algorithms share a key mechanism: they simultaneously operate on images and on partitions, and therefore they can be described as operations on a joint image-partition model. As a result, we analyze both segmentation algorithms and connected operators by defining operators on complete product lattices, that explicitly model gray-level and partition attributes. In the first place, starting with a complete lattice of partitions, we initially define the concept of the segmentation model as a mapping in a product lattice, whose elements are three-tuples consisting of a partition, an image that models the partition attributes, and an image that represents the gray-level model associated to the segmentation. Then, assuming a conditional ordering relation, we show that any region merging/classification segmentation algorithm can be defined as an extensive operator in such a complete product lattice, in the second place, we proposed a very similar lattice-based extended representation of gray-level functions in the context of connected operators, that highlights the mathematical analogy with segmentation algorithms, but in which the ordering relation is different. We use this framework to show that every region merging/classification segmentation algorithm indeed corresponds to a connected operator. While this result provides an explanation to previous work in the area, it also opens possibilities for further analysis in the two domains. From this perspective, we additionally study some theoretical properties of a general region merging segmentation algorithm.
Daniel Gatica-Perez, Chuang Gu, Ming-Ting Sun, Salvador Ruiz-Correa
IEEE Trans. Image Process.3
2000 Generating Video Objects by Multiple-View Extensive Partition Lattice Operators
abstract
Occlusion/disocclusion is one of the fundamental problems for semantic video object (SVO) extraction because the tracking degradation from occlusion/disocclusion creates difficulty for the pixel-wise accuracy requirement and adds to the amount of user interaction required in off-line video editing applications. We present an approach based on the application of a new extensive operator in a lattice of partitions, that exploits information from various views of the scene to address the disocclusion problem by using a regional Bayesian formulation. Our multiview operator builds on the application of the maximum a posteriori (MAP) principle, by integrating a single view region classification stage and a multiview stage that solves for disoccluded regions labeled as uncertain. Results on several real sequences show that this approach improves the quality of the extracted SVOs and reduces the total amount of user interaction.
Daniel Gatica-Perez, Ming-Ting Sun, Chuang Gu
ICIP2
2000 Semiautomatic video object generation using multivalued watershed and partition lattice operators
abstract
We conceive the problem of multiple semantic wide object (SVO) extraction as an issue of designing extensive operators on the lattice of partitions. As a result, we propose a framework based on spatial partition generation and application of optimal operators on the generated partitions. The first stage is obtained with a multivalued morphological spatial segmentation method that incorporates an edge-driven marker extraction algorithm and a growing method which integrates both color and edge information. For the second stage, we propose a new spatio-temporal regional maximum likelihood partition operator for extraction purposes. Subjective and objective evaluation of the experimental results obtained with our approach on several MPEG-4 test video sequences show that it accurately tracks multiple SVOs in several scenarios, while improving the SVO extraction precision compared to traditional watershed techniques.
Daniel Gatica-Perez, Ming-Ting Sun, Chuang Gu
ISCAS2
2000 Fast video transcoding architectures for networked multimedia applications
abstract
In various networked multimedia applications, it is often necessary to change the bit-rate and format of a pre-encoded bit-stream. This can be achieved using a cascaded pixel-domain transcoder, which fully decodes an incoming bit-stream and then re-encodes the decoded pictures with the desired bit-rate or format. However, the cascaded pixel-domain transcoder is computationally expensive. To reduce the computations, several fast architectures have been proposed in the literature. However, these fast transcoder architectures introduce new limitations and are not drift-free. In this paper, we propose new techniques to implement a fast cascaded pixel-domain transcoder. We further discuss the limitation and speed of the different transcoder architectures. We also discuss the methods for adding watermark or company logo, using the different transcoder architectures.
Jeongnam Youn, Jun Xin, Ming-Ting Sun
ISCAS3
2000 Coding scheme for wireless video transport with reduced frame skipping
Supavadee Aramvith, Ming-Ting Sun
VCIP2
2000 Video transcoding for multiple clients
Jeongnam Youn, Jun Xin, Ming-Ting Sun, Ya-Qin Zhang
VCIP3
2000 Region-Based Video Coding Using a Geometric Motion Compensation
Chih-Shoung Huang, Yinyi Lin, Ming-Ting Sun
J. Vis. Commun. Image Represent.3
2000 Video Transcoding with H.263 Bit-Streams
Jeongnam Youn, Ming-Ting Sun
J. Vis. Commun. Image Represent.2
1999 Semantic Video Object Extraction Based on Backward Tracking of Multivalued Watershed
abstract
We present a novel algorithm for semantic video object (SVO) extraction, based on a new multivalued morphological spatial segmentation that integrates color and edge information, and object tracking by backward region classification. Our proposed spatial segmentation incorporates a new marker extraction method based on intensity edge information that improves the definition of the real borders of the scene objects, and a new distance criterion based on color and edge information to guide the watershed algorithm. Experimental results on several MPEG-4 test video sequences show that our algorithm improves the precision of the extracted SVO boundaries compared to the traditional water-shed technique, and that it is capable of tracking multiple SVOs in static and moving camera scenarios.
Daniel Gatica-Perez, Ming-Ting Sun, Chuang Gu
ICIP (2)2
1999 Video transcoder architectures for bit rate scaling of H.263 bit streams
abstract
Video transcoding is one of the key technologies in implementing dynamic adaptation of the bit rate of a pre-encoded video stream to the available bandwidth over various networks. Many different transcoder architectures have been proposed to achieve fast processing. However, they suffer from quality degradation due to the drift error. In this paper, we investigate the drift caused by the fast transcoder architectures for transcoding H.263 bitstreams. We also discuss the limitations of the existing fast transcoder architectures and the flexibility that can be offered by a cascaded pixel-domain transcoder. In favor of drift-free performance and transcoding flexibility, we propose methods to reduce the computational complexity of the cascaded pixel-domain transcoder.
Jeongnam Youn, Ming-Ting Sun, Jun Xin
ACM Multimedia (1)2
1999 Special Issue On Representation And Coding Of Images And Video II
King Ngi Ngan, Sethuraman Panchanathan, Thomas Sikora, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.4
1999 Modeling DCT coefficients for fast video encoding
abstract
Digital video coding standards such as H.263 and MPEG are becoming more and more important for multimedia applications. Due to the huge amount of computations required, there are significant efforts to speed up the processing of video encoders. Previously, the efforts were mainly focused on the fast motion-estimation algorithm. However, as the motion-estimation algorithm becomes optimized, to speed up the video encoders further we also need to optimize other functions such as the discrete cosine transform (DCT) and inverse DCT (IDCT). In this paper, we propose a theoretical model for DCT coefficients. Based on the model, we develop an adaptive algorithm to reduce the computations of DCT, IDCT, quantization, and inverse quantization. We also present a fast DCT algorithm to speed up the calculations of DCT further when the quantization step size is large. We show, by simulations, that significant improvement in the processing speed can be achieved with negligible video-quality degradation, We also implement the algorithm in a real-time PC-based platform to show that it is effective and practical.
I-Ming Pao, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.2
1999 Motion Vector Refinement for High-Performance Transcoding
abstract
In transcoding, simply reusing the motion vectors extracted from an incoming video bit stream may not result in the best quality. In this paper, we show that the incoming motion vectors become nonoptimal due to the reconstruction errors. To achieve the best video quality possible, a new motion estimation should be performed in the transcoder. We propose a fast-search adaptive motion vector refinement scheme that is capable of providing video quality comparable to that can be achieved by performing a new full-scale motion estimation but with much less computation. We discuss the case when some incoming frames are dropped for frame-rate conversions, and propose motion vector composition method to compose a motion vector from the incoming motion vectors. The composed motion vector can also be refined using the proposed motion vector refinement scheme to achieve better results.
Jeongnam Youn, Ming-Ting Sun, Chia-Wen Lin
IEEE Trans. Multim.2
1998 Adaptive Motion Vector Refinement for High Performance Transcoding
Jeongnam Youn, Ming-Ting Sun
ICIP (3)2
1998 Statistical Computation of Discrete Cosine Transform in Video Encoders
Ming-Ting Sun, I-Ming Pao
J. Vis. Commun. Image Represent.1
1998 Guest Editorial
King Ngi Ngan, Sethuraman Panchanathan, Thomas Sikora, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.4
1998 Special Issue On Representation And Coding Of Images And Video I [Guest Editorial]
King Ngi Ngan, Sethuraman Panchanathan, Thomas Sikora, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.4
1998 Approximation of calculations for forward discrete cosine transform
abstract
This paper presents new schemes to reduce the computation of the discrete cosine transform (DCT) with negligible peak-signal-to-noise ratio (PSNR) degradation. The methods can be used in the software implementation of current video standard encoders, for example, H.26x and MPEG. We investigated the relationship between the quantization parameters and the position of the last nonzero DCT coefficient after quantization. That information is used to adaptively make the decision of calculating all 8/spl times/8 DCT coefficients or only part of the coefficients. To further reduce the computation, instead of using the exact DCT coefficients, we propose a method to approximate the DCT coefficients which leads to significant computation savings. The results show that for practical situations, significant computation reductions can be achieved while causing negligible PSNR degradation. The proposed method also results in computation savings in the quantization calculations.
I-Ming Pao, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.2
1997 Video Browsing for Course-on-Demand in Distance Learning
abstract
The rapid evolution of communication technologies, contributed by the advanced data networks, which enable the integration of multimedia information processing and have greatly impacted our approach to distance learning. Several new multimedia features are developed and reported in this paper. These features allow students to browse the lecture contents directly on digital video sequences while also providing access to multimedia information like lecture captions, databases or handbooks with additional information or another video sequence to explain the current contents. Using these multimedia features, students can learn the course more effectively as compared to learning from conventional analog media in the distance learning environments.
Jenq-Neng Hwang, Jeongnam Youn, Sachin G. Deshpande, Ming-Ting Sun
ICIP (2)4
1997 Multimedia features for course-on-demand in distance learning
abstract
The rapid evolution of communication technologies, contributed by the advanced data networks and the integration of multimedia information processing, has greatly impacted our approach to distance learning. Most existing research work emphasizes the communication aspects of multimedia roles in distance learning [3, 4]. On the other hand, we are initiating a new multimedia research activity which emphasises the development of tools for creating constructive multimedia features for course-on-demand in distance learning. In this research, we investigated the development of a course-on-demand system which allows students to access a multimedia database containing recorded video courses over various mechanisms, such as TCP/IP LANs, PSTN, and ISDN. The system is based on a client-server model. The recorded courses are stored in the server which allows multiple clients to simultaneously access the multimedia database over the networks. Students can access these pre-recorded sequences through Web browsers (e.g., Netscape, Mosaic, and Microsoft Internet Explorer) by clicking on the specific course numbers (e.g., EE440, EE505, etc) or on the specific contents (e.g., Z-transform, FIR Filters, etc). To help course instructors to create this multimedia features we have developed a hyper video editor tool. This tool allows instructor to mark various portions of the class video and create the corresponding hyper links and multimedia features by editing on line forms.
Jenq-Neng Hwang, Sachin G. Deshpande, Ming-Ting Sun
MMSP3
1997 A coded-domain video combiner for multipoint continuous presence video conferencing
abstract
This paper presents a coded-domain combining technique which can accommodate up to six users for continuous presence multipoint video conferencing. This technique is based on a group-of-block (GOB) video unit in the H.261 video standard. Three important technical issues including frame rate synchronization, combiner delay accumulation, and potential quality degradation at the GOB boundary, are addressed. A frame rate synchronization scheme is proposed for the coded-domain combiner. It is observed that delay accumulation will not occur at the combiner with the proposed multiplexing strategy. Further, simulation results indicate that the output perceptual quality is very comparable even without special motion search handling at the GOB boundary. This implies that the proposed scheme can be applied to standard terminals. Finally, we demonstrate the viability of this software-based coded-domain combiner for continuous-presence multipoint video conferencing with a fully functional experimental prototype built around a low-cost personal computer.
Ming-Ting Sun, Alexander C. Loui, Ting-Chung Chen
IEEE Trans. Circuits Syst. Video Technol.1
1994 Video bridging based on H.261 standard
abstract
Multi-point ISDN videoconferencing with video bridging in network-based servers represents a viable new network service. The paper presents a detailed technical analysis of a continuous presence video bridge using the H.261 video coding standard. The authors first compare the pros and cons of coded domain versus pel domain video bridges. The architecture and the required operations of a coded domain bridge using H.261 are then investigated. They derive the bounds of the bridge delay and the required buffer size for the implementation of the bridge. The delay and the buffer occupancy of the video bridge depend on the order, complexity, and the bit-distribution of the input video sources. To investigate a typical case, the authors simulate the delay and the buffer occupancy of a video bridge. They also provide a heuristic method to estimate the delay in a typical case. Several techniques are discussed to minimize the bridge delay and the buffer size. Finally, they simulate intra slice coding and show that the delay and the buffer size can be reduced significantly using this technique.>
Shawmin Lei, Ting-Chung Chen, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.3
1993 Coding and interworking for videotelephony
Ming-Ting Sun, Karel Rijkse, Dolf Schinkel, Arthur Meijboom
ISCAS1
1993 Continuous presence video bridging based on H.261 standard
abstract
Multi-point videoconferencing provides the full benefits of teleconference but also incurs more involved technical issues. This paper does a detailed analysis of a continuous presence video bridge using the H.261 video coding standard. We first describe the architecture and the required operations of a coded domain bridge using H.261. We then derive the bounds of the bridge delay and the required buffer size for the implementation of the bridge. The delay and the buffer occupancy of the video bridge depend on the order, complexity, and the bit- distribution of the input video sources. To investigate a typical case, we simulate the delay and the buffer occupancy of a video bridge. We also provide a heuristic method to estimate the delay in a typical case. Several techniques were discussed to minimize the bridge delay and the buffer size. Finally, we simulate an intra slice coding and show that the delay and the buffer size can be reduced significantly using this technique.
Ting-Chung Chen, Shawmin Lei, Ming-Ting Sun
VCIP3
1992 An all-ASIC implementation of a low bit-rate video codec
abstract
After many years of intensive deliberation, an international low-bit-rate video coding standard, known as CCITT (International Telegraph and Telephone Consultative Committee) Recommendation H.261, has been completed. The H.261 covers a wide range of bit rates at p*64 kbs, where p=1, 2, . . ., 30. A great deal of real-time signal processing power is required to compress an NTSC or other similar video signals to these rates for transport and to reconstruct the original signal back for display. In order to demonstrate the video quality of the newly established standard and the feasibility of a cost-effective VLSI solution, a real-time video codec based on H.261 has been constructed using ASICs (application specific integrated circuits). A single-board research prototype consisting of 11 ASICs with an aggregate signal processing power of approximately two billion operations per second is presented.>
Hiroshi Fujiwara, Ming Lei Liou, Ming-Ting Sun, Kun-Min Yang, Masanori Maruyama, Kazuyoshi Shomura, Koichi Ohyama
IEEE Trans. Circuits Syst. Video Technol.3
1992 Design and hardware architecture of high-order conditional entropy coding for images
abstract
High-order conditional entropy coding has not been practical due to its high complexity and lack of hardware to extract the conditioning state efficiently. The authors adopt the recently developed incremental-tree-extension technique to design the conditional tree for high-order conditional entropy coding. In order to make the high-speed conditional entropy coder feasible, they introduce several key innovations in the areas of complexity reduction and hardware architecture. For complexity reduction, they develop two techniques: code table reduction and nonlinear quantization of conditioning pixels. For hardware architecture, they propose a pattern-matching technique for fast conditioning state extraction and a multistage pipelined structure to handle the case of a large number of conditioning pixels. Using the complexity reduction techniques and the hardware structures. the authors demonstrate that it is possible to implement practical high-order conditional entropy codecs using current low-cost very large-scale integration (VLSI) technology.>
Shawmin Lei, Ming-Ting Sun, Kou-Hu Tzou
IEEE Trans. Circuits Syst. Video Technol.2
1991 An entropy coding system for digital HDTV applications
abstract
Run-length coding (RLC) and variable-length coding (VLC) are widely used techniques for lossless data compression. A high-speed entropy coding system using these two techniques is considered for digital high definition television (HDTV) applications. Traditionally, VLC decoding is implemented through a tree-searching algorithm as the input bits are received serially. For HDTV applications, it is very difficult to implement a real-time VLC decoder of this kind due to the very high data rate required. A parallel structured VLC decoder which decodes each codeword in one clock cycle regardless of its length is introduced. The required clock rate of the decoder is thus lower, and parallel processing architectures become easy to adopt in the entropy coding system. The parallel entropy coder and decoder are designed for implementation in two experimental prototype chips which are designed to encode and decode more than 52 million samples/s. Some related system issues, such as the synchronization of variable-length codewords and error concealment, are also discussed.>
Shawmin Lei, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.2
1990 VLSI architecture and implementation of a multifunction, forward/inverse discrete cosine transform processor
abstract
The Discrete Cosine Transform (DCT) is considered to be the most effective transform coding technique for image and video compression. In this paper, a new implementation of an experimental prototype multi-function DCT/IDCT (Inverse DCT) chip is reported. The chip is based on a distributed arithmetic architecture. The main features of the chip include: 1) The DCT and the IDCT are integrated in the same chip, 2) the chip achieves high accuracy, exceeding the stringent requirements of a proposed CCITF standard, 3) it achieves a high operating speed of 27 MHz, and is thus applicable to a wide-range of real-time image and video applications, 4) the internal clock frequency is the same as the pixel rate, and 5) with an on-chip zigzag scan converter and an adder/subtractor, it is multifunctional and useful in a DPCM configuration. The chip is implemented with standard cells and contains about 156k transistors.
Masanori Maruyama, H. Uwabu, I. Iwasaki, Hiroshi Fujiwara, Toshifumi Sakaguchi, Ming-Ting Sun, Ming Lei Liou
VCIP6
1989 Very high efficiency VLSI chip-pair for full search block matching with fractional precision
abstract
VLSI architecture design and implementation of a chip pair for the motion compensation full search block matching algorithm are described. This pair of ASICs (application-specific integrated circuits) is motivated by the intensive computational demands for performing motion compensation in real time. They have been developed to calculate fractional motion vectors with quarter-pel precision. The VLSI architecture is based on some special data-flow designs that allow sequential inputs but perform parallel processing with 100% efficiency for integer motion vector estimation and nearly 100% for fractional motion vector estimation. The chip-pair design has been laid out and simulated using a silicon compiler tool, and the chip statistics are summarized. Testing circuitry is included to increase the observability of the chips.>
Kun-Min Yang, Ming-Ting Sun, Lance Wu, I-Fei G. Chuang
ICASSP2