Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shipeng Li 0001

dblp:31/3974-1 · DBLP profile ↗
← Back
221ranked-venue papers
7as first author
0since 2021 · last 2019
0000-0001-5368-4256ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 185 · 6 first-authorSystems, architecture and hardware · 16 · 1 first-authorArtificial intelligence and machine learning · 12Computer networks · 8Databases, data management, data science and information retrieval · 8Applied, interdisciplinary, general and emerging computing · 4Security and privacy · 3Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
38 papers
Information retrieval · 78% Recommender systems · 9% Data mining · 7%
Computer graphics and multimedia
33 papers
Multimedia analysis and retrieval · 66% Image and video coding · 17% Computational photography and imaging · 7%
Artificial intelligence
13 papers
Image recognition and object detection · 27% Segmentation and scene understanding · 16% Transfer learning and domain adaptation · 14%
Human-computer interaction and pervasive computing
10 papers
Interaction techniques and input · 53% Ubiquitous computing and smart environments · 40% Haptics and multimodal interaction · 4%

Topics — the 30 heaviest of 136, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
image retrieval
1.7132014
Browse-to-Search: Interactive Exploratory Search with Visual Entities · ACM Trans. Inf. Syst. 2014
Image Relevance Prediction Using Query-Context Bag-of-Object Retrieval Model · IEEE Trans. Multim. 2014
A bag-of-objects retrieval model for web image search · ACM Multimedia 2012
Information retrieval › similarity search › nearest neighbor search
approximate nearest neighbor search
0.852015
Optimized Cartesian K-Means · IEEE Trans. Knowl. Data Eng. 2015
Order preserving hashing for approximate nearest neighbor search · ACM Multimedia 2013
Fast Neighborhood Graph Search Using Cartesian Concatenation · ICCV 2013
Information retrieval
retrieval models
0.552014
Image Relevance Prediction Using Query-Context Bag-of-Object Retrieval Model · IEEE Trans. Multim. 2014
A bag-of-objects retrieval model for web image search · ACM Multimedia 2012
Query-driven iterated neighborhood graph search for large scale indexing · ACM Multimedia 2012
Information retrieval › image retrieval › web image search
image re-ranking
0.532014
Image Relevance Prediction Using Query-Context Bag-of-Object Retrieval Model · IEEE Trans. Multim. 2014
A bag-of-objects retrieval model for web image search · ACM Multimedia 2012
Image search results refinement via outlier detection using deep contexts · CVPR 2012
Information retrieval › image retrieval
content-based image retrieval
0.432014
Browse-to-Search: Interactive Exploratory Search with Visual Entities · ACM Trans. Inf. Syst. 2014
Color filter for image search · ACM Multimedia 2012
Descriptive visual words and visual phrases for image applications · ACM Multimedia 2009
Data mining
clustering
0.422015
Optimized Cartesian K-Means · IEEE Trans. Knowl. Data Eng. 2015
Fast approximate k-means via cluster closures · CVPR 2012
Computer vision › Segmentation and scene understanding › saliency detection
salient object detection
0.432013
Salient Object Detection: A Discriminative Regional Feature Integration Approach · CVPR 2013
Salient object detection for searched web images via global saliency · CVPR 2012
Color filter for image search · ACM Multimedia 2012
Recommender systems
video recommendation
0.332012
SocialTransfer: cross-domain transfer learning from social streams for media applications · ACM Multimedia 2012
Contextual Video Recommendation by Multimodal Relevance and User Feedback · ACM Trans. Inf. Syst. 2011
VideoReach: an online video recommendation system · SIGIR 2007
Information retrieval › evaluation
query performance prediction
0.322014
Predicting Failing Queries in Video Search · IEEE Trans. Multim. 2014
When video search goes wrong: predicting query failure using search engine logs and visual search results · ACM Multimedia 2012
Information retrieval › similarity search
nearest neighbor search
0.322013
Order preserving hashing for approximate nearest neighbor search · ACM Multimedia 2013
Fast Neighborhood Graph Search Using Cartesian Concatenation · ICCV 2013
Machine learning › Transfer learning and domain adaptation
cross-domain learning
0.322013
Towards Cross-Domain Learning for Social Video Popularity Prediction · IEEE Trans. Multim. 2013
SocialTransfer: cross-domain transfer learning from social streams for media applications · ACM Multimedia 2012
Multimedia analysis and retrieval
contextual advertising
0.342008
Contextual in-image advertising · ACM Multimedia 2008
ImageSense · ACM Multimedia 2008
VideoSense: a contextual video advertising system · ACM Multimedia 2007
Multimedia analysis and retrieval
image retrieval
0.322014
Salable Image Search with Reliable Binary Code · ACM Multimedia 2014
Image search by concept map · SIGIR 2010
Information retrieval
hashing
0.322013
Order preserving hashing for approximate nearest neighbor search · ACM Multimedia 2013
Complementary hashing for approximate nearest neighbor search · ICCV 2011
Information retrieval › image retrieval
mobile visual search
0.322012
Local visual words coding for low bit rate mobile visual search · ACM Multimedia 2012
JIGSAW: interactive mobile visual search with multimodal queries · ACM Multimedia 2011
Information retrieval › image retrieval
web image search
0.322012
A bag-of-objects retrieval model for web image search · ACM Multimedia 2012
The role of attractiveness in web image search · ACM Multimedia 2011
Multimedia analysis and retrieval
image annotation
0.232008
A comprehensive human computation framework: with application to image labeling · ACM Multimedia 2008
ImageSense · ACM Multimedia 2008
Coherent image annotation by learning semantic distance · CVPR 2008
Information retrieval
visual word representation
0.222011
Million-scale near-duplicate video retrieval system · ACM Multimedia 2011
Descriptive visual words and visual phrases for image applications · ACM Multimedia 2009
Data mining › clustering
k-means clustering
0.212015
Optimized Cartesian K-Means · IEEE Trans. Knowl. Data Eng. 2015
Information retrieval › similarity search › vector quantization
product quantization
0.212015
Optimized Cartesian K-Means · IEEE Trans. Knowl. Data Eng. 2015
Image and video coding › video compression › 3d video coding
depth map coding
0.212015
Layered Compression for High-Precision Depth Data · IEEE Trans. Image Process. 2015
Multimedia analysis and retrieval
video retrieval
0.222013
Listen, look, and gotcha: instant video search with mobile phones by layered audio-video indexing · ACM Multimedia 2013
Video-based image retrieval · ACM Multimedia 2011
Information retrieval › interactive information retrieval
exploratory search
0.212014
Browse-to-Search: Interactive Exploratory Search with Visual Entities · ACM Trans. Inf. Syst. 2014
Information retrieval › hashing › hashing-based retrieval
hash code ranking
0.212014
Optimized Distances for Binary Code Ranking · ACM Multimedia 2014
Information retrieval › retrieval models
language model
0.212014
Image Relevance Prediction Using Query-Context Bag-of-Object Retrieval Model · IEEE Trans. Multim. 2014
Machine learning and data management
metric learning
0.212014
Optimized Distances for Binary Code Ranking · ACM Multimedia 2014
Computational photography and imaging
photography assistance
0.212014
Socialized Mobile Photography: Learning to Photograph With Social Context via Mobile Devices · IEEE Trans. Multim. 2014
Ubiquitous computing and smart environments › context-aware computing
context-aware mobile computing
0.212014
Socialized Mobile Photography: Learning to Photograph With Social Context via Mobile Devices · IEEE Trans. Multim. 2014
Interaction techniques and input
gesture input
0.222012
Browse-to-search · ACM Multimedia 2012
TapTell: understanding visual intents on-the-go · ACM Multimedia 2011
Multimedia analysis and retrieval › visual search
mobile visual search
0.222013
TapTell: understanding visual intents on-the-go · ACM Multimedia 2011
Interaction Design for Mobile Visual Search · IEEE Trans. Multim. 2013

Methods — techniques the papers use, named apart from their topics

inverted index · 0.5best-first search · 0.5product quantization · 0.4codebook learning · 0.4social media mining · 0.4crowdsourcing · 0.4context modeling · 0.4discriminative regional feature integration · 0.3backgroundness descriptor · 0.3transfer learning · 0.3speech recognition · 0.3bag-of-words · 0.3error-controllable pixel domain encoding · 0.28-bit image/video encoding · 0.2relevance feedback · 0.2attention fusion · 0.2visual entity encapsulation · 0.2visual concept consistency · 0.2
YearPublicationVenuePosition
2019 Introduction of New Associate Editors
Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 Editor-in-Chief Message
abstract
Time really flies! I can hardly believe that almost two years have passed since I assumed the Editor-in-Chief (EiC) position of the IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) on January 1, 2018. As planned, I will not pursue another term due to other commitments at work and to my family and will retire from the EiC position at the end of this year. The good news is that the TCSVT will be in the capable hands of the new EiC, Prof. Feng Wu, from the University of Science and Technology of China, and his new Editorial Board (EB). Prof. Wu is the current Deputy EiC of the TCSVT, and has many years of experience and leadership in the TCSVT community. I have the full confidence that he will take the Transactions to the next level!
Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Editor-in-Chief Message: Embracing the Era of Intelligent Visual Technology and Systems
Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2017 Delay-Rate-Distortion Optimization for Cloud Gaming With Hybrid Streaming
abstract
Cloud gaming as the emerging game service has attracted significant attention. However, traditional video streaming approach suffers from high bandwidth consumption, and traditional graphics streaming approach requires a long initial period to download game models. In this paper, we propose a novel hybrid streaming framework, jointly applying video streaming and graphics streaming to provide a high-quality gaming experience. In the proposed framework, cloud servers not only transmit the encoded video frames but also progressively transmit the graphics data, which are used to render a game frame to provide an additional reference to the video encoder. Based on the proposed framework, we investigate the delay-rate-distortion optimization problem, where the source rate between the video stream and the graphics stream is optimized to minimize the overall distortion under the bandwidth and response delay constraints. The experimental results demonstrate that the proposed hybrid streaming can achieve the lowest distortion under the constraints of bandwidth and response delay, compared with the traditional video streaming and graphics streaming.
Xiaoming Nan, Xun Guo 0002, Yan Lu 0001, Ling Guan, Shipeng Li 0001, Baining Guo
IEEE Trans. Circuits Syst. Video Technol.6
2016 A High-Fidelity and Low-Interaction-Delay Screen Sharing System
abstract
The pervasive computing environment and wide network bandwidth provide users more opportunities to share screen content among multiple devices. In this article, we introduce a remote display system to enable screen sharing among multiple devices with high fidelity and responsive interaction. In the developed system, the frame-level screen content is compressed and transmitted to the client side for screen sharing, and the instant control inputs are simultaneously transmitted to the server side for interaction. Even if the screen responds immediately to the control messages and updates at a high frame rate on the server side, it is difficult to update the screen content with low delay and high frame rate in the client side due to non-negligible time consumption on the whole screen frame compression, transmission, and display buffer updating. To address this critical problem, we propose a layered structure for screen coding and rendering to deliver diverse screen content to the client side with an adaptive frame rate. More specifically, the interaction content with small region screen update is compressed by a blockwise screen codec and rendered at a high frame rate to achieve smooth interaction, while the natural video screen content is compressed by standard video codec and rendered at a regular frame rate for a smooth video display. Experimental results with real applications demonstrate that the proposed system can successfully reduce transmission bandwidth cost and interaction delay during screen sharing. Especially for user interaction in small regions, the proposed system can achieve a higher frame rate than most previous counterparts.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
ACM Trans. Multim. Comput. Commun. Appl.4
2016 Automatic Generation of Visual-Textual Presentation Layout
abstract
Visual-textual presentation layout (e.g., digital magazine cover, poster, Power Point slides, and any other rich media), which combines beautiful image and overlaid readable texts, can result in an eye candy touch to attract users’ attention. The designing of visual-textual presentation layout is therefore becoming ubiquitous in both commercially printed publications and online digital magazines. However, handcrafting aesthetically compelling layouts still remains challenging for many small businesses and amateur users. This article presents a system to automatically generate visual-textual presentation layouts by investigating a set of aesthetic design principles, through which an average user can easily create visually appealing layouts. The system is attributed with a set of topic-dependent layout templates and a computational framework integrating high-level aesthetic principles (in a top-down manner) and low-level image features (in a bottom-up manner). The layout templates, designed with prior knowledge from domain experts, define spatial layouts, semantic colors, harmonic color models, and font emotion and size constraints. We formulate the typography as an energy optimization problem by minimizing the cost of text intrusion, the utility of visual space, and the mismatch of information importance in perception and semantics, constrained by the automatically selected template and further preserving color harmonization. We demonstrate that our designs achieve the best reading experience compared with the reimplementation of parts of existing state-of-the-art designs through a series of user studies.
Xuyong Yang, Tao Mei 0001, Ying-Qing Xu, Yong Rui, Shipeng Li 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2015 Region-of-interest based coding scheme for synthesized video
abstract
In many multimedia applications, such as online speech, video chat and online conference, multiple source videos are synthesized in a single scene for explicit presentation and the synthesized video is compressed for transmission. The source video with important contents deserves more compression resources for quality preservation under the bandwidth constraint. To address this problem, a region-of-interest (ROI) based coding scheme for synthesized video is proposed in this paper aiming at achieve better and consistent quality for ROI source videos with the bitrate meeting the constraint bandwidth. In the proposed coding scheme, ROI based rate-distortion (R-D) models are established, in which different R-D models are built for different source video. Then an objective function is defined with respect to the video quality and the consistency of video quality. By minimizing the objective function, the optimal quantization parameters for the ROI and non-ROI source videos are obtained. The experimental results show that the proposed coding scheme achieves better and consistent quality for ROI source videos.
Wenbo Zhao 0004, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Debin Zhao
VCIP4
2015 TapTell: Interactive visual search for mobile task recommendation
Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001
J. Vis. Commun. Image Represent.5
2015 Exploratory Product Image Search With Circle-to-Search Interaction
abstract
Exploratory search is emerging as a new form of information-seeking activity in the research community, which generally combines browsing and searching content together to help users gain additional knowledge and form accurate queries, thereby assisting the users with their seeking and investigation activities. However, there have been few attempts at addressing integrated exploratory search solutions when image browsing is incorporated into the exploring loop. In this paper, we investigate the challenges of understanding users' search interests from the product images being browsed and inferring their actual search intentions. We propose a novel interactive image exploring system for allowing users to lightly switch between browse and search processes, and naturally complete visual-based exploratory search tasks in an effective and efficient way. This system enables users to specify their visual search interests in product images by circling any visual objects in web pages, and then the system automatically infers users' underlying intent by analyzing the browsing context and by analyzing the same or similar product images obtained by large-scale image search technology. Users can then utilize the recommended queries to complete intent-specific exploratory tasks. The proposed solution is one of the first attempts to understand users' interests for a visual-based exploratory product search task by integrating the browse and search activities. We have evaluated our system performance based on five million product images. The evaluation study demonstrates that the proposed system provides accurate intent-driven search results and fast response to exploratory search demands compared with the conventional image search methods, and also, provides users with robust results to satisfy their exploring experience.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.6
2015 MoVieUp: Automatic Mobile Video Mashup
abstract
With the proliferation of mobile devices, people are taking videos of the same events anytime and anywhere. Even though these crowdsourced videos are uploaded to the cloud and shared, the viewing experience is very limited due to monotonous viewing, visual redundancy, and bad audio-video quality. In this paper, we present a fully automatic mobile video mashup system that works in the cloud to combine recordings captured by multiple devices from different view angles and at different time slots into a single yet enriched and professional looking video-audio stream. We summarize a set of computational filming principles for multicamera settings from a formal focus study. Based on these principles, given a set of recordings of the same event, our system is able to synchronize these recordings with audio fingerprints, assess audio and video quality, detect video cut points, and generate video and audio mashups. The audio mashup is the maximization of audio quality under the less switching principle, while the video mashup is formalized as maximizing video quality and content diversity, constrained by the summarized filming principles. Our system is different from any existing work in this field in three ways: 1) our system is fully automatic; 2) the system incorporates a set of computational domain-specific filming principles summarized from a formal focus study; and 3) in addition to video, we also consider audio mashup that is a key factor of user experience (UX) yet often overlooked in existing research. Evaluations show that our system achieves performance results that are superior to state-of-the-art video mashup techniques, thus providing a better UX.
Tao Mei 0001, Ying-Qing Xu, Nenghai Yu, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2015 Layered Compression for High-Precision Depth Data
abstract
With the development of depth data acquisition technologies, access to high-precision depth with more than 8-b depths has become much easier and determining how to efficiently represent and compress high-precision depth is essential for practical depth storage and transmission systems. In this paper, we propose a layered high-precision depth compression framework based on an 8-b image/video encoder to achieve efficient compression with low complexity. Within this framework, considering the characteristics of the high-precision depth, a depth map is partitioned into two layers: 1) the most significant bits (MSBs) layer and 2) the least significant bits (LSBs) layer. The MSBs layer provides rough depth value distribution, while the LSBs layer records the details of the depth value variation. For the MSBs layer, an error-controllable pixel domain encoding scheme is proposed to exploit the data correlation of the general depth information with sharp edges and to guarantee the data format of LSBs layer is 8 b after taking the quantization error from MSBs layer. For the LSBs layer, standard 8-b image/video codec is leveraged to perform the compression. The experimental results demonstrate that the proposed coding scheme can achieve real-time depth compression with satisfactory reconstruction quality. Moreover, the compressed depth data generated from this scheme can achieve better performance in view synthesis and gesture recognition applications compared with the conventional coding schemes because of the error control algorithm.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
IEEE Trans. Image Process.4
2015 Optimized Cartesian K-Means
abstract
Product quantization-based approaches are effective to encode high-dimensional data points for approximate nearest neighbor search. The space is decomposed into a Cartesian product of low-dimensional subspaces, each of which generates a sub codebook. Data points are encoded as compact binary codes using these sub codebooks, and the distance between two data points can be approximated efficiently from their codes by the precomputed lookup tables. Traditionally, to encode a subvector of a data point in a subspace, only one sub codeword in the corresponding sub codebook is selected, which may impose strict restrictions on the search accuracy. In this paper, we propose a novel approach, named optimized cartesian K-means (ock-means), to better encode the data points for more accurate approximate nearest neighbor search. In ock-means, multiple sub codewords are used to encode the subvector of a data point in a subspace. Each sub codeword stems from different sub codebooks in each subspace, which are optimally generated with regards to the minimization of the distortion errors. The high-dimensional data point is then encoded as the concatenation of the indices of multiple sub codewords from all the subspaces. This can provide more flexibility and lower distortion errors than traditional methods. Experimental results on the standard real-life data sets demonstrate the superiority over state-of-the-art approaches for approximate nearest neighbor search.
Jingdong Wang 0001, Jingkuan Song, Xin-Shun Xu, Heng Tao Shen, Shipeng Li 0001
IEEE Trans. Knowl. Data Eng.6
2014 Content adaptive screen image scaling
abstract
This paper proposes an efficient content adaptive screen image scaling scheme for the real-time screen applications like remote desktop and screen sharing. In the proposed screen scaling scheme, a screen content classification step is first introduced to classify the screen image into text and pictorial regions. Afterward, we propose an adaptive shift linear interpolation algorithm to predict the new pixel values with the shift offset adapted to the content type of each pixel. The shift offset for each screen content type is offline optimized by minimizing the theoretical interpolation error based on the training samples respectively. The proposed content adaptive screen image scaling scheme can achieve good visual quality and also keep the low complexity for realtime applications.
Yao Zhai, Qifei Wang, Yan Lu 0001, Shipeng Li 0001
ICIP4
2014 A novel cloud gaming framework using joint video and graphics streaming
abstract
As the popularity of smart phones and tablets, users have an increasing desire to enjoy ubiquitous game playing. The emerging cloud gaming turns this desire into reality, enabling users to play games at anywhere on any devices. However, due to the huge amount of data transmission, it is challenging to provide a high quality game experience under the limited bandwidth capacity. In this paper, we propose a novel cloud gaming framework, in which we introduce two synchronized graphics buffers at both the server and the client sides. The server not only streams the compressed frames captured from game scenes, but also progressively transmits graphics data. The received graphics data is used to generate reference frames. When compressing the next frame, the cloud server will choose the reference frame with a lower residual error, from the previous frame and the current frame rendered from the graphics buffer. With the accumulation of graphics data, the frame rendered from the graphics buffer is close to the captured frame, which greatly reduces the transmission bit rates. Based on the proposed framework, we study the rate allocation problem, in which we optimize the allocated bit rates between the compressed frame and the graphics data to minimize the total distortion under the bandwidth constraint. Experimental results demonstrate that the proposed framework can optimally allocate bit rates to achieve a minimal distortion for cloud gaming compared to the traditional video streaming and graphics streaming approaches.
Xiaoming Nan, Xun Guo 0002, Yan Lu 0001, Ling Guan, Shipeng Li 0001, Baining Guo
ICME6
2014 A low latency cloud gaming system using edge preserved image homography
abstract
The emerging cloud gaming technology has been growing fast, driving up huge mobile consumer demands. The video streaming based cloud gaming scenario renders the game scenes in the cloud servers, and streams the encoded sequences to the thin clints where the game scenes are decoded and displayed to the players. However, current existing clouding gaming services have some problems, such as the latency and bandwidth limitation. The size of the video stream is usually quite large which requires heavy transmission. Worse still, the frame data rate will burst when the game scenes contain fast translation or rotation, resulting in strong latency problem. In this paper, we propose a novel video streaming based cloud gaming algorithm which reduces the burst of the frame rate significantly. There are mainly two innovations in this paper. Firstly, based on the analysis of the motion estimation strategy in the video codec, we introduce image homography technique for better motion prediction. Meanwhile, according to the rasterization rules of the game engine, we present a special designed interpolation algorithm named Edge Preserved Interpolation (EPI), for more accurate edge interpolation and further reduce the residues in the edge regions. The proposed algorithm is implemented on the x264 platform. Experimental results show that our algorithm has 18.0% BD-rate reduction compared with x264.
Lingfeng Xu, Xun Guo 0002, Yan Lu 0001, Shipeng Li 0001, Oscar C. Au, Lu Fang 0001
ICME4
2014 High frame rate screen video coding for screen sharing applications
abstract
In this paper, we propose a high frame rate screen video compression scheme aiming at improving the interactive user experience on screen sharing applications. The proposed screen video compression is performed as two-layer coding: a base layer coding using the conventional video codec and an enhancement layer coding using the proposed open-loop coding scheme. For efficient frame level layer selection and compression, the content update of each frame is evaluated through global motion detection. The screen frame with significant content update is fed to the conventional video encoder in base layer. In contrast, the frame with little update is compressed in enhancement layer in which the duplicate content is indicated by global motion vector and skip flag while the updated content is encoded by distinct intra modes in terms of inherent local features. The experimental results demonstrate that for the screen video containing interaction, the proposed coding scheme can achieve 3.09ms/frame encoding rate and 2.33ms/frame decoding rate with efficient rate distortion performance.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
ISCAS4
2014 Salable Image Search with Reliable Binary Code
abstract
In many existing image retrieval algorithms, Bag-of-Words (BoW) model has been widely adopted for image representation. To achieve accurate indexing and efficient retrieval, local features such as the SIFT descriptor are extracted and quantized to visual words. One of the most popular quantization scheme is scalar quantization, which generates binary signature with an empirical threshold value. However, such binarization strategy inevitably suffers from the quantization loss induced by each quantized bit and impairs the effectiveness of search performance. In this paper, we investigate the reliability of each bit in scalar quantization and propose a novel reliable binary SIFT feature. We move one step ahead to incorporate the reliability in both index word expansion and feature similarity. Our proposed approach not only accelerates the search speed by narrowing search space, but also improves the retrieval accuracy by alleviating the impact of unreliable quantized bits. Experimental results demonstrate that the proposed approach achieves significant improvement in retrieval efficiency and accuracy.
Guangxin Ren, Shipeng Li 0001, Nenghai Yu, Qi Tian 0001
ACM Multimedia3
2014 Optimized Distances for Binary Code Ranking
abstract
Binary encoding on high-dimensional data points has attracted much attention due to its computational and storage efficiency. While numerous efforts have been made to encode data points into binary codes, how to calculate the effective distance on binary codes to approximate the original distance is rarely addressed. In this paper, we propose an effective distance measurement for binary code ranking. In our approach, the binary code is firstly decomposed into multiple sub codes, each of which generates a query-dependent distance lookup table. Then the distance between the query and the binary code is constructed as the aggregation of the distances from all sub codes by looking up their respective tables. The entries of the lookup tables are optimized by minimizing the misalignment between the approximate distance and the original distance. Such a scheme is applied to both the symmetric distance and the asymmetric distance. Extensive experimental results show superior performance of the proposed approach over state-of-the-art methods on three real-world high-dimensional datasets for binary code ranking.
Heng Tao Shen, Shuicheng Yan, Nenghai Yu, Shipeng Li 0001, Jingdong Wang 0001
ACM Multimedia5
2014 Image tag refinement by regularized latent Dirichlet allocation
Jingdong Wang 0001, Jiazhen Zhou, Tao Mei 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
Comput. Vis. Image Underst.6
2014 Paradigm Shifts in Video Technologies: Introduction to T-CSVT Future Special Issues
abstract
Recently, we have witnessed many exciting technology and methodology advances that will deeply impact the way we do research and development. These advances imply many important new research directions in the next 5–10 years, and potentially introduce paradigm shifts in video technologies—visual signal processing, communication, and computing. We have observed new trends ranging from breakthroughs in information theory and machine learning, big data, crowd sourcing, cloud computing, low-cost yet powerful computing end points, Internet of things (IoT), mobile and wireless computing, augmented reality, green computing, to integrated system research. Let us elaborate on each of these trends.
Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2014 Predicting Failing Queries in Video Search
abstract
The ability to predict when a video search query is not likely to deliver satisfying search results is expected to enable more effective search results optimizations and improved search experience for users. In this paper, we propose a novel context-aware query failure prediction approach that predicts whether a particular query submitted in a user's search session is likely to fail. The approach builds on the well-known concept of query performance prediction introduced in conventional text-based Web search to estimate the query's retrieval performance, but extends this concept with two novel characteristics, user indicators and engine indicators. User indicators are derived from transaction logs, capture the patterns of user interactions with the video search engine, and exploit the context in which a particular query was submitted. Engine indicators are derived from the search results list and measure the consistency of visual search results at the level of visual concepts and textual metadata associated with videos. Extensive evaluation of the approach on a test set containing over one million video search queries shows its effectiveness and demonstrates a significant improvement over traditional and state-of-the-art baseline approaches.
Christoph Kofler, Linjun Yang, Martha A. Larson, Tao Mei 0001, Alan Hanjalic, Shipeng Li 0001
IEEE Trans. Multim.6
2014 Image Relevance Prediction Using Query-Context Bag-of-Object Retrieval Model
abstract
Image search reranking and image research result summarization are two effective approaches which enhance text-based image search results using visual information. Since the existing approaches optimize search relevance in terms of average performance, they usually cannot achieve satisfactory results for some particular classes of queries, like “object queries,” which is defined as the queries with the intent of searching for some kinds of objects. One possible reason is that the generic approaches such as , , are mostly built based on the global statistics of images as features while ignoring the fact that the relevance between the image and the query sometimes depends on an image patch instead of the whole image. In this paper, we therefore design a novel bag-of-object retrieval model to predict image relevance, which is particularly effective for object queries. First, we construct an object vocabulary containing query-relative objects by mining frequent object patches from the result image collection of the expanded query set. After representing each image as a bag of objects, our retrieval model can be derived from a risk-minimization framework for language modeling. To demonstrate the effectiveness of the proposed model, this paper also present two related applications: for image search reranking, we adopt a supervised framework to combine multiple ranking features from different assumptions; for image search result summarization, we propose a two-step ranking process which optimizes not only representativeness but also image attractiveness. The experimental results show that the proposed methods can significantly outperform the existing approaches.
Yang Yang 0222, Linjun Yang, Gangshan Wu, Shipeng Li 0001
IEEE Trans. Multim.4
2014 Socialized Mobile Photography: Learning to Photograph With Social Context via Mobile Devices
abstract
The popularity of mobile devices equipped with various cameras has revolutionized modern photography. People are able to take photos and share their experiences anytime and anywhere. However, taking a high quality photograph via mobile device remains a challenge for mobile users. In this paper we investigate a photography model to assist mobile users in capturing high quality photos by using both the rich context available from mobile devices and crowdsourced social media on the Web. The photography model is learned from community-contributed images on the Web, and dependent on user's social context. The context includes user's current geo-location, time (i.e., time of the day), and weather (e.g., clear, cloudy, foggy, etc.). Given a wide view of scene, our socialized mobile photography system is able to suggest the optimal view enclosure (composition) and appropriate camera parameters (aperture, ISO, and exposure time). Extensive experiments have been performed for eight well-known hot spot landmark locations where sufficient context tagged photos can be obtained. Through both objective and subjective evaluations, we show that the proposed socialized mobile photography system can indeed effectively suggest proper composition and camera parameters to help the user capture high quality photos.
Wenyuan Yin, Tao Mei 0001, Chang Wen Chen, Shipeng Li 0001
IEEE Trans. Multim.4
2014 Browse-to-Search: Interactive Exploratory Search with Visual Entities
abstract
With the development of image search technology, users are no longer satisfied with searching for images using just metadata and textual descriptions. Instead, more search demands are focused on retrieving images based on similarities in their contents (textures, colors, shapes etc.). Nevertheless, one image may deliver rich or complex content and multiple interests. Sometimes users do not sufficiently define or describe their seeking demands for images even when general search interests appear, owing to a lack of specific knowledge to express their intents. A new form of information seeking activity, referred to as exploratory search, is emerging in the research community, which generally combines browsing and searching content together to help users gain additional knowledge and form accurate queries, thereby assisting the users with their seeking and investigation activities. However, there have been few attempts at addressing integrated exploratory search solutions when image browsing is incorporated into the exploring loop. In this work, we investigate the challenges of understanding users' search interests from the images being browsed and infer their actual search intentions. We develop a novel system to explore an effective and efficient way for allowing users to seamlessly switch between browse and search processes, and naturally complete visual-based exploratory search tasks. The system, called Browse-to-Search enables users to specify their visual search interests by circling any visual objects in the webpages being browsed, and then the system automatically forms the visual entities to represent users' underlying intent. One visual entity is not limited by the original image content, but also encapsulated by the textual-based browsing context and the associated heterogeneous attributes. We use large-scale image search technology to find the associated textual attributes from the repository. Users can then utilize the encapsulated visual entities to complete search tasks. The Browse-to-Search system is one of the first attempts to integrate browse and search activities for a visual-based exploratory search, which is characterized by four unique properties: (1) in session—searching is performed during browsing session and search results naturally accompany with browsing content; (2) in context—the pages being browsed provide text-based contextual cues for searching; (3) in focus—users can focus on the visual content of interest without worrying about the difficulties of query formulation, and visual entities will be automatically formed; and (4) intuitiveness—a touch and visual search-based user interface provides a natural user experience. We deploy the Browse-to-Search system on tablet devices and evaluate the system performance using millions of images. We demonstrate that it is effective and efficient in facilitating the user's exploratory search compared to the conventional image search methods and, more importantly, provides users with more robust results to satisfy their exploring experience.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, Shipeng Li 0001
ACM Trans. Inf. Syst.6
2013 Salient Object Detection: A Discriminative Regional Feature Integration Approach
abstract
Salient object detection has been attracting a lot of interest, and recently various heuristic computational models have been designed. In this paper, we regard saliency map computation as a regression problem. Our method, which is based on multi-level image segmentation, uses the supervised learning approach to map the regional feature vector to a saliency score, and finally fuses the saliency scores across multiple levels, yielding the saliency map. The contributions lie in two-fold. One is that we show our approach, which integrates the regional contrast, regional property and regional background ness descriptors together to form the master saliency map, is able to produce superior saliency maps to existing algorithms most of which combine saliency maps heuristically computed from different types of features. The other is that we introduce a new regional feature vector, background ness, to characterize the background, which can be regarded as a counterpart of the objectness descriptor [2]. The performance evaluation on several popular benchmark data sets validates that our approach outperforms existing state-of-the-arts.
Huaizu Jiang, Jingdong Wang 0001, Zejian Yuan, Yang Wu 0001, Nanning Zheng 0001, Shipeng Li 0001
CVPR6
2013 Supervised Kernel Descriptors for Visual Recognition
abstract
In visual recognition tasks, the design of low level image feature representation is fundamental. The advent of local patch features from pixel attributes such as SIFT and LBP, has precipitated dramatic progresses. Recently, a kernel view of these features, called kernel descriptors (KDES), generalizes the feature design in an unsupervised fashion and yields impressive results. In this paper, we present a supervised framework to embed the image level label information into the design of patch level kernel descriptors, which we call supervised kernel descriptors (SKDES). Specifically, we adopt the broadly applied bag-of-words (BOW) image classification pipeline and a large margin criterion to learn the low-level patch representation, which makes the patch features much more compact and achieve better discriminative ability than KDES. With this method, we achieve competitive results over several public datasets comparing with state-of-the-art methods.
Peng Wang 0001, Jingdong Wang 0001, Weiwei Xu 0003, Hongbin Zha, Shipeng Li 0001
CVPR6
2013 Fast Neighborhood Graph Search Using Cartesian Concatenation
abstract
In this paper, we propose a new data structure for approximate nearest neighbor search. This structure augments the neighborhood graph with a bridge graph. We propose to exploit Cartesian concatenation to produce a large set of vectors, called bridge vectors, from several small sets of subvectors. Each bridge vector is connected with a few reference vectors near to it, forming a bridge graph. Our approach finds nearest neighbors by simultaneously traversing the neighborhood graph and the bridge graph in the best-first strategy. The success of our approach stems from two factors: the exact nearest neighbor search over a large number of bridge vectors can be done quickly, and the reference vectors connected to a bridge (reference) vector near the query are also likely to be near the query. Experimental results on searching over large scale datasets (SIFT, GISTand HOG) show that our approach outperforms state-of-the-art ANN search algorithms in terms of efficiency and accuracy. The combination of our approach with the IVFADC system [18] also shows superior performance over the BIGANN dataset of 1 billion SIFT features compared with the best previously published result.
Jing Wang 0068, Jingdong Wang 0001, Rui Gan, Shipeng Li 0001, Baining Guo
ICCV5
2013 Arbitrary-sized motion detection in screen video coding
abstract
In real-time screen remoting system, frame rate is one of essential factors that affect user experience. Therefore, how to compress diversity of screen contents fast and efficiently is a key issue. Existing video codecs such as H.264 are always used in such a system for screen compression. However, arbitrary-sized regions with large motion always exist in typical screen content videos, which lead to a lower encoding speed and higher bit-rate, thereby decrease the frame rate. This paper proposes an efficient motion detection algorithm, which is fast and efficient for large motion regions. In specific, a region-based motion detection is used to find motion vectors instead of traditional block based motion estimation. The motion vectors are then utilized by H.264 encoder for normal motion compensated prediction. Experimental results show that the proposed algorithm can reduce both encoding time and bit-rate significantly.
Tao Zhang 0013, Xun Guo 0002, Yan Lu 0001, Shipeng Li 0001, Siwei Ma 0001, Debin Zhao
ICIP4
2013 Effective hand segmentation and gesture recognition for browsing web pages on a large screen
abstract
Modern digital family technology enables people surf the Internet and watch videos via a large screen. This paper proposes an effective scheme for using hand gestures rather than the common remote controllers to browse the web pages on a large TV screen. The proposed scheme models four gesture modes: mouse mode, scroll mode, zoom mode and input mode to help the user browse web pages naturally and comfortably. Then we combine RGB, depth, motion information and face detection to achieve accurate and real-time hand segmentation and gesture recognition for enabling the four gesture modes. The experiments show the proposed scheme works well in various illumination environments and complicated backgrounds with multiple moving humans. The recognition accuracy of hand shapes in the proposed scheme arrives at 98.50%, and the successful rate for visual digits input reaches 89.00%. Furthermore, the frame rate of the hand-gesture detection and recognition is about 18 fps. Thus the scheme is accurate, real-time and natural.
Zhanghui Chen, Huifeng Shen, Yan Lu 0001, Shipeng Li 0001
ICME4
2013 Layered screen video coding leveraging hardware video codec
abstract
In this paper, we propose a layered screen video coding scheme based on existing video codecs to leverage hardware video codec for efficient screen video compression. In this scheme, the screen video compression is performed as two-layer coding: base layer coding and enhancement layer coding. The screen video is first analyzed in both frame and block levels for useful temporal and spatial information extraction to assist coding content selection in each layer. The non-skip screen frames are directly compressed by the conventional video codec in the base layer, while the screen contents sensitive to the video quality degradation are selected for improved coding in the enhancement layer. For contents to be enhanced, two intra coding modes are designed to improve the quality of the compressed text/graphics contents and suppress the artifacts introduced by chroma downsampling. The experimental results demonstrate that the screen video quality is improved objectively and subjectively by the proposed scheme with low cost on bitrate and computation complexity. Moreover, an average of 2.95dB coding gain is achieved in high bitrate.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
ICME4
2013 Rate-distortion optimized block classification and bit allocation in screen video compression
abstract
Due to the divergent characteristics of image contents and text contents in screen videos, how to make the joint optimization leveraging rate-distortion (R-D) optimized block classification and bit allocation is critical to the compression performance. In this paper, a general model-based solution is proposed as an attempt to solve this problem. The contributions of this paper are twofold: First, the rate and distortion characteristics of image blocks and text blocks in block-based content-adaptive screen video encoder (BASC) are carefully studied, and the rate and distortion models are proposed. Second, with the proposed rate and distortion models, the R-D optimized block classification and bit allocation are derived using bisection searched Lagrange multiplier method. Experimental results demonstrate that the proposed R-D optimized block classification and bit allocation algorithms are able to adapt to diverse screen contents, which results in a significant gain of up to 4.5dB in PSNR.
Oscar C. Au, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001
ISCAS5
2013 Listen, look, and gotcha: instant video search with mobile phones by layered audio-video indexing
abstract
Mobile video is quickly becoming a mass consumer phenomenon. More and more people are using their smartphones to search and browse video content while on the move. In this paper, we have developed an innovative instant mobile video search system through which users can discover videos by simply pointing their phones at a screen to capture a very few seconds of what they are watching. The system is able to index large-scale video data using a new layered audio-video indexing approach in the cloud, as well as extract light-weight joint audio-video signatures in real time and perform progressive search on mobile devices. Unlike most existing mobile video search applications that simply send the original video query to the cloud, the proposed mobile system is one of the first attempts at instant and progressive video search leveraging the light-weight computing capacity of mobile devices. The system is characterized by four unique properties: 1) a joint audio-video signature to deal with the large aural and visual variances associated with the query video captured by the mobile phone, 2) layered audio-video indexing to holistically exploit the complementary nature of audio and video signals, 3) light-weight fingerprinting to comply with mobile processing capacity, and 4) a progressive query process to significantly reduce computational costs and improve the user experience---the search process can stop anytime once a confident result is achieved. We have collected 1,400 query videos captured by 25 mobile users from a dataset of 600 hours of video. The experiments show that our system outperforms state-of-the-art methods by achieving 90.79% precision when the query video is less than 10 seconds and 70.07% even when the query video is less than 5 seconds.
Wu Liu 0005, Tao Mei 0001, Yongdong Zhang 0001, Jintao Li 0001, Shipeng Li 0001
ACM Multimedia5
2013 Order preserving hashing for approximate nearest neighbor search
abstract
In this paper, we propose a novel method to learn similarity-preserving hash functions for approximate nearest neighbor (NN) search. The key idea is to learn hash functions by maximizing the alignment between the similarity orders computed from the original space and the ones in the hamming space. The problem of mapping the NN points into different hash codes is taken as a classification problem in which the points are categorized into several groups according to the hamming distances to the query. The hash functions are optimized from the classifiers pooled over the training points. Experimental results demonstrate the superiority of our approach over existing state-of-the-art hashing techniques.
Jingdong Wang 0001, Nenghai Yu, Shipeng Li 0001
ACM Multimedia4
2013 Annotation for free: video tagging by mining user search behavior
abstract
The problem of tagging is mostly considered from the perspectives of machine learning and data-driven philosophy. A fundamental issue that underlies the success of these approaches is the visual similarity, ranging from the nearest neighbor search to manifold learning, to identify similar instances of an example for tag completion. The need to searching for millions of visual examples in high-dimensional feature space, however, makes the task computationally expensive. Moreover, the results can suffer from robustness problem, when the underlying data, such as online videos, are rich of semantics and the similarity is difficult to be learnt from low-level features. This paper studies the exploration of user searching behavior through click-through data, which is largely available and freely accessible by search engines, for learning video relationship and applying the relationship for economic way of annotating online videos. We demonstrated that, by a simple approach using co-click statistics, promising results were obtained in contrast to feature-based similarity measurement. Furthermore, considering the long tail effect that few videos dominate most clicks, a new method based on~polynomial~semantic indexing is proposed to learn a latent space~for alleviating the sparsity problem of click-through data. The proposed approaches are then applied for three major tasks in tagging: tag assignment, ranking, and enrichment. On~a bipartite graph constructed from click-through data with~over 15 million queries and 20 million video URL clicks,~we showed that annotation can be performed for free with competitive performance and minimum computing resource, representing a new and promising paradigm for video tagging in addition to machine learning and data-driven methodologies.
Ting Yao 0003, Tao Mei 0001, Chong-Wah Ngo, Shipeng Li 0001
ACM Multimedia4
2013 Depth sensor assisted real-time gesture recognition for interactive presentation
Hanjie Wang, Jingjing Fu, Yan Lu 0001, Xilin Chen 0001, Shipeng Li 0001
J. Vis. Commun. Image Represent.5
2013 A Low-Complexity Screen Compression Scheme for Interactive Screen Sharing
abstract
Interactive screen sharing requires extremely low latency end-to-end transmission, which in turn requires highly efficient and low-complexity screen compression. In this paper, we present a block-based low-complexity screen compression scheme, in which multiple block modes are adopted to exploit the intra- and inter-frame redundancies. In particular, we classify the intra-coded blocks to pictorial blocks and textual blocks using a proposed fast block classification algorithm, which exploits the discriminative features between the pictorial and the textual blocks. Then, we design a low-complexity, yet efficient, algorithm to compress the textual blocks. We use base colors and escape colors to represent and quantize the textual pixels, which not only achieves high compression ratios but also preserves a high quality on textual pixels. The two-dimensionally predictive index coding and hierarchical pattern coding technologies are used to exploit local spatial correlations and global pattern correlation, respectively. To further utilize the correlation between the luminance and chrominance channels, we propose a joint-channel index coding method. We compare the coding efficiency and the computational complexity of the proposed scheme against the standard image coding schemes such as JPEG, JPEG2000, and PNG, the compound image compressor HJPC, and the popular video coding standard H.264. We also compare the visual quality of the proposed scheme against H.264 intra coding, JPEG2000, and HJPC. The evaluation results show that the proposed scheme achieves superior or comparable compression efficiency with much lower complexity than other schemes in most of the cases.
Zhaotai Pan, Huifeng Shen, Yan Lu 0001, Shipeng Li 0001, Nenghai Yu
IEEE Trans. Circuits Syst. Video Technol.4
2013 Kinect-Like Depth Data Compression
abstract
Unlike traditional RGB video, Kinect-like depth is characterized by its large variation range and instability. As a result, traditional video compression algorithms cannot be directly applied to Kinect-like depth compression with respect to coding efficiency. In this paper, we propose a lossy Kinect-like depth compression framework based on the existing codecs, aiming to enhance the coding efficiency while preserving the depth features for further applications. In the proposed framework, the Kinect-like depth is reformed first by divisive normalized bilateral filter (DNBL) to suppress the depth noises caused by disparity normalization, and then block-level depth padding is implemented for invalid depth region compensation in collaboration with mask coding to eliminate the sharp variation caused by depth measurement failures. Before the traditional video coding, the inter-frame correlation of reformed depth is explored by proposed 2D+T prediction, in which depth volume is developed to simulate 3D volume to generate pseudo 3D prediction reference for depth uniqueness detection. The unique depth region, called active region is fed into the video encoder for traditional intra and inter prediction with residual coding, while the inactive region is skipped during depth coding. The experimental results demonstrate that our compression scheme can save 55%-85% in terms of bit cost and reduce coding complexity by 20%-65% in comparison with the traditional video compression algorithms. The visual quality of the 3D reconstruction is also improved after employing our compression scheme.
Jingjing Fu, Dan Miao, Weiren Yu, Shiqi Wang 0001, Yan Lu 0001, Shipeng Li 0001
IEEE Trans. Multim.6
2013 Interactive Multimodal Visual Search on Mobile Device
abstract
This paper describes a novel multimodal interactive image search system on mobile devices. The system, the Joint search with ImaGe, Speech, And Word Plus(JIGSAW+), takes full advantage of the multimodal input and natural user interactions of mobile devices. It is designed for users who already have pictures in their minds but have no precise descriptions or names to address them. By describing it using speech and then refining the recognized query by interactively composing a visual query using exemplary images, the user can easily find the desired images through a few natural multimodal interactions with his/her mobile device. Compared with our previous work JIGSAW, the algorithm has been significantly improved in three aspects: 1) segmentation-based image representation is adopted to remove the artificial block partitions; 2) relative position checking replaces the fixed position penalty; and 3) inverted index is constructed instead of brute force matching. The proposed JIGSAW+ is able to achieve 5% gain in terms of search performance and is ten times faster.
Houqiang Li, Tao Mei 0001, Jingdong Wang 0001, Shipeng Li 0001
IEEE Trans. Multim.5
2013 Towards Cross-Domain Learning for Social Video Popularity Prediction
abstract
Previous research on online media popularity prediction concluded that the rise in popularity of online videos maintains a conventional logarithmic distribution. However, recent studies have shown that a significant portion of online videos exhibit bursty/sudden rise in popularity, which cannot be accounted for by video domain features alone. In this paper, we propose a novel transfer learning framework that utilizes knowledge from social streams (e.g., Twitter) to grasp sudden popularity bursts in online content. We develop a transfer learning algorithm that can learn topics from social streams allowing us to model the social prominence of video content and improve popularity predictions in the video domain. Our transfer learning framework has the ability to scale with incoming stream of tweets, harnessing physical world event information in real-time. Using data comprising of 10.2 million tweets and 3.5 million YouTube videos, we show that social prominence of the video topic (context) is responsible for the sudden rise in its popularity where social trends have a ripple effect as they spread from the Twitter domain to the video domain. We envision that our cross-domain popularity prediction model will be substantially useful for various media applications that could not be previously solved by traditional multimedia techniques alone.
Suman Deb Roy, Tao Mei 0001, Wenjun Zeng 0001, Shipeng Li 0001
IEEE Trans. Multim.4
2013 Interaction Design for Mobile Visual Search
abstract
Mobile devices are becoming ubiquitous. People take pictures via their phone cameras to explore the world on the go. In many cases, they are concerned with the picture-related information. Understanding user intent conveyed by those pictures therefore becomes important. Existing mobile applications employ visual search to connect the captured picture with the physical world. However, they only achieve limited success due to the ambiguity nature of user intent in the picture-one picture usually contains multiple objects. By taking advantage of multitouch interactions on mobile devices, this paper presents a prototype of interactive mobile visual search, named TapTell, to help users formulate their visual intent more conveniently. This kind of search leverages limited yet natural user interactions on the phone to achieve more effective visual search while maintaining a satisfying user experience. We make three contributions in this work. First, we conduct a focus study on the usage patterns and concerned factors for mobile visual search, which in turn leads to the interactive design of expressing visual intent by gesture. Second, we introduce four modes of gesture-based interactions (crop, line, lasso, and tap) and develop a mobile prototype. Third, we perform an in-depth usability evaluation on these different modes, which demonstrates the advantage of interactions and shows that lasso is the most natural and effective interaction mode. We show that TapTell provides a natural user experience to use phone camera and gesture to explore the world. Based on the observation and conclusion, we also suggest some design principles for interactive mobile visual search in the future.
Jitao Sang 0001, Tao Mei 0001, Ying-Qing Xu, Changsheng Xu, Shipeng Li 0001
IEEE Trans. Multim.6
2013 Robust and accurate mobile visual localization and its applications
abstract
Mobile applications are becoming increasingly popular. More and more people are using their phones to enjoy ubiquitous location-based services (LBS). The increasing popularity of LBS creates a fundamental problem: mobile localization. Besides traditional localization methods that use GPS or wireless signals, using phone-captured images for localization has drawn significant interest from researchers. Photos contain more scene context information than the embedded sensors, leading to a more precise location description. With the goal being to accurately sense real geographic scene contexts, this article presents a novel approach to mobile visual localization according to a given image (typically associated with a rough GPS position). The proposed approach is capable of providing a complete set of more accurate parameters about the scene geo-context including the real locations of both the mobile user and perhaps more importantly the captured scene, as well as the viewing direction. To figure out how to make image localization quick and accurate, we investigate various techniques for large-scale image retrieval and 2D-to-3D matching. Specifically, we first generate scene clusters using joint geo-visual clustering, with each scene being represented by a reconstructed 3D model from a set of images. The 3D models are then indexed using a visual vocabulary tree structure. Taking geo-tags of the database image as prior knowledge, a novel location-based codebook weighting scheme proposed to embed this additional information into the codebook. The discriminative power of the codebook is enhanced, thus leading to better image retrieval performance. The query image is aligned with the models obtained from the image retrieval results, and eventually registered to a real-world map. We evaluate the effectiveness of our approach using several large-scale datasets and achieving estimation accuracy of a user's location within 13 meters, viewing direction within 12 degrees, and viewing distance within 26 meters. Of particular note is our showcase of three novel applications based on localization results: (1) an on-the-spot tour guide, (2) collaborative routing, and (3) a sight-seeing guide. The evaluations through user studies demonstrate that these applications are effective in facilitating the ideal rendezvous for mobile users.
Tao Mei 0001, Houqiang Li, Jiebo Luo 0001, Shipeng Li 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2012 Image search results refinement via outlier detection using deep contexts
abstract
Visual reranking has become a widely-accepted method to improve traditional text-based image search results. The main principle is to exploit the visual aggregation property of relevant images among top results so as to boost ranking scores of relevant images, by explicitly or implicitly detecting the confident relevant images, and propagating ranking scores among visually similar images. However, such a visual aggregation property does not always hold, and thus these schemes may fail. In this paper, we instead propose to filter out the most probable irrelevant images using deep contexts, which is the extra information that is not limited in the current search results. The deep contexts for each image consist of sets of images that are returned by searches using the queries formed by the textual context of this image. We compare the popularity of this image in the current search results and the deep contexts to check the irrelevance score. Then the irrelevance scores are propagated to the images whose useful textual context is missed. We formulate the two schemes together to reach a Markov random field, which is effectively solved by graph cuts. The key is that our scheme does not rely on the assumption that relevant images are visually aggregated among top results and is based on the observation that an outlier under the current query is likely to be more popular under some other query. After that, we perform graph reranking over filtered results to reorder them. Experimental results on the INRIA dataset show that our proposed method achieves significant improvements over previous approaches.
Junyang Lu, Jiazhen Zhou, Jingdong Wang 0001, Tao Mei 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
CVPR6
2012 Fast approximate k-means via cluster closures
abstract
K-means, a simple and effective clustering algorithm, is one of the most widely used algorithms in computer vision community. Traditional k-means is an iterative algorithm - in each iteration new cluster centers are computed and each data point is re-assigned to its nearest center. The cluster re-assignment step becomes prohibitively expensive when the number of data points and cluster centers are large. In this paper, we propose a novel approximate k-means algorithm to greatly reduce the computational complexity in the assignment step. Our approach is motivated by the observation that most active points changing their cluster assignments at each iteration are located on or near cluster boundaries. The idea is to efficiently identify those active points by pre-assembling the data into groups of neighboring points using multiple random spatial partition trees, and to use the neighborhood information to construct a closure for each cluster, in such a way only a small number of cluster candidates need to be considered when assigning a data point to its nearest cluster. Using complexity analysis, real data clustering, and applications to image retrieval, we show that our approach out-performs state-of-the-art approximate k-means algorithms in terms of clustering quality and efficiency.
Jing Wang 0068, Jingdong Wang 0001, Qifa Ke, Shipeng Li 0001
CVPR5
2012 Salient object detection for searched web images via global saliency
abstract
In this paper, we deal with the problem of detecting the existence and the location of salient objects for thumbnail images on which most search engines usually perform visual analysis in order to handle web-scale images. Different from previous techniques, such as sliding window-based or segmentation-based schemes for detecting salient objects, we propose to use a learning approach, random forest in our solution. Our algorithm exploits global features from multiple saliency indicators to directly predict the existence and the position of the salient object. To validate our algorithm, we constructed a large image database collected from Bing image search, that contains hundreds of thousands of manually labeled web images. The experimental results using this new database and the resized MSRA database [16] demonstrate that our algorithm outperforms previous state-of-the-art methods.
Peng Wang 0001, Jingdong Wang 0001, Jie Feng 0012, Hongbin Zha, Shipeng Li 0001
CVPR6
2012 Scalable k-NN graph construction for visual descriptors
abstract
The k-NN graph has played a central role in increasingly popular data-driven techniques for various learning and vision tasks; yet, finding an efficient and effective way to construct k-NN graphs remains a challenge, especially for large-scale high-dimensional data. In this paper, we propose a new approach to construct approximate k-NN graphs with emphasis in: efficiency and accuracy. We hierarchically and randomly divide the data points into subsets and build an exact neighborhood graph over each subset, achieving a base approximate neighborhood graph; we then repeat this process for several times to generate multiple neighborhood graphs, which are combined to yield a more accurate approximate neighborhood graph. Furthermore, we propose a neighborhood propagation scheme to further enhance the accuracy. We show both theoretical and empirical accuracy and efficiency of our approach to k-NN graph construction and demonstrate significant speed-up in dealing with large scale visual data.
Jing Wang 0068, Jingdong Wang 0001, Zhuowen Tu, Rui Gan, Shipeng Li 0001
CVPR6
2012 Probabilistic sequential POIs recommendation via check-in data
abstract
While on the go, people are using their phones as a personal concierge discovering what is around and deciding what to do. Mobile phone has become a recommendation terminal customized for individuals. While existing research predominantly focuses on one-step recommendation---recommending the next single activity according to current context, this work moves one step beyond by recommending a series of activities, which is a package of sequential Points of Interest (POIs). The recommended POIs are not only relevant to user context (i.e., current location, time, and check-in), but also personalized to his/her check-in history. We presents a probabilistic approach, which is highly motivated from a large-scale commercial mobile check-in data analysis, to ranking a list of sequential POI categories (e.g., "Japanese food" and "bar") and POIs (e.g., "I love sushi"). The approach enables users to plan consecutive activities on the move. Specifically, the probabilistic recommendation approach estimates the transition probability from one POI to another, conditioned on current context and check-in history in a Markov chain. To alleviate the discritization error and sparsity problem, we further introduce context collaboration and integrate prior information. Experiments on over 100k real-world check-in records and 20k POIs validate the effectiveness of the proposed approach.
Jitao Sang 0001, Tao Mei 0001, Jian-Tao Sun, Changsheng Xu, Shipeng Li 0001
SIGSPATIAL/GIS5
2012 Empowering Cross-Domain Internet Media with Real-Time Topic Learning from Social Streams
abstract
This paper aims to connect social media from disparate sources on the Internet by building a common topic space in-between, using which cross domain media recommendations can be realized on the web. The topic space is built and updated in real time by extending the Latent Dirichlet Allocation (LDA) model to cater to streaming online data. Our topical model, named Online Streaming LDA (OSLDA), is able to extract, learn, populate, and update the topic space in real time, scaling with streaming tweets. Based on the proposed topic space learned in real time, we present media recommendation applications that cannot be achieved by conventional media analysis techniques: (1) tweet enrichment by recommending related videos, and (2) popular video recommendation for featuring socially trending topical videos. We conduct experiments over a collection of 3.6 million tweets and 1.2 million click-through data from a video search engine. Our results show that the learned topic model plays a natural role connecting cross-domain social media, leading to a better user experience consuming social media.
Suman Deb Roy, Tao Mei 0001, Wenjun Zeng 0001, Shipeng Li 0001
ICME4
2012 Kinect-like depth denoising
abstract
Accuracy and stability of Kinect-like depth data is limited by its generating principle. In order to serve further applications with high quality depth, the preprocessing on depth data is essential. In this paper, we analyze the characteristics of the Kinect-like depth data by examing its generation principle and propose a spatial-temporal denoising algorithm taking into account its special properties. Both the intra-frame spatial correlation and the inter-frame temporal correlation are exploited to fill the depth hole and suppress the depth noise. Moreover, a divisive normalization approach is proposed to assist the noise filtering process. The 3D rendering results of the processed depth demonstrates that the lost depth is recovered in some hole regions and the noise is suppressed with depth features preserved.
Jingjing Fu, Shiqi Wang 0001, Yan Lu 0001, Shipeng Li 0001, Wenjun Zeng 0001
ISCAS4
2012 Texture-assisted Kinect depth inpainting
abstract
The emergence of Kinect facilitates the possibility of depth capture in real-time and with low cost by consumers. It also provides powerful tool and inspiration for researchers to engage in new array of technology development. However, the quality of the depth map captured from Kinect is still inadequate for many applications due to holes, noises and artifacts existing within the depth information. In this paper, we present a texture assisted Kinect depth inpainting framework, aiming at obtaining improved depth information. In this framework, the relationship between texture and depth is investigated, and the characteristics of depth are also exploited. More specifically, texture edge information is extracted to assist the depth inpainting. Furthermore, filtering and diffusion are designed for hole-filling and edge alignment. Experiment results demonstrate that the Kinect depth can be appropriately repaired in both smooth and edge region. Comparing with the original depth, the inpainted depth information enhances the quality of advanced processing such as 3D reconstruction.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
ISCAS4
2012 A low-latency transmission scheme for interactive screen sharing
abstract
Screen is becoming a new dimension in cloud computing platforms, and low latency screen sharing in unreliable networks is becoming more and more important. Due to the different characteristics between the screen codecs and video codecs, current transmission technologies on low-latency video streaming cannot be directly applied to screen sharing. So in this paper we first theoretically analyze the difference in latency performance between ARQ and FEC for the UDP-based screen sharing. Then, considering the characteristics of the main-stream screen codecs, we propose an improved ARQ scheme to decrease the transmission latency. The experimental results show that the proposed system achieves better latency performance than the popular systems.
Zhaotai Pan, Huifeng Shen, Yan Lu 0001, Shipeng Li 0001
ISCAS4
2012 Content-aware layered compound video compression
abstract
Compound video compression is crucial for remote control and data assessment. In this paper, we propose a content-aware layered video coding scheme as an attempt to efficiently compress the compound video. In this scheme, the compound video is analyzed and processed progressively at three pyramid levels: block, object and layer. Firstly, the compound video is analyzed by a block type classification technique to access each block's spatial and temporal properties. Secondly, the natural video object is detected adaptively in each frame based on the block type. Finally, the compound video content is distributed into different layers and specifically designed video coding algorithms are employed to compress each layer. Experiments demonstrate that our proposed scheme can preserve the advantages of the employed compression algorithms for each layer and outperform each of them in the compound video compression.
Shiqi Wang 0001, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Wen Gao 0001
ISCAS4
2012 When video search goes wrong: predicting query failure using search engine logs and visual search results
abstract
The recent increase in the volume and variety of video content available online presents growing challenges for video search. Users face increased difficulty in formulating effective queries and search engines must deploy highly effective algorithms to provide relevant results. Although lately much effort has been invested in optimizing video search engine results, relatively little attention has been given to predicting for which queries results optimization is most useful, i.e., predicting which queries will fail. Being able to predict when a video search query would fail is likely to make the video search result optimization more efficient and effective, improve the search experience for the user by providing support in the query formulation process and in this way boost the development of video search engines in general. While insight about a query's performance in general could be obtained using the well-known concept of query performance prediction (QPP), we propose a novel approach for predicting a failure of a video search query in the specific context of a search session. Our 'context-aware query failure' prediction approach uses a combination of 'user indicators' and 'engine indicators' to predict whether a particular query is likely to fail in the context of a particular search session. User indicators are derived from the search log and capture the patterns of query (re)formulation behavior and the click-through data of a user during a typical video search session. Engine indicators are derived from the video search results list and capture the visual variance of search results that would be offered to the user for the given query. We validate our approach experimentally on a test set containing 1+ million video search queries and show its effectiveness compared to a set of conventional QPP baselines. Our approach achieves a 13% relative improvement over the baseline.
Christoph Kofler, Linjun Yang, Martha A. Larson, Tao Mei 0001, Alan Hanjalic, Shipeng Li 0001
ACM Multimedia6
2012 Finding perfect rendezvous on the go: accurate mobile visual localization and its applications to routing
abstract
While on the go, more and more people are using their phones to enjoy ubiquitous location-based services (LBS). One of the fundamental problems of LBS is localization. Researchers are now investigating ways to use a phone-captured image for localization as it contains more scene context information than the embedded sensors. In this paper, we present a novel approach to mobile visual localization that accurately senses geographic scene context according to the current image (typically associated with a rough GPS position). Unlike most existing visual localization methods, the proposed approach is capable of providing a complete set of more accurate parameters about the scene geo---including the actual locations of both the mobile user and perhaps more importantly the captured scene along with the viewing direction. Our approach takes advantage of advanced techniques for large-scale image retrieval and 3D model reconstruction from photos. Specifically, we first perform joint geo-visual clustering in the cloud to generate scene clusters, with each scene represented by a 3D model. The 3D scene models are then indexed using a visual vocabulary tree structure. The phone-captured image is used to retrieve the relevant scene models, then aligned with the models, and further registered to the real-world map. Our approach achieves an estimation accuracy of user location within 14 meters, viewing direction within 9 degrees, and scene location within 21 meters. Such a complete set of accurate geo-parameters can lead to various LBS applications for routing that cannot be achieved with most existing methods. In particular, we showcase three novel applications: 1) accurate self-localization, 2) collaborative localization for rendezvous routing, and 3) routing for photographing. The evaluations through user studies indicate these applications are effective for facilitating the perfect rendezvous for mobile users.
Tao Mei 0001, Jiebo Luo 0001, Houqiang Li, Shipeng Li 0001
ACM Multimedia5
2012 Browse-to-search
abstract
This demonstration presents a novel interactive online shopping application based on visual search technologies. When users want to buy something on a shopping site, they usually have the requirement of looking for related information from other web sites. Therefore users need to switch between the web page being browsed and other websites that provide search results. The proposed application enables users to naturally search products of interest when they browse a web page, and make their even causal purchase intent easily satisfied. The interactive shopping experience is characterized by: 1) in session---it allows users to specify the purchase intent in the browsing session, instead of leaving the current page and navigating to other websites; 2) in context---the browsed web page provides implicit context information which helps infer user purchase preferences; 3) in focus---users easily specify their search interest using gesture on touch devices and do not need to formulate queries in search box; 4) natural-gesture inputs and visual-based search provides users a natural shopping experience. The system is evaluated against a data set consisting of several millions commercial product images.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, David Dagan Feng, Jian-Tao Sun, Shipeng Li 0001
ACM Multimedia8
2012 SocialTransfer: cross-domain transfer learning from social streams for media applications
abstract
The usage and applications of social media have become pervasive. This has enabled an innovative paradigm to solve multimedia problems (e.g., recommendation and popularity prediction), which are otherwise hard to address purely by traditional approaches. In this paper, we investigate how to build a mutual connection among the disparate social media on the Internet, using which cross-domain media recommendation can be realized. We accomplish this goal through SocialTransfer---a novel cross-domain real-time transfer learning framework. While existing transfer learning methods do not address how to utilize the real time social streams, our proposed SocialTransfer is able to effectively learn from social streams to help multimedia applications, assuming an intermediate topic space can be built across domains. It is characterized by two key components: 1) a topic space learned in real time from social streams via Online Streaming Latent Dirichlet Allocation (OSLDA), and 2) a real-time cross-domain graph spectra analysis based transfer learning method that seamlessly incorporates learned topic models from social streams into the transfer learning framework. We present as use cases of \emph{SocialTransfer} two video recommendation applications that otherwise can hardly be achieved by conventional media analysis techniques: 1) socialized query suggestion for video search, and 2) socialized video recommendation that features socially trending topical videos. We conduct experiments on a real-world large-scale dataset, including 10.2 million tweets and 5.7 million YouTube videos and show that \emph{SocialTransfer} outperforms traditional learners significantly, and plays a natural and interoperable connection across video and social domains, leading to a wide variety of cross-domain applications.
Suman Deb Roy, Tao Mei 0001, Wenjun Zeng 0001, Shipeng Li 0001
ACM Multimedia4
2012 Query-driven iterated neighborhood graph search for large scale indexing
abstract
In this paper, we address the approximate nearest neighbor (ANN) search problem over large scale visual descriptors. We investigate a simple but very effective approach, neighborhood graph search, which constructs a neighborhood graph to index the data points and conducts a local search, expanding neighborhoods with a best-first manner, for ANN search. Our empirical analysis shows that neighborhood expansion is very efficient, with O(1) cost, for a new NN candidate location, and has high chances to locate true NNs and hence it usually performs well. However, it often gets sub-optimal solutions since local search only checks the neighborhood of the current solution, or conducts exhaustive and continuous neighborhood expansions to find better solutions, which deteriorates the query efficiency.
Jingdong Wang 0001, Shipeng Li 0001
ACM Multimedia2
2012 Scalable similar image search by joint indices
abstract
Text-based image search is able to return desired images for simple queries, but has limited capabilities in finding images with additional visual requirements. As a result, an image is usually used to help describe the appearance requirements. In this demonstration, we show a similar image search system that can support the joint textual and visual query. We present an efficient and effective indexing algorithm, neighborhood graph index, which is suitable for millions of images, and use it to organize joint inverted indices to search over billions of images.
Jing Wang 0068, Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia4
2012 Color filter for image search
abstract
Image search relying on surrounding texts can return reliably relevant images to some extent. Most recent efforts are focusing on utilizing visual contents to help users find images with specific visual requirements. In this demonstration, we show a color filter scheme for image search, which enables users to find images containing objects or scenes with their interested color. Color is one of the most crucial cues in describing visual contents and has been frequently used in various applications. The key components in developing our demo include salient object detection and perceptual color naming. The color filter in Microsoft Bing image search is developed using these techniques.
Peng Wang 0001, Dongqing Zhang, Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia6
2012 Local visual words coding for low bit rate mobile visual search
abstract
Mobile visual search has attracted extensive attention for its huge potential for numerous applications. Research on this topic has been focused on two schemes: sending query images, and sending compact descriptors extracted on mobile phones. The first scheme requires about 30-40KB data to transmit, while the second can reduce the bit rate by 10 times. In this paper, we propose a third scheme for extremely low bit rate mobile visual search, which sends compressed visual words consisting of vocabulary tree histogram and descriptor orientations rather than descriptors. This scheme can further reduce the bit rate with few extra computational costs on the client. Specifically, we store a vocabulary tree and extract visual descriptors on the mobile client. A light-weight pre-retrieval is performed to obtain the visited leaf nodes in the vocabulary tree. The orientation of each local descriptor and the tree histogram are then encoded to be transmitted to server. Our new scheme transmits less than 1KB data, which reduces the bit rate in the second scheme by 3 times, and obtains about 30% improvement in terms of search accuracy over the traditional Bag-of-Words baseline. The time cost is only 1.5 secs on the client and 240 msecs on the server.
Shiyang Lu, Tao Mei 0001, Jian Zhang 0002, Shipeng Li 0001
ACM Multimedia5
2012 A bag-of-objects retrieval model for web image search
abstract
Image search reranking has been an active research topic in recent years to boost the performance of the existing web image search engine which is mostly based on textual metadata of images. Various approaches have been proposed to rerank images for general queries and argue that, they may not necessarily be optimal for queries in specific domain, e.g., object queries, since the reranking algorithms are operated on whole images, instead of the relevant parts of images. In this paper, we propose a novel bag-of-objects retrieval model for image search reranking of object queries. Firstly, we employ a common object discovery algorithm to discover query-relevant objects from the search results returned by text-based image search engine. Then, the query and its result images are represented as a language model on the query relevant object vocabulary, based on which the ranking function can be derived. As the common object discovery is unreliable and may introduce noises, we propose to incorporate the attributes of the discovered objects, e.g., size, position, etc., into the ranking function through a linear model, and the weights on the object attributes can be learned. The experiments on two subsets of Web Queries dataset comprising object queries demonstrate that our approach can significantly outperform the existing reranking methods on object queries.
Yang Yang 0222, Linjun Yang, Gangshan Wu, Shipeng Li 0001
ACM Multimedia4
2012 Interactive mobile visual search for social activities completion using query image contextual model
abstract
Mobile devices are ubiquitous. People use their phones as a personal concierge not only discovering information but also searching for particular interest on-the-go and making decisions. This brings a new horizon for multimedia retrieval on mobile. While existing efforts have predominantly focused on understanding textual or a voice query, this paper presents a new perspective which understands visual queries captured by the built-in camera such that mobile-based social activities can be recommended for users to complete. In this work, a query image-based contextual model is proposed for visual search. A mobile user can take a photo and naturally indicate an object-of-interest within the photo via circle based gesture called “O” gesture. Both selected object-of-interest region as well as surrounding visual context in photo are used in achieving a search-based recognition by retrieving similar images based on a large-scale of visual vocabulary tree. Consequently, social activities such as visiting contextually relevant entities (i.e., local businesses) are recommended to the users based on their visual queries and GPS location. Along with the proposed method, an exemplary real application has been developed on Windows Phone 7 devices and evaluated with a wide variety of scenarios on million-scale image database. To test the performance of proposed mobile visual search model, extensive experimentation has been conducted and compared with state-of-the-art algorithms in content-based image retrieval (CBIR) domain.
Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001
MMSP5
2012 Layered compression for high dynamic range depth
abstract
With the rapid development of depth data acquisition technology, the high precision depth becomes much easier to access in real-time by depth sensors, and the generated high dynamic range (HDR) depth is widely adopted to benefit the depth assistant applications. Accordingly, the HDR depth compression becomes essential for the efficient depth storage and transmission. In this paper, we introduce a layered compression framework for HDR depth to achieve efficient and low-complexity depth compression. To leverage the state-of-art 8-bit image/video encoders, the HDR depth is partitioned into two layers: most significant bit (MSB) layer and least significant bit (LSB) layer. For MSB layer, an error controllable pixel domain encoding scheme is proposed to guarantee the compatibility for existing 8-bit codec by controlling quantization errors added back to LSB layer. Meanwhile, the efficient major color extraction and adaptive quantization enhance the coding performance of MSB layer. For LSB layer, the layer data with limited dynamic range is compressed by normal 8-bit image/video based encoding scheme. The experimental results demonstrate that our coding scheme can achieve real-time depth compression with the satisfactory reconstruction quality. The encoding time is less than 31ms/frame and the decoding time is around 20ms/frame in average. Our compression scheme can be easily integrated into the real-time depth transmission system.
Dan Miao, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001, Chang Wen Chen
VCIP4
2012 A low-complexity screen compression scheme
abstract
This paper presents a block-based low-complexity screen compression scheme. In this scheme, the input screen is split into non-overlapping blocks which are classified as pictorial blocks and textual blocks. We design a low-complexity yet efficient algorithm to compress the textual blocks. We use base colors plus escape pixels to represent and quantize the text pixels, and such quantization mechanism not only achieves high compression efficiency but also keeps low encoding/decoding complexity. We also propose the two-direction predictive index coding and hierarchical pattern coding technologies to utilize the local spatial correlation and the global pattern correlation for text pixels. In addition, to utilize the correlation between the luminance and chrominance channels, we propose a joint-channel index coding method to further improve the compression efficiency. The compression efficiency and complexity of the proposed method is evaluated against the popular image codecs JPEG/JPEG200 and PNG, a recent published screen compression scheme HJPC as well as the popular video codec H.264.
Zhaotai Pan, Huifeng Shen, Yan Lu 0001, Nenghai Yu, Shipeng Li 0001
VCIP5
2012 Flickr Distance: A Relationship Measure for Visual Concepts
abstract
This paper proposes the Flickr Distance (FD) to measure the visual correlation between concepts. For each concept, a collection of related images are obtained from the Flickr website. We assume that each concept consists of several states, e.g., different views, different semantics, etc., which are considered as latent topics. Then a latent topic visual language model (LTVLM) is built to capture these states. The Flickr distance between two concepts is defined as the Jensen-Shannon (J-S) divergence between their LTVLM. Differently from traditional conceptual distance measurements, which are based on Web textual documents, FD is based on the visual information. Comparing with the WordNet distance, FD can easily scale up with the increasing size of the conceptual corpus. Comparing with the Google Distance (NGD) and Tag Concurrence Distance (TCD), FD uses the visual information and can properly measure the conceptual relations. We apply FD to multimedia-related tasks and find methods based on FD significantly outperform those based on NGD and TCD. With the FD measurement, we also construct a large-scale visual conceptual network (VCNet) to store the knowledge of conceptual relationship. Experiments show that FD is more coherent to human cognition and it also outperforms text-based distances in real-world applications.
Lei Wu 0017, Xian-Sheng Hua 0001, Nenghai Yu, Wei-Ying Ma, Shipeng Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2012 ImageSense: Towards contextual image advertising
abstract
The daunting volumes of community-contributed media contents on the Internet have become one of the primary sources for online advertising. However, conventional advertising treats image and video advertising as general text advertising by displaying relevant ads based on the contents of the Web page, without considering the inherent characteristics of visual contents. This article presents a contextual advertising system driven by images, which automatically associates relevant ads with an image rather than the entire text in a Web page and seamlessly inserts the ads in the nonintrusive areas within each individual image. The proposed system, called ImageSense , supports scalable advertising of, from root to node, Web sites, pages, and images. In ImageSense, the ads are selected based on not only textual relevance but also visual similarity, so that the ads yield contextual relevance to both the text in the Web page and the image content. The ad insertion positions are detected based on image salience, as well as face and text detection, to minimize intrusiveness to the user. We evaluate ImageSense on a large-scale real-world images and Web pages, and demonstrate the effectiveness of ImageSense for online image advertising.
Tao Mei 0001, Lusong Li, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2011 When recommendation meets mobile: contextual and personalized recommendation on the go
abstract
Mobile devices are becoming ubiquitous. People use their phones as a personal concierge discovering and making decisions anywhere and anytime. Understanding user intent on the go therefore becomes important for task completion on the phone. While existing efforts have predominantly focused on understanding the explicit user intent expressed by a textual or voice query, this paper presents an approach to context-aware and personalized entity recommendation which understands the implicit intent without any explicit user input on the phone. The approach, highly motivated from a large-scale mobile click-through analysis, is able to rank both the entity types and the entities within each type (here an entity is a local business, e.g., "I love sushi," while an entity type is a category, e.g., "restaurant"). The recommended entity types and entities are relevant to both user context (past behaviors) and sensor context (time and geo-location). Specifically, it estimates the generation probability of an entity by a given user conditioned on the current context in a probabilistic framework. A random-walk propagation is then employed to refine the estimated probability by mining the temporal patterns among entities. We deploy a recommendation application based on the proposed approach on Window Phone 7 devices. We evaluate recommendation performance on 10 thousand mobile clicks, as well as user experience through subjective user studies. We show that the application is effective to facilitate the exploration and discovery of surroundings for mobile users.
Jinfeng Zhuang, Tao Mei 0001, Steven C. H. Hoi, Ying-Qing Xu, Shipeng Li 0001
UbiComp5
2011 Complementary hashing for approximate nearest neighbor search
abstract
Recently, hashing based Approximate Nearest Neighbor (ANN) techniques have been attracting lots of attention in computer vision. The data-dependent hashing methods, e.g., Spectral Hashing, expects better performance than the data-blind counterparts, e.g., Locality Sensitive Hashing (LSH). However, most data-dependent hashing methods only employ a single hash table. When higher recall is desired, they have to retrieve exponentially growing number of hash buckets around the bucket containing the query, which may drag down the precision rapidly. In this paper, we propose a so-called complementary hashing approach, which is able to balance the precision and recall in a more effective way. The key idea is to employ multiple complementary hash tables, which are learned sequentially in a boosting manner, so that, given a query, its true nearest neighbors missed from the active bucket of one hash table are more likely to be found in the active bucket of the next hash table. Compared with LSH that also can exploit multiple hash tables, our approach is more effective to find true NNs, thanks to the complementarity property of the hash tables from our approach. Experimental results on large scale ANN search show that the proposed method significantly improves the performance and outperforms the state-of-the-art.
Jingdong Wang 0001, Shipeng Li 0001, Nenghai Yu
ICCV5
2011 Browser-friendly hybrid codec for compound image compression
abstract
In this paper, we present a browser-friendly hybrid JPEG/PNG codec for compound images. First we employ a simple yet efficient block classification algorithm to identify the blocks to pictorial and textural ones. And then JPEG is used to encode the pictorial blocks and PNG is used for textual blocks. Our evaluation results show that our codec significantly outperforms JPEG and PNG in terms of rate-distortion performance, and also it outperforms JPEG, JPEG2000 and DjVu in terms of visual quality. Moreover, since JPEG and PNG are naturally supported by modern browsers, the coded images generated from our proposed coder can be natively supported by browsers and possible to be widely deployed in Web applications.
Zhaotai Pan, Huifeng Shen, Yan Lu 0001, Shipeng Li 0001
ISCAS4
2011 Million-scale near-duplicate video retrieval system
abstract
In this paper, we present a novel near-duplicate video retrieval system serving one million web videos. To achieve both the effectiveness and efficiency, a visual word based approach is proposed, which quantizes each video frame into a word and represents the whole video as a bag of words. The system can respond to a query in 41ms with 78.4% MAP on average.
Linjun Yang, Wei Ping, Tao Mei 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia7
2011 The role of attractiveness in web image search
abstract
Existing web image search engines are mainly designed to optimize topical relevance. However, according to our user study, attractiveness is becoming a more and more important factor for web image search engines to satisfy users' search intentions. Important as it can be, web image attractiveness from the search users' perspective has not been sufficiently recognized in both the industry and the academia. In this paper, we present a definition of web image attractiveness with three levels according to the end users' feedback, including perceptual quality, aesthetic sensitivity and affective tune. Corresponding to each level of the definition, various visual features are investigated on their applicability to attractiveness estimation of web images. To further deal with the unreliability of visual features induced by the large variations of web images, we propose a contextual approach to integrate the visual features with contextual cues mined from image EXIF information and the associated web pages. We explore the role of attractiveness by applying it to various stages of a web image search engine, including the online ranking and the interactive reranking, as well as the offline index selection. Experimental results on three large-scale web image search datasets demonstrate that the incorporation of attractiveness can bring more satisfaction to 80% of the users for ranking/reranking search results and 30.5% index coverage improvement for index selection, compared to the conventional relevance based approaches.
Bo Geng, Linjun Yang, Chao Xu 0006, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia5
2011 Contextual image search
abstract
In this paper, we propose a novel image search scheme, contextual image search. Different from conventional image search schemes that present a separate interface (e.g., text input box) to allow users to submit a query, the new search scheme enables users to search images by only masking a few words when they are reading through Web pages or other documents. Rather than merely making use of the explicit query input that is often not sufficient to express user's search intent, our approach explores the context information to better understand the search intent with two key steps: query augmenting and search results reranking using context, and expects to obtain better search results. Beyond contextual Web search, the context in our case is much richer and includes images besides texts. In addition to this type of search scheme, called contextual image search with text input, we also present another type of scheme, called contextual image search with image input, to allow users to select an image as the search query from Web pages or other documents they are reading. The key idea is to use the search-to-annotation technique and the contextual textual query mining scheme to determine the corresponding textual query, to finally get semantically similar search results. Experiments show that the proposed schemes make image search more convenient and the search results are more relevant to user intention.
Wenhao Lu, Jingdong Wang 0001, Xian-Sheng Hua 0001, Shengjin Wang, Shipeng Li 0001
ACM Multimedia5
2011 JIGSAW: interactive mobile visual search with multimodal queries
abstract
The traditional text-based visual search has not been sufficiently improved over the years to accommodate the new emerging demand of mobile users. While on the go, searching on one's phone is becoming pervasive. This paper presents an innovative application for mobile phone users to facilitate their visual search experience. By taking advantage of smart phone functionalities such as multi-modal and multi-touch interactions, users can more conveniently formulate their search intent, and thus search performance can be significantly improved. The system, called JIGSAW (Joint search with ImaGe, Speech, And Words), represents one of the first attempts to create an interactive and multi-modal mobile visual search application. The key of JIGSAW is the composition of an exemplary image query generated from the raw speech via multi-touch user interaction, as well as the visual search based on the exemplary image. Through JIGSAW, users can formulate their search intent in a natural way like playing a jigsaw puzzle on the phone screen: 1) a user speaks a natural sentence as the query, 2) the speech is recognized and transferred to text which is further decomposed to keywords through entity extraction, 3) the user selects preferred exemplary images that can visually represent his/her intent and composes a query image via multi-touch, and 4) the composite image is then used as a visual query to search similar images. We have deployed JIGSAW on a real-world phone system, evaluated the performance on one million images, and demonstrated that it is an effective complement to existing mobile visual search applications.
Tao Mei 0001, Jingdong Wang 0001, Houqiang Li, Shipeng Li 0001
ACM Multimedia5
2011 Hybrid image summarization
abstract
In this paper, we address a problem of managing tagged images with hybrid summarization. We formulate this problem as finding a few image exemplars to represent the image set semantically and visually and solve it in a hybrid way by exploiting both visual and textual information associated with images. We propose a novel approach, called Homogeneous and Heterogeneous Message Propagation (H2MP), which extends affinity propagation that only works over homogeneous relations to heterogeneous relations. The summary obtained by our approach is both visually and semantically satisfactory. The experimental results demonstrate the effectiveness and efficiency of the proposed approach.
Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia4
2011 Video-based image retrieval
abstract
Likely variations in the capture conditions (e.g. light, blur, scale, occlusion) and in the viewpoint between the query image and the images in the collection are the factors due to which image retrieval based on the Query-by-Example (QBE) principle is still not reliable enough. In this paper, we propose a novel QBE-based image retrieval system where users are allowed to submit a short video clip as a query to improve the retrieval reliability. Improvement is achieved by integrating the information about different viewpoints and conditions under which object and scene appearances can be captured across different video frames. Rich information extracted from a video can be exploited to generate a more complete query representation than in the case of a single-image query and to improve the relevance of the retrieved results. Our experimental results show that video-based image retrieval (VBIR) is significantly more reliable than the retrieval using a single image as a query.
Linjun Yang, Alan Hanjalic, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia5
2011 TapTell: understanding visual intents on-the-go
abstract
This demonstration presents a mobile-based visual recognition and recommendation application on Windows Phone 7 called TapTell. This is different from other mobile-based visual search mechanisms which merely focus on the search process. TapTell firstly discovers and understands users' visual intents via a circle based natural user interaction called "O" gestures. Following, a Tap action is operated to choose the "O" gestured regions. The context-aware visual search mechanism is utilized for recognizing the intents and associating them with indexed metadata. Finally, the "Tell" action recommends relevant entities utilizing contextual information. The TapTell system has been evaluated at different scenarios on million scale images.
Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001
ACM Multimedia5
2011 Modeling social strength in social media community via kernel-based learning
abstract
Modeling continuous social strength rather than conventional binary social ties in the social network can lead to a more precise and informative description of social relationship among people. In this paper, we study the problem of social strength modeling (SSM) for the users in a social media community, who are typically associated with diverse form of data. In particular, we take Flickr---the most popular online photo sharing community---as an example, in which users are sharing their experiences through substantial amounts of multimodal contents (e.g., photos, tags, geo-locations, friend lists) and social behaviors (e.g., commenting and joining interest groups). Such heterogeneous data in Flickr bring opportunities yet challenges to the research community for SSM. One of the key issues in SSM is how to effectively explore the heterogeneous data and how to optimally combine them to measure the social strength. In this paper, we present a kernel-based learning to rank framework for inferring the social strength of Flickr users, which involves two learning stages. The first stage employs a kernel target alignment algorithm to integrate the heterogeneous data into a holistic similarity space. With the learned kernel, the second stage rectifies the pair-wise learning to rank approach to estimating the social strength. By learning the social strength graph, we are able to conduct collaborative recommendation and collective classification. The promising results show that the learning-based approach is effective for SSM. Despite being focused on Flickr, our technique can be applied to model social strength of users in any other social media community.
Jinfeng Zhuang, Tao Mei 0001, Steven C. H. Hoi, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia5
2011 Tap-to-search: Interactive and contextual visual search on mobile devices
abstract
Mobile visual search has been an emerging topic for both researching and industrial communities. Among various methods, visual search has its merit in providing an alternative solution, where text and voice searches are not applicable. This paper proposes an interactive “tap-to-search” approach utilizing both individual's intention in selecting interested regions via “tap” actions on the mobile touch screen, as well as a visual recognition by search mechanism in a large-scale image database. Automatic image segmentation technique is applied in order to provide region candidates. Visual vocabulary tree based search is adopted by incorporating rich contextual information which are collected from mobile sensors. The proposed approach has been conducted on an image dataset with the scale of two million. We demonstrated that using GPS contextual information, such an approach can further achieve satisfactory results with the standard information retrieval evaluation.
Ning Zhang 0023, Tao Mei 0001, Xian-Sheng Hua 0001, Ling Guan, Shipeng Li 0001
MMSP5
2011 Contextual Video Recommendation by Multimodal Relevance and User Feedback
abstract
With Internet delivery of video content surging to an unprecedented level, video recommendation, which suggests relevant videos to targeted users according to their historical and current viewings or preferences, has become one of most pervasive online video services. This article presents a novel contextual video recommendation system, called VideoReach, based on multimodal content relevance and user feedback. We consider an online video usually consists of different modalities (i.e., visual and audio track, as well as associated texts such as query, keywords, and surrounding text). Therefore, the recommended videos should be relevant to current viewing in terms of multimodal relevance. We also consider that different parts of videos are with different degrees of interest to a user, as well as different features and modalities have different contributions to the overall relevance. As a result, the recommended videos should also be relevant to current users in terms of user feedback (i.e., user click-through). We then design a unified framework for VideoReach which can seamlessly integrate both multimodal relevance and user feedback by relevance feedback and attention fusion. VideoReach represents one of the first attempts toward contextual recommendation driven by video content and user click-through, without assuming a sufficient collection of user profiles available. We conducted experiments over a large-scale real-world video data and reported the effectiveness of VideoReach.
Tao Mei 0001, Bo Yang 0008, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Trans. Inf. Syst.4
2010 Low-cost realtime screen sharing to multiple clients
abstract
In this paper, we propose an efficient encoding solution for the screen-sharing applications with multiple clients connected. We first present a lightweight screen codec to compress the complicated screen content. Based on the architecture of the proposed codec, we propose a one-pass encoding algorithm for multiple bit-rates. The one-pass encoding algorithm enables the host to only involve in one-pass encoding process for multiple bitrates, and as a result the computation complexity of the host is decreased significantly. Specially, in the limited computing-resource case, the one-pass encoding algorithm can improve the screen-sharing performance by about 43%, in terms of framerates the clients can get. In addition, based on our codec, we propose a fast screen transcoding scheme for the data-center based screen-sharing applications.
Huifeng Shen, Yan Lu 0001, Feng Wu 0001, Shipeng Li 0001
ICME4
2010 A proxy-based mobile web browser
abstract
In this paper, we present a proxy-based mobile web browser with rich experiences. We use the server-side web parsing and rendering to leverage the browser computing logic. We use a composite screen format to represent the display of the web content, incorporating the web background screen and the dynamic web objects. And then we employ a slice-based screen encoding scheme to efficiently compress the web background screen. Besides the display screen of the web content, we also send the side information of the web objects to enable the designed object-level interaction mechanisms. The experimental results show that our browser can achieve the superior browsing speed, compared with the native browser and yield much better visual quality than the existing proxy-based browser
Huifeng Shen, Zhaotai Pan, Haicheng Sun, Yan Lu 0001, Shipeng Li 0001
ACM Multimedia5
2010 ReDi: an interactive virtual display system for ubiquitous devices
abstract
In this paper, we present an interactive virtual display system to facilitate the ubiquitous user interaction with heterogeneous devices. By using small-size programmable hardware and wearable sensors, any display device (referred to as display surface) can act as a thin client for users to interact with the different remote devices. Under a flexible system architecture for local and remote devices' communication and collaboration, several techniques, such as adaptive screen compression, interactive ROI control, and accelerometer-based pointing input, are developed to improve the system performance and user experience. Evaluations show that the proposed system can efficiently utilize the remote computing resources and local display capabilities of ubiquitous devices, which will greatly benefit interactive multimedia applications.
Wen Sun 0001, Yan Lu 0001, Shipeng Li 0001
ACM Multimedia3
2010 Media 2.0 - The New Media Revolution?
Shipeng Li 0001
MMM1
2010 Image search by concept map
abstract
In this paper, we present a novel image search system, image search by concept map. This system enables users to indicate not only what semantic concepts are expected to appear but also how these concepts are spatially distributed in the desired images. To this end, we propose a new image search interface to enable users to formulate a query, called concept map, by intuitively typing textual queries in a blank canvas to indicate the desired spatial positions of the concepts. In the ranking process, by interpreting each textual concept as a set of representative visual instances, the concept map query is translated into a visual instance map, which is then used to evaluate the relevance of the image in the database. Experimental results demonstrate the effectiveness of the proposed system.
Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
SIGIR4
2010 Interactive image search by 2D semantic map
abstract
In this demo, we present a novel interactive image search system, image search by 2D semantic map. This system enables users to indicate what semantic concepts are expected to appear and even how these concepts are spatially distributed in the desired images. To this end, we design an intuitive interface for users to formulate a query in the form of 2D semantic map, called concept map, by typing textual queries in a blank canvas. In the ranking process, by interpreting each textual concept as a set of representative visual instances, the concept map query is translated into a visual instance map, which is then used for comparison with the images in the database. Besides, in this demo, we also show an image search system with a simplest semantic map, a 2D color map, where the concepts are limited from the colors.
Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
WWW4
2010 High-Dynamic-Range Texture Compression for Rendering Systems of Different Capacities
abstract
In this paper, we propose a novel approach for high-dynamic-range (HDR) texture compression (TC) suitable for rendering systems of different capacities. Based on the previously proposed DHTC scheme, we first work out an improved joint-channel compression framework, which is robust and flexible enough to provide compressed HDR textures at different bit rates. Then, two compressed HDR texture formats based on the proposed framework are developed. The 8 bpp format is of near lossless visual quality, improving upon known state-of-the-art algorithms. And, to our knowledge, the 4 bpp format is the first workable 4 bpp solution with good quality. We also show that HDR textures in the proposed 4 bpp and 8 bpp formats can compose a layered architecture in the texture consumption pipeline, to significantly save the memory bandwidth and storage in real-time rendering. In addition, the 8 bpp format can also be used to handle traditional low dynamic range (LDR) RGBA textures. Our scheme exhibits a practical solution for compressing HDR textures at different rates and LDR textures with alpha maps.
Wen Sun 0001, Yan Lu 0001, Feng Wu 0001, Shipeng Li 0001, John Tardif
IEEE Trans. Vis. Comput. Graph.4
2009 How Can Intra Correlation Be Exploited Better?
abstract
Summary form only given. This paper studies how to better exploit intra correlation to compress images. In general, edge and texture areas of images exhibit strong anisotropic property. The correlation among samples is determined by not only their distance but also the link orientation. Traditional transforms are not efficient on handling this anisotropic correlation. Therefore, in this paper we propose a directional filtering transform (dFT, in order to distinguish from the common usage on DFT) to exploit local anisotropic correlation among samples. Similar to directional prediction in H.264 intra-frame coding, but it adopts the hierarchal structure to decrease the distance between samples to be predicted and that are used for prediction. From another viewpoint, the dFT prediction resembles the directional wavelet transform without update, which can take both intra-block and inter-block correlations into account.
Feng Wu 0001, Xiulian Peng, Jizheng Xu, Shipeng Li 0001
DCC4
2009 Level embedded medical image compression based on value of interest
abstract
In this paper, we propose a novel compression scheme for medical images with extended bit depth. Our scheme efficiently utilizes the VOI (value of interest) settings of the medical images, so that only a small part of the coded bit-stream is needed at the decoder to lossless display the original image with recommended VOI parameters. Gray-Golomb coding is used to support VOI progressive transmission. In addition, the pixel domain bit-plane coding enables accurate distortion estimation at pixel level for any truncated bit-stream. Compared to JPEG2000, our scheme achieves much higher VOI performance with very slight overall bit-rate increase.
Wen Sun 0001, Yan Lu 0001, Feng Wu 0001, Shipeng Li 0001
ICIP4
2009 Real-time screen image scaling and its GPU acceleration
abstract
In this paper, we propose a simple and effective scheme for screen image scaling, targeting at real-time applications like remote desktop and screen sharing. To balance visual quality and complexity, we build our scheme upon bilinear interpolation and devise a content adaptive post-processing method. Our scheme keeps the text/graphics regions of screen images as sharp as the original while avoiding magnifying noise and artifacts in the picture regions. Further, the involved operations are quite suitable for hardware acceleration. When implemented on commodity GPUs, our scheme achieves a frame rate more than 300 fps, which is fast enough for practical use.
Wen Sun 0001, Yan Lu 0001, Feng Wu 0001, Shipeng Li 0001
ICIP4
2009 Utilizing affective analysis for efficient movie browsing
abstract
Because of the fast increasing number of movies and long time span each movie lasts, novel methods should be developed to help users browse movies and find their desired clips effectively. Affective information in movies is closely related with users' experiences and preferences. Therefore, in this paper, we analyze the affective states of movies and propose affective information based movie browsing. Affective movie content analysis is challenging due to the great variety of movie contents and styles. To address this challenge, we first extract rich audio-visual features. Then, feature selection and affective modeling are carried out to select and map effective features into corresponding affective states. Finally, we propose novel Affective Visualization techniques which intuitively visualize affective states to achieve efficient and user-friendly movie browsing. Experiments on representative movie dataset demonstrate the effectiveness of our proposed methods.
Shiliang Zhang, Qi Tian 0001, Qingming Huang, Wen Gao 0001, Shipeng Li 0001
ICIP5
2009 A lexica family with small semantic gap
abstract
Defining a lexicon of high-level concepts is the first step for data collection and model construction in concept-based image retrieval. Differences of semantic gaps among concepts are well worth considering. By measuring consistency in visual space and textual space, concepts with small semantic gap can be obtained. Considering so many diverse concepts in large-scale image dataset, we construct a lexica family of high-level concepts with small semantic gap based on different low-level features and different consistency measurements. In this lexica family, the lexica are independent to each other and mutually complementary. It provides helpful suggestions about data collection, feature selection and search model construction for large-scale image retrieval.
Jiemin Liu, Qi Tian 0001, Yijuan Lu, Changhu Wang, Lei Zhang 0001, Xiaokang Yang 0001, Shipeng Li 0001
ICME7
2009 Summarizing tagged image collections by cross-media representativeness voting
abstract
In this paper, we address the problem of generating both visual and textual summaries for tagged image collections simultaneously. The visual and textual summaries consist of representative images and tags of the collection, which are selected through a proposed cross-media voting scheme. In the voting scheme, the likelihood of an image to be a representative is voted by not only other images but also the tags, according to the intra-media and cross-media affinities. The likelihood of a tag to be a representative is obtained in similar manner at the same time. We demonstrate that the proposed scheme produces more informative textual and visual summaries than summarizing images and tags separately.
Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ICME4
2009 Tag refinement by regularized LDA
abstract
Tagging is nowadays the most prevalent and practical way to make images searchable. However, in reality many tags are irrelevant to image content. To refine the tags, previous solutions usually mine tag relevance relying on the tag similarity estimated right from the corpus to be refined. The calculation of tag similarity is affected by the noisy tags in the corpus, which is not conducive to estimate accurate tag relevance. In this paper, we propose to do tag refinement from the angle of topic modeling. In the proposed scheme, tag similarity and tag relevance are jointly estimated in an iterative manner, so that they can benefit from each other. Specifically, a novel graphical model, regularized Latent Dirichlet Allocation (rLDA), is presented. It facilitates the topic modeling by exploiting both the statistics of tags and visual affinities of images in the corpus. The experiments on tag ranking and image retrieval demonstrate the advantages of the proposed method.
Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia4
2009 Descriptive visual words and visual phrases for image applications
abstract
The Bag-of-visual Words (BoW) image representation has been applied for various problems in the fields of multimedia and computer vision. The basic idea is to represent images as visual documents composed of repeatable and distinctive visual elements, which are comparable to the words in texts. However, massive experiments show that the commonly used visual words are not as expressive as the text words, which is not desirable because it hinders their effectiveness in various applications. In this paper, Descriptive Visual Words (DVWs) and Descriptive Visual Phrases (DVPs) are proposed as the visual correspondences to text words and phrases, where visual phrases refer to the frequently co-occurring visual word pairs. Since images are the carriers of visual objects and scenes, novel descriptive visual element set can be composed by the visual words and their combinations which are effective in representing certain visual objects or scenes. Based on this idea, a general framework is proposed for generating DVWs and DVPs from classic visual words for various applications. In a large-scale image database containing 1506 object and scene categories, the visual words and visual word pairs descriptive to certain scenes or objects are identified as the DVWs and DVPs. Experiments show that the DVWs and DVPs are compact and descriptive, thus are more comparable with the text words than the classic visual words. We apply the identified DVWs and DVPs in several applications including image retrieval, image re-ranking, and object recognition. The DVW and DVP combination outperforms the classic visual words by 19.5% and 80% in image retrieval and object recognition tasks, respectively. The DVW and DVP based image re-ranking algorithm: DWPRank outperforms the state-of-the-art VisualRank by 12.4% in accuracy and about 11 times faster in efficiency.
Shiliang Zhang, Qi Tian 0001, Gang Hua 0001, Qingming Huang, Shipeng Li 0001
ACM Multimedia5
2009 VideoSense: A Contextual In-Video Advertising System
abstract
With Internet delivery of video content surging to an unprecedented level, video has become one of the primary sources for online advertising. In this paper, we presentVideoSenseas a novel contextual in-video advertising system, which automatically associates the relevant video ads and seamlessly inserts the ads at the appropriate positions within each individual video. Unlike most video sites which treat video advertising as general text advertising by displaying video ads at the beginning or the end of a video or around a video, VideoSense aims to embed more contextually relevant ads at less intrusive positions within the video stream. Specifically, given a Web page containing an online video, VideoSense is able to extract the surrounding text related to this video, detect a set of candidate ad insertion positions based on video contentdiscontinuityandattractiveness, select a list of relevant candidate ads according tomultimodal relevance. To support contextual advertising, we formulate this task as a nonlinear 0-1 integer programming problem by maximizing contextual relevance while minimizing content intrusiveness at the same time. The experiments proved the effectiveness of VideoSense for online video service.
Tao Mei 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2009 Complexity-Constrained H.264 Video Encoding
abstract
In this paper, a joint complexity-distortion optimization approach is proposed for real-time H.264 video encoding under the power-constrained environment. The power consumption is first translated to the encoding computation costs measured by the number of scaled computation units consumed by basic operations. The solved problem is then specified to be the allocation and utilization of the computational resources. A computation allocation model (CAM) with virtual computation buffers is proposed to optimally allocate the computational resources to each video frame. In particular, the proposed CAM and the traditional hypothetical reference decoder model have the same temporal phase in operations. Further, to fully utilize the allocated computational resources, complexity-configurable motion estimation (CAME) and complexity-configurable mode decision (CAMD) algorithms are proposed for H.264 video encoding. In particular, the CAME is performed to select the path of motion search at the frame level, and the CAMD is performed to select the order of mode search at the macroblock level. Based on the hierarchical adjusting approach, the adaptive allocation of computational resources and the fine scalability of complexity control can be achieved.
Li Su 0003, Yan Lu 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2008 Coherent image annotation by learning semantic distance
abstract
Conventional approaches to automatic image annotation usually suffer from two problems: (1) They cannot guarantee a good semantic coherence of the annotated words for each image, as they treat each word independently without considering the inherent semantic coherence among the words; (2) They heavily rely on visual similarity for judging semantic similarity. To address the above issues, we propose a novel approach to image annotation which simultaneously learns a semantic distance by capturing the prior annotation knowledge and propagates the annotation of an image as a whole entity. Specifically, a semantic distance function (SDF) is learned for each semantic cluster to measure the semantic similarity based on relative comparison relations of prior annotations. To annotate a new image, the training images in each cluster are ranked according to their SDF values with respect to this image and their corresponding annotations are then propagated to this image as a whole entity to ensure semantic coherence. We evaluate the innovative SDF-based approach on Corel images compared with Support Vector Machine-based approach. The experiments show that SDF-based approach outperforms in terms of semantic coherence, especially when each training image is associated with multiple words.
Tao Mei 0001, Xian-Sheng Hua 0001, Shaogang Gong, Shipeng Li 0001
CVPR5
2008 Learning to video search rerank via pseudo preference feedback
abstract
Conventional approaches to video search reranking only care whether search results are relevant or irrelevant to the given query, while the ranking order of these results indicating the level of relevance or typicality are usually neglected. This paper presents a novel learning-based approach to video search reranking by investigating the ranking order information. The proposed approach, called pseudo preference feedback (PPF), automatically discovers an optimal set of pseudo preference pairs from the initial ranked list and learns a reranking model by ranking support vector machines (ranking SVM) based on the selected pairs. We have proved that PPF can be used for any reranking purpose such as video search and concept detection. We conducted comprehensive experiments for both automatic search and concept detection tasks over TRECVID 2006-2007 benchmark, and showed that PPF could gain significant improvements over the baselines.
Yuan Liu 0017, Tao Mei 0001, Xian-Sheng Hua 0001, Jinhui Tang 0001, Xiuqing Wu, Shipeng Li 0001
ICME6
2008 Continuous Network Coding in Wireless Relay Networks
abstract
Network coding has recently been applied to wireless networks and has achieved some initial success. Researches in wireless network coding have been mostly focusing on utilizing the broadcast nature of the wireless networks. In this paper, we propose a novel network coding framework for wireless relay networks that also takes into consideration the fading and error prone nature of the wireless networks. First, we extend the traditional network coding in lossless networks which operates on 0-1 bits, to a new framework which defines network coding on the posterior probability of each bit. This new framework allows an imperfect decode-recode process at a relay node and avoids possible error propagation when a hard decision is made at the relay node. It implicitly integrates decode-and-forward and estimate-and-forward strategies for wireless network coding to address the technical issues of channel fading and transmission errors. The proposed approach is validated through both theoretical analysis and extensive simulations. Both analysis and simulation confirm that this new framework is able to achieve significant gain over traditional network coding. This new framework also enables the introduction of adaptive scheme into network coding. We demonstrate a basic adaptation scheme and present some preliminary experimental results. The proposed adaptive scheme will lay down an essential foundation in this emerging field of wireless network coding in order to address issues related to link heterogeneity.
Chong Luo 0001, Shipeng Li 0001, Chang Wen Chen
INFOCOM3
2008 ImageSense
abstract
This demonstration presents an innovative contextual advertising platform for online image service, called ImageSense. Unlike most current ad-networks which treat image advertising as general text advertising by displaying relevant ads based on the contents of the Web page, ImageSense aims to embed more contextually relevant ads at less intrusive positions within each suitable image. Given a Web page containing images, ImageSense is able to decompose the page into a set of semantic blocks, select the suitable images from these blocks for advertising, rank the ads according to the relevance derived from surrounding text and visual similarity, and insert the relevant ads into the nonintrusive areas within the selected images. ImageSense represents one of the first attempts towards contextual image advertising which enables both the publishers and advertisers deliver more effective ads carried through image contents.
Lusong Li, Tao Mei 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia4
2008 Contextual in-image advertising
abstract
The community-contributed media contents over the Internet have become one of the primary sources for online advertising. However, conventional ad-networks such as Google AdSense treat image and video advertising as general text advertising without considering the inherent characteristics of visual contents. In this work, we propose an innovative contextual advertising system driven by images, which automatically associates relevant ads with an image rather than the entire text in a Web page and seamlessly inserts the ads in the nonintrusive areas within each individual image. The proposed system, called ImageSense, represents the first attempt towards contextual in-image advertising. The relevant ads are selected based on not only textual relevance but also visual similarity so that the ads yield contextual relevance to both the text in the Web page and the image content. The ad insertion positions are detected based on image saliency to minimize intrusiveness to the user. We evaluate ImageSense on three photo-sharing sites with around one million images and 100 Web pages collected from several major sites, and demonstrate the effectiveness of ImageSense.
Tao Mei 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia3
2008 Flickr distance
abstract
This paper presents Flickr distance, which is a novel measurement of the relationship between semantic concepts (objects, scenes) in visual domain. For each concept, a collection of images are obtained from Flickr, based on which the improved latent topic based visual language model is built to capture the visual characteristic of this concept. Then Flickr distance between different concepts is measured by the square root of Jensen-Shannon (JS) divergence between the corresponding visual language models. Comparing with WordNet, Flickr distance is able to handle far more concepts existing on the Web, and it can scale up with the increase of concept vocabularies. Comparing with Google distance, which is generated in textual domain, Flickr distance is more precise for visual domain concepts, as it captures the visual relationship between the concepts instead of their co-occurrence in text search results. Besides, unlike Google distance, Flickr distance satisfies triangular inequality, which makes it a more reasonable distance metric. Both subjective user study and objective evaluation show that Flickr distance is more coherent to human perception than Google distance. We also design several application scenarios, such as concept clustering and image annotation, to demonstrate the effectiveness of this proposed distance in image related applications.
Lei Wu 0017, Xian-Sheng Hua 0001, Nenghai Yu, Wei-Ying Ma, Shipeng Li 0001
ACM Multimedia5
2008 A comprehensive human computation framework: with application to image labeling
abstract
Image and video labeling is important for computers to understand images and videos and for image and video search. Manual labeling is tedious and costly. Automatically image and video labeling is yet a dream. In this paper, we adopt a Web 2.0 approach to labeling images and videos efficiently: Internet users around the world are mobilized to apply their "common sense" to solve problems that are hard for today's computers, such as labeling images and videos. We first propose a general human computation framework that binds problem providers, Web sites, and Internet users together to solve large-scale common sense problems efficiently and economically. The framework addresses the technical challenges such as preventing a malicious party from attacking others, removing answers from bots, and distilling human answers to produce high-quality solutions to the problems. The framework is then applied to labeling images. Three incremental refinement stages are applied. The first stage collects candidate labels of objects in an image. The second stage refines the candidate labels using multiple choices. Synonymic labels are also correlated in this stage. To prevent bots and lazy humans from selecting all the choices, trap labels are generated automatically and intermixed with the candidate labels. Semantic distance is used to ensure that the selected trap labels would be different enough from the candidate labels so that no human users would mistakenly select the trap labels. The last stage is to ask users to locate an object given a label from a segmented image. The experimental results are also reported in this paper. They indicate that our proposed schemes can successfully remove spurious answers from bots and distill human answers to produce high-quality image labels.
Yang Yang 0059, Bin B. Zhu, Linjun Yang, Shipeng Li 0001, Nenghai Yu
ACM Multimedia5
2008 When multimedia advertising meets the new Internet era
abstract
The advent of media-sharing sites, especially along with the so called Web 2.0 wave, has led to the unprecedented Internet delivery of community-contributed media contents such as images and videos, which have become the primary sources for online advertising. However, conventional ad-networks such as Google Adwords and AdSense treat image and video advertising as general text advertising by displaying the ads either relevant to the queries or the Web page content, without considering automatically monetizing the rich contents of individual images and videos. In this paper, we summarize the trends of online advertising and propose an innovative advertising model driven by the compelling contents of images and videos. We present recently developed ImageSense and VideoSense as two exemplary applications dedicated to images and videos, respectively, in which the most contextually relevant ads are embedded at the most appropriate positions within the images or videos. The ads are selected based on not only textual relevance but also visual similarity so that the ads yield contextual relevance to both the text in the Web page and the visual content. The ad insertion positions are detected based on visual saliency analysis to minimize the intrusiveness to the user. We also envision that the next trend of multimedia advertising would be game-alike advertising.
Xian-Sheng Hua 0001, Tao Mei 0001, Shipeng Li 0001
MMSP3
2007 Temporally Consistent Gaussian Random Field for Video Semantic Analysis
abstract
As a major family of semi-supervised learning, graph based semi-supervised learning methods have attracted lots of interests in the machine learning community as well as many application areas recently. However, for the application of video semantic annotation, these methods only consider the relations among samples in the feature space and neglect an intrinsic property of video data: the temporally adjacent video segments (e.g., shots) usually have similar semantic concept. In this paper, we adapt this temporal consistency property of video data into graph based semi-supervised learning and propose a novel method named temporally consistent Gaussian random field (TCGRF) to improve the annotation results. Experiments conducted on the TREC VID data set have demonstrated its effectiveness.
Jinhui Tang 0001, Xian-Sheng Hua 0001, Tao Mei 0001, Guo-Jun Qi, Shipeng Li 0001, Xiuqing Wu
ICIP (4)5
2007 Image Coding with Parameter-Assistant Inpainting
abstract
This paper carves out an image compression approach that integrates our parameter-assistant inpainting (PAI) technique to exploit the visual redundancy inherent in color-gradation image regions. In our scheme, an input image is first classified at block level according to the degree of edge content as well as chromatic variation in each block. An exemplar selection approach is then adopted to skip a majority of the gradation blocks during encoding. Only their positions and certain parameters extracted for condensed description are encoded along with the reserved blocks. At the decoder side, the skipped regions are recovered through image inpainting, relying on both the delivered parameters and reserved regions. Experimental results show that our proposed method outperforms baseline JPEG at color-gradation regions by nearly 80% bits-saving, at similar visual quality levels.
Zhiwei Xiong, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001
ICIP (2)4
2007 EMS: Energy Minimization Based Video Scene Segmentation
abstract
This paper proposes a novel energy minimization based approach to video scene segmentation. In video content analysis, scene is defined as a set of adjacent shots related to a particular setting or a continuous action in one place. This indicates that not only the global distribution of time and content, but also the local temporal continuity should be taken into account for scene segmentation. Motivated from this fact, we formulate the segmentation procedure as a unified energy minimization framework, in which the global and local constraint is represented by content and context energy, respectively. This energy minimization problem is optimized by two steps in an iterative fashion: first find an initial estimation of scene label for content energy by a generative model; and then iterated conditional modes (ICM) is used for context energy to find the global optimization. Furthermore, a boundary voting procedure is devised to decide the optimal scene boundaries. We apply EMS on an extensive set of home videos and feature movies, and report superior performance compared with several existing key approaches to scene segmentation.
Zhiwei Gu, Tao Mei 0001, Xian-Sheng Hua 0001, Xiuqing Wu, Shipeng Li 0001
ICME5
2007 Distributed Video Coding with Spatial Correlation Exploited Only at the Decoder
abstract
A new pixel-domain distributed video coding (DVC) scheme is proposed in this paper, in which both the temporal and the spatial correlations are exploited only at the decoder. A video is treated as a collection of data correlated in temporal and spatial directions. Besides splitting a video into frames at different time instants, a frame is further split by spatially sub-sampling. Each yielded part is then encoded individually. At the decoder, the side information signals are from both adjacent frames and the spatially decoupled signals. To utilize these multiple side information signals, a new probability model is proposed, in which the transitional probability is calculated from the conditional probabilities on the multiple side information signals. The coding efficiency is enhanced by further removing the spatial redundancy, while the encoding complexity remains the same as the previous pixel-domain DVC techniques that only consider the temporal correlation
Mei Guo, Yan Lu 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001
ISCAS4
2007 Macroblock-Based Adaptive In-Scale Prediction for Scalable Video Coding
abstract
This paper extends the so called generalized in-scale coding technique developed in our previous work to a macroblock-based adaptive in-scale motion compensation framework for H.264/MPEG-4 SVC (scalable video coding). In this framework, the in-scale technique is employed to generate predictions for video frames at high resolution layers, which can utilize both temporal correlation and cross-resolution correlation simultaneously. The lowpass content of a high resolution frame is predicted from its lower resolution layer and the highpass content is predicted from neighboring frames at the same resolution layer. For quality and resolution combined scalability, the prediction from lower resolution layer is not always better than temporal prediction. The in-scale technique is thereby employed adaptively at macroblock level according to the local characteristics of video signal. Techniques for motion estimation and mode decision are also investigated. Experimental results demonstrate that the scheme proposed in this paper can outperform H.264/MPEG-4 SVC significantly. The coding performance is quite promising especially for high-fidelity video coding
Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001
ISCAS4
2007 Video search re-ranking via multi-graph propagation
abstract
This paper1 is concerned with the problem of multimodal fusion in video search. First, we employ an object-sensitive approach to query analysis to improve the baseline result of text-based video search. Then, we propose a PageRank-like graph-based approach to text-based search result re-ranking. To better exploit the underlying relationship between video shots, the proposed re-ranking scheme simultaneously leverages textual relevancy, semantic concept relevancy, and low-level-feature-based visual similarity. In this PageRank-like scheme, we construct a set of graphs with the video shots as vertexes, and the conceptual and visual similarity between video shots as "hyperlinks". A modified topic-sensitive PageRank algorithm is then applied on these graphs to propagate the relevance scores through all related video shots. Experimental results verify the effectiveness of the graph-based propagation approach combined with the object-sensitive query analysis approach, which brings significant improvement to the baseline of text-based video search. Our experimental analysis also indicates that the proposed re-ranking method is highly generic and independent of different query classes, training data, and human interference.
Xian-Sheng Hua 0001, Yalou Huang, Shipeng Li 0001
ACM Multimedia5
2007 VideoSense: towards effective online video advertising
abstract
With Internet delivery of video content surging to an unprecedented level, online video advertising is becoming increasingly pervasive. In this paper, we present a novel advertising system for online video service called VideoSense, which automatically associates the most relevant video ads with online videos and seamlessly inserts the ads at the most appropriate positions within each individual video. Unlike most current video-oriented sites that only display a video ad at the beginning or the end of a video, VideoSense aims to embed more contextually relevant ads at less intrusive positions within the video stream. Given an online video, VideoSense is able to detect a set of candidate ad insertion points based on content discontinuity and attractiveness, select a list of relevant candidate ads ranked according to global textual relevance, and compute local visual-aural relevance between each pair of insertion points and ads. To support contextually relevant and less intrusive advertising, the ads are expected to be inserted at the positions with highest discontinuity and lowest attractiveness, while the overall global and local relevance is maximized. We formulate this task as a nonlinear 0-1 integer programming problem and embed these rules as constraints. The experiments have proved the effectiveness of VideoSense for online video advertising.
Tao Mei 0001, Xian-Sheng Hua 0001, Linjun Yang, Shipeng Li 0001
ACM Multimedia4
2007 VideoSense: a contextual video advertising system
abstract
This demonstration presents a novel contextual advertising platform for online video service, called VideoSense. Unlike most current video-oriented sites that only display a video ad at the beginning or the end of a video, VideoSense aims to embed more contextually relevant ads at less intrusive positions within the video stream. Given an online video, VideoSense is able to detect a set of candidate ad insertion points based on content analysis, select a list of relevant candidate ads ranked according to textual relevance, and find the best match between insertion points and ads which maximizes the overall multimodal relevance. The effectiveness of VideoSense supporting contextually relevant and less intrusive advertising is validated by the user studies conducted on a variety of online video documents.
Tao Mei 0001, Linjun Yang, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia5
2007 Object-Sensitive Query Analysis for Video Search
abstract
This paper is concerned with the problem of improving the performance of text search baseline in video retrieval, specifically for the search tasks in TRECVID. Given a query in plain text, we first implement syntactic segmentation and semantic expansion of the query, then identify the underlying "targeted objects" which should appear in the retrieved video shots, and scale up the weights of the video shots retrieved by the query terms that represent these targeted objects. We name the approaches as "object-sensitive query analysis" for video search. Specifically, we propose a set of methods to identify the specific terms representing the "targeted objects" in a video search query, and a modified object-centric BM25 algorithm to emphasize the impact of these specific object-terms. In practice, we place the process of object-sensitive query analysis before the text search stage, and verify the effectiveness of the proposed approaches with the TRECVID 2005 and 2006 datasets. The experimental results indicate that the proposed object-sensitive approaches to query analysis bring significant improvement upon the raw text search baseline of video search.
Xian-Sheng Hua 0001, Shipeng Li 0001
MMSP3
2007 VideoReach: an online video recommendation system
abstract
This paper presents a novel online video recommendation system called VideoReach, which alleviates users' efforts on finding the most relevant videos according to current viewings without a sufficient collection of user profiles as required in traditional recommenders. In this system, video recommendation is formulated as finding a list of relevant videos in terms of multimodal relevance (i.e. textual, visual, and aural relevance) and user click-through. Since different videos have different intra-weights of relevance within an individual modality and inter-weights among different modalities, we adopt relevance feedback to automatically find optimal weights by user click-though, as well as an attention fusion function to fuse multimodal relevance. We use 20 clips as the representative test videos, which are searched by top 10 queries from more than 13k online videos, and report superior performance compared with an existing video site.
Tao Mei 0001, Bo Yang 0008, Xian-Sheng Hua 0001, Linjun Yang, Shiqiang Yang, Shipeng Li 0001
SIGIR6
2007 Rate-distortion optimized color quantization for compound image compression
abstract
In this paper, we present a new image compression scheme, which is specially designed for computer generated compound color images. First we classify the image content into two kinds: text/graphic content and picture content. Then two different compression schemes are applied blocks of different types. We propose a two stage segmentation scheme which combines thresholding block features and rate-distortion optimization. The text/graphics blocks compression scheme consists of two parts: color quantization and lossless coding of quantized images. The input images will first be color quantized and converted to codebooks and labels, introducing constraint distortion to the color quantization images. Then generated labels and codebooks are lossless compressed respectively. We proposed a rate-distortion optimized color quantization algorithm for text/graphic content, which introduces distortion to text content and minimizes the bit rate produced by the following lossless entropy compression algorithm. The picture content is compressed using conventional image algorithms like JPEG. The results show that the proposed scheme achieves better coding performance than other images compression algorithms such as JPEG2000 and DjVu.
Wenpeng Ding, Yan Lu 0001, Feng Wu 0001, Shipeng Li 0001
VCIP4
2007 Real-time video coding under power constraint based on H.264 codec
abstract
In this paper, we propose a joint power-distortion optimization scheme for real-time H.264 video encoding under the power constraint. Firstly, the power constraint is translated to the complexity constraint based on DVS technology. Secondly, a computation allocation model (CAM) with virtual buffers is proposed to facilitate the optimal allocation of constrained computational resource for each frame. Thirdly, the complexity adjustable encoder based on optimal motion estimation and mode decision is proposed to meet the allocated resource. The proposed scheme takes the advantage of some new features of H.264/AVC video coding tools such as early termination strategy in fast ME. Moreover, it can avoid suffering from the high overhead of the parametric power control algorithms and achieve fine complexity scalability in a wide range with stable rate-distortion performance. The proposed scheme also shows the potential of a further reduction of computation and power consumption in the decoding without any change on the existing decoders.
Li Su 0003, Yan Lu 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001
VCIP4
2007 Generalized in-scale motion compensation framework for spatial scalable video coding
abstract
In existing video coding schemes with spatial scalability based on pyramid frame representation, such as the ongoing H.264/MPEG-4 SVC (scalable video coding) standard, video frame at a high resolution is mainly predicted either from the lower-resolution image of the same frame or from the temporal neighboring frames at the same resolution. Most of these prediction techniques fail to exploit the two correlations simultaneously and efficiently. This paper extends the in-scale prediction technique developed for wavelet video coding to a generalized in-scale motion compensation framework for H.264/MPEG-4 SVC. In this framework, for a video frame at a high resolution layer, the lowpass content is predicted from the information already coded in lower resolution layer, but the highpass content is predicted by exploiting the neighboring frames at current resolution. In this way, both the cross-resolution correlation and temporal correlation are exploited simultaneously, which leads to much higher efficiency in prediction. Preliminary experimental results demonstrate that the proposed framework improves the spatial scalability performance of current H.264/MPEG-4 SVC. The improvement is significant especially for high-fidelity video coding. In addition, another advantage over wavelet-based in-scale scheme is achieved that the proposed framework can support arbitrary down-sampling and up-sampling filters.
Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001
VCIP4
2007 Efficient and Syntax-Compliant JPEG 2000 Encryption Preserving Original Fine Granularity of Scalability
Yang Yang 0059, Bin B. Zhu, Shipeng Li 0001, Neng H. Yu
EURASIP J. Inf. Secur.3
2007 Optimized streaming media proxy and its applications
Guobin Shen, Zhiguang Wang, Shipeng Li 0001
J. Netw. Comput. Appl.4
2007 Image Compression With Edge-Based Inpainting
abstract
In this paper, image compression utilizing visual redundancy is investigated. Inspired by recent advancements in image inpainting techniques, we propose an image compression framework towards visual quality rather than pixel-wise fidelity. In this framework, an original image is analyzed at the encoder side so that portions of the image are intentionally and automatically skipped. Instead, some information is extracted from these skipped regions and delivered to the decoder as assistant information in the compressed fashion. The delivered assistant information plays a key role in the proposed framework because it guides image inpainting to accurately restore these regions at the decoder side. Moreover, to fully take advantage of the assistant information, a compression-oriented edge-based inpainting algorithm is proposed for image restoration, integrating pixel-wise structure propagation and patch-wise texture synthesis. We also construct a practical system to verify the effectiveness of the compression approach in which edge map serves as assistant information and the edge extraction and region removal approaches are developed accordingly. Evaluations have been made in comparison with baseline JPEG and standard MPEG-4 AVC/H.264 intra-picture coding. Experimental results show that our system achieves up to 44% and 33% bits-savings, respectively, at similar visual quality levels. Our proposed framework is a promising exploration towards future image and video compression.
Dong Liu 0002, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.4
2007 Home Video Visual Quality Assessment With Spatiotemporal Factors
abstract
Compared with the video programs taken by professionals, home videos are always with low quality content resulted from non-professional capture skills. In this paper, we present a novel spatiotemporal quality assessment scheme in terms of low-level content features for home videos. In contrast to existing frame-level-based quality assessment approaches, a type of temporal segment of video, subshot, is selected as the basic unit for quality assessment. A set of spatiotemporal visual artifacts, regarded as the key factors affecting the overall perceived quality (i.e., unstableness and jerkiness as temporal factors; infidelity, blurring, brightness, and orientation as spatial factors), are mined from each subshot based on particular characteristics of home videos. The relationship between the overall quality metric and these factors are exploited by three different methods, including user study-based, rule-based and learning-based. To validate the proposed scheme, we present a scalable quality-based home video summarization system from a novel perspective-achieving the best visual quality while simultaneously preserving the most informative content. A comparison user study between this system and the attention model-based video skimming approach demonstrated the effectiveness of the proposed quality assessment scheme
Tao Mei 0001, Xian-Sheng Hua 0001, Cai-Zhi Zhu, He-Qin Zhou, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2007 Barbell-Lifting Based 3-D Wavelet Coding Scheme
abstract
This paper provides an overview of the Barbell lifting coding scheme that has been adopted as common software by the MPEG ad hoc group on further exploration of wavelet video coding. The core techniques used in this scheme, such as Barbell lifting, layered motion coding, 3D entropy coding and base layer embedding, are discussed. The paper also analyzes and compares the proposed scheme with the oncoming scalable video coding (SVC) standard because the hierarchical temporal prediction technique used in SVC has a close relationship with motion compensated temporal lifting (MCTF) in wavelet coding. The commonalities and differences between these two schemes are exhibited for readers to better understand modern scalable video coding technologies. Several challenges that still exist in scalable video coding, e.g., performance of spatial scalable coding and accurate MC lifting, are also discussed. Two new techniques are presented in this paper although they are not yet integrated into the common software. Finally, experimental results demonstrate the performance of the Barbell-lifting coding scheme and compare it with SVC and another well-known 3D wavelet coding scheme, MC embedded zero block coding (MC-EZBC).
Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2007 Subband Coupling Aware Rate Allocation for Spatial Scalability in 3-D Wavelet Video Coding
abstract
The motion compensated temporal filtering (MCTF) technique, which is extensively used in 3-D wavelet video coding schemes nowadays, leads to signal coupling among various spatial subbands because motion alignment is introduced in the temporal filtering. Using all spatial subbands as a reference enables MCTF to fully take advantage of temporal correlation across frames but inevitably brings drifting problem in supporting spatial scalability. This paper first analyzes the signal coupling phenomenon and then proposes a quantitative model to describe signal propagation across spatial subbands during the MCTF process. The signal propagation is modeled for a single MC step based on the shifting effect of wavelet synthesis filters and then it is extended to multilevel MCTF. This model is called subband coupling aware signal propagation (SCASP) model in this paper. Based on the model, we further propose a subband coupling aware rate allocation scheme as one possible solution to the above dilemma in supporting spatial scalability. To find the optimal rate allocation among all subbands for a specified reconstruction resolution, the SCASP model is used to approximate the reconstruction process and derive the synthesis gain of each subband with regard to that reconstruction. Experimental results have fully demonstrated the advantages of our proposed rate allocation scheme in improving both objective and subjective qualities of reconstructed low-resolution video, especially at middle bit rates and high bit rates.
Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.4
2007 Adaptive Directional Lifting-Based Wavelet Transform for Image Coding
abstract
We present a novel 2-D wavelet transform scheme of adaptive directional lifting (ADL) in image coding. Instead of alternately applying horizontal and vertical lifting, as in present practice, ADL performs lifting-based prediction in local windows in the direction of high pixel correlation. Hence, it adapts far better to the image orientation features in local windows. The ADL transform is achieved by existing 1-D wavelets and is seamlessly integrated into the global wavelet transform. The predicting and updating signals of ADL can be derived even at the fractional pixel precision level to achieve high directional resolution, while still maintaining perfect reconstruction. To enhance the ADL performance, a rate-distortion optimized directional segmentation scheme is also proposed to form and code a hierarchical image partition adapting to local features. Experimental results show that the proposed ADL-based image coding technique outperforms JPEG 2000 in both PSNR and visual quality, with the improvement up to 2.0 dB on images with rich orientation features.
Wenpeng Ding, Feng Wu 0001, Xiaolin Wu 0001, Shipeng Li 0001, Houqiang Li
IEEE Trans. Image Process.4
2007 Modeling and Mining of Users' Capture Intention for Home Videos
abstract
With the rapid adoption of consumer digital video recorders and an increase of home video data, content analysis has become an interesting and key research issue to provide personalized experiences and services for both camcorder users and viewers. In this paper, we present a novel view to tackle this issue, which aims at modeling and mining of the capture intention of camcorder users. Based on the study of intention mechanism in psychology, a set of domain-specific capture intention concepts is defined. A comprehensive and extensible scheme consisting of video structure decomposition, intention-oriented feature analysis, as well as singular-value-decomposition-based intention segmentation and learning-based intention classification is proposed to mine the users' capture intention. Experiments were carried on home video sequences of 90 h in total, taken by 16 persons over the past 20 years. Both the user study and objective evaluations indicate that our proposed intention-based approach is an effective complement to existing home video content analysis schemes
Tao Mei 0001, Xian-Sheng Hua 0001, He-Qin Zhou, Shipeng Li 0001
IEEE Trans. Multim.4
2006 An efficient key scheme for multiple access of JPEG 2000 and motion JPEG 2000 enabling truncations
abstract
JPEG 2000 provides multiple scalable accesses to a single codestream. Digital Rights Management of a JPEG 2000 codestream should preserve the original flexibility of scalability yet provide a mechanism to ensure what you see is what you pay: a low resolution version displayed on a smart phone should pay less than a high resolution version displayed on a PC. We present an efficient key scheme for multi-type, multilevel scalable access control for JPEG 2000 and motion JPEG 2000 codestreams. The scheme is based on a poset representation of the scalable access control and a hash based hierarchical access key scheme, both proposed elsewhere. The proposed key scheme exploits the information contained in a codestream and the features invariant under truncations to minimize the file size overhead for DRM applications yet preserve correct derivation of keys for descendants even when an encrypted codestream is truncated.
Bin B. Zhu, Yang Yang 0059, Shipeng Li 0001
CCNC3
2006 Wyner-Ziv Video Coding Based on Set Partitioning in Hierarchical Tree
abstract
In this paper, we propose a Wyner-Ziv video coding scheme based on set-partitioning in hierarchical trees (SPIHT) which can utilizing not only the spatial and temporal correlations but also the high-order statistical correlations. Wyner-Ziv theory on source coding with side information is employed as the basic coding principle, which makes the independent encoding and joint decoding become possible. In the proposed scheme, wavelet transform is first used to de-correlate the spatial dependency of a Wyner-Ziv frame. Then the quantized transform coefficients are organized by using magnitude with a set partitioning sorting algorithm. The ordered bit planes are coded using the Wyner-Ziv coding based on turbo codes. At the decoder, side information generated by motion compensated interpolation is used to conditionally decode the Wyner-Ziv frame. Since the high order statistical correlation is used, the proposed algorithm owns advantages over the traditional pixel-domain and transform-domain Wyner-ziv video coding schemes.
Xun Guo 0002, Yan Lu 0001, Feng Wu 0001, Wen Gao 0001, Shipeng Li 0001
ICIP5
2006 Transcoding to FGS Streams from H.264/AVC Hierarchical B-Pictures
abstract
This paper presents a transcoder which transcodes to FGS streams from H.264/AVC hierarchical B-pictures. First, the DCT-domain architecture is designed for fast FGS transcoding. Then, we propose a mode decision method in DCT domain to achieve a trade-off between the performances at low bit-rate and high bit-rate. Experimental results demonstrated that our method can improve the coding performance up to 1 dB at low rate and only lose at worst 0.5 dB at high rate.
Huifeng Shen, Xiaoyan Sun 0001, Feng Wu 0001, Houqiang Li, Shipeng Li 0001
ICIP5
2006 Automatic Video Genre Categorization using Hierarchical SVM
abstract
This paper presents an automatic video genre categorization scheme based on the hierarchical ontology on video genres. Ten computable spatio-temporal features are extracted to distinguish the different genres using a hierarchical support vector machines (SVM) classifier built by cross-validation, which consists of a series of SVM classifiers united in a binary-tree form. As the order and genre partition strategy of the SVM classifier series affect the over performance of the united classifier, two optimal SVM binary trees, local and global, are constructed aiming at finding the best categorization orders, i.e., the best tree structure, of the genre ontology. Experimental results show that the proposed scheme outperforms C4.5 decision tree, typical 1-vs-1 SVM scheme, as well as the hierarchical SVM built by K-means.
Xun Yuan 0001, Tao Mei 0001, Xian-Sheng Hua 0001, Xiuqing Wu, Shipeng Li 0001
ICIP6
2006 Motion Aligned Spatial Scalable Video Coding
abstract
A motion aligned spatial scalable video coding scheme (MA-SSC) is proposed in this paper. Different from the traditional spatial scalable coding schemes derived from MPEG-2, in the proposed scheme only one set of intra or inter prediction modes are optimally selected by jointly considering the base and enhancement layers. Thus, it saves one set of macroblock (MB) mode and motion vectors. Moreover, the combined motion estimation can reduce the residual coding bits of the base layer. The MA-SSC and traditional spatial scalable coding schemes are both implemented based on H.264 reference software to evaluate their performance. Simulation results show that the enhancement layer coding efficiency of MA-SSC is up to 0.6dB better than that of the traditional scheme, while the base layer coding efficiency of MA-SSC decreases less than 0.3db compared with the single-layer coding
Debing Liu, Yuwen He, Shipeng Li 0001, Debin Zhao, Wen Gao 0001
ICME3
2006 Probabilistic Multimodality Fusion for Event based Home Photo Clustering
abstract
This paper presents a novel probabilistic approach to fusing multimodal metadata for event based home photo clustering. Photo events are characterized by the coherence of multimodality including time, content and camera settings. We incorporate these multimodal metadata into a unified probabilistic framework, in which event is taken as a latent semantic concept and discovered by fitting a generative model through an expectation-maximization (EM) algorithm. This approach is general and unsupervised, without any training procedure or predefined threshold. The experimental evaluations on 14 k photos taken by 10 amateur photographers have indicated the effectiveness and efficiency of the proposed framework in browsing and searching personal photo collections
Tao Mei 0001, Xian-Sheng Hua 0001, He-Qin Zhou, Shipeng Li 0001
ICME5
2006 Complexity Scalable 2 : 1 Resolution Downscaling MPEG-2 to WMV Transcoder with Adaptive Error Compensation
abstract
In this paper, we focus on 2:1 spatial resolution downscaling transcoding from MPEG-2 to WMV. We propose two architectures (for sequences with or without B-frames respectively) that are unique in their complexity scalability and efficient control over the drifting error, which in return provide a flexible mechanism to achieve desired tradeoff between the complexity and the quality. We achieve resolution downscaling completely in the DCT domain and show that the standard IDCT (as in all the MPEG series standards) can be merged with other DCT-like transform (e.g., the integer transform in WMV) with proper one-time per-element scaling. Extensive experimental results verified the effectiveness of proposed structures against several design objectives such as complexity scalability and performance tradeoffs
Guobin Shen, Yuwen He, Wanyong Cao, Shipeng Li 0001
ICME4
2006 A Fast Downsizing Video Transcoder for H.264/AVC with Rate-Distortion Optimal Mode Decision
abstract
This paper focuses on the mode decision and motion selection problem when H.264/AVC video streams are transcoded in spatial resolution. A fast downsizing transcoding scheme is developed in which a new rate-distortion (R-D) optimal mode decision mechanism is presented for high speed transcoding as well as high coding efficiency. A model for estimating relative prediction errors is applied in this paper, which is free from computation of interpolation and SAD/SSD computation. Based on the selected model, a motion refinement within a distance of 1 pixel is performed after mode decision. Experimental results demonstrate that our method can significantly speed up the spatial resolution reduction process, while achieving high coding efficiency
Huifeng Shen, Xiaoyan Sun 0001, Feng Wu 0001, Houqiang Li, Shipeng Li 0001
ICME5
2006 Off-Line Motion Description for Fast Video Stream Generation in MPEG-4 AVC/H.264
abstract
The rate-distortion optimal mode decision as well as motion estimation adopted in H.264 brings a big challenge to real-time encoding and transcoding due to the high computation complexity. In this paper, we propose a hierarchical motion description model to present the motion data of each macroblock (MB) from coarsely to finely. A preprocessing approach is developed to estimate the motion data for each MB at each quality level with regard to its reference quality, its adjacent MBs and the target bit-rate. The resulting motion data can be coded and stored as metadata in a media file or a stream. Moreover, we propose a method to readily extract the specific motion data from the model for each MB at given bit-rates. Experimental results have shown the effectiveness of our proposed motion description model in terms of coding efficiency as well as fast bit-rate adaptation in comparison with that of H.264
Yi Wang 0037, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Houqiang Li, Zhengkai Liu
ICME4
2006 Adaptive MCTF based on Correlation Noise Model for SNR Scalable Video Coding
abstract
This paper proposes a subband adaptive motion compensated temporal filtering (MCTF) technique for scalable video coding and introduces a revised synthesis gain model for the quantization in this adaptive MCTF scheme. In scalable video coding, hierarchical MCTF is extensively adopted to exploit the temporal correlation across video frames. In this hierarchical MCTF structure, the strength of temporal correlation varies with the level of temporal transform and varies with the various spatial frequency components in a frame. The reconstruction noises also have diverse strength at various subbands. According to the correlation and noise characteristics of various subbands, we can adjust the strength of motion compensated prediction step in MCTF to maximally take the advantage of temporal correlation but restrict the propagation of reconstruction noise. The quantization step of each subband is also adjusted according to synthesis gain determined by the MCTF structure. In this way an adaptive MCTF scheme is formed and the proposed technique improves the coding performance of scalable video coding
Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001
ICME4
2006 Complexity scalable MPEG-2 to WMV transcoder with adaptive error compensation
abstract
In this paper, we study the problem of video transcoding from MPEG-2 to WMV format, together with several desired functionalities such as bit rate reduction etc. We propose two architectures (for different typical application scenarios) that are unique in their complexity scalability and adaptive drifting error control, which in return provide a mechanism to achieve desired trade-off between the complexity and the quality. A simple model-based rate control algorithm is also presented. We performed extensive experiments for various design targets such as complexity scalability, performance tradeoff, drifting control effect etc. The proposed transcoding architectures can be straightforwardly applied to the MPEG-2 to MPEG-4 transcoding applications due to the significant overlap between the MPEG-4 and WMV coding technology.
Guobin Shen, Yuwen He, Wanyong Cao, Shipeng Li 0001
ISCAS4
2006 Rate-distortion optimization for fast hierarchical B-picture transcoding
abstract
An efficient rate-distortion (R-D) optimal method for transcoding hierarchical B-pictures is proposed in this paper. A new R-D model is presented for fast transcoding hierarchical B-pictures in DCT domain, in which the total R-D optimization problem is adjusted to motion R-D optimization and texture R-D optimization separately. Accordingly, a mechanism for fast mode decision is proposed to enable the mode and motion adjustment in hierarchical B-picture DCT-domain transcoding. Experimental results show that the proposed transcoding scheme with the new R-D optimization can achieve 4-dB average PNSR improvement at low bit-rate, and similar performance at high bit-rate, compared to the transcoding by reusing the input motion information. Moreover, as the transcoding is performed in DCT domain, it is fast and simple enough for many real applications.
Huifeng Shen, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001
ISCAS4
2006 Automatic video annotation based on co-adaptation and label correction
abstract
As there is a large gap between high-level semantics and low-level features, it is difficult to obtain high-accuracy video semantic annotation through automatic methods. In this paper, we propose a novel automatic video annotation method, which greatly improves the annotation performance by learning from unlabeled video data, as well as exploring temporal consistency of video sequences. To effectively learn from unlabeled data, a scheme called co-adaptation is proposed to progressively refine two pre-trained complementary classifiers, and then a minimum entropy based method is applied to sufficiently explore the video temporal consistency, which further improves the annotation accuracy. Experiments show that the proposed automatic video annotation method performs superior than both general learning-based and co-training-based methods
Meng Wang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Li-Rong Dai 0001, Shipeng Li 0001
ISCAS5
2006 In-scale motion aligned temporal filtering
abstract
To handle the mismatch problems of spatial-domain motion aligned temporal filtering (MATF) in providing spatial scalability, this paper presents a novel in-scale motion aligned temporal filtering (ISMATF) for highly scalable video coding. For each lifting step of ISMATF, the predict or update is carried out at various resolutions in a scale-by-scale way and the operation at higher resolution is designed to embed the operations at lower resolutions to prevent redundancy in frame representation. The proposed ISMATF has several main features, which as a whole differentiates it from other schemes (e.g. spatial-domain MATF, in-band MATF, straightforward multi-resolution MATF) and makes it effective in video coding with spatial scalabilities. These features include 1) Multi-resolution lifting with perfect reconstruction capability; 2) non-redundancy in frame representation, which favors highly scalable video coding; 3) temporal filtering performed in the image-domain of corresponding resolution, which is different from in-band schemes and preserves high efficiency of motion compensation; 4) no mismatch between the encoder and decoder, no matter which resolution is decoded. The proposed ISMATF scheme solves the mismatch problems of SDMATF at low resolution while maintains a good performance at high resolution. It can outperform redundant multi-resolution MATF schemes by up to 1.0 dB at high resolution.
Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001
ISCAS4
2006 Signed MSB-Set Comb Method for Elliptic Curve Point Multiplication
Min Feng 0002, Bin B. Zhu, Cunlai Zhao, Shipeng Li 0001
ISPEC4
2006 Automatic video annotation by semi-supervised learning with kernel density estimation
abstract
Insufficiency of labeled training data is a major obstacle for automatically annotating large-scale video databases with semantic concepts. Existing semi-supervised learning algorithms based on parametric models try to tackle this issue by incorporating the information in a large amount of unlabeled data. However, they are based on a assumption that the assumed generative model is correct, which usually cannot be satisfied in automatic video annotation due to the large variations of video semantic concepts. In this paper, we propose a novel semi-supervised learning algorithm, named Semi Supervised Learning by Kernel Density Estimation (SSLKDE), which is based on a non-parametric method, and therefore the assumption is avoided. While only labeled data are utilized in the classical Kernel Density Estimation (KDE) approach, in SSLKDE both labeled and unlabeled data are leveraged to estimate class conditional probability densities based on an extended form of KDE. We also investigate the connection between SSLKDE and existing graph-based semi-supervised learning algorithms. Experiments prove that SSLKDE significantly outperforms existing supervised methods for video annotation.
Meng Wang 0001, Yan Song 0001, Xun Yuan 0001, HongJiang Zhang, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia6
2006 An extension of direct macroblock coding in Predictive (P) slices of the H.264 standard
Alexis M. Tourapis, Feng Wu 0001, Shipeng Li 0001
J. Vis. Commun. Image Represent.3
2006 MPEG-2 to WMV Transcoder With Adaptive Error Compensation and Dynamic Switches
abstract
In this paper, we study the problem of video transcoding from MPEG-2 to Windows Media Video (WMV) format, together with several desired functionalities such as bit-rate reduction and spatial resolution downscaling. Based on in-depth analysis of error propagation behavior, we propose two architectures (for different typical application scenarios) that are unique in their complexity scalability and adaptive drifting error control, which in return provide a mechanism to achieve a desired tradeoff between complexity and quality. We perform extensive experiments for various design targets such as complexity, scalability, performance tradeoff, and drifting control effect. The proposed transcoding architectures can be straightforwardly applied to the MPEG-2 to MPEG-4 transcoding applications due to the significant overlap between the MPEG-4 and WMV coding technology.
Guobin Shen, Yuwen He, Wanyong Cao, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2006 Drift-free switching of compressed video bitstreams at predictive frames
abstract
Two schemes are proposed to efficiently compress video contents into bitstreams that support drift-free switching at predictive frames. They are inspired by the original SP coding scheme presented in the early H.26L. First, we propose a Flex SP coding scheme in which the prediction signal of the SP frame is directly subtracted from the input without quantization and de-quantization. The decoded video quality of the Flex SP scheme is significantly improved when additional inverse discrete cosine transform (DCT) and post-filter are provided. Then, the Hybrid SP scheme is presented to further improve the quality of the display image, as well as the reconstructed reference, by defining two coding modes for each DCT coefficient. Moreover, a rate-distortion algorithm is proposed to determine the coding mode for each coefficient. The bitstreams generated by the two proposed schemes can be decoded successfully by a decoder that complies with MPEG-4 AVC/H.264. In addition, we also investigate how to choose the quantization parameters for switching. An empirical method is proposed to achieve a good tradeoff between high coding efficiency of SP frames and small size of switching bits.
Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Guobin Shen, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2006 4-D Wavelet-Based Multiview Video Coding
abstract
The conventional multiview video coding (MVC) schemes, utilizing both neighboring temporal frames and view frames as possible references, have only shown a slight gain over those using temporal frames alone in terms of coding efficiency. The reason for this is that the neighboring temporal frames exhibit stronger correlation with the current frame and the view frames often fail to be selected as references. This paper proposes an elegant MVC framework using high dimensional wavelet, which rightly matches the inherent high dimension property of multiview video. It also makes a better usage of both temporal and view correlations thanks to the hierarchical decomposition. Besides the proposed framework, this paper also investigates MVC coding from the following aspects. First, a disparity-compensated view filter (DCVF) with pixel alignment is proposed, which can accommodate both global and local view disparities among view frames. The proposed DCVF and the existing motion-compensated temporal filter (MCTF) unify the view and temporal decompositions as a generic lifting transform. Second, an adaptive decomposition structure based on the analysis of the temporal and view correlations is proposed. A Lagrangian cost function is derived to determine the optimum decomposition structure. Third, the major components of the proposed MVC coding are figured out, including macroblock type design, subband coefficient coding, and rate allocation. Extensive experiments are carried out on the MPEG 3DAV test sequences and the superior performance of the proposed MVC coding is demonstrated. In addition, the proposed MVC framework can easily support temporal, spatial, SNR, as well as view scalabilities
Yan Lu 0001, Feng Wu 0001, Jianfei Cai 0001, King Ngi Ngan, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.6
2005 Prediction-based directional fractional pixel motion estimation for H.264 video coding
abstract
In an H.264 video encoder, motion estimation (ME) is the most time-consuming component. The ME process consists of two stages, integer pixel search and fractional pixel search. Since the complexity of integer pixel search has been greatly reduced by numerous fast ME algorithms, the computation overhead required by fractional pixel ME has become relatively significant. To reduce the complexity of fractional pixel ME, we propose a prediction-based directional fractional pixel ME algorithm. We utilize more accurate motion vector predictions and directional search to achieve better computation reduction. We further propose an early termination method to decrease the amount of search. Experimental results show that, compared to the full search sub-pel ME and the fast sub-pel ME proposed in H.264, the proposed method can reduce up to 84% and 74% of fractional pixel search points respectively, with a negligible degradation in quality.
Libo Yang, Keman Yu, Jiang Li 0008, Shipeng Li 0001
ICASSP (2)4
2005 ThresPassport - A Distributed Single Sign-On Service
Tierui Chen, Bin B. Zhu, Shipeng Li 0001, Xueqi Cheng 0001
ICIC (2)3
2005 Spatio-temporal video error concealment using priority-ranked region-matching
abstract
When transmitted over error-prone networks, compressed video sequences may be received with errors. In this paper, we propose a priority-ranked region-matching algorithm to recover the "lost" area of the decoded frames, in which both temporal and spatial correlations of the video sequence are exploited. In the proposed scheme, we first calculate the priorities of all edge pixels of the "lost" area and generate a priority-ranked region group. Then according to their priorities, the regions in the group will search their best matching regions temporally and spatially. Finally, the "lost" area is recovered progressively by the corresponding pixels in the matching regions. Experimental results show that the proposed scheme achieves higher PSNR as well as better video quality in comparison with the method adopted in H.264.
Yan Chen 0007, Xiaoyan Sun 0001, Feng Wu 0001, Zhengkai Liu, Shipeng Li 0001
ICIP (2)5
2005 H.264-compatible spatially scalable video coding with in-band prediction
abstract
In this paper, a H.264 compatible spatially scalable video coding method with in-band prediction is proposed which taking advantages from both the high coding efficiency of H.264 coding scheme and the attractive performance of in-band overcomplete discrete wavelet transform (ODWT) in wavelet-domain motion estimation and motion compensation. Four MV prediction modes are proposed for INTER prediction of high frequency subbands. The intra prediction modes of H.264 are also simplified for each high band according to the directional features inherited inside. Finally, a H.264 compatible scheme based on one of the MV prediction modes is presented to provide better tradeoff among standard compatibility, low complexity and high performance.
Xin Jin 0014, Xiaoyan Sun 0001, Feng Wu 0001, Guangxi Zhu, Shipeng Li 0001
ICIP (1)5
2005 Efficient video mosaicing based on motion analysis
abstract
Presenting more comprehensive information than key frame and any subset of frames, mosaic has attracted a growing attention in recent years as a useful element for a variety of vision tasks. In this paper, we present an efficient approach to video mosaicing based on motion analysis with investigating two important issues that significantly affect the performance of mosaicing. First, given a video shot, a method based on motion entropy analysis is introduced to determine whether the visual content in a shot or video segment is suited to be represented by a mosaic. The second is to decide which subset of frames in the shot is the best to be selected for efficient and effective mosaicing by constructing a global motion path. Experimental comparisons with typical mosaicing method approve that with this approach, the visual quality of mosaic is significantly improved, as well as the computation time is remarkably reduced.
Tao Mei 0001, Xian-Sheng Hua 0001, He-Qin Zhou, Shipeng Li 0001, HongJiang Zhang
ICIP (1)4
2005 Optimal packetization of fine granularity scalability codestreams for error-prone channels
abstract
An optimal source-channel packetization scheme for MPEG-4 fine granularity scalability (FGS) codestreams is proposed in this paper. The channel is modeled with a uniform error distribution to the enhancement layer transmission. A cost function that models the error expansion for a MPEG-4 FGS stream is derived, and then used in the optimal packetization problem subject to the same overhead as the conventional packetization scheme. An efficient scheme to find the optimal solution is described, which takes time similar to encoding an MPEG-4 FGS codestream. Experiments show that our scheme has up to 1.96 dB gain over the conventional packetization scheme.
Bin B. Zhu, Yang Yang 0059, Chang Wen Chen, Shipeng Li 0001
ICIP (2)4
2005 JPEG 2000 syntax-compliant encryption preserving full scalability
abstract
An efficient syntax-compliant encryption scheme for JPEG 2000 and motion JPEG 2000 is proposed in this paper. Compressed visual data is completely encrypted yet the full scalability of the unencrypted codestream is completely preserved to allow near RD-optimal truncations and other manipulations securely without decryption. Compared with other reported schemes, our scheme shows advantages on syntax compliance, compression overhead, scalable granularity, and error resilience. In addition to preserving the original scalability, a JPEG 2000 codestream encrypted with our scheme has the same error resilience capability as the unencrypted codestream. The encrypted codestream is still syntax-compliant so that an encryption-unaware decoder can still decode the encrypted codestream, although the decoded visual data is completely garbled and meaningless. Our scheme has virtually no adverse impact on the compression efficiency.
Bin B. Zhu, Yang Yang 0059, Shipeng Li 0001
ICIP (3)3
2005 Video booklet
abstract
In this paper, we propose a novel system, video booklet, which enables efficient and nature video browsing and searching. In the system, a set of selected thumbnails excerpted from the original videos are printed out on a real physical booklet or album. When users plan to browse the content of or search a specific clip in their digital video library, they can firstly browse their booklets in such a manner as browsing a typical papery family album. When he wants to watch the segment represented by a certain thumbnail in the booklet, he is able to use his camera phone to capture the corresponding thumbnail. Then the captured image is sent to computer or other devices connected to the monitor via wireless network, and last the video booklet system will automatically find the corresponding segment in the video library for the user and begin to play the segment. Video booklet builds a seamless bridge between digital media library and analog papery albums.
Xian-Sheng Hua 0001, Shipeng Li 0001, HongJiang Zhang
ICME2
2005 Camera notes
abstract
Taking notes is frequently required in daily life. The rapid development of consumer devices provides new ways to achieve this goal such as taking notes by digital cameras, camcorders, camera phones, or voice recorders. In this paper, a novel camera-based note taking system is presented, which enables efficient and nature notes taking, management, searching and exporting. Firstly, the system automatically classifies photos and video clips taken by a variety of capturing devices into two classes, notes and non-notes, and then further classifies the class of notes into more fractional classes such as document, map, slides, bulletin board, and so on. Next the visual quality of the note photos and video clips are enhanced and adjusted by a set of image and video processing algorithms, as well as textual, color and texture information are extracted from them. And last, based on these analyses, a management system enabling efficient note importing, searching, browsing and exporting is built.
Xian-Sheng Hua 0001, Shipeng Li 0001, HongJiang Zhang
ICME2
2005 Online End Detection for Live-Broadcast Sports TV Programs
abstract
In this paper, a method for automatically detecting the end of lively broadcasted sports programs is proposed, which enables users to record the full TV programs when they run over time. Taking advantage of the property of relatively high content consistency within sports programs, this method is based on checking the break point of this content consistency. A scalable video segment similarity measure is proposed to measure content consistency of video segments in different similarity levels. Based on this measure, a two-round end detection scheme is applied, in which the first round, Candidate End Point Finding, is able to find a coarse candidate end point, and the second round, End Point Refining, is able to find a more accurate end. Experiments show that the proposed end detection scheme is able to detect the real ends with high accuracy.
Hao-Da Huang, Xian-Sheng Hua 0001, Shipeng Li 0001, HongJiang Zhang
ICME3
2005 Advanced Motion Search and Adaptation Techniques for Deinterlacing
abstract
Unlike video coding, video deinterlacing relies heavily on the correctness of motion. To obtain more reliable motion, we propose a new motion search criterion that imposes constraints on the motion diversity among neighboring blocks, and improve the symmetric ME method by dynamic block splitting and single direction ME. To further improve the visual quality, adaptive deinterlacing algorithm based on block variances is also proposed. Extensive experimental results demonstrate that the proposed techniques significantly improved the deinterlaced video quality, both PSNR-wise and visually.
Kefei Ouyang, Guobin Shen, Shipeng Li 0001
ICME3
2005 Joint Sender/Receiver Optimization Algorithm for Multi-Path Video Streaming Using High Rate Erasure Resilient Code
abstract
In this paper we present a joint sender/receiver optimization algorithm and a seamless rate adjustment protocol to reduce the total number of packets over different paths in a streaming framework with a variety of constraints such as target throughput, dynamic packet loss ratio, and available bandwidths. We exploit the high rate erasure resilient code for the ease of packet loss adaption and seamless rate adjustment. The proposed algorithm and adjustment protocol can be applied at an arbitrary scale. Simulation results demonstrate that the overall traffic is significantly reduced with the proposed algorithm and protocol.
Changxi Zheng, Guobin Shen, Shipeng Li 0001, Qianni Deng
ICME3
2005 Personal media sharing and authoring on the web
abstract
In this paper, we propose a novel system working on the Web for personal media sharing and authoring. Three primary technologies enable this end-to-end system, including scalable video coding, intelligent multimedia content analysis, and template-based media authoring. Scalable video codec tackles the issue of huge data transmission, multimedia content analysis facilitates automatic video editing, and template-based media authoring scheme further reduces the workload of media sharing and authoring. Experiment and a demo system on a real Internet environment show that this novel system is effective.
Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia2
2005 LazyCut: content-aware template-based video authoring
abstract
Though there are many commercial video authoring tools available today, video authoring remains as a tedious and extremely time consuming task that often requires trained professional skills. To tackle this issue, this demonstration presents a novel end-to-end system, called LazyCut, which enables fast, flexible and personalized video authoring and sharing. LazyCut provides a semi-automatic video authoring and sharing system that significantly reduces users' efforts in video editing while preserving sufficient flexibility and personalization.
Xian-Sheng Hua 0001, Zengzhi Wang, Shipeng Li 0001
ACM Multimedia3
2005 Secure Key Management for Flexible Digital Rights Management of Scalable Codestreams
abstract
The key management for multi-type, multi-level scalable access control of a fine granularity scalability codestream is addressed in this paper. We first present an efficient partially ordered set (poset) to represent scalable access control so that any access control schemes for a poset can be applied and a single secret key is needed to transfer to a user. We then present a secure key scheme modified from our previous scalable access key scheme for the resulting poset. The key scheme is based on the group Diffie-Hellman key agreement, and is secure
Bin B. Zhu, Min Feng 0002, Shipeng Li 0001
MMSP3
2005 Fine Granularity Scalability Encryption of MPEG-4 FGS Bitstreams
abstract
In this paper, we present an encryption scheme for MPEG-4 FGS which provides the same or a little coarser granularity of scalability after encryption. The scheme encrypts compressed data of each video packet or block independently. Initialization vectors are generated with a method to minimize the overhead. The scalability provided in an encrypted codestream using this scheme enables intermediate nodes to truncate an encrypted bitstream at near R-D optimality directly without decryption, which enhances system security. The scheme has virtually negligible overhead, and produces encrypted codestream with virtually the same error resilience performance as the unencrypted case. These features are very desirable in many applications
Bin B. Zhu, Yang Yang 0059, Chang Wen Chen, Shipeng Li 0001
MMSP4
2005 Accelerate Video Decoding With Generic GPU
abstract
Most modern computers or game consoles are equipped with powerful yet cost-effective graphics processing units (GPUs) to accelerate graphics operations. Though the graphics engines in these GPUs are specially designed for graphics operations, can we harness their computing power for more general nongraphics operations? The answer is positive. In this paper, we present our study on leveraging the GPUs graphics engine to accelerate the video decoding. Specifically, a video decoding framework that involves both the central processing unit (CPU) and the GPU is proposed. By moving the whole motion compensation feedback loop of the decoder to the GPU, the CPU and GPU have been made to work in parallel in a pipelining fashion. Several techniques are also proposed to overcome the GPUs constraints or to optimize the GPU computation. Initial experimental results show that significant speed-up can be achieved by utilizing the GPU power. We have achieved real-time playback of high definition video on a PC with an Intel Pentium III 667-MHz CPU and an nVidia GeForce3 GPU.
Guobin Shen, Guang-ping Gao, Shipeng Li 0001, Harry Shum, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.3
2005 Direct mode coding for bipredictive slices in the H.264 standard
abstract
Abstract—The new H.264 (MPEG-4 AVC) video coding standard can achieve considerably higher coding efficiency compared to previous standards. This is accomplished mainly due to the consideration of variable block sizes for motion compensation, multiple reference frames, intra prediction, but also due to better exploitation of the spatiotemporal correlation that may exist between adjacent Macroblocks, with the SKIP mode in predictive ( ) slices and the two DIRECT modes in bipredictive ( ) slices. These modes, when signaled, could in effect represent the motion of a macroblock (MB) or block without having to transmit any additional motion information required by other inter-MB types. This property also allows these modes to be highly compressible especially due to the consideration of run length coding strategies. Although spatial correlation of motion vectors from adjacent MBs is used for SKIP mode to predict its motion parameters, until recently, DIRECT mode considered only temporal correlation of adjacent pictures. In this letter, we introduce alternative methods for the generation of the motion information for the DIRECT mode using spatial or combined spatiotemporal correlation. Considering that temporal correlation requires that the motion and timestamp information from previous pictures are available in both the encoder and decoder, it is shown that our spatial-only method can reduce or eliminate such requirements while, at the same time, achieving similar performance. The combined methods, on the other hand, by jointly exploiting spatial and temporal correlation either at the MB or slice/picture level, can achieve even higher coding efficiency. Finally, improvements on the existing Rate Distortion Optimization related to slices within the H.264 codec are also presented, which can lead to improvements of up to 16 % in bit rate reduction or, equivalently, more than 0.7 dB in PSNR. Index Terms—Biprediction, DIRECT mode, H.264, motion compensation, MPEG-4 AVC, spatial correlation, temporal correlation,
Alexis M. Tourapis, Feng Wu 0001, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2005 An effective variable block-size early termination algorithm for H.264 video coding
abstract
The H.264 video coding standard provides considerably higher coding efficiency than previous standards do, whereas its complexity is significantly increased at the same time. In an H.264 encoder, the most time-consuming component is variable block-size motion estimation. To reduce the complexity of motion estimation, an early termination algorithm is proposed in this paper. It predicts the best motion vector by examining only one search point. With the proposed method, some of the motion searches can be stopped early, and then a large number of search points can be skipped. The proposed method can work with any fast motion estimation algorithm. Experiments are carried out with a fast motion estimation algorithm that has been adopted by H.264. Results show that significant complexity reduction is achieved while the degradation in video quality is negligible.
Libo Yang, Keman Yu, Jiang Li 0008, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2005 A novel model-based rate-control method for portrait video coding
abstract
The rapid development of wireless networks and mobile devices has made mobile video communication a particularly promising service. We previously proposed an effective video form, scalable portrait video. In low-bandwidth conditions, portrait video possesses clearer shape, smoother motion, and much cheaper computational cost than discrete cosine transform (DCT)-based schemes. However, the bit rate of portrait video cannot be accurately modeled by a rate-distortion function as in DCT-based schemes. How to effectively control the bit rate is a hard challenge for portrait video. In this paper, we propose a novel model-based rate-control method. Although the coding parameters cannot be directly calculated from the target bit rate, we build a model between the bit-rate reduction and the percentage of less probable symbols (LPS) based on the principle of entropy coding, which is referred to as the LPS-rate model. We use this model to obtain the desired coding parameters. Experimental results show that the proposed method not only effectively controls the bit rate, but also significantly reduces the number of skipped frames. The principle of this method can also be applied to general bit plane coding in other image processing and video compression technologies.
Keman Yu, Jiang Li 0008, Cuizhu Shi, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2005 Scalable protection for MPEG-4 fine granularity scalability
abstract
The newly adopted MPEG-4 fine granularity scalability (FGS) video coding standard offers easy and flexible adaptation to varying network bandwidths and different application needs. Encryption for FGS should preserve such adaptation capabilities and enable intermediate stages to process encrypted data directly without decryption. In this paper, we propose two novel encryption algorithms for MPEG-4 FGS that meet these requirements. The first algorithm encrypts an FGS stream (containing both the base and the enhancement layers) into a single access layer and preserves the original fine granularity scalability and error resilience performance in an encrypted stream. The second algorithm encrypts an FGS stream into multiple quality layers divided according to either peak signal-to-noise ratio (PSNR) or bit rates, with lower quality layers being accessible and reusable by a higher quality layer of the same type, but not vice versa. Both PSNR and bit-rate layers are supported simultaneously so a layer of either type can be selected on the fly without decryption. The base layer for the second algorithm may be unencrypted to allow free view of the content at low-quality or content-based search of a video database without decryption. Both algorithms are fast, error-resilient, and have negligible compression overhead. The same approach can be applied to other scalable multimedia formats.
Bin B. Zhu, Chun Yuan 0003, Shipeng Li 0001
IEEE Trans. Multim.4
2004 DigiMetro - an application-level multicast system for multi-party video conferencing
abstract
The increasing demand for multi-party videoconferencing has aroused the research interest in the underlying multicast support. In this paper, we propose DigiMetro, an application-level multicast system tailored to small and impromptu videoconferencing. Breaking through the conventional wisdom to use shared overlay to handle multiple data sources, DigiMetro organizes the data delivery routes as source-specific trees, which are first constructed by a local greedy algorithm and then gradually improved by a global refinement procedure. Extensive simulation experiments demonstrate the efficiency of both algorithms. Moreover, DigiMetro is able to handle different video bit rates and provide different services over voice/video streams.
Chong Luo 0001, Jiang Li 0008, Shipeng Li 0001
GLOBECOM3
2004 Automatic image quality improvement for videoconferencing
abstract
In videoconferencing, the image quality is significantly affected by the illumination condition. Unsatisfactory illumination conditions may lead to underexposure or overexposure of the area of interest, in particular a human face. To resolve this issue, we propose a solution to improve image quality automatically by correcting exposure and enhancing contrast. Our work is characterized by a method for automatically building a skin-color model and a novel contrast enhancement approach. Some techniques that can reduce the computational cost are also introduced. Experimental results show that obvious improvement in image quality is achieved while the computation overhead is very small. The proposed solution can be integrated into videoconferencing systems and is especially suitable for scenarios where low-complexity computing is required.
Cuizhu Shi, Keman Yu, Jiang Li 0008, Shipeng Li 0001
ICASSP (3)4
2004 Color space compatible coding framework for YUV422 video coding
abstract
A new video coding framework for YUV422 video sources is proposed in this paper. It features color space compatibility to the more popular YUV420 syntax. Specifically, the chrominance components are separately into two parts and coded differently. The first part, together with the luminance component, conforms to the YUV420 layout and is coded the same way as a normal YUV420 video to produce a YUV420-compatible base bit stream. The second part, i.e., the remaining chrominance components, is coded to generate an enhancement chrominance bitstream for improving the chrominance quality. This is in sharp contrast to the YUV422 coding method of MPEG-2/4 standards where all the chrominance are coded together and in the same way. Consequently, the resulting YUV422 bitstream can be easily converted to a YUV420 bitstream by simple truncation instead of undergoing an expensive transcoding process. New coding modes are also introduced for more efficient coding of the enhancing chrominance components. Performance-wise, the new framework also outperforms existing methods thanks to the new coding modes introduced.
Lujun Yuan, Guobin Shen, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001
ICASSP (3)4
2004 Weighted motion estimation for efficiently coding scene transition video
abstract
Scene transition video has brought great challenges to current video coding methods because the traditional motion model of block displacement cannot efficiently represent transition motion. The paper first analyzes the features of static transitions. By incorporating the transition filter information into the coding scheme, the weighted motion estimation (WTME) technique is proposed to get accurate motion parameters for the transition video, thereby efficiently compensating both normal and transition motions among frames. Experimental results show that the proposed technique can significantly improve the coding performance of H.264 by up to 2.0 dB while coding scene transition video.
Xiaoyan Sun 0001, Hong Bao, Shipeng Li 0001
ICASSP (3)4
2004 Error-resilient unequal protection of fine granularity scalable video bitstreams
abstract
This paper deals with the optimal packet loss protection issue for streaming the fine granularity scalable (FGS) video bitstreams over IP networks. Unlike many other existing protection schemes, we develop an error-resilient unequal protection (ER-UEP) method that adds redundant information optimally for loss protection and, at the same time, cancels completely the dependency among bitstream after loss recovery. In our ER-UEP method, the FGS enhancement-layer bitstream is first packetized into a group of independent data packets, while each packet can be truncated to represent the original video signal at any fidelity (i.e., scalability). Parity packets are then created with intrinsic UEP capabilities that can easily adapt to the current channel conditions. Unlike conventional UEP schemes that suffer from bitstream contamination due to the dependency among packets, our method guarantees the successful decoding of all received bits, thus leading to a better error resilience as well as higher robustness (under varying and/or unclean channel conditions).
Hua Cai, Bing Zeng 0001, Guobin Shen, Shipeng Li 0001
ICC4
2004 Energy distributed update steps(edu) in lifting based motion compensated video coding
abstract
Subband video coding is an elegant scheme to fulfill high performance scalable video coding. In this paper, a new update scheme, energy distributed update steps (EDU), is proposed for the temporal transform in lifting based motion compensated video coding. The idea is to update where predict is made by distributing high-pass signals to the low-pass frame. The scheme avoids complex and inaccurate inversion of the motion information that used in the traditional update steps, thus it reduces computations in temporal transform. Experimental results show that the coding performances can be improved up to 0.77 dB.
Jizheng Xu, Feng Wu 0001, Shiqiang Yang, Shipeng Li 0001
ICIP5
2004 Variable block-size transform and entropy coding at the enhancement layer of FGS
abstract
This paper proposes a variable block-size transform and context-based entropy coding techniques for the enhancement layer of FGS (fine granularity scalable) video coding. First, the variable block-size transform is introduced into the enhancement layer to improve the performance of FGS in terms of both visual quality and PSNR. Different from that used in the traditional single layer coding, an R-D selection algorithm is proposed to optimally decide the transform size of each block, under consideration of consistent performance at a range of bit rates. Furthermore, to fully take advantage of the characteristics and correlations of symbols coded in the FGS enhancement layer, different context models are designed for the arithmetic coding according to symbol type and transform size. Experimental results show that the coding efficiency of FGS can be increased by 0.2-0.90 dB with the proposed techniques.
Jungong Han, Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Zhaoyang Lu
ICIP4
2004 A secure image authentication algorithm with pixel-level tamper localization
Jinhai Wu, Bin B. Zhu, Shipeng Li 0001, Fuzong Lin
ICIP3
2004 Layered motion estimation and coding for fully scalable 3d wavelet video coding
abstract
This paper proposes a framework of scalable motion estimation and coding with the structure of multilayers for 3D wavelet video coding. The motion representation consists of multiple layers. The encoder uses motion of all layers to perform analysis, while the decoder may receive only part of motion for synthesis. Different from other schemes, each layer of motion is a point optimized at a certain range of bit-rate. We observe that the distortion introduced by motion mismatch is highly independent with the rate for texture in a wide range. Therefore, to make the best trade-off between motion and texture under the constraint of a given bit rate, a motion layer decision algorithm is used to find the appropriate number of motion layers to be included into the bit-stream. The proposed framework also supports the spatial and temporal scalabilities of motion. Experimental results show significant improvement at low bit-rates and nearly no loss at high bit-rates with layered motion coding and optimal motion decision. The performance is approaching to the convex hull of those with multiple sets of nonscalable motion.
Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang
ICIP4
2004 Flexible p-picture (FLEXP) coding for tue efficient fine-granular scalabilitv (FGS)
Xiaoyan Sun 0001, Feng Wu 0001, Hong Bao, Shipeng Li 0001
ICIP5
2004 DigiParty - a decentralized multi-party video conferencing system
abstract
The increased speeds of PCs and networks have made media communication possible on the Internet. However, nearly ten years after the first release of Microsoft NetMeeting, Internet video telephony is still limited to the point-to-point communication mode. Today, people have a need for an easy-to-use multi-party video conferencing tool that can connect families and friends around the world over the Internet. We present DigiParty, a fully distributed multi-party video conferencing system. DigiParty employs a full mesh conferencing architecture and adopts a loosely coupled conferencing mode. A novel conference control protocol is designed with the system. DigiParty can be integrated with any existing instant messaging services and is applicable to all types of Internet connections.
Ling Chen 0001, Chong Luo 0001, Jiang Li 0008, Shipeng Li 0001
ICME4
2004 Efficient oracle attacks on Yeung-Mintzer and variant authentication schemes
abstract
The Yeung-Mintzer (Y-M) image authentication scheme has been well studied. Several vulnerabilities and modified schemes to fix them have been reported. We propose a novel oracle attack on the Y-M scheme and its variations. Our attack is very different from the previously proposed attacks. A single authenticated image plus access to a verifier (oracle) is enough in our attack. The verifier returns if a testing image is authentic or not. Locations of tampered pixels are not needed. To launch the attack, a single pixel is modified and the resulting image is sent to the verifier. Observation of outputs of the verifier is used to deduce the secret mapping functions and the embedded logo within an uncertainty of two possibilities. The deduced mapping functions are then used to modify the content of an authenticated image without detection or to authenticate an arbitrary image of the same size. Note that the logo is not used in the forgery so sophisticated protection of the logo cannot thwart the attack. Our attack is very efficient. Only 255 trials are needed to attack an 8-bit grayscale image and 765 trials for a 24-bit color image. The proposed attack can also be applied to attack pixel-wise variations of the Y-M scheme proposed to fix the previously reported vulnerabilities.
Jinhai Wu, Bin B. Zhu, Shipeng Li 0001, Fuzong Lin
ICME3
2004 An efficient key scheme for layered access control of MPEG-4 FGS video
abstract
The recently proposed scalable multi-layer FGS (fine granularity scalability) encryption (SMLFE) encrypts an MPEG-4 FGS stream into multiple PSNR and bitrate quality layers for layered access control. Both layer types are supported simultaneously. A simple key scheme was used in SMLFE. We propose a novel key scheme for SMLFE that reduces the number of keys maintained and managed by a license server for each protected MPEG-4 FGS stream to two. The new key scheme needs only one key contained in a license to be sent to a consumer. This scheme is based on a cryptographic secure hash function and the Diffie-Hellman key agreement. It satisfies all the requirements of SMLFE and can be used to replace the original simple key scheme for SMLFE. The secure one-way hash and intractability of the Diffie-Hellman and the related problems of computing discrete logarithms ensure the security of the new key scheme.
Bin B. Zhu, Min Feng 0002, Shipeng Li 0001
ICME3
2004 Direct macroblock coding for predictive (P) pictures in the H.264 standard
abstract
In this paper we introduce a new Inter Macroblock type within the H.264 (or MPEG-4 AVC) video coding standard that can further improve coding efficiency by exploiting the temporal correlation of motion within a sequence. This leads to a reduction in the bits required for encoding motion information, while retaining or even improving quality under a Rate Distortion Optimization Framework. An extension of this concept within the skip macroblock type of the same standard is also presented. Simulation results show that the proposed semantic changes can lead to up to 7.6% average bitrate reduction or equivalently 0.39dB quality improvement over the current H.264 standard.
Alexis M. Tourapis, Feng Wu 0001, Shipeng Li 0001
VCIP3
2004 Exploiting temporal correlation with adaptive block-size motion alignment for 3D wavelet coding
abstract
This paper proposes an adaptive block-size motion alignment technique in 3D wavelet coding to further exploit temporal correlations across pictures. Similar to B picture in traditional video coding, each macroblock can motion align from forward and/or backward for temporal wavelet de-composition. In each direction, a macroblock may select its partition from one of seven modes - 16x16, 8x16, 16x8, 8x8, 8x4, 4x8 and 4x4 - to allow accurate motion alignment. Furthermore, the rate-distortion optimization criterions are proposed to select motion mode, motion vectors and partition mode. Although the proposed technique greatly improves the accuracy of motion alignment, it does not directly bring the coding efficiency gain because of smaller block size and more block boundaries. Therefore, an overlapped block motion alignment is further proposed to cope with block boundaries and to suppress spatial high-frequency components. The experimental results show the proposed adaptive block-size motion alignment with the overlapped block motion alignment can achieve up to 1.0 dB gain in 3D wavelet video coding. Our 3D wavelet coder outperforms the MC-EZBC for most sequences by 1~2dB and we are doing up to 1.5 dB better than H.264.
Ruiqin Xiong, Feng Wu 0001, Shipeng Li 0001, Zixiang Xiong, Ya-Qin Zhang
VCIP3
2004 Advanced motion threading for 3D wavelet video coding
Lin Luo 0004, Feng Wu 0001, Shipeng Li 0001, Zixiang Xiong, Zhenquan Zhuang
Signal Process. Image Commun.3
2004 Seamless switching of scalable video bitstreams for efficient streaming
abstract
Efficient adaptation to channel bandwidth is broadly required for effective streaming video over the Internet. To address this requirement, a novel seamless switching scheme among scalable video bitstreams is proposed in this paper. It can significantly improve the performance of video streaming over a broad range of bit rates by fully taking advantage of both the high coding efficiency of nonscalable bitstreams and the flexibility of scalable bitstreams, where small channel bandwidth fluctuations are accommodated by the scalability of a single scalable bitstream, whereas large channel bandwidth fluctuations are tolerated by flexible switching between different scalable bitstreams. Two main techniques for switching between video bitstreams are proposed. Firstly, a novel coding scheme is proposed to enable drift-free switching at any frame from the current scalable bitstream to one operated at lower rates without sending any overhead bits. Secondly, a switching-frame coding scheme is proposed to greatly reduce the number of extra bits needed for switching from the current scalable bitstream to one operated at higher rates. Compared with existing approaches, such as switching between nonscalable bitstreams and streaming with a single scalable bitstream, our experimental results clearly show that the proposed scheme brings higher efficiency and more flexibility in video streaming.
Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001, Ya-Qin Zhang
IEEE Trans. Multim.3
2003 PLI: A New Framework to Protect Digital Content for P2P Networks
Guofei Gu, Bin B. Zhu, Shipeng Li 0001, Shiyong Zhang
ACNS3
2003 Wavelet video coding via a spatially adaptive lifting structure
abstract
We present a spatially adaptive wavelet video coding technique with an update-first lifting structure. A common problem in many adaptive-transform frameworks is the introduction of a large overhead to address side information. We demonstrate that our structure does not need to transmit any side information to synchronize the encoder and decoder. We incorporate this technique in a motion compensated wavelet video codec. The experimental results confirm the performance improvement.
Zhen Li 0008, Feng Wu 0001, Shipeng Li 0001, Edward J. Delp
ICASSP (3)3
2003 Accelerating video decoding using GPU
abstract
Most modern computers or game consoles are equipped with powerful graphics processing units (GPU) to accelerate graphics operations. There is a trend that the power of GPU outgrows that of the CPU (central processing unit). However, the GPU engines are specially designed for graphics operations. Can we take advantage of the powerful GPU engines for more general operations other than pure graphics operations? The answer is positive. In this study, we present schemes that map other non-graphics operations into graphics engines with an example application of accelerating video decoding with the assistance of GPU. Our results show that significant speed-up can be achieved by leveraging the GPU power. Specifically, we have achieved real-time playback of high definition video on a PC with an Intel Pentium III 667 MHz CPU and an nVidia GeForce3 GPU.
Guobin Shen, Lihua Zhu, Shipeng Li 0001, Harry Shum, Ya-Qin Zhang
ICASSP (4)3
2003 Video object extraction using extended intelligent scissors
abstract
An approach to user-assisted video object extraction at pixel-level accuracy is proposed. This approach is based on a novel framework that aims at reducing users' workload while making full use of user interaction. The framework takes over multipass scan on a video sequence and bidirectional tracking on the sequence's sub-shots during each pass. Systems based on this framework can not only free users from laborious work, but also be allowed to embed various complex but effective tracking algorithms. To fulfill a system, we propose an extended intelligent scissors for machine to precisely track VOs. The approach to tracking is a multilevel process, locating VO boundaries from a coarse level to a fine one and yielding pixel-accurate results of extraction. Experiments on MPEG test sequences show that this approach together with the proposed framework performs well in practice.
Shipeng Li 0001
ICIP (2)2
2003 Layer-correlated motion estimation and motion vector coding for the 3D-wavelet video coding
abstract
This paper proposes an efficient layer-correlated scheme to educe the bit cost for motion vectors in the 3D wavelet coding, previous works show that incorporating motion alignment into lifting structure enables the 3D wavelet coding to provide a competitive performance to the state-of-the-art JVT standard. In general, the temporal wavelet decomposition consists of multiple layers, while each layer adopts one set of motion vectors to achieve high coding efficiency and temporal scalability. Since the current schemes code these MVs independently, this greatly increases the bit cost for coding MV. In order to reduce the motion cost, the proposed scheme performs motion estimation considering the MV correlation among layers. Several modes are proposed to describe the different local correlations at the macroblock level. By an R-D optimized mode selection engine, the proposed scheme can save up to 33% bits of MVs at the similar texture quality.
Lin Luo 0004, Feng Wu 0001, Shipeng Li 0001, Zhenquan Zhuang
ICIP (2)3
2003 The improved SP frame coding technique for the JVT standard
abstract
An efficient and flexible coding technique is proposed in this paper inspired by the SP frame in the H.26L standard, which can achieve a drift-free bitstream switching at the predicted frame. The proposed scheme improves the coding efficiency of the SP frames in the H.26L standard by limiting the mismatch between the references for the prediction and reconstruction with two DCT coefficient coding modes and the rate-distortion optimization. Furthermore, the proposed scheme allows independent quantization parameters for up-switching and down-switching bitstreams. It further reduces the switching bitstream size while keeping the coding efficiency of the normal bitstreams. More rapid and frequent down-switching than up-switching and much smaller size of down-switching bitstream can be achieved with the proposed SP technique. These are very desirable features for any TCP-friendly protocols. Compared with the original SP method for H.26L, the proposed SP method improves the coding efficiency up to 1.0 dB. This SP technique has been officially accepted by the JVT standard.
Xiaoyan Sun 0001, Shipeng Li 0001, Feng Wu 0001, Guobin Shen, Wen Gao 0001
ICIP (3)2
2003 Rate distortion optimized mode decision in the scalable video coding
abstract
In this paper, we discuss how to apply the rate distortion technique to select the optimal mode in the scalable coding. Firstly, we analyze this problem from a general scalable model and point out that the complicated dependencies among the different macroblocks and layers make the original independency assumption no longer a right approximation. Secondly, we propose an EOD function to estimate this dependency and derive a simple formula of this function. We apply the proposed algorithm to the H.26L PFGS, and the experimental results show that the algorithm significantly improves the coding efficiency of the H.26L PFGS. Further studies on how to design the EOD function more accurately is quite interesting and significant.
Feng Wu 0001, Shipeng Li 0001
ICIP (3)3
2003 Layered access control for MPEG-4 FGS video
abstract
MPEG-4 has recently adopted the fine granularity scalability (FGS) video coding technology which enables easy and flexible adaptation to bandwidth fluctuations and device capabilities. Encryption for FGS should preserve such adaptation capabilities and allow intermediate stages in the delivery to process the media on the ciphertext directly. In this paper, we propose a novel scalable access control scheme with this property for the MPEG-4 FGS format. It offers free browsing of the low-quality base layer video but controls the access to the enhancement layer at different service levels based on either PSNR or bitrates. Both types of service levels are supported simultaneously without jeopardizing each other's security. The scheme is fast and degrades neither compression efficiency nor error resilience of the MPEG-4 FGS. The approach is also applicable to other scalable multimedia.
Chun Yuan 0003, Bin B. Zhu, Ming Su, Shipeng Li 0001, Yuzhuo Zhong
ICIP (1)5
2003 L-TFRC: an end-to-end congestion control mechanism for video streaming over the Internet
abstract
Real-time multimedia applications over the Internet have posed a lot of challenges due to the lack of quality of service (QoS) guarantees, frequent fluctuations in channel bandwidth, and packet losses. To address these issues, a great deal of research has been done in both video coding and video transmission fields. In this paper we present a logarithm-based TCP-friendly rate control (L-TFRC) mechanism, which can estimate the available bandwidth more accurately and improve the smoothness of the multimedia streaming significantly. We also apply it to a progressive fine granularity scalable (PFGS)-based video streaming. Both simulations and experiments over the Internet confirm the performance of L-TFRC.
Zhen Li 0008, Guobin Shen, Shipeng Li 0001, Edward J. Delp
ICME3
2003 Practical real-time video codec for mobile devices
abstract
Real-time software-based video codec is widely used on PCs with relatively strong computing capability. However, mobile devices, such as pocket PCs and handheld PCs, still suffer from weak computational power, short battery lifetime and limited display capability. We developed a practical low-complexity real-time video codec for mobile devices. Several methods that can significantly reduce the computational cost are adopted in this codec and described in this paper, including a predictive algorithm for motion estimation, the integer discrete cosine transform (IntDCT), and a DCT/quantizer bypass technique. A real-time video communication implementation of the proposed coded is also introduced. Experiments show that substantial computation reduction is achieved while the loss in video quality is negligible. The proposed codec is very suitable for scenarios where low-complexity computing is required.
Keman Yu, Jiangbo Lu, Jiang Li 0008, Shipeng Li 0001
ICME4
2003 Multimetric evaluation protocol for user-assisted video object extraction systems
Shipeng Li 0001
VCIP2
2003 Advanced lifting-based motion-threading (MT) technique for 3D wavelet video coding
Lin Luo 0004, Feng Wu 0001, Shipeng Li 0001, Zhenquan Zhuang
VCIP3
2003 Bitstream switching for progressive fine granularity scalable video coding
Jizheng Xu, Feng Wu 0001, Shipeng Li 0001
VCIP3
2003 A Motion Compensated Lifting Wavelet Codec for 3D Video Coding
Lin Luo 0004, Jin Li 0001, Shipeng Li 0001, Zhenquan Zhuang
J. Comput. Sci. Technol.3
2003 Scalable portrait video for mobile video communication
abstract
Wireless networks have been rapidly developing in recent years. General Packet Radio Service (GPRS) and Code Division Multiple Access (CDMA 1X) for wide areas, and 802.11 and Bluetooth for local areas have already emerged. Broadband wireless networks urgently call for rich contents for consumers. Among various possible applications, video communication is one of the most promising for mobile devices on wireless networks. This paper describes the generation, coding, and transmission of an effective video form, scalable portrait video for mobile video communication. As an expansion to bilevel video, portrait video is composed of more gray levels, and therefore possesses higher visual quality while it maintains a low bit rate and low computational costs. Portrait video is a scalable video in that each video with a higher level always contains all the information of the video with a lower level. The bandwidths of 2-4-level portrait videos fit into the bandwidth range of 20-40 kbps that GPRS and CDMA 1X can stably provide; therefore, portrait video is very promising for video broadcast and communication on 2.5-G wireless networks. With portrait video technology, we are the first to enable two-way video communication on pocket PCs and handheld PCs.
Jiang Li 0008, Keman Yu, Tielin He, Yunfeng Lin, Shipeng Li 0001, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.5
2002 Optimal rate allocation for macroblock-based progressive fine granularity scalable video coding
abstract
This paper addresses the problem of optimal rate allocation for the macroblock-based progressive fine granularity scalable (PFGS) video coding. To solve this complicated problem, the error propagation pattern in the macroblock-based PFGS is first investigated. An effective drifting model is established subsequently for estimating the drifting for each enhancement bit stream segment encoded by the macroblock-based PFGS. The distortion reduction for the current frame and the estimated drifting suppression for the subsequent frames form the actual contribution of the enhancement layer bitstream. The equal-slope argument is then applied to select the best bit stream segments for the given bandwidth. Experiments show that our optimal rate allocation outperforms the uniform rate allocation by 0.3-1.4 dB.
Hua Cai, Guobin Shen, Shipeng Li 0001, Bing Zeng 0001
ICIP (3)3
2002 Enhancing multimedia streaming performance through peer-paired collaboration
abstract
In this paper, a novel multimedia streaming framework called peer-paired pyramid streaming (P/sup 3/S) is proposed. The philosophy of P/sup 3/S is to enable collaboration between clients so as to bring in better performance. The structure of P/sup 3/S is basically a hybrid client/server and peer-to-peer structure and exhibits a triangle-cell based hierarchy. Based on P/sup 3/S, performance enhancement techniques are designed to increase the aggregated bandwidth of all participants. We present an optimal data allocation algorithm, which maximizes the overall throughput of the whole streaming session. We also present a greedy data allocation algorithm that is slightly suboptimal but much simpler. Extensive simulations were performed to demonstrate the effectiveness of proposed techniques.
Guobin Shen, Shipeng Li 0001, Yuzhuo Zhong
ICIP (3)3
2002 Efficient and universal scalable video coding
abstract
This paper proposes a unified efficient and universal scalable video coding framework that supports different scalabilities, such as fine granularity quality, temporal, spatial and complexity scalabilities. The proposed framework is established upon the recent studies in fine granularity scalable (FGS) video coding. It contains two key points. Firstly, in order to improve the coding efficiency of the proposed framework, more than one motion compensation loop is used. Since high quality references are introduced into the enhancement layer coding, the proposed framework can efficiently compress different-resolution video at different layers for the purpose of the complexity and spatial scalability. Secondly, the drifting reduction techniques are studied in this paper. This helps the proposed framework to maintain good performance at lower enhancement bit rates. By defining coding modes, a macroblock level control mechanism is developed to achieve a better trade-off between low drifting errors and high coding efficiency.
Feng Wu 0001, Shipeng Li 0001, Xiaoyan Sun 0001, Ya-Qin Zhang
ICIP (2)2
2002 Error concealment for fine granularity scalable video transmission
abstract
In this paper we present an efficient error concealment (EC) method for the fine granularity scalable (FGS) video transmission. The proposed EC method exploits both the temporal and spatial correlations in an FGS encoded bitstream. In our scheme, the temporal redundancy is used to improve the quality of contaminated regions, and the intensity of the temporal correlation in contaminated regions is estimated by exploiting the spatial correlation in the surrounding high-quality regions. To maximally utilize the spatial correlation for estimation, we also propose two interleaving patterns that can avoid packetizing the neighboring regions into the same packet. Experiments show that our EC method achieves very good performance and is robust to different bandwidths and different sequences.
Hua Cai, Guobin Shen, Feng Wu 0001, Shipeng Li 0001, Bing Zeng 0001
ICME (1)4
2002 Efficient and flexible drift-free video bitstream switching at predictive frames
abstract
We propose an efficient and flexible coding scheme inspired by the SP picture technique in H.26L TML; it can achieve drift-free bitstream switching at predictive frames. Firstly, the proposed scheme improves the coding efficiency of the SP frames in H.26L TML by (1) reducing the number of quantization modules in the encoding path; (2) eliminating the mismatch between references for the prediction and the reconstruction; (3) outputting a high quality image for display purpose before the quantization step in the reconstruction loop. Secondly, the proposed scheme allows independent quantization parameters for up-switching and down-switching bitstreams. It can further reduce the switching bitstream size while keeping the coding efficiency of the normal bitstreams. It allows more rapid and frequent down-switching than up-switching. Furthermore, the size of the down-switching bitstream can be much smaller than that of the up-switching one. This is a very desirable feature for any TCP-friendly protocols currently used in most existing streaming systems.
Xiaoyan Sun 0001, Shipeng Li 0001, Feng Wu 0001, Guobin Shen, Wen Gao 0001
ICME (1)2
2002 Efficient video coding with hybrid spatial and fine-grain SNR scalabilities
Feng Wu 0001, Shipeng Li 0001, Ran Tao 0003, Yue Wang 0001
VCIP3
2002 Optimal rate allocation for progressive fine granularity scalable video coding
abstract
We examine the enhancement-layer rate allocation problem in progressive fine granularity scalable (PFGS) video coding. The problem arises from the fact that different frames in the enhancement layer have different rates in PFGS coding. A rate-distortion (R-D) function for a multiframe group is first established for enhancement-layer PFGS coding, followed by experiments to verify its validity using real test sequences. Optimal rate allocation among frames in the group is then given based on the R-D function, together with a simple implementation that is suitable for applications such as streaming video. Experiments show that, compared with uniform bit allocation, optimal bit allocation not only makes the quality variation in decoded video much smoother, but also improves the average PSNR of PFGS coding by 0.3-0.5 dB.
Zixiang Xiong, Feng Wu 0001, Shipeng Li 0001
IEEE Signal Process. Lett.4
2002 Memory-constrained 3D wavelet transform for video coding without boundary effects
abstract
Three-dimensional (3D) wavelet-based scalable video coding provides a viable alternative to standard MC-DCT coding. However, many current 3D wavelet coders experience severe boundary effects across group of pictures (GOP) boundaries. This paper proposes a memory-efficient transform technique via lifting that effectively computes wavelet transforms of a video sequence continuously on the fly, thus eliminating the boundary effects due to limited length of individual GOPs. Coding results show that the proposed scheme completely eliminates the boundary effects and gives superb video playback quality.
Jizheng Xu, Zixiang Xiong, Shipeng Li 0001, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.3
2001 Fine-granularity spatially scalable video coding
abstract
We propose a novel architecture for spatially scalable video coding, namely, fine-granularity spatially scalable (FGSS) coding. The traditional layered spatially scalable coding provides only coarse scalability in which the bit-stream can be decoded only at a few fixed resolutions, but not something in between. The proposed FGSS scheme provides a fine-granularity property to the spatial scalability. In this scheme, the bit plane technique is combined with spatial scalability, thus a fine granularity increase in the image quality from low-resolution to high-resolution can be obtained. In addition, the proposed scheme provides a flexible embedded bitstream that can be decoded up to any point in the enhancement layer bitstream from low-resolution to high-resolution. This feature further enables efficient video streaming over the Internet where the scalable bitstream, can adapt to the widely fluctuating bandwidth. The FGSS coding scheme extends new functionalities such as multi-resolution, fine granularity, channel adaptation and error-recovery properties to scalable video coding, thus it can satisfy different user clients with a wide range of channel bandwidth and screen resolution.
Feng Wu 0001, Shipeng Li 0001, Yuzhuo Zhong, Ya-Qin Zhang
ICASSP3
2001 Macroblock-based progressive fine granularity scalable (PFGS) video coding with flexible temporal-SNR scalablilities
abstract
We proposed a flexible and efficient architecture for scalable video coding, namely, the macroblock (MB)-based progressive fine granularity scalable video coding with temporal-SNR scalabilities (PFGST). The proposed architecture can provide not only much improved coding efficiency but also simultaneous SNR scalability and temporal scalability. Building upon the original frame-based progressive fine granularity scalable (PFGS) coding approach, the MB-based PFGS scheme is first proposed. Three INTER modes and the corresponding mode selection mechanism are presented for coding the SNR enhancement MBs in order to make a good trade-off between low drifting errors and high compression efficiency. Furthermore, temporal scalability is introduced into the MB-based PFGS, which forms the MB-based PFGST scheme. Two coding modes are proposed for coding the temporal enhancement MBs. Since it would not cause any error propagation if using the high quality reference in the temporal enhancement MB coding, the coding efficiency of the PFGST is highly improved by always choosing the most suitable reference for the temporal scalable coding. Experimental results show that the MB-based PFGST video coding scheme can significantly improve the coding efficiency up to 2.8 dB compared with the FGST scheme adopted in MPEG-4, while supporting full SNR, full temporal, and hybrid SNR-temporal scalabilities according to the different requirements from the channels, the clients or the servers.
Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001, Ya-Qin Zhang
ICIP (2)3
2001 Compression of M-FISH images using 3-D ESCOT
abstract
This paper introduces a lossy to lossless coding technique for compression of multitarget fluorescence in situ hybridization (M-FISH) images using 3-D embedded subband coding with optimal truncation (3-D ESCOT) (Xu et al.). With a lifting-based integer wavelet decomposition, 3-D ESCOT achieves about twice as much compression as Lempel-Ziv (WinZip) coding-the current method for archiving M-FISH images. The lossy coding performance of 3-D ESCOT is significantly better than that of 2-D based JPEG-2000.
Jizheng Xu, Zixiang Xiong, Qiang Wu 0007, Shipeng Li 0001
ICIP (2)4
2001 Motion Compensated Lifting Wavelet And Its Application In Video Coding
abstract
A motion compensated lifting (MCLIFT) framework is proposed for the 3D wavelet video coder. By using bi-directional motion compensation in each lifting step of the temporal direction, the video frames are effectively de-correlated. With proper entropy coding and bitstream packaging schemes, the MCLIFT wavelet video coder can be scalable in frame rate and quality level. Experimental results show that the MCLIFT video coder outperforms the 3D wavelet video coder with the same entropy coding scheme by an average of 1.1-1.6dB, and outperforms MPEG-4 coder by an average of 0.9-1.4dB.
Lin Luo 0004, Jin Li 0001, Shipeng Li 0001, Zhenquan Zhuang, Ya-Qin Zhang
ICME3
2001 Macroblock-Based Progressive Fine Granularity Scalable Video Coding
abstract
In this paper, we proposed a flexible and efficient architecture for scalable video coding, namely, the macroblock (MB)-based progressive fine granularity scalable video coding with temporal-SNR scalabilities (PFGST in short). The proposed architecture can provide not only much improved coding efficiency but also simultaneous SNR scalability and temporal scalability. Building upon the original frame-based progressive fine granularity scalable (PFGS) coding approach, the MB-based PFGS scheme is first proposed. Three INTER modes and the corresponding mode selection mechanism are presented for coding the SNR enhancement MBs in order to make a good trade-off between low drifting errors and high compression efficiency. Furthermore, temporal scalability is introduced into the MB-based PFGS, which forms the MB-based PFGST scheme. Two coding modes are proposed for coding the temporal enhancement MBs. Since it would not cause any error propagation if using the high quality reference in the temporal enhancement MB coding, the coding efficiency of the PFGST is highly improved by always choosing the most suitable reference for the temporal scalable coding. Experimental results show that the MB-based PFGST video coding scheme can significantly improve the coding efficiency up to 2.8dB compared with the FGST scheme adopted in MPEG-4, while supporting full SNR, full temporal, and hybrid SNR-temporal scalabilities according to the different requirements from the channels, the clients or the servers. 1.
Xiaoyan Sun 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001, Ya-Qin Zhang
ICME3
2001 Interactive Tracker - A Semi-Automatic
Wenyin Liu, Shipeng Li 0001
ICME3
2001 A framework for efficient progressive fine granularity scalable video coding
abstract
A basic framework for efficient scalable video coding, namely progressive fine granularity scalable (PFGS) video coding is proposed. Similar to the fine granularity scalable (PGS) video coding in MPEG-4, the PFGS framework has all the features of FGS, such as fine granularity bit-rate scalability, channel adaptation, and error recovery. On the other hand, different from the PGS coding, the PFGS framework uses multiple layers of references with increasing quality to make motion prediction more accurate for improved video-coding efficiency. However, using multiple layers of references with different quality also introduces several issues. First, extra frame buffers are needed for storing the multiple reconstructed reference layers. This would increase the memory cost and computational complexity of the PFGS scheme. Based on the basic framework, a simplified and efficient PFGS framework is further proposed. The simplified PPGS framework needs only one extra frame buffer with almost the same coding efficiency as in the original framework. Second, there might be undesirable increase and fluctuation of the coefficients to be coded when switching from a low-quality reference to a high-quality one, which could partially offset the advantage of using a high-quality reference. A further improved PFGS scheme can eliminate the fluctuation of enhancement-layer coefficients when switching references by always using only one high-quality prediction reference for all enhancement layers. Experimental results show that the PFGS framework can improve the coding efficiency up to more than 1 dB over the FGS scheme in terms of average PSNR, yet still keeps all the original properties, such as fine granularity, bandwidth adaptation, and error recovery. A simple simulation of transmitting the PFGS video over a wireless channel further confirms the error robustness of the PFGS scheme, although the advantages of PFGS have not been fully exploited.
Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.2
2001 Arbitrarily shaped video-object coding by wavelet
abstract
Video-object coding is one of the most important functionalities proposed by MPEG-4. We propose a new wavelet method to encode the texture of an arbitrarily shaped object, both for the still and for the video-object. The method uses the shape adaptive wavelet transform (SA-DWT) in MPEG-4 still object coding, but with a computationally more efficient lifting implementation. The transformed object coefficients are then quantized and entropy encoded with a partial bit-plane embedded coder, which greatly improves the coding efficiency. We denote the coding algorithm as the video-object wavelet (VOW) coder. Experimental results show that VOW significantly outperforms MPEG-4 in still-object coding, and achieves a comparable performance in video-object coding in terms of PSNR. Moreover, the VOW decoded object looks better subjectively, with less annoying blocking artifacts than that of MPEG-4.
Guiwei Xing, Jin Li 0001, Shipeng Li 0001, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.3
2000 DCT-Prediction Based Progressive Fine Granularity Scalability Coding
abstract
We propose a novel architecture for scalable video coding, namely, progressive fine granularity scalable (PEGS) coding, which can provide a high coding efficiency along with good bandwidth adaptation and error recovery properties. Unlike the fine granularity scalable (FGS) coding in the MPEG-4 proposal, some of the enhancement layers in a current frame are predicted from a high quality enhancement layer in a reference frame, rather than always from the base layer. Using a high quality enhancement layer as the reference makes the motion prediction more accurate to improve the coding efficiency. On the other hand, the use of multiple layers of different quality references may also result in increases and fluctuations of the prediction residues to be coded when switching the references, which may limit the coding efficiency improvement. A multiple-layer conditional replenishment approach is used to eliminate this kind of fluctuation. Experimental results show that our coding scheme can improve the coding efficiency up to 0.5 dB compared with fine granularity scalability coding.
Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang
ICIP2
2000 Generic, scalable and efficient shape coding for visual texture objects in MPEG-4
abstract
This paper presents a generic, scalable, efficient shape coding scheme for scalable object-oriented visual texture. The base-layer coding scheme used is similar to the binary CAE coding scheme adopted in MPEG-1 video. The proposed scheme introduces a new set of generic contexts to efficiently encode (predict) enhanced shape layer based on the lower spatial layer using a context-based arithmetic coder. It is not dependent on any specific sub-sampling filters, thus it can generate the exact shape matching the texture decomposed using any wavelet filters. The proposed shape coding scheme was adopted in MPEG-4 Version 2 Standard to enable the visual texture wavelet coding to be a true spatially-scalable object-based texture coding technique. In addition, when operated in macro-block mode, the proposed scheme provides full backward compatibility with the MPEG-3 scalable shape coding for video objects. A simple solution to solve the chroma shape mismatch is also presented and was adopted in MPEG-4 Version 1 visual texture coding part. The results show that the proposed scalable shape coding scheme also achieves significant better coding efficiency than the non-scalable shape coding and the other competing shape coding scheme.
Shipeng Li 0001, Iraj Sodagar
ISCAS1
2000 Automatic extraction of moving objects using multiple features and multiple frames
abstract
This paper introduces a novel automatic video object extraction algorithm based on combination of color and motion segmentation results. The algorithm includes five parts: preprocessing, color segmentation, motion segmentation, combination of color and motion segmentation of multiple frames, post-processing. The performance of this algorithm is very promising, resulting in pixel-wise accuracy of extracted objects. Since it is an automatic extraction algorithm, it can be very useful in some real time video processing system based on video objects.
Jinhui Pan, Shipeng Li 0001, Ya-Qin Zhang
ISCAS2
2000 Arbitrarily shaped video object coding by wavelet
abstract
Video object coding is one of the most important functionalities proposed by MPEG4. In this paper, we propose a new wavelet method to encode the texture of an arbitrarily shaped object, both for the still and video object. The method uses the shape adaptive wavelet transform (SA-DWT) in MPEG4 still object coding, but with a computationally more efficient lifting implementation. The transformed object coefficients are then quantized and entropy encoded with a partial bitplane embedded coder, which greatly improves the coding efficiency. We denote the coding algorithm as a video object wavelet (VOW) coder. Experimental results show that VOW significantly outperforms MPEG4 in still object coding, and achieves a comparable performance in video object coding in terms of PSNR. Moreover, the VOW decoded object looks better subjectively, with less annoying blocking artifacts than that of MPEG4.
Guiwei Xing, Jin Li 0001, Shipeng Li 0001, Ya-Qin Zhang
ISCAS3
2000 Three-dimensional shape-adaptive discrete wavelet transforms for efficient object-based video coding
Jizheng Xu, Shipeng Li 0001, Ya-Qin Zhang
VCIP2
2000 Shape-adaptive discrete wavelet transforms for arbitrarily shaped visual object coding
abstract
This paper presents a shape-adaptive wavelet coding technique for coding arbitrarily shaped still texture. This technique includes shape-adaptive discrete wavelet transforms (SA-DWTs) and extensions of zerotree entropy (ZTE) coding and embedded zerotree wavelet (EZW) coding. Shape-adaptive wavelet coding is needed for efficiently coding arbitrarily shaped visual objects, which is essential for object-oriented multimedia applications. The challenge is to achieve high coding efficiency while satisfying the functionality of representing arbitrarily shaped visual texture. One of the features of the SA-DWTs is that the number of coefficients after SA-DWTs is identical to the number of pixels in the original arbitrarily shaped visual object. Another feature of the SA-DWT is that the spatial correlation, locality properties of wavelet transforms, and self-similarity across subbands are well preserved in the SA-DWT. Also, for a rectangular region, the SA-DWT becomes identical to the conventional wavelet transforms. For the same reason, the extentions of ZTE and EZW to coding arbitrarily shaped visual objects carefully treat "don't care" nodes in the wavelet trees. Comparison of shape-adaptive wavelet coding with other coding schemes for arbitrarily shaped visual objects shows that shape-adaptive wavelet coding always achieves better coding efficiency than other schemes. One implementation of the shape-adaptive wavelet coding technique has been included in the new multimedia coding standard MPEG-4 for coding arbitrarily shaped still texture. Software implementation is also available.
Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
1997 A video coding algorithm using vector-based techniques
abstract
This paper presents an algorithm proposal submitted to MPEG-4 for video coding. The proposed algorithm addresses the functionality of improved coding efficiency for compression. It uses vector-based techniques for coding intraframes (the first frame and subsequent refreshing key frames) and motion-compensated difference frames. It uses the same motion estimation and motion compensation techniques as H.263. A video frame (I or P frame) is first decomposed into a set of vector bands using a vector wavelet transform. This stage of vector-based signal processing makes subsequent vector quantization in the vector bands very efficient. Lattice vector quantization is then used in the vector bands. A 100% labeling efficiency is achieved for lattice vector quantization by using a set of generalized labeling algorithms for various important lattices with pyramid and sphere boundaries. Finally, entropy coding is used to code the indexes generated from lattice vector quantization. Our coding results have shown that a gain in peak signal-to-noise ratio (PSNR) up to 8 dB for intraframe coding and up to 6 dB for interframe coding can be achieved over H.263. Subjective quality improvement of the proposed algorithm over H.263 can be easily observed.
Hugh Q. Cao, Shipeng Li 0001, Fan Ling, Scott A. Segan, Hongqiao Sun, John Wus, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.3
1996 Very low bit rate video coding using vector-based techniques
abstract
This paper reports the advances of using vector-based techniques for very low bit rate video coding. High efficiency has been achieved for coding intraframes (the first frame and subsequent refreshing key frames) and motion compensated difference frames of a video sequence. A video frame (I or P frame) is first decomposed into a set of vector bands using a vector wavelet transform. Adaptive lattice vector quantization is then used in the vector bands. A 100% labeling efficiency is achieved for lattice vector quantization by using a set of generalized labeling algorithms for various important lattices with pyramid and sphere boundaries. An adaptive algorithm is used to determine the type of lattice vector quantisation according to the statistical distribution of the vectors in the vector wavelet domain. Finally, entropy coding is used to code the indexes generated from lattice vector quantization. Our coding results show that a gain of 3 to 8 dBs in PSNR for intraframe coding can be achieved at a bitrate level of 16 Kbpf over H.263. Subjective quality improvement of the proposed algorithm over H.263 can be observed.
Hugh Q. Cao, Shipeng Li 0001, Fan Ling, Scott A. Segan, H. Q. Sun, John Wus, Ya-Qin Zhang
ICIP (1)3