EDBT 2026 Demo / reviewers in the wild / expert
Shiqiang Yang
dblp:y/ShiqiangYang · also Shi-Qiang Yang
· DBLP profile ↗
190ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0001-5356-4094ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 111 · 1 since 2021Databases, data management, data science and information retrieval · 39 · 1 since 2021Artificial intelligence and machine learning · 34 · 1 since 2021Computer networks · 16Applied, interdisciplinary, general and emerging computing · 12 · 3 since 2021Systems, architecture and hardware · 9Human-computer interaction and ubiquitous computing · 5Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
33 papers |
Data mining · 47% Recommender systems · 22% Web and social media mining · 19% | |
| Artificial intelligence
15 papers |
Probabilistic and Bayesian machine learning · 26% Transfer learning and domain adaptation · 20% Video understanding and tracking · 10% | |
| Computer networks
14 papers |
Content delivery and video streaming · 72% Network management and operations · 11% Network measurement and analytics · 5% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Computational social science and digital humanities · 100% | |
| Computer graphics and multimedia
10 papers |
Multimedia analysis and retrieval · 47% Image and video processing · 42% Geometric modeling and processing · 5% | |
| Theoretical computer science
3 papers |
Mathematical optimization · 63% Coding theory · 37% | |
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Cloud and datacenter computing · 63% Distributed systems · 17% Parallel and multicore computing · 17% |
Topics — the 30 heaviest of 134, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
confounder balancing |
0.9 | 2 | 2022 | Data-Driven Variable Decomposition for Treatment Effect Estimation · IEEE Trans. Knowl. Data Eng. 2022 Estimating Treatment Effect in the Wild via Differentiated Confounder Balancing · KDD 2017 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.8 | 2 | 2020 | Learning to Select Base Classes for Few-Shot Classification · CVPR 2020 Learning to Learn Image Classifiers With Visual Analogy · CVPR 2019 |
Data mining
anomaly detection |
0.6 | 3 | 2016 | Spotting Suspicious Behaviors in Multimodal Data: A General Metric and Algorithms · IEEE Trans. Knowl. Data Eng. 2016 A General Suspiciousness Metric for Dense Blocks in Multimodal Data · ICDM 2015 Cascading outbreak prediction in networks: a data-driven approach · KDD 2013 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
treatment effect estimation |
0.6 | 2 | 2017 | Estimating Treatment Effect in the Wild via Differentiated Confounder Balancing · KDD 2017 Treatment Effect Estimation with Data-Driven Variable Decomposition · AAAI 2017 |
Computational social science and digital humanities
causal inference |
0.6 | 1 | 2022 | Data-Driven Variable Decomposition for Treatment Effect Estimation · IEEE Trans. Knowl. Data Eng. 2022 |
Computational social science and digital humanities › causal inference
treatment effect estimation |
0.6 | 1 | 2022 | Data-Driven Variable Decomposition for Treatment Effect Estimation · IEEE Trans. Knowl. Data Eng. 2022 |
Security and privacy of machine learning
adversarial attack |
0.6 | 1 | 2022 | Adversarial Eigen Attack on BlackBox Models · CVPR 2022 |
Security and privacy of machine learning › adversarial attack
black-box attack |
0.6 | 1 | 2022 | Adversarial Eigen Attack on BlackBox Models · CVPR 2022 |
Security and privacy of machine learning › adversarial attack
transferable adversarial attack |
0.6 | 1 | 2022 | Adversarial Eigen Attack on BlackBox Models · CVPR 2022 |
Recommender systems
social recommendation |
0.6 | 3 | 2015 | Social Recommendation with Cross-Domain Transferable Knowledge · IEEE Trans. Knowl. Data Eng. 2015 Scalable Recommendation with Social Contextual Information · IEEE Trans. Knowl. Data Eng. 2014 Joint Social and Content Recommendation for User-Generated Videos in Online Social Network · IEEE Trans. Multim. 2013 |
Data mining › structured data mining
graph mining |
0.5 | 2 | 2018 | Power-law Distribution Aware Trust Prediction · IJCAI 2018 CatchSync: catching synchronized behavior in large directed graphs · KDD 2014 |
Content delivery and video streaming
content delivery network |
0.5 | 3 | 2015 | CPCDN: Content Delivery Powered by Context and User Intelligence · IEEE Trans. Multim. 2015 A Joint Online Transcoding and Delivery Approach for Dynamic Adaptive Streaming · IEEE Trans. Multim. 2015 Joint online transcoding and geo-distributed delivery for dynamic adaptive streaming · INFOCOM 2014 |
Data mining › anomaly detection
dense block detection |
0.5 | 2 | 2016 | Spotting Suspicious Behaviors in Multimodal Data: A General Metric and Algorithms · IEEE Trans. Knowl. Data Eng. 2016 A General Suspiciousness Metric for Dense Blocks in Multimodal Data · ICDM 2015 |
Mathematical optimization
submodular optimization |
0.4 | 1 | 2020 | Learning to Select Base Classes for Few-Shot Classification · CVPR 2020 |
Information retrieval
image retrieval |
0.4 | 3 | 2014 | Social-Sensed Image Search · ACM Trans. Inf. Syst. 2014 Social Embedding Image Distance Learning · ACM Multimedia 2014 Auto-cut for web images · ACM Multimedia 2009 |
Content delivery and video streaming › video coding
video transcoding |
0.4 | 2 | 2015 | A Joint Online Transcoding and Delivery Approach for Dynamic Adaptive Streaming · IEEE Trans. Multim. 2015 Joint online transcoding and geo-distributed delivery for dynamic adaptive streaming · INFOCOM 2014 |
Data mining
pattern mining |
0.4 | 2 | 2015 | A General Suspiciousness Metric for Dense Blocks in Multimodal Data · ICDM 2015 Cascading outbreak prediction in networks: a data-driven approach · KDD 2013 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.3 | 1 | 2018 | Scalable Optimization for Embedding Highly-Dynamic and Recency-Sensitive Data · KDD 2018 |
Web and social media mining › social network analysis
trust prediction |
0.3 | 1 | 2018 | Power-law Distribution Aware Trust Prediction · IJCAI 2018 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.3 | 1 | 2017 | Treatment Effect Estimation with Data-Driven Variable Decomposition · AAAI 2017 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
0.3 | 1 | 2017 | Estimating Treatment Effect in the Wild via Differentiated Confounder Balancing · KDD 2017 |
Machine learning › Graph learning
network embedding |
0.3 | 1 | 2017 | Community Preserving Network Embedding · AAAI 2017 |
Data mining › structured data mining › graph mining
community detection |
0.3 | 1 | 2017 | Community Preserving Network Embedding · AAAI 2017 |
Data mining › clustering › graph clustering
modularity-based clustering |
0.3 | 1 | 2017 | Community Preserving Network Embedding · AAAI 2017 |
Data mining › statistical analysis
survival analysis |
0.3 | 1 | 2017 | A Temporally Heterogeneous Survival Framework with Application to Social Behavior Dynamics · KDD 2017 |
Web and social media mining › social influence analysis
social influence prediction |
0.2 | 2 | 2011 | Who should share what?: item-level social influence prediction for users and posts ranking · SIGIR 2011 Item-Level Social Influence Prediction with Probabilistic Hybrid Factor Matrix Factorization · AAAI 2011 |
Web and social media mining
social media marketing |
0.2 | 1 | 2016 | Steering Social Media Promotions with Effective Strategies · ICDM 2016 |
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis › tensor factorization
tensor mining |
0.2 | 1 | 2016 | Spotting Suspicious Behaviors in Multimodal Data: A General Metric and Algorithms · IEEE Trans. Knowl. Data Eng. 2016 |
Content delivery and video streaming › video sharing
social video sharing |
0.2 | 1 | 2016 | Dispersing Instant Social Video Service Across Multiple Clouds · IEEE Trans. Parallel Distributed Syst. 2016 |
Cloud and datacenter computing › cloud deployment
multi-cloud deployment |
0.2 | 1 | 2016 | Dispersing Instant Social Video Service Across Multiple Clouds · IEEE Trans. Parallel Distributed Syst. 2016 |
Methods — techniques the papers use, named apart from their topics
variable decomposition · 1.4propensity score · 1.4measurement study · 0.9submodular optimization · 0.9similarity ratio · 0.9nested segment tree · 0.7diffused stochastic gradient descent · 0.7trace-driven measurement · 0.6zeroth-order optimization · 0.6singular value decomposition · 0.6tensor decomposition · 0.6heuristic algorithm · 0.5graph partitioning · 0.5optimization · 0.5data-driven monitoring · 0.4IP address analysis · 0.4non-negative matrix factorization · 0.4hybrid random walk · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Novel Grasping Robot Control Method Using Motion Execution BCI Combining Knowledge ReasoningabstractRecently, with the growing number of disabled people, brain-controlled technology offers a novel way to help patients restore their daily abilities. However, the conventional brain-controlled system based on the motion related task lacks intelligence in real-world environments. To address above problem, this study proposed a share-controlled system combining a precise hand movement (PHM)-based brain computer interface (BCI) system and knowledge-driven reasoning method. Six types of precise hand movements were selected to design novel motion execution paradigm for BCI system. A feature intermediate fusion convolutional neural network was employed to accurately decode electroencephalogram. Furthermore, a shared control grasping technology based on knowledge-based reasoning combined PHM-based BCI system was designed for grasping robot, which enhancing the system's intelligence and versatility in selecting objects. To verify the improvement of proposed method, experiments were conducted with 15 healthy subjects and 2 patients. The proposed method achieved an average accuracy of 82.80 ± 6.08%, with the highest accuracy reaching 94.27%. All the experimental results demonstrate the effectiveness of the proposed shared control method. Jinli Liu, Shiqiang Yang |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Group-based social diffusion in recommendation
Xumin Chen, Ruobing Xie, Zhijie Qiu, Peng Cui 0001, Ziwei Zhang 0001, Shiqiang Yang, Bo Zhang 0056, Leyu Lin |
World Wide Web (WWW) | 7 |
| 2022 | Adversarial Eigen Attack on BlackBox ModelsabstractBlack-box adversarial attack has aroused much research attention for its difficulty on nearly no available information of the attacked model and the additional constraint on the query budget. A common way to improve attack efficiency is to transfer the gradient information of a white-box substitute model trained on an extra dataset. In this paper, we deal with a more practical setting where a pre-trained white-box model with network parameters is provided without extra training data. To solve the model mismatch problem between the white-box and black-box models, we propose a novel algorithm EigenBA by systematically integrating gradient-based white-box method and zeroth-order optimization in black-box methods. We theoretically show the optimal directions of perturbations for each step are closely related to the right singular vectors of the Jacobian matrix of the pretrained white-box model. Extensive experiments on ImageNet, CIFAR-10 and WebVision show that EigenBA can consistently and significantly outperform state-of-the-art baselines in terms of success rate and attack efficiency. Linjun Zhou, Peng Cui 0001, Xingxuan Zhang, Yinan Jiang, Shiqiang Yang |
CVPR | 5 |
| 2022 | AIDEDNet: anti-interference and detail enhancement dehazing network for real-world scenes
Fazhi He, Yansong Duan, Shiqiang Yang |
Frontiers Comput. Sci. | 4 |
| 2022 | Data-Driven Variable Decomposition for Treatment Effect EstimationabstractCausal Inference plays an important role in decision making in many fields, such as social marketing, healthcare, and public policy. One fundamental problem in causal inference is the treatment effect estimation in observational studies when variables are confounded. Controlling for confounding effects is generally handled by propensity score. But it treats all observed variables as confounders and ignores the adjustment variables, which have no influence on treatment but are predictive of the outcome. Recently, it has been demonstrated that the adjustment variables are effective in reducing the variance of the estimated treatment effect. However, how to automatically separate the confounders and adjustment variables in observational studies is still an open problem, especially in the scenarios of high dimensional variables, which are common in the big data era. In this paper, we first propose a Data-Driven Variable Decomposition (D$^2$VD) algorithm, which can 1) automatically separate confounders and adjustment variables with a data-driven approach, and 2) simultaneously estimate treatment effect in observational studies with high dimensional variables. Under standard assumptions, we theoretically prove that our D$^2$VD algorithm can unbiased estimate treatment effect and achieve lower variance than traditional propensity score based methods. Moreover, to address the challenges from high-dimensional variables and nonlinear, we extend our D$^2$VD to a non-linear version, namely Nonlinear-D$^2$VD (N-D$^2$VD) algorithm. To validate the effectiveness of our proposed algorithms, we conduct extensive experiments on both synthetic and real-world datasets. The experimental results demonstrate that our D$^2$VD and N-D$^2$VD algorithms can automatically separate the variables precisely, and estimate treatment effect more accurately and with tighter confidence intervals than the state-of-the-art methods. We also demonstrated that the top-ranked features by our algorithm have the best prediction performance on an online advertising dataset. Kun Kuang 0001, Peng Cui 0001, Hao Zou 0001, Bo Li 0064, Jianrong Tao, Fei Wu 0001, Shiqiang Yang |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2020 | Learning to Select Base Classes for Few-Shot ClassificationabstractFew-shot learning has attracted intensive research attention in recent years. Many methods have been proposed to generalize a model learned from provided base classes to novel classes, but no previous work studies how to select base classes, or even whether different base classes will result in different generalization performance of the learned model. In this paper, we utilize a simple yet effective measure, the Similarity Ratio, as an indicator for the generalization performance of a few-shot model. We then formulate the base class selection problem as a submodular optimization problem over Similarity Ratio. We further provide theoretical analysis on the optimization lower bound of different optimization methods, which could be used to identify the most appropriate algorithm for different experimental settings. The extensive experiments on ImageNet, Caltech256 and CUB-200-2011 demonstrate that our proposed method is effective in selecting a better base dataset. Linjun Zhou, Peng Cui 0001, Xu Jia 0012, Shiqiang Yang, Qi Tian 0001 |
CVPR | 4 |
| 2020 | Treatment Effect Estimation via Differentiated Confounder Balancing and RegressionabstractTreatment effect plays an important role on decision making in many fields, such as social marketing, healthcare, and public policy. The key challenge on estimating treatment effect in the wild observational studies is to handle confounding bias induced by imbalance of the confounder distributions between treated and control units. Traditional methods remove confounding bias by re-weighting units with supposedly accurate propensity score estimation under the unconfoundedness assumption. Controlling high-dimensional variables may make the unconfoundedness assumption more plausible, but poses new challenge on accurate propensity score estimation. One strand of recent literature seeks to directly optimize weights to balance confounder distributions, bypassing propensity score estimation. But existing balancing methods fail to do selection and differentiation among the pool of a large number of potential confounders, leading to possible underperformance in many high-dimensional settings. In this article, we propose a data-driven Differentiated Confounder Balancing (DCB) algorithm to jointly select confounders, differentiate weights of confounders and balance confounder distributions for treatment effect estimation in the wild high-dimensional settings. Besides, under some settings with heavy confounding bias, in order to further reduce the bias and variance of estimated treatment effect, we propose a Regression Adjusted Differentiated Confounder Balancing (RA-DCB) algorithm based on our DCB algorithm by incorporating outcome regression adjustment. The synergistic learning algorithms we proposed are more capable of reducing the confounding bias in many observational studies. To validate the effectiveness of our DCB and RA-DCB algorithms, we conduct extensive experiments on both synthetic and real-world datasets. The experimental results clearly demonstrate that our algorithms outperform the state-of-the-art methods. By incorporating regression adjustment, our RA-DCB algorithm achieves more precise estimation on treatment effect than DCB algorithm, especially under the settings with heavy confounding bias. Moreover, we show that the top features ranked by our algorithm generate accurate prediction of online advertising effect. Kun Kuang 0001, Peng Cui 0001, Bo Li 0064, Meng Jiang 0001, Yashen Wang, Fei Wu 0001, Shiqiang Yang |
ACM Trans. Knowl. Discov. Data | 7 |
| 2019 | Learning to Learn Image Classifiers With Visual AnalogyabstractHumans are far better learners who can learn a new concept very fast with only a few samples compared with machines. The plausible mystery making the difference is two fundamental learning mechanisms: learning to learn and learning by analogy. In this paper, we attempt to investigate a new human-like learning method by organically combining these two mechanisms. In particular, we study how to generalize the classification parameters from previously learned concepts to a new concept. we first propose a novel Visual Analogy Graph Embedded Regression (VAGER) model to jointly learn a low-dimensional embedding space and a linear mapping function from the embedding space to classification parameters for base classes. We then propose an out-of-sample embedding method to learn the embedding of a new class represented by a few samples through its visual analogy with base classes and derive the classification parameters for the new class. We conduct extensive experiments on ImageNet dataset and the results show that our method could consistently and significantly outperform state-of-the-art baselines. Linjun Zhou, Peng Cui 0001, Shiqiang Yang, Wenwu Zhu 0001, Qi Tian 0001 |
CVPR | 3 |
| 2019 | Towards QoS-Aware Cloud Live Transcoding: A Deep Reinforcement Learning ApproachabstractVideo transcoding is widely adopted in live streaming services to bridge the format and resolution gap between content producers and consumers (i.e., broadcasters and viewers). Meanwhile, the cloud has been recognized as one of the most reliable and cost-effective ways for video transcoding. However, due to the dynamic and uncertainty of the transcoding workloads in live streaming, it is very challenging for cloud service providers to provision computing resources and schedule transcoding tasks while guaranteeing the Service Level Agreement (SLA). To this end, we propose a joint resource provisioning and task scheduling approach for transcoding live streams in the cloud. We adopt Deep Reinforcement Learning (DRL) to train a neural network model for resource provisioning under dynamic workloads. Moreover, we design a QoS-aware task scheduling algorithm that maps transcoding tasks to Virtual Machines (VMs) by considering the real-time QoS requirement. We evaluate our approach with trace-driven experiments and the results demonstrate that our approach outperforms heuristic baselines by up to 89% improvements on average QoS with 4% extra resource overhead at most. Zhengyuan Pang, Lifeng Sun, Tianchi Huang, Zhi Wang 0001, Shiqiang Yang |
ICME | 5 |
| 2018 | Power-law Distribution Aware Trust PredictionabstractTrust prediction, aiming to predict the trust relations between users in a social network, is a key to helping users discover the reliable information. Many trust prediction methods are proposed based on the low-rank assumption of a trust network. However, one typical property of the trust network is that the trust relations follow the power-law distribution, i.e., few users are trusted by many other users, while most tail users have few trustors. Due to these tail users, the fundamental low-rank assumption made by existing methods is seriously violated and becomes unrealistic. In this paper, we propose a simple yet effective method to address the problem of the violated low-rank assumption. Instead of discovering the low-rank component of the trust network alone, we learn a sparse component of the trust network to describe the tail users simultaneously. With both of the learned low-rank and sparse components, the trust relations in the whole network can be better captured. Moreover, the transitive closure structure of the trust relations is also integrated into our model. We then derive an effective iterative algorithm to infer the parameters of our model, along with the proof of correctness. Extensive experimental results on real-world trust networks demonstrate the superior performance of our proposed method over the state-of-the-arts. Xiao Wang 0017, Ziwei Zhang 0001, Jing Wang 0023, Peng Cui 0001, Shiqiang Yang |
IJCAI | 5 |
| 2018 | Scalable Optimization for Embedding Highly-Dynamic and Recency-Sensitive DataabstractA dataset which is highly-dynamic and recency-sensitive means new data are generated in high volumes with a fast speed and of higher priority for the subsequent applications. Embedding technique is a popular research topic in recent years which aims to represent any data into low-dimensional vector space, which is widely used in different data types and have multiple applications. Generating embeddings on such data in a high-speed way is a challenging problem to consider the high dynamics and the recency sensitiveness together with both effectiveness and efficient. Popular embedding methods are usually time-consuming. As well as the common optimization methods are limited since it may not have enough time to converge or deal with recency-sensitive sample weights. This problem is still an open problem. In this paper, we propose a novel optimization method named Diffused Stochastic Gradient Descent for such highly-dynamic and recency-sensitive data. The notion of our idea is to assign recency-sensitive weights to different samples, and select samples according to their weights in calculating gradients. And after updating the embedding of the selected sample, the related samples are also updated in a diffusion strategy. We propose a Nested Segment Tree to improve the recency-sensitive weight method and the diffusion strategy into a complexity no slower than the iteration step in practice. We also theoretically prove the convergence rate of D-SGD for independent data samples, and empirically prove the efficacy of D-SGD in large-scale real datasets. Xumin Chen, Peng Cui 0001, Lingling Yi, Shiqiang Yang |
KDD | 4 |
| 2018 | Accelerating HEVC Encoding Using Early-SplitabstractThe increase in coding efficiency and complexity of high efficiency video coding (HEVC) over H.264 is due to, among other factors, the time needed to find the optimal partition structure among the more flexible encoding modes for the coding units (CUs) and prediction units (PUs). Although many classification-based algorithms have been proposed to expedite the partition decision, the features that can be acquired from current HEVC encoding order are not sufficient to minimize the loss in coding efficiency. In this letter, we proposed an early-split (ES) order for HEVC CU-level encoding, where the encoder checks the split mode before the nonsquare PU partition modes and utilizes the encoding output of the subCUs to expedite subsequent encoding. Experiments show that the proposed algorithm can save 48% of encoding time on average with only about 0.8% loss in coding performance. Minhao Tang, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen, Shiqiang Yang |
IEEE Signal Process. Lett. | 6 |
| 2018 | Effective Promotional Strategies Selection in Social Media: A Data-Driven ApproachabstractNowdays, many companies, organizations and individuals are using the function of sharing or retweeting information to promote their products, policies, and ideas on social media. While a growing body of research has focused on identifying the promoters from millions of users, the promoters themselves are seeking to know which strategy can improve promotional effectiveness, which is rarely studied in the literature. In this work, we investigate an open problem of effective promotional strategy selection via causal analysis which is challenging in identifying and quantifying promotional strategies as well as the selection bias when estimating the causal effect of promotional strategies from observational data. We study the promotional strategies not only on the content level (what to promote) but also on the context level (when and how to promote). To alleviate the issue of selection bias in observational studies, we propose a data-driven approach that is a Propensity Score Matching (PSM) based method, which helps to evaluate the causal effect of each promotional strategy and discover the set of effective strategies to predict the promotional effectiveness (i.e., the number of users infected by the promotion). We evaluate our proposed method on a real social dataset including 194 million users and 5 million promoted messages. Experimental results show that (1) the top-ranked strategies by our PSM based method significantly and consistently outperform the correlation based feature selection methods in predicting promotional effectiveness; (2) we conclude our observations from the real data with three interpretable and practical ideas for steering social media promotion. Kun Kuang 0001, Meng Jiang 0001, Peng Cui 0001, Hengliang Luo, Shiqiang Yang |
IEEE Trans. Big Data | 5 |
| 2017 | Treatment Effect Estimation with Data-Driven Variable DecompositionabstractOne fundamental problem in causal inference is the treatment effect estimation in observational studies when variables are confounded. Control for confounding effect is generally handled by propensity score. But it treats all observed variables as confounders and ignores the adjustment variables, which have no influence on treatment but are predictive of the outcome. Recently, it has been demonstrated that the adjustment variables are effective in reducing the variance of the estimated treatment effect. However, how to automatically separate the confounders and adjustment variables in observational studies is still an open problem, especially in the scenarios of high dimensional variables, which are common in big data era. In this paper, we propose a Data-Driven Variable Decomposition (D$^2$VD) algorithm, which can 1) automatically separate confounders and adjustment variables with a data driven approach, and 2) simultaneously estimate treatment effect in observational studies with high dimensional variables. Under standard assumptions, we show experimentally that the proposed D$^2$VD algorithm can automatically separate the variables precisely, and estimate treatment effect more accurately and with tighter confidence intervals than the state-of-the-art methods on both synthetic data and real online advertising dataset. Kun Kuang 0001, Peng Cui 0001, Bo Li 0064, Meng Jiang 0001, Shiqiang Yang, Fei Wang 0001 |
AAAI | 5 |
| 2017 | Community Preserving Network EmbeddingabstractNetwork embedding, aiming to learn the low-dimensional representations of nodes in networks, is of paramount importance in many real applications. One basic requirement of network embedding is to preserve the structure and inherent properties of the networks. While previous network embedding methods primarily preserve the microscopic structure, such as the first- and second-order proximities of nodes, the mesoscopic community structure, which is one of the most prominent feature of networks, is largely ignored. In this paper, we propose a novel Modularized Nonnegative Matrix Factorization (M-NMF) model to incorporate the community structure into network embedding. We exploit the consensus relationship between the representations of nodes and community structure, and then jointly optimize NMF based representation learning model and modularity based community detection model in a unified framework, which enables the learned representations of nodes to preserve both of the microscopic and community structures. We also provide efficient updating rules to infer the parameters of our model, together with the correctness and convergence guarantees. Extensive experimental results on a variety of real-world networks show the superior performance of the proposed method over the state-of-the-arts. Xiao Wang 0017, Peng Cui 0001, Jing Wang 0023, Jian Pei 0001, Wenwu Zhu 0001, Shiqiang Yang |
AAAI | 6 |
| 2017 | HEVC-based motion compensated joint temporal-spatial video denoisingabstractA novel HEVC-based efficient video denoising algorithm is proposed in this paper. It uses a spatial Gaussian filter for the chrominance components and then utilizes the HEVC motion estimation process to find the best temporal correspondence for low-pass filtering. Other HEVC tools such as quantization, the interpolation and the in-loop filters are also used. Experiments implementing the proposed algorithm in the open-source HEVC encoder ×265 showed a good denoising performance with a much lower computing complexity than the competitors. The performance was comparable to those highly sophisticated algorithms such as the VBM4D, which is 200 times slower. The proposed algorithm can be easily integrated into the real-world video processing systems due to its compatibility with the HEVC standard. Minhao Tang, Yuxing Han 0001, Jiangtao Wen, Shiqiang Yang |
ICASSP | 4 |
| 2017 | CP-operated dash caching via reinforcement learningabstractIn recent years, Dynamic Adaptive Streaming over HTTP (DASH) has gained momentum as an effective solution for delivering videos on the Internet. This trend is further driven by the deployment of existing HTTP cache infrastructures in DASH systems to reduce the traffic load as well as to serve clients better. However, deploying conventional cache servers in DASH systems still suffers from low cache hit ratio and bitrate oscillations, which makes it challenging for content providers (CPs) to balance the user-perceived quality-of-experience (QoE) and the operating cost in cache-enabled DASH systems. To address this challenge, we propose a CP-operated DASH caching framework to provide good user QoE with low cost. In particular, we first formulate the caching decision problem as a stochastic optimization problem over a finite time horizon. The objective of this problem is to maximize a weighted sum of the user QoE and the operating cost, termed as the utility. Then we design a reinforcement learning based online algorithm which can obtain approximately optimal solution of this problem. Through extensive trace-driven experiments, we show that our approach not only achieves 40% average improvement of the overall utility compared to baseline approaches, but also adapts to the server load. Zhengyuan Pang, Lifeng Sun, Zhi Wang 0001, Wen Hu 0003, Shiqiang Yang |
ICME | 5 |
| 2017 | Optimized video coding for omnidirectional videosabstractThe ever widening application of virtual reality requires the ultra high resolution omnidirectional videos (OVs) to be transmitted over the wired and wireless Internet at low cost (i.e. bitrate). Various solutions have been proposed to intelligently reduce the bitrate, e.g. adapting the spatial resolution of the video for different directions of the panorama with regard to current direction that the viewer is looking at and the distribution of the probability of each direction to be watched according to the video content. Due to various reasons, spatial resolution adaptation may often cause perceivable quality degradation to user experience. In this paper, we proposed two adaptive encoding techniques to reduce the bitrate of OVs after compression. The first is a content adaptive temporal resolution adaptation scheme for OVs using cube map projection. The second is a quantization and rate-distortion optimization scheme for equirectangular projection. Experiments implementing the proposed algorithms in the open source HEVC encoder x265 show that the proposed algorithms can save 14.69% and 13.01% in BD-rate on average for the cube map and equirectangular projected OVs respectively, while no degradation to the visual experience was reported in subjective tests. The proposed algorithms are compatible with and therefore can collaborate with the currently used adaptive spatial resolution scheme for further bitrate reduction. Minhao Tang, Jiangtao Wen, Shiqiang Yang |
ICME | 4 |
| 2017 | Estimating Treatment Effect in the Wild via Differentiated Confounder BalancingabstractEstimating treatment effect plays an important role on decision making in many fields, such as social marketing, healthcare, and public policy. The key challenge on estimating treatment effect in the wild observational studies is to handle confounding bias induced by imbalance of the confounder distributions between treated and control units. Traditional methods remove confounding bias by re-weighting units with supposedly accurate propensity score estimation under the unconfoundedness assumption. Controlling high-dimensional variables may make the unconfoundedness assumption more plausible, but poses new challenge on accurate propensity score estimation. One strand of recent literature seeks to directly optimize weights to balance confounder distributions, bypassing propensity score estimation. But existing balancing methods fail to do selection and differentiation among the pool of a large number of potential confounders, leading to possible underperformance in many high dimensional settings. In this paper, we propose a data-driven Differentiated Confounder Balancing (DCB) algorithm to jointly select confounders, differentiate weights of confounders and balance confounder distributions for treatment effect estimation in the wild high dimensional settings. The synergistic learning algorithm we proposed is more capable of reducing the confounding bias in many observational studies. To validate the effectiveness of our DCB algorithm, we conduct extensive experiments on both synthetic and real datasets. The experimental results clearly demonstrate that our DCB algorithm outperforms the state-of-the-art methods. We further show that the top features ranked by our algorithm generate accurate prediction of online advertising effect. Kun Kuang 0001, Peng Cui 0001, Bo Li 0064, Meng Jiang 0001, Shiqiang Yang |
KDD | 5 |
| 2017 | A Temporally Heterogeneous Survival Framework with Application to Social Behavior DynamicsabstractSocial behavior dynamics is one of the central building blocks in understanding and modeling complex social dynamic phenomena, such as information spreading, opinion formation, and social mobilization. While a wide range of models for social behavior dynamics have been proposed in recent years, the essential ingredients and the minimum model for social behavior dynamics is still largely unanswered. Here, we find that human interaction behavior dynamics exhibit rich complexities over the response time dimension and natural time dimension by exploring a large scale social communication dataset. To tackle this challenge, we develop a temporal Heterogeneous Survival framework where the regularities in response time dimension and natural time dimension can be organically integrated. We apply our model in two online social communication datasets. Our model can successfully regenerate the interaction patterns in the social communication datasets, and the results demonstrate that the proposed method can significantly outperform other state-of-the-art baselines. Meanwhile, the learnt parameters and discovered statistical regularities can lead to multiple potential applications. Linyun Yu, Peng Cui 0001, Chaoming Song, Tianyang Zhang 0001, Shiqiang Yang |
KDD | 5 |
| 2017 | Understanding Performance of Edge Prefetching
Zhengyuan Pang, Lifeng Sun, Zhi Wang 0001, Yuan Xie 0005, Shiqiang Yang |
MMM (1) | 5 |
| 2017 | CELoF: WiFi Dwell Time Estimation in Free Environment
Peng Wang 0012, Haitian Pang, Lifeng Sun, Shiqiang Yang |
MMM (1) | 5 |
| 2017 | A Dataset for Exploring User Behaviors in VR Spherical Video StreamingabstractWith Virtual Reality (VR) devices and content getting increasingly popular, understanding user behaviors in virtual environment is important for not only VR product design but also user experience improvement. In VR applications, the head movement is one of the most important user behaviors, which can reflect a user's visual attention, preference, and even unique motion pattern. However, to the best of our knowledge, no dataset containing this information is publicly available. In this paper, we present a head tracking dataset composed of 48 users (24 males and 24 females) watching 18 sphere videos from 5 categories. We carefully record how users watch the videos, how their heads move in each session, what directions they focus, and what content they can remember after each session. Based on this dataset, we show that people share certain common patterns in VR spherical video streaming, which are different from conventional video streaming. We believe the dataset can serve good resource for exploring user behavior patterns in VR applications. Chenglei Wu, Zhihao Tan, Zhi Wang 0001, Shiqiang Yang |
MMSys | 4 |
| 2017 | Training-free indexing refinement for visual media via multi-semantics
Peng Wang 0012, Lifeng Sun, Shiqiang Yang, Alan F. Smeaton |
Neurocomputing | 3 |
| 2017 | Uncovering and predicting the dynamic process of information cascades with survival model
Linyun Yu, Peng Cui 0001, Fei Wang 0001, Chaoming Song, Shiqiang Yang |
Knowl. Inf. Syst. | 5 |
| 2017 | Cross-Scale Cost Aggregation for Stereo MatchingabstractThis paper proposes a generic framework that enables a multiscale interaction in the cost aggregation step of stereo matching algorithms. Inspired by the formulation of image filters, we first reformulate cost aggregation from a weighted least-squares (WLS) optimization perspective and show that different cost aggregation methods essentially differ in the choices of similarity kernels. Our key motivation is that while the human stereo vision system processes information at both coarse and fine scales interactively for the correspondence search, state-of-the-art approaches aggregate costs at the finest scale of the input stereo images only, ignoring inter-consistency across multiple scales. This motivation leads us to introduce an inter-scale regularizer into the WLS optimization objective to enforce the consistency of the cost volume among the neighboring scales. The new optimization objective with the inter-scale regularization is convex, and thus, it is easily and analytically solved. Minimizing this new objective leads to the proposed framework. Since the regularization term is independent of the similarity kernel, various cost aggregation approaches, including discrete and continuous parameterization methods, can be easily integrated into the proposed framework. We show that the cross-scale framework is important as it effectively and efficiently expands state-of-the-art cost aggregation methods and leads to significant improvements, when evaluated on Middlebury, Middlebury Third, KITTI, and New Tsukuba data sets. Kang Zhang 0004, Yuqiang Fang, Dongbo Min, Lifeng Sun, Shiqiang Yang, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | comeNgo: A Dynamic Model for Social Group EvolutionabstractHow do social groups, such as Facebook groups and Wechat groups, dynamically evolve over time? How do people join the social groups, uniformly or with burst? What is the pattern of people quitting from groups? Is there a simple universal model to depict the come-and-go patterns of various groups? In this article, we examine temporal evolution patterns of more than 100 thousands social groups with more than 10 million users. We surprisingly find that the evolution patterns of real social groups goes far beyond the classic dynamic models like SI and SIR. For example, we observe both diffusion and non-diffusion mechanism in the group joining process, and power-law decay in group quitting process, rather than exponential decay as expected in SIR model. Therefore, we propose a new modelcomeNgo, a concise yet flexible dynamic model for group evolution. Our model has the following advantages: (a) Unification power: it generalizes earlier theoretical models and different joining and quitting mechanisms we find from observation. (b) Succinctness and interpretability: it contains only six parameters with clear physical meanings. (c) Accuracy: it can capture various kinds of group evolution patterns preciously, and the goodness of fit increases by 58% over baseline. (d) Usefulness: it can be used in multiple application scenarios, such as forecasting and pattern discovery. Furthermore, our model can provide insights about different evolution patterns of social groups, and we also find that group structure and its evolution has notable relations with temporal patterns of group evolution. Tianyang Zhang 0001, Peng Cui 0001, Christos Faloutsos, Yunfei Lu, Wenwu Zhu 0001, Shiqiang Yang |
ACM Trans. Knowl. Discov. Data | 7 |
| 2016 | Little Is Much: Bridging Cross-Platform Behaviors through Overlapped CrowdsabstractPeople often use multiple platforms to fulfill their different information needs. With the ultimate goal of serving people intelligently, a fundamental way is to get comprehensive understanding about user needs. How to organically integrate and bridge cross-platform information in a human-centric way is important. Existing transfer learning assumes either fully-overlapped or non-overlapped among the users. However, the real case is the users of different platforms are partially overlapped. The number of overlapped users is often small and the explicitly known overlapped users is even less due to the lacking of unified ID for a user across different platforms. In this paper, we propose a novel semi-supervised transfer learning method to address the problem of cross-platform behavior prediction, called XPTrans. To alleviate the sparsity issue, it fully exploits the small number of overlapped crowds to optimally bridge a user's behaviors in different platforms. Extensive experiments across two real social networks show that XPTrans significantly outperforms the state-of-the-art. We demonstrate that by fully exploiting 26% overlapped users, XPTrans can predict the behaviors of non-overlapped users with the same accuracy as overlapped users, which means the small overlapped crowds can successfully bridge the information across different platforms. Meng Jiang 0001, Peng Cui 0001, Nicholas Jing Yuan, Xing Xie 0001, Shiqiang Yang |
AAAI | 5 |
| 2016 | Crowdsourced Live Streaming over Aggregated Edge NetworksabstractRecent years have witnessed a dramatic increase of user-generated video services. In such user-generated video services, crowdsourced live streaming (e.g., Periscope, Twitch) has significantly challenged today's content delivery infrastructure: today's edge networks (e.g., 4G, Wi-Fi) have limited uplink capacity support, making high-bitrate live streaming over such links fundamentally impossible. In this paper, we propose to let broadcasters (i.e., users who generate the video) upload crowdsourced video streams using aggregated network resources from multiple edge networks. There are several challenges in the proposal: First, how to design a framework that aggregates bandwidth from multiple edge networks? Second, how to make this framework transparent to today's crowdsourced live stream- ing services? Third, how to maximize the streaming quality for the whole system? We design a multi-objective and deployable bandwidth aggregation system BASS to address these challenges: (1) We propose an aggregation framework transparent to today's crowdsourced live streaming services, using an edge proxy box and aggregation cloud paradigm; (2) We dynamically allocate geo- distributed cloud aggregation servers to enable MPTCP (i.e., multi- path TCP), according to location and network characteristics of both broadcasters and the original streaming servers; (3) We maximize the overall performance gain for the whole system, by matching streams with the best aggregation paths. Chenglei Wu, Zhi Wang 0001, Jiangchuan Liu, Shiqiang Yang |
GLOBECOM | 4 |
| 2016 | Steering Social Media Promotions with Effective StrategiesabstractOn social media platforms, companies, organizations and individuals are using the function of sharing or retweeting information to promote their products, policies, and ideas. While a growing body of research has focused on identifying the promoters from millions of users, the promoters themselves are seeking to know what strategies can improve promotional effectiveness, which is rarely studied in literature. In this work, we study a new problem of promotional strategy effect estimation which is challenging in identifying and quantifying promotional strategies, as well as estimating effectiveness of promotional strategies with selection bias in observational data. Here we study a series of strategies on both context and content levels. To alleviate the selection bias issue, we propose a method based on Propensity Score Matching (PSM) to evaluate the effect of each promotional strategy. Our data study provides three interpretable and insightful ideas on steering social media promotions, including (1) three significant and stable strategies, (2) a critical trade-off, and (3) different concerns for promoters of different popularity. These results provided comprehensive suggestions to the practitioners to steer social media promotions with effective strategies. Kun Kuang 0001, Meng Jiang 0001, Peng Cui 0001, Shiqiang Yang |
ICDM | 4 |
| 2016 | Learning-based quality assessment of retargeted stereoscopic imagesabstractStereoscopic image retargeting techniques aim to flexibly display 3D images with different aspect ratios and simultaneously preserve salient regions and comfortable depth perception. Various stereoscopic image retargeting techniques have been proposed recently. However, there is still no effective objective metric for visual quality assessment of retargeted stereoscopic images. In this paper, we build a stereoscopic image retargeting database and propose a learning-based objective method to evaluate the stereoscopic image retargeting quality. The perception quality of the database are evaluated by subjects. We extract new features of quality assessment and fuse them to assess stereoscopic image retargeting quality using neural network. Experiments conducted with above-mentioned database confirm the effectiveness of the proposed method. The results show the good consistency between the objective assessments and subjective rankings. Lifeng Sun, Shiqiang Yang |
ICME | 3 |
| 2016 | Come-and-Go Patterns of Group Evolution: A Dynamic ModelabstractHow do social groups, such as Facebook groups and Wechat groups, dynamically evolve over time? How do people join the social groups, uniformly or with burst? What is the pattern of people quitting from groups? Is there a simple universal model to depict the come-and-go patterns of various groups? Tianyang Zhang 0001, Peng Cui 0001, Christos Faloutsos, Yunfei Lu, Wenwu Zhu 0001, Shiqiang Yang |
KDD | 7 |
| 2016 | Towards Training-Free Refinement for Semantic Indexing of Visual Media
Peng Wang 0012, Lifeng Sun, Shiqiang Yang, Alan F. Smeaton |
MMM (1) | 3 |
| 2016 | What are the Limits to Time Series Based Recognition of Semantic Concepts?
Peng Wang 0012, Lifeng Sun, Shiqiang Yang, Alan F. Smeaton |
MMM (2) | 3 |
| 2016 | Characterizing everyday activities from visual lifelogs based on enhancing concept representation
Peng Wang 0012, Lifeng Sun, Shiqiang Yang, Alan F. Smeaton, Cathal Gurrin |
Comput. Vis. Image Underst. | 3 |
| 2016 | Inferring lockstep behavior from connectivity pattern in large graphs
Meng Jiang 0001, Peng Cui 0001, Alex Beutel, Christos Faloutsos, Shiqiang Yang |
Knowl. Inf. Syst. | 5 |
| 2016 | Catching Synchronized Behaviors in Large Networks: A Graph Mining ApproachabstractGiven a directed graph of millions of nodes, how can we automatically spot anomalous, suspicious nodes judging only from their connectivity patterns? Suspicious graph patterns show up in many applications, from Twitter users who buy fake followers, manipulating the social network, to botnet members performing distributed denial of service attacks, disturbing the network traffic graph. We propose a fast and effective method, C atch S ync , which exploits two of the tell-tale signs left in graphs by fraudsters: (a) synchronized behavior: suspicious nodes have extremely similar behavior patterns because they are often required to perform some task together (such as follow the same user); and (b) rare behavior: their connectivity patterns are very different from the majority. We introduce novel measures to quantify both concepts (“synchronicity” and “normality”) and we propose a parameter-free algorithm that works on the resulting synchronicity-normality plots. Thanks to careful design, C atch S ync has the following desirable properties: (a) it is scalable to large datasets, being linear in the graph size; (b) it is parameter free ; and (c) it is side-information-oblivious : it can operate using only the topology, without needing labeled data, nor timing information, and the like., while still capable of using side information if available. We applied C atch S ync on three large, real datasets, 1-billion-edge Twitter social graph, 3-billion-edge, and 12-billion-edge Tencent Weibo social graphs, and several synthetic ones; C atch S ync consistently outperforms existing competitors, both in detection accuracy by 36% on Twitter and 20% on Tencent Weibo, as well as in speed. Meng Jiang 0001, Peng Cui 0001, Alex Beutel, Christos Faloutsos, Shiqiang Yang |
ACM Trans. Knowl. Discov. Data | 5 |
| 2016 | Spotting Suspicious Behaviors in Multimodal Data: A General Metric and AlgorithmsabstractMany commercial products and academic research activities are embracing behavior analysis as a technique for improving detection of attacks of many sorts-from retweet boosting, hashtag hijacking to link advertising. Traditional approaches focus on detecting dense blocks in the adjacency matrix of graph data, and recently, the tensors of multimodal data. No method gives a principled way to score the suspiciousness of dense blocks with different numbers of modes and rank them to draw human attention accordingly. In this paper, we first give a list of axioms that any metric of suspiciousness should satisfy; we propose an intuitive, principled metric that satisfies the axioms, and is fast to compute; moreover, we propose CrossSpot, an algorithm to spot dense blocks that are worth inspecting, typically indicating fraud or some other noteworthy deviation from the usual, and sort them in the order of importance (“suspiciousness”). Finally, we apply CrossSpot to the real data, where it improves the F1 score over previous techniques by 68 percent and finds suspicious behavioral patterns in social datasets spanning 0.3 billion posts. Meng Jiang 0001, Alex Beutel, Peng Cui 0001, Bryan Hooi, Shiqiang Yang, Christos Faloutsos |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2016 | Dispersing Instant Social Video Service Across Multiple CloudsabstractInstant social video sharing which combines the online social network and user-generated short video streaming services, has become popular in today’s Internet. Cloud-based hosting of such instant social video contents has become a norm to serve the increasing users with user-generated contents. A fundamental problem of cloud-based social video sharing service is that users are located globally, who cannot be served with good service quality with a single cloud provider. In this paper, we investigate the feasibility of dispersing instant social video contents to multiple cloud providers. The challenge is that inter-cloud socialpropagationis indispensable with such multi-cloud social video hosting, yet such inter-cloud traffic incurs substantial operational cost. We analyze and formulate the multi-cloud hosting of an instant social video system as an optimization problem. We conduct large-scale measurement studies to show the characteristics of instant social video deployment, and demonstrate the trade-off between satisfying users with their ideal cloud providers, and reducing the inter-cloud data propagation. Our measurement insights of the social propagation allow us to propose a heuristic algorithm with acceptable complexity to solve the optimization problem, by partitioning a propagation-weighted social graph in two phases: a preference-aware initial cloud provider selection and a propagation-aware re-hosting. Our simulation experiments driven by real-world social network traces show the superiority of our design. Zhi Wang 0001, Baochun Li, Lifeng Sun, Wenwu Zhu 0001, Shiqiang Yang |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2015 | A General Suspiciousness Metric for Dense Blocks in Multimodal DataabstractWhich seems more suspicious: 5,000 tweets from 200 users on 5 IP addresses, or 10,000 tweets from 500 users on 500 IP addresses but all with the same trending topic and all in 10 minutes? The literature has many methods that try to find dense blocks in matrices, and, recently, tensors, but no method gives a principled way to score the suspiciouness of dense blocks with different numbers of modes and rank them to draw human attention accordingly. Dense blocks are worth inspecting, typically indicating fraud, emerging trends, or some other noteworthy deviation from the usual. Our main contribution is that we show how to unify these methods and how to give a principled answer to questions like the above. Specifically, (a) we give a list of axioms that any metric of suspicousness should satisfy, (b) we propose an intuitive, principled metric that satisfies the axioms, and is fast to compute, (c) we propose CROSSSPOT, an algorithm to spot dense regions, and sort them in importance ("suspiciousness") order. Finally, we apply CROSSSPOT to real data, where it improves the F1 score over previous techniques by 68% and finds retweet-boosting in a real social dataset spanning 0.3 billion posts. Meng Jiang 0001, Alex Beutel, Peng Cui 0001, Bryan Hooi, Shiqiang Yang, Christos Faloutsos |
ICDM | 5 |
| 2015 | From Micro to Macro: Uncovering and Predicting Information Cascading Process with Behavioral DynamicsabstractCascades are ubiquitous in various network environments. How to predict these cascades is highly nontrivial in several vital applications, such as viral marketing, epidemic prevention and traffic management. Most previous works mainly focus on predicting the final cascade sizes. As cascades are typical dynamic processes, it is always interesting and important to predict the cascade size at any time, or predict the time when a cascade will reach a certain size (e.g. an threshold for outbreak). In this paper, we unify all these tasks into a fundamental problem: cascading process prediction. That is, given the early stage of a cascade, how to predict its cumulative cascade size of any later time? For such a challenging problem, how to understand the micro mechanism that drives and generates the macro phenomena (i.e. cascading process) is essential. Here we introduce behavioral dynamics as the micro mechanism to describe the dynamic process of a node's neighbors getting infected by a cascade after this node getting infected (i.e. one-hop subcascades). Through data-driven analysis, we find out the common principles and patterns lying in behavioral dynamics and propose a novel Networked Weibull Regression model for behavioral dynamics modeling. After that we propose a novel method for predicting cascading processes by effectively aggregating behavioral dynamics, and present a scalable solution to approximate the cascading process with a theoretical guarantee. We extensively evaluate the proposed method on a large scale social network dataset. The results demonstrate that the proposed method can significantly outperform other state-of-the-art baselines in multiple tasks including cascade size prediction, outbreak time prediction and cascading process prediction. Linyun Yu, Peng Cui 0001, Fei Wang 0001, Chaoming Song, Shiqiang Yang |
ICDM | 5 |
| 2015 | Improvement of re-sample template matching for lossless screen content videoabstractScreen Content (SC) video coding becomes more important for screen sharing and screen broadcasting applications. There are many easy to see different characters between screen content video and camera-captured video. We proposed a template matching prediction method for lossless SC intra picture coding with higher compression ratio. The pixels are re-sampled to form the Virtual Largest Coding Unit (VLCU) firstly. About 80% pixels in VLCU can be predicted exactly by template matching with zero error. Then, pixels with non-zero prediction error should be coded with three information, index, position and value. Among these three, position will consume the most bits than the other two. In order to handle this challenge, we propose to apply the similarity of non-zero prediction error pixel positions of neighbor VLCU which can greatly help to improve the compression performance. The VLCU can be divided into sub CU as the same as the division in the standard High Efficiency Video Coding(HEVC) intra coding, and RDO is applied to find the best coding efficiency. Pin Tao, Lixin Feng, Sichao Song 0002, Jiangtao Wen, Shiqiang Yang |
ICME | 5 |
| 2015 | Learning Socially Embedded Visual Representation from ScratchabstractLearning image representation by deep model has recently made remarkable achievements for semantic-oriented applications, such as image classification. However, for user-centric tasks, such as image search and recommendation, simply employing the representation learnt from semantic-oriented tasks may fail to capture user intentions. In this paper, we propose a novel Socially Embedded VIsual Representation Learning (SEVIR) approach, where an Asymmetric Multi-task CNN (amtCNN) model is proposed to embed user intention learning task into semantic learning task. Specifically, to address the sparsity and unreliability problems in social behavioral data, we propose to use user clustering, reliability evaluation, random dropout in output layer in our amtCNN. With its the partially shared network architecture, the learnt representation can capture both semantics and user intentions. Comprehensive experiments are conducted to investigate the effectiveness of our approach in applications of user favoring prediction, personalized image recommendation, and image reranking. Compared to the state-of-the-art image representation techniques, our approach achieves significant improvement in performance. Peng Cui 0001, Wenwu Zhu 0001, Shiqiang Yang |
ACM Multimedia | 4 |
| 2015 | QOEYE: A Data Driven Platform for QoE Visualization and System Performance MonitoringabstractThe stunning increase of video streaming has been a major part of the network flow in the past few years. It is essential for content providers to manage more flow servers to satisfy the demands of users. Therefore, to effectively manage large-scale servers and to identify problems, service node become crucial to guarantee video user experience. The rule-based approach that the traditional service providers use is effective but could hardly be applied to large-scale server clusters, let alone in a joint service resources and the maximization of user experience. Unlike traditional rule-based QoS monitoring platform, we design a framework on a basis of user experience metrics to detect the quality of service. In line with the characteristics of IP addresses, we also design a method to quickly pinpoint the fault location. Lifeng Sun, Wenming Shi, Shiqiang Yang |
ACM Multimedia | 4 |
| 2015 | A retargeting method for stereoscopic 3D videoabstractWe propose a disparity-constrained retargeting method for stereoscopic 3D video, which simultaneously resizes a binocular video to a new aspect ratio and remaps the depth to the perceptual comfort zone. First, we model distortion energies to prevent important video contents from deforming. Then, to maintain depth mapping stability, we model disparity variation energies to constraint the disparity range both in spatial and temporal domains. The last component of our method is a non-uniform, pixel-wise warp to the target resolution based on these energy models. Using this method, we can process the original stereoscopic video to generate new, high-perceptual-quality versions at different display resolutions. For evaluation, we conduct a user study; we also discuss the performance of our method. Lifeng Sun, Shiqiang Yang |
Comput. Vis. Media | 3 |
| 2015 | Social Recommendation with Cross-Domain Transferable KnowledgeabstractRecommender systems can suffer from data sparsity and cold start issues. However, social networks, which enable users to build relationships and create different types of items, present an unprecedented opportunity to alleviate these issues. In this paper, we represent a social network as a star-structured hybrid graph centered on a social domain, which connects with other item domains. With this innovative representation, useful knowledge from an auxiliary domain can be transferred through the social domain to a target domain. Various factors of item transferability, including popularity and behavioral consistency, are determined. We propose a novel Hybrid Random Walk (HRW) method, which incorporates such factors, to select transferable items in auxiliary domains, bridge cross-domain knowledge with the social domain, and accurately predict user-item links in a target domain. Extensive experiments on a real social dataset demonstrate that HRW significantly outperforms existing approaches. Meng Jiang 0001, Peng Cui 0001, Xumin Chen, Fei Wang 0001, Wenwu Zhu 0001, Shiqiang Yang |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2015 | A Joint Online Transcoding and Delivery Approach for Dynamic Adaptive StreamingabstractDynamic adaptive streaming has emerged as a popular approach for video services in today's Internet. To date, the two important components in dynamic adaptive streaming, video transcoding that generates the adaptive bitrates of a video and video delivery that streams the videos to users, have been separately studied, resulting in a huge waste of computation and storage resource due to producing and caching different versions of videos regardless of their demands. We conduct extensive measurement studies of video sharing systems, including an IPTV service which streams regular, professionally made videos and an instant video clip sharing service which provides extremely short user-generated videos, as well as the availability of computation resource in conventional content delivery networks (CDNs). Based on the measurement insights, we propose an online joint transcoding and delivery approach for adaptive video streaming. We formulate optimization problems to enable high streaming quality for the users, and low computation and replication costs for the system. In particular, our strategy connects video transcoding and video delivery based on users' preferences of CDN regions and regional preferences of video versions. We analyze hardness of these problems and design distributed solutions. Extensive trace-driven experiments further demonstrate the superiority of our design. Zhi Wang 0001, Lifeng Sun, Chuan Wu 0001, Wenwu Zhu 0001, Qidong Zhuang, Shiqiang Yang |
IEEE Trans. Multim. | 6 |
| 2015 | CPCDN: Content Delivery Powered by Context and User IntelligenceabstractThere is an unprecedented trend that content providers (CPs) are building their own content delivery networks (CDNs) to provide a variety of content services to their users. By exploiting powerful CP-level information in content distribution, these CP-built CDNs open up a whole new design space and are changing the content delivery landscape. In this paper, we adopt a measurement-based approach to understanding why, how, and how much CP-level intelligences can help content delivery. We first present a measurement study of the CDN built by Tencent, a largest content provider based in China. We observe new characteristics and trends in content delivery which pose great challenges to the conventional content delivery paradigm and motivate the proposal of CPCDN, a CDN powered by CP-aware information. We then reveal the benefits obtained by exploiting two indispensable CP-level intelligences, namely context intelligence and user intelligence, in content delivery. Inspired by the insights learnt from the measurement studies, we systematically explore the design space of CPCDN and present the novel architecture and algorithms to address the new content delivery challenges that have arisen. Our results not only demonstrate the potential of CPCDN in pushing content delivery performance to the next level, but also identify new research problems calling for further investigation. Zhi Wang 0001, Wenwu Zhu 0001, Minghua Chen 0001, Lifeng Sun, Shiqiang Yang |
IEEE Trans. Multim. | 5 |
| 2015 | Enhancing Internet-Scale Video Service Deployment Using Microblog-Based PredictionabstractOnline microblogging has been very popular in today's Internet, where users follow other people they are interested in and exchange information between themselves. Among these exchanges, video links are a representative type on a microblogging site. The impact is fundamental-not only are viewers in a video service directly coming from the microblog sharing and recommendation, but also are the users in the microblogging site representing a promising sample to all the viewers. It is intriguing to study a proactive service deployment for such videos, using the propagation patterns of microblogs. Based on extensive traces from Youku and Tencent Weibo, a popular video sharing site and a favored microblogging system, we explore how video propagation patterns in the microblogging system are correlated with video popularity on the video sharing site. Using influential factors summarized from the measurement studies, we further design a neural network-based learning framework to predict the number of potential viewers and their geographic distribution. We then design proactive video deployment algorithms based on the prediction framework, which not only determines the upload capacities of servers in different regions, but also strategically replicates videos to these regions to serve users. Our PlanetLab-based experiments verify the effectiveness of our design. Zhi Wang 0001, Lifeng Sun, Chuan Wu 0001, Shiqiang Yang |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2014 | Cross-Scale Cost Aggregation for Stereo MatchingabstractHuman beings process stereoscopic correspondence across multiple scales. However, this bio-inspiration is ignored by state-of-the-art cost aggregation methods for dense stereo correspondence. In this paper, a generic cross-scale cost aggregation framework is proposed to allow multi-scale interaction in cost aggregation. We firstly reformulate cost aggregation from a unified optimization perspective and show that different cost aggregation methods essentially differ in the choices of similarity kernels. Then, an inter-scale regularizer is introduced into optimization and solving this new optimization problem leads to the proposed framework. Since the regularization term is independent of the similarity kernel, various cost aggregation methods can be integrated into the proposed general framework. We show that the cross-scale framework is important as it effectively and efficiently expands state-of-the-art cost aggregation methods and leads to significant improvements, when evaluated on Middlebury, KITTI and New Tsukuba datasets. Kang Zhang 0004, Yuqiang Fang, Dongbo Min, Lifeng Sun, Shiqiang Yang, Shuicheng Yan, Qi Tian 0001 |
CVPR | 5 |
| 2014 | Find you from your friends: Graph-based residence location prediction for users in social mediaabstractAs a bridge between social media and physical space, location information will potentially make the internet smarter, and release the real power of social media to address the serious and significant problems in the real world. However, in terms of privacy and security, most of the users are unwilling to make their locations public. To address the problem, an algorithm is necessary to predict the users' residence locations based on the public profiles. We define location propagation probability of users, leverage a semi-supervised learning algorithm, and introduce a novel method of location propagation to predict users' residence locations based on users' social relationships, textual and visual contents and a small amount of known users' residence locations. The experimental results on a large scale real data set in Tencent Weibo demonstrate that our location propagation algorithm outperforms the state-of-the-art approaches in both accuracy and scalability. Dan Xu 0002, Peng Cui 0001, Wenwu Zhu 0001, Shiqiang Yang |
ICME | 4 |
| 2014 | Joint online transcoding and geo-distributed delivery for dynamic adaptive streamingabstractDynamic adaptive video streaming has emerged as a popular approach for video streaming in today's Internet. To date the two important components in dynamic adaptive streaming, video transcoding which generates the adaptive bitrates of a video and video delivery which streams the videos to users, have been separately studied, resulting in a huge waste of computation and storage resource due to transcoding useless videos and suboptimal streaming quality due to homogeneous video replication. In this paper, we propose to jointly perform video transcoding and video delivery for adaptive streaming in an online manner. We conduct extensive measurement studies of a video sharing system and a CDN to motivate our design. We formulate and solve optimization problems to enable high streaming quality for the users, and low computation and replication costs for the system. In particular, our design connects video transcoding and video delivery based on users' preferences of CDN regions and regional preferences of video versions. Extensive trace-driven experiments further confirm the superiority of our design. Zhi Wang 0001, Lifeng Sun, Chuan Wu 0001, Wenwu Zhu 0001, Shiqiang Yang |
INFOCOM | 5 |
| 2014 | CatchSync: catching synchronized behavior in large directed graphsabstractGiven a directed graph of millions of nodes, how can we automatically spot anomalous, suspicious nodes, judging only from their connectivity patterns? Suspicious graph patterns show up in many applications, from Twitter users who buy fake followers, manipulating the social network, to botnet members performing distributed denial of service attacks, disturbing the network traffic graph. We propose a fast and effective method, CatchSync, which exploits two of the tell-tale signs left in graphs by fraudsters: (a) synchronized behavior: suspicious nodes have extremely similar behavior pattern, because they are often required to perform some task together (such as follow the same user); and (b) rare behavior: their connectivity patterns are very different from the majority. We introduce novel measures to quantify both concepts ("synchronicity" and "normality") and we propose a parameter-free algorithm that works on the resulting synchronicity-normality plots. Thanks to careful design, CatchSync has the following desirable properties: (a) it is scalable to large datasets, being linear on the graph size; (b) it is parameter free; and (c) it is side-information-oblivious: it can operate using only the topology, without needing labeled data, nor timing information, etc., while still capable of using side information, if available. We applied CatchSync on two large, real datasets 1-billion-edge Twitter social graph and 3-billion-edge Tencent Weibo social graph, and several synthetic ones; CatchSync consistently outperforms existing competitors, both in detection accuracy by 36% on Twitter and 20% on Tencent Weibo, as well as in speed. Meng Jiang 0001, Peng Cui 0001, Alex Beutel, Christos Faloutsos, Shiqiang Yang |
KDD | 5 |
| 2014 | FEMA: flexible evolutionary multi-faceted analysis for dynamic behavioral pattern discoveryabstractBehavioral pattern discovery is increasingly being studied to understand human behavior and the discovered patterns can be used in many real world applications such as web search, recommender system and advertisement targeting. Traditional methods usually consider the behaviors as simple user and item connections, or represent them with a static model. In real world, however, human behaviors are actually complex and dynamic: they include correlations between user and multiple types of objects and also continuously evolve along time. These characteristics cause severe data sparsity and computational complexity problem, which pose great challenge to human behavioral analysis and prediction. In this paper, we propose a Flexible Evolutionary Multi-faceted Analysis (FEMA) framework for both behavior prediction and pattern mining. FEMA utilizes a flexible and dynamic factorization scheme for analyzing human behavioral data sequences, which can incorporate various knowledge embedded in different object domains to alleviate the sparsity problem. We give approximation algorithms for efficiency, where the bound of approximation loss is theoretically proved. We extensively evaluate the proposed method in two real datasets. For the prediction of human behaviors, the proposed FEMA significantly outperforms other state-of-the-art baseline methods by 17.4%. Moreover, FEMA is able to discover quite a number of interesting multi-faceted temporal patterns on human behaviors with good interpretability. More importantly, it can reduce the run time from hours to minutes, which is significant for industry to serve real-time applications. Meng Jiang 0001, Peng Cui 0001, Fei Wang 0001, Xinran Xu, Wenwu Zhu 0001, Shiqiang Yang |
KDD | 6 |
| 2014 | Emotionally Representative Image Discovery for Social EventsabstractWith the emerging social networks, images have become a major medium for emotion delivery in social events due to their infectious and vivid characteristics. Discovering the emotionally representative images can help people intuitively understand the emotional aspects of social events. Prior works focus on finding the most visually representative images for the target queries or social events. However, the emotionally representative image should not only visually relevant with the social event, but also has a strong emotional appeal among people. In this paper, we propose an emotionally representative image discovery framework by jointly considering textual, visual and social factors. In particular, we build a hybrid link graph for images of each social event, where the weight of each link is measured by textual emotion information, visual similarity and social similarity. Then we propose the Visual-Social-Textual Rank (VSTRank) algorithm to calculate the importance score for each image, so that the emotionally representative images can be discovered under the constraint of textual, visual and social representativeness. To evaluate the effectiveness of our approach, we conduct a series of experiments with 15 social events extracted from real social media dataset, and evaluate the proposed method with both quantitative criterions and user study. Peng Cui 0001, Wenwu Zhu 0001, H. Vicky Zhao, Shiqiang Yang |
ICMR | 6 |
| 2014 | Social Embedding Image Distance LearningabstractImage distance (similarity) is a fundamental and important problem in image processing. However, traditional visual features based image distance metrics usually fail to capture human cognition. This paper presents a novel Social embedding Image Distance Learning (SIDL) approach to embed the similarity of collective social and behavioral information into visual space. The social similarity is estimated according to multiple social factors. Then a metric learning method is especially designed to learn the distance of visual features from the estimated social similarity. In this manner, we can evaluate the cognitive image distance based on the visual content of images. Comprehensive experiments are designed to investigate the effectiveness of SIDL, as well as the performance in the image recommendation and reranking tasks. The experimental results show that the proposed approach makes a marked improvement compared to the state-of-the-art image distance metrics. An interesting observation is given to show that the learned image distance can better reflect human cognition. Peng Cui 0001, Wenwu Zhu 0001, Shiqiang Yang, Qi Tian 0001 |
ACM Multimedia | 4 |
| 2014 | Inferring Strange Behavior from Connectivity Pattern in Social Networks
Meng Jiang 0001, Peng Cui 0001, Alex Beutel, Christos Faloutsos, Shiqiang Yang |
PAKDD (1) | 5 |
| 2014 | Social-oriented visual image search
Peng Cui 0001, Huan-Bo Luan, Wenwu Zhu 0001, Shiqiang Yang, Qi Tian 0001 |
Comput. Vis. Image Underst. | 5 |
| 2014 | A data-driven study of image feature extraction and fusion
Peng Cui 0001, Fangtao Li, Edward Y. Chang, Shiqiang Yang |
Inf. Sci. | 5 |
| 2014 | Scalable Recommendation with Social Contextual InformationabstractExponential growth of information generated by online social networks demands effective and scalable recommender systems to give useful results. Traditional techniques become unqualified because they ignore social relation data; existing social recommendation approaches consider social network structure, but social contextual information has not been fully considered. It is significant and challenging to fuse social contextual factors which are derived from users' motivation of social behaviors into social recommendation. In this paper, we investigate the social recommendation problem on the basis of psychology and sociology studies, which exhibit two important factors: individual preference and interpersonal influence. We first present the particular importance of these two factors in online behavior prediction. Then we propose a novel probabilistic matrix factorization method to fuse them in latent space. We further provide a scalable algorithm which can incrementally process the large scale data. We conduct experiments on both Facebook style bidirectional and Twitter style unidirectional social network data sets. The empirical results and analysis on these two large data sets demonstrate that our method significantly outperforms the existing approaches.approaches. Meng Jiang 0001, Peng Cui 0001, Fei Wang 0001, Wenwu Zhu 0001, Shiqiang Yang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2014 | Social-Sensed Image SearchabstractAlthough Web search techniques have greatly facilitate users’ information seeking, there are still quite a lot of search sessions that cannot provide satisfactory results, which are more serious in Web image search scenarios. How to understand user intent from observed data is a fundamental issue and of paramount significance in improving image search performance. Previous research efforts mostly focus on discovering user intent either from clickthrough behavior in user search logs (e.g., Google), or from social data to facilitate vertical image search in a few limited social media platforms (e.g., Flickr). This article aims to combine the virtues of these two information sources to complement each other, that is, sensing and understanding users’ interests from social media platforms and transferring this knowledge to rerank the image search results in general image search engines. Toward this goal, we first propose a novel social-sensed image search framework, where both social media and search engine are jointly considered. To effectively and efficiently leverage these two kinds of platforms, we propose an example-based user interest representation and modeling method, where we construct a hybrid graph from social media and propose a hybrid random-walk algorithm to derive the user-image interest graph. Moreover, we propose a social-sensed image reranking method to integrate the user-image interest graph from social media and search results from general image search engines to rerank the images by fusing their social relevance and visual relevance. We conducted extensive experiments on real-world data from Flickr and Google image search, and the results demonstrated that the proposed methods can significantly improve the social relevance of image search results while maintaining visual relevance well. Peng Cui 0001, Wenwu Zhu 0001, Huan-Bo Luan, Tat-Seng Chua, Shiqiang Yang |
ACM Trans. Inf. Syst. | 6 |
| 2014 | Bilateral Correspondence Model for Words-and-Pictures Association in Multimedia-Rich MicroblogsabstractNowadays, the amount of multimedia contents in microblogs is growing significantly. More than 20% of microblogs link to a picture or video in certain large systems. The rich semantics in microblogs provides an opportunity to endow images with higher-level semantics beyond object labels. However, this raises new challenges for understanding the association between multimodal multimedia contents in multimedia-rich microblogs. Disobeying the fundamental assumptions of traditional annotation, tagging, and retrieval systems, pictures and words in multimedia-rich microblogs are loosely associated and a correspondence between pictures and words cannot be established. To address the aforementioned challenges, we present the first study analyzing and modeling the associations between multimodal contents in microblog streams, aiming to discover multimodal topics from microblogs by establishing correspondences between pictures and words in microblogs. We first use a data-driven approach to analyze the new characteristics of the words, pictures, and their association types in microblogs. We then propose a novel generative model called the Bilateral Correspondence Latent Dirichlet Allocation (BC-LDA) model. Our BC-LDA model can assign flexible associations between pictures and words and is able to not only allow picture-word co-occurrence with bilateral directions, but also single modal association. This flexible association can best fit the data distribution, so that the model can discover various types of joint topics and generate pictures and words with the topics accordingly. We evaluate this model extensively on a large-scale real multimedia-rich microblogs dataset. We demonstrate the advantages of the proposed model in several application scenarios, including image tagging, text illustration, and topic discovery. The experimental results demonstrate that our proposed model can significantly and consistently outperform traditional approaches. Peng Cui 0001, Lexing Xie, Wenwu Zhu 0001, Yong Rui, Shiqiang Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2013 | Comparisons reducing for local stereo matching using hierarchical structureabstractWe propose a method for local stereo matching that reduces the number of matching comparisons, which also reduces the running time. By using hierarchical structure with suitable disparity candidate set, the proposed method maintains accuracy of local matching methods. Theoretical analysis and experimental results show that the number of comparisons of the proposed method is independent of the disparity range, while the accuracy is nearly the same as the corresponding local method. WeiDong Hu, Kang Zhang 0004, Lifeng Sun, Shiqiang Yang |
ICME | 4 |
| 2013 | Cascading outbreak prediction in networks: a data-driven approachabstractCascades are ubiquitous in various network environments such as epidemic networks, traffic networks, water distribution networks and social networks. The outbreaks of cascades will often bring bad or even devastating effects. How to accurately predict the cascading outbreaks in early stage is of paramount importance for people to avoid these bad effects. Although there have been some pioneering works on cascading outbreaks detection, how to predict, rather than detect, the cascading outbreaks is still an open problem. In this paper, we attempt harnessing historical cascade data, propose a novel data driven approach to select important nodes as sensors, and predict the outbreaks based on the cascading behaviors of these sensors. In particular, we propose Orthogonal Sparse LOgistic Regression (OSLOR) method to jointly optimize node selection and outbreak prediction, where the prediction loss are combined with an orthogonal regularizer and L1 regularizer to guarantee good prediction accuracy, as well as the sparsity and low-redundancy of selected sensors. We evaluate the proposed method on a real online social network dataset including 182.7 million information cascades. The experimental results show that the proposed OSLOR significantly and consistently outperform topological measure based method and other data driven methods in prediction performances. Peng Cui 0001, Shifei Jin, Linyun Yu, Fei Wang 0001, Wenwu Zhu 0001, Shiqiang Yang |
KDD | 6 |
| 2013 | Comparing apples to oranges: a scalable solution with heterogeneous hashingabstractAlthough hashing techniques have been popular for the large scale similarity search problem, most of the existing methods for designing optimal hash functions focus on homogeneous similarity assessment, i.e., the data entities to be indexed are of the same type. Realizing that heterogeneous entities and relationships are also ubiquitous in the real world applications, there is an emerging need to retrieve and search similar or relevant data entities from multiple heterogeneous domains, e.g., recommending relevant posts and images to a certain Facebook user. In this paper, we address the problem of ``comparing apples to oranges'' under the large scale setting. Specifically, we propose a novel Relation-aware Heterogeneous Hashing (RaHH), which provides a general framework for generating hash codes of data entities sitting in multiple heterogeneous domains. Unlike some existing hashing methods that map heterogeneous data in a common Hamming space, the RaHH approach constructs a Hamming space for each type of data entities, and learns optimal mappings between them simultaneously. This makes the learned hash codes flexibly cope with the characteristics of different data domains. Moreover, the RaHH framework encodes both homogeneous and heterogeneous relationships between the data entities to design hash functions with improved accuracy. To validate the proposed RaHH method, we conduct extensive evaluations on two large datasets; one is crawled from a popular social media sites, Tencent Weibo, and the other is an open dataset of Flickr(NUS-WIDE). The experimental results clearly demonstrate that the RaHH outperforms several state-of-the-art hashing methods with significant performance gains. Mingdong Ou, Peng Cui 0001, Fei Wang 0001, Jun Wang 0006, Wenwu Zhu 0001, Shiqiang Yang |
KDD | 6 |
| 2013 | User interest and social influence based emotion prediction for individualsabstractEmotions are playing significant roles in daily life, making emotion prediction important. To date, most of state-of-the-art methods make emotion prediction for the masses which are invalid for individuals. In this paper, we propose a novel emotion prediction method for individuals based on user interest and social influence. To balance user interest and social influence, we further propose a simple yet efficient weight learning method in which the weights are obtained from users' behaviors. We perform experiments in real social media network, with 4,257 users and 2,152,037 microblogs. The experimental results demonstrate that our method outperforms traditional methods with significant performance gains. Peng Cui 0001, Wenwu Zhu 0001, Shiqiang Yang |
ACM Multimedia | 4 |
| 2013 | Social Visual Image Ranking for Web Image Search
Peng Cui 0001, Huan-Bo Luan, Wenwu Zhu 0001, Shiqiang Yang, Qi Tian 0001 |
MMM (1) | 5 |
| 2013 | Joint Social and Content Recommendation for User-Generated Videos in Online Social NetworkabstractOnline social network is emerging as a promising alternative for users to directly access video contents. By allowing users to import videos and re-share them through the social connections, a large number of videos are available to users in the online social network. The rapid growth of the user-generated videos provides enormous potential for users to find the ones that interest them; while the convergence of online social network service and online video sharing service makes it possible to perform recommendation using social factors and content factors jointly. In this paper, we design a joint social-content recommendation framework to suggest users which videos to import or re-share in the online social network. In this framework, we first propose a user-content matrix update approach which updates and fills in cold user-video entries to provide the foundations for the recommendation. Then, based on the updated user-content matrix, we construct a joint social-content space to measure the relevance between users and videos, which can provide a high accuracy for video importing and re-sharing recommendation. We conduct experiments using real traces from Tencent Weibo and Youku to verify our algorithm and evaluate its performance. The results demonstrate the effectiveness of our approach and show that our approach can substantially improve the recommendation accuracy. Zhi Wang 0001, Lifeng Sun, Wenwu Zhu 0001, Shiqiang Yang, Dapeng Oliver Wu |
IEEE Trans. Multim. | 4 |
| 2013 | Peer-Assisted Social Media Streaming with Social ReciprocityabstractOnline video sharing and social networking are cross-pollinating rapidly in today's Internet: Online social network users are sharing more and more media contents among each other, while online video sharing sites are leveraging social connections among users to promote their videos. An intriguing development as it is, the operational challenge in previous video sharing systems persists, em i.e., the large server cost demanded for scaling of the systems. Peer-to-peer video sharing could be a rescue, only if the video viewers' mutual resource contribution has been fully incentivized and efficiently scheduled. Exploring the unique advantages of a social network based video sharing system, we advocate to utilize social reciprocities among peers with social relationships for efficient contribution incentivization and scheduling, so as to enable high-quality video streaming with low server cost. We exploit social reciprocity with two give-and-take ratios at each peer: (1) peer contribution ratio (em PCR), which evaluates the reciprocity level between a pair of social friends, and (2) system contribution ratio (em SCR), which records the give-and-take level of the user to and from the entire system. We design efficient peer-to-peer mechanisms for video streaming using the two ratios, where each user optimally decides which other users to seek relay help from and help in relaying video streams, respectively, based on combined evaluations of their social relationship and historical reciprocity levels. Our design achieves effective incentives for resource contribution, load balancing among relay peers, as well as efficient social-aware resource scheduling. We also discuss practical implementation and implement our design in a prototype social media sharing system. Our extensive evaluations based on PlanetLab experiments verify that high-quality large-scale social media sharing can be achieved with conservative server costs. Zhi Wang 0001, Chuan Wu 0001, Lifeng Sun, Shiqiang Yang |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2013 | Propagation-based social-aware multimedia content distributionabstractOnline social networks have reshaped how multimedia contents are generated, distributed, and consumed on today's Internet. Given the massive number of user-generated contents shared in online social networks, users are moving to directly access these contents in their preferred social network services. It is intriguing to study the service provision of social contents for global users with satisfactory quality of experience. In this article, we conduct large-scale measurement of a real-world online social network system to study the social content propagation. We have observed important propagation patterns, including social locality, geographical locality, and temporal locality. Motivated by the measurement insights, we propose a propagation-based social-aware delivery framework using a hybrid edge-cloud and peer-assisted architecture. We also design replication strategies for the architecture based on three propagation predictors designed by jointly considering user, content, and context information. In particular, we design a propagation region predictor and a global audience predictor to guide how the edge-cloud servers backup the contents, and a local audience predictor to guide how peers cache the contents for their friends. Our trace-driven experiments further demonstrate the effectiveness and superiority of our design. Zhi Wang 0001, Wenwu Zhu 0001, Xiangwen Chen, Lifeng Sun, Jiangchuan Liu, Minghua Chen 0001, Peng Cui 0001, Shiqiang Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2013 | Find where you are: a new try in place recognition
Mingying Gong, Lifeng Sun, Shiqiang Yang |
Vis. Comput. | 3 |
| 2012 | Social contextual recommendationabstractExponential growth of information generated by online social networks demands effective recommender systems to give useful results. Traditional techniques become unqualified because they ignore social relation data; existing social recommendation approaches consider social network structure, but social context has not been fully considered. It is significant and challenging to fuse social contextual factors which are derived from users' motivation of social behaviors into social recommendation. In this paper, we investigate social recommendation on the basis of psychology and sociology studies, which exhibit two important factors: individual preference and interpersonal influence. We first present the particular importance of these two factors in online item adoption and recommendation. Then we propose a novel probabilistic matrix factorization method to fuse them in latent spaces. We conduct experiments on both Facebook style bidirectional and Twitter style unidirectional social network datasets in China. The empirical result and analysis on these two large datasets demonstrate that our method significantly outperform the existing approaches. Meng Jiang 0001, Peng Cui 0001, Rui Liu 0014, Qiang Yang 0001, Fei Wang 0001, Wenwu Zhu 0001, Shiqiang Yang |
CIKM | 7 |
| 2012 | Social recommendation across multiple relational domainsabstractSocial networks enable users to create different types of personal items. In dealing with serious information overload, the major problems of social recommendation are sparsity and cold start. In existing approaches, relational and heterogeneous domains can not be effectively utilized for social recommendation, which brings a challenge to model users and multiple types of items together on social networks. In this paper, we consider how to represent social networks with multiple relational domains and alleviate the major problems in an individual domain by transferring knowledge from other domains. We propose a novel Hybrid Random Walk (HRW), which can integrate multiple heterogeneous domains including directed/undirected links, signed/unsigned links and within-domain/cross-domain links into a star-structured hybrid graph with user graph at the center. We perform random walk until convergence and use the steady state distribution for recommendation. We conduct experiments on a real social network dataset and show that our method can significantly outperform existing social recommendation approaches. Meng Jiang 0001, Peng Cui 0001, Fei Wang 0001, Qiang Yang 0001, Wenwu Zhu 0001, Shiqiang Yang |
CIKM | 6 |
| 2012 | Cloud-based social application deployment using local processing and global distributionabstractSocial applications represent a paradigm shift on how the Internet is to be used, and have already changed the way we work, live, and play. When it comes to deploying social applications, cloud computing platforms are used to meet the Internet-scale, self-propagating, and fast-growing demands from these applications. Yet, to deploy social media applications in the most effective and economic fashion, we need to strategically design and follow a set of theoretical and practical principles. In this paper, we seek to design a set of new principles to guide social application deployment. Learning from large-scale measurement-based observations using a real-world social application, the gist of our principles is to detach the typically integrated "collection → processing → distribution" work ows in social applications into separate local processing and global distribution procedures, which can be effectively deployed using different cloud services. Moreover, based on a predictive model of regional propagation, we formulate the resource allocation problems in the processes of collecting/processing and distributing content as two optimization problems, which can be solved by efficient algorithms. Finally, based on our theoretical design, we have implemented an example social application on Amazon EC2 and Google AppEngine, where IaaS-based computation instances perform content collection and processing, and the PaaS-based platform is employed to distribute the contents that are widely propagating. Our PlanetLab-based trace-driven experiments have further confirmed the superiority of our design. Zhi Wang 0001, Baochun Li, Lifeng Sun, Shiqiang Yang |
CoNEXT | 4 |
| 2012 | Robust Place Recognition by Avoiding Confusing Features and Fast Geometric Re-ranking
Mingying Gong, Lifeng Sun, Shiqiang Yang |
CVM | 3 |
| 2012 | Binary stereo matching
Kang Zhang 0004, Jiyang Li, WeiDong Hu, Lifeng Sun, Shiqiang Yang |
ICPR | 6 |
| 2012 | Guiding internet-scale video service deployment using microblog-based predictionabstractOnline microblogging has been very popular in today's Internet, where users exchange short messages and follow various contents shared by people that they are interested in. Among the variety of exchanges, video links are a representative type on a microblogging site. More and more viewers of an Internet video service are coming from microblog recommendations. It is intriguing research to explore the connections between the patterns of microblog exchanges and the popularity of videos, in order to potentially use the propagation patterns of microblogs to guide proactive service deployment of a video sharing system. Based on extensive traces from Youku and Tencent Weibo, a popular video sharing site and a favored microblogging system in China, we explore how patterns of video link propagation in the microblogging system are correlated with video popularity on the video sharing site, at different times and in different geographic regions. Using influential factors summarized from the measurement studies, we further design neural network-based learning frameworks to predict the number of potential viewers of different videos and the geographic distribution of viewers. Experiments show that our neural network-based frameworks achieve better prediction accuracy, as compared to a classical approach that relies on historical numbers of views. We also briefly discuss how proactive video service deployment can be effectively enabled by our prediction frameworks. Zhi Wang 0001, Lifeng Sun, Chuan Wu 0001, Shiqiang Yang |
INFOCOM | 4 |
| 2012 | Analyzing social media via event facetsabstractMicroblog is a prominent information platform for sharing experiences, discussing current events, and exchanging ideas. Many events are first reported in social media, and increasing amounts of rich-media content are associated with the posts, making them more credible and attractive. We design a rich-media analysis system to address the important challenge of sensing and exploring events from social media in real-time. The system includes a novel bilateral correspondence topic model to extract representative content and meaningful facets about events over time. It also includes a digital magazine that anchors user interactions with event facets. We demonstrate several examples from more than 4 million rich media microblogs, showing the effectiveness of key content extraction and natrual interactions with facets. Peng Cui 0001, Lexing Xie, Wenwu Zhu 0001, Shiqiang Yang |
ACM Multimedia | 6 |
| 2012 | Propagation-based social-aware replication for social video contentsabstractOnline social network has reshaped the way how video contents are generated, distributed and consumed on today's Internet. Given the massive number of videos generated and shared in online social networks, it has been popular for users to directly access video contents in their preferred social network services. It is intriguing to study the service provision of social video contents for global users with satisfactory quality-of-experience. In this paper, we conduct large-scale measurement of a real-world online social network system to study the propagation of the social video contents. We have summarized important characteristics from the video propagation patterns, including social locality, geographical locality and temporal locality. Motivated by the measurement insights, we propose a propagation-based social-aware replication framework using a hybrid edge-cloud and peer-assisted architecture, namely PSAR, to serve the social video contents. Our replication strategies in PSAR are based on the design of three propagation-based replication indices, including a geographic influence index and a content propagation index to guide how the edge-cloud servers backup the videos, and a social influence index to guide how peers cache the videos for their friends. By incorporating these replication indices into our system design, PSAR has significantly improved the replication performance and the video service quality. Our trace-driven experiments further demonstrate the effectiveness and superiority of PSAR, which improves the local download ratio in the edge-cloud replication by 30%, and the local cache hit ratio in the peer-assisted replication by 40%, against traditional approaches. Zhi Wang 0001, Lifeng Sun, Xiangwen Chen, Wenwu Zhu 0001, Jiangchuan Liu, Minghua Chen 0001, Shiqiang Yang |
ACM Multimedia | 7 |
| 2012 | Learning influence from heterogeneous social networks
Lu Liu 0005, Jie Tang 0001, Jiawei Han 0001, Shiqiang Yang |
Data Min. Knowl. Discov. | 4 |
| 2012 | A probabilistic graphical model for topic and preference discovery on social media
Lu Liu 0005, Feida Zhu 0001, Lei Zhang 0001, Shiqiang Yang |
Neurocomputing | 4 |
| 2012 | Guest editorial: Special issue on information retrieval for social media
Fei Wang 0001, Peng Cui 0001, Gordon Sun, Tat-Seng Chua, Shiqiang Yang |
Inf. Retr. | 5 |
| 2012 | Mining diversity on social media networks
Lu Liu 0005, Feida Zhu 0001, Meng Jiang 0001, Jiawei Han 0001, Lifeng Sun, Shiqiang Yang |
Multim. Tools Appl. | 6 |
| 2012 | Highly Scalable Parallel Arithmetic Coding on Multi-Core Processors Using LDPC CodesabstractWe describe a highly scalable parallel arithmetic coder for Markov inputs suitable for implementation on modern multi-core processors. The algorithm divides the input into interleaved sub-sequences which can be then processed independently on different processing units using LDPC-based Slepian-Wolf coding. Experimental simulations show good scalability of the proposed algorithm while also maintaining good compression performance. Notably, when compared with traditional parallel arithmetic coding, the proposed method maintains a much higher efficiency both respect to the entropy limit as well as in terms of the ability to distribute computations across multiple cores without performance loss. WeiDong Hu, Jiangtao Wen, Weiyi Wu, Yuxing Han 0001, Shiqiang Yang, John D. Villasenor |
IEEE Trans. Commun. | 5 |
| 2012 | A Fast and Robust Sparse Approach for Hyperspectral Data Classification Using a Few Labeled SamplesabstractThe classification of high-dimensional data with too few labeled samples is a major challenge which is difficult to meet unless some special characteristics of the data can be exploited. In remote sensing, the problem is particularly serious because of the difficulty and cost factors involved in assignment of labels to high-dimensional samples. In this paper, we exploit certain special properties of hyperspectral data and propose an$\ell^{1}$-minimization -based sparse representation classification approach to overcome this difficulty in hyperspectral data classification. We assume that the data within each hyperspectral data class lies in a very low-dimensional subspace. Unlike traditional supervised methods, the proposed method does not have separate training and testing phases and, therefore, does not need a training procedure for model creation. Further, to prove the sparsity of hyperspectral data and handle the computational intensiveness and time demand of general-purpose linear programming (LP) solvers, we propose a Homotopy-based sparse classification approach, which works efficiently when data is highly sparse. The approach is not only time efficient, but it also produces results, which are comparable to the traditional methods. The proposed approaches are tested for our difficult classification problem of hyperspectral data with few labeled samples. Extensive experiments on four real hyperspectral data sets prove that hyperspectral data is highly sparse in nature, and the proposed approaches are robust across different databases, offer more classification accuracy, and are more efficient than state-of-the-art methods. Qazi Sami ul Haq, Linmi Tao, Fuchun Sun 0001, Shiqiang Yang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2012 | Analyzing Image Deblurring Through Three ParadigmsabstractTo recover a sharp version from a blurred image is a long-standing inverse problem. In this paper, we analyze the research on this topic both theoretically and experimentally through three paradigms: 1) the deterministic filter; 2) Bayesian estimation; and 3) the conjunctive deblurring algorithm (CODA), which performs the deterministic filter and Bayesian estimation in a conjunctive manner. We point out the weaknesses of the deterministic filter and unify the limitation latent in two kinds of Bayesian estimators. We further explain why the CODA is able to handle quite large blurs beyond Bayesian estimation. Finally, we propose a novel method to overcome several unreported limitations of the CODA. Although extensive experiments demonstrate that our method outperforms state-of-the-art methods with a large margin, some common problems of image deblurring still remain unsolved and should attract further research efforts. Chao Wang 0063, Lifeng Sun, Peng Cui 0001, Jianwei Zhang 0001, Shiqiang Yang |
IEEE Trans. Image Process. | 5 |
| 2012 | A Matrix-Based Approach to Unsupervised Human Action CategorizationabstractHuman action, as the basic unit of most human-relevant video content, bridges the gap between low-level visual features and high-level semantics. Human action recognition is of great significance in the applications of human-computer interaction, intelligent video surveillance, video retrieval and search. In this paper, we propose a novel unsupervised approach to mining categories from action video sequences, which consists of two modules: action representation for video data structurization and learning model for unsupervised categorization. In action representation, a novel view of video decomposition is presented. Videos are regarded as spatially distributed dynamic pixel time series, and these dynamic pixels are first quantized into pixel prototypes. After replacing the pixel time series with their corresponding prototype labels, the video sequences are compressed into two-dimensional action matrices. In the learning model, we put these matrices together to form an multi-action tensor, and propose the joint matrix factorization method to simultaneously cluster the pixel prototypes into pixel signatures, and matrices into action classes with the consideration of the duality between pixel clustering and action clustering. The approach is tested on public and popular Weizmann, and KTH datasets, and promising results are achieved. Peng Cui 0001, Fei Wang 0001, Lifeng Sun, Jianwei Zhang 0001, Shiqiang Yang |
IEEE Trans. Multim. | 5 |
| 2011 | Item-Level Social Influence Prediction with Probabilistic Hybrid Factor Matrix FactorizationabstractSocial influence has become the essential factor which drives the dynamic evolution process of social network structure and user behaviors. Previous research often focus on social influence analysis in network-level or topic-level. In this paper, we concentrate on predicting item-level social influence to reveal the users' influences in a more fine-grained level. We formulate the social influence prediction problem as the estimation of a user-post matrix, where each entry in the matrix represents the social influence strength the corresponding user has given the corresponding web post. To deal with the sparsity and complex factor challenges in the research, we model the problem by extending the probabilistic matrix factorization method to incorporate rich prior knowledge on both user dimension and web post dimension, and propose the Probabilistic Hybrid Factor Matrix Factorization (PHF-MF) approach. Intensive experiments are conducted on a real world online social network to demonstrate the advantages and characteristics of the proposed method. Peng Cui 0001, Fei Wang 0001, Shiqiang Yang, Lifeng Sun |
AAAI | 3 |
| 2011 | Peer-assisted online games with social reciprocityabstractOnline games and social networks are cross-pollinating rapidly in today's Internet: Online social network sites are deploying more and more games in their systems, while online game providers are leveraging social networks to power their games. An intriguing development as it is, the operational challenge in the previous game persists, i.e., the large server operational cost remains a non-negligible obstacle for deploying high-quality multi-player games. Peer-to-peer based game network design could be a rescue, only if the game players' mutual resource contribution has been fully incentivized and efficiently scheduled. Exploring the unique advantage of social network based games (social games), we advocate to utilize social reciprocities among peers with social relationships for efficient contribution incentivization and scheduling, so as to power a high-quality online game with low server cost. In this paper, social reciprocity is exploited with two give-and-take ratios at each peer: (1) peer contribution ratio (PCR), which evaluates the reciprocity level between a pair of social friends, and (2) system contribution ratio (SCR), which records the give-and-take level of the player to and from the entire network. We design efficient peer-to-peer mechanisms for game state distribution using the two ratios, where each player optimally decides which other players to seek relay help from and help in relaying game states, respectively, based on combined evaluations of their social relationship and historical reciprocity levels. Our design achieves effective incentives for resource contribution, load balancing among relay peers, as well as efficient social-aware resource scheduling. We also discuss practical implementation concerns and implement our design in a prototype online social game. Our extensive evaluations based on experiments on PlanetLab verify that high-quality large-scale social games can be achieved with conservative server costs. Zhi Wang 0001, Chuan Wu 0001, Lifeng Sun, Shiqiang Yang |
IWQoS | 4 |
| 2011 | ACM international workshop on social and behavioral networked media access (SBNMA'11)abstractIn an endeavour to speak and prevail over some of the open problems that obstruct efficient networked media, this workshop will fetch together folks from a number of research communities, including but not limited to Multimedia Distribution and Access, Social Network Analysis, Multimedia Content Analysis, Behavioral Analysis, User Modelling Adaptation and Personalization. It is our credence that a synergetic approach involving the above mentioned research areas can surpass their individual potentials, leading to improved networked media access. The main objective of this workshop is to provide a forum to disseminate work that explicitly exploit the synergy between multimedia content analysis, behavioral modelling, personalisation, and next generation networking and community aspects of social networks. This synergetic methodology could produce high quality of experience for personalized multimedia access in networking environment. Naeem Ramzan, Fei Wang 0001, Charalampos Z. Patrikakis, Peng Cui 0001, Nikolaos D. Doulamis, Shiqiang Yang, Gordon Sun |
ACM Multimedia | 6 |
| 2011 | Prefetching strategy in peer-assisted social video streamingabstractOnline social network has emerged as the most popular approach for people to directly access multimedia contents. Among these contents, video sharing is a challenging task due to the demand on a large amount of uplink bandwidth at the dedicated server. We leverage a P2P paradigm to alleviate the server to distribute shared videos. By investigating traces obtained from a popular online social network in China, we observe that users' preferences can be predicted. We design a user preference guided prefetching strategy to reduce video startup delays, enabling smooth playback. Simulation experiments show that our design achieves high prefetch accuracy and short startup delay with conservative storage and bandwidth capacities at peers. Zhi Wang 0001, Lifeng Sun, Shiqiang Yang, Wenwu Zhu 0001 |
ACM Multimedia | 3 |
| 2011 | Who should share what?: item-level social influence prediction for users and posts rankingabstractPeople and information are two core dimensions in a social network. People sharing information (such as blogs, news, albums, etc.) is the basic behavior. In this paper, we focus on predicting item-level social influence to answer the question Who should share What, which can be extended into two information retrieval scenarios: (1) Users ranking: given an item, who should share it so that its diffusion range can be maximized in a social network; (2) Web posts ranking: given a user, what should she share to maximize her influence among her friends. We formulate the social influence prediction problem as the estimation of a user-post matrix, in which each entry represents the strength of influence of a user given a web post. We propose a Hybrid Factor Non-Negative Matrix Factorization (HF-NMF) approach for item-level social influence modeling, and devise an efficient projected gradient method to solve the HF-NMF problem. Intensive experiments are conducted and demonstrate the advantages and characteristics of the proposed method. Peng Cui 0001, Fei Wang 0001, Mingdong Ou, Shiqiang Yang, Lifeng Sun |
SIGIR | 5 |
| 2011 | Virtual support window for adaptive-weight stereo matchingabstractIt attracts many researchers' attention to find a stereo matching algorithm both accurate and fast. Yoon and Kweon's Adaptive Support-Weight(ASW) method is supposed to be a very successful algorithm in both accuracy and speed. However, it is very time consuming for ASW to take a large number of pixels into consideration for computing a disparity. In this paper, we present a stereo matching algorithm based on the ASW method to solve this problem. We introduce a time efficient technique, partial sum, to reduce the information of pixels in a large window into a small one so that disparities can be computed using the large window of pixels with low time complexity. Our method can obtain better representation of local properties of a pixel than the ASW method. Experimental results show that our method produces more accurate disparity maps than ASW method. WeiDong Hu, Kang Zhang 0004, Lifeng Sun, Jiyang Li, Shiqiang Yang |
VCIP | 6 |
| 2011 | Hierarchical visual event pattern mining and its applications
Peng Cui 0001, Lifeng Sun, Shiqiang Yang |
Data Min. Knowl. Discov. | 4 |
| 2010 | Mining topic-level influence in heterogeneous networksabstractInfluence is a complex and subtle force that governs the dynamics of social networks as well as the behaviors of involved users. Understanding influence can benefit various applications such as viral marketing, recommendation, and information retrieval. However, most existing works on social influence analysis have focused on verifying the existence of social influence. Few works systematically investigate how to mine the strength of direct and indirect influence between nodes in heterogeneous networks. Lu Liu 0005, Jie Tang 0001, Jiawei Han 0001, Meng Jiang 0001, Shiqiang Yang |
CIKM | 5 |
| 2010 | Mining Diversity on Networks
Lu Liu 0005, Feida Zhu 0001, Chen Chen 0005, Xifeng Yan, Jiawei Han 0001, Philip S. Yu, Shiqiang Yang |
DASFAA (1) | 7 |
| 2010 | Image Compression Using the DCT and Noiselets: A New Algorithm and Its Rate Distortion PerformanceabstractWe describe an image coding algorithm combining the DCT and noiselet information. The algorithm first transmits DCT information sufficient to reproduce a "low-quality" version of the image at the decoder. This image is then used both at the decoder and encoder to create a mutually known list of locations of likely significant noiselet coefficients. The coefficient values themselves are then transmitted to the decoder differentially, by subtracting, at the encoder, the low-quality image from the original image, obtaining the noiselet values and subjecting them to quantization and entropy coding. There remain significant opportunities for further work combining CS-inspired information theoretic techniques with the rate-distortion considerations that are critical in practical image communications. Zhuoyuan Chen, Jiangtao Wen, Shiqiang Yang, Yuxing Han 0001, John D. Villasenor |
DCC | 3 |
| 2010 | Reconstruction of Sparse Binary Signals Using Compressive SensingabstractSummary form only given. This paper has described an improved algorithm for reconstructing sparse binary signals using compressive sensing. The algorithm is based on the reweighted lqnorm optimization algorithm, but with the important additional operation of bounding in each round of the interior-point method iteration, and progressive reduction of q. Experimental results confirm that the algorithm performs well both in terms of the ability to recover an input signal as well as in terms of speed. We also found that both the progressive reduction and the bounding are integral to the improvement in performance. Future work includes extending this approach to Gaussian distributed, as opposed to binary inputs. Jiangtao Wen, Zhuoyuan Chen, Shiqiang Yang, Yuxing Han 0001, John D. Villasenor |
DCC | 3 |
| 2010 | Strategies of Collaboration in Multi-Channel P2P VoD StreamingabstractAs compared to live peer-to-peer (P2P) streaming, modern P2P video-on-demand (VoD) systems have brought much larger volumes of videos and more interactive controls to the Internet users. Nevertheless, the larger number of available videos and the flexibility of allowing users to jump back and forth in a video, have led to much fewer numbers of concurrent peers watching at a similar pace, that reduces the chance for collaborative chunk supply among peers and thus significantly increases the server bandwidth cost. Towards the ultimate goal of maximizing peer resource utilization, in this paper, we design effective strategies for both cross-channel and intra-channel collaborations in multi- channel P2P VoD systems, such that individual peer's resources, including download/upload bandwidths and the cache capacity, are effectively utilized to maximize the streaming qualities in all the channels. In particular, each peer actively and strategically determines the supply-and-demand imbalance in different channels, as well as that among different chunks within each video, makes use of its surplus download capacity to fetch chunks with the most need, and then serves those chunks using its idle upload bandwidth, all without impairing its own streaming quality. Our extensive trace-driven simulations show the effectiveness of our strategies in reducing the server cost while guaranteeing high streaming qualities in the entire system, even during extreme scenarios such as unexpected flash crowds. Zhi Wang 0001, Chuan Wu 0001, Lifeng Sun, Shiqiang Yang |
GLOBECOM | 4 |
| 2010 | A compressive sensing image compression algorithm using quantized DCT and noiselet informationabstractInspired by recent theoretical advances in compressive sensing (CS), we propose a new framework that combines the classical local discrete cosine transform used in image compression algorithms such as JPEG with a global noiselet measure which is solved using second order cone programming (SOCP). Jiangtao Wen, Zhuoyuan Chen, Yuxing Han 0001, John D. Villasenor, Shiqiang Yang |
ICASSP | 5 |
| 2010 | HFAG: Hierarchical Frame Affinity Group for video retrieval on very large video datasetabstractContent-based video retrieval systems are desired to fast and accurately find the nearest-neighbors of user input examples from very large video datasets. This poses a great challenge since exhaustive and redundant computation of similarities is required. Cluster based index approaches can be used to address this problem, but the similarity computation and clustering methods for videos are very time-consuming, thus preventing it from indexing very large video datasets. In this paper, we propose the Hierarchical Frame Affinity Group (HFAG), which is a hierarchy of frame clusters built using affinity propagation (AP) method, to represent video clusters. Our proposed video similarity metric and AP method guarantee the high performance of forming HFAG. We then build the cluster-based index structure to support retrieval of the nearest-neighbors of video sequences. The experiments on real large video datasets prove the effectiveness and efficiency of our approach. Yin-Jun Miao, Chao Wang 0063, Peng Cui 0001, Lifeng Sun, Pin Tao, Shiqiang Yang |
ICIP | 6 |
| 2010 | Adaptive Server Bandwidth Allocation for Multi-channel P2P Live Streaming
Lifeng Sun, Shiqiang Yang |
MMM | 3 |
| 2010 | Strategies of buffering schedule in P2P VoD streamingabstractAs compared to live peer-to-peer (P2P) streaming, modern P2P video-on-demand (VoD) systems have brought much larger volumes of videos and more interactive controls to the Internet users. As the increase of bitrate of the videos and the full VCR controls of P2P VoD, the behavior “buffering” motivates us to design different schedule and service strategies for peers, to improve the playback performance, and the alleviation of the dedicated streaming server, by making best use of the bandwidth and cache capacities of these buffering peers. In our design, peers strategically decide which segments in the video to download first, and which requests to serve first. We conduct extended simulations to evaluate the performance of the strategies, and the results show our design outperforms the conventional sequential scheme, with respect to improving the playback quality and reducing the server load. Zhi Wang 0001, Lifeng Sun, Shiqiang Yang |
MMSP | 3 |
| 2010 | iGridMedia: The system to provide low delay peer-to-peer live streaming service over internet
Meng Zhang 0001, Lifeng Sun, Yechang Fang, Shiqiang Yang |
Peer-to-Peer Netw. Appl. | 4 |
| 2009 | Overlay collaboration towards reduced bandwidth costs in multi-view streamingabstractDelivering high-quality multiview video (MV) through Internet is very challenging due to its excessive consumption on server bandwidth resources. Existing solutions encode video contents independently for each view and deliver them separately in isolated view channels, without leveraging the features of MV and taking advantage of multiview video coding(MVC). To minimize the server bandwidth costs, we introduce a novel overlay collaboration framework that unifies all view channels to cooperate in delivering MV: I pictures in MVC are shared among them instead of requesting from server respectively; Surplus resources of hotspot view channels are effectively utilized to help channels with insufficient resources, both of which contribute to remarkable reduction in server bandwidth costs. Simulation experiments show that our method achieves more than 40% reduced bandwidth costs on server while maintaining scalability and resilience to user dynamics. Zhibo Chen 0001, Meng Zhang 0001, Lifeng Sun, Shiqiang Yang |
ICASSP | 4 |
| 2009 | Frame-level heuristic scheduling Multi-view Video Coding on symmetric multi-core architectureabstractIn this paper, we propose a frame-level heuristic scheduling parallel emerging Multi-view Video Coding (MVC) using Directed Acyclic Graph (DAG) on Intel multi-core processor. We illustrate the reason to choose heuristic scheduling and formulate the problem. Through defining dependent degree and concurrent degree, we demonstrate why to choose frame as parallel granularity. Experimental results demonstrate the effectiveness and scalability of our parallel MVC. Yi Pang, Jiangtao Wen, Lifeng Sun, WeiDong Hu, Shiqiang Yang |
ICIP | 5 |
| 2009 | Robust inter-scale non-blind image motion deblurringabstractKernel estimate errors and image noise are major causes of visual artifacts in image motion deblurring. We propose an inter-scale non-blind image motion deblurring approach that significantly reduces those artifacts. We use Gaussian Scale Mixture Field of Experts (GSM FOE) model as image prior. The inter-scale smoothness constraint is adopted to suppress the ringing artifacts. In each scale, image details are recovered by the residual deconvolution and the cross bilateral filter (CBF). We further propose a std-controlled CBF to denoise the result. The experimental results are much better than those of previous methods. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Shiqiang Yang, Jianwei Zhang 0001 |
ICIP | 4 |
| 2009 | High-quality non-blind motion deblurringabstractTraditional non-blind motion deblurring methods are sensitive to kernel estimate errors and image noise, thus suffering from either ringing artifacts, enlarged image noise, or over-smoothed image details. We introduce a robust non-blind deblurring algorithm that produces high quality results even from many challenging images with noisy kernels. We adopt the Gaussian Scale Mixture Fields of Experts (GSM FOE) model and the smoothness constraint as image prior, and use the iterative re-weight least-square (IRLS) algorithm to produce the temporal result. The residual deconvolution suite is used to restore the lost image details. We denoise the result using our std-controlled cross bilateral filter. The experimental results are much better than those of previous approaches. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Shiqiang Yang, Jianwei Zhang 0001 |
ICIP | 4 |
| 2009 | Data parallelization of Kd-tree ray tracing on the Cell Broadband EngineabstractRay tracing is a widely used rendering technique in computer graphics, and its intense computational requirement prohibits real-time ray tracing applications wide use in consumer markets. One main feature of ray tracing is parallelism and the mainstream computer market is switching to systems with multi-core. In this paper, in order to accelerate ray tracing to achieve real-time processing, we propose a data parallel kd-tree ray tracing algorithm on Cell Broadband Engine (Cell/B.E.) which is a state-of-the-art multi-core processor. This paper expounds the feasibility and key issues of kd-tree ray tracing on the Cell/B.E. processor and introduces the implementation and Cell-specific in ray tracing program design. The results highlight that our parallel algorithm for kd-tree ray tracing is scalable with a matrix of cores, resolutions or different object models. The execution time reduces from several minutes to several seconds. Yi Pang, Lifeng Sun, Shiqiang Yang |
ICME | 3 |
| 2009 | Delay-guaranteed Interactive Multiview Video StreamingabstractMultiview video is well known to involve interactions with audience and offer better view experience than conventional single-view video. The interactive nature of multiview video has made the service very delay-sensitive and thus imposes great challenges for providing multiview streaming service in Internet. Currently, almost none of existing works have addressed the delay issue in multiview streaming. To ensure the quality of service and offer preferable view experience for users, we propose a novel streaming framework to provide delay-guaranteed service for interactive multiview video. The basic tradeoff between consumed server bandwidth and the required delay is carefully studied. We leverage the features of MVC to keep the server bandwidth costs low while guarantee the delay constraint. In addition, we introduce a neighbor-assisted view switching scheme: peer's neighbors are involved in its switching process and their upload bandwidth resources are effectively utilized to reduce the bandwidth costs on server. Simulation results show that our proposed framework meet the delay-guaranteed requirement with restrained server bandwidth costs while maintaining scalability and resilience to users dynamics. Zhibo Chen 0001, Meng Zhang 0001, Lifeng Sun, Shiqiang Yang |
ISCAS | 4 |
| 2009 | A Cascade SVM Approach for Head-shoulder Detection using Histograms of Oriented GradientsabstractThis paper presents a head-shoulder detection approach using cascade SVM and histograms of oriented gradients (HOG). The HOG features which are extracted from variable-size blocks can capture salient features of head-shoulder automatically. A two stage cascade using SVM approach is designed to be the classifier. During detection, the majority of negative windows are rejected at the first stage, leaving a relatively small number of windows to be classified at the second stage, which improves the speed and precision of the detector. Due to the large number of possible target locations in an image, we applied camera self-calibration approach to facilitate the estimation for the size and location of the detection window. The experiments on surveillance videos from Trecvid 2008 proved that our approach can achieve fast and accurate head-shoulder detection. Xi-Feng Ding, Peng Cui 0001, Lifeng Sun, Shiqiang Yang |
ISCAS | 5 |
| 2009 | The Bilateral Wavelet Pyramid (BWP): A Novel Image RepresentationabstractIn this paper, we study the properties of the bilateral pyramid, based on which a novel image representation-the bilateral wavelet pyramid (BWP) is proposed. The BWP provides a multiresolution technique to manipulate image variances simultaneously in radiometric, spatial and frequency domains. Application to texture retrieval is discussed to demonstrate the effectiveness of the BWP. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Jianwei Zhang 0001, Shiqiang Yang |
ISCAS | 5 |
| 2009 | Statistics in Bilateral Domain: Novel Statistics of Natural ImagesabstractRecently, there has been a great deal of interest in statistics of natural images in linear domain. However, a detail study of image statistics in non-linear domain on a very large dataset is still missing. In this paper, we analyze natural images' statistics in ldquobilateral domainrdquo on different sub bands which are generated by the classical bilateral filter. Compared with those in linear domain, the statistics in bilateral domain are significantly superior in generalization, de-correlation, and local feature extraction. We also observe some entirely new but interesting features. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Jianwei Zhang 0001, Shiqiang Yang |
ISCAS | 5 |
| 2009 | A network coding scheme for SVC streaming over wireless mesh networkabstractWe argue that the conventional forward error correcting code (FEC) and retransmission solutions will become ineffective to transmit scalable video coding streaming in wireless mesh network (WMN) for the reason that video frame errors appear in bursts rather than random and the bad channel condition usually persists for a period in wireless channels. A novel SVC streaming cooperative transmission scheme is proposed in this paper. Our cooperative transmission scheme includes two stages among nodes in WMN. The first stage is selecting appropriate paths for the relevant SVC sub-layer streams and the second one is selecting nearby nodes to cache the network coding results of the adjacent sub-layer streams and retransmit the network coded packets to the user terminal for recovering the error packets. Simulation results are given to demonstrate the significant improvement of the SVC streaming reconstructive video quality. Liu Yong, Lifeng Sun, Shiqiang Yang |
IWCMC | 3 |
| 2009 | Auto-cut for web imagesabstractIn this paper, we propose a novel automatic algorithm for foreground/background labeling. We aim to generate ROI cutout automatically for further processing such as image editing, classification and information retrieval. Different from traditional semi-supervised segmentation method, we use a rather weak prior on boundary label. Accordingly, a global cost function is proposed to combine our prior knowledge with pixel-level feature. We compute fuzzy matting components as building blocks to construct semantically meaningful mattes. Finally, these mattes are hierarchically clustered and ranked by central preference. Experimental results on a large benchmark data set demonstrate the performance of our algorithm. Zhuoyuan Chen, Lifeng Sun, Shiqiang Yang |
ACM Multimedia | 3 |
| 2009 | Background subtraction in dynamic scenes with adaptive spatial fusingabstractBackground subtraction in highly dynamic scenes has been a critical challenge for traditional pixel-wise background models which perform poorly when the background has dynamic textures. In this paper, we consider background modelling in a spatial perspective and make an attempt to exploit more information from the outputs of pixel-wise model. We propose a background subtraction scheme using adaptive spatial fusing to refine the output of typical pixel-wise background model - mixture of Gaussians (MoG) and employ a MRF-MAP scheme to make foreground-background classification using the spatial correlation. Experiments on several challenge sequences show that our method is able to yield significantly better results than the traditional ones and is compelling with existing state of the art background subtraction algorithms. Additionally, we proved our algorithm has linear running time complexity and any pixel-wise background model could be easily integrated into our spatial fusing scheme which greatly enhanced its scalabilities and applications. Lifeng Sun, Shiqiang Yang |
MMSP | 4 |
| 2009 | Adaptive mixture observation models for multiple object tracking
Peng Cui 0001, Lifeng Sun, Shiqiang Yang |
Sci. China Ser. F Inf. Sci. | 3 |
| 2009 | Adaptive data-driven parallelization of multi-view video coding on multi-core processor
Yi Pang, WeiDong Hu, Lifeng Sun, Shiqiang Yang |
Sci. China Ser. F Inf. Sci. | 4 |
| 2009 | A Framework for Heuristic Scheduling for Parallel Processing on Multicore Architecture: A Case Study With Multiview Video CodingabstractIn this paper, using the Intel multicore architectures and the emerging multiview video coding standard, we introduce a framework for performing analysis, simulation, and evaluation of heuristics scheduling algorithms for implementing computationally intensive algorithms on multicore processors. The framework allows for accurate and quantitative characterization of the performance of dynamic scheduling algorithms for multimedia applications on different multicore processors without actual implementation of the scheduling algorithm and application on the actual platform. Experimental results demonstrate the effectiveness and scalability of our framework. Yi Pang, Lifeng Sun, Jiangtao Wen, Fengyan Zhang, WeiDong Hu, Shiqiang Yang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2009 | Contextual Mixture TrackingabstractMultiple Object Tracking (MOT) poses three challenges to conventional well-studied Single Object Tracking (SOT) algorithms: 1) Multiple targets lead the configuration space to be exponential to the number of targets; 2) Multiple motion conditions due to multiple targets' entering, exiting and intersection make the prediction process degrade in precision; 3) Visual ambiguities among nearby targets make the trackers error prone. In this paper, we address the MOT problem by embedding contextual proposal distributions and contextual observation models into a mixture tracker which is implemented in a Particle Filter framework. The proposal distributions are adaptively selected by motion conditions of targets which are determined by context information, and the multiple features are combined according to their discriminative power between ambiguity prone objects. The induction of contextual proposal distribution and observation model can help to surmount the incapability of conventional mixture tracker in handling object occlusions, meanwhile retain its merits of flexibility and high efficiency. The final experiments show significant improvement in variable number objects tracking scenarios compared with other methods. Peng Cui 0001, Lifeng Sun, Fei Wang 0001, Shiqiang Yang |
IEEE Trans. Multim. | 4 |
| 2009 | A Trace-Driven Approach to Evaluate the Scalability of P2P-Based Video-on-Demand ServiceabstractPeer-to-Peer (P2P) networks have emerged as one of the most promising approaches to improve the scalability of Video-on-Demand (VoD) service over Internet. However, despite a number of architectures and streaming protocols have been proposed in past years, there is few work to study the practical performance of P2P-based VoD service especially in consideration of real user behavior which actually has significant impact on system scalability. Therefore, in this paper, we first characterize the user behavior by analyzing a large amount of real traces from a popular VoD system supported by the biggest television station in China, cctv.com. Then we ex-amine the practical scalability of P2P-based VoD service through extensive trace-driven simula-tion under a general system framework. The results show that P2P networks scale well in provid-ing VoD service under real user behavior by obtaining a considerable good cache hit ratio. Moreover, it is observed that adopting hard cache at client side help achieves better system scal-ability than that with soft cache. We also identify the impact of various aspects of user behavior upon system scalability through detailed simulation. We believe that our study will shine insight-ful light on the understanding of practical scalability of P2P-based VoD service and be helpful to future system design and optimization. Jian-Guang Luo, Qian Zhang 0001, Shiqiang Yang |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2009 | Optimizing the Throughput of Data-Driven Peer-to-Peer StreamingabstractDuring recent years, the Internet has witnessed a rapid growth in deployment of data-driven (or swarming based) peer-to-peer (P2P) media streaming. In these applications, each node independently selects some other nodes as its neighbors (i.e. gossip-style overlay construction), and exchanges streaming data with the neighbors (i.e. data scheduling). To improve the performance of such protocol, many existing works focus on the gossip-style overlay construction issue. However, few of them concentrate on optimizing the streaming data scheduling to maximize the throughput of a constructed overlay. In this paper, we analytically study the scheduling problem in data-driven streaming system and model it as a classical min-cost network flow problem. We then propose both the global optimal scheduling scheme and distributed heuristic algorithm to optimize the system throughput. Furthermore, we introduce layered video coding into data-driven protocol and extend our algorithm to deal with the end-host heterogeneity. The results of simulation with the real world traces indicate that our distributed algorithm significantly outperforms conventional ad hoc scheduling strategies especially in stringent buffer and bandwidth constraints. Meng Zhang 0001, Yongqiang Xiong, Qian Zhang 0001, Lifeng Sun, Shiqiang Yang |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2009 | Video collage: presenting a video sequence using a single image
Tao Mei 0001, Bo Yang 0008, Shiqiang Yang, Xian-Sheng Hua 0001 |
Vis. Comput. | 3 |
| 2008 | Adding Confidentiality to Pull-Based Peer-to-Peer Live StreamingabstractWhile pull-based peer-to-peer (P2P) live streaming prevails in recent years, an important but less studied problem is security. To achieve security, confidentiality of the streaming data is essential and the corresponding key management scheme must scale well without degrading the system performance too much. In this paper, we thus propose P2P-SMP -a novel secure multicast protocol specifically designed for pull-based P2P live streaming -to achieve good reliability in key distribution and guarantee that the decryption key is received by legitimate peers before the corresponding encrypted streaming data even in dynamic P2P environments. Through extending the pull-based P2P live streaming protocol to distribute rekeying messages and further adopting a batch rekeying method, P2P-SMP dramatically reduces the communication and computing overhead in key distribution. The efficiency and reliability of P2P-SMP is evaluated through extensive simulations. Jian-Guang Luo, Shiqiang Yang |
CCNC | 3 |
| 2008 | iGridMedia: Providing Delay-Guaranteed Peer-to-Peer Live Streaming Service on InternetabstractAlmost all existing peer-to-peer (P2P) streaming systems can only provide non-interactive service on Internet. In this paper, we focus on providing delay-guaranteed service to support delay-tolerant interactive applications (such as online auction, person interview, video sharing & commenting, etc.) using P2P streaming technology. In an interactive channel, all or part of the participants can interact with the presenter or publisher at the source. Unlike existing P2P streaming service, there is different level of delay and synchronization requirement in the interactive applications, meanwhile, the participant number in each channel is relatively small but the concurrent channel number is large. With such challenges, how much server bandwidth can be saved in P2P streaming is an interesting problem. We propose a very practical protocol iGridMedia and fully implement it. The basic tradeoff between consumed server bandwidth and the required delay is carefully studied. Simulation and real-world experiment show that our system consumes only about 4 times streaming rate with 5 sec guaranteed delay even if the peer churn rate is high. So far as we know, our work represents the state-of- the-art approach that using P2P technology to support emergent interactive live applications on Internet with delay-guaranteed requirement. Meng Zhang 0001, Lifeng Sun, Xiaolu Xi, Shiqiang Yang |
GLOBECOM | 4 |
| 2008 | Topic mining on web-shared videosabstractInternet videos have grown exponentially with the help from video sharing Websites. Automatic topic mining is therefore increasingly important for organizing and navigating such large video databases. Most of current solutions of topic detection and mining were done on news videos and cannot be directly applied on Web videos, because of their limited and noisy semantic information. In this paper, we will try to address this problem and propose an automatic topic mining framework on Web videos. We develop an iterative weight-updated co-clustering scheme to filter "noisy" tags and mine the "hot" topics. We then propose a visual-based clustering approach to further group the videos with similar content, and rank the visual-similar groups by their similarity to the topic center. Experiments on a large Web video database demonstrate the superior performance of our weight-updated co-clustering to both of the traditional co-clustering and k-means. The experiments also demonstrate significant improvement of users' experience by our visual-based clustering and ranking. Lu Liu 0005, Yong Rui, Lifeng Sun, Bo Yang 0008, Jianwei Zhang 0001, Shiqiang Yang |
ICASSP | 6 |
| 2008 | A Joint Matrix Factorization Approach to Unsupervised Action CategorizationabstractIn this paper, a novel unsupervised approach to mining categories from action video sequences is presented. This approach consists of two modules: action representation and learning model. Videos are regarded as spatially distributed dynamic pixel time series, which are quantized into pixel prototypes. After replacing the pixel time series with their corresponding prototype labels, the video sequences are compressed into 2D action matrices. We put these matrices together to form an multi-action tensor, and propose the joint matrix factorization method to simultaneously cluster the pixel prototypes into pixel signatures, and matrices into action classes. The approach is tested on public and popular Weizmann data set, and promising results are achieved. Peng Cui 0001, Fei Wang 0001, Lifeng Sun, Shiqiang Yang |
ICDM | 4 |
| 2008 | Fast and robust detection of near-duplicates in web video database
Lu Liu 0005, Lifeng Sun, Shiqiang Yang |
ICME | 4 |
| 2008 | Free-Shaped Video Collage
Bo Yang 0008, Tao Mei 0001, Lifeng Sun, Shiqiang Yang, Xian-Sheng Hua 0001 |
MMM | 4 |
| 2008 | Highlight detection in soccer video using web-casting textabstractHighlight detection is a challenge task in soccer video analysis. Using Web-casting text as external knowledge is proved to be a short cut to achieve both efficiency and effectiveness. Based on the previous framework using Web-casting text, we have improved the processes of video time detection and highlight boundary detection. Our method can detect the transparent time bar and can achieve acceptable precision in highlight boundary detection though the Web text time is not accurate at all. This progress can make the framework more robust in practice. Xi-Feng Ding, Yin-Jun Miao, Lifeng Sun, Shiqiang Yang |
MMSP | 5 |
| 2008 | CCL-SVC: Optimizing user experience of broadcasting video on computation capability limited handheld devicesabstractIn this paper, we propose a novel scheme using computing complexity layered scalable video coding (CCLSVC) to optimize the user experience of broadcasting video in the computing capability limited handheld terminals. To address the heterogeneity of computing capability among different handheld devices, we employ hierarchal B reference structure of SVC to divide the frames into multiple computing complexity layers (CC Layers) in server side. The handheld clients simply choose to decode the frames in their corresponding layers in terms of their computation capability to maximize the video PSNR. We have proved that the optimal CC Layers division problem is a precedence constrained scheduling problem, which is an NP-complete problem. And we further propose our fast greedy method to approximately get optimized broadcasting video playback PSNR. The simulation shows that our method is superior to temporal SVC and random frame discarding method. Jingyuan Wang 0001, Lifeng Sun, Bin Li 0088, Meng Zhang 0001, Shiqiang Yang |
MMSP | 5 |
| 2008 | An adaptive frame interpolation algorithm using statistic analysis of motions and residual energyabstractIn this paper, a new motion-compensated frame interpolation (MCFI) algorithm using adaptive criterions to correct the motion vector field is proposed. First, an effective pre-processing scheme is done to the transmitted motion vectors. Then, unlike the conventional MCFI algorithms using fixed criterion to analyse the reliability of motion vectors, we proposed to select the parameters and thresholds by analysing the statistical characterization of motion vectors and residual energy, thus thresholds can be changed adaptively during the decoding process. Meanwhile, our new criterions consider both reliability of motion vectors and the smoothness of the region, which avoid the unnecessary motion estimation and reduce the complexity. Experimental results show that the proposed algorithm has 1.4~5.5 dB increase comparing with vector median filter algorithm in average PSNR, and greatly improves the subjective visual quality. Chunbo Yang, Pin Tao, Shiqiang Yang |
MMSP | 3 |
| 2008 | Web video topic discovery and tracking via bipartite graph reinforcement modelabstractAutomatic topic discovery and tracking on web-shared videos can greatly benefit both web service providers and end users. Most of current solutions of topic detection and tracking were done on news and cannot be directly applied on web videos, because the semantic information of web videos is much less than that of news videos. In this paper, we propose a bipartite graph model to address this issue. The bipartite graph represents the correlation between web videos and their keywords, and automatic topic discovery is achieved through two steps - coarse topic filtering and fine topic re-ranking. First, a weight-updating co-clustering algorithm is employed to filter out topic candidates at a coarse level. Then the videos on each topic are re-ranked by analyzing the link structures of the corresponding bipartite graph. After the topics are discovered, the interesting ones can also be tracked over a period of time using the same bipartite graph model. The key is to propagate the relevant scores and keywords from the videos of interests to other relevant ones through the bipartite graph links. Experimental results on real web videos from YouKu, a YouTube counterpart in China, demonstrate the effectiveness of the proposed methods. We report very promising results. Lu Liu 0005, Lifeng Sun, Yong Rui, Shiqiang Yang |
WWW | 5 |
| 2007 | Automatic Player Detection, Labeling and Tracking in Broadcast Soccer VideoabstractAutomatic player detection, labeling and tracking in broadcast soccer video are significant while quite challenging tasks. In this paper, we present a solution to perform automatic multiple player detection, unsupervised labeling and efficient tracking. Players ’ position and scale are determined by a boosting based detector. Players ’ appearance models are unsupervised learned from hundreds of samples automatically collected by detection. Thereafter, these models can be utilized for player labeling (Team A, Team B and Referee). Player tracking is achieved by Markov Chain Monte Carlo (MCMC) data association. Some data driven dynamics are proposed to improve the Markov chain’s efficiency. The testing results on FIFA World Cup 2006 video demonstrate that our method can reach high detection and labeling precision, and reliably tracking in cases of scenes such as multiple player occlusion, moderate camera motion and pose variation. 1 Jia Liu 0001, Xiaofeng Tong, Wenlong Li 0003, Tao Wang 0003, Yimin Zhang 0002, Bo Yang 0008, Lifeng Sun, Shiqiang Yang |
BMVC | 9 |
| 2007 | A Sequential Monte Carlo Approach to Anomaly Detection in Tracking Visual EventsabstractIn this paper we propose a technique to detect anomalies in individual and interactive event sequences. We categorize anomalies into two classes: abnormal event, and abnormal context, and model them in the Sequential Monte Carlo framework which is extended by Markov Random Field for tracking interactive events. Firstly, we propose a novel pixel-wise event representation method to construct feature images, in which each blob corresponds to a visual event. Then we transform the original blob-level features into subspaces to model probabilistic appearance manifolds for each event-class. With the probability of an observation associated with each event-class (or state) derived from probabilistic manifolds, and state transitional probability, the prior and posterior state distributions can be estimated. We demonstrate in experiments that the approach can reliably detect such anomalies with low false alarm rates. Peng Cui 0001, Lifeng Sun, Shiqiang Yang |
CVPR | 4 |
| 2007 | On Real-Time Detecting Duplicate Web VideosabstractWith the rapid development of telecommunication techniques and digital devices, it is quite easy to copy, modify and republish videos in digital format, resulting in large volume of duplicate videos on the Web in recent years. In this paper we mainly investigate the problem of detecting excessive content duplication, so as to facilitate video search and intelligence propriety protection. A real-time detection method is hence proposed, which first selects videos' representative frames and then reduces each to a 64 bit hash code. Then the similarity of any two videos can be estimated by the proportion of their similar hash codes. The experiments demonstrate that our approach is both efficient and effective in terms of real-time applications. Lu Liu 0005, Xian-Sheng Hua 0001, Shiqiang Yang |
ICASSP (1) | 4 |
| 2007 | An Experimental Study on Cheating and Anti-Cheating in Gossip-Based ProtocolabstractThe Internet has witnessed a rapid growth in deployment of gossip-based protocol in many multicast applications. In a typical gossip-based protocol, each node independently exchanges data with its neighbors, acting as dual roles of receiver and sender to facilitate scalability and resilience. However, most of previous work in this literature seldom considered cheating issue of end users, which is also very important in face of the fact that the mutual cooperation inherently determines overall system performance. In this paper, we mainly investigate the dishonest behaviors in decentralized gossip-based protocol through extensive experimental study. Our original contributions come in two-fold: In the first part of cheating study, we analytically discuss two typical cheating strategies, that is, intentionally increasing subscription requests and untruthfully calculating forwarding probability, and further evaluate their negative impacts. The results indicate that more attention should be paid on defending cheating behaviors in gossip-based protocol. In the second part of anti-cheating study, we propose a simple receiver-driven measurement mechanism, which evaluates individual forwarding traffic from the perspective of receivers and thus identifies cheating nodes with high incoming/outgoing ratio. The experiments under various conditions show that it performs quite well in case of serious cheating and achieves considerable performance in other cases. Yuanchun Shi, Shiqiang Yang, Yuzhuo Zhong |
ICC | 4 |
| 2007 | Generation of Layered Depth Images from Multi-View VideoabstractThe feature of rendering arbitrary view make layered depth image (LDI) be suitable to present multi-view video data and provide interactivity form such as free view video(FVV). However, visual artifacts occurred in the rendered result limit the practicality of LDI. In this paper, we proposed an approach to deal with this problem during generating of LDI, which refines colors of the depth pixels by choosing proper candidate depth pixels, removes matting effects by projecting LDI backward to reference views, and eliminate the gaps or holes due to disocclusion and undersample by local background interpolation. Experimental results show that our approach is practicable and efficient. Lifeng Sun, Shiqiang Yang |
ICIP (5) | 3 |
| 2007 | A Novel Event-Oriented Segment-of-Interest Discovery Method for Surveillance VideoabstractDuring recent years, the quick development of computer techniques has witnessed the ever-increasing surveillance video data, which essentially pose great challenge on the data storage, management, analysis and even retrieval. Considering that most of the high volume of data is with no interest, we mainly investigate the problem of effectively and efficiently discovering segments-of-interest (SoI) in this paper. To do so, we propose a novel event-oriented Sol discovery method in two steps: first, we represent an event by modeling pixels' change in temporal-spatial space, aiming to unify both the inter-frames and frames-background changes; second, with the benefit of unsupervised learning, the prototype-event models could be learned from these detected events and in turn exploited to measure the interest factor of each prototype-event. The experiment results demonstrate that the proposed method precisely discriminate different events and effectively discover Sols. Peng Cui 0001, Lifeng Sun, Zhi Wang 0001, Shiqiang Yang |
ICME | 4 |
| 2007 | Longer, Better: On Extending User Online Duration to Improve Quality of Streaming Service in P2P NetworksabstractDuring recent years, the Internet has witnessed a rapid growth in deployment of peer to peer (P2P) based live media streaming systems. In this paper, we are motivated to investigate the problem of improving the streaming quality of service (QoS) of those systems by two universal recognitions: one is that the understanding to practical service experiences would benefit the performance enhancement, while the other is "the more, the better" design philosophy in P2P networks. As an interesting approach, we first retrieve the practical online duration information of end users from service traces in our GridMedia system, and then intentionally extend the online duration of each peer, which would result in a more stable peer overlay. The comparative simulations with original real-world traces and modified ones demonstrate that the quality of streaming service would be better if peers stay longer in the community. We claim that, this simple but positive result essentially validates both the influence of end users' behaviors and the need of an incentive to encourage users to spend more time in P2P live streaming systems. Lifeng Sun, Kaiyun Zhang, Shiqiang Yang, Yuzhuo Zhong |
ICME | 4 |
| 2007 | Video Histogram: A Novel Video Signature for Efficient Web Video Duplicate Detection
Lu Liu 0005, Xian-Sheng Hua 0001, Shiqiang Yang |
MMM (2) | 4 |
| 2007 | Characterizing User Behavior Model to Evaluate Hard Cache in Peer-to-Peer Based Video-on-Demand Service
Jian-Guang Luo, Meng Zhang 0001, Shiqiang Yang |
MMM (2) | 4 |
| 2007 | Optimization of System Performance for DVC Applications with Energy Constraints over Ad Hoc Networks
Lifeng Sun, Ke Liang 0001, Shiqiang Yang, Yuzhuo Zhong |
MMM (2) | 3 |
| 2007 | Improving Quality of Live Streaming Service over P2P Networks with User Behavior Model
Lifeng Sun, Jian-Guang Luo, Shiqiang Yang, Yuzhuo Zhong |
MMM (2) | 4 |
| 2007 | Optimizing the Throughput of Data-Driven Based Streaming in Heterogeneous Overlay Network
Meng Zhang 0001, Chunxiao Chen, Yongqiang Xiong, Qian Zhang 0001, Shiqiang Yang |
MMM (1) | 5 |
| 2007 | A multi-view video coding approach using Layered Depth ImageabstractMulti-view video introduces more forms of interactivity and can potentially be used for a variety of applications, such as free-view video (FVV) and three-dimensional video (3DTV). However, the huge volume video data and the synthesis of virtual view limit the practicality of multi-view video. In this paper, we propose an approach to multi-view video coding, in which the mutli-view video sequences are converted to layered depth images (LDIs) to represent scenes, and compressed layer by layer. To achieve better compression efficiency, we restructure layered images according depth data in LDIs. Experimental results show that our approach is practicable and efficient to meet the compression efficiency and flexible interactivity. Lifeng Sun, Shiqiang Yang |
MMSP | 3 |
| 2007 | Power-Rate-Distortion Optimization for Multi-Source Video Streaming under Energy Constraints over Ad Hoc NetworksabstractWe propose a dynamic rate allocation scheme based on power-rate-distortion (PRD) optimization model among multiple video sources over ad hoc networks. This work is an extension of the PRD model for single source video streaming. With a total rate constraint and different power consumption constraints for each node, our optimization algorithm minimizes the average video distortion for all sources. The optimization is performed at the receiver of the video streams. Experimental results for a video surveillance scenario demonstrate that the proposed scheme outperforms a simple fixed-QP scheme. The proposed scheme promotes the average PSNR by 0.32-0.45 dB without shortening the system lifetime, or prolongs system lifetime by more than 20% without cutting down the overall PSNR. Juntao Ouyang, Lifeng Sun, Yuzhuo Zhong, Shiqiang Yang |
MMSP | 4 |
| 2007 | Spatial and Temporal Data Parallelization of Multi-view Video Encoding AlgorithmabstractMulti-view video coding technology is proposed to resolve the problem of huge data storage and transmission for free-view and 3D interactive video. How to support real time multi-view video encoding which has high computing complexity with sharply increased multi-view video data is essential In this paper, we proposed a solution of spatial and temporal data parallelization for multi-view video encoding algorithm based on IBM cell multiprocessor system using selections of optimal theories & methods. The performance of our tasks distributing scheme is eight times faster than the serial algorithm, speedup is notable. Yi Pang, Lifeng Sun, Songliu Guo, Shiqiang Yang |
MMSP | 4 |
| 2007 | VideoReach: an online video recommendation systemabstractThis paper presents a novel online video recommendation system called VideoReach, which alleviates users' efforts on finding the most relevant videos according to current viewings without a sufficient collection of user profiles as required in traditional recommenders. In this system, video recommendation is formulated as finding a list of relevant videos in terms of multimodal relevance (i.e. textual, visual, and aural relevance) and user click-through. Since different videos have different intra-weights of relevance within an individual modality and inter-weights among different modalities, we adopt relevance feedback to automatically find optimal weights by user click-though, as well as an attention fusion function to fuse multimodal relevance. We use 20 clips as the representative test videos, which are searched by top 10 queries from more than 13k online videos, and report superior performance compared with an existing video site. Tao Mei 0001, Bo Yang 0008, Xian-Sheng Hua 0001, Linjun Yang, Shiqiang Yang, Shipeng Li 0001 |
SIGIR | 5 |
| 2007 | Exploiting multi-scale support vector regression for image compression
Bin Li 0088, Danian Zheng, Lifeng Sun, Shiqiang Yang |
Neurocomputing | 4 |
| 2007 | Understanding the Power of Pull-Based Streaming Protocol: Can We Do Better?abstractMost of the real deployed peer-to-peer streaming systems adopt pull-based streaming protocol. In this paper, we demonstrate that, besides simplicity and robustness, with proper parameter settings, when the server bandwidth is above several times of the raw streaming rate, which is reasonable for practical live streaming system, simple pull-based P2P streaming protocol is nearly optimal in terms of peer upload capacity utilization and system throughput even without intelligent scheduling and bandwidth measurement. We also indicate that whether this near optimality can be achieved depends on the parameters in pull-based protocol, server bandwidth and group size. Then we present our mathematical analysis to gain deeper insight in this characteristic of pull-based streaming protocol. On the other hand, the optimality of pull-based protocol comes from a cost -tradeoff between control overhead and delay, that is, the protocol has either large control overhead or large delay. To break the tradeoff, we propose a pull-push hybrid protocol. The basic idea is to consider pull-based protocol as a highly efficient bandwidth-aware multicast routing protocol and push down packets along the trees formed by pull-based protocol. Both simulation and real-world experiment show that this protocol is not only even more effective in throughput than pull-based protocol but also has far lower delay and much smaller overhead. And to achieve near optimality in peer capacity utilization without churn, the server bandwidth needed can be further relaxed. Furthermore, the proposed protocol is fully implemented in our deployed GridMedia system and has the record to support over 220,000 users simultaneously online. Meng Zhang 0001, Qian Zhang 0001, Lifeng Sun, Shiqiang Yang |
IEEE J. Sel. Areas Commun. | 4 |
| 2007 | A comparative study of Minimax Probability Machine-based approaches for face recognition
Johnny K. C. Ng, Yuzhuo Zhong, Shiqiang Yang |
Pattern Recognit. Lett. | 3 |
| 2007 | Investigation on unsupervised clustering algorithms for video shot categorization
Peng Wang 0003, Shiqiang Yang |
Soft Comput. | 3 |
| 2006 | TPOD: A Trust-Based Incentive Mechanism for Peer-to-Peer Live Broadcasting
Lifeng Sun, Jian-Guang Luo, Shiqiang Yang, Yuzhuo Zhong |
ATC | 4 |
| 2006 | On the Optimal Scheduling for Media Streaming in Data-driven Overlay NetworksabstractThe Internet has witnessed a rapid growth in deployment of data-driven overlay network (DON) based streaming applications during recent years. In these applications, each node independently selects some other nodes as its neighbors (i.e. overlay construction), and exchanges streaming data with these neighbors (i.e. data scheduling). This scheme improves the robustness of the system. However, most of the work in the literature focused on the construction problem, and very few addressed its scheduling problem which is also very important for the overall performance. In this paper, we analytically study the scheduling problem in DON and model it as a classical min-cost network flow problem. We then propose both the global optimal scheduling scheme and distributed heuristic algorithm to maximize the system throughput. Experimental results indicate that our algorithms outperform other schemes and the throughput gain is up to 80%. Meng Zhang 0001, Yongqiang Xiong, Qian Zhang 0001, Shiqiang Yang |
GLOBECOM | 4 |
| 2006 | Optimized Rate Allocation for Unbalanced Multiple Description Video Coding Over Unreliable Packet NetworkabstractVideo transmission over unreliable packet network is in general hampered by the packet losses and constraint by stringent playback deadline. With these two key factors in consideration, multiple description coding (MDC), comprising balanced and unbalanced MDC has been proposed as an error-robust source coding technique. Recently, transmitting multiple descriptions over a single path is interesting due to the unavailability of multiple independent paths. Therefore, in this paper, we investigate the problem of rate allocation for the high-resolution (HR) and low-resolution (LR) descriptions in UMDC transmission over single path. We first propose an approximate but efficient rate allocation model with the aid of two-state Markov link model and a simple distortion model at the sender side. Then we conduct extensive experiments to verity the proposed model and more excitedly the simulation results clearly demonstrate the effectiveness of proposed model Bin Li 0088, Lifeng Sun, Shiqiang Yang |
ICME | 4 |
| 2006 | An Unbalanced Multiple Description Coding Scheme for Video Transmission Over Wireless Ad Hoc NetworksabstractVideo transmission over wireless ad hoc networks is hampered by packet losses. Even a single packet loss may cause error propagation until an intra-coded frame is received. Indeed, packet losses greatly degrade the video quality. In this paper, we propose an unbalanced multiple description coding (UMDC) scheme over a single path which requires only one single path as additional links are difficult to be guaranteed in reality and is capable of quickly recovering from packet losses and ensuring continuous playback. The proposed scheme uses two descriptions, the high-resolution (HR) description and the low-resolution (LR) one. It uses the 'peg frames' to limit error propagation in the HR description. The two descriptions can help each other recover from packet losses. The simulation results show that the proposed UMDC scheme over a single path has a comparable performance with our UMDC scheme with multiple path transmission (MPT) and has a better viewing experience than the state-of-the-art FEC-based scheme Bin Li 0088, Lifeng Sun, Shiqiang Yang |
ICME | 4 |
| 2006 | Evaluation of Practical Scalability of Overlay Networks in Providing Video-on-Demand ServiceabstractRecently, overlay networks have been proposed to address the problem of scalability in providing video-on-demand (VoD) service. However, from the perspective of service providing, their efficiency has not been carefully studied and still remains far from clear, especially considering the impacts of user interactivities and in the case of multiple files with different and varying popularities on sharing. Towards this end, in this paper, by analyzing more than 20,000,000 real workload traces, we first identify two practical factors which we believe have determinant impacts on the scalability: user interactivities and popularity differences among files. Then we further evaluate cache-and-relay (CR), a representative scheme of overlay networks, with the real workload traces. Simulation results show that CR only saves about half of the server bandwidth even when there is no buffer constraint at clients, not so scalable as our original expectation Jian-Guang Luo, Shiqiang Yang |
ICME | 4 |
| 2006 | A Novel Distributed and Practical Incentive Mechanism for Peer to Peer Live Video StreamingabstractThe successful deployment of peer-to-peer (P2P) live video streaming systems has practically demonstrated that it can scale to reliably support a large population of peers. However, peers, representing rational end users, tend to be non-cooperative when it comes to the duty rather than the self-interests, running counter to the fundamental design philosophy of P2P concept. The objective of this paper is to investigate the problem of encouraging users to balance what they take from the system with what they contribute. We first make a statistical analysis to the service logs of a practical P2P live video streaming system and reveal intrinsic characteristic of users' online duration. Second, we thus propose a novel incentive mechanism based on the composite contributions which consist of two objective metrics, i.e. the on-line duration and effective upstream traffic. This mechanism offers service differentiation to users with different contributions and has some desirable properties: (1) distributed nature upon gossip-based overlay structure and (2) practical oriented evaluation criteria. The experiment results over PlanetLab further verify its effectiveness Lifeng Sun, Meng Zhang 0001, Shiqiang Yang, Yuzhuo Zhong |
ICME | 4 |
| 2006 | SIKAS: A Scalable Distributed Key Management Scheme for Dynamic Collaborative GroupsabstractThe increasing popularity of distributed and collaborative applications prompts the need for secure communication in collaborative groups. Some distributed collaborative key management protocols have been proposed to provide group communication privacy and data confidentiality for collaborative groups. However, most of them rekey on each member change, and the costs of group rekeying can be quite substantial for large groups with frequent membership changes. In this paper, we propose a scalable distributed key management scheme using a distributed one-way function tree named SIKAS which can significantly reduce the computation and communication costs of maintaining the group key based upon period-based group rekeying. A comparison with previous work has shown that SIKAS provides scalability and rekeying efficiency while preserving both distributed and collaborative properties Jian-Guang Luo, Bin Li 0088, Shiqiang Yang |
ICME | 4 |
| 2006 | On deployment of overlay network for live video streamingabstractOverlay network has emerged as a popular alternative to traditional client-server architecture for the distribution of media content. Intrinsic to the notion is a 'self-growing' community of end users. It is thus important to investigate practical issues in reality rather than following prescribed theoretical assumptions. This paper presents our research experiences on deployment of overlay network for live video streaming over Internet. In our implementation we adopt a gossip-based overlay construction protocol to accommodate topology changes and a composite scheduling mechanism to stream video contents. More importantly, with the service to totally more than 500,000 users and maximum 15,239 concurrent users provided by one common streaming server, we then exhibit insightful statistical results to reveal system performance and nodes properties in terms of capacity evolution, quality of service, request rate and online duration Lifeng Sun, Meng Zhang 0001, Shiqiang Yang, Yuzhuo Zhong |
ISCAS | 4 |
| 2006 | Bit rate reduction of H.264/AVC video coding for mobile applicationsabstractThe increasing use of mobile applications requires higher compression efficiency for video coding. However, in mobile applications, the capability of the processor is limited and the channel is unreliable. Since the emerging H.264/AVC video coding standard provides superior compression performance compared to all the existing video coding standards, many research efforts concentrate on the compression efficiency of intra-frame coding in H.264/AVC. In this paper, we propose a novel algorithm to reduce the bit rate compressed by H.264/AVC intra-frame coding. We focus on the method to code the intra-frame prediction modes. Our proposed method significantly reduces the bits used to code the modes and do not change the syntax of the bit stream. Experimental results demonstrate it is effective and efficient Bin Li 0088, Lifeng Sun, Shiqiang Yang |
MMM | 3 |
| 2006 | Flexible video coding scheme for content-based video storage and retrievalabstractIn this paper, we proposed a flexible video coding scheme for content-based video storage and retrieval. The proposed flexible video coding scheme applies shot clustering on the uncompressed video in the first pass and encodes the video in multi-clusters based on different video contents in the second pass. The compressed video can be flexibly decoding to meet the different requirements such as video indexing, video summary and structure extraction without fully decoding and re-analyzing. Also the flexible encoder can achieve global coding optimization by encoding similarity video shots continuously. It is very important to reduce the process complexity at the decoder and provide the flexible video content access on the compressed videos in many off-line video storage and retrieval applications. The experiments on sport and film videos illustrate the coding performance increase and the effect of the proposed coding scheme for content-based video storage. Jianning Zhang, Lifeng Sun, Shiqiang Yang, Yuzhuo Zhong |
MMM | 3 |
| 2006 | Chasing: An Efficient Streaming Mechanism for Scalable and Resilient Video-on-Demand Service over Peer-to-Peer Networks
Jian-Guang Luo, Shiqiang Yang |
Networking | 3 |
| 2006 | Enhancements of Representation and Interactivity for Multi-view Video Based on Layered Depth ImageabstractThe feature of rendering arbitrary view make layered depth image (LDI) be able to present multi-view video data and provide enhanced interactivity, such as free view video (FVV). However, visual artifacts occurred in the rendered result limit the practicality of LDI. In this paper, we proposed an approach to deal with this problem, which refines colors of the depth pixels by clustering to remove matting effects during generating of LDI, and applies median filter to rendered result to remove the gaps or holes caused by disocclusion and undersample. Experimental results show that our approach is practicable and efficient. Lifeng Sun, Shiqiang Yang |
SMC | 3 |
| 2006 | Fast Arc Detection Algorithm for Play Field Registration in Soccer Video MiningabstractThis paper presents an LSF-based framework for detecting arcs in broadcast soccer video. The successful identification of the arcs will evidently facilitate the soccer video analysis. The existing methods are not available for all playfield arcs including both the middle field circle and penalty box arcs. A new algorithm is proposed in the paper improved from LSF, called ALSF (advanced least square fitting), which can be used to detect arcs even though they are only 1/4 of eclipses (the same as penalty box arcs). With the improvement, we proposed a new framework to detect the arcs in broadcast soccer video. The proposed method first removes the points of straight lines. Then, all points of each connective area are transformed into a new reference frame to use LSF to get the equations of arcs. Experiments on more than 3 hours broadcast soccer video show the proposed method is effective with above 98% precision and 70% recall. Fei Wang 0001, Lifeng Sun, Bo Yang 0008, Shiqiang Yang |
SMC | 4 |
| 2006 | Mid-Level Descriptors Extraction of Soccer Video with Domain KnowledgeabstractIn this paper, we propose a method for two typical mid-level descriptors extraction of soccer video: view type and playfield position. The output reveals high-level structure of the soccer video. After the robust playfield segmentation, we use the grass-area-ratio and the boundary approximate lines between playfield and non-playfield to classify each frame into three kinds of view types according to a series of rules. We then classify frames captured by the main camera into fifteen kinds of positions by two steps: first, we classify the frames into five kinds according to the orientation of boundary lines; secondly, we use lines detection to further determine the specific playfield position. Experiments on three hours soccer videos show promising results. Bo Yang 0008, Lifeng Sun, Fei Wang 0001, Peng Wang 0003, Shiqiang Yang |
SMC | 5 |
| 2006 | The Analysis of Offloading H.264 Video Encoder on Mobile Devices for Energy SavingabstractAs more and more people want to communicate with each other vividly, video encoder application will be very popular on mobile devices. But it is also a resource-killer application. Meanwhile the mobile devices are powered by battery, and energy resource is very important for mobile devices. So when using this application on mobile devices, we must consider energy conserving. The computation offloading method is to offload some computations of application from mobile device to powerful server around mobile device via wireless network. So it can relieve the resource constraint and save energy. In this paper, we use the computation offloading method to H.264 video encoder on mobile device and propose three offloading schemes. Then we apply them to encode different video sequences. Finally, we analyze the efficiency of offloading schemes with different encoder parameters and different video sequences. Pin Tao, Shiqiang Yang |
SMC | 3 |
| 2005 | Position prediction motion-compensated interpolation for frame rate up-conversion using temporal modelingabstractIn this paper, a new motion-compensated interpolation method (MCI) for frame rate up-conversion is proposed. In conventional MCI algorithms using transmitted motion vectors (MVs), the block artifacts are caused by the unreliable MVs from the encoder. In our proposed method, motion vectors correction (MVC) is applied on the transmitted MVs before interpolation considering the spatial correlation. In consideration of the temporal correlation, generative model is used to generate the temporal-predicted MVs. Then position prediction method is applied on the transmitted MVs and the temporal-predicted MVs separately to predict the positions the interpolated blocks really move to, which makes the MVs used for interpolation nearer to the true motion. The final MVs are selected from the two candidates of position prediction by the minimum sum of absolute difference (SAD). Applied to the H.264 decoder, our proposed method achieves significant increase on PSNR and obvious decrease of the block artifacts. Jianning Zhang, Lifeng Sun, Shiqiang Yang, Yuzhuo Zhong |
ICIP (1) | 3 |
| 2005 | License management scheme with anonymous trust for digital rights managementabstractOne of the major issues raised by digital rights management systems concerns the protection of the user's privacy and anonymous consumption of content. However, most existing license management schemes for DRM systems do not support the protection of user privacy. Moreover, some other schemes such as PrecePt can only bind the license with a specified device though they concern privacy protection. In this paper, we propose a license management scheme named LMSAT (License Management Scheme with Anonymous Trust) which provides a more powerful and flexible license acquisition and usage tracking scheme to allow the user access the contents anytime, anywhere, and on any compliant devices anonymously. Bin Li 0088, Li Zhao 0006, Shiqiang Yang |
ICME | 4 |
| 2005 | Joint Inter and Intra Shot Modeling for Spectral Video Shot ClusteringabstractThis paper proposed a novel video shot clustering algorithm using spectral method by joint modeling of inter and intra shot. Gauss Mixture Model (GMM) is used for probabilistic space-time modeling of intra-shot pixels. The spectral clustering method is applied on the GMM parameters. The problem of automatic model selection is currently an open issue for spectral method. Here we propose a novel automatic model selection based on joint inter-intra model optimization to achieve the global optimization in both model parameters and cluster numbers. We compare the proposed method with the conventional spectral method for sports video clustering. The simulation results show more accuracy on how many clusters and clustering results. Jianning Zhang, Lifeng Sun, Shiqiang Yang, Yuzhuo Zhong |
ICME | 3 |
| 2005 | Gridmedia: A Multi-Sender Based Peer-to-Peer Multicast System for Video StreamingabstractWe present a novel single source peer-to-peer multicast architecture called GridMedia which mainly consists of 1) multi-sender based overlay multicast protocol (MSOMP) and 2) multi-sender based redundancy-retransmitting algorithm (MSRRA). The MSOMP deploys mesh-based two-layer structure and groups all the peers into clusters with multiple distinct paths from the source root to each peer. To address the problem of long burst packet loss, the MSRRA is proposed at the sender peers to patch the lost packets by using receiver peer loss pattern prediction. Consequently, GridMedia provides a scalable and reliable video streaming system for a large and highly dynamic population of end hosts, and ensures the quality of service in terms of continuous playback, bandwidth demanding and low latency. A real experimental system based on GridMedia architecture has been constructed over CERNET and broadcasting TV programs for seven months. More than 140,000 end users have been attracted with almost 600 simultaneously being online at Aug 2004 during Athens Olympic Games Meng Zhang 0001, Li Zhao 0006, Jian-Guang Luo, Shiqiang Yang |
ICME | 5 |
| 2005 | A probabilistic template-based approach to discovering repetitive patterns in broadcast videosabstractThere are usually repetitive sub-segments in broadcast videos, which may be associated with high-level concepts or events, e.g., news footage, repeated scores in basketball. Unsupervised mining techniques provide generic solutions to discovering such temporal patterns in various video genres, which are currently the subject of great interests to researchers working on multimedia content analysis. In this paper, we propose a novel approach to automatically detecting repetitive patterns in a video stream. In this approach, a video stream is first transformed to a symbol sequence via the spectral clustering algorithm. After computing the transition probabilities of any two symbols in temporal evolution, we produce a set of probabilistic templates to characterize the patterns of potential interest. Finally, we verify each probabilistic template by measuring the similarities between the video sub-segments and the template. Evaluations on various sports videos show promising results. Peng Wang 0003, Shiqiang Yang |
ACM Multimedia | 3 |
| 2005 | A peer-to-peer network for live media streaming using a push-pull approachabstractIn this paper, we present an unstructured peer-to-peer network called GridMedia for live media streaming employing a push-pull approach. Each node in GridMedia randomly selects its neighbors in the overlay and uses push-pull method to fetch data from the neighbors. The pull mode in the unstructured overlay which is inherently robust can work well with the high churn rate in P2P environment while the push mode can efficiently reduce the accumulated latency observed at user nodes. A practical system based on this framework has been developed. And the performance evaluation of our system which is established on PlanetLab [8] demonstrates that the pull-push method in GridMedia achieves good qualities even in high group change rate. Furthermore, our system was adopted by CCTV to broadcast the Gala Evening for Spring Festival 2005 through the Internet and attracted more than 500,000 users all over the world at that night with the incredibly maximum concurrent users of 15,239. Meng Zhang 0001, Jian-Guang Luo, Li Zhao 0006, Shiqiang Yang |
ACM Multimedia | 4 |
| 2005 | Gridmedia: A Practical Peer-to-Peer Based Live Video Streaming SystemabstractIn this paper, we describe the design and implementation of a peer-to-peer based live video streaming system called Gridmedia. Gridmedia organizes the nodes into an unstructured overlay, and adopts a novel push-pull streaming mechanism to fetch data from the partner nodes. The pull mode in the unstructured overlay can work well with the high churn rate in P2P environment while the push mode can efficiently reduce the accumulated latency at user side. We also depict our practical solution of traversal over network address translators (NATs) and firewalls. Gridmedia was adopted by CCTV to broadcast the CCTV Spring Festival Gala 2005 through Internet and attracted more than 500,000 users all over the world during that night, with the maximum amount of concurrent users reaches as high as 15,239 Li Zhao 0006, Jian-Guang Luo, Meng Zhang 0001, Wen-Jie Fu 0001, Yi-Fei Zhang, Shiqiang Yang |
MMSP | 7 |
| 2005 | An HMM-based framework for video semantic analysisabstractVideo semantic analysis is essential in video indexing and structuring. However, due to the lack of robust and generic algorithms, most of the existing works on semantic analysis are limited to specific domains. In this paper, we present a novel hidden Markove model (HMM)-based framework as a general solution to video semantic analysis. In the proposed framework, semantics in different granularities are mapped to a hierarchical model space, which is composed of detectors and connectors. In this manner, our model decomposes a complex analysis problem into simpler subproblems during the training process and automatically integrates those subproblems for recognition. The proposed framework is not only suitable for a broad range of applications, but also capable of modeling semantics in different semantic granularities. Additionally, we also present a new motion representation scheme, which is robust to different motion vector sources. The applications of the proposed framework in basketball event detection, soccer shot classification, and volleyball sequence analysis have demonstrated the effectiveness of the proposed framework on video semantic analysis. Gu Xu, Yufei Ma 0006, HongJiang Zhang, Shiqiang Yang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2004 | Energy distributed update steps(edu) in lifting based motion compensated video codingabstractSubband video coding is an elegant scheme to fulfill high performance scalable video coding. In this paper, a new update scheme, energy distributed update steps (EDU), is proposed for the temporal transform in lifting based motion compensated video coding. The idea is to update where predict is made by distributing high-pass signals to the low-pass frame. The scheme avoids complex and inaccurate inversion of the motion information that used in the traditional update steps, thus it reduces computations in temporal transform. Experimental results show that the coding performances can be improved up to 0.77 dB. Jizheng Xu, Feng Wu 0001, Shiqiang Yang, Shipeng Li 0001 |
ICIP | 4 |
| 2004 | A pinhole camera modeling of motion vector field for tennis video analysisabstractThe motion vector field (MVF) represents motion characteristics in video sequences, and has been widely used and been proved to be effective in sports video analysis. However, in tennis video analysis, MVF-based methods are seldom utilized for the reasons that: (i) the players' true motion is not accurately represented by the extracted motion vector, due to the deformation caused by the diagonal shooting of the camera; and (ii) the motion vector's magnitude is not prominent enough and is prone to be disturbed by noises. In this paper, a pinhole camera modeling of the motion vector field is proposed to revise the deformed motion vector. In this modeling, a foreground object mask is adopted and global motion compensation is incorporated as the pre-processing steps. Evaluation of the proposed modeling using four hours of tennis videos shows very encouraging results. Peng Wang 0003, Rui Cai 0002, Bin Li 0088, Shiqiang Yang |
ICIP | 4 |
| 2004 | Contextual browsing for highlights in sports videoabstractA novel highlight browsing approach for sports videos, based on contextual cue categorization, is proposed in this paper. Contextual cues are those salient audio and visual cues around highlight segments, and they are also informative and attractive to audiences. With further analysis and categorization of these contextual cues, the user could access highlights more flexibly and effectively. An experimental system is built for broadcasting diving videos in our current work, in which contextual cues such as slow-motion replays and game statistics information are extracted and categorized for highlight browsing. An efficient algorithm for slow-motion replay detection and categorization in diving videos is also proposed. Experimental results on 2 hours broadcasting diving videos have demonstrated the effectiveness and efficiency of the proposed approach. Peng Wang 0003, Rui Cai 0002, Shiqiang Yang |
ICME | 3 |
| 2004 | An unsymmetrical-cross multi-resolution motion search algorithm for MPEG4-AVC/H.264 codingabstractThe new H.264 (MPEG-4 AVC) video-coding standard has a significant performance benefit compared to former standards. Unfortunately, those advanced coding features, including variable block-size motion compensation and multiple reference frames, incur a considerable increase in encoder complexity, especially when using the straightforward full search (FS) algorithm, mainly with regards to motion estimation (ME) and mode decision. We propose a new ME algorithm for fast H.264 coding, utilizing the technique of the generalized motion vector (MV) predictors, named hybrid unsymmetrical-cross multi-resolution grid search (UCMRGS), which takes advantage of the correlation among variable blocks and can considerably reduce the computational cost of ME at the encoder, while at the same time give similar, and, in some cases, better, visual quality compared with the brute force full search algorithm. The proposed algorithms mainly rely upon very robust and reliable predictive techniques with parameters adapted to the local characteristics combined with the UCMRGS pattern Yuwen He, Shiqiang Yang |
ICME | 3 |
| 2003 | A people similarity based approach to video indexingabstractThis paper presents a new approach to people-based video indexing. In this approach, we define a people-based similarity measure according to both clothing similarity and speaking voice similarity. Such similarity depicts how perceptually similar two people appearing in different scenes are and if they belong to an identical person. Instead of computing in feature space, the proposed people-based similarity is computed in distance space. The extended support vector machines (SVM) are employed to map a serial of low-level feature distances to a perceived people similarity. In order to build people-based video indexing, a novel unsupervised clustering algorithm is also proposed, which can more correctly identify an individual person according to mutual people similarities between two people. The experiments on large video testing data have demonstrated the effectiveness and efficiency of the proposed people-based similarity, unsupervised clustering and video indexing. Peng Wang 0003, Yufei Ma 0006, HongJiang Zhang, Shiqiang Yang |
ICASSP (3) | 4 |
| 2003 | A HMM based semantic analysis framework for sports game event detectionabstractVideo events detection or recognition is one of important tasks in semantic understanding of video content. Sports game video should be considered as a rule-based sequential signal. Therefore, it is reasonable to model sports events using hidden Markov models. In this paper, we present a generic, scalable and multilayer framework based on HMMs, called SG-HMMs (sports game HMMs), for sports game event detection. At the bottom layer of this framework, event HMMs output basic hypotheses based on low-level features. The upper layers are composed of composition HMMs, which add constraints on those hypotheses of the lower layer. Instead of isolated event recognition, the hypotheses at different layers are optimized in a bottom-up manner and the optimal semantics are determined by top-down process. The experimental results on basketball and volleyball videos have demonstrated the effectiveness of the proposed framework for sports game analysis. Gu Xu, Yufei Ma 0006, HongJiang Zhang, Shiqiang Yang |
ICIP (1) | 4 |
| 2003 | Robust video transmission over lossy packet networks using block-based fine granularity scalable coding
Yuwen He, Shiqiang Yang |
VCIP | 2 |
| 2002 | Improved Fine Granular Scalable Coding with Inter-Layer PredictionabstractThis paper proposes an improved fine granular scalable (FGS) coding method with interlayer prediction. There are two important aspects to improving FGS coding efficiency. One is a low bit-rate video coding method and the other is interlayer prediction with enhancement layer reference at base-layer coding. The whole scalable coding performance with the proposed method is greatly enhanced over a wide bandwidth. New spatial and temporal prediction methods are investigated in order to increase FGS base-layer low bit-rate coding efficiency. There are nine kinds of spatial prediction modes for intra coding to exploit the pixels' spatial correlation, including DC and eight directional predictions, and multiple model-based motion prediction is utilized to predict a large irregular motion for inter coding, including a 6-parameter affine model. The reference for motion compensation is adaptively selected from base layer or enhancement layer according to their prediction error. The references are reconstructed at two layers. The references for prediction and reconstruction can be different in order to decrease drifting error due to bit-stream truncation. The coding efficiency of our base layer coding can be comparable with that of latest draft H.26L and holds a compelling improvement compared to MPEG-4. With our proposed scalable coding scheme, the whole FGS coding efficiency can be improved by about 2.0 dB at low bit-rate and 3.0-4.0 dB at medium or high bit-rate. The visual quality of the decoded video is also impressively improved at all decoded bit-rates. Yuwen He, Xuejun Zhao, Yuzhuo Zhong, Shiqiang Yang |
DCC | 4 |
| 2002 | Bilock-based fine granularity scalable video coding for content-aware streamingabstractVideo streaming is becoming more and more popular with widely used hybrid networks. The prime challenge of such applications is to deal with varying transmission bandwidth. This paper proposes a block-based fine granularity scalable (FGS) coding structure, which is a more flexible scalable video coding structure supporting content-aware streaming compared with the MPEG-4 FGS coding structure. The streaming server can conveniently implement content-aware rate allocation or content-based selective enhancement dynamically through user's interaction with the proposed scalable coding structure. Thus the streaming server can have a differentiated delivery strategy according to user's preference. However the uniform rate allocation for bit-stream truncation of the proposed block-based FGS will result in more than 2dB loss by PSNR compared with MPEG-4 FGS within quite a wide range of bit rates. A fast optimal rate allocation method is also proposed to solve this problem in this paper. The coding efficiency is improved, which can be comparable with MPEG-4 FGS coding and is even better (0.5dB) with some sequences at some bit rates. Yuwen He, Shiqiang Yang, Yuzhuo Zhong |
ICIP (2) | 2 |
| 2002 | Content-based selective enhancement layer dropping algorithm for FGS streaming using nearest feature line method
Li Zhao 0006, Shiqiang Yang, Yuzhuo Zhong |
VCIP | 3 |
| 2001 | Content-based retrieval of video shot using the-improved nearest feature line methodabstractShot-based classification and retrieval is very important for video database organization and access. We present a new approach: 'nearest feature line - NFL' used in shot retrieval. We look at key-frames in a shot as feature points to represent the shot in feature space. Lines connecting the feature points are further used to approximate the variations in the whole shot. The similarity between the query image and the shots in video database are measured by calculating the distance between the query image and the feature lines in feature space. To make it more suited to video data, we improved the original NFL method by adding constraints on the feature lines. Experimental results show that our improved NFL method is better than the traditional classification methods such as nearest neighbor (NN) and nearest center (NC). Li Zhao 0006, Stan Z. Li, Shiqiang Yang, HongJiang Zhang |
ICASSP | 4 |
| 2001 | Method Based On Temporal Constrain Shot Method Based On Temporal Constrain Shot Similarity
Li Zhao 0006, Shiqiang Yang |
ICME | 2 |
| 2000 | Region-Based Tracking in Video Sequences Using Planar Perspective Models
Yuwen He, Li Zhao 0006, Shiqiang Yang, Yuzhuo Zhong |
ICMI | 3 |
| 1998 | Fast road classification and orientation estimation using omni-view images and neural networksabstractThis paper presents the results of integrating omnidirectional view image analysis and a set of adaptive backpropagation networks to understand the outdoor road scene by a mobile robot. Both the road orientations used for robot heading and the road categories used for robot localization are determined by the integrated system, the road understanding neural networks (RUNN). Classification is performed before orientation estimation so that the system can deal with road images with different types effectively and efficiently. An omni-view image (OVI) sensor captures images with 360 degree view around the robot in real-time. The rotation-invariant image features are extracted by a series of image transformations, and serve as the inputs of a road classification network (RCN). Each road category has its own road orientation network (RON), and the classification result (the road category) activates the corresponding RON to estimate the road orientation of the input image. Several design issues, including the network model, the selection of input data, the number of the hidden units, and learning problems are studied. The internal representations of the networks are carefully analyzed. Experimental results with real scene images show that the method is fast and robust. Zhigang Zhu 0001, Shiqiang Yang, Guangyou Xu, Xueyin Lin, Dingji Shi |
IEEE Trans. Image Process. | 2 |