VLDB 2026 Research / reviewers in the wild / expert
Yanwei Liu 0001
dblp:49/3816-1
· DBLP profile ↗
69ranked-venue papers
16as first author
24since 2021 · last 2026
0000-0002-0626-1901ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 13 first-author · 13 since 2021Computer networks · 20 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DRGW: Learning Disentangled Representations for Robust Graph WatermarkingabstractGraph-structured data is foundational to numerous web applications, and watermarking is crucial for protecting their intellectual property and ensuring data provenance. Existing watermarking methods primarily operate on graph structures or entangled graph representations, which compromise the transparency and robustness of watermarks due to the information coupling in representing graphs and uncontrollable discretization in transforming continuous numerical representations into graph structures. This motivates us to propose DRGW, the first graph watermarking framework that addresses these issues through disentangled representation learning. Specifically, we design an adversarially trained encoder that learns an invariant structural representation against diverse perturbations and derives a statistically independent watermark carrier, ensuring both robustness and transparency of watermarks. Meanwhile, we devise a graph-aware invertible neural network to provide a lossless channel for watermark embedding and extraction, guaranteeing high detectability and transparency of watermarks. Additionally, we develop a structure-aware editor that resolves the issue of latent modifications into discrete graph edits, ensuring robustness against structural perturbations. Experiments on diverse benchmark datasets demonstrate the superior effectiveness of DRGW. Jiasen Li, Yanwei Liu 0001, Zhuoyi Shang, Xiaoyan Gu 0001, Weiping Wang 0005 |
WWW | 2 |
| 2026 | Sparse adversarial attack via robust attack points selection
Yanwei Liu 0001, Jianing Li 0001, Enci Liu, Yao Zhu 0003, Antonios Argyriou |
Pattern Recognit. | 2 |
| 2025 | Know2Vec: A Black-Box Proxy for Neural Network RetrievalabstractFor general users, training a neural network from scratch is usually challenging and labor-intensive. Fortunately, neural network zoos enable them to find a well-performing model for directly use or fine-tuning it in their local environments. Although current model retrieval solutions attempt to convert neural network models into vectors to avoid complex multiple inference processes required for model selection, it is still difficult to choose a suitable model due to inaccurate vectorization and biased correlation alignment between the query dataset and models. From the perspective of knowledge consistency, i.e., whether the knowledge possessed by the model can meet the needs of query tasks, we propose a model retrieval scheme, named Know2Vec, that acts as a black-box retrieval proxy for model zoo. Know2Vec first accesses to models via a black-box interface in advance, capturing vital decision knowledge from models while ensuring their privacy. Next, it employs an effective encoding technique to transform the knowledge into precise model vectors. Secondly, it maps the user's query task to a knowledge vector by probing the semantic relationships within query samples. Furthermore, the proxy ensures the knowledge-consistency between query vector and model vectors within their alignment space, which is optimized through the supervised learning with diverse loss functions, and finally it can identify the most suitable model for a given task during the inference stage. Extensive experiments show that our Know2Vec achieves superior retrieval accuracy against the state-of-the-art methods in diverse neural network retrieval tasks. Zhuoyi Shang, Yanwei Liu 0001, Jinxia Liu, Xiaoyan Gu 0001, Xiangyang Ji |
AAAI | 2 |
| 2025 | Towards Understanding How Knowledge Evolves in Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) are gradually becoming the foundation for many artificial intelligence applications. However, understanding their internal working mechanisms has continued to puzzle researchers, which in turn limits the further enhancement of their capabilities. In this paper, we seek to investigate how multimodal knowledge evolves and eventually induces natural languages in LVLMs. We design a series of novel strategies for analyzing internal knowledge within LVLMs, and delve into the evolution of multimodal knowledge from three levels, including single token probabilities, token probability distributions, and feature encodings. In this process, we identify two key nodes in knowledge evolution: the critical layers and the mutation layers, dividing the evolution process into three stages: rapid evolution, stabilization, and mutation. Our research is the first to reveal the trajectory of knowledge evolution in LVLMs, providing a fresh perspective for understanding their underlying mechanisms. Our codes are avaiable at https://github.com/XIAO4579/Vlm-Interpretability. Sudong Wang, Yao Zhu 0003, Jianing Li 0001, Zizhe Wang, Yanwei Liu 0001, Xiangyang Ji |
CVPR | 6 |
| 2025 | PlugMark: A Plug-In Zero-Watermarking Framework for Diffusion Models
Pengzhen Chen, Yanwei Liu 0001, Xiaoyan Gu 0001, Enci Liu, Zhuoyi Shang, Xiangyang Ji, Wu Liu 0005 |
ICCV | 2 |
| 2025 | SHIFT: Smoothing Hallucinations by Information Flow Tuning for Multimodal Large Language Models
Sudong Wang, Yao Zhu 0003, Enci Liu, Jianing Li 0001, Yanwei Liu 0001, Xiangyang Ji |
ICCV | 6 |
| 2025 | Cross-Layer-Optimized Link Selection for Hologram Video Streaming Over Millimeter Wave NetworksabstractHolographic-type communication brings an immersive tele-holography experience by delivering holographic contents to users. As the direct representation of holographic contents, hologram videos are naturally three-dimensional representation, which consist of a huge volume of data. Advanced multi-connectivity (MC) millimeter-wave (mmWave) networks are now available to transmit hologram videos by providing the necessary bandwidth. However, the existing link selection schemes in MC-based mmWave networks neglect the source content characteristics of hologram videos and the coordination among the parameters of different protocol layers in each link, leading to sub-optimal streaming performance. To address this issue, we propose a cross-layer-optimized link selection scheme for hologram video streaming over mmWave networks. This scheme optimizes link selection by jointly adjusting the video coding bitrate, the modulation and channel coding schemes (MCS), and link power allocation to minimize the end-to-end hologram distortion while guaranteeing the synchronization and quality balance between real and imaginary components of the hologram. Results show that the proposed scheme can effectively improve the hologram video streaming performance in terms of PSNR by 1.2 dB to$\mathbf{6. 4 d B}$against the non-cross-layer scheme. Yiming Jiang 0002, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou |
WCNC | 2 |
| 2024 | BBR-based and fairness-guaranteed congestion control and packet scheduling for MPQUIC over heterogeneous networks
Zhenjie Deng, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Dacai Liu |
Comput. Commun. | 2 |
| 2023 | Hierarchical Multi-Scale Adaptive Conv-LSTM Network for Human Action Recognition Based on Wearable SensorsabstractRecently, human action recognition has been widely used in the fields of health monitoring, human-robot interaction, medical treatment, and sports. Due to the availability of various wearable devices on the market, we can easily access sensor data for human action recognition. However, it is still a challenge to capture minute action processes as well as extract spatio-temporal motion patterns from serial sensor data. Therefore, we propose a novel hierarchical multi-scale adaptive Conv-LSTM network structure called HMA Conv-LSTM. The finer-grained spatial information in the sensor signals is extracted by hierarchical multi-scale convolution. The multi-channel feature fusion through adaptive channel feature fusion retains important information and improves model efficiency. We capture temporal context information by dynamic channel selection-LSTM based on the attention mechanism. Extensive experiments on the Opportunity and PAMAP2 public datasets show that our proposed model achieves competitive performance compared to several state-of-the-art approaches. Weiliang Xie, Qian Huang 0008, Yanfang Wang 0005, Yanwei Liu 0001 |
MMAsia | 5 |
| 2023 | Perspectively Equivariant Keypoint Learning for Omnidirectional ImagesabstractRobust keypoint detection on omnidirectional images against large perspective variations, is a key problem in many computer vision tasks. In this paper, we propose a perspectively equivariant keypoint learning framework named OmniKL for addressing this problem. Specifically, the framework is composed of a perspective module and a spherical module, each one including a keypoint detector specific to the type of the input image and a shared descriptor providing uniform description for omnidirectional and perspective images. In these detectors, we propose a differentiable candidate position sorting operation for localizing keypoints, which directly sorts the scores of the candidate positions in a differentiable manner and returns the globally top-K keypoints on the image. This approach does not break the differentiability of the two modules, thus they are end-to-end trainable. Moreover, we design a novel training strategy combining the self-supervised and co-supervised methods to train the framework without any labeled data. Extensive experiments on synthetic and real-world 360° image datasets demonstrate the effectiveness of OmniKL in detecting perspectively equivariant keypoints on omnidirectional images. Our source code are available online at https://github.com/vandeppce/sphkpt. Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Liming Wang 0001, Zhen Xu 0009, Xiangyang Ji |
IEEE Trans. Image Process. | 2 |
| 2022 | 360-Attack: Distortion-Aware Perturbations from Perspective-ViewsabstractThe application of deep neural networks (DNNs) on 360-degree images has achieved remarkable progress in the recent years. However, DNNs have been demonstrated to be vulnerable to well-crafted adversarial examples, which may trigger severe safety problems in the real-world applications based on 360-degree images. In this paper, we propose an adversarial attack targeting spherical images, called 360-attactk, that transfers adversarial perturbations from perspective-view (PV) images to a final adversarial spherical image. Given a target spherical image, we first represent it with a set of planar PV images, and then perform 2D attacks on them to obtain adversarial PV images. Considering the issue of the projective distortion between spherical and PV images, we propose a distortion-aware attack to reduce the negative impact of distortion on attack. Moreover, to reconstruct the final adversarial spherical image with high aggressiveness, we calculate the spherical saliency map with a novel spherical spectrum method and next propose a saliency-aware fusion strategy that merges multiple inverse perspective projections for the same position on the spherical image. Extensive experimental results show that 360-attack is effective for disturbing spherical images in the black-box setting. Our attack also proves the presence of adversarial transferability from Z2 to SO(3) groups. Yanwei Liu 0001, Jinxia Liu, Jingbo Miao, Antonios Argyriou, Liming Wang 0001, Zhen Xu 0009 |
CVPR | 2 |
| 2022 | RAPMiner: A Generic Anomaly Localization Mechanism for CDN System with Multi-dimensional KPIsabstractAs essential work in IT operations, anomaly localization, aiming to identify the affected scope of Internet infrastructure once an anomaly alarm occurs, is challenging due to the huge search space. The existing solutions usually show limited performances in the CDN scenario since they take the desirable assumptions that do not match with the practical anomaly pattern features. To address this issue, in this paper, we propose RAPMiner, which first uses a classification power-based redundant attribute deletion to prune the non-root cause attribute combinations, and then adopts an anomaly confidence-guided layer-by-layer top-down search to avoid searching for anomaly but non-root patterns. Both of them are effective in narrowing the search space. Experimental results show that RAPMiner can achieve comparable performance with the SOTA approach on the published Squeeze dataset according to F1-score and efficiency, as well as the best RC@k with stable parameter sensitivity on the RAPMD. Yanwei Liu 0001, Zhen Xu 0009 |
DSN | 2 |
| 2022 | SP Attack: Single-Perspective Attack for Generating Adversarial Omnidirectional ImagesabstractThe safety of Deep Neural Networks (DNNs) processing omnidirectional images (ODIs) is an under-researched topic. In this paper, we propose a novel sparse attack, named Single-Perspective (SP) Attack, towards fooling these models by perturbing only one perspective image (PI) rendered from the target ODI. The attack is launched from the perspective domain, and finally the perturbation is transferred to the original ODI. To this end, we propose an effective PI position searching algorithm based on Bayesian Optimization, and then corrupt the PI centered on the desirable position with unconstrained/constrained perturbations. Extensive experiments on synthetic and real-world omnidirectional datasets demonstrate that SP Attack can overcome the projection deformation of ODIs, and mislead the neural networks by limiting the perturbations in a single patch on the target ODI. Yanwei Liu 0001, Jinxia Liu, Pengwei Zhan, Liming Wang 0001, Zhen Xu 0009 |
ICASSP | 2 |
| 2022 | Variational Depth Estimation on Hypersphere for PanoramaabstractDepth estimation for panorama is a key part of 3D scene understanding, and adopting discriminative models is the most common solution. However, due to the rectangular convolution kernel, these existing learning methods cannot efficiently extract the distorted features in panoramas. To this end, we propose OmniVAE, a generative model based on Conditional Variational Auto-Encoder (CVAE) and von Mises-Fisher (vMF) distribution, to strengthen the exclusive generative ability for spherical signals by mapping panoramas to hypersphere space. Further, to alleviate the side effects of manifold-mismatching caused by non-planar distribution, we put forward the Atypical Receptive Field (ARF) module to slightly shift the receptive field of the network and even take the distribution difference into account in the reconstruction loss. The quantitative and qualitative evaluations are performed on real-world and synthetic datasets, and the results show that OmniVAE outperforms the state-of-the-art methods. Jingbo Miao, Yanwei Liu 0001, Jinxia Liu, Zhen Xu 0009 |
ICIP | 2 |
| 2022 | Viewport-Oriented Panoramic Image InpaintingabstractPanoramic images are usually viewed through Head Mounted Displays (HMDs), which renders only a narrow field of view from the raw panoramic image. This distinctive viewing feature has largely been ignored when inpainting panoramic images. To address this issue, we propose a viewport-oriented generative adversarial panoramic image inpainting network in this paper. For capturing the distorted features accurately in the generating process of equirectangular projection (ERP) panoramic image, a latitude-adaptive feature fusion module is devised to aggregate the latitude-level features in ERP image and less-distorted patch-level viewport-domain features. Furthermore, a novel cross-domain discriminator is proposed to force the inpainting network to generate more plausible results in viewports. Extensive experiments show that our model achieves better performance compared to the baseline methods, especially in the viewport images. Zhuoyi Shang, Yanwei Liu 0001, Guoyi Li, Jingbo Miao, Jinxia Liu, Liming Wang 0001 |
ICIP | 2 |
| 2022 | Six-to-one: Cubemap-guided Feature Calibration for Panorama Object DetectionabstractObject detection methods for perspective images have proven increasingly efficient, but the techniques for equirectangular projection (ERP) panoramas from inherently spherical imaging cannot still achieve satisfactory performance. Due to the various degrees of distortion at different pixel locations, current algorithms cannot adapt to the changes in shape and contour caused by stretching, which results in performance degradation when migrating them from perspective images to spherical ones. In this paper, we improve the network for panorama object detection and introduce the cube-domain information with discontinuity but low distortion to correct the panorama features. Unlike previous works, we consider the impact of semantic discontinuity from all tangent planes instead of overlaying features when needed. Considering the six facets as unified, i.e., six-to-one for extraction, the proposed Facet-Link module enhances the long-range sensing capability at the facet level in the frequency domain. Moreover, the position alignment packs different facets, i.e., six-to-one for calibration, to preserve more global signals during the correction stage, which establishes semantic pathways for feature interactions between panorama and cubemap in the two dimensions, facet-facet and cube-pano, respectively. Extensive experiments on synthetic and real-world datasets verify the effectiveness and robustness of our proposed method. Jingbo Miao, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Yanni Han, Zhen Xu 0009 |
ICTAI | 2 |
| 2022 | Switching Gaussian Mixture Variational RNN for Anomaly Detection of Diverse CDN WebsitesabstractTo conduct service quality management of industry devices or Internet infrastructures, various deep learning approaches have been used for extracting the normal patterns of multivariate Key Performance Indicators (KPIs) for unsupervised anomaly detection. However, in the scenario of Content Delivery Networks (CDN), KPIs that belong to diverse websites usually exhibit various structures at different timesteps and show the non-stationary sequential relationship between them, which is extremely difficult for the existing deep learning approaches to characterize and identify anomalies. To address this issue, we propose a switching Gaussian mixture variational recurrent neural network (SGmVRNN) suitable for multivariate CDN KPIs. Specifically, SGmVRNN introduces the variational recurrent structure and assigns its latent variables into a mixture Gaussian distribution to model complex KPI time series and capture the diversely structural and dynamical characteristics within them, while in the next step it incorporates a switching mechanism to characterize these diversities, thus learning richer representations of KPIs. For efficient inference, we develop an upward-downward autoencoding inference method which combines the bottom-up likelihood and up-bottom prior information of the parameters for accurate posterior approximation. Extensive experiments on real-world data show that SGmVRNN significantly outperforms the state-of-the-art approaches according to F1-score on CDN KPIs from diverse websites. Yanwei Liu 0001, Antonios Argyriou, Tao Lin 0001, Zhen Xu 0009, Bo Chen 0001 |
INFOCOM | 3 |
| 2021 | Improved Face Detector on Fisheye Images via Spherical-Domain AttentionabstractAs one type of omnidirectional projection, fisheye images have been widely used in automatic driving and visual surveillance. However, they cannot be processed well by the traditional algorithms designed for the planar rectilinear images since they usually suffer from severe geometric distortion during image formation. In this paper, the conventional face detection algorithm is enhanced to fit the fisheye images via combining with the spherical convolution block by learning rotation-invariant features from the spherical domain. The learned features from both planar and spherical domains are subsequently mixed by the spatial attention mechanism. Consequently, the whole network can automatically learn the distorted features directly from different positions on the target image. Experimental results verify that our network can detect distorted faces on fisheye images effectively and maintain the performance on traditional planar images. Jingbo Miao, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Zhen Xu 0009, Yanni Han |
ISCC | 2 |
| 2021 | Output Security for Multi-user Augmented Reality using Federated Reinforcement LearningabstractWith the rapid advancements in Augmented Reality, the number of AR users is gradually increasing and the multiuser AR ecosystem is on the rise. Currently, AR applications usually present results without limitations, which causes great latent danger to users, so it is necessary to apply strategies to ensure the safe output of AR. Due to the environmental diversities among the distributed users, the traditional approaches designed for single-user AR are not efficient for multi-user AR applications. Considering the characteristics of multi-user AR scenarios, we propose a multi-user AR output strategy model based on Federated Reinforcement Learning. With the device-fog-cloud hierarchical architecture, the proposed models are obtained first by Reinforcement Learning on users' devices, and are then hierarchically aggregated on the fog nodes and cloud server. We performed extensive AR simulations in Unity and obtained the results that show our method can avoid several security problems existent in multi-user AR applications. Fengchao Wang, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Liming Wang 0001, Zhen Xu 0009 |
ISCC | 2 |
| 2021 | D2D-Assisted Federated Learning in Mobile Edge Computing NetworksabstractWith the proliferation of edge intelligence and the breakthroughs in machine learning, Federated Learning (FL) is capable of learning a shared model across several edge devices by preserving their private data from being exposed to external adversaries. However, the distributed architecture of FL naturally introduces communication between the central parameter server and the distributed learning nodes. The huge communication cost poses a challenge to practical FL, especially for FL in mobile edge computing (MEC) networks. Existing communication-efficient FL systems predominantly optimize their intrinsic learning process and are not concerned with the implications on the network. In this paper we propose a FL scheme that leverages Device-to-Device (D2D) communication (hence called D2D-FedAvg) and is suitable for mobile edge networks. D2D-FedAvg creates a two-tier learning model where D2D learning groups communicate their results as a single entity to the MEC server leading to traffic reduction. We propose the schemes for D2D grouping, master UE selection, and also D2D exit in the learning process and then form a complete D2D-assisted federated averaging algorithm. Via extensive simulations on the Federated Extended MNIST dataset, the feasibility and convergence of D2D-FedAvg scheme are evaluated. Our results show that D2D-FedAvg lowers the communication cost relative to the typical Federated Averaging (FedAvg) in cellular networks as the number of users is increased (for 100 cellular users 37% traffic reduction), while keeping the same learning accuracy with FedAvg across the board. Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Yanni Han |
WCNC | 2 |
| 2021 | SDFVAE: Static and Dynamic Factorized VAE for Anomaly Detection of Multivariate CDN KPIsabstractContent Delivery Networks (CDNs) are critical for providing good user experience of cloud services. CDN providers typically collect various multivariate Key Performance Indicators (KPIs) time series to monitor and diagnose system performance. State-of-the-art anomaly detection methods mostly use deep learning to extract the normal patterns of data, due to its superior performance. However, KPI data usually exhibit non-additive Gaussian noise, which makes it difficult for deep learning models to learn the normal patterns, resulting in degraded performance in anomaly detection. In this paper, we propose a robust and noise-resilient anomaly detection mechanism using multivariate KPIs. Our key insight is that different KPIs are constrained by certain time-invariant characteristics of the underlying system, and that explicitly modelling such invariance may help resist noise in the data. We thus propose a novel anomaly detection method called SDFVAE, short for Static and Dynamic Factorized VAE, that learns the representations of KPIs by explicitly factorizing the latent variables into dynamic and static parts. Extensive experiments using real-world data show that SDFVAE achieves a F1-score ranging from 0.92 to 0.99 on both regular and noisy dataset, outperforming state-of-the-art methods by a large margin. Tao Lin 0001, Bo Jiang 0003, Yanwei Liu 0001, Zhen Xu 0009, Zhi-Li Zhang |
WWW | 5 |
| 2021 | Tile caching for scalable VR video streaming over 5G mobile networks
Kedong Liu, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Cross-layer DASH-based multipath video streaming over LTE and 802.11ac networks
Zhenjie Deng, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou |
Multim. Tools Appl. | 2 |
| 2021 | 360-Degree VR Video Watermarking Based on Spherical Wavelet TransformabstractSimilar to conventional video, the increasingly popular 360 virtual reality (VR) video requires copyright protection mechanisms. The classic approach for copyright protection is the introduction of a digital watermark into the video sequence. Due to the nature of spherical panorama, traditional watermarking schemes that are dedicated to planar media cannot work efficiently for 360 VR video. In this article, we propose a spherical wavelet watermarking scheme to accommodate 360 VR video. With our scheme, the watermark is first embedded into the spherical wavelet transform domain of the 360 VR video. The spherical geometry of the 360 VR video is used as the host space for the watermark so that the proposed watermarking scheme is compatible with the multiple projection formats of 360 VR video. Second, the just noticeable difference model, suitable for head-mounted displays (HMDs), is used to control the imperceptibility of the watermark on the viewport. Third, besides detecting the watermark from the spherical projection, the proposed watermarking scheme also supports detecting watermarks robustly from the viewport projection. The watermark in the spherical domain can protect not only the 360 VR video but also its corresponding viewports. The experimental results show that the embedded watermarks are reliably extracted both from the spherical and the viewport projections of the 360 VR video, and the robustness of the proposed scheme to various copyright attacks is significantly better than that of the competing planar-domain approaches when detecting the watermark from viewport projection. Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Siwei Ma 0001, Liming Wang 0001, Zhen Xu 0009 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | RegionSparse: Leveraging Sparse Coding and Object Localization to Counter Adversarial AttacksabstractAlthough deep neural networks have demonstrated exceptional performance in substantial computer vision tasks, they can be easily confused by carefully generated adversarial examples. Via a novel technique we call activation visualization, the particular characteristics of adversarial examples are analyzed in this paper. Observing that the dominant features of adversarial examples are distributed over a high-dimensional space, we propose a defense framework named RegionSparse that projects the images into a low-dimensional space to remove the influence of the adversarial perturbations on the performance of deep neural networks. In RegionSparse, after training a robust global dictionary, the region where pixels are highly related to classification is firstly located by an object localization mechanism, then the sparse coding is performed on the located object region, together with a perturbation suppression for the remaining region. Extensive experiments on ImageNet dataset for gray-box, black-box, and transferred attacks are performed and the results show that RegionSparse can eliminate up to 90% attacks delivered by strong attacks including Momentum Iterative Fast Gradient Sign Method and Carlini-Wagner's L2attack. Yanwei Liu 0001, Liming Wang 0001, Zhen Xu 0009, Qiuqing Jin |
IJCNN | 2 |
| 2019 | Joint EPC and RAN Caching of Tiled VR Videos for Mobile Networks
Kedong Liu, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou |
MMM (1) | 2 |
| 2019 | MEC-Assisted Panoramic VR Video Streaming Over Millimeter Wave Mobile NetworksabstractPanoramic virtual reality video (PVRV) is becoming increasingly popular since it offers a true immersive experience. However, the ultra-high resolution of PVRV requires significant bandwidth and ultra-low latency for PVRV streaming, something that makes challenging the extension of this application to mobile networks. Besides bandwidth, the frequent perspective viewport rendering induces a heavy computational load on battery-constrained mobile devices. To attack these problems jointly, this paper proposes a PVRV streaming system that is designed for modern multiconnectivity-based millimeter wave (mmWave) cellular networks in conjunction with mobile edge computing (MEC). First, mmWave is deployed to support the high bandwidth needs of PVRV streaming. Next, the multiple mmWave links that tend to suffer from outages are coupled with a sub-6 GHz link to ensure disruption-free wireless communication. With the help of an MEC server, the tradeoff among link adaptation, transcoding-based chunk quality adaptation, and viewport rendering offloading is sought to improve the wireless bandwidth utilization and mobile device's energy efficiency. Simulation results show that the proposed scheme can improve the streaming performance in both energy efficiency and the quality of received viewport over the state-of-the-art schemes. Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Song Ci |
IEEE Trans. Multim. | 1 |
| 2018 | Binocular-Combination-Oriented Perceptual Rate-Distortion Optimization for Stereoscopic Video CodingabstractIn the chain of stereoscopic video processing, stereoscopic video coding and viewing are usually two independent stages. Conventional stereoscopic video coding puts an emphasis on improving the coding efficiency by seeking the optimal tradeoff between the coding bit rate and the signal-based distortion, while neglecting the perceptual behaviors of binocular combination when stereoscopic video is viewed by human beings. In this paper, we propose to utilize binocular combination to optimize the stereoscopic video coding from the perspective of perceptual quality measurement. Specifically, we propose a novel binocular-combination-oriented measurement for visual distortion and then derive the Lagrange multiplier for the binocular-combination-oriented rate-distortion optimization (RDO). Via extensive subjective tests, the results show that the proposed perceptual RDO can save more than 5% BD rate over the traditional RDO in multiview extension of High Efficiency Video Coding for stereoscopic video coding. Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Song Ci |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | 3DQoE-Oriented and Energy-Efficient 2D plus Depth Based 3D Video Streaming Over Centrally Controlled NetworksabstractIP networks have become the dominant platform for video delivery. However, bandwidth-hungry video is pushing networks to their limits: costs are rising for the operators and the viewing experience is not always satisfactory for the users. When considering 3D video delivery, the previous problems are exacerbated because of the higher volume of data that must be communicated, and the difficulty in characterizing the viewing experience of the end user. Consequently, network operators may be reluctant to deliver 3D video due to costs and unclear quality improvements to their users. In this setting, the true immersive experience of 3D video remains elusive. In this paper, we focus on the efficient delivery of 3D video in terms of quality and energy cost over centrally controlled networks. As a representative example of a centrally controlled network, a software-defined network (SDN) is assumed. Our approach is based on a comprehensive network-dependent 3D quality of experience (3DQoE) model and an energy cost model for 3D video streaming. By using the developed models, we formulate the problem of energy-efficient and 3DQoE-optimized 3D video flow path routing. The particular characteristic of video/depth rate allocation presented in 3D video is embedded seamlessly into the selection of the optimal routing paths for multiple 3D video streams. The formulated problem is NP-hard and is solved with a heuristic algorithm based on the branch-and-bound method after significant reduction of the solution search space. Extensive 3D video streaming experiments are conducted over an OpenFlow-based SDN with subjective and objective evaluations and they highlight the significant benefits of the proposed approach. Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Song Ci |
IEEE Trans. Multim. | 1 |
| 2017 | Joint Source Encoding and Networking Optimization for Panoramic Video Streaming over LTE-A DownlinkabstractWith the increasing capacity of wireless networks, more people would like to consume the 360-degree panoramic video (PV) in virtual reality (VR) applications as its immersive experience. However, due to the super-high resolution of the PV and the dynamic features of wireless networks, it is very difficult to efficiently deliver PVs over wireless networks. The traditionally independent PV encoding and networking sometimes also results in the PV quality deterioration since it neglects the harmony between the source encoding and networking. In this paper, a joint source encoding and networking optimization scheme is proposed to transmit the PV over LTE-A downlink. The PV encoding parameters during the source compression, the modulation and coding scheme (MCS), and relay selection during the networking are jointly considered to optimize the end-to-end PV quality. In addition, the video quality for region of interest (RoI, the possible viewport region) is enhanced by allowing a larger latency bound in the joint source encoding and networking optimization. Experimental results show that the proposed scheme achieves significant performance improvement for the quality of the received PV over traditional PV streaming approaches. Kedong Liu, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Xinghua Yang |
GLOBECOM | 2 |
| 2017 | Cross-layer optimized authentication and error control for wireless 3D medical video streaming over LTE
Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Song Ci |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | Cross-network and cross-layer optimized video streaming over LTE and WCDMA downlinkabstractVideo services are proliferating over today's mobile Internet and efforts have been made to improve their performance. The cross-layer video streaming that jointly optimizes the parameters at different protocol layers is a feasible solution but without bandwidth aggregation of multiple wireless networks. The cross-network video streaming can realize the bandwidth aggregation by using multiple overlapping wireless networks simultaneously. However, parameters at different protocol layers are optimized independently in the cross-network optimization. In this paper, we propose a joint cross-network and cross-layer optimized video streaming scheme that utilizes the bandwidth aggregation of cross-network video streaming and further improve the performance by jointly optimizing the parameters of different protocol layers in each network with a cross-layer manner. In the proposed scheme, the LTE and WCDMA networks are adopted. The bit-rate of the video at application layer, the rate allocation among networks and the parameters of physical layers in each network are jointly optimized. Experimental results show that the proposed scheme gains higher quality of experience in terms of PSNR than the state-of-the-art schemes. Zhenjie Deng, Yanwei Liu 0001, Jinxia Liu, Xin Chen 0019, Antonios Argyriou, Zhen Xu 0009, Song Ci |
ISCC | 2 |
| 2016 | Choquet integral based QoS-to-QoE mapping for mobile VoD applicationsabstractToday, how to accurately predict the quality of experience (QoE) of the networking service is a very important issue for the network operator to optimize the service. However, due to the complex multi-dimensional characteristics of QoE, QoE estimation is extremely challenging. With utilizing the advantages of quality of service (QoS) in evaluating the networking performance, we exploit QoS/QoE correlation to predict QoE by building a QoSto-QoE mapping relationship. To fully consider the inter-dependency among QoS parameters towards forming the QoE, a Choquet integral based fuzzy measurement method is used to map QoS to QoE. Via extensive experiments in mobile VoD applications, the advancement and effectiveness of the proposed method are verified. Yanwei Liu 0001, Jinxia Liu, Zhen Xu 0009, Song Ci |
IWQoS | 1 |
| 2016 | Perceptual rate-distortion optimization for H.264/AVC video coding from both signal and vision perspectives
Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Ruixiao Yao, Song Ci |
Multim. Tools Appl. | 2 |
| 2016 | SSIM-based error-resilient cross-layer optimization for wireless video streaming
Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Song Ci |
Signal Process. Image Commun. | 2 |
| 2016 | Achieving energy-neutral data transmission by adjusting transmission power for energy-harvesting wireless sensor networksabstractAbstract Recently, benefiting from rapid development of energy harvesting technologies, the research trend of wireless sensor networks has shifted from the battery‐powered network to the one that can harvest energy from ambient environments. In such networks, a proper use of harvested energy poses plenty of challenges caused by numerous influence factors and complex application environments. Although numerous works have been based on the energy status of sensor nodes, no work refers to the issue of minimizing the overall data transmission cost by adjusting transmission power of nodes in energy‐harvesting wireless sensor networks. In this paper, we consider the optimization problem of deriving the energy‐neutral minimum cost paths between the source nodes and the sink node. By introducing the concept of energy‐neutral operation, we first propose a polynomial‐time optimal algorithm for finding the optimal path from a single source to the sink by adjusting the transmission powers. Based on the work earlier, another polynomial‐time algorithm is further proposed for finding the approximated optimal paths from multiple sources to the sink node. Also, we analyze the network capacity and present a near‐optimal algorithm based on the Ford–Fulkerson algorithm for approaching the maximum flow in the given network. We have validated our algorithms by various numerical results in terms of path capacity, least energy of nodes, energy ratio, and path cost. Simulation results show that the proposed algorithms achieve significant performance enhancements over existing schemes. Copyright © 2016 John Wiley & Sons, Ltd. Wei An 0002, Yanni Han, Haiyan Luo, Yanwei Liu 0001, Song Ci, Hui Tang 0001 |
Wirel. Commun. Mob. Comput. | 5 |
| 2015 | Video-aware time-domain resource partitioning in heterogeneous cellular networksabstractHeterogenous cellular networks (HCN) consist of macrocells and small cells that are overlaid in the same geographical area. Hence, is critical that the high power macrocell shuts off its transmissions for a fraction of the time to allow the low power small cells to transmit without interference. This is the time-domain resource partitioning (TDRP) mechanism. In this paper we investigate video communication in HCNs when TDRP is employed. More specifically we consider the problem of maximizing the average video quality of all users, by jointly optimizing the rate allocated to each specific video stream and the quality that it is streamed. The resulting mixed integer linear program (MILP) formulation is solved numerically. Simulation results indicate clearly that as the small cells and the users are increased the proposed system can improve significantly the video quality. Antonios Argyriou, Dimitrios Kosmanos, Leandros Tassiulas, Yanwei Liu 0001, Song Ci |
ICC | 4 |
| 2015 | Transmit power aware cross-layer optimization for LTE uplink video streamingabstractThe rapid developments of advanced wireless communication technologies and mobile devices are boosting the uplink multimedia applications. In this paper, a transmit power aware cross-layer optimization scheme is proposed to achieve a good trade-off between the transmit power and the perceived video quality for Long Term Evolution uplink video streaming. Specifically, the video coding quantization parameter and encoding mode at the application layer, and the uplink transmit power as well as modulation and coding scheme at the physical layer are jointly adjusted in the cross-layer optimization. To further improve the perceptual video experience for the end user with limited transmission resources, unequal quality control is performed by enhancing the video quality of region of interest. Additionally, the structural similarity is adopted as the video quality measurement metric to make the optimized video properly preserve the structural information during the cross-layer optimization process. Experimental results show that significant performance improvements in terms of the transmit power reduction and the perceptual video quality are achieved for the proposed transmit power aware cross-layer optimization scheme. Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Ruixiao Yao, Song Ci, Antonios Argyriou |
ICC | 2 |
| 2015 | Energy harvesting aware topology control with power adaptation in wireless sensor networks
Wei An 0002, Yanni Han, Yanwei Liu 0001, Song Ci, Fang-Ming Shao, Hui Tang 0001 |
Ad Hoc Networks | 4 |
| 2015 | Scalable 3D video streaming over P2P networks with playback length changeable chunk segmentation
Yanwei Liu 0001, Jinxia Liu, Junping Song, Antonios Argyriou |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | A cooperative protocol for video streaming in dense small cell wireless relay networks
Dimitrios Kosmanos, Antonios Argyriou, Yanwei Liu 0001, Leandros Tassiulas, Song Ci |
Signal Process. Image Commun. | 3 |
| 2015 | Utility-Based H.264/SVC Video Streaming Over Multi-Channel Cognitive Radio NetworksabstractIn a cognitive radio (CR) network, a secondary user (SU) with multiple interfaces is capable of accessing multiple CR channels in an opportunistic fashion. Therefore, the available channel resources may change dramatically, and the reliabilities of the multiple accessed CR channels are also time-varying. Video streaming in such a multi-channel CR network faces great challenges in guaranteeing the quality of the received video. To deal with these challenges, we adopt H.264/SVC encoded video as the source and firstly optimize the video streaming from the perspective of exploiting more channel resources for the SU by developing a flexible sensing-transmission scheme for opportunistic spectrum access (OSA). In this scheme, the primary user activity, channel sensing result, and channel sensing accuracy are all considered in reducing the unnecessary channel sensings and correspondingly extending the transmission duration for the SU. Based on the flexible sensing-transmission scheme, we then propose a utility-based H.264/SVC video transmission scheme to further improve the expected video quality at the receiver. Specifically , the network abstraction layer units (NALUs) in the SVC video are assigned utilities which accurately reflect their contributions to the video quality, and the total effective utility of expected received video is maximized through perfectly dispatching the NALUs over the multiple CR channels. Both analytical studies and experimental results validate the effectiveness and efficiency of the proposed method. Ruixiao Yao, Yanwei Liu 0001, Jinxia Liu, Pinghua Zhao, Song Ci |
IEEE Trans. Multim. | 2 |
| 2014 | Computation scalable disparity estimation for delay sensitive 3D video surveillance systemabstractDisparity estimation is an important task in many 3D video surveillance applications. How to generate the disparity information at the front end under limited delay budget is very challenging in a real-time surveillance system. In this paper, we tackle this problem through adopting multiresolution strategy in the disparity estimation process. Our contribution is twofold. First, unlike existing coarse-to-fine strategies based on uniform sampling, we present a fast disparity estimation algorithm based on nonuniform sampling at the coarse level. The disparity values of the non-samples are initially interpolated through Delaunay Triangulation, and then are refined through bilateral filtering. The content-aware nonuniform sampling provides better disparity approximation in the triangulated interpolation, and consequently leads to faster convergence in the refinement procedure. Second, we model the computation time in the disparity estimation process for execution assistance in the delay sensitive surveillance system. Both the sampling cell size and the data resolution are considered in this model, in order to accommodate different runtime requirements. Experimental results demonstrate the efficiency of the proposed scheme. Song Ci, Yanwei Liu 0001 |
AVSS | 3 |
| 2014 | A novel low-complexity method for determining nonadditive interaction measures based on least-norm learningabstractNumerous research works have been done on the Choquet integral model due to the tremendous usage in many fields. However, the application is still significantly restricted by the curse of dimensionality, involved in determining the non-additive interaction measures, that can properly reflect the interactions among predictive attributes toward the objective. To this end, in this paper we propose a novel determination method for non-additive interaction measures by the way of solving a sequence of least norm problems and iteratively updating the values of interaction measures, namely least norm learning. This method can achieve a significant reduction on the computation time complexity from O(m × 2n) to O(mn) for solving the Choquet integral model, where ra and n are the numbers of observations and attributes, respectively. Also we achieve to reduce the computation space complexity from O(m × 2n) to 0(2n). A case study on cross-layer optimized wireless multimedia communications is adopted to validate the proposed method. Both analytical and experimental results show the effectiveness of the proposed method. Wei An 0002, Chunxiao Ren, Song Ci, Dalei Wu, Haiyan Luo, Yanwei Liu 0001 |
FUZZ-IEEE | 6 |
| 2014 | Hierarchical-matching based scalable video streaming over multi-channel cognitive radio networksabstractWith its ability to promote the wireless spectrum utilization, cognitive radio (CR) is a promising technology for various broadband wireless video applications. A secondary user (SU) with multiple wireless interfaces is capable to sense and access multiple CR channels, each of which can only be accessed in an opportunistic fashion. Therefore, the amount of available channel resources may change dramatically, and the reliabilities of the multiple accessed CR channels are time-varying and different from each other. To improve the quality of the received video in such a multi-channel CR network, we first develop a flexible sensing-transmission structure for dynamic spectrum access according to primary user activity and channel sensing accuracy. Based on the proposed sensing-transmission structure, we then propose a hierarchical-matching scheme to adapt the scalable video stream to the multiple time-varying and reliability-different CR channels, where we consider the priorities and validity of the network abstraction layer units (NALUs) in the transmission scheduling. Under this scheme, the more important NALUs at the global Group Of Pictures (GOP) scope will be dispatched over the more reliable channels. Both analytical studies and experimental results show the effectiveness and efficiency of the proposed method. Ruixiao Yao, Yanwei Liu 0001, Jinxia Liu, Pinghua Zhao, Song Ci |
GLOBECOM | 2 |
| 2014 | SSIM-based cross-layer optimized video streaming over LTE downlinkabstractMany research efforts have been done to guarantee the quality of service for video streaming over LTE. Among them, cross-layer optimized video delivery is one of the most commonly adopted approaches. However, most existing schemes of cross-layer optimized video delivery adopt PSNR as the optimization target, but do not well consider human vision characteristics. In this paper, we adopt a new metric - Structural Similarity (SSIM) in video quality evaluation and develop a SSIM-based cross-layer optimization scheme with the suppression of propagated error for video delivery over LTE downlink. The modulation and coding scheme at the physical layer of LTE downlink is selected by jointly considering the characteristics of video packets and time-varying channel states. Correspondingly, the quantization parameter of video codec at the application layer is adjusted to make the bit rate of the video stream adapt to the varying throughput of the physical link. Moreover, the error-resilient rate-distortion optimization is adopted in the proposed cross-layer optimization to suppress the effect of error propagation on the video quality degradation. Experimental results show that the proposed cross-layer optimization scheme can maintain more structural information than the conventional schemes in the received video, which correspondingly improves the video quality perceived by the end user. Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Ruixiao Yao, Song Ci |
GLOBECOM | 2 |
| 2014 | Intrinsic flexibility exploiting for scalable video streaming over multi-channel wireless networksabstractScalable video has natural advantages in adapting to the multi-channel wireless networks. And some existing works tried to further optimize the scalable video transmission by combining the crude layer-importance mapping with some extrinsic techniques, such as Forward Error Correction (FEC) and Adaptive Modulation and Coding (AMC). However, the intrinsic flexibility of scalable video streaming over the multichannel wireless networks was neglected. In this paper, we try to exploit the intrinsic flexibility by firstly analyzing the priorities of H.264/SVC video data at the network abstraction layer unit (NALU) level, and then designing the priority-validity delivery scheme for the scalable video streaming. With this strategy, the sub-stream extraction is intelligently adjusted according to the delivery history, and the more important data in a group of pictures (GOP) will be delivered through the more reliable channels. Experimental results also validate the strategy's effectiveness in improving the objective quality and perceptual experience of the received video. Ruixiao Yao, Yanwei Liu 0001, Jinxia Liu, Pinghua Zhao, Song Ci |
VCIP | 2 |
| 2014 | Energy harvesting aware topology control with power adaptation in wireless sensor networksabstractRecently, energy harvesting technology has been introduced into wireless sensor networks to solve the traditional battery-powered energy bottleneck problem. However, due to battery capacity limitation, the harvested energy would overflow while the nodes are in the energy saturation status. Aiming at this, we consider topology control approach in the EHWSNs that allows each node to adaptively adjust its transmission power level to utilize the harvested energy efficiently. Specifically, we first model nodes' behaviors as an ordinal potential game where the high harvesting power nodes interact with the low harvesting power nodes to collaboratively maintain the whole network's connectivity. And we theoretically prove the existence of Nash equilibrium in this game. Then, a polynomial-time algorithm has been proposed to achieve Nash equilibrium. Simulation results show that our algorithm exploits the available energy resources in an efficient way and outperforms existing energy-aware algorithm in terms of energy conservation and equilibrium distribution. Yanwei Liu 0001, Yanni Han, Wei An 0002, Song Ci, Hui Tang 0001 |
WCNC | 2 |
| 2014 | 3D visual experience oriented cross-layer optimized scalable texture plus depth based 3D video streaming over wireless networks
Jinxia Liu, Yanwei Liu 0001, Song Ci, Ruixiao Yao |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | SSIM-based error-resilient rate-distortion optimization of H.264/AVC video coding for wireless streaming
Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Song Ci, Ruixiao Yao |
Signal Process. Image Commun. | 2 |
| 2013 | Binocular video object tracking with fast disparity estimationabstractThis paper presents a binocular PTU (pan-tilt unit) camera video object tracking scheme using the MeanShift algorithm and the runtime disparity estimation. The proposed method is to accommodate the requirement of 3D content generation and accurate tracking in more advanced video surveillance applications. The disparity estimation process for each stereoscopic pair is formulated as an energy minimization problem. The iterative solution procedure is implemented in a course-to-fine manner. The estimated disparity is used to scale the tracking window by the MeanShift algorithm, i.e. the size of the tracking area is adjustable according to its inner disparity, and thus the moving object can be better located by the camera. The program maintains the semi-real-time performance and acceptable accuracy as evaluated on a set of standard test data. In our experiment, two PointGrey cameras are controlled through a PTU device. The disparity estimation process on the recorded tracking video (640×480) achieves 6fps on an ordinary PC (2.66GHz CPU, 4GB RAM). Song Ci, Yanwei Liu 0001, Haohong Wang, Aggelos K. Katsaggelos |
AVSS | 3 |
| 2013 | Perceptual quality driven cross-layer optimization for wireless video streamingabstractThe traditional cross-layer model for video streaming can achieve significant improvement to the end-to-end video quality. However, the video quality measurement in terms of sum of squared error (SSE) in the model does not always correlate well with the perception of the human visual system. In this paper, we propose a perceptual quality driven cross-layer optimization scheme based on the structural similarity (SSIM) index for wireless video streaming. The H.264/AVC encoding parameters on the application layer and the modulation and channel coding modes on the physical layer are mainly considered to optimize the end-to-end perceptual video quality. In order to decrease the computation complexity, a low complexity optimization algorithm based on the rate-quantization (R-Q) model is proposed to reduce the range of the candidate optimization parameters. Experimental results show that the proposed SSIM-based cross-layer optimization scheme achieves better perceptual quality than the SSE-based cross-layer optimization scheme, and the proposed low complexity optimization algorithm can achieve about 62% computation decrease compared to the conventional exhaustive searching algorithm with little loss in perceptual quality. Pinghua Zhao, Yanwei Liu 0001, Ruixiao Yao, Song Ci, Hui Tang 0001 |
CCNC | 2 |
| 2013 | Perceptual experience oriented transmission scheduling for scalable video streaming over cognitive radio networksabstractCognitive radio (CR) promotes the utilization of the wireless spectrum through admitting the secondary users to access the primary channels in an opportunistic manner. Therefore, the real-time video streaming over the CR network is challenging due to the time-varying channels. In this paper, we propose an adaptive scalable video transmission scheduling method for video streaming over time-varying channels in CR network, in which the end-to-end perceptual visual experience is optimized systematically by considering various experience-influencing factors of the scalable video source, channel, and receiving buffer. Moreover, an enhanced content-based adaptive video playout scheme is developed to further optimize the end-to-end perceptual visual experience by decreasing the receiving buffer underflow probability. The efficiency and effectiveness of the proposed method has been validated by both analytical study and extensive experimental results. Ruixiao Yao, Yanwei Liu 0001, Jinxia Liu, Pinghua Zhao, Song Ci |
GLOBECOM | 2 |
| 2013 | A playback length changeable 3D data segmentation algorithm for scalable 3D video P2P streaming systemabstractScalable 3D video P2P streaming systems can supply diverse 3D experiences for heterogeneous clients with high efficiencies. Data characteristics of the scalable 3D video make the P2P streaming efficiency more depends on the data segmentation algorithm. However, traditional data segmentation algorithm is not very appropriate for scalable 3D video P2P streaming systems. In this paper, we propose a Playback Length Changeable 3D video Segmentation (PLC3DS) algorithm. It considers the particular source-data characteristics of scalable 3D video, and provides different error resilience strengths to video and depth as well as layers with different importance levels in the transmission. The simulation results show that the proposed PLC3DS algorithm can increase the success delivery rates of chunks in more important layers, and further improve the 3D experiences of the client. Moreover, it improves the network utilization ratio remarkably. Junping Song, Yanwei Liu 0001, Jinxia Liu, Song Ci, Yan Zhang 0013 |
ICME | 2 |
| 2013 | A multi-camera motion capture system for remote healthcare monitoringabstractThis paper presents a multi-camera motion capture system aiming to provide caregivers with timely access to the patient's health status through mobile communication devices. The major components include video capture, object detection, video coding and transmission, error concealment, and video analysis. Our contribution is twofold. First, several novel ideas are developed, including fast object detection, and content-aware and adaptive video coding and transmission. Second, all components are seamlessly integrated in a unified optimization framework dedicated for online data transmission. In the scenario, the subject walked on a treadmill with four tripod cameras capturing the video from different viewpoints. After video compression and transmission over a wireless sensor network, the remote receiver recovered the videos and performed multi-view motion capture for gait analysis. Experimental results show that the presented system design achieves better video quality than traditional video coding and transmission scheme, while the requirement for a low-cost, noninvasive and real-time healthcare monitoring system is accommodated. Song Ci, Aggelos K. Katsaggelos, Yanwei Liu 0001 |
ICME | 4 |
| 2013 | Low-complexity content-adaptive Lagrange multiplier decision for SSIM-based RD-optimized video codingabstractThe SSIM-based rate distortion optimization (R-DO) has been proved to be an effective way to promote the perceptual video coding performance, and the Lagrange multiplier decision is the key to the SSIM-based RD-optimized video coding. Through extensively analyzing the characteristics of SSIM-based and SSE-based video distortions, this paper presents a low-complexity content-adaptive Lagrange multiplier decision method. The proposed method first estimates frame-level SSIM-based Lagrange multiplier by scaling the traditional SSE-based Lagrange multiplier with the ratio of SSE-based distortion to SSIM-based distortion. Via predicting the macroblock's perceptual importance in the whole frame, the macroblock-level Lagrange multiplier is further refined to promote the accuracy of the Lagrange multiplier decision. Experimental results show that the proposed method can obtain almost the same rate-SSIM performance and subjective quality as the state-of-the-art SSIM-based RD-optimized video coding methods with lower computation overheads. Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Ruixiao Yao, Song Ci, Hui Tang 0001 |
ISCAS | 2 |
| 2013 | Joint video/depth/FEC rate allocation with considering 3D visual saliency for scalable 3D video streamingabstractFor robust video plus depth based 3D video streaming, video, depth and packet-level forward error correction (FEC) can provide many rate combinations with various 3D visual qualities to adapt to the dynamic channel conditions. Video/depth/FEC rate allocation under the channel bandwidth constraint is an important optimization problem for robust 3D video streaming. This paper proposes a joint video/depth/FEC rate allocation method by maximizing the receiver's 3D visual quality. Through predicting the perceptual 3D visual qualities of the different video/depth/FEC rate combinations, the optimal GOP-level video/depth/FEC rate combination can be found. Further, the selected FEC rates are unequally assigned to different levels of 3D saliency regions within each video/depth frame. The effectiveness of the proposed 3D saliency based joint video/depth/FEC rate allocation method for scalable 3D video streaming is validated by extensive experimental results. Yanwei Liu 0001, Jinxia Liu, Song Ci |
VCIP | 1 |
| 2012 | A transcoding framework with error-resilient video/depth rate allocation for mobile 3D video streamingabstractWith the development of wireless multimedia communication, mobile 3D video will be gradually popular as it can make people enjoy the natural 3D experience anywhere and anytime. In the current stage, mobile 3D video is usually distributed over the error-prone and heterogeneous network which consists of wire-line and wireless channels. In such applications, the transcoding plays an important role to provide the adaptive and error-resilient bit-stream to the wireless channel. To guarantee the high quality mobile 3D video streaming, this paper proposes a transcoding framework with error-resilient video/depth rate allocation, which utilizes the multi-pass re-encodings to build a rate allocation table (Rate-QP-PLR (packet loss rate) table). Through looking up the table, the proposed framework can select the optimal transcoding quantization parameters (QPs) of video and depth to obtain the error-resilient 3D video transmission and therefore it can make the encoded stream adapt to the transition from wire-line to wireless network. Experimental results show that the proposed 3D video transcoding framework can provide the superior rate-distortion performance for the mobile 3D video streaming. Yanwei Liu 0001, Song Ci, Hui Tang 0001 |
ICC | 1 |
| 2012 | Integrating stereoscopic image transcoding with retargeting for mobile streamingabstractWith the progresses in 3D imaging technologies, mobile stereoscopic images are gradually popular since it can provide the anytime and anywhere 3D viewing effects. Due to the heterogeneous transmission environments and various screen sizes of mobile devices, the stereoscopic image transcoding and retargeting are two completely independent adaptation techniques for mobile 3D image streaming. These two processing operations both deal with the disparity estimation. In this work, we propose to integrate stereoscopic 3D image transcoding with retargeting for mobile streaming. The stereoscopic image transcoding and retargeting are coupled through a straightforward link that provides the pixel-based disparity map from the retargeting stage to guide the transcoding for saving some computations while keeping efficient bit allocation. Our experimental results clearly demonstrate the advantage of coupling the transcoding with retargeting in terms of faster transcoding and saliency-based bit allocation for low bit-rate mobile streaming. Yanwei Liu 0001, Song Ci, Jinxia Liu |
VCIP | 1 |
| 2012 | A wireless video surveillance system with an active cameraabstractThis paper introduced a camera surveillance system in wireless communications. The system contains three major modules, PTU (pan-tilt unit) camera control for surveillance video capture, cross-layer control for data compression and transmission, and error concealment for video quality enhancement. Our contribution is twofold. First, a system design for data collection and transmission over wireless networks is presented and is evaluated with physical surveillance equipments. The camera is capable of following the moving target according to the control information. The end-to-end distortion estimation in the delay constrained video coding process takes into account the dynamic channel condition and physical layer modulation and coding scheme (MCS) to determine optimal coding and transmission parameters. Second, multiple error concealment strategies, including interleaving, boundary match and video up-sampling, are applied utilizing the special property of the PTU camera motion. Song Ci, Yanwei Liu 0001, Dalei Wu, Haohong Wang, Aggelos K. Katsaggelos |
VCIP | 3 |
| 2012 | QoE-oriented 3D video transcoding for mobile streamingabstractWith advance in mobile 3D display, mobile 3D video is already enabled by the wireless multimedia networking, and it will be gradually popular since it can make people enjoy the natural 3D experience anywhere and anytime. In current stage, mobile 3D video is generally delivered over the heterogeneous network combined by wired and wireless channels. How to guarantee the optimal 3D visual quality of experience (QoE) for the mobile 3D video streaming is one of the important topics concerned by the service provider. In this article, we propose a QoE-oriented transcoding approach to enhance the quality of mobile 3D video service. By learning the pre-controlled QoE patterns of 3D contents, the proposed 3D visual QoE inferring model can be utilized to regulate the transcoding configurations in real-time according to the feedbacks of network and user-end device information. In the learning stage, we propose a piecewise linear mean opinion score (MOS) interpolation method to further reduce the cumbersome manual work of preparing QoE patterns. Experimental results show that the proposed transcoding approach can provide the adapted 3D stream to the heterogeneous network, and further provide superior QoE performance to the fixed quantization parameter (QP) transcoding and mean squared error (MSE) optimized transcoding for mobile 3D video streaming. Yanwei Liu 0001, Song Ci, Hui Tang 0001, Jinxia Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2011 | Fast disparity estimation utilizing depth information for multiview video codingabstractMultiview video coding (MVC) improves the coding efficiency by motion estimation (ME) and disparity estimation (DE). ME and DE at encoder side involve in heavy computation, which needs to be further reduced for practical applications. This paper presents fast disparity estimation by utilizing depth information to reduce DE's computational complexity. First, the coordinate offset of the encoding block is calculated by 3D warping in the reference view frame utilizing depth map. Then the search region for DE is adjusted according to the coordinate offset. Experimental results verify that the proposed algorithm can save roughly 50% coding time with negligible loss of coding efficiency. Zhongmei Qiao, Debin Zhao, Yanwei Liu 0001, Wen Gao 0001 |
ISCAS | 4 |
| 2011 | End-to-end distortion optimized error control for real-time wireless video streamingabstractIn wireless video streaming, the packet loss often occurs and affects the end-user visual quality. To alleviate the transmission error effects, intra refresh coding is usually used to improve the streaming error resilience ability from the view of source coding. At the physical layer, the adaptive modulation and coding (AMC) is also used to promote the transmission reliability at the transporting level. Both the error control components have their own influences on the received video quality. To achieve the best video transmission performance, it is crucial to make an error control tradeoff between intra refresh coding and AMC. In this paper, we propose an end-to-end video distortion optimized cross-layer error control method which jointly considers the video quantization parameter (QP) and intra refresh rate at the application layer, and AMC at the physical layer for delay-constraint real-time video streaming. The experimental results show that the proposed cross-layer error control streaming method can achieve the superior objective and subjective performances to the layer-independent error control streaming methods with and without cross-layer optimization. Guangchao Peng, Yanwei Liu 0001, Yahui Hu, Song Ci, Hui Tang 0001 |
MMSP | 2 |
| 2011 | Dynamic video object detection with single PTU cameraabstractThis paper presents a video object detection method under dynamic background with single PTU (pantilt unit) camera. The overall procedure contains two steps. First the moving object is tracked through interactive Mean Shift estimation and PTU control. Camera angular speed is updated online according to the estimated object position, whereby the object can remain in the center area of the image plane. In the second step, the foreground is detected through background subtraction and shape contour. An adaptive correlation thresholding method is applied to mitigate the detection distortion in boundary areas. Variational level set contour is then applied to further remove noises and locate moving areas. Our contribution is to incorporate in this procedure the special property of camera movement on a PTU to realize automatic tracking and efficient background matching for motion detection. As demonstrated in the experiment, the proposed method well outlines the foreground objects. Song Ci, Yanwei Liu 0001, Hui Tang 0001 |
VCIP | 3 |
| 2010 | View synthesis error analysis for selecting the optimal QP of depth map coding in 3D video applicationabstractIn 3D video communication, how to select the appropriate quantization parameter (QP) for depth map coding is very important for obtaining the optimal view synthesis quality. This paper first analyzes the depth uncertainty induced two kinds of view synthesis errors, namely the original depth error induced view synthesis error and the depth compression induced view synthesis error, and then proposes a quadratic model to characterize the relationship between the view synthesis quality and the depth quantization step size. The proposed model can find the inflexion point in the curve of the view synthesis quality with the increasing depth quantization step size. Experimental results show that, given the rate constraint for depth map, the proposed model can accurately find the optimal QP for depth map coding. Yanwei Liu 0001, Song Ci, Hui Tang 0001 |
PCS | 1 |
| 2010 | RD-optimized interactive streaming of multiview video with multiple encodings
Yanwei Liu 0001, Qingming Huang, Siwei Ma 0001, Debin Zhao, Wen Gao 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2009 | Compression-Induced Rendering Distortion Analysis for Texture/Depth Rate Allocation in 3D Video CompressionabstractIn 3D video applications, the virtual view is generally rendered by the compressed texture and depth. The texture and depth compression with different bit-rate overheads can lead to different virtual view rendering qualities. In this paper, we analyze the compression-induced rendering distortion for the virtual view. Based on the 3D warping principle, we first address how the texture and depth compression affects the virtual view quality, and then derive an upper bound for the compression-induced rendering distortion. The derived distortion bound depends on the compression-induced depth error and texture intensity error. Simulation results demonstrate that the theoretical upper bound is an approximate indication of the rendering quality and can be used to guide sequence-level texture/depth rate allocation for 3D video compression. Yanwei Liu 0001, Siwei Ma 0001, Qingming Huang, Debin Zhao, Wen Gao 0001, Nan Zhang 0015 |
DCC | 1 |
| 2009 | Joint video/depth rate allocation for 3D video coding based on view synthesis distortion model
Yanwei Liu 0001, Qingming Huang, Siwei Ma 0001, Debin Zhao, Wen Gao 0001 |
Signal Process. Image Commun. | 1 |
| 2007 | Low-delay View Random Access for Multi-view Video CodingabstractMulti-view video coding is becoming a very active research topic, as multi-view video system provides the interactive feature which makes viewers experience the free viewpoint navigation within the range covered by the shooting cameras compared with the traditional single view video. Multi-view video coding aims at compressing the redundancy between views besides temporal correlations within each view. Rapid view random access is a basic requirement for multi-view video communication. Though inter-view prediction enhances the coding efficiency, it limits the rapid view random access capability. In this paper, we propose three approaches to provide low-delay view random access capability while keeping the high rate-distortion performance. The proposed techniques, including SP/SI frame coding, interleaved view coding and secondary representation coding, vastly reduce the decoding delays while view random access occurs and greatly improve the ability of switching between views. Yanwei Liu 0001, Qingming Huang, Debin Zhao, Wen Gao 0001 |
ISCAS | 1 |