Jinxia Liu

dblp:54/2871 · DBLP profile ↗
← Back
40ranked-venue papers
1as first author
15since 2021 · last 2025
0000-0002-6552-9795ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 9 since 2021Computer networks · 14 · 6 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Know2Vec: A Black-Box Proxy for Neural Network Retrieval
abstract
For general users, training a neural network from scratch is usually challenging and labor-intensive. Fortunately, neural network zoos enable them to find a well-performing model for directly use or fine-tuning it in their local environments. Although current model retrieval solutions attempt to convert neural network models into vectors to avoid complex multiple inference processes required for model selection, it is still difficult to choose a suitable model due to inaccurate vectorization and biased correlation alignment between the query dataset and models. From the perspective of knowledge consistency, i.e., whether the knowledge possessed by the model can meet the needs of query tasks, we propose a model retrieval scheme, named Know2Vec, that acts as a black-box retrieval proxy for model zoo. Know2Vec first accesses to models via a black-box interface in advance, capturing vital decision knowledge from models while ensuring their privacy. Next, it employs an effective encoding technique to transform the knowledge into precise model vectors. Secondly, it maps the user's query task to a knowledge vector by probing the semantic relationships within query samples. Furthermore, the proxy ensures the knowledge-consistency between query vector and model vectors within their alignment space, which is optimized through the supervised learning with diverse loss functions, and finally it can identify the most suitable model for a given task during the inference stage. Extensive experiments show that our Know2Vec achieves superior retrieval accuracy against the state-of-the-art methods in diverse neural network retrieval tasks.
Zhuoyi Shang, Yanwei Liu 0001, Jinxia Liu, Xiaoyan Gu 0001, Xiangyang Ji
AAAI3
2025 Cross-Layer-Optimized Link Selection for Hologram Video Streaming Over Millimeter Wave Networks
abstract
Holographic-type communication brings an immersive tele-holography experience by delivering holographic contents to users. As the direct representation of holographic contents, hologram videos are naturally three-dimensional representation, which consist of a huge volume of data. Advanced multi-connectivity (MC) millimeter-wave (mmWave) networks are now available to transmit hologram videos by providing the necessary bandwidth. However, the existing link selection schemes in MC-based mmWave networks neglect the source content characteristics of hologram videos and the coordination among the parameters of different protocol layers in each link, leading to sub-optimal streaming performance. To address this issue, we propose a cross-layer-optimized link selection scheme for hologram video streaming over mmWave networks. This scheme optimizes link selection by jointly adjusting the video coding bitrate, the modulation and channel coding schemes (MCS), and link power allocation to minimize the end-to-end hologram distortion while guaranteeing the synchronization and quality balance between real and imaginary components of the hologram. Results show that the proposed scheme can effectively improve the hologram video streaming performance in terms of PSNR by 1.2 dB to$\mathbf{6. 4 d B}$against the non-cross-layer scheme.
Yiming Jiang 0002, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou
WCNC3
2024 BBR-based and fairness-guaranteed congestion control and packet scheduling for MPQUIC over heterogeneous networks
Zhenjie Deng, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Dacai Liu
Comput. Commun.3
2023 Perspectively Equivariant Keypoint Learning for Omnidirectional Images
abstract
Robust keypoint detection on omnidirectional images against large perspective variations, is a key problem in many computer vision tasks. In this paper, we propose a perspectively equivariant keypoint learning framework named OmniKL for addressing this problem. Specifically, the framework is composed of a perspective module and a spherical module, each one including a keypoint detector specific to the type of the input image and a shared descriptor providing uniform description for omnidirectional and perspective images. In these detectors, we propose a differentiable candidate position sorting operation for localizing keypoints, which directly sorts the scores of the candidate positions in a differentiable manner and returns the globally top-K keypoints on the image. This approach does not break the differentiability of the two modules, thus they are end-to-end trainable. Moreover, we design a novel training strategy combining the self-supervised and co-supervised methods to train the framework without any labeled data. Extensive experiments on synthetic and real-world 360° image datasets demonstrate the effectiveness of OmniKL in detecting perspectively equivariant keypoints on omnidirectional images. Our source code are available online at https://github.com/vandeppce/sphkpt.
Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Liming Wang 0001, Zhen Xu 0009, Xiangyang Ji
IEEE Trans. Image Process.3
2022 360-Attack: Distortion-Aware Perturbations from Perspective-Views
abstract
The application of deep neural networks (DNNs) on 360-degree images has achieved remarkable progress in the recent years. However, DNNs have been demonstrated to be vulnerable to well-crafted adversarial examples, which may trigger severe safety problems in the real-world applications based on 360-degree images. In this paper, we propose an adversarial attack targeting spherical images, called 360-attactk, that transfers adversarial perturbations from perspective-view (PV) images to a final adversarial spherical image. Given a target spherical image, we first represent it with a set of planar PV images, and then perform 2D attacks on them to obtain adversarial PV images. Considering the issue of the projective distortion between spherical and PV images, we propose a distortion-aware attack to reduce the negative impact of distortion on attack. Moreover, to reconstruct the final adversarial spherical image with high aggressiveness, we calculate the spherical saliency map with a novel spherical spectrum method and next propose a saliency-aware fusion strategy that merges multiple inverse perspective projections for the same position on the spherical image. Extensive experimental results show that 360-attack is effective for disturbing spherical images in the black-box setting. Our attack also proves the presence of adversarial transferability from Z2 to SO(3) groups.
Yanwei Liu 0001, Jinxia Liu, Jingbo Miao, Antonios Argyriou, Liming Wang 0001, Zhen Xu 0009
CVPR3
2022 SP Attack: Single-Perspective Attack for Generating Adversarial Omnidirectional Images
abstract
The safety of Deep Neural Networks (DNNs) processing omnidirectional images (ODIs) is an under-researched topic. In this paper, we propose a novel sparse attack, named Single-Perspective (SP) Attack, towards fooling these models by perturbing only one perspective image (PI) rendered from the target ODI. The attack is launched from the perspective domain, and finally the perturbation is transferred to the original ODI. To this end, we propose an effective PI position searching algorithm based on Bayesian Optimization, and then corrupt the PI centered on the desirable position with unconstrained/constrained perturbations. Extensive experiments on synthetic and real-world omnidirectional datasets demonstrate that SP Attack can overcome the projection deformation of ODIs, and mislead the neural networks by limiting the perturbations in a single patch on the target ODI.
Yanwei Liu 0001, Jinxia Liu, Pengwei Zhan, Liming Wang 0001, Zhen Xu 0009
ICASSP3
2022 Variational Depth Estimation on Hypersphere for Panorama
abstract
Depth estimation for panorama is a key part of 3D scene understanding, and adopting discriminative models is the most common solution. However, due to the rectangular convolution kernel, these existing learning methods cannot efficiently extract the distorted features in panoramas. To this end, we propose OmniVAE, a generative model based on Conditional Variational Auto-Encoder (CVAE) and von Mises-Fisher (vMF) distribution, to strengthen the exclusive generative ability for spherical signals by mapping panoramas to hypersphere space. Further, to alleviate the side effects of manifold-mismatching caused by non-planar distribution, we put forward the Atypical Receptive Field (ARF) module to slightly shift the receptive field of the network and even take the distribution difference into account in the reconstruction loss. The quantitative and qualitative evaluations are performed on real-world and synthetic datasets, and the results show that OmniVAE outperforms the state-of-the-art methods.
Jingbo Miao, Yanwei Liu 0001, Jinxia Liu, Zhen Xu 0009
ICIP4
2022 Viewport-Oriented Panoramic Image Inpainting
abstract
Panoramic images are usually viewed through Head Mounted Displays (HMDs), which renders only a narrow field of view from the raw panoramic image. This distinctive viewing feature has largely been ignored when inpainting panoramic images. To address this issue, we propose a viewport-oriented generative adversarial panoramic image inpainting network in this paper. For capturing the distorted features accurately in the generating process of equirectangular projection (ERP) panoramic image, a latitude-adaptive feature fusion module is devised to aggregate the latitude-level features in ERP image and less-distorted patch-level viewport-domain features. Furthermore, a novel cross-domain discriminator is proposed to force the inpainting network to generate more plausible results in viewports. Extensive experiments show that our model achieves better performance compared to the baseline methods, especially in the viewport images.
Zhuoyi Shang, Yanwei Liu 0001, Guoyi Li, Jingbo Miao, Jinxia Liu, Liming Wang 0001
ICIP6
2022 Six-to-one: Cubemap-guided Feature Calibration for Panorama Object Detection
abstract
Object detection methods for perspective images have proven increasingly efficient, but the techniques for equirectangular projection (ERP) panoramas from inherently spherical imaging cannot still achieve satisfactory performance. Due to the various degrees of distortion at different pixel locations, current algorithms cannot adapt to the changes in shape and contour caused by stretching, which results in performance degradation when migrating them from perspective images to spherical ones. In this paper, we improve the network for panorama object detection and introduce the cube-domain information with discontinuity but low distortion to correct the panorama features. Unlike previous works, we consider the impact of semantic discontinuity from all tangent planes instead of overlaying features when needed. Considering the six facets as unified, i.e., six-to-one for extraction, the proposed Facet-Link module enhances the long-range sensing capability at the facet level in the frequency domain. Moreover, the position alignment packs different facets, i.e., six-to-one for calibration, to preserve more global signals during the correction stage, which establishes semantic pathways for feature interactions between panorama and cubemap in the two dimensions, facet-facet and cube-pano, respectively. Extensive experiments on synthetic and real-world datasets verify the effectiveness and robustness of our proposed method.
Jingbo Miao, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Yanni Han, Zhen Xu 0009
ICTAI4
2021 Improved Face Detector on Fisheye Images via Spherical-Domain Attention
abstract
As one type of omnidirectional projection, fisheye images have been widely used in automatic driving and visual surveillance. However, they cannot be processed well by the traditional algorithms designed for the planar rectilinear images since they usually suffer from severe geometric distortion during image formation. In this paper, the conventional face detection algorithm is enhanced to fit the fisheye images via combining with the spherical convolution block by learning rotation-invariant features from the spherical domain. The learned features from both planar and spherical domains are subsequently mixed by the spatial attention mechanism. Consequently, the whole network can automatically learn the distorted features directly from different positions on the target image. Experimental results verify that our network can detect distorted faces on fisheye images effectively and maintain the performance on traditional planar images.
Jingbo Miao, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Zhen Xu 0009, Yanni Han
ISCC3
2021 Output Security for Multi-user Augmented Reality using Federated Reinforcement Learning
abstract
With the rapid advancements in Augmented Reality, the number of AR users is gradually increasing and the multiuser AR ecosystem is on the rise. Currently, AR applications usually present results without limitations, which causes great latent danger to users, so it is necessary to apply strategies to ensure the safe output of AR. Due to the environmental diversities among the distributed users, the traditional approaches designed for single-user AR are not efficient for multi-user AR applications. Considering the characteristics of multi-user AR scenarios, we propose a multi-user AR output strategy model based on Federated Reinforcement Learning. With the device-fog-cloud hierarchical architecture, the proposed models are obtained first by Reinforcement Learning on users' devices, and are then hierarchically aggregated on the fog nodes and cloud server. We performed extensive AR simulations in Unity and obtained the results that show our method can avoid several security problems existent in multi-user AR applications.
Fengchao Wang, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Liming Wang 0001, Zhen Xu 0009
ISCC3
2021 D2D-Assisted Federated Learning in Mobile Edge Computing Networks
abstract
With the proliferation of edge intelligence and the breakthroughs in machine learning, Federated Learning (FL) is capable of learning a shared model across several edge devices by preserving their private data from being exposed to external adversaries. However, the distributed architecture of FL naturally introduces communication between the central parameter server and the distributed learning nodes. The huge communication cost poses a challenge to practical FL, especially for FL in mobile edge computing (MEC) networks. Existing communication-efficient FL systems predominantly optimize their intrinsic learning process and are not concerned with the implications on the network. In this paper we propose a FL scheme that leverages Device-to-Device (D2D) communication (hence called D2D-FedAvg) and is suitable for mobile edge networks. D2D-FedAvg creates a two-tier learning model where D2D learning groups communicate their results as a single entity to the MEC server leading to traffic reduction. We propose the schemes for D2D grouping, master UE selection, and also D2D exit in the learning process and then form a complete D2D-assisted federated averaging algorithm. Via extensive simulations on the Federated Extended MNIST dataset, the feasibility and convergence of D2D-FedAvg scheme are evaluated. Our results show that D2D-FedAvg lowers the communication cost relative to the typical Federated Averaging (FedAvg) in cellular networks as the number of users is increased (for 100 cellular users 37% traffic reduction), while keeping the same learning accuracy with FedAvg across the board.
Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Yanni Han
WCNC3
2021 Tile caching for scalable VR video streaming over 5G mobile networks
Kedong Liu, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou
J. Vis. Commun. Image Represent.3
2021 Cross-layer DASH-based multipath video streaming over LTE and 802.11ac networks
Zhenjie Deng, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou
Multim. Tools Appl.3
2021 360-Degree VR Video Watermarking Based on Spherical Wavelet Transform
abstract
Similar to conventional video, the increasingly popular 360 virtual reality (VR) video requires copyright protection mechanisms. The classic approach for copyright protection is the introduction of a digital watermark into the video sequence. Due to the nature of spherical panorama, traditional watermarking schemes that are dedicated to planar media cannot work efficiently for 360 VR video. In this article, we propose a spherical wavelet watermarking scheme to accommodate 360 VR video. With our scheme, the watermark is first embedded into the spherical wavelet transform domain of the 360 VR video. The spherical geometry of the 360 VR video is used as the host space for the watermark so that the proposed watermarking scheme is compatible with the multiple projection formats of 360 VR video. Second, the just noticeable difference model, suitable for head-mounted displays (HMDs), is used to control the imperceptibility of the watermark on the viewport. Third, besides detecting the watermark from the spherical projection, the proposed watermarking scheme also supports detecting watermarks robustly from the viewport projection. The watermark in the spherical domain can protect not only the 360 VR video but also its corresponding viewports. The experimental results show that the embedded watermarks are reliably extracted both from the spherical and the viewport projections of the 360 VR video, and the robustness of the proposed scheme to various copyright attacks is significantly better than that of the competing planar-domain approaches when detecting the watermark from viewport projection.
Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Siwei Ma 0001, Liming Wang 0001, Zhen Xu 0009
ACM Trans. Multim. Comput. Commun. Appl.2
2019 Joint EPC and RAN Caching of Tiled VR Videos for Mobile Networks
Kedong Liu, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou
MMM (1)3
2019 Location Recognition Algorithm for Vision-Based Industrial Sorting Robot via Deep Learning
abstract
In this paper, the deep convolutional neural network (DCNN) is applied to locating and recognizing complex workpieces automatically for the vision-based sorting robot in industrial production process. Firstly, in order to obtain the location of workpieces, the pixel projection algorithm (PPA), which consists of pre-procession and pixel projection operation, is presented to eliminate uneven illumination, and locate and segment workpieces images. Then, we get the objective information and identify the object by training DCNN, which is used to recognize the rational degree and type of workpieces at a high rate of speed. Finally, experimental results prove the validity of the location-recognition algorithms for the vision-based sorting robot. The location error and recognition accuracy can be significantly improved in the experimental environment.
Xiru Wu, Xingyu Ling, Jinxia Liu
Int. J. Pattern Recognit. Artif. Intell.3
2019 MEC-Assisted Panoramic VR Video Streaming Over Millimeter Wave Mobile Networks
abstract
Panoramic virtual reality video (PVRV) is becoming increasingly popular since it offers a true immersive experience. However, the ultra-high resolution of PVRV requires significant bandwidth and ultra-low latency for PVRV streaming, something that makes challenging the extension of this application to mobile networks. Besides bandwidth, the frequent perspective viewport rendering induces a heavy computational load on battery-constrained mobile devices. To attack these problems jointly, this paper proposes a PVRV streaming system that is designed for modern multiconnectivity-based millimeter wave (mmWave) cellular networks in conjunction with mobile edge computing (MEC). First, mmWave is deployed to support the high bandwidth needs of PVRV streaming. Next, the multiple mmWave links that tend to suffer from outages are coupled with a sub-6 GHz link to ensure disruption-free wireless communication. With the help of an MEC server, the tradeoff among link adaptation, transcoding-based chunk quality adaptation, and viewport rendering offloading is sought to improve the wireless bandwidth utilization and mobile device's energy efficiency. Simulation results show that the proposed scheme can improve the streaming performance in both energy efficiency and the quality of received viewport over the state-of-the-art schemes.
Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Song Ci
IEEE Trans. Multim.2
2018 Binocular-Combination-Oriented Perceptual Rate-Distortion Optimization for Stereoscopic Video Coding
abstract
In the chain of stereoscopic video processing, stereoscopic video coding and viewing are usually two independent stages. Conventional stereoscopic video coding puts an emphasis on improving the coding efficiency by seeking the optimal tradeoff between the coding bit rate and the signal-based distortion, while neglecting the perceptual behaviors of binocular combination when stereoscopic video is viewed by human beings. In this paper, we propose to utilize binocular combination to optimize the stereoscopic video coding from the perspective of perceptual quality measurement. Specifically, we propose a novel binocular-combination-oriented measurement for visual distortion and then derive the Lagrange multiplier for the binocular-combination-oriented rate-distortion optimization (RDO). Via extensive subjective tests, the results show that the proposed perceptual RDO can save more than 5% BD rate over the traditional RDO in multiview extension of High Efficiency Video Coding for stereoscopic video coding.
Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Song Ci
IEEE Trans. Circuits Syst. Video Technol.2
2018 3DQoE-Oriented and Energy-Efficient 2D plus Depth Based 3D Video Streaming Over Centrally Controlled Networks
abstract
IP networks have become the dominant platform for video delivery. However, bandwidth-hungry video is pushing networks to their limits: costs are rising for the operators and the viewing experience is not always satisfactory for the users. When considering 3D video delivery, the previous problems are exacerbated because of the higher volume of data that must be communicated, and the difficulty in characterizing the viewing experience of the end user. Consequently, network operators may be reluctant to deliver 3D video due to costs and unclear quality improvements to their users. In this setting, the true immersive experience of 3D video remains elusive. In this paper, we focus on the efficient delivery of 3D video in terms of quality and energy cost over centrally controlled networks. As a representative example of a centrally controlled network, a software-defined network (SDN) is assumed. Our approach is based on a comprehensive network-dependent 3D quality of experience (3DQoE) model and an energy cost model for 3D video streaming. By using the developed models, we formulate the problem of energy-efficient and 3DQoE-optimized 3D video flow path routing. The particular characteristic of video/depth rate allocation presented in 3D video is embedded seamlessly into the selection of the optimal routing paths for multiple 3D video streams. The formulated problem is NP-hard and is solved with a heuristic algorithm based on the branch-and-bound method after significant reduction of the solution search space. Extensive 3D video streaming experiments are conducted over an OpenFlow-based SDN with subjective and objective evaluations and they highlight the significant benefits of the proposed approach.
Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Song Ci
IEEE Trans. Multim.2
2017 Joint Source Encoding and Networking Optimization for Panoramic Video Streaming over LTE-A Downlink
abstract
With the increasing capacity of wireless networks, more people would like to consume the 360-degree panoramic video (PV) in virtual reality (VR) applications as its immersive experience. However, due to the super-high resolution of the PV and the dynamic features of wireless networks, it is very difficult to efficiently deliver PVs over wireless networks. The traditionally independent PV encoding and networking sometimes also results in the PV quality deterioration since it neglects the harmony between the source encoding and networking. In this paper, a joint source encoding and networking optimization scheme is proposed to transmit the PV over LTE-A downlink. The PV encoding parameters during the source compression, the modulation and coding scheme (MCS), and relay selection during the networking are jointly considered to optimize the end-to-end PV quality. In addition, the video quality for region of interest (RoI, the possible viewport region) is enhanced by allowing a larger latency bound in the joint source encoding and networking optimization. Experimental results show that the proposed scheme achieves significant performance improvement for the quality of the received PV over traditional PV streaming approaches.
Kedong Liu, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Xinghua Yang
GLOBECOM3
2017 Cross-layer optimized authentication and error control for wireless 3D medical video streaming over LTE
Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Song Ci
J. Vis. Commun. Image Represent.2
2016 Cross-network and cross-layer optimized video streaming over LTE and WCDMA downlink
abstract
Video services are proliferating over today's mobile Internet and efforts have been made to improve their performance. The cross-layer video streaming that jointly optimizes the parameters at different protocol layers is a feasible solution but without bandwidth aggregation of multiple wireless networks. The cross-network video streaming can realize the bandwidth aggregation by using multiple overlapping wireless networks simultaneously. However, parameters at different protocol layers are optimized independently in the cross-network optimization. In this paper, we propose a joint cross-network and cross-layer optimized video streaming scheme that utilizes the bandwidth aggregation of cross-network video streaming and further improve the performance by jointly optimizing the parameters of different protocol layers in each network with a cross-layer manner. In the proposed scheme, the LTE and WCDMA networks are adopted. The bit-rate of the video at application layer, the rate allocation among networks and the parameters of physical layers in each network are jointly optimized. Experimental results show that the proposed scheme gains higher quality of experience in terms of PSNR than the state-of-the-art schemes.
Zhenjie Deng, Yanwei Liu 0001, Jinxia Liu, Xin Chen 0019, Antonios Argyriou, Zhen Xu 0009, Song Ci
ISCC3
2016 Choquet integral based QoS-to-QoE mapping for mobile VoD applications
abstract
Today, how to accurately predict the quality of experience (QoE) of the networking service is a very important issue for the network operator to optimize the service. However, due to the complex multi-dimensional characteristics of QoE, QoE estimation is extremely challenging. With utilizing the advantages of quality of service (QoS) in evaluating the networking performance, we exploit QoS/QoE correlation to predict QoE by building a QoSto-QoE mapping relationship. To fully consider the inter-dependency among QoS parameters towards forming the QoE, a Choquet integral based fuzzy measurement method is used to map QoS to QoE. Via extensive experiments in mobile VoD applications, the advancement and effectiveness of the proposed method are verified.
Yanwei Liu 0001, Jinxia Liu, Zhen Xu 0009, Song Ci
IWQoS2
2016 Perceptual rate-distortion optimization for H.264/AVC video coding from both signal and vision perspectives
Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Ruixiao Yao, Song Ci
Multim. Tools Appl.3
2016 SSIM-based error-resilient cross-layer optimization for wireless video streaming
Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Antonios Argyriou, Song Ci
Signal Process. Image Commun.3
2015 Transmit power aware cross-layer optimization for LTE uplink video streaming
abstract
The rapid developments of advanced wireless communication technologies and mobile devices are boosting the uplink multimedia applications. In this paper, a transmit power aware cross-layer optimization scheme is proposed to achieve a good trade-off between the transmit power and the perceived video quality for Long Term Evolution uplink video streaming. Specifically, the video coding quantization parameter and encoding mode at the application layer, and the uplink transmit power as well as modulation and coding scheme at the physical layer are jointly adjusted in the cross-layer optimization. To further improve the perceptual video experience for the end user with limited transmission resources, unequal quality control is performed by enhancing the video quality of region of interest. Additionally, the structural similarity is adopted as the video quality measurement metric to make the optimized video properly preserve the structural information during the cross-layer optimization process. Experimental results show that significant performance improvements in terms of the transmit power reduction and the perceptual video quality are achieved for the proposed transmit power aware cross-layer optimization scheme.
Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Ruixiao Yao, Song Ci, Antonios Argyriou
ICC3
2015 Scalable 3D video streaming over P2P networks with playback length changeable chunk segmentation
Yanwei Liu 0001, Jinxia Liu, Junping Song, Antonios Argyriou
J. Vis. Commun. Image Represent.2
2015 Utility-Based H.264/SVC Video Streaming Over Multi-Channel Cognitive Radio Networks
abstract
In a cognitive radio (CR) network, a secondary user (SU) with multiple interfaces is capable of accessing multiple CR channels in an opportunistic fashion. Therefore, the available channel resources may change dramatically, and the reliabilities of the multiple accessed CR channels are also time-varying. Video streaming in such a multi-channel CR network faces great challenges in guaranteeing the quality of the received video. To deal with these challenges, we adopt H.264/SVC encoded video as the source and firstly optimize the video streaming from the perspective of exploiting more channel resources for the SU by developing a flexible sensing-transmission scheme for opportunistic spectrum access (OSA). In this scheme, the primary user activity, channel sensing result, and channel sensing accuracy are all considered in reducing the unnecessary channel sensings and correspondingly extending the transmission duration for the SU. Based on the flexible sensing-transmission scheme, we then propose a utility-based H.264/SVC video transmission scheme to further improve the expected video quality at the receiver. Specifically , the network abstraction layer units (NALUs) in the SVC video are assigned utilities which accurately reflect their contributions to the video quality, and the total effective utility of expected received video is maximized through perfectly dispatching the NALUs over the multiple CR channels. Both analytical studies and experimental results validate the effectiveness and efficiency of the proposed method.
Ruixiao Yao, Yanwei Liu 0001, Jinxia Liu, Pinghua Zhao, Song Ci
IEEE Trans. Multim.3
2014 Hierarchical-matching based scalable video streaming over multi-channel cognitive radio networks
abstract
With its ability to promote the wireless spectrum utilization, cognitive radio (CR) is a promising technology for various broadband wireless video applications. A secondary user (SU) with multiple wireless interfaces is capable to sense and access multiple CR channels, each of which can only be accessed in an opportunistic fashion. Therefore, the amount of available channel resources may change dramatically, and the reliabilities of the multiple accessed CR channels are time-varying and different from each other. To improve the quality of the received video in such a multi-channel CR network, we first develop a flexible sensing-transmission structure for dynamic spectrum access according to primary user activity and channel sensing accuracy. Based on the proposed sensing-transmission structure, we then propose a hierarchical-matching scheme to adapt the scalable video stream to the multiple time-varying and reliability-different CR channels, where we consider the priorities and validity of the network abstraction layer units (NALUs) in the transmission scheduling. Under this scheme, the more important NALUs at the global Group Of Pictures (GOP) scope will be dispatched over the more reliable channels. Both analytical studies and experimental results show the effectiveness and efficiency of the proposed method.
Ruixiao Yao, Yanwei Liu 0001, Jinxia Liu, Pinghua Zhao, Song Ci
GLOBECOM3
2014 SSIM-based cross-layer optimized video streaming over LTE downlink
abstract
Many research efforts have been done to guarantee the quality of service for video streaming over LTE. Among them, cross-layer optimized video delivery is one of the most commonly adopted approaches. However, most existing schemes of cross-layer optimized video delivery adopt PSNR as the optimization target, but do not well consider human vision characteristics. In this paper, we adopt a new metric - Structural Similarity (SSIM) in video quality evaluation and develop a SSIM-based cross-layer optimization scheme with the suppression of propagated error for video delivery over LTE downlink. The modulation and coding scheme at the physical layer of LTE downlink is selected by jointly considering the characteristics of video packets and time-varying channel states. Correspondingly, the quantization parameter of video codec at the application layer is adjusted to make the bit rate of the video stream adapt to the varying throughput of the physical link. Moreover, the error-resilient rate-distortion optimization is adopted in the proposed cross-layer optimization to suppress the effect of error propagation on the video quality degradation. Experimental results show that the proposed cross-layer optimization scheme can maintain more structural information than the conventional schemes in the received video, which correspondingly improves the video quality perceived by the end user.
Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Ruixiao Yao, Song Ci
GLOBECOM3
2014 Intrinsic flexibility exploiting for scalable video streaming over multi-channel wireless networks
abstract
Scalable video has natural advantages in adapting to the multi-channel wireless networks. And some existing works tried to further optimize the scalable video transmission by combining the crude layer-importance mapping with some extrinsic techniques, such as Forward Error Correction (FEC) and Adaptive Modulation and Coding (AMC). However, the intrinsic flexibility of scalable video streaming over the multichannel wireless networks was neglected. In this paper, we try to exploit the intrinsic flexibility by firstly analyzing the priorities of H.264/SVC video data at the network abstraction layer unit (NALU) level, and then designing the priority-validity delivery scheme for the scalable video streaming. With this strategy, the sub-stream extraction is intelligently adjusted according to the delivery history, and the more important data in a group of pictures (GOP) will be delivered through the more reliable channels. Experimental results also validate the strategy's effectiveness in improving the objective quality and perceptual experience of the received video.
Ruixiao Yao, Yanwei Liu 0001, Jinxia Liu, Pinghua Zhao, Song Ci
VCIP3
2014 3D visual experience oriented cross-layer optimized scalable texture plus depth based 3D video streaming over wireless networks
Jinxia Liu, Yanwei Liu 0001, Song Ci, Ruixiao Yao
J. Vis. Commun. Image Represent.1
2014 SSIM-based error-resilient rate-distortion optimization of H.264/AVC video coding for wireless streaming
Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Song Ci, Ruixiao Yao
Signal Process. Image Commun.3
2013 Perceptual experience oriented transmission scheduling for scalable video streaming over cognitive radio networks
abstract
Cognitive radio (CR) promotes the utilization of the wireless spectrum through admitting the secondary users to access the primary channels in an opportunistic manner. Therefore, the real-time video streaming over the CR network is challenging due to the time-varying channels. In this paper, we propose an adaptive scalable video transmission scheduling method for video streaming over time-varying channels in CR network, in which the end-to-end perceptual visual experience is optimized systematically by considering various experience-influencing factors of the scalable video source, channel, and receiving buffer. Moreover, an enhanced content-based adaptive video playout scheme is developed to further optimize the end-to-end perceptual visual experience by decreasing the receiving buffer underflow probability. The efficiency and effectiveness of the proposed method has been validated by both analytical study and extensive experimental results.
Ruixiao Yao, Yanwei Liu 0001, Jinxia Liu, Pinghua Zhao, Song Ci
GLOBECOM3
2013 A playback length changeable 3D data segmentation algorithm for scalable 3D video P2P streaming system
abstract
Scalable 3D video P2P streaming systems can supply diverse 3D experiences for heterogeneous clients with high efficiencies. Data characteristics of the scalable 3D video make the P2P streaming efficiency more depends on the data segmentation algorithm. However, traditional data segmentation algorithm is not very appropriate for scalable 3D video P2P streaming systems. In this paper, we propose a Playback Length Changeable 3D video Segmentation (PLC3DS) algorithm. It considers the particular source-data characteristics of scalable 3D video, and provides different error resilience strengths to video and depth as well as layers with different importance levels in the transmission. The simulation results show that the proposed PLC3DS algorithm can increase the success delivery rates of chunks in more important layers, and further improve the 3D experiences of the client. Moreover, it improves the network utilization ratio remarkably.
Junping Song, Yanwei Liu 0001, Jinxia Liu, Song Ci, Yan Zhang 0013
ICME3
2013 Low-complexity content-adaptive Lagrange multiplier decision for SSIM-based RD-optimized video coding
abstract
The SSIM-based rate distortion optimization (R-DO) has been proved to be an effective way to promote the perceptual video coding performance, and the Lagrange multiplier decision is the key to the SSIM-based RD-optimized video coding. Through extensively analyzing the characteristics of SSIM-based and SSE-based video distortions, this paper presents a low-complexity content-adaptive Lagrange multiplier decision method. The proposed method first estimates frame-level SSIM-based Lagrange multiplier by scaling the traditional SSE-based Lagrange multiplier with the ratio of SSE-based distortion to SSIM-based distortion. Via predicting the macroblock's perceptual importance in the whole frame, the macroblock-level Lagrange multiplier is further refined to promote the accuracy of the Lagrange multiplier decision. Experimental results show that the proposed method can obtain almost the same rate-SSIM performance and subjective quality as the state-of-the-art SSIM-based RD-optimized video coding methods with lower computation overheads.
Pinghua Zhao, Yanwei Liu 0001, Jinxia Liu, Ruixiao Yao, Song Ci, Hui Tang 0001
ISCAS3
2013 Joint video/depth/FEC rate allocation with considering 3D visual saliency for scalable 3D video streaming
abstract
For robust video plus depth based 3D video streaming, video, depth and packet-level forward error correction (FEC) can provide many rate combinations with various 3D visual qualities to adapt to the dynamic channel conditions. Video/depth/FEC rate allocation under the channel bandwidth constraint is an important optimization problem for robust 3D video streaming. This paper proposes a joint video/depth/FEC rate allocation method by maximizing the receiver's 3D visual quality. Through predicting the perceptual 3D visual qualities of the different video/depth/FEC rate combinations, the optimal GOP-level video/depth/FEC rate combination can be found. Further, the selected FEC rates are unequally assigned to different levels of 3D saliency regions within each video/depth frame. The effectiveness of the proposed 3D saliency based joint video/depth/FEC rate allocation method for scalable 3D video streaming is validated by extensive experimental results.
Yanwei Liu 0001, Jinxia Liu, Song Ci
VCIP2
2012 Integrating stereoscopic image transcoding with retargeting for mobile streaming
abstract
With the progresses in 3D imaging technologies, mobile stereoscopic images are gradually popular since it can provide the anytime and anywhere 3D viewing effects. Due to the heterogeneous transmission environments and various screen sizes of mobile devices, the stereoscopic image transcoding and retargeting are two completely independent adaptation techniques for mobile 3D image streaming. These two processing operations both deal with the disparity estimation. In this work, we propose to integrate stereoscopic 3D image transcoding with retargeting for mobile streaming. The stereoscopic image transcoding and retargeting are coupled through a straightforward link that provides the pixel-based disparity map from the retargeting stage to guide the transcoding for saving some computations while keeping efficient bit allocation. Our experimental results clearly demonstrate the advantage of coupling the transcoding with retargeting in terms of faster transcoding and saliency-based bit allocation for low bit-rate mobile streaming.
Yanwei Liu 0001, Song Ci, Jinxia Liu
VCIP3
2012 QoE-oriented 3D video transcoding for mobile streaming
abstract
With advance in mobile 3D display, mobile 3D video is already enabled by the wireless multimedia networking, and it will be gradually popular since it can make people enjoy the natural 3D experience anywhere and anytime. In current stage, mobile 3D video is generally delivered over the heterogeneous network combined by wired and wireless channels. How to guarantee the optimal 3D visual quality of experience (QoE) for the mobile 3D video streaming is one of the important topics concerned by the service provider. In this article, we propose a QoE-oriented transcoding approach to enhance the quality of mobile 3D video service. By learning the pre-controlled QoE patterns of 3D contents, the proposed 3D visual QoE inferring model can be utilized to regulate the transcoding configurations in real-time according to the feedbacks of network and user-end device information. In the learning stage, we propose a piecewise linear mean opinion score (MOS) interpolation method to further reduce the cumbersome manual work of preparing QoE patterns. Experimental results show that the proposed transcoding approach can provide the adapted 3D stream to the heterogeneous network, and further provide superior QoE performance to the fixed quantization parameter (QP) transcoding and mean squared error (MSE) optimized transcoding for mobile 3D video streaming.
Yanwei Liu 0001, Song Ci, Hui Tang 0001, Jinxia Liu
ACM Trans. Multim. Comput. Commun. Appl.5