Kebin Jia

dblp:61/5185 · also Ke-Bin Jia, Ke-bin Jia · DBLP profile ↗
← Back
12ranked-venue papers in the field
0as first author
4since 2021 · last 2023
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 9Data Mining & Knowledge Discovery · 2Other / Interdisciplinary · 1
YearPublicationVenuePosition
2023 Spatio-Temporal Information Fusion Network for Compressed Video Quality Enhancement
abstract
Video is often compressed by standard compression algorithms to facilitate storage and transmission. Compressed video will produce artifacts that affect the video quality. How to improve the quality of compressed video during video post-processing has become an important topic in the multimedia field. This paper proposes a Spatio-Temporal Information Fusion Network for quality enhancement of compressed video, as shown in Fig. 1. The algorithm comprises two parts. In the first part, we use 3D convolution to build a U-shaped network to model the temporal dynamics between input frames. We concatenate features with the same spatial resolution from shallow layers to deep layers by skip-connecting merging channels, which helps the local information of shallow- generated features to reach the output. In the second part, we designed a quality enhancement module to fully mine the spatio-temporal information extracted in the first part, cut the feature map on the time dimension t, and then extract the feature map separately on the spatial dimension and refine the feature. The network is trained in an end-to-end manner, and the data sets are selected from the database Xiph (Xiph.org) and VQEG. We use the H.265/HEVC reference software HM16.5 to compress the video to evaluate the performance of the model under different compression levels. The experimental results show that the average PSNR of 18 HEVC standard sequences is improved by 0.88 dB, 0.86 dB, 0.81 dB and 0.72 dB when the quantization parameters(QP) are equal to 37, 32, 27 and 22, respectively, and the number of parameters are only 0.66 million.
Kebin Jia, Pengyu Liu 0001
DCC2
2022 VIMTS: Variational-based Imputation for Multi-modal Time Series
abstract
Multi-modal time series data in real applications often contain data of different dimensionalities, e.g., high-dimensional modality such as image data series, and low-dimensional univariate time series. Multi-modal time series data with missing high-dimensional modal values are ubiquitous in real-world classification and regression applications. To accurately predict the target labels, it is important to appropriately impute the high-dimensional modal missing values. However, most existing imputation methods focus on multivariate time series, fail to simultaneously consider temporal dependencies within each series and the correlations across the series, and also lack a probabilistic interpretation. In this paper, we propose a novel method, which uses a new structured variational approximation technique for the imputation of missing values in multi-modal time series. Instead of directly imputing high-dimensional modal missing values, we use the variational approximation technique to impute intermediate lower-dimensional feature representations of high-dimensional modal missing values from simple modalities related to high-dimensional modality and then feed them into a dynamical model. The dynamical model captures the temporal dependencies of the feature representations and finally predicts the target labels. In order to address the optimization difficulties caused by the lack of ground truth values of lower-dimensional feature representations, we also propose a two-stage isolated optimization strategy for better convergence. We evaluate our method on a real-world stream monitoring dataset. Our extensive experiments demonstrate that the proposed method outperforms several state-of-the-art methods in both data imputation and prediction performance.
Kebin Jia, Benjamin H. Letcher, Jennifer H. Fair, Yiqun Xie, Xiaowei Jia
IEEE Big Data2
2022 SAQENet: A Quality Enhancement Network for Compressed Video with Self-attention
abstract
Existing block-based encoding frameworks often use inaccurate quantification and motion compensation techniques, which result in many compression artifacts due to the loss of high-frequency information. In particular, the blurring of content edges and significant compression distortion can negatively impact the subjective video quality given limited coding resources. Hence, there is an urgent need to build a quality enhancement method for improving the quality of the compressed video at the receiving end given the same coding resources.
Pengyu Liu 0001, Kebin Jia, Shanji Chen
DCC3
2021 3D-CVQE: An Effective 3D-CNN Quality Enhancement for Compressed Video Using Limited Coding Information
abstract
How to obtain higher quality reconstructed video within limited coding resources is a research focus for video coding. First, the fluctuating quality, missing pixel and position fluctuation characteristics of the compressed video are found in this paper based on the block-based coding frameworks. Then, in order to reduce the degradation of compressed video quality caused by the above three characteristics, an non-aligned 3D-CNNs for compressed video quality enhancement called "3D-CVQE" is proposed, which can preserve and utilize limited input information in both temporal and spatial domains effectively. Finally, the experiments results validate the effectiveness and generalization ability of the proposed 3D-CVQE approach in the quality enhancement of compressed video.
Pengyu Liu 0001, Kebin Jia
DCC3
2020 Sequence Matching with Discriminative Binary Features for Robust and Fast Light-Rail Localization at High Frame Rate
abstract
As an essential part of the advanced driving assistant system, visual localization technology has drawn much attention. To solve the interference caused by the high similarity of scenes at a high frame rate, this paper proposes a localization method based on global-local features and keyframe retrieval. In this method, the semantic features of the scene extracted by the fused semantic segmentation model are used to guide the detection of the key regions with discriminative information. By using the screening strategy based on pixel position cues and unsupervised learning algorithm, the binary feature descriptor with strong discrimination power is extracted, which can improve the matching accuracy and reduce the computational complexity. Secondly, combining the keyframes obtained by calculating scene discriminative score with a sequence matching algorithm can promote localization performance. The experiments conducted on the Nordland dataset and the challenging Hong Kong light rail dataset. The experiment results show that our method has higher computing efficiency and locating accuracy than the state-of-the-art ConvNet featurebased localization method named SeqCNNSLAM under the condition of drastic changes in appearance. Moreover, compared with SeqSLAM which is based on global features, the precision of our approach has significant improvement while meeting real-time requirements.
Tingxian Wang, Kebin Jia
IEEE BigData2
2020 Fast Depth Intra Coding Based on Layer-Classification and CNN for 3D-HEVC
abstract
View synthesis optimization (VSO) introduces heavy computational complexity caused by the VSO-based iterative search of all possible quad-tree partitions. To reduce the complexity, this paper proposes a convolutional neural network (CNN) scheme based on layer-classification for fast depth intra coding. First, a layer-classification model based on texture smoothness is proposed to determine the smoothest depth map. Then, a CNN network incorporating SENet (CNN-SENet) structure is designed and trained. Finally, the layer-classification model and the CNN-SENet network are combined to predict the coding unit (CU) partition of all coding units (CUs) for depth map at a specific view.
Kebin Jia, Pengyu Liu 0001, Zhonghua Sun 0003
DCC2
2020 Fast CU Size Decision Using Machine Learning for Depth Map Coding in 3D-HEVC
abstract
3D-High Efficiency Video Coding (3D-HEVC) is a video compression standard developed for multi-view video plus depth map coding based on the latest HEVC coding standard. We propose an eXtreme Gradient Boosting (XGBoost) system based fast coding unit (CU) level decision for depth maps, which is used to solve the problem of high coding complexity caused by the addition of depth maps and new coding tools in 3D-HEVC. We explore the application of data mining and machine learning in video coding by using texture feature attributes that are highly correlated with CU size. The algorithm is mainly divided into three parts as shown in Figure 1. The algorithm comprises two parts: Models training and fast CU segmentation decision. In the first part, we use data mining and machine learning to construct the decision models by using the texture information of the depth maps as the feature attribute vectors and whether the current CU continues to be divided into sub-CUs as class labels. In the second part, feature attributes were extracted from the coding process, and the trained models were used to decide if the CU continues to partition. Experimental results demonstrated that proposed algorithm yields average 43.52% encoding time reduction with 0.12% BD-rate decrease on V/T and 0.37% increase on S/T under the all intra configuration, compared with the reference software HTM-16.0. In addition, compared with related work, the proposed method achieves different degrees of improvement in coding performance.
Ruyi Zhang 0001, Kebin Jia, Pengyu Liu 0001
DCC2
2019 Fast PU Intra Mode Decision in Intra HEVC Coding
abstract
As an upgrade for H.264/AVC, high efficiency video coding (HEVC) achieves 50% bitrate reduction under the equivalent visual quality. However, high computational complexity increases dramatically for adopting up to 35 intra prediction modes. To deal with this issue, we explore spatial-temporal correlation between PUs to narrow rough mode decision (RMD) candidate list and rate distortion optimization (RDO) candidate list respectively. A fast PU intra mode decision scheme is proposed for HEVC fast intra encoding. For RMD, the proposed scheme early determines the impossible range of the PU optimal mode in line with spatial statistical analysis theory. It would narrow RMD candidate list and speed up RMD process. First, 4 subsets (Sb1, Sb2, Sb3 and Sb4) are defined in mode set with 33 directional prediction modes as shown in Figure 1. And then, on the basis of spatial statistical analysis theory, the impossible range of the parent PU optimal mode would be judged by its known optimal mode. Last, low probability subsets of current PU are removed, and candidate subsets are determined for current PU. Where ModeParBest and ModeCurBest are the parent PU optimal mode and the current PU optimal mode respectively. For RDO, the proposed scheme tries to combine temporal correlation with intra coding, which adds co-located optimal mode of the previous frame into the RDO list of the current PU. At the same time, the modes in RDO list are decreased focusing on most time-consuming PUs (4 4 PUs, 8 8 PUs) from 8 candidate modes to 3 candidate modes. Experimental results demonstrate that the proposed scheme yields average 29% encoding time reduction with average 1.19% BDBR gain and 0.06dB BDPSNR loss compared with HM16.9. Further, the proposed method implements fast CU encoding without additional computation during the encoding process.
Kun Duan, Pengyu Liu 0001, Zeqi Feng, Kebin Jia
DCC4
2018 MuVAN: A Multi-view Attention Network for Multivariate Temporal Data
abstract
Recent advances in attention networks have gained enormous interest in time series data mining. Various attention mechanisms are proposed to soft-select relevant timestamps from temporal data by assigning learnable attention scores. However, many real-world tasks involve complex multivariate time series that continuously measure target from multiple views. Different views may provide information of different levels of quality varied over time, and thus should be assigned with different attention scores as well. Unfortunately, the existing attention-based architectures cannot be directly used to jointly learn the attention scores in both time and view domains, due to the data structure complexity. Towards this end, we propose a novel multi-view attention network, namely MuVAN, to learn fine-grained attentional representations from multivariate temporal data. MuVAN is a unified deep learning model that can jointly calculate the two-dimensional attention scores to estimate the quality of information contributed by each view within different timestamps. By constructing a hybrid focus procedure, we are able to bring more diversity to attention, in order to fully utilize the multi-view information. To evaluate the performance of our model, we carry out experiments on three real-world benchmark datasets. Experimental results show that the proposed MuVAN model outperforms the state-of-the-art deep representation approaches in different real-world tasks. Analytical results through a case study demonstrate that MuVAN can discover discriminative and meaningful attention scores across views over time, which improves the feature representation of multivariate temporal data.
Ye Yuan 0006, Guangxu Xun, Fenglong Ma, Yaqing Wang 0001, Nan Du 0001, Kebin Jia, Lu Su 0001, Aidong Zhang 0001
ICDM6
2017 Wave2Vec: Learning Deep Representations for Biosignals
abstract
Time series data mining has gained increasing attention in health domain. Recently, researchers attempt to employ Natural Language Processing (NLP) to health data mining, in order to learn proper representations of discrete medical concepts from Electronic Health Records (EHRs). However, existing models do not take continuous physiological records into account, which are naturally existed in EHRs. The major challenges for this task are to model non-obvious representations from observed high dimensional biosignals, and to interpret the learned features. To address these issues, we propose Wave2Vec, an end-to-end deep learning model, to bridge the gap between biosignal processing and language modeling. Wave2Vec jointly learns both inherent and embedding representations of biosignals at the same time. To evaluate the performance of our model in clinical task, we carry out experiments on two real world benchmark biosignal datasets. Experimental results show that the proposed Wave2Vec model outperforms the six feature leaning baselines in biosignal processing.
Ye Yuan 0006, Guangxu Xun, Qiuling Suo, Kebin Jia, Aidong Zhang 0001
ICDM4
2016 HEVC Fast CU Encoding Based Quadtree Prediction
abstract
Quadtree brings extremely high computational complexity in high efficiency video coding (HEVC). Innovative works for improving quadtree structures to further reduce encoding time are stated in this paper. A novel quadtree probability mechanism is proposed for HEVC fast coding unit (CU) encoding. Firstly, this paper makes an in-depth study of the relationship among CU distribution, quantization parameter (QP) and video content change. Secondly, a CU quadtree probability model is proposed for modeling and predicting CU partition based on the group of picture (GOP). Eventually, a CU quadtree probability update is proposed, aiming to address probabilistic model distortion problems caused by video content change. Experimental results have shown that the proposed CU quadtree probability mechanism significantly outperforms HEVC by considerably reducing encoding time by 27% for lossy coding and 42% for (visually) lossless coding, without compromising rate-distortion (RD) performance.
Pengyu Liu 0001, Yueying Wu 0003, Kebin Jia
DCC4
2016 Development and Application of Mobile Nursing System in Obstetrics
abstract
With the development of healthcare technology and medical standard, the demands of the quality and efficiency of mobile nursing are increased significantly, hence it is necessary to improve hospital nursing mechanism dealing with massive tasks. In this paper, we develop an optimized mobile nursing system based on the specific nurse workflow analysis in obstetrics. An efficient integration system framework is proposed combined with existing hospital common systems and WLAN network environment, which implements automatically execute data interaction. A combination model of C/S using PDA on hand and B/S using PC working at nurse workstation is implemented which ensure the mobility and centrality. A lighten AJAX-SSH2 development framework is used to enhance the function expansibility. The proposed system has been applied in a regional obstetrics hospital in China. Practical clinical application results prove that the proposed system can raise nursing efficiency and reduce medical error rate, make it possible to nurses paying more attention on patients. The healthcare big data collected by this system has considerable value for further research.
Ye Yuan 0006, Kebin Jia, Zhonghua Sun 0003
WI2