Yao Lu 0001

dblp:26/5662-1 · DBLP profile ↗
← Back
63ranked-venue papers
1as first author
13since 2021 · last 2024
0000-0003-3760-1638ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 10 since 2021Artificial intelligence and machine learning · 29 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Revisiting Large Kernel Convolution for Light Field Image Angular Super-Resolution
abstract
Light field (LF) image angular super-resolution (LFASR) aims to reconstruct densely-sampled LF images from a sparsely-sampled sub-aperture image array. However, mainstream LFASR methods based on convolutional neural networks (CNNs) usually suffer from long-range modeling capability due to the limited receptive field of standard convolutions. This leads to insufficient utilization of non-local spatial and angular information under large disparity situations. More recently, some Transformer-based methods have emerged to tackle this issue and achieve promising performance, however, the self-attention mechanism is suboptimal to recover detailed textures and computationally expensively. To tackle these issues, we propose a novel architecture Large Kernel Convolution Attention (LKCA) which combines large kernel convolutions with the transformer architecture to enhance the ability to capture long-range dependencies of LF images. We design a Muti-scale Hybrid Attention Block (MHAB) based on LKCA to fully utilize non-local spatial-angular information to exploit spatial, angular, and epipolar plane image features. Extensive experiments are carried out to show that our method can achieve more accurate reconstruction effects and generate more evident texture compared with the state-of-the-art methods.
Peiqi Xia, Yao Lu 0001, Shunzhou Wang
ICME2
2024 Focal Aggregation Transformer for Light Field Image Super-Resolution
Shunzhou Wang, Yao Lu 0001
PRCV (8)2
2024 How to set safety boundary in virtual reality: A dynamic approach based on user motion prediction
abstract
Abstract Virtual reality (VR) interaction safety is a prerequisite for all user activities in the virtual environment. While seeking a deep sense of immersion with little concern about surrounding obstacles, users may have limited ability to perceive the real‐world space, resulting in possible collisions with real‐world objects. Nowadays, recent works and rendering techniques such as the Chaperone can provide safety boundaries to users but confines them in a small static space and lack of immediacy. To solve this problem, we propose a dynamic approach based on user motion prediction named SCARF, which uses Spearman's correlation analysis, rule learning, and few‐shot learning to achieve prediction of user movements in specific VR tasks. Specifically, we study the relationship between user characteristics, human motion, and categories of VR tasks and provides an approach that uses biomechanical analysis to define the interaction space in VR dynamically.We report on a user study with 58 volunteers and establish a three dimensional kinematic dataset from a VR game. The experiments validate that our few‐shot learning model is effective and can improve the performance of motion prediction. Finally, we implement SCARF in VR environment for dynamic safety boundary adjustment.
Xiaoping Che, Enyao Chang, Chenxin Qu, Yao Lu 0001, Zhenlin Wei
Comput. Animat. Virtual Worlds5
2022 Detail-Preserving Transformer for Light Field Image Super-resolution
abstract
Recently, numerous algorithms have been developed to tackle the problem of light field super-resolution (LFSR), i.e., super-resolving low-resolution light fields to gain high-resolution views. Despite delivering encouraging results, these approaches are all convolution-based, and are naturally weak in global relation modeling of sub-aperture images necessarily to characterize the inherent structure of light fields. In this paper, we put forth a novel formulation built upon Transformers, by treating LFSR as a sequence-to-sequence reconstruction task. In particular, our model regards sub-aperture images of each vertical or horizontal angular view as a sequence, and establishes long-range geometric dependencies within each sequence via a spatial-angular locally-enhanced self-attention layer, which maintains the locality of each sub-aperture image as well. Additionally, to better recover image details, we propose a detail-preserving Transformer (termed as DPT), by leveraging gradient maps of light field to guide the sequence learning. DPT consists of two branches, with each associated with a Transformer for learning from an original or gradient image sequence. The two branches are finally fused to obtain comprehensive feature representations for reconstruction. Evaluations are conducted on a number of light field datasets, including real-world scenes and synthetic data. The proposed method achieves superior performance comparing with other state-of-the-art schemes. Our code is publicly available at: https://github.com/BITszwang/DPT.
Shunzhou Wang, Tianfei Zhou, Yao Lu 0001, Huijun Di
AAAI3
2022 Local-Global Feature Aggregation for Light Field Image Super-Resolution
abstract
Deep convolutional neural networks (CNNs) have been widely explored in light field (LF) image super-resolution (SR) to achieve remarkable progress. However, most of the existing CNNs-based methods ignore the similarity of local neighbor views in the 4D LF data. Besides, due to the limitations of CNNs, these methods can’t fully model the global spatial properties of the whole LF images. In this paper, we propose a network with Local-Global Feature Aggregation (LF-LGFA) to handle these problems for LF image SR. Specifically, the Local Aggregation Module is designed to incorporate the local angular information by utilizing the similarity of the local neighbor views’ features in LF images. Moreover, the Global Aggregation Module is designed to capture long-range spatial information via row-wise and column-wise self-attention. Extensive experimental results on five public LF datasets demonstrate that our method achieves comparable results against state-of-the-art techniques.
Yao Lu 0001, Shunzhou Wang, Zijian Wang 0007
ICASSP2
2022 Multi-Granularity Aggregation Transformer for Light Field Image Super-Resolution
abstract
Many light field image super-resolution methods aim to improve the quality of light-field image super-resolution reconstruction by exploiting the complementary information between sub-aperture images. Although these methods have achieved good results, the mutual information learning strategies between sub-aperture images are mostly handcrafted, limiting the super-resolution method to deal with the reconstruction of different scenes. We design a Transformer-based network named Multi-granularity Aggregation Transformer (MAT) to dynamically learn the complementary information between sub-aperture images in this paper. MAT is mainly implemented with the proposed multi-granularity aggregation blocks, which process sub-aperture images with three different granularity aggregation approaches and generate comprehensive spatial-angular representations for light field image super-resolution reconstruction. Extensive experiments are carried out on the mainstream light field image super-resolution datasets. MAT achieves new state-of-the-art results compared with other light field image super-resolution methods.
Zijian Wang 0007, Yao Lu 0001
ICIP2
2022 Multimodal Unsupervised Image-to-Image Translation Without Independent Style Encoder
Yanbei Sun, Yao Lu 0001, Haowei Lu 0002, Qingjie Zhao, Shunzhou Wang
MMM (1)2
2022 Contextual Transformation Network for Lightweight Remote-Sensing Image Super-Resolution
abstract
Current super-resolution networks typically reduce network parameters and multiadds operations by designing lightweight structures, but lightening the convolution layer is often ignored. In this work, we observe that$3 \times 3$convolutions occupy a high percentage of network parameters in most lightweight super-resolution networks. This motivates us to consider lightening super-resolution networks by replacing$3 \times 3$convolutions with lightweight convolutions, while maintaining the performance. To achieve this, we propose a lightweight convolution layer named contextual transformation layer (CTL). It can yield efficient contextual features through a context feature extraction module and enrich extracted contextual features through a context feature transformation module. Based on CTLs, we build a lightweight super-resolution network called contextual transformation network (CTN) for remote-sensing image super-resolution. Specifically, we use two CTLs to construct a contextual transformation block (CTB) for hierarchical feature learning. Interleaved with a CTB, a context enhancement module (CEM) is employed to enhance the extracted feature representations. All extracted features are processed by a contextual feature aggregation module for final remote-sensing image super-resolution. Extensive experiments are performed on a remote-sensing image super-resolution benchmark named UC Merced. Our method achieves superior results to the other state-of-the-art methods. To demonstrate the generalization ability of our CTL, we extend our CTN to two relevant tasks: natural image super-resolution and natural image denoising. Experimental results on natural image super-resolution benchmarks (i.e., Set5, Set14, B100, Urban100, and Manga109) and natural image denoising benchmarks (i.e., SIDD and DND) further prove the superiority of our method. Our code is publicly available athttps://github.com/BITszwang/CTNet.
Shunzhou Wang, Tianfei Zhou, Yao Lu 0001, Huijun Di
IEEE Trans. Geosci. Remote. Sens.3
2021 Edge Guided Attention Based Densely Connected Network for Single Image Super-Resolution
Zijian Wang 0007, Yao Lu 0001, Qingxuan Shi
ICONIP (3)2
2021 Learning Multi-level Interaction Relations and Feature Representations for Group Activity Recognition
Yao Lu 0001, Shunzhou Wang
MMM (1)2
2021 Scale-Aware Distillation Network for Lightweight Image Super-Resolution
Haowei Lu 0002, Yao Lu 0001, Gongping Li, Yanbei Sun, Shunzhou Wang, Yugang Li
PRCV (3)2
2021 LF-MAGNet: Learning Mutual Attention Guidance of Sub-Aperture Images for Light Field Image Super-Resolution
Zijian Wang 0007, Yao Lu 0001, Haowei Lu 0002, Shunzhou Wang, Binglu Wang
PRCV (3)2
2021 Single image super-resolution with attention-based densely connected module
Zijian Wang 0007, Yao Lu 0001, Shunzhou Wang, Xuebo Wang, Xiaozhen Chen
Neurocomputing2
2020 RSANet: Deep Recurrent Scale-Aware Network for Crowd Counting
abstract
Most recent works have made significant progress in crowd counting by fusing multi-scale features directly with weighted sum or concatenation to handle large scale variation problems. Meanwhile, there is very little attention paid on the prediction of high-resolution density maps and predicted low-resolution density maps lead to inaccurate counting results. In this paper, we present a novel recurrent scale-aware network(RSANet) to generate a high-resolution density map with scale-aware feature fusion approach. Within this network, we introduce a coarse-to-fine scheme restoring the high-resolution feature map from a low-resolution feature map progressively with stacked dilated convolution blocks. Then, we incorporate recurrent modules to capture dynamic scale-aware information and to benefit the restoration of high-resolution feature maps through multi-scale feature fusion to generate a high-resolution density map. We also use a multi-resolution supervision strategy for training to improve the performance of our network. Extensive experiments on three challenging crowd counting datasets demonstrate the effectiveness of the proposed method.
Yujun Xie 0002, Yao Lu 0001, Shunzhou Wang
ICIP2
2020 Blind Super-Resolution with Kernel-Aware Feature Refinement
Yao Lu 0001, Gongping Li, Shunzhou Wang, Xuebo Wang, Zijian Wang 0007
PRCV (1)2
2020 RBPNET: An asymptotic Residual Back-Projection Network for super-resolution of very low-resolution face image
Xiaozhen Chen, Xuebo Wang, Yao Lu 0001, Zijian Wang 0007, Zhuowei Huang
Neurocomputing3
2020 SCLNet: Spatial context learning network for congested crowd counting
Shunzhou Wang, Yao Lu 0001, Tianfei Zhou, Huijun Di, Lin Zhang 0033
Neurocomputing2
2020 GAIM: Graph Attention Interaction Model for Collective Activity Recognition
abstract
Unbalanced interaction relationships at personal and group levels play a pivotal role in collective activity recognition, which has not been adaptively and jointly explored by previous approaches. In this paper, we propose a graph attention interaction model (GAIM) embedded with the graph attention block (GAB) to explicitly and adaptively infer unbalanced interaction relations at personal and group levels in a unified architecture, and further to learn the spatial and temporal evolutions of the collective activity from these interactions to predict the activity labels. We first design the spatiotemporal graphs tailored to the collective activity where the concurrent person and group nodes, respectively, represent individuals' actions and the collective activity. The graphs provide both spatial structures and semantic appearance features for the collective activity. Then, GAB performs convolution-like filters on the graphs to infer unequal and two-level interaction relations in the collective activity by implementing graph convolutional networks with a shared attention mechanism. At the personal level, the GAB learns different levels of interactions for each person node from its neighbor person nodes under the guidance from the group node. At the group level, the GAB assesses various degrees of interactions to the group node contributed by person nodes. Equipped with the GRUs network, the GAIM learns the spatial and temporal evolutions of individuals' actions as well as the collective activity from the captured interactions, and finally predicts the label of the collective activity. Experiments on four publicly available datasets and ablation studies are conducted to evaluate the performance of our GAIM, and the improved performance demonstrates the effectiveness of our model.
Yao Lu 0001, Ruizhe Yu, Huijun Di, Lin Zhang 0033, Shunzhou Wang
IEEE Trans. Multim.2
2019 Cross-Domain Scene Text Detection via Pixel and Image-Level Adaptation
Danlu Chen, Yao Lu 0001, Ruizhe Yu, Shunzhou Wang, Lin Zhang 0033, Tingxi Liu
ICONIP (5)3
2019 Accurate Single Image Super-Resolution Using Deep Aggregation Network
Xiaozhen Chen, Yao Lu 0001, Xuebo Wang, Zijian Wang 0007
ICONIP (2)2
2019 RBPNET: An Asymptotic Residual Back-Projection Network for Super Resolution of Very Low Resolution Face Image
Xuebo Wang, Yao Lu 0001, Xiaozhen Chen, Zijian Wang 0007
ICONIP (2)2
2019 ADSRNet: Attention-Based Densely Connected Network for Image Super-Resolution
Yao Lu 0001, Xuebo Wang, Xiaozhen Chen, Zijian Wang 0007
PRCV (3)2
2019 Asymmetric Pyramid Based Super Resolution from Very Low Resolution Face Image
Xuebo Wang, Yao Lu 0001, Xiaozhen Chen, Zijian Wang 0007
PRCV (2)2
2019 SAF: Semantic Attention Fusion Mechanism for Pedestrian Detection
Ruizhe Yu, Shunzhou Wang, Yao Lu 0001, Huijun Di, Lin Zhang 0033
PRICAI (2)3
2019 Refined video segmentation through global appearance regression
Lin Zhang 0033, Yao Lu 0001, Tianfei Zhou
Neurocomputing2
2019 Spatio-temporal attention mechanisms based model for collective activity recognition
Huijun Di, Yao Lu 0001, Lin Zhang 0033, Shunzhou Wang
Signal Process. Image Commun.3
2019 Objectness-based smoothing stochastic sampling and coherence approximate nearest neighbor for visual tracking
Jimmy T. Mbelwa, Qingjie Zhao, Yao Lu 0001, Fasheng Wang, Mercy Mbise
Vis. Comput.3
2018 Soccer Video Super-Resolution via Sub-Pixel Convolutional Neural Network
abstract
In this paper, we consider the problem of the soccer video super-resolution (SR). We propose an end-to-end framework based on the sub-pixel convolution neural network and pre-train our network on the image datasets to improve the resolution of soccer video and the speed of the SR algorithm. Different from the general SR methods, our method does not need to interpolate the low-resolution (LR) images first. It means that we directly extract the features from the LR images so the computational cost is low. We also compare our proposed algorithm with the current SR methods, and the experimental results show that our network has better performance. In order to evaluate the effectiveness of SR, we apply the proposed method to the object detection on the soccer videos. The experiments show that integrating the SR technique into the detecting field can improve the accuracy of detection.
Yao Lu 0001
IJCNN2
2018 RC-CNN: Reverse Connected Convolutional Neural Network for Accurate Player Detection
Yao Lu 0001, Hanfeng Zheng
PRICAI2
2018 A two-level attention-based interaction model for multi-person activity recognition
Huijun Di, Yao Lu 0001, Lin Zhang 0033, Shunzhou Wang
Neurocomputing3
2017 Learning to generate video object segment proposals
abstract
This paper proposes a fully automatic pipeline to generate accurate object segment proposals in realistic videos. Our approach first detects generic object proposals for all video frames and then learns to rank them using a Convolutional Neural Networks (CNN) descriptor built on appearance and motion cues. The ambiguity of the proposal set can be reduced while the quality can be retained as highly as possible Next, high-scoring proposals are greedily tracked over the entire sequence into distinct tracklets. Observing that the proposal tracklet set at this stage is noisy and redundant, we perform a tracklet selection scheme to suppress the highly overlapped tracklets, and detect occlusions based on appearance and location information. Finally, we exploit holistic appearance cues for refinement of video segment proposals to obtain pixel-accurate segmentation. Our method is evaluated on two video segmentation datasets i.e. SegTrack v1 and FBMS-59 and achieves competitive results in comparison with other state-of-the-art methods.
Jianwu Li, Tianfei Zhou, Yao Lu 0001
ICME3
2017 RGB-D Tracking Based on Kernelized Correlation Filter with Deep Features
Yao Lu 0001, Lin Zhang 0033, Jian Zhang 0002
ICONIP (3)2
2017 Soccer Video Event Detection Using 3D Convolutional Networks and Shot Boundary Detection via Deep Feature Distance
Tingxi Liu, Yao Lu 0001, Xiaoyu Lei, Wei Huang 0030, Zijian Wang 0007
ICONIP (2)2
2017 Robust Visual Tracking Based on Multi-channel Compressive Features
Yao Lu 0001
MMM (1)2
2017 Video pose estimation with global motion cues
Qingxuan Shi, Huijun Di, Yao Lu 0001, Feng Lv, Xuedong Tian
Neurocomputing3
2017 Region-based Mixture Models for human action recognition in low-resolution videos
Huijun Di, Jian Zhang 0002, Yao Lu 0001, Feng Lv
Neurocomputing4
2017 Locality-Constrained Collaborative Model for Robust Visual Tracking
abstract
This paper presents a novel discriminative, generative, and collaborative appearance model for robust object tracking. In contrast to existing methods, we use different appearance manifolds to represent the target in the discriminative and generative appearance models and propose a novel collaborative scheme to combine these two components. In particular: 1) for the discriminative component, we develop a graph regularized discriminant analysis (GRDA) algorithm that can find a projection to more effectively distinguish the target from the background; 2) for the generative component, we introduce a simple yet effective coding method for object representation. The method involves no optimization, and thus better efficiency can be achieved; and 3) for the collaborative model, we apply GRDA again to find a subspace for discriminating the likelihood features (generated from the discriminative and generative appearance models) and use the nearest neighbor criterion to determine the final likelihood. Besides, all the components are online updated so that our tracker can deal with appearance changes effectively. The experimental results over 23 challenging image sequences demonstrate that the proposed algorithm achieves better performance compared with other state-of-the-art methods.
Tianfei Zhou, Yao Lu 0001, Huijun Di
IEEE Trans. Circuits Syst. Video Technol.2
2016 Face hallucination scheme based on singular value content metric for K-NN selection and an iterative refining in a modified feature space
abstract
Numbers of neighbor embedding (NE) methods have been proposed, which use the image content metric based on the distance values such as Euclidean distance between the input image patch and the image patches in the training set to find the nearest neighbors. In contrast to these approaches we propose to use image content metric that uses the most effective singular values of the patch of interest. Singular value content metric give the effective and quantitative measure of the true image content and can search the most similar patches from the training set which possess the local similarity with the input patch. First we find the K most similar low resolution (LR) and corresponding high resolution (HR) patches by using the proposed image content metric. Secondly we project the K neighbor onto a modified feature space by employing easy partial least square estimation (EZ-PLS). In modified feature space we propose to explore the data structure of both LR and HR manifold and iteratively update Z nearest neighbors and reconstruction weights based on the results from previous iteration. The Rigorous experimentation with application to face hallucination demonstrate the effectiveness of the proposed method.
Javaria Ikram, Yao Lu 0001, Jianwu Li, Nie Hui
ICIP2
2016 Instance significance guided multiple instance boosting for robust visual tracking
abstract
Multiple Instance Learning (MIL) recently provides an appealing way to alleviate the drifting problem in visual tracking. Following the tracking-by-detection framework, an online MILBoost approach is developed that sequentially chooses weak classifiers by maximizing the bag likelihood. In this paper, we extend this idea towards incorporating the instance significance estimation into the online MILBoost framework. First, instead of treating all instances equally, with each instance we associate a significance-coefficient that represents its contribution to the bag likelihood. The coefficients are estimated by a Bayesian formula that jointly considers the predictions from multiple randomized MILBoost classifiers. Next, we incorporate the estimates within a new boosting procedure for more effectively selection of weak classifiers. Experiments with challenging public datasets show that the proposed method outperforms both existing MIL based and boosting based trackers.
Jinwu Liu, Yao Lu 0001, Tianfei Zhou
ICIP2
2016 Video pose estimation via medium granularity graphical model with spatial-temporal symmetric constraint part model
abstract
We address the problem of full body human pose estimation in video. Most previous work consider body part, pose or trajectory of body part as basic unit to compose the pose sequence. In contrast, we consider tracklet of body part as the basic unit. Based on this medium granularity representation we develop a spatio-temporal graphical model to select an optimal tracklet for each part in each video segment. In our model, tracklet nodes of symmetric parts are coupled to one node to overcome the double counting problem. Through iterative spatial and temporal parsing, optimal solution is achieved in polynomial time. We apply our model on three publicly available datasets and show remarkable quantitative and qualitative improvements over the state-of-the-art approaches.
Qingxuan Shi, Huijun Di, Yao Lu 0001, Ming Qin, Xuedong Tian
ICIP3
2016 Recognizing human actions from low-resolution videos by region-based mixture models
abstract
Recognizing human action from low-resolution (LR) videos is essential for many applications including large-scale video surveillance, sports video analysis and intelligent aerial vehicles. Currently, state-of-the-art performance in action recognition is achieved by the use of dense trajectories which are extracted by optical flow algorithms. However, the optical flow algorithms are far from perfect in LR videos. In addition, the spatial and temporal layout of features is a powerful cue for action discrimination. While, most existing methods encode the layout by previously segmenting body parts which is not feasible in LR videos. Addressing the problems, we adopt the Layered Elastic Motion Tracking (LEMT) method to extract a set of long-term motion trajectories and a long-term common shape from each video sequence, where the extracted trajectories are much denser than those of sparse interest points(SIPs); then we present a hybrid feature representation to integrate both of the shape and motion features; and finally we propose a Region-based Mixture Model (RMM) to be utilized for action classification. The RMM models the spatial layout of features without any needs of body parts segmentation. Experiments are conducted on two publicly available LR human action datasets. Among which, the UT-Tower dataset is very challenging because the average height of human figures is only about 20 pixels. The proposed approach attains near-perfect accuracy on both of the datasets.
Huijun Di, Jian Zhang 0002, Yao Lu 0001, Feng Lv
ICME4
2016 Video object segmentation aggregation
abstract
We present an approach for unsupervised object segmentation in unconstrained videos. Driven by the latest progress in this field, we argue that segmentation performance can be largely improved by aggregating the results generated by state-of-the-art algorithms. Initially, objects in individual frames are estimated through a per-frame aggregation procedure using majority voting. While this can predict relatively accurate object location, the initial estimation fails to cover the parts that are wrongly labeled by more than half of the algorithms. To address this, we build a holistic appearance model using non-local appearance cues by linear regression. Then, we integrate the appearance priors and spatio-temporal information into an energy minimization framework to refine the initial estimation. We evaluate our method on challenging benchmark videos and demonstrate that it outperforms state-of-the-art algorithms.
Tianfei Zhou, Yao Lu 0001, Huijun Di, Jian Zhang 0002
ICME2
2016 Removing Ring Artifacts in CBCT Images Using Smoothing Based on Relative Total Variation
Qirun Huo, Jianwu Li, Yao Lu 0001, Ziye Yan
ICONIP (1)3
2016 Face Hallucination Using Correlative Residue Compensation in a Modified Feature Space
Javaria Ikram, Yao Lu 0001, Jianwu Li, Nie Hui
ICONIP (2)2
2016 Face Hallucination via Convolution Neural Network
abstract
Deep learning methods have been successfully used in many areas of computer vision, including super resolution. However, all of the previous deep learning methods have been proposed for generic image super resolution. In this paper, we proposed to use convolutional neural network for face hallucination (FH) by combining the domain specific prior knowledge of face images and properties of deep learning. In the proposed method, an end to end mapping is learned as a deep convolutional network between the low resolution (LR) images and their corresponding high resolution (HR) images to upscale the input face image directly. In order to achieve larger magnification factor, we consider to cascade several convolution neural networks each of which is with a fixed up-scaling factor and upscales the LR image step by step. Experimental result shows that our proposed method can achieve better performance comparing to the traditional face hallucination methods.
Nie Hui, Yao Lu 0001, Javaria Ikram
ICTAI2
2016 Automatic Soccer Video Event Detection Based on a Deep Neural Network Combined CNN and RNN
abstract
Soccer video semantic analysis has attracted a lot of researchers in the last few years. Many methods of machine learning have been applied to this task and have achieved some positive results, but the neural network method has not yet been used to this task from now. Taking into account the advantages of Convolution Neural Network(CNN) in fully exploiting features and the ability of Recurrent Neural Network(RNN) in dealing with the temporal relation, we construct a deep neural network to detect soccer video event in this paper. First we determine the soccer video event boundary which we used Play-Break(PB) segment by the traditional method. Then we extract the semantic features of key frames from PB segment by pre-trained CNN, and at last use RNN to map the semantic features of PB to soccer event types, including goal, goal attempt, card and corner. Because there is no suitable and effective dataset, we classify soccer frame images into nine categories according to their different semantic views and then construct a dataset called Soccer Semantic Image Dataset(SSID) for training CNN. The sufficient experiments evaluated on 30 soccer match videos demonstrate the effectiveness of our method than state-of-art methods.
Haohao Jiang, Yao Lu 0001
ICTAI2
2016 A Background Basis Selection-Based Foreground Detection Method
abstract
Foreground detection plays a fundamental role in video analysis. Frames with only background information are usually beneficial for many foreground detection algorithms, especially for regression-based methods where the background is recovered from a background basis matrix. However, many regression-based methods ignore the basis selection process or select bases by simple sampling, which may limit their performance. In this paper, a regression-based foreground detection method with a novel background basis selection process is proposed. The proposed basis selection method, which includes basis matrix construction and basis matrix update processes, aims to build an effective background basis matrix which helps to boost the performance of our foreground detection method. In our algorithm, the basis matrix construction process first builds the basis matrix locally with a multiple clustering evaluation process. With the locally constructed basis matrix, a modified linear regression-based foreground detection method is proposed for separating foreground and background globally. To further increase the representativeness and the adaptiveness of the background basis matrix, a basis matrix update algorithm is designed to incrementally replace the ineffective bases with new selected ones. Extensive experiments on challenging sequences demonstrate the effectiveness and the advantages of our method.
Ming Qin, Yao Lu 0001, Huijun Di, Wei Huang 0030
IEEE Trans. Multim.2
2015 Contour Flow: Middle-Level Motion Estimation by Combining Motion Segmentation and Contour Alignment
abstract
Our goal is to estimate contour flow (the contour pairs with consistent point correspondence) from inconsistent contours extracted independently in two video frames. We formulate the contour flow estimation locally as a motion segmentation problem where motion patterns grouped from optical flow field are exploited for local correspondence measurement. To solve local ambiguities, contour flow estimation is further formulated globally as a contour alignment problem. We propose a novel two-staged strategy to obtain global consistent point correspondence under various contour transitions such as splitting, merging and branching. The goal of the first stage is to obtain possible accurate contour-to-contour alignments, and the second stage aims to make a consistent fusion of many partial alignments. Such a strategy can properly balance the accuracy and the consistency, which enables a middle-level motion representation to be constructed by just concatenating frame-by-frame contour flow estimation. Experiments prove the effectiveness of our method.
Huijun Di, Qingxuan Shi, Feng Lv, Ming Qin, Yao Lu 0001
ICCV5
2015 Human pose estimation with global motion cues
abstract
We present a novel method to estimate full-body human pose in video sequence by incorporating global motion cues. It has been demonstrated that temporal constraints can largely enhance the pose estimation. Most current approaches typically employ local motion to propagate pose detections to supplement the pose candidates. However, the local motion estimation is often inaccurate under fast movements of body parts and unhelpful when no strong detections achieved in adjacent frames. In this paper, we propose to propagate the detection in each frame using the global motion estimation. Benefiting from the strong detections, our algorithm first produces reasonable trajectory hypotheses for each body part. Then, we cast pose estimation as an optimization problem defined on these trajectories with spatial links between body parts. In the optimization process, we select body part trajectory rather than body part candidate to infer the human pose. Experimental results demonstrate significant performance improvement in comparison with the state-of-the-art methods.
Qingxuan Shi, Huijun Di, Yao Lu 0001, Feng Lv
ICIP3
2015 Graph regularized discriminant analysis and its application to face recognition
abstract
Linear Discriminant Analysis (LDA) is a powerful technology for supervised dimensionality reduction, however, it only captures the extrinsic (or global) structure in the data and fails to discover the intrinsic structure of the data manifold. In this paper, we develop a new linear supervised dimensionality reduction method, called Graph Regularized Discriminant Analysis(GRDA), which respects both extrinsic and intrinsic structure in the data. In particular, a regularization term, incorporating the manifold structure, is introduced into the objective function of LDA. The formulation allows us to achieve a more discriminative subspace by simultaneously considering the graph preserving and the global LDA criteria. We then apply the proposed GRDA algorithm to face recognition by exploiting the local dissimilarity of face images in different classes. Experimental results clearly show that the proposed GRDA method outperforms many state-of-the-art face recognition algorithms.
Tianfei Zhou, Yao Lu 0001
ICIP2
2015 Background basis selection from multiple clustering on local neighborhood structure
abstract
Foreground detection with dynamic background is a challenging task in video surveillance analysis. When clean background bases are constructed, regression based foreground detection usually becomes more effective. In this paper, a novel basis selection method based on local neighborhood structure is proposed. The present method first constructs local neighborhood relationships among the basis candidates in a reconstruction manner. Then a multiple clustering strategy is designed to evaluate these basis candidates on local neighborhood structure. According to the evaluation score given by multiple clustering process, clean background bases (including dynamic background) are separated from candidates corrupted by foreground. By adding the proposed basis selection process to a modified linear regression framework, the foreground detection can be implemented in a more effective way. Experimental results on multiple videos show that the modified framework with basis selection is competitive with the state of the art.
Ming Qin, Yao Lu 0001, Huijun Di, Wei Huang 0030
ICME2
2015 Abrupt motion tracking via nearest neighbor field driven stochastic sampling
Tianfei Zhou, Yao Lu 0001, Feng Lv, Huijun Di, Qingjie Zhao, Jian Zhang 0002
Neurocomputing2
2014 Robust tracking via weighted spatio-temporal context learning
abstract
Designing a robust visual tracker is a challenging problem due to many disturbed factors such as illumination changes, appearance changes, rotation, partial or full occlusions, etc. Among numerous existed trackers, correlation filter based tracker is a fast and robust method with resistance to the above-mentioned factors. Motivated by that, spatio-temporal context (STC) learning algorithm is proposed, which considers the information of the context around the target and achieved better performance. However, STC treats the whole region of the context equally, which weakens the effectiveness of the context information. In this paper, we propose a novel weighted spatio-temporal context (WSTC) learning algorithm. Our algorithm considers the surrounding context discriminatively and integrates a weighted map by evaluating the importance of different regions. Extensive experimental results on various benchmark databases show that our algorithm outperforms the STC algorithm and the other state-of-the-art algorithms.
Yao Lu 0001, Jinwu Liu
ICIP2
2014 Nearest neighbor field driven stochastic sampling for abrupt motion tracking
abstract
Stochastic sampling based trackers have shown good performance for abrupt motion tracking so that they have gained popularity in recent years. However, the existing methods tend to explore the whole state space uniformly with an inefficiency preliminary sampling phase. In this paper, we propose a nearest neighbor field(NNF) driven stochastic sampling framework for abrupt motion tracking in which NNF provides us promising regions the target may exist, and thus can help to explore the state space more effectively. Our approach firstly computes NNF to determine the promising regions; subsequently, we adopt Smoothing Stochastic Approximate Monte Carlo(SSAMC) sampling scheme to accurately localize the target. SSAMC is robust to handle the noises in NNF by propagating a sample's information to its neighboring regions. Finally, we refine the result with sparse representation based template matching technique. The experimental results on challenging sequences show that our tracker outperforms other related methods by better accuracy and higher robustness.
Tianfei Zhou, Yao Lu 0001, Huijun Di
ICME2
2013 Effective two-step method for face hallucination based on sparse compensation on over-complete patches
abstract
Sparse representation has been successfully applied to image d using low‐ and high‐resolution training face images based on sparse representation. In this study, the sparse residual compensation is adopted to face hallucination. Firstly, a global face image is constructed by optimal coefficients of the interpolated training images. Secondly, the high‐resolution residual image (local face image) is found by using an over‐complete patch dictionary and the sparse representation. Finally, a hallucinated face image is obtained by combining these two steps. In addition, the more details of the face image in high frequency parts are recovered using a residual compensation strategy. In the authors’ experimental work, it is observed that balance sparsity parameter ( λ ) has affected the residual compensation. Further, the proposed algorithm can acquire a high‐resolution image even though the number of training image pairs is comparatively smaller. The experiments show that the authors’ method is more effective than the other existing two‐step face hallucination methods.
M. Naleer Haju Mohamed, Yao Lu 0001, Feng Lv
IET Image Process.2
2012 Heart Sounds Classification with a Fuzzy Neural Network Method with Structure Learning
Lijuan Jia, Linmi Tao, Yao Lu 0001
ISNN (2)4
2011 Super Resolution of Text Image by Pruning Outlier
Ziye Yan, Yao Lu 0001, Jianwu Li
ICONIP (3)2
2010 Spatial-Temporal Motion Compensation Based Video Super Resolution
Yaozu An, Yao Lu 0001, Ziye Yan
ACCV (2)2
2010 Reducing the spiral ct slice thickness using super resolution
abstract
An approach for improving the z-axis resolution of spiral CT using super resolution (SR) technology in the post processing step is proposed in this work. As the spiral CT has the ability to produce over-lapped slice images without over-lapped scanning, the sub-slice thickness shifts in z-axis require neither change in hardware nor any additional radiation dose. The thinner slice is obtained by combining those over-lapped slices and applying the SR algorithm. It is secure and convenient to be used in clinical cases. We use an algorithm which introduces the Papoulis-Gerchberg extrapolation within the Iterative back-projection method to process the reconstructed image serial. Experiments on phantom and patient demonstrate effectiveness of this approach. And the slice thickness limit of spiral CT system is broken by this approach.
Ziye Yan, Yao Lu 0001, Hongxia Yan
ICIP2
2010 Refining Kernel Matching Pursuit
Jianwu Li, Yao Lu 0001
ISNN (2)2
2009 Spatially Varying Regularization of Image Sequences Super-Resolution
Yaozu An, Yao Lu 0001, Zhengang Zhai
ACCV (3)2
2008 Adapting radial basis function neural networks for one-class classification
abstract
One-class classification (OCC) is to describe one class of objects, called target objects, and discriminate them from all other possible patterns. In this paper, we propose to adapt radial basis function neural networks (RBFNNs) for OCC. First, target objects are mapped into a feature space by using neurons in the hidden layer of the RBFNNs. Then, we perform support vector domain description (SVDD) with linear kernel functions in the feature space to realize OCC. In addition, we also model, in the feature space, the closed sphere centered on the mean of target objects for OCC. Compared to the SVDD with nonlinear kernel functions, our methods can use flexible nonlinear mappings, which do not necessarily satisfy Mercerpsilas conditions. Moreover, we can also control the complexity of solutions easily by setting the number of neurons in the hidden layer of RBFNNs. Experimental results show that the classification accuracies of our methods can be close to, and even can reach those of the SVDD for most of results, but with typically much sparser models.
Jianwu Li, Zhanyong Mao, Yao Lu 0001
IJCNN3
2002 Spatial resolution improvement of spatial shift multi-observation images by neural network
abstract
In this paper, improvements on spatial resolution of the spatial shift multi-observation images are discussed. And a block of pixels-based artificial neural network is proposed for this purpose. This system makes full use of the spatial information to implement the superposition of multiple images. Its convergence and learning problems are also discussed. The effectiveness and the high performance of the proposed neural network are demonstrated by computer experiments, error calculation and comparison with other methods.
Yao Lu 0001, Minoru Inamura
IGARSS1