EDBT 2026 Demo / reviewers in the wild / expert
Yingli Tian
dblp:54/8250 · also Ying-li Tian
· DBLP profile ↗
122ranked-venue papers
19as first author
21since 2021 · last 2026
0000-0003-4458-360XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 73 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 56 · 12 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 3 first-authorHuman-computer interaction and ubiquitous computing · 7 · 3 first-authorDatabases, data management, data science and information retrieval · 5Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SepPrune: Structured Pruning for Efficient Deep Speech SeparationabstractAlthough deep learning has substantially advanced speech separation in recent years, most existing studies continue to prioritize separation quality while overlooking computational efficiency, an essential factor for low-latency speech processing in real-time applications. In this paper, we propose SepPrune, the first structured pruning framework specifically designed to compress deep speech separation models and reduce their computational cost. SepPrune begins by analyzing the computational structure of a given model to identify layers with the highest computational burden. It then introduces a differentiable masking strategy to enable gradient-driven channel selection. Based on the learned masks, SepPrune prunes redundant channels and fine-tunes the remaining parameters to recover performance. Extensive experiments demonstrate that this learnable pruning paradigm yields substantial advantages for channel pruning in speech separation models, outperforming existing methods. Notably, a model pruned with SepPrune can recover 85% of the performance of a pre-trained model (trained over hundreds of epochs) with only one epoch of fine-tuning, and achieves convergence 36x faster than training from scratch. Zhifei Yang 0004, Zeyu Dong, Zhengtao Yao, Haoyan Xu, Yingli Tian, Yao Lu 0041 |
AAAI | 8 |
| 2026 | GaitProtector: Impersonation-Driven Gait De-Identification via Training-Free Diffusion Latent Optimization
Huiran Duan, Qian Zhou 0001, Zhongliang Guo 0001, Junhao Dong 0001, Guoying Zhao 0001, Yingli Tian |
FG | 7 |
| 2026 | Towards robust medical image segmentation: Spectro-spatial domain generalization with MRAM and DMIR
Junhao Dong 0001, Hansheng Zeng, Fuyan Zhang, Zeyu Dong, Chuanguang Yang, Yingli Tian |
Comput. Vis. Image Underst. | 7 |
| 2026 | AMFOR: Adaptive Multi-Granularity Fusion and Occlusion Reconstruction for Person Re-IdentificationabstractOccluded person re-identification (ReID) poses substantial challenges in computer vision, primarily due to incomplete information and occlusion interference. Although Transformer architectures have become dominant in ReID due to their strong feature modeling capabilities, their lack of an adaptive weight allocation mechanism for multi-granularity feature processing limits their ability to extract generalizable and robust features. Recently, Masked Image Modeling (MIM) has demonstrated considerable promise in visual tasks, but its integration into ReID models remains underexplored. This paper presents AMFOR (Adaptive Multi-granularity feature Fusion and Occlusion Reconstruction), a novel framework combining MIM and Transformer architectures. AMFOR consists of three key components: AMFF-Encoder, HPR-Decoder, and Teacher-Student Decoder. The AMFF-Encoder enables adaptive fusion of multi-granularity features through learnable queries, allowing interaction between text-visual features and visual features extracted from multiple Transformer layers. The HPR-Decoder conceptualizes occluded regions in pedestrian images as reconstructable patches, guiding the encoder to extract more discriminative features through reconstruction. Additionally, the self-distillation teacher-student decoder is employed to refine pedestrian part features, further optimized by the proposed AMGDLoss. This paper represents the first successful implementation of the MIM mechanism in person ReID models. Empirical evaluations on five benchmark datasets, covering both occluded (Occluded-DukeMTMC, Occluded-REID, and P-DukeMTMC-reID) and complete (Market-1501 and DukeMTMC-reID) scenarios, demonstrate that AMFOR outperforms existing state-of-the-art methods in person ReID. Our code is available athttps://github.com/Guangdeng-Li/AMFOR Zhi Liu 0004, Guangdeng Li, Yingli Tian |
IEEE Trans. Multim. | 3 |
| 2025 | Enhancing Facial Privacy Protection via Weakening Diffusion PurificationabstractThe rapid growth of social media has led to the widespread sharing of individual portrait images, which pose serious privacy risks due to the capabilities of automatic face recognition (AFR) systems for mass surveillance. Hence, protecting facial privacy against unauthorized AFR systems is essential. Inspired by the generation capability of the emerging diffusion models, recent methods employ diffusion models to generate adversarial face images for privacy protection. However, they suffer from the diffusion purification effect, leading to a low protection success rate (PSR). In this paper, we first propose learning unconditional embeddings to increase the learning capacity for adversarial modifications and then use them to guide the modification of the adversarial latent code to weaken the diffusion purification effect. Moreover, we integrate an identity-preserving structure to maintain structural consistency between the original and generated images, allowing human observers to recognize the generated image as having the same identity as the original. Extensive experiments conducted on two public datasets, i.e., CelebA-HQ and LADN, demonstrate the superiority of our approach. The protected faces generated by our method outperform those produced by existing facial privacy protection approaches in terms of transferability and natural appearance. The code is available at https://github.com/parham1998/FacialPrivacy-Protection Ali Salar, Qing Liu 0003, Yingli Tian, Guoying Zhao 0001 |
CVPR | 3 |
| 2025 | Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal ForecastingabstractSpatiotemporal forecasting tasks, such as traffic flow, combustion dynamics, and weather forecasting, often require complex models that suffer from low training efficiency and high memory consumption. This paper proposes a lightweight framework, Spectral Decoupled Knowledge Distillation (termed SDKD), which transfers the multi-scale spatiotemporal representations from a complex teacher model to a more efficient lightweight student network. The teacher model follows an encoder-latent evolution-decoder architecture, where its latent evolution module decouples high-frequency details and low-frequency trends using convolution and Transformer (global low-frequency modeler). However, the multi-layer convolution and deconvolution structures result in slow training and high memory usage. To address these issues, we propose a frequency-aligned knowledge distillation strategy, which extracts multi-scale spectral features from the teacher's latent space, including both high and low frequency components, to guide the lightweight student model in capturing both local fine-grained variations and global evolution patterns. Experimental results show that SDKD significantly improves performance, achieving reductions of up to 81.3% in MSE and in MAE 52.3% on the Navier-Stokes equation dataset. The framework effectively captures both high-frequency variations and long-term trends while reducing computational complexity. Our codes are available at https://github.com/itsnotacie/SDKD Chuanguang Yang, Hansheng Zeng, Zeyu Dong, Zhulin An, Yongjun Xu 0001, Yingli Tian, Hao Wu 0094 |
ICCV | 7 |
| 2024 | POTLoc: Pseudo-label Oriented Transformer for point-supervised temporal Action Localization
Elahe Vahdani, Yingli Tian |
Comput. Vis. Image Underst. | 2 |
| 2024 | Sequential Point Clouds: A SurveyabstractPoint clouds have garnered increasing research attention and found numerous practical applications. However, many of these applications, such as autonomous driving and robotic manipulation, rely on sequential point clouds, essentially adding a temporal dimension to the data (i.e., four dimensions) because the information of the static point cloud data could provide is still limited. Recent research efforts have been directed towards enhancing the understanding and utilization of sequential point clouds. This paper offers a comprehensive review of deep learning methods applied to sequential point cloud research, encompassing dynamic flow estimation, object detection & tracking, point cloud segmentation, and point cloud forecasting. This paper further summarizes and compares the quantitative results of the reviewed methods over the public benchmark datasets. Ultimately, the paper concludes by addressing the challenges in current sequential point cloud research and pointing towards promising avenues for future research. Haiyan Wang 0019, Yingli Tian |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Learning transformer-based attention region with multiple scales for occluded person re-identification
Zhi Liu 0013, Xingyu Mu, Yunhua Lu, Yingli Tian |
Comput. Vis. Image Underst. | 5 |
| 2023 | Deep Learning-Based Action Detection in Untrimmed Videos: A SurveyabstractUnderstanding human behavior and activity facilitates advancement of numerous real-world applications, and is critical for video analysis. Despite the progress of action recognition algorithms in trimmed videos, the majority of real-world videos are lengthy and untrimmed with sparse segments of interest. The task of temporal activity detection in untrimmed videos aims to localize the temporal boundary of actions and classify the action categories. Temporal activity detection task has been investigated in full and limited supervision settings depending on the availability of action annotations. This article provides an extensive overview of deep learning-based algorithms to tackle temporal action detection in untrimmed videos with different supervision levels including fully-supervised, weakly-supervised, unsupervised, self-supervised, and semi-supervised. In addition, this article reviews advances in spatio-temporal action detection where actions are localized in both temporal and spatial dimensions. Action detection in online setting is also reviewed where the goal is to detect actions in each frame without considering any future context in a live video stream. Moreover, the commonly used action detection benchmark datasets and evaluation metrics are described, and the performance of the state-of-the-art methods are compared. Finally, real-world applications of temporal action detection in untrimmed videos and a set of future directions are discussed. Elahe Vahdani, Yingli Tian |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | PSMNet: Position-aware Stereo Merging Network for Room Layout EstimationabstractIn this paper, we propose a new deep learning-based method for estimating room layout given a pair of 360° panoramas. Our system, called Position-aware Stereo Merging Network or PSMNet, is an end-to-end joint layout-pose estimator. PSMNet consists of a Stereo Pano Pose (SP2) transformer and a novel Cross-Perspective Projection (CP2) layer. The stereo-view SP2 transformer is used to implicitly infer correspondences between views, and can handle noisy poses. The pose-aware CP2layer is designed to render features from the adjacent view to the anchor (reference) view, in order to perform view fusion and estimate the visible layout. Our experiments and analysis validate our method, which significantly outperforms the state-of-the-art layout estimators, especially for large and complex room spaces. Haiyan Wang 0019, Will Hutchcroft, Yuguang Li, Zhiqiang Wan, Ivaylo Boyadzhiev, Yingli Tian, Sing Bing Kang |
CVPR | 6 |
| 2022 | Disentangling Object Motion and Occlusion for Unsupervised Multi-frame Monocular Depth
Ziyue Feng, Longlong Jing, Haiyan Wang 0019, Yingli Tian, Bing Li 0008 |
ECCV (32) | 5 |
| 2022 | Unambiguous Text Localization, Retrieval, and Recognition for Cluttered ScenesabstractText instance as one category of self-described objects provides valuable information for understanding and describing cluttered scenes. The rich and precise high-level semantics embodied in the text could drastically benefit the understanding of the world around us. While most recent visual phrase grounding approaches focus on general objects, this paper explores extracting designated texts and predicting unambiguous scene text information, i.e., to accurately localize and recognize a specific targeted text instance in a cluttered image from natural language descriptions (referring expressions). To address this issue, first a novel recurrent dense text localization network (DTLN) is proposed to sequentially decode the intermediate convolutional representations of a cluttered scene image into a set of distinct text instance detections. Our approach avoids repeated text detections at multiple scales by recurrently memorizing previous detections, and effectively tackles crowded text instances in close proximity. Second, we propose a context reasoning text retrieval (CRTR) model, which jointly encodes text instances and their context information through a recurrent network, and ranks localized text bounding boxes by a scoring function of context compatibility. Third, a recurrent text recognition module is introduced to extend the applicability of aforementioned DTLN and CRTR models, via text verification or transcription. Quantitative evaluations on standard scene text extraction benchmarks and a newly collected scene text retrieval dataset demonstrate the effectiveness and advantages of our models for the joint scene text localization, retrieval, and recognition task. Xuejian Rong, Chucai Yi, Yingli Tian |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Multimodal Semi-Supervised Learning for 3D Objects
Longlong Jing, Yingli Tian, Bing Li 0008 |
BMVC | 4 |
| 2021 | Cross-Modal Center Loss for 3D Cross-Modal RetrievalabstractCross-modal retrieval aims to learn discriminative and modal-invariant features for data from different modalities. Unlike the existing methods which usually learn from the features extracted by offline networks, in this paper, we propose an approach to jointly train the components of cross-modal retrieval framework with metadata, and enable the network to find optimal features. The proposed end-to-end framework is updated with three loss functions: 1) a novel cross-modal center loss to eliminate cross-modal discrepancy, 2) cross-entropy loss to maximize inter-class variations, and 3) mean-square-error loss to reduce modality variations. In particular, our proposed cross-modal center loss minimizes the distances of features from objects belonging to the same class across all modalities. Extensive experiments have been conducted on the retrieval tasks across multi-modalities including 2D image, 3D point cloud and mesh data. The proposed framework significantly out-performs the state-of-the-art methods for both cross-modal and in-domain retrieval for 3D objects on the ModelNet10 and ModelNet40 datasets. Longlong Jing, Elahe Vahdani, Jiaxing Tan, Yingli Tian |
CVPR | 4 |
| 2021 | FESTA: Flow Estimation via Spatial-Temporal Attention for Scene Point CloudsabstractScene flow depicts the dynamics of a 3D scene, which is critical for various applications such as autonomous driving, robot navigation, AR/VR, etc. Conventionally, scene flow is estimated from dense/regular RGB video frames. With the development of depth-sensing technologies, precise 3D measurements are available via point clouds which have sparked new research in 3D scene flow. Nevertheless, it remains challenging to extract scene flow from point clouds due to the sparsity and irregularity in typical point cloud sampling patterns. One major issue related to irregular sampling is identified as the randomness during point set abstraction/feature extraction—an elementary process in many flow estimation scenarios. A novel Spatial Abstraction with Attention (SA2) layer is accordingly proposed to alleviate the unstable abstraction problem. Moreover, a Temporal Abstraction with Attention (TA2) layer is proposed to rectify attention in temporal domain, leading to benefits with motions scaled in a larger range. Extensive analysis and experiments verified the motivation and significant performance gains of our method, dubbed as Flow Estimation via Spatial-Temporal Attention (FESTA), when compared to several state-of-the-art benchmarks of scene flow estimation. Haiyan Wang 0019, Jiahao Pang, Muhammad Asad Lodhi, Yingli Tian, Dong Tian |
CVPR | 4 |
| 2021 | Subsurface Pipes Detection Using DNN-based Back Projection on GPR DataabstractLocalization and reconstruction of underground targets, the problem of estimating the position and geometry of the objects from Ground Penetration Radar (GPR), still lies at the core of non-destructive testing (NDT). In this paper, we present MigrationNet, a learning-based approach to detect and visualize subsurface objects. Compared with the existing learning-based method of GPR, our proposed approach could not only detect the hyperbola feature in the raw B-scan image but also interpret hyperbola features into the cross-section image of subsurface pipes. Furthermore, to compare the proposed method with the conventional back-projection methods for GPR data interpretation, a synthetic GPR dataset that mimics the real NDT environment is also introduced in this work. The study indicates the effectiveness of our method, it uses less GPR data for underground pipes reconstruction, produces better GPR imaging results with less computation, and shows the robustness to noise. Jinglun Feng, Haiyan Wang 0019, Yingli Tian, Jizhong Xiao |
WACV | 4 |
| 2021 | VideoSSL: Semi-Supervised Learning for Video ClassificationabstractWe propose a semi-supervised learning approach for video classification, VideoSSL, using convolutional neural networks (CNN). Like other computer vision tasks, existing supervised video classification methods demand a large amount of labeled data to attain good performance. However, annotation of a large dataset is expensive and time consuming. To minimize the dependence on a large annotated dataset, our proposed semi-supervised method trains from a small number of labeled examples and exploits two regulatory signals from unlabeled data. The first signal is the pseudo-labels of unlabeled examples computed from the confidences of the CNN being trained. The other is the normalized probabilities, as predicted by an image classifier CNN, that captures the information about appearances of the interesting objects in the video. We show that, under the supervision of these guiding signals from unlabeled examples, a video classification CNN can achieve impressive performances utilizing a small fraction of annotated examples on three publicly available datasets: UCF101, HMDB51, and Kinetics. Longlong Jing, Toufiq Parag, Yingli Tian |
WACV | 4 |
| 2021 | Self-supervised 4D Spatio-temporal Feature Learning via Order Prediction of Sequential Point Cloud ClipsabstractRecently 3D scene understanding attracts attention for many applications, however, annotating a vast amount of 3D data for training is usually expensive and time consuming. To alleviate the needs of ground truth, we propose a self-supervised schema to learn 4D spatio-temporal features (i.e. 3 spatial dimensions plus 1 temporal dimension) from dynamic point cloud data by predicting the temporal order of sampled and shuffled point cloud clips. 3D sequential point cloud contains precious geometric and depth information to better recognize activities in 3D space compared to videos. To learn the 4D spatio-temporal features, we introduce 4D convolution neural networks to predict the temporal order on a self-created large scale dataset, NTU-PCLs, derived from the NTU-RGB+D dataset. The efficacy of the learned 4D spatio-temporal features is verified on two tasks: 1) Self-supervised 3D nearest neighbor retrieval; and 2) Self-supervised representation learning transferred for action recognition on a smaller 3D dataset. Our extensive experiments prove the effectiveness of the proposed self-supervised learning method which achieves comparable results w.r.t. the fully-supervised methods on action recognition on MSRAction3D dataset. Haiyan Wang 0019, Xuejian Rong, Jinglun Feng, Yingli Tian |
WACV | 5 |
| 2021 | Self-Supervised Visual Feature Learning With Deep Neural Networks: A SurveyabstractLarge-scale labeled data are generally required to train deep neural networks in order to obtain better performance in visual feature learning from images or videos for computer vision applications. To avoid extensive cost of collecting and annotating large-scale datasets, as a subset of unsupervised learning methods, self-supervised learning methods are proposed to learn general image and video features from large-scale unlabeled data without using any human-annotated labels. This paper provides an extensive review of deep learning-based self-supervised general visual feature learning methods from images or videos. First, the motivation, general pipeline, and terminologies of this field are described. Then the common deep neural network architectures that used for self-supervised learning are summarized. Next, the schema and evaluation metrics of self-supervised learning methods are reviewed followed by the commonly used datasets for images, videos, audios, and 3D data, as well as the existing self-supervised visual feature learning methods. Finally, quantitative performance comparisons of the reviewed methods on benchmark datasets are summarized and discussed for both image and video feature learning. At last, this paper is concluded and lists a set of promising future directions for self-supervised visual feature learning. Longlong Jing, Yingli Tian |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | A Heuristic Neural Network Structure Relying on Fuzzy Logic for Images ScoringabstractTraditional deep learning methods are sub-optimal in classifying ambiguity features, which often arise in noisy and hard to predict categories, especially, to distinguish semantic scoring. Semantic scoring, depending on semantic logic to implement evaluation, inevitably contains fuzzy description and misses some concepts, for example, the ambiguous relationship between normal and probably normal always presents unclear boundaries (normal - more likely normal - probably normal). Thus, human error is common when annotating images. Differing from existing methods that focus on modifying kernel structure of neural networks, this study proposes a dominant fuzzy fully connected layer (FFCL) for Breast Imaging Reporting and Data System (BI-RADS) scoring and validates the universality of this proposed structure. This proposed model aims to develop complementary properties of scoring for semantic paradigms, while constructing fuzzy rules based on analyzing human thought patterns, and to particularly reduce the influence of semantic conglutination. Specifically, this semantic-sensitive defuzzier layer projects features occupied by relative categories into semantic space, and a fuzzy decoder modifies probabilities of the last output layer referring to the global trend. Moreover, the ambiguous semantic space between two relative categories shrinks during the learning phases, as the positive and negative growth trends of one category appearing among its relatives were considered. We first used the Euclidean Distance (ED) to zoom in the distance between the real scores and the predicted scores, and then employed two sample t test method to evidence the advantage of the FFCL architecture. Extensive experimental results performed on the CBIS-DDSM dataset show that our FFCL structure can achieve superior performances for both triple and multiclass classification in BI-RADS scoring, outperforming the state-of-the-art methods. Cheng Kang, Shuihua Wang, David S. Guttery, Hari Mohan Pandey, Yingli Tian, Yudong Zhang 0001 |
IEEE Trans. Fuzzy Syst. | 6 |
| 2020 | Burst Denoising via Temporally Shifted Wavelet Transforms
Xuejian Rong, Denis Demandolx, Kevin Matzen, Priyam Chatterjee, Yingli Tian |
ECCV (13) | 5 |
| 2020 | Recognizing American Sign Language Nonmanual Signal Grammar Errors in Continuous VideosabstractAs part of the development of an educational tool that can help students achieve fluency in American Sign Language (ASL) through independent and interactive practice with immediate feedback, this paper introduces a near real-time system to recognize grammatical errors in continuous signing videos without necessarily identifying the entire sequence of signs. Our system automatically recognizes if a performance of ASL sentences contains grammatical errors made by ASL students. We first recognize the ASL grammatical elements including both manual gestures and nonmanual signals independently from multiple modalities (i.e. hand gestures, facial expressions, and head movements) by 3D-ResNet networks. Then the temporal boundaries of grammatical elements from different modalities are examined to detect ASL grammatical mistakes by using a sliding window-based approach. We have collected a dataset of continuous sign language, ASL-HW-RGBD, covering different aspects of ASL grammars for training and testing. Our system is able to recognize grammatical elements on ASL-HW-RGBD from manual gestures, facial expressions, and head movements and successfully detect 8 ASL grammatical mistakes. Elahe Vahdani, Longlong Jing, Yingli Tian, Matt Huenerfauth |
ICPR | 3 |
| 2020 | Towards Efficient 3D Point Cloud Scene Completion via Novel Depth View Synthesisabstract3D point cloud completion has been a long-standing challenge at scale, and corresponding per-point supervised training strategies suffered from cumbersome annotations. 2D supervision has recently emerged as a promising alternative for 3D tasks, but specific approaches for 3D point cloud completion still remain to be explored. To overcome these limitations, we propose an end-to-end method that directly lifts a single depth map to a completed point cloud. With one depth map as input, a multiway novel depth view synthesis network (NDVNet) is designed to infer coarsely completed depth maps under various viewpoints. Meanwhile, a geometric depth perspective rendering module is introduced to utilize the raw input depth map to generate a re-projected depth map for each view. Therefore, the two parallelly generated depth maps for each view are further concatenated and refined by a depth completion network (DCNet). The final completed point cloud is fused from all refined depth views. Experimental results demonstrate the effectiveness of our proposed approach composed of aforementioned components, to produce high-quality, state-of-the-art results on the popular SUNCG benchmark. Haiyan Wang 0019, Xuejian Rong, Yingli Tian |
ICPR | 4 |
| 2020 | Deep Active Learning for Effective Pulmonary Nodule Detection
Jingya Liu, Liangliang Cao, Yingli Tian |
MICCAI (6) | 3 |
| 2020 | Monocular human pose estimation: A survey of deep learning-based methods
Yingli Tian, Mingyi He |
Comput. Vis. Image Underst. | 2 |
| 2020 | Coarse-to-Fine Semantic Segmentation From Image-Level LabelsabstractDeep neural network-based semantic segmentation generally requires large-scale cost extensive annotations for training to obtain better performance. To avoid pixel-wise segmentation annotations that are needed for most methods, recently some researchers attempted to use object-level labels (e.g., bounding boxes) or image-level labels (e.g., image categories). In this paper, we propose a novel recursive coarse-to-fine semantic segmentation framework based on only image-level category labels. For each image, an initial coarse mask is first generated by a convolutional neural network-based unsupervised foreground segmentation model and then is enhanced by a graph model. The enhanced coarse mask is fed to a fully convolutional neural network to be recursively refined. Unlike the existing image-level label-based semantic segmentation methods, which require labeling of all categories for images that contain multiple types of objects, our framework only needs one label for each image and can handle images that contain multi-category objects. Only trained on ImageNet, our framework achieves comparable performance on the PASCAL VOC dataset with other image-level label-based state-of-the-art methods of semantic segmentation. Furthermore, our framework can be easily extended to foreground object segmentation task and achieves comparable performance with the state-of-the-art supervised methods on the Internet object dataset. Longlong Jing, Yingli Tian |
IEEE Trans. Image Process. | 3 |
| 2020 | Unambiguous Scene Text Segmentation With Referring Expression ComprehensionabstractText instance provides valuable information for the understanding and interpretation of natural scenes. The rich, precise high-level semantics embodied in the text could be beneficial for understanding the world around us, and empower a wide range of real-world applications. While most recent visual phrase grounding approaches focus on general objects, this paper explores extracting designated texts and predicting unambiguous scene text segmentation mask, i.e. scene text segmentation from natural language descriptions (referring expressions) like orange text on a little boy in black swinging a bat. The solution of this novel problem enables accurate segmentation of scene text instances from the complex background. In our proposed framework, a unified deep network jointly models visual and linguistic information by encoding both region-level and pixel-level visual features of natural scene images into spatial feature maps, and then decode them into saliency response map of text instances. To conduct quantitative evaluations, we establish a new scene text referring expression segmentation dataset: COCO-CharRef. Experimental results demonstrate the effectiveness of the proposed framework on the text instance segmentation task. By combining image-based visual features with language-based textual explanations, our framework outperforms baselines that are derived from state-of-the-art text localization and natural language object retrieval methods on COCO-CharRef dataset. Xuejian Rong, Chucai Yi, Yingli Tian |
IEEE Trans. Image Process. | 3 |
| 2019 | Towards Weakly Supervised Semantic Segmentation in 3D Graph-Structured Point Clouds of Wild Scenes
Haiyan Wang 0019, Xuejian Rong, Shuihua Wang, Yingli Tian |
BMVC | 5 |
| 2019 | Towards Accurate Instance-Level Text Spotting with Guided AttentionabstractWe tackle the text detection problem from the instance-aware segmentation perspective, in which text bounding boxes are directly extracted from segmentation results without location regression. Specifically, a text-specific attention model and a global enhancement block are introduced to enrich the semantics of text detection features. The attention model is trained with a weakly segmentation supervision signal and enforces the detector to focus on the text regions, while also suppressing the influence of neighboring background clutters. In conjunction with the attention model, a global enhancement block (GEB) is adapted to reason the relationship among different channels with channel-wise weights calibration. Our method achieves comparable performance with the recent state-of-the-arts on ICDAR2013, ICDAR2015, and ICDAR2017-MLT benchmark datasets. Haiyan Wang 0019, Xuejian Rong, Yingli Tian |
ICME | 3 |
| 2019 | 3DFPN-HS ^2 2 : 3D Feature Pyramid Network Based High Sensitivity and Specificity Pulmonary Nodule Detection
Jingya Liu, Liangliang Cao, Oguz Akin, Yingli Tian |
MICCAI (6) | 4 |
| 2019 | Incremental Scene SynthesisabstractWe present a method to incrementally generate complete 2D or 3D scenes with the following properties: (a) it is globally consistent at each step according to a learned scene prior, (b) real observations of a scene can be incorporated while observing global consistency, (c) unobserved regions can be hallucinated locally in consistence with previous observations, hallucinations and global priors, and (d) hallucinations are statistical in nature, i.e., different scenes can be generated from the same observations. To achieve this, we model the virtual scene, where an active agent at each step can either perceive an observed part of the scene or generate a local hallucination. The latter can be interpreted as the agent's expectation at this step through the scene and can be applied to autonomous navigation. In the limit of observing real data at each point, our method converges to solving the SLAM problem. It can otherwise sample entirely imagined scenes from prior distributions. Besides autonomous agents, applications include problems where large data is required for building robust real-world applications, but few samples are available. We demonstrate efficacy on various 2D as well as 3D data. Benjamin Planche, Xuejian Rong, Ziyan Wu 0001, Srikrishna Karanam, Harald Kosch, Yingli Tian, Jan Ernst, Andreas Hutter |
NeurIPS | 6 |
| 2019 | Multimodal clothing recognition for semantic search in unconstrained surveillance imagery
Michael Halstead, Simon Denman, Sridha Sridharan, Yingli Tian, Clinton Fookes |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Discovering spatio-temporal action tubes
Yuancheng Ye, Xiaodong Yang 0001, Yingli Tian |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Prediction of Sea Ice Motion With Convolutional Long Short-Term Memory NetworksabstractPrediction of sea ice motion is important for safeguarding human activities in polar regions, such as ship navigation, fisheries, and oil and gas exploration, as well as for climate and ocean-atmosphere interaction models. Numerical prediction models used for sea ice motion prediction often require a large number of data from diverse sources with varying uncertainties. In this paper, a deep learning approach is proposed to predict sea ice motion for several days in the future, given only a series of past motion observations. The proposed approach consists of an encoder-decoder network with convolutional long short-term memory (LSTM) units. Optical flow is calculated from satellite passive microwave and scatterometer daily images covering the entire Arctic and used in the network. The network proves able to learn long-time dependencies within the motion time series, whereas its convolutional structure effectively captures spatial correlations among neighboring motion vectors. The approach is unsupervised and end-to-end trainable, requiring no manual annotation. Experiments demonstrate that the proposed approach is effective in predicting sea ice motion of up to 10 days in the future, outperforming previous deep learning networks and being a promising alternative or complementary approach to resource-demanding numerical prediction methods. Zisis I. Petrou, Yingli Tian |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Self-Guiding Multimodal LSTM - When We Do Not Have a Perfect Training Dataset for Image CaptioningabstractIn this paper, a self-guiding multimodal LSTM (sgLSTM) image captioning model is proposed to handle an uncontrolled imbalanced real-world image-sentence dataset. We collect a FlickrNYC dataset from Flickr as our testbed with 306,165 images and the original text descriptions uploaded by the users are utilized as the ground truth for training. Descriptions in the FlickrNYC dataset vary dramatically ranging from short term-descriptions to long paragraph-descriptions and can describe any visual aspects, or even refer to objects that are not depicted. To deal with the imbalanced and noisy situation and to fully explore the dataset itself, we propose a novel guiding textual feature extracted utilizing a multimodal LSTM (mLSTM) model. Training of mLSTM is based on the portion of data in which the image content and the corresponding descriptions are strongly bonded. Afterward, during the training of sgLSTM on the rest training data, this guiding information serves as additional input to the network along with the image representations and the ground-truth descriptions. By integrating these input components into a multimodal block, we aim to form a training scheme with the textual information tightly coupled with the image content. The experimental results demonstrate that the proposed sgLSTM model outperforms the traditional state-of-the-art multimodal RNN captioning framework in successfully describing the key components of the input images. Yang Xian, Yingli Tian |
IEEE Trans. Image Process. | 2 |
| 2019 | Vision-Based Mobile Indoor Assistive Navigation Aid for Blind PeopleabstractThis paper presents a new holistic vision-based mobile assistive navigation system to help blind and visually impaired people with indoor independent travel. The system detects dynamic obstacles and adjusts path planning in real-time to improve navigation safety. First, we develop an indoor map editor to parse geometric information from architectural models and generate a semantic map consisting of a global 2D traversable grid map layer and context-aware layers. By leveraging the visual positioning service (VPS) within the Google Tango device, we design a map alignment algorithm to bridge the visual area description file (ADF) and semantic map to achieve semantic localization. Using the on-board RGB-D camera, we develop an efficient obstacle detection and avoidance approach based on a time-stamped map Kalman filter (TSM-KF) algorithm. A multi-modal human-machine interface (HMI) is designed with speech-audio interaction and robust haptic interaction through an electronic SmartCane. Finally, field experiments by blindfolded and blind subjects demonstrate that the proposed system provides an effective tool to help blind individuals with indoor navigation and wayfinding. Bing Li 0008, Juan Pablo Muñoz, Xuejian Rong, Qingtian Chen, Jizhong Xiao, Yingli Tian, Aries Arditi, Mohammed Yousuf |
IEEE Trans. Mob. Comput. | 6 |
| 2018 | Semantic Person Retrieval in Surveillance Using Soft Biometrics: AVSS 2018 Challenge IIabstractIn surveillance and security today it is a common goal to locate a subject of interest purely from a semantic description; think of an offender description form handed into a law enforcement agency. To date, these tasks are primarily undertaken by operators on the ground either by manually searching a premises or by combing through hours of video footage. Using computer vision to attempt to partially or fully automate these tasks has been gathering interest within the research community in recent years, however, to date there has been little coordinated effort to advance the field. This has motivated the challenge that is presented in this paper: the AVSS Challenge on Semantic Person Retrieval in Surveillance Using Soft Biometrics. This challenge consists of two related tasks: person re-identification from a semantic query and person search within a video from a query. In this paper, we present the publicly available data for this challenge, the evaluation framework, and the challenge results. It is our hope that the outcomes of this challenge and the availability of the data used in this challenge will expedite research and development in this societal field. Michael Halstead, Simon Denman, Clinton Fookes, Yingli Tian, Mark S. Nixon |
AVSS | 4 |
| 2018 | DAAL: Deep activation-based attribute learning for action recognition in depth videos
Chenyang Zhang 0001, Yingli Tian, Xiaojie Guo 0001, Jingen Liu |
Comput. Vis. Image Underst. | 2 |
| 2018 | Video you only look once: Overall temporal convolutions for action recognition
Longlong Jing, Xiaodong Yang 0001, Yingli Tian |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Unambiguous Text Localization and Retrieval for Cluttered ScenesabstractText instance as one category of self-described objects provides valuable information for understanding and describing cluttered scenes. In this paper, we explore the task of unambiguous text localization and retrieval, to accurately localize a specific targeted text instance in a cluttered image given a natural language description that refers to it. To address this issue, first a novel recurrent Dense Text Localization Network (DTLN) is proposed to sequentially decode the intermediate convolutional representations of a cluttered scene image into a set of distinct text instance detections. Our approach avoids repeated detections at multiple scales of the same text instance by recurrently memorizing previous detections, and effectively tackles crowded text instances in close proximity. Second, we propose a Context Reasoning Text Retrieval (CRTR) model, which jointly encodes text instances and their context information through a recurrent network, and ranks localized text bounding boxes by a scoring function of context compatibility. Quantitative evaluations on standard scene text localization benchmarks and a newly collected scene text retrieval dataset demonstrate the effectiveness and advantages of our models for both scene text localization and retrieval. Xuejian Rong, Chucai Yi, Yingli Tian |
CVPR | 3 |
| 2017 | 3D convolutional neural network with multi-model framework for action recognitionabstractIn this paper, we propose an efficient and effective action recognition framework by combining multiple feature models from dynamic image, optical flow and raw frame, with 3D convolutional neural network (CNN). Dynamic image preserves the long-term temporal information, while optical flow captures short-term temporal information, and raw frame represents the appearance information. Experiments demonstrate that dynamic image provides complementary information to raw frame feature and optical flow feature. Furthermore, with the approximate rank pooling, the computation of dynamic images is about 360 times faster than optical flow, and the dynamic image requires far less memory than optical flow and raw frame. Longlong Jing, Yuancheng Ye, Xiaodong Yang 0001, Yingli Tian |
ICIP | 4 |
| 2017 | Prediction of sea ice motion with recurrent neural networksabstractPrediction of sea ice motion is important for ocean-atmosphere interaction modeling and safe naval operations in polar regions. In this study, we investigate the potential of Recurrent Neural Networks (RNNs) in predictions of motion for several days in the future based only on previously observed satellite image data. We collect a large dataset of daily Advanced Microwave Scanning Radiometer - Earth Observing System (AMSR-E) images that cover the entire Arctic. Optical flow is employed to calculate dense sea ice motion between images of each consecutive-day pair. The optical flow images are then used to train an encoder-decoder Long Short-Term Memory (LSTM) RNN and estimate motion for several days in the future. Experiments demonstrate that the proposed method is successful in predicting short-term sea ice motion with accuracy close to motion calculated from the original images and buoys, and proves promising for further applications and research. Zisis I. Petrou, Yingli Tian |
IGARSS | 2 |
| 2017 | Increasing spatial resolution of sea ice motion estimationabstractEstimation of sea ice motion at a fine scale is essential for climate modeling and naval operations in polar regions. This study proposes an approach for increasing the spatial resolution of the motion estimated from passive microwave satellite images. A hierarchical pattern matching approach, based on normalized cross-correlation and phase correlation, is applied to calculate sea ice drifts between Advanced Microwave Scanning Radiometer 2 (AMSR2) pairs of images captured at different time. Contrary to the widely used approach of embedding oversampling within the pattern matching framework that increases the resolution of the minimum detectable drift, in this study a nearest-neighbor interpolation upscaling is applied directly to the images. This additionally increases the density of the estimated motion vectors. Experiments demonstrate that the proposed method can lead to motion vector field with increased resolution by up to eight times, while outperforming the original images. Image upscaling by four times also outperforms corresponding previous examples with embedded oversampling. Zisis I. Petrou, Yang Xian, Yingli Tian |
IGARSS | 3 |
| 2017 | Super Normal Vector for Human Activity Recognition with Depth CamerasabstractThe advent of cost-effectiveness and easy-operation depth cameras has facilitated a variety of visual recognition tasks including human activity recognition. This paper presents a novel framework for recognizing human activities from video sequences captured by depth cameras. We extend the surface normal to polynormal by assembling local neighboring hypersurface normals from a depth sequence to jointly characterize local motion and shape information. We then propose a general scheme of super normal vector (SNV) to aggregate the low-level polynormals into a discriminative representation, which can be viewed as a simplified version of the Fisher kernel representation. In order to globally capture the spatial layout and temporal order, an adaptive spatio-temporal pyramid is introduced to subdivide a depth video into a set of space-time cells. In the extensive experiments, the proposed approach achieves superior performance to the state-of-the-art methods on the four public benchmark datasets, i.e., MSRAction3D, MSRDailyActivity3D, MSRGesture3D, and MSRActionPairs3D. Xiaodong Yang 0001, Yingli Tian |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Evaluation of Low-Level Features for Real-World Surveillance Event DetectionabstractEvent detection targets at recognizing and localizing specified spatio-temporal patterns in videos. Most research of human activity recognition in the past decades experimented on relatively clean scenes with limited actors performing explicit actions. Recently, more efforts have been paid to the real-world surveillance videos in which the human activity recognition is more challenging due to large variations caused by factors, such as scaling, resolution, viewpoint, cluttered background, and crowdedness. In this paper, we systematically evaluate seven different types of low-level spatio-temporal features in the context of surveillance event detection (SED) using a uniform experimental setup. Fisher vector is employed to aggregate low-level features as the representation of each video clip. A set of random forests is then learned as the classification models. To bridge the research efforts and real-world applications, we utilize the NIST TRECVID SED as our testbed in which seven events are predefined involving different levels of human activity analysis. Strengths and limitations for each low-level feature type are analyzed and discussed. Yang Xian, Xuejian Rong, Xiaodong Yang 0001, Yingli Tian |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | High-Resolution Sea Ice Motion Estimation With Optical Flow Using Satellite Spectroradiometer DataabstractThe purpose of this paper is twofold: 1) to propose an approach based on optical flow for the estimation of sea ice motion as an accurate, dense, and computationally efficient alternative to state-of-the-art pattern matching approaches and 2) to investigate the potential of moderate resolution imaging spectroradiometer (MODIS) optical satellite data for combined daily and high-resolution motion estimation. A series of MODIS image pairs for a selected region in Arctic Ocean is employed, and sparse pairwise correspondences between nonrigid patches are calculated. An edge-preserving sparse-to-dense interpolation is applied followed by variational energy minimization to compute the final optical flow. A state-of-the-art multiresolution pattern matching method based on phase correlation and normalized cross correlation is also implemented and evaluated for comparison. Thorough experimentation with different settings of input data preprocessing as well as varying scales of motion is performed. The derived motion vectors are compared with coarser resolution operational sea ice motion vector products from combined buoy and microwave satellite data. The proposed optical flow approach clearly outperforms the pattern matching method in most cases, both in terms of accuracy and motion vector consistency. The estimated motion vectors from the MODIS images highly correlate with the operational vector products. In addition, MODIS provides a vector field with spatial resolution two orders of magnitude higher than the operational products that are able to detect significantly smaller sea ice drifts. Zisis I. Petrou, Yingli Tian |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Super-Resolved Fine-Scale Sea Ice Motion TrackingabstractMonitoring sea ice activities is particularly critical to safe naval operations in the Arctic Ocean. Accurately tracking sea ice motions is essential to validate or even improve sea ice models for ice hazard forecasts at a fine scale. Fine-scale motions can be tracked from high-resolution radar or optical satellite imagery but with limited coverage. Daily motions over the entire Arctic are retrievable from passive microwave data, but at a much lower spatial resolution. Thus, providing motions at the passive microwave spatial and temporal coverage, but at an enhanced spatial resolution, will be a significant benefit. To break the resolution limitation and to boost tracking accuracy, a sequential super-resolved fine-scale sea ice motion tracking framework is proposed in which a hybrid example-based single image super-resolution algorithm is employed before the tracking procedure. Experiments demonstrate that the proposed framework significantly improves the tracking performance in both accuracy and robustness for a benchmark algorithm and a recently proposed state-of-the-art tracking algorithm. Yang Xian, Zisis I. Petrou, Yingli Tian, Walter N. Meier |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | Multi-modality American Sign Language recognitionabstractAmerican Sign Language (ASL) is a visual gestural language which is used by many people who are deaf or hard-of-hearing. In this paper, we design a visual recognition system based on action recognition techniques to recognize individual ASL signs. Specifically, we focus on recognition of words in videos of continuous ASL signing. The proposed framework combines multiple signal modalities because ASL includes gestures of both hands, body movements, and facial expressions. We have collected a corpus of RBG + depth videos of multi-sentence ASL performances, from both fluent signers and ASL students; this corpus has served as a source for training and testing sets for multiple evaluation experiments reported in this paper. Experimental results demonstrate that the proposed framework can automatically recognize ASL. Chenyang Zhang 0001, Yingli Tian, Matt Huenerfauth |
ICIP | 2 |
| 2016 | BCA: Bi-symmetric component analysis for temporal symmetry in human actionsabstractIn the past, many research efforts are invested into discriminative action recognition task but the general temporal structure of human actions is overlooked. In this paper, we focus on a specific yet common structure of human actions: temporal symmetry. The key contribution is that we model the temporal symmetry property of human action and separate this signal out of original action sequences without specifying which action category. Based on this modeling, a novel and effective method is proposed to detect the temporal symmetric part of any given human action sequence. Experimental results on two popular human action datasets verify that the temporal symmetry benefits both action detection and action recognition. Chenyang Zhang 0001, Yingli Tian |
ICME | 2 |
| 2016 | Automatic video description generation via LSTM with joint two-stream encodingabstractIn this paper, we propose a novel two-stream framework based on combinational deep neural networks. The framework is mainly composed of two components: one is a parallel two-stream encoding component which learns video encoding from multiple sources using 3D convolutional neural networks and the other is a long-short-term-memory (LSTM)-based decoding language model which transfers the input encoded video representations to text descriptions. The merits of our proposed model are: 1) It extracts both temporal and spatial features by exploring the usage of 3D convolutional networks on both raw RGB frames and motion history images. 2) Our model can dynamically tune the weights of different feature channels since the network is trained end-to-end from learning combinational encoding of multiple features to LSTM-based language model. Our model is evaluated on three public video description datasets: one YouTube clips dataset (Microsoft Video Description Corpus) and two large movie description datasets (MPII Corpus and Montreal Video Annotation Dataset) and achieves comparable or better performance than the state-of-the-art approaches in video caption generation. Chenyang Zhang 0001, Yingli Tian |
ICPR | 2 |
| 2016 | Demo: Assisting Visually Impaired People Navigate Indoors
Juan Pablo Muñoz, Bing Li 0008, Xuejian Rong, Jizhong Xiao, Yingli Tian, Aries Arditi |
IJCAI | 5 |
| 2016 | Region Trajectories for Video Semantic Concept DetectionabstractRecently, with the advent of the convolutional neural network (CNN), many CNN-based object detection algorithms have been proposed and achieved encouraging results. In this paper, we introduce an algorithm based on region trajectories to establish the connections between object localizations in individual frames and video sequences. To detect object regions in the individual frames of a video, we enhance the region-based convolutional neural network (R-CNN), by incorporating EdgeBox with the Selective Search to generate candidate region proposals and combining the GoogLeNet with the AlexNet to improve the discriminability of the feature representations. The DeepMatching algorithm is employed in our proposed region trajectory method to track the points in the detected object regions. The experiments are conducted on the validation split of the TRECVID 2015 Localization dataset. As demonstrated by the experimental results, our proposed approach improves the object detection accuracy in both temporal and spatial measurements. Yuancheng Ye, Xuejian Rong, Xiaodong Yang 0001, Yingli Tian |
ICMR | 4 |
| 2016 | Resolution enhancement in single depth map and aligned imageabstractDepth resolution enhancement aims to recover a high quality depth map from one or multiple low-resolution depth input(s) with missing pixels. While a registered high-resolution intensity image is often utilized to assist, little attention has been paid to the circumstances when there is only one pair of low-resolution depth map and aligned intensity image available. In this paper, we propose a novel resolution enhancement approach that targets at improving the quality of both the input depth map and the low-resolution RGB image. By exploiting the statistical dependency between the input pairs, a label matrix is generated utilizing the support vector machine classifier. Guided by the constructed label matrix and the aligned intensity image, the missing values in the depth map are well predicted in a manner consistent with the embedded structure. After that, the completed depth map and the intensity image are super-resolved through a set of regression models trained via external exemplars. Extensive experiments demonstrate that our framework is effective with satisfying performance. Yingli Tian, Yang Xian |
WACV | 1 |
| 2016 | 3D-based Deep Convolutional Neural Network for action recognition with depth sequences
Zhi Liu 0013, Chenyang Zhang 0001, Yingli Tian |
Image Vis. Comput. | 3 |
| 2016 | Single image super-resolution via internal gradient similarity
Yang Xian, Yingli Tian |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Robust internal exemplar-based image enhancementabstractImage enhancement aims to modify images to achieve a better perception for human visual system or a more suitable representation for further analysis. Based on different attributes of given input images, tasks vary, e.g., noise removal, deblur-ring, resolution enhancement, prediction of missing pixels, etc. The latter two are usually referred to as image super-resolution and image inpainting. There exist complicated circumstances where low-quality input images suffer from insufficient resolution with missing regions. In this paper, we propose a novel uniform framework to accomplish both image super-resolution and inpainting simultaneously. The proposed approach adopts internal exemplar similarities in image level and gradient level where later enhancement results from both levels are fed into a pre-defined cost function to restore the final output. Experimental results demonstrate that our method is capable of generating visually plausible, natural-looking results with clear edges and realistic textures. Yang Xian, Yingli Tian |
ICIP | 2 |
| 2015 | Exploring Pooling Strategies based on Idiosyncrasies of Spatio-Temporal Interest PointsabstractRecent studies have demonstrated that the implementation of local space-time interest points has good competence and robustness in the area of human action recognition, which has become one of the challenging problems in multimedia analysis. While most research focuses on the techniques of detecting feature points or capturing spatial and temporal information around those points, there has been very limited research on delving into the pooling strategies which are also important components of action recognition algorithms. In this paper, we propose a novel pooling framework by categorizing the interest points with respect to their idiosyncrasies. Specifically, we discuss three pooling strategies based on the optical flow orientation, foreground weight and spatio-temporal locations respectively and further investigate the fusion of different pooling strategies. For the encoding process, instead of the popular bag-of-visual words (BoV) method, we adopt the improved Fisher Vector (FV) approach. Our proposed methods are evaluated on a benchmark dataset with controlled settings (KTH), and two more challenging datasets with realistic background (HMDB51 and UCF101). The experimental results demonstrate that pooling strategies based on the appropriate idiosyncrasies of individual interest points can improve the performance of action classification. Yuancheng Ye, Xiaodong Yang 0001, Yingli Tian |
ICMR | 3 |
| 2015 | A SLAM Based Semantic Indoor Navigation System for Visually Impaired UsersabstractThis paper proposes a novel assistive navigation system based on simultaneous localization and mapping (SLAM) and semantic path planning to help visually impaired users navigate in indoor environments. The system integrates multiple wearable sensors and feedback devices including a RGB-D sensor and an inertial measurement unit (IMU) on the waist, a head mounted camera, a microphone and an earplug/speaker. We develop a visual odometry algorithm based on RGB-D data to estimate the user's position and orientation, and refine the orientation error using the IMU. We employ the head mounted camera to recognize the door numbers and the RGB-D sensor to detect major landmarks such as corridor corners. By matching the detected landmarks against the corresponding features on the digitalized floor map, the system localizes the user, and provides verbal instruction to guide the user to the desired destination. The software modules of our system are implemented in Robotics Operating System (ROS). The prototype of the proposed assistive navigation system is evaluated by blindfolded sight persons. The field tests confirm the feasibility of the proposed algorithms and the system prototype. Bing Li 0008, Samleo L. Joseph, Jizhong Xiao, Yi Sun 0005, Yingli Tian, Juan Pablo Muñoz, Chucai Yi |
SMC | 6 |
| 2015 | An effective view and time-invariant action recognition method based on depth videosabstractLittle progress has been achieved in hand-crafted feature based human action recognition (HAR) for RGB videos in recent years. The emergence of low price depth camera presents more information for action recognition. Compared to RGB videos, depth video sequences are more insensitive to light changes and more discriminative in many vision tasks such as segmentation and activity recognition. In this paper, we propose an effective and straightforward HAR method by using skeleton joints information of the depth sequence. First, we calculate three feature vectors which capture angle and position information between joints. Then, the obtained vectors are used as the inputs of three separate support vector machine (SVM) classifiers. Finally, the action recognition is conducted by fusing the SVM classification results. Our features are viewinvariant because the extracted vectors contain only angle and normalized position information based on joint coordinates. By normalizing action videos with different temporal lengths to a fixed size using interpolation, the extracted features have the same dimension for different videos and can still keep the principal movement patterns which make the proposed method timeinvariant. Experimental results demonstrate that our method performs comparable results on the UTKinect-Action3D dataset, and is more efficient and simpler than state-of-the-art methods. Zhi Liu 0013, Yingli Tian |
VCIP | 3 |
| 2015 | Histogram of 3D Facets: A depth descriptor for human action and hand gesture recognition
Chenyang Zhang 0001, Yingli Tian |
Comput. Vis. Image Underst. | 2 |
| 2015 | Pyramid of Spatial Relatons for Scene-Level Land Use ClassificationabstractLocal feature with bag-of-words (BOW) representation has become one of the most popular approaches in object classification and image retrieval applications in the computer vision community. The recent efforts in the remote sensing community have demonstrated that the BOW approach can also effectively apply to geographic images for the applications of classification and retrieval. However, the BOW representation discards spatial information, which is critical for the remotely sensed land use classification. Several algorithms have incorporated spatial information into the BOW representation by hard encoding coordinates of local features. Such rigid spatial encoding is not robust to translation and rotation variations, which are common characteristics of geographic images. To effectively incorporate spatial information into the BOW model for the land use classification, we propose a pyramid-of-spatial-relatons (PSR) model to capture both absolute and relative spatial relationships of local features. Unlike the conventional cooccurrence approach to describe pairwise spatial relationships between local features, the PSR model employs a novel concept of spatial relation to describe relative spatial relationship of a group of local features. As the result, the storage cost of the PSR model only linearly increases with the visual word codebook size instead of the quadratic relationship as in the cooccurrence approach. The PSR model is robust to translation and rotation variations and demonstrates excellent performance for the application of remotely sensed land use classification. On the Land Use and Land Cover image database, the PSR achieves 8% higher in the classification accuracy than the state of the art. If using only gray images, it outperforms the state of the art by more than 11%. Shizhi Chen, Yingli Tian |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | Discriminative Hierarchical K-Means Tree for Large-Scale Image ClassificationabstractA key challenge in large-scale image classification is how to achieve efficiency in terms of both computation and memory without compromising classification accuracy. The learning-based classifiers achieve the state-of-the-art accuracies, but have been criticized for the computational complexity that grows linearly with the number of classes. The nonparametric nearest neighbor (NN)-based classifiers naturally handle large numbers of categories, but incur prohibitively expensive computation and memory costs. In this brief, we present a novel classification scheme, i.e., discriminative hierarchical K-means tree (D-HKTree), which combines the advantages of both learning-based and NN-based classifiers. The complexity of the D-HKTree only grows sublinearly with the number of categories, which is much better than the recent hierarchical support vector machines-based methods. The memory requirement is the order of magnitude less than the recent Naïve Bayesian NN-based approaches. The proposed D-HKTree classification scheme is evaluated on several challenging benchmark databases and achieves the state-of-the-art accuracies, while with significantly lower computation cost and memory requirement. Shizhi Chen, Xiaodong Yang 0001, Yingli Tian |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Super Normal Vector for Activity Recognition Using Depth SequencesabstractThis paper presents a new framework for human activity recognition from video sequences captured by a depth camera. We cluster hypersurface normals in a depth sequence to form the polynormal which is used to jointly characterize the local motion and shape information. In order to globally capture the spatial and temporal orders, an adaptive spatio-temporal pyramid is introduced to subdivide a depth video into a set of space-time grids. We then propose a novel scheme of aggregating the low-level polynormals into the super normal vector (SNV) which can be seen as a simplified version of the Fisher kernel representation. In the extensive experiments, we achieve classification results superior to all previous published results on the four public benchmark datasets, i.e., MSRAction3D, MSRDailyActivity3D, MSRGesture3D, and MSRActionPairs3D. Xiaodong Yang 0001, Yingli Tian |
CVPR | 2 |
| 2014 | Action Recognition Using Super Sparse Coding Vector with Spatio-temporal Awareness
Xiaodong Yang 0001, Yingli Tian |
ECCV (2) | 2 |
| 2014 | Scene text recognition in multiple frames based on text trackingabstractText signage as visual indicators in natural scene plays an important role in navigation and notification in our daily life. Most previous methods of scene text extraction are developed from a single scene image. In this paper, we propose a multi-frame based scene text recognition method by tracking text regions in a video captured by a moving camera. The main contributions of this paper are as follows. First, we present a framework of scene text recognition in multiple frames based on feature representation of scene text character (STC) for character prediction and conditional random field (CRF) model for word configuration. Second, a feature representation of STC is employed from dense sampled SIFT descriptors and Fisher Vector. Third, we collect a dataset for text information extraction from natural scene videos. Our proposed multi-frame scene text recognition is more compatible with image/video-based mobile applications. The experimental results demonstrate that STC prediction and word configuration in multiple frames based on text tracking significantly improves the performance of scene text recognition. Xuejian Rong, Chucai Yi, Xiaodong Yang 0001, Yingli Tian |
ICME | 4 |
| 2014 | RGB-D image-based detection of stairs, pedestrian crosswalks and traffic signs
Shuihua Wang, Hangrong Pan, Chenyang Zhang 0001, Yingli Tian |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | Effective 3D action recognition using EigenJoints
Xiaodong Yang 0001, Yingli Tian |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Assistive Clothing Pattern Recognition for Visually Impaired PeopleabstractChoosing clothes with complex patterns and colors is a challenging task for visually impaired people. Automatic clothing pattern recognition is also a challenging research problem due to rotation, scaling, illumination, and especially large intraclass pattern variations. We have developed a camera-based prototype system that recognizes clothing patterns in four categories (plaid, striped, patternless, and irregular) and identifies 11 clothing colors. The system integrates a camera, a microphone, a computer, and a Bluetooth earpiece for audio description of clothing patterns and colors. A camera mounted upon a pair of sunglasses is used to capture clothing images. The clothing patterns and colors are described to blind users verbally. This system can be controlled by speech input through microphone. To recognize clothing patterns, we propose a novel Radon Signature descriptor and a schema to extract statistical properties from wavelet subbands to capture global features of clothing patterns. They are combined with local features to recognize complex clothing patterns. To evaluate the effectiveness of the proposed approach, we used the CCNY Clothing Pattern dataset. Our approach achieves 92.55% recognition accuracy which significantly outperforms the state-of-the-art texture analysis methods on clothing pattern recognition. The prototype was also used by ten visually impaired participants. Most thought such a system would support more independence in their daily life but they also made suggestions for improvements. Xiaodong Yang 0001, Yingli Tian |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2014 | Scene Text Recognition in Mobile Applications by Character Descriptor and Structure ConfigurationabstractText characters and strings in natural scene can provide valuable information for many applications. Extracting text directly from natural scene images or videos is a challenging task because of diverse text patterns and variant background interferences. This paper proposes a method of scene text recognition from detected text regions. In text detection, our previously proposed algorithms are applied to obtain text regions from scene image. First, we design a discriminative character descriptor by combining several state-of-the-art feature detectors and descriptors. Second, we model character structure at each character class by designing stroke configuration maps. Our algorithm design is compatible with the application of scene text extraction in smart mobile devices. An Android-based demo system is developed to show the effectiveness of our proposed method on scene text information extraction from nearby objects. The demo system also provides us some insight into algorithm design and performance improvement of scene text extraction. The evaluation results on benchmark data sets demonstrate that our proposed scheme of text recognition is comparable with the best existing methods. Chucai Yi, Yingli Tian |
IEEE Trans. Image Process. | 2 |
| 2013 | Visual speech learning from an e-tutor via dynamic lip movement-based video segmentation and comparisonabstractThis paper is motivated by the difficulties that deaf students encounter when learning speechreading and speaking; the skills that enable them to effectively communicate with hearing people. In this paper, we propose a speech learning prototype system based on the analysis and comparison of lip movements of an E-Tutor and those of a deaf student in a video. The main framework of our proposed system can be divided into two stages: lip movement segmentation and speech comparison. Lip movement segmentation fragments the frames of each word from a visual speech video sequence by analyzing the movement and shape of lips. Comparison determines whether a student is producing a correct word utterance or not, this is accomplished by comparing the lip shape and movements according to that of an e-tutor. To model lip movement, we compute two dynamic-based features by using a lip tracking method, which employs landmark points to define lip shapes. We utilize these dynamic features along with Space-Time Interest Points (STIP) to capture lip movements. In order to evaluate the effectiveness of our proposed methods, we collect a visual speech learning dataset consisting of 220 videos and 1100 word utterances. The proposed system achieves promising performances in both visual speech segmentation and visual speech comparison on this dataset. Carol Mazuera, Xiaodong Yang 0001, Yingli Tian |
BIBM | 3 |
| 2013 | Detecting good quality frames in videos captured by a wearable camera for blind navigationabstractRecent technology developments in computer vision, digital cameras, and portable computers make it possible to assist blind individuals by developing camera-based object recognition products. However, motion blur caused by a moving camera limits the real-world application of wayfinding for blind users. In this paper, we propose a new method to detect good quality frames from videos captured by cameras, which are taken by blind users. In our proposed method, both gradient and intensity statistics are extracted from video frames. Then a support vector machine (SVM) based classifier is applied to identify the frames with good quality (Unblurred) from those blurred frames. The Unblurred frames will be further processed to extract essential information for blind wayfinding and navigation such as signage recognition and text extraction. Experimental results demonstrate that our proposed method is able to robustly handle video motions in both indoor and outdoor environments. Yingli Tian, Chucai Yi |
BIBM | 2 |
| 2013 | Feature Representations for Scene Text Character Recognition: A Comparative StudyabstractRecognizing text character from natural scene images is a challenging problem due to background interferences and multiple character patterns. Scene Text Character (STC) recognition, which generally includes feature representation to model character structure and multi-class classification to predict label and score of character class, mostly plays a significant role in word-level text recognition. The contribution of this paper is a complete performance evaluation of image-based STC recognition, by comparing different sampling methods, feature descriptors, dictionary sizes, coding and pooling schemes, and SVM kernels. We systematically analyze the impact of each option in the feature representation and classification. The evaluation results on two datasets CHARS74K and ICDAR2003 demonstrate that Histogram of Oriented Gradient (HOG) descriptor, soft-assignment coding, max pooling, and Chi-Square Support Vector Machines (SVM) obtain the best performance among local sampling based feature representations. To improve STC recognition, we apply global sampling feature representation. We generate Global HOG (GHOG) by computing HOG descriptor from global sampling. GHOG enables better character structure modeling and obtains better performance than local sampling based feature representations. The GHOG also outperforms existing methods in the two benchmark datasets. Chucai Yi, Xiaodong Yang 0001, Yingli Tian |
ICDAR | 3 |
| 2013 | Reading labels of cylinder objects for blind personsabstractWe propose a camera-based assistive framework to help blind persons to read text labels from cylinder objects in their daily life. First, the object is detected from the background or other surrounding objects in the camera view by shaking the object. Then we propose a mosaic model to unwarp the text label on the cylinder object surface and reconstruct the whole label for recognizing text information. This model can handle cylinder objects in any orientations and scales. The text information is then extracted from the unwarped and flatted labels. The recognized text codes are then output to blind users in speech. Experimental results demonstrate the efficiency and effectiveness of the proposed framework from different cylinder objects with complex backgrounds. Ze Ye, Chucai Yi, Yingli Tian |
ICME | 3 |
| 2013 | Semantic Indoor Navigation with a Blind-User Oriented Augmented RealityabstractThe aim of this paper is to design an inexpensive conceivable wearable navigation system that can aid in the navigation of a visually impaired user. A novel approach of utilizing the floor plan map posted on the buildings is used to acquire a semantic plan. The extracted landmarks such as room numbers, doors, etc act as a parameter to infer the way points to each room. This provides a mental mapping of the environment to design a navigation framework for future use. A human motion model is used to predict a path based on how real humans ambulate towards a goal by avoiding obstacles. We demonstrate the possibilities of augmented reality (AR) as a blind user interface to perceive the physical constraints of the real world using haptic and voice augmentation. The haptic belt vibrates to direct the user towards the travel destination based on the metric localization at each step. Moreover, travel route is presented using voice guidance, which is achieved by accurate estimation of the user's location and confirmed by extracting the landmarks, based on landmark localization. The results show that it is feasible to assist a blind user to travel independently by providing the constraints required for safe navigation with user oriented augmented reality. Samleo L. Joseph, Ivan Dryanovski, Jizhong Xiao, Chucai Yi, Yingli Tian |
SMC | 6 |
| 2013 | Text extraction from scene images by character appearance and structure modeling
Chucai Yi, Yingli Tian |
Comput. Vis. Image Underst. | 2 |
| 2013 | Recognizing expressions from face and body gesture by temporal normalized motion and appearance features
Shizhi Chen, Yingli Tian, Qingshan Liu 0001, Dimitris N. Metaxas |
Image Vis. Comput. | 2 |
| 2013 | Toward a computer vision-based wayfinding aid for blind persons to access unfamiliar indoor environments
Yingli Tian, Xiaodong Yang 0001, Chucai Yi, Aries Arditi |
Mach. Vis. Appl. | 1 |
| 2013 | Texture representations using subspace embeddings
Xiaodong Yang 0001, Yingli Tian |
Pattern Recognit. Lett. | 2 |
| 2012 | Towards a Visual Speech Learning System for the Deaf by Matching Dynamic Lip Shapes
Shizhi Chen, D. Michael Quintian, Yingli Tian |
ICCHP (1) | 3 |
| 2012 | Visual Nouns for Indoor/Outdoor Navigation
Edgardo Molina, Zhigang Zhu 0001, Yingli Tian |
ICCHP (2) | 3 |
| 2012 | Camera-Based Signage Detection and Recognition for Blind Persons
Shuihua Wang, Yingli Tian |
ICCHP (2) | 2 |
| 2012 | Privacy Preserving Automatic Fall Detection for Elderly Using RGBD Cameras
Chenyang Zhang 0001, Yingli Tian, Elizabeth Capezuti |
ICCHP (1) | 2 |
| 2012 | Recognizing actions using depth motion maps-based histograms of oriented gradientsabstractIn this paper, we propose an effective method to recognize human actions from sequences of depth maps, which provide additional body shape and motion information for action recognition. In our approach, we project depth maps onto three orthogonal planes and accumulate global activities through entire video sequences to generate the Depth Motion Maps (DMM). Histograms of Oriented Gradients (HOG) are then computed from DMM as the representation of an action video. The recognition results on Microsoft Research (MSR) Action3D dataset show that our approach significantly outperforms the state-of-the-art methods, although our representation is much more compact. In addition, we investigate how many frames are required in our framework to recognize actions on the MSR Action3D dataset. We observe that a short sub-sequence of 30-35 frames is sufficient to achieve comparable results to that operating on entire video sequences. Xiaodong Yang 0001, Chenyang Zhang 0001, Yingli Tian |
ACM Multimedia | 3 |
| 2012 | Robust and efficient foreground analysis in complex surveillance videos
Yingli Tian, Andrew W. Senior, Max Lu |
Mach. Vis. Appl. | 1 |
| 2012 | Localizing Text in Scene Images by Boundary Clustering, Stroke Segmentation, and String Fragment ClassificationabstractIn this paper, we propose a novel framework to extract text regions from scene images with complex backgrounds and multiple text appearances. This framework consists of three main steps: boundary clustering (BC), stroke segmentation, and string fragment classification. In BC, we propose a new bigram-color-uniformity-based method to model both text and attachment surface, and cluster edge pixels based on color pairs and spatial positions into boundary layers. Then, stroke segmentation is performed at each boundary layer by color assignment to extract character candidates. We propose two algorithms to combine the structural analysis of text stroke with color assignment and filter out background interferences. Further, we design a robust string fragment classification based on Gabor-based text features. The features are obtained from feature maps of gradient, stroke distribution, and stroke width. The proposed framework of text localization is evaluated on scene images, born-digital images, broadcast video images, and images of handheld objects captured by blind persons. Experimental results on respective datasets demonstrate that the framework outperforms state-of-the-art localization algorithms. Chucai Yi, Yingli Tian |
IEEE Trans. Image Process. | 2 |
| 2012 | Robust and Effective Component-Based Banknote Recognition for the BlindabstractWe develop a novel camera-based computer vision technology to automatically recognize banknotes for assisting visually impaired people. Our banknote recognition system is robust and effective with the following features: 1) high accuracy: high true recognition rate and low false recognition rate, 2) robustness: handles a variety of currency designs and bills in various conditions, 3) high efficiency: recognizes banknotes quickly, and 4) ease of use: helps blind users to aim the target for image capture. To make the system robust to a variety of conditions including occlusion, rotation, scaling, cluttered background, illumination change, viewpoint variation, and worn or wrinkled bills, we propose a component-based framework by using Speeded Up Robust Features (SURF). Furthermore, we employ the spatial relationship of matched SURF features to detect if there is a bill in the camera view. This process largely alleviates false recognition and can guide the user to correctly aim at the bill to be recognized. The robustness and generalizability of the proposed system is evaluated on a dataset including both positive images (with U.S. banknotes) and negative images (no U.S. banknotes) collected under a variety of conditions. The proposed algorithm, achieves 100% true recognition rate and 0% false recognition rate. Our banknote recognition system is also tested by blind users. Faiz M. Hasanuzzaman, Xiaodong Yang 0001, Yingli Tian |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2012 | Hierarchical Filtered Motion for Action Recognition in Crowded VideosabstractAction recognition with cluttered and moving background is a challenging problem. One main difficulty lies in the fact that the motion field in an action region is contaminated by the background motions. We propose a hierarchical filtered motion (HFM) method to recognize actions in crowded videos by the use of motion history image (MHI) as basic representations of motion because of its robustness and efficiency. First, we detect interest points as the two-dimensional Harris corners with recent motion, e.g., locations with high intensities in the MHI. Then, a global spatial motion smoothing filter is applied to the gradients of the MHI to eliminate isolated unreliable or noisy motions. At each interest point, a local motion field filter is applied to the smoothed gradients of the MHI by computing structure proximity between any pixel in the local region and the interest point. Thus, the motion at a pixel is enhanced or weakened based on its structure proximity with the interest point. To validate its effectiveness, we characterize the spatial and temporal features by histograms of oriented gradient in the intensity image and the MHI, respectively, and use a Gaussian-mixture-model-based classifier for action recognition. The performance of the proposed approach achieves the state-of-the-art results on the KTH dataset that has clean background. More importantly, we perform cross-dataset action classification and detection experiments, where the KTH dataset is used for training, while the microsoft research (MSR) action dataset II that consists of crowded videos with people moving in the background is used for testing. Our experiments show that the proposed HFM method significantly outperforms existing techniques. Yingli Tian, Liangliang Cao, Zicheng Liu 0001, Zhengyou Zhang |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2011 | Segment and recognize expression phase by fusion of motion area and neutral divergence featuresabstractAn expression can be approximated by a sequence of temporal segments called neutral, onset, offset and apex. However, it is not easy to accurately detect such temporal segments only based on facial features. Some researchers try to temporally segment expression phases with the help of body gesture analysis. The problem of this approach is that the expression temporal phases from face and gesture channels are not synchronized. Additionally, most previous work adopted facial key points tracking or body tracking to extract motion information, which is unreliable in practice due to illumination variations and occlusions. In this paper, we present a novel algorithm to overcome the above issues, in which two simple and robust features are designed to describe face and gesture information, i.e., motion area and neutral divergence features. Both features do not depend on motion tracking, and they can be easily calculated too. Moreover, it is different from previous work in that we integrate face and body gesture together in modeling the temporal dynamics through a single channel of sensorial source, so it avoids the unsynchronized issue between face and gesture channels. Extensive experimental results demonstrate the effectiveness of the proposed algorithm. Shizhi Chen, Yingli Tian, Qingshan Liu 0001, Dimitris N. Metaxas |
FG | 2 |
| 2011 | Text Detection in Natural Scene Images by Stroke Gabor WordsabstractIn this paper, we propose a novel algorithm, based on stroke components and descriptive Gabor filters, to detect text regions in natural scene images. Text characters and strings are constructed by stroke components as basic units. Gabor filters are used to describe and analyze the stroke components in text characters or strings. We define a suitability measurement to analyze the confidence of Gabor filters in describing stroke component and the suitability of Gabor filters on an image window. From the training set, we compute a set of Gabor filters that can describe principle stroke components of text by their parameters. Then a K-means algorithm is applied to cluster the descriptive Gabor filters. The clustering centers are defined as Stroke Gabor Words (SGWs) to provide a universal description of stroke components. By suitability evaluation on positive and negative training samples respectively, each SGW generates a pair of characteristic distributions of suitability measurements. On a testing natural scene image, heuristic layout analysis is applied first to extract candidate image windows. Then we compute the principle SGWs for each image window to describe its principle stroke components. Characteristic distributions generated by principle SGWs are used to classify text or non-text windows. Experimental results on benchmark datasets demonstrate that our algorithm can handle complex backgrounds and variant text patterns (font, color, scale, etc.). Chucai Yi, Yingli Tian |
ICDAR | 2 |
| 2011 | Recognizing clothes patterns for blind people by confidence margin based feature combinationabstractClothes pattern recognition is a challenging task for blind or visually impaired people. Automatic clothes pattern recognition is also a challenging problem in computer vision due to the large pattern variations. In this paper, we present a new method to classify clothes patterns into 4 categories: stripe, lattice, special, and patternless. While existing texture analysis methods mainly focused on textures varying with distinctive pattern changes, they cannot achieve the same level of accuracy for clothes pattern recognition because of the large intra-class variations in each clothes pattern category. To solve this problem, we extract both structural feature and statistical feature from image wavelet subbands. Furthermore, we develop a new feature combination scheme based on the confidence margin of a classifier to combine the two types of features to form a novel local image descriptor in a compact and discriminative format. The recognition experiment is conducted on a database with 627 clothes images of 4 categories of patterns. Experimental results demonstrate that the proposed method significantly outperforms the state-of-the-art texture analysis methods in the context of clothes pattern recognition. Xiaodong Yang 0001, Yingli Tian |
ACM Multimedia | 3 |
| 2011 | Text String Detection From Natural Scenes by Structure-Based Partition and GroupingabstractText information in natural scene images serves as important clues for many image-based applications such as scene understanding, content-based image retrieval, assistive navigation, and automatic geocoding. However, locating text from a complex background with multiple colors is a challenging task. In this paper, we explore a new framework to detect text strings with arbitrary orientations in complex natural scene images. Our proposed framework of text string detection consists of two steps: 1) image partition to find text character candidates based on local gradient features and color uniformity of character components and 2) character candidate grouping to detect text strings based on joint structural features of text characters in each text string such as character size differences, distances between neighboring characters, and character alignment. By assuming that a text string has at least three characters, we propose two algorithms of text string detection: 1) adjacent character grouping method and 2) text line grouping method. The adjacent character grouping method calculates the sibling groups of each character candidate as string segments and then merges the intersecting sibling groups into text string. The text line grouping method performs Hough transform to fit text line among the centroids of text candidates. Each fitted text line describes the orientation of a potential text string. The detected text string is presented by a rectangle region covering all characters whose centroids are cascaded in its text line. To improve efficiency and accuracy, our algorithms are carried out in multi-scales. The proposed methods outperform the state-of-the-art results on the public Robust Reading Dataset, which contains text only in horizontal orientation. Furthermore, the effectiveness of our methods to detect text strings with arbitrary orientations is evaluated on the Oriented Scene Text Dataset collected by ourselves containing text strings in nonhorizontal orientations. Chucai Yi, Yingli Tian |
IEEE Trans. Image Process. | 2 |
| 2011 | Robust Detection of Abandoned and Removed Objects in Complex Surveillance VideosabstractTracking-based approaches for abandoned object detection often become unreliable in complex surveillance videos due to occlusions, lighting changes, and other factors. We present a new framework to robustly and efficiently detect abandoned and removed objects based on background subtraction (BGS) and foreground analysis with complement of tracking to reduce false positives. In our system, the background is modeled by three Gaussian mixtures. In order to handle complex situations, several improvements are implemented for shadow removal, quick-lighting change adaptation, fragment reduction, and keeping a stable update rate for video streams with different frame rates. Then, the same Gaussian mixture models used for BGS are employed to detect static foreground regions without extra computation cost. Furthermore, the types of the static regions (abandoned or removed) are determined by using a method that exploits context information about the foreground masks, which significantly outperforms previous edge-based techniques. Based on the type of the static regions and user-defined parameters (e.g., object size and abandoned time), a matching method is proposed to detect abandoned and removed objects. A person-detection process is also integrated to distinguish static objects from stationary people. The robustness and efficiency of the proposed method is tested on IBM Smart Surveillance Solutions for public safety applications in big cities and evaluated by several public databases, such as The Image library for intelligent detection systems (i-LIDS) and IEEE Performance Evaluation of Tracking and Surveillance Workshop (PETS) 2006 datasets. The test and evaluation demonstrate our method is efficient to run in real-time, while being robust to quick-lighting changes and occlusions in complex environments. Yingli Tian, Rogério Feris, Arun Hampapur, Ming-Ting Sun |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2010 | Clothes Matching for Blind and Color Blind People
Yingli Tian |
ICCHP (2) | 1 |
| 2010 | Improving Computer Vision-Based Indoor Wayfinding for Blind Persons with Context Information
Yingli Tian, Chucai Yi, Aries Arditi |
ICCHP (2) | 1 |
| 2010 | Computer Vision-Based Door Detection for Accessibility of Unfamiliar Environments to Blind Persons
Yingli Tian, Xiaodong Yang 0001, Aries Arditi |
ICCHP (2) | 1 |
| 2010 | Action detection using multiple spatial-temporal interest point featuresabstractThis paper considers the problem of detecting actions from cluttered videos. Compared with the classical action recognition problem, this paper aims to estimate not only the scene category of a given video sequence, but also the spatial-temporal locations of the action instances. In recent years, many feature extraction schemes have been designed to describe various aspects of actions. However, due to the difficulty of action detection, e.g., the cluttered background and potential occlusions, a single type of features cannot solve the action detection problems perfectly in cluttered videos. In this paper, we attack the detection problem by combining multiple Spatial-Temporal Interest Point (STIP) features, which detect salient patches in the video domain, and describe these patches by feature of local regions. The difficulty of combining multiple STIP features for action detection is two folds: First, the number of salient patches detected by different STIP methods varies across different salient patches. How to combine such features is not considered by existing fusion methods. Second, the detection in the videos should be efficient, which excludes many slow machine learning algorithms. To handle these two difficulties, we propose a new approach which combines Gaussian Mixture Model with Branch-and-Bound search to efficiently locate the action of interest. We build a new challenging dataset for our action detection task, and our algorithm obtains impressive results. On classical KTH dataset, our method outperforms the state-of-the-art methods. Liangliang Cao, Yingli Tian, Zicheng Liu 0001, Benjamin Z. Yao, Zhengyou Zhang, Thomas S. Huang |
ICME | 2 |
| 2010 | Context-based indoor object detection as an aid to blind persons accessing unfamiliar environmentsabstractIndependent travel is a well known challenge for blind or visually impaired persons. In this paper, we propose a computer vision-based indoor wayfinding system for assisting blind people to independently access unfamiliar buildings. In order to find different rooms (i.e. an office, a lab, or a bathroom) and other building amenities (i.e. an exit or an elevator), we incorporate door detection with text recognition. First we develop a robust and efficient algorithm to detect doors and elevators based on general geometric shape, by combining edges and corners. The algorithm is generic enough to handle large intra-class variations of the object model among different indoor environments, as well as small inter-class differences between different objects such as doors and elevators. Next, to distinguish an office door from a bathroom door, we extract and recognize the text information associated with the detected objects. We first extract text regions from indoor signs with multiple colors. Then text character localization and layout analysis of text strings are applied to filter out background interference. The extracted text is recognized by using off-the-shelf optical character recognition (OCR) software products. The object type, orientation, and location can be displayed as speech for blind travelers. Xiaodong Yang 0001, Yingli Tian, Chucai Yi, Aries Arditi |
ACM Multimedia | 2 |
| 2008 | Facial image analysis using local feature adaptation prior to learningabstractMany facial image analysis methods rely on learning-based techniques such as Adaboost or SVMs to project classifiers based on the selection of local image filters (e.g., Haar and Gabor filters) from large sets of training data. In general, the learning process consists of selecting discriminative image filters from a large feature pool that contains filters uniformly sampled from the parameter space. In this paper, we argue that we are able to improve these methods by incorporating a local feature adaptation technique prior to learning, which generates a more compact and meaningful pool of image filters, consequently reducing both learning and detection/recognition computational costs, while at the same time improving accuracies. In the first stage of our approach, local feature adaptation is carried out by a nonlinear optimization method that determines image filter parameters (such as position, orientation and scale) in order to match the geometrical structure of each training sample. In the second stage, Adaboost feature selection technique is applied to the adapted feature pool to obtain the final set of discriminative local image filters. We demonstrate the effectiveness and efficiency of the proposed framework in the face detection domain. In the experiments, we have applied our method using a pool of wavelet features, including Haar and Gabor filters. The results showed that with local feature adaptation, significant improvements in terms of detection accuracy and computational cost reduction are achieved over learning based on the same features sampled uniformly from the parameter space. Rogério Feris, Yingli Tian, Yun Zhai, Arun Hampapur |
FG | 2 |
| 2008 | IBM smart surveillance system (S3): event based video surveillance system with an open and extensible framework
Yingli Tian, Lisa M. Brown, Arun Hampapur, Max Lu, Andrew W. Senior, Chiao-Fe Shu |
Mach. Vis. Appl. | 1 |
| 2007 | Searching surveillance videoabstractSurveillance video is used in two key modes, watching for known threats in real-time and searching for events of interest after the fact. Typically, real-time alerting is a localized function, e.g. airport security center receives and reacts to a “perimeter breach alert”, while investigations often tend to encompass a large number of geographically distributed cameras like the London bombing, or Washington sniper incidents. Enabling effective search of surveillance video for investigation & preemption, involves indexing the video along multiple dimensions. This paper presents a framework for surveillance search which includes, video parsing, indexing and query mechanisms. It explores video parsing techniques which automatically extract index data from video, indexing which stores data in relational tables, retrieval which uses SQL queries to retrieve events of interest and the software architecture that integrates these technologies. Arun Hampapur, Lisa M. Brown, Rogério Feris, Andrew W. Senior, Chiao-Fe Shu, Yingli Tian, Yun Zhai, Max Lu |
AVSS | 6 |
| 2007 | Video analytics for retailabstractWe describe a set of tools for retail analytics based on a combination of video understanding and transaction-log. Tools are provided for loss prevention (returns fraud and cashier fraud), store operations (customer counting) and merchandising (display effectiveness). Results are presented on returns fraud and customer counting. Andrew W. Senior, Lisa M. Brown, Arun Hampapur, Chiao-Fe Shu, Yun Zhai, Rogério Feris, Yingli Tian, Sergio Borger, Christopher R. Carlson |
AVSS | 7 |
| 2007 | Capturing People in Surveillance VideoabstractThis paper presents reliable techniques for detecting, tracking, and storing keyframes of people in surveillance video. The first component of our system is a novel face detector algorithm, which is based on first learning local adaptive features for each training image, and then using Adaboost learning to select the most general features for detection. This method provides a powerful mechanism for combining multiple features, allowing faster training time and better detection rates. The second component is a face tracking algorithm that interleaves multiple view-based classifiers along the temporal domain in a video sequence. This interleaving technique, combined with a correlation-based tracker, enables fast and robust face tracking over time. Finally, the third component of our system is a keyframe selection method that combines a person classifier with a face classifier. The basic idea is to generate a person keyframe in case the face is not visible, in order to reduce the number of false negatives. We performed quantitatively evaluation of our techniques on standard datasets and on surveillance videos captured by a camera over several days. Rogério Feris, Yingli Tian, Arun Hampapur |
CVPR | 2 |
| 2007 | An end-to-end eChronicling System for Mobile Human SurveillanceabstractRapid advances in mobile computing devices and sensor technologies are enabling the capture of unprecedented volumes of data by individuals involved in field operations in a variety of applications. As capture becomes ever more rich and pervasive the biggest challenge is in developing information processing and representation tools that maximize the utility of the captured multi-sensory data. The right tools hold the promise of converting captured data into actionable intelligence resulting in improved memory, enhanced situational understanding, and more efficient execution of operations. These tools need to be at least as rich and diverse as the sensors used for capture, and need to be unified within an effective system architecture. This paper presents our initial attempt at such a system and architecture that combines several emerging sensor technologies, state of the art analytic engines, and multi-dimensional navigation tools, into an end-to-end electronic chronicling solution for mobile surveillance by humans. Gopal Sarma Pingali, Yingli Tian, Shahram Ebadollahi, Jason W. Pelecanos, Mark Podlaseck, Harry Stavropoulos |
CVPR | 2 |
| 2007 | S3: The IBM Smart Surveillance System: From Transactional Systems to Observational SystemsabstractPervasive sensor based systems are transforming Information Technology systems from being transactional in nature to being observational in nature. Observational systems are inherently distributed and capture information at a much finer grain of space and time. Enabling and building such systems also poses many technology challenges, extracting information from sensor signals, indexing and searching sensor meta-data, data mining and scalability. In this paper we use S3: the IBM smart surveillance system as an example of an observational system to explore several of these issues through real world deployment examples. Arun Hampapur, Sergio Borger, Lisa M. Brown, Christopher R. Carlson, Jonathan H. Connell, Max Lu, Andrew W. Senior, V. Reddy, Chiao-Fe Shu, Yingli Tian |
ICASSP (4) | 10 |
| 2007 | Answering relationship queries on the webabstractFinding relationships between entities on the Web, e.g., the connections between different places or the commonalities of people, is a novel and challenging problem. Existing Web search engines excel in keyword matching and document ranking, but they cannot well handle many relationship queries. This paper proposes a new method for answering relationship queries on two entities. Our method first respectively retrieves the top Web pages for either entity from a Web search engine. It then matches these Web pages and generates an ordered list of Web page pairs. Each Web page pair consists of one Web page for either entity. The top ranked Web page pairs are likely to contain the relationships between the two entities. One main challenge in the ranking process is to effectively filter out the large amount of noise in the Web pages without losing much useful information. To achieve this, our method assigns appropriate weights to terms in Web pages and intelligently identifies the potential connecting terms that capture the relationships between the two entities. Only those top potential connecting terms with large weights are used to rank Web page pairs. Finally, the top ranked Web page pairs are presented to the searcher. For each such pair, the query terms and the top potential connecting terms are properly highlighted so that the relationships between the two entities can be easily identified. We implemented a prototype on top of the Google search engine and evaluated it under a wide variety of query scenarios. The experimental results show that our method is effective at finding important relationships with low overhead. Gang Luo 0001, Chunqiang Tang, Yingli Tian |
WWW | 3 |
| 2006 | Automatic Counting of Interacting People by using a Single Uncalibrated CameraabstractAutomatic counting of people, entering or exiting a region of interest, is very important for both business and security applications. This paper introduces an automatic and robust people counting system which can count multiple people who interact in the region of interest, by using only one camera. Two-level hierarchical tracking is employed. For cases not involving merges or splits, a fast blob tracking method is used. In order to deal with interactions among people in a more thorough and reliable way, the system uses the mean shift tracking algorithm. Using the first-level blob tracker in general, and employing the mean shift tracking only in the case of merges and splits saves power and makes the system computationally efficient. The system setup parameter can be automatically learned in a new environment from a 3 to 5 minute-video with people going in or out of the target region one at a time. With a 2 GHz Pentium machine, the system runs at about 33 fps on 320times240 images without code optimization. Average accuracy rates of 98.5% and 95% are achieved on videos with normal traffic flow and videos with many cases of merges and splits, respectively Senem Velipasalar, Yingli Tian, Arun Hampapur |
ICME | 2 |
| 2006 | Appearance models for occlusion handling
Andrew W. Senior, Arun Hampapur, Yingli Tian, Lisa M. Brown, Sharath Pankanti, Ruud M. Bolle |
Image Vis. Comput. | 3 |
| 2005 | IBM smart surveillance system (S3): a open and extensible framework for event based surveillanceabstractAs smart surveillance technology becomes a critical component in security infrastructures, the system architecture assumes a critical importance. This paper considers the example of smart surveillance in an airport environment. We start with a threat model for airports and use this to derive the security requirements. These requirements are used to motivate an open-standards based architecture for surveillance. We discuss the critical aspects of this architecture and its implementation in the IBM S3 smart surveillance system. Demo results from a pilot deployment in Hawthorne, NY are presented. Chiao-Fe Shu, Arun Hampapur, Max Lu, Lisa M. Brown, Jonathan H. Connell, Andrew W. Senior, Yingli Tian |
AVSS | 7 |
| 2005 | Robust and Efficient Foreground Analysis for Real-Time Video SurveillanceabstractWe present a new method to robustly and efficiently analyze foreground when we detect background for a fixed camera view by using mixture of Gaussians models and multiple cues. The background is modeled by three Gaussian mixtures as in the work of Stauffer and Grimson (1999). Then the intensity and texture information are integrated to remove shadows and to enable the algorithm working for quick lighting changes. For foreground analysis, the same Gaussian mixture model is employed to detect the static foreground regions without using any tracking or motion information. Then the whole static regions are pushed back to the background model to avoid a common problem in background subtraction /spl times/ fragmentation (one object becomes multiple parts). The method was tested on our real time video surveillance system. It is robust and run about 130 fpsfor color images and 150 fps for grayscale images at size 160/spl times/120 on a 2GB Pentium IV machine with MMX optimization. Yingli Tian, Max Lu, Arun Hampapur |
CVPR (1) | 1 |
| 2004 | Detection and tracking in the IBM PeopleVision systemabstractThe detection and tracking of people lie at the heart of many current and near-future applications of computer vision. We describe a background subtraction system designed to detect moving objects in a wide variety of conditions, and a second system to detect objects moving in front of moving backgrounds. Detected foreground regions are tracked with a tracking system which can initiate real-time alarms and generate a smart surveillance index which can be searched to find interesting events in stored video Jonathan H. Connell, Andrew W. Senior, Arun Hampapur, Yingli Tian, Lisa M. Brown, Sharath Pankanti |
ICME | 4 |
| 2003 | Face Cataloger: Multi-Scale Imaging for Relating Identity to LocationabstractThe level of security at a facility is directly related to how well the facility can keep track of "who is where". The "who" part of this question is typically addressed through the use of face images for recognition either by a person or a computer face recognition system. The "where" part of this question can be addressed through 3D position tracking. The "who is where" problem is inherently multi-scale, wide angle views are needed for location estimation and high resolution face images for identification. A number of other people tracking challenges like activity understanding are multiscale in nature. An effective system to answer "who is where?" must acquire face images without constraining the users and must closely associate the face images with the 3D path of the person. Our solution to this problem uses computer controlled pan-tilt-zoom cameras driven by a 3D wide-baseline stereo tracking system. The pan-tilt-zoom cameras automatically acquire zoomed-in views of a person's head, while the person is in motion within the monitored space. Arun Hampapur, Sharath Pankanti, Andrew W. Senior, Yingli Tian, Lisa M. Brown, Ruud M. Bolle |
AVSS | 4 |
| 2001 | Recognizing Action Units for Facial Expression AnalysisabstractMost automatic expression analysis systems attempt to recognize a small set of prototypic expressions, such as happiness, anger, surprise, and fear. Such prototypic expressions, however, occur rather infrequently. Human emotions and intentions are more often communicated by changes in one or a few discrete facial features. In this paper, we develop an Automatic Face Analysis (AFA) system to analyze facial expressions based on both permanent facial features (brows, eyes, mouth) and transient facial features (deepening of facial furrows) in a nearly frontal-view face image sequence. The AFA system recognizes fine-grained changes in facial expression into action units (AUs) of the Facial Action Coding System (FACS), instead of a few prototypic expressions. Multistate face and facial component models are proposed for tracking and modeling the various facial features, including lips, eyes, brows, cheeks, and furrows. During tracking, detailed parametric descriptions of the facial features are extracted. With these parameters as the inputs, a group of action units (neutral expression, six upper face AUs and 10 lower face AUs) are recognized whether they occur alone or in combinations. The system has achieved average recognition rates of 96.4 percent (95.4 percent if neutral expressions are excluded) for upper face AUs and 96.7 percent (95.6 percent with neutral expressions excluded) for lower face AUs. The generalizability of the system has been tested by using independent image databases collected and FACS-coded for ground-truth by different research teams. Yingli Tian, Takeo Kanade, Jeffrey F. Cohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Recognizing Upper Face Action Units for Facial Expression AnalysisabstractWe develop an automatic system to analyze subtle changes in upper face expressions based on both permanent facial features (brows, eyes, mouth) and transient facial features (deepening of facial furrows) in a nearly frontal image sequence. Our system recognizes fine-grained changes in facial expression based on Facial Action Coding System (FACS) action units (AUs). Multi-state facial component models are proposed for tracting and modeling different facial features, including eyes, brews, cheeks, and furrows. Then we convert the results of tracking to detailed parametric descriptions of the facial features. These feature parameters are fed to a neural network which recognizes 7 upper face action units. A recognition rate of 95% is obtained for the test data that include both single action units and AU combinations. Yingli Tian, Takeo Kanade, Jeffrey F. Cohn |
CVPR | 1 |
| 2000 | Comprehensive Database for Facial Expression AnalysisabstractWithin the past decade, significant effort has occurred in developing methods of facial expression analysis. Because most investigators have used relatively limited data sets, the generalizability of these various methods remains unknown. We describe the problem space for facial expression analysis, which includes level of description, transitions among expressions, eliciting conditions, reliability and validity of training and test data, individual differences in subjects, head orientation and scene complexity image characteristics, and relation to non-verbal behavior. We then present the CMU-Pittsburgh AU-Coded Face Expression Image Database, which currently includes 2105 digitized image sequences from 182 adult subjects of varying ethnicity, performing multiple tokens of most primary FACS action units. This database is the most comprehensive testbed to date for comparative studies of facial expression analysis. Takeo Kanade, Yingli Tian, Jeffrey F. Cohn |
FG | 2 |
| 2000 | Dual-State Parametric Eye TrackingabstractMost eye trackers work well for open eyes. However blinking is a physiological necessity for humans. More over, for applications such as facial expression analysis and driver awareness systems, we need to do more than tracking of the locations of the person's eyes but obtain their detailed description. We need to recover the state of the eyes (i.e., whether they are open or closed), and the parameters of an eye model (e.g., the location and radius of the iris, and the corners and height of the eye opening). We develop a dual-state model-based system for tracking eye features that uses convergent tracking techniques and show how it can be used to detect whether the eyes are open or closed, and to recover the parameters of the eye model. Processing speed on a Pentium II 400 MHz PC is approximately 3 frames/second. In experimental tests on 500 image sequences from child and adult subjects with varying colors of skin and eye, accurate tracking results are obtained in 98% of image sequences. Yingli Tian, Takeo Kanade, Jeffrey F. Cohn |
FG | 1 |
| 2000 | Recognizing Lower Face Action Units for Facial Expression AnalysisabstractMost automatic expression analysis systems attempt to recognize a small set of prototypic expressions (e.g., happiness and anger). Such prototypic expressions, however, occur infrequently. Human emotions and intentions are communicated more often by changes in one or two discrete facial features. We develop an automatic system to analyze subtle changes in facial expressions based on both permanent (e.g., mouth, eye, and brow) and transient (e.g., furrows and wrinkles) facial features in a nearly frontal image sequence. Multi-state facial component models are proposed for tracking and modeling different facial features. Based on these multi-state models, and without artificial enhancement, we detect and track the facial features, including mouth, eyes, brow, cheeks, and their related wrinkles and facial furrows. Moreover we recover detailed parametric descriptions of the facial features. With these features as the inputs, 11 individual action units or action unit combinations are recognized by a neural network algorithm. A recognition rate of 96.7% is obtained. The recognition results indicate that our system can identify action units regardless of whether they occur singly or in combinations. Yingli Tian, Takeo Kanade, Jeffrey F. Cohn |
FG | 1 |
| 2000 | Eye-State Action Unit Detection by Gabor Wavelets
Yingli Tian, Takeo Kanade, Jeffrey F. Cohn |
ICMI | 1 |
| 1998 | Shape Recovery from One Image under Multiple Light Sources
Yingli Tian, Hung-Tat Tsui, S. Y. Yeung |
ACCV (1) | 1 |
| 1998 | Spherical and Cylindrical Light Source Models for Shape Recovery
Yingli Tian, Hung-Tat Tsui, S. Y. Yeung |
ACCV (1) | 1 |
| 1996 | Shape from shading for non-Lambertian surfaces from one color imageabstractIn this paper a robust approach of shape recovery from one color image for non-Lambertian surfaces is presented. A dichromatic reflection model is assumed to describe the color hybrid surface. An extended light source model is used so that both specular and diffuse reflection components can be detected simultaneously in one image. The color of the light source is assumed different from that of the object. This is a very mild restriction which can almost be satisfied. Using color information, regions of the image with hybrid reflection can be segmented from regions with only diffuse reflection. A new reflectance map for non-Lambertian surface is developed so that surface shape can be inferred using the shape from shading technique. For regions with only diffuse reflection, the conventional reflectance map for Lambertian surface can be used. Some results based on real experiments and simulations are presented to show the validity of our method. Yingli Tian, Hung-Tat Tsui |
ICPR | 1 |
| 1996 | 3D shape recovery from two-color image sequences using a genetic algorithmabstractWe propose a new method to recover the 3D shape and surface reflectance of objects by color image analysis. We introduce a new lighting system with an extended light source moving in a plane which contains the optical axis of the camera. An image sequence is acquired with each image taken at a different light source position. A dichromatic reflection model is assumed The diffuse and specular reflection components are separated by a pseudoinverse technique. A fast generic algorithm is then used for the shape and reflectance recovery from the disuse reflection component extracted from a pair of color image sequences. As a result, the normal of a surface patch nor in the light source moving plane can also be estimated robustly. Our method is verified by both real and simulated experiments. Yingli Tian, Hung-Tat Tsui |
ICPR | 1 |