EDBT 2026 Demo / reviewers in the wild / expert
Matti Pietikäinen
dblp:95/4955
· DBLP profile ↗
177ranked-venue papers
8as first author
24since 2021 · last 2025
0000-0003-2263-6731ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 136 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 91 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-authorHuman-computer interaction and ubiquitous computing · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Security and privacy · 3Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rapid Salient Object Detection With Difference Convolutional Neural NetworksabstractThis paper addresses the challenge of deploying salient object detection (SOD) on resource-constrained devices with real-time performance. While recent advances in deep neural networks have improved SOD, existing top-leading models are computationally expensive. We propose an efficient network design that combines traditional wisdom on SOD and the representation power of modern CNNs. Like biologically-inspired classical SOD methods relying on computing contrast cues to determine saliency of image regions, our model leverages Pixel Difference Convolutions (PDCs) to encode the feature contrasts. Differently, PDCs are incorporated in a CNN architecture so that the valuable contrast cues are extracted from rich feature maps. For efficiency, we introduce a difference convolution reparameterization (DCR) strategy that embeds PDCs into standard convolutions, eliminating computation and parameters at inference. Additionally, we introduce SpatioTemporal Difference Convolution (STDC) for video SOD, enhancing the standard 3D convolution with spatiotemporal contrast capture. Our models, SDNet for image SOD and STDNet for video SOD, achieve significant improvements in efficiency-accuracy trade-offs. On a Jetson Orin device, our models with $< $< 1M parameters operate at 46 FPS and 150 FPS on streamed images and videos, surpassing the second-best lightweight models in our experiments by more than $2\times$2× and $3\times$3× in speed with superior accuracy. Zhuo Su 0002, Li Liu 0002, Matthias Müller 0011, Diana Wofk, Ming-Ming Cheng, Matti Pietikäinen |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Few-Shot Class-Incremental Learning for Classification and Object Detection: A SurveyabstractFew-shot Class-Incremental Learning (FSCIL) presents a unique challenge in Machine Learning (ML), as it necessitates the Incremental Learning (IL) of new classes from sparsely labeled training samples without forgetting previous knowledge. While this field has seen recent progress, it remains an active exploration area. This paper aims to provide a comprehensive and systematic review of FSCIL. In our in-depth examination, we delve into various facets of FSCIL, encompassing the problem definition, the discussion of the primary challenges of unreliable empirical risk minimization and the stability-plasticity dilemma, general schemes, and relevant problems of IL and Few-shot Learning (FSL). Besides, we offer an overview of benchmark datasets and evaluation metrics. Furthermore, we introduce the Few-shot Class-incremental Classification (FSCIC) methods from data-based, structure-based, and optimization-based approaches and the Few-shot Class-incremental Object Detection (FSCIOD) methods from anchor-free and anchor-based approaches. Beyond these, we present several promising research directions within FSCIL that merit further investigation. Li Liu 0002, Olli Silvén, Matti Pietikäinen, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Advancing Segment Anything Model for Efficient Salient Object Detection in Remote Sensing ImagesabstractSalient object detection in optical remote sensing images (ORSI-SOD) often relies on leveraging pre-trained knowledge from natural images to achieve high accuracy with limited training data. Traditional methods typically employ vision backbones (e.g., Convolutional Neural Networks (CNNs) or Vision Transformers (ViTs)) pre-trained on ImageNet to extract features from ORSI scenes. However, these backbones exhibit limited generalization across diverse scenarios compared to recent vision foundation models. To this end, we propose ORSI-SAM, a novel ORSI-SOD framework based on the Segment Anything Model (SAM), leveraging its superior generalization capabilities to achieve an exceptional efficiency-accuracy trade-off. Specifically, ORSI-SAM adopts lightweight SAM as the backbone, effectively reducing parameter size and computational overhead to enable efficient deployment on satellite devices while retaining the rich knowledge learned from large-scale natural image datasets. To mitigate the impact of unavailable prompts in ORSI-SOD on the prediction capability of the SAM decoder, we introduce a Hierarchical Interaction Prompt Generator (HIPG), which aggregates hierarchical features and generates mask prompts tailored for salient objects to guide the decoder in producing high-quality saliency maps. Furthermore, to address the recognition challenges caused by the inherent characteristics of ORSIs, we propose a Semantic-Aware Refinement Decoder (SARD). SARD integrates structural details from low-level features to enrich fine-grained object information while leveraging high-level features to suppress redundant interference in shallow layers, thereby improving the detailed information in the predicted saliency map. ORSI-SAM is the first work to explore the accuracy-efficiency trade-offs for ORSI-SOD based on SAM architecture. Extensive experiments on benchmark datasets show that ORSI-SAM achieves superior performance compared to recent state-of-the-art methods with 12.2M parameters and 8.9G FLOPs. Li Liu 0002, Zhuo Su 0002, Tianpeng Liu, Zhen Liu 0004, Matti Pietikäinen |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Boosting Convolutional Neural Networks With Middle Spectrum Grouped ConvolutionabstractThis article proposes a novel module called middle spectrum grouped convolution (MSGC) for efficient deep convolutional neural networks (DCNNs) with the mechanism of grouped convolution. It explores the broad "middle spectrum" area between channel pruning and conventional grouped convolution. Compared with channel pruning, MSGC can retain most of the information from the input feature maps due to the group mechanism; compared with grouped convolution, MSGC benefits from the learnability, the core of channel pruning, for constructing its group topology, leading to better channel division. The middle spectrum area is unfolded along four dimensions: groupwise, layerwise, samplewise, and attentionwise, making it possible to reveal more powerful and interpretable structures. As a result, the proposed module acts as a booster that can reduce the computational cost of the host backbones for general image recognition with even improved predictive accuracy. For example, in the experiments on the ImageNet dataset for image classification, MSGC can reduce the multiply-accumulates (MACs) of ResNet-18 and ResNet-50 by half but still increase the Top-1 accuracy by more than 1%. With a 35% reduction of MACs, MSGC can also increase the Top-1 accuracy of the MobileNetV2 backbone. Results on the MS COCO dataset for object detection show similar observations. Our code and trained models are available at https://github.com/hellozhuo/msgc. Zhuo Su 0002, Tianpeng Liu, Zhen Liu 0004, Shuanghui Zhang, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Editorial: Learning With Fewer Labels in Computer VisionabstractUndoubtedly, Deep Neural Networks (DNNs), from AlexNet to ResNet to Transformer, have sparked revolutionary advancements in diverse computer vision tasks. The scale of DNNs has grown exponentially due to the rapid development of computational resources. Despite the tremendous success, DNNs typically depend on massive amounts of training data (especially the recent various foundation models) to achieve high performance and are brittle in that their performance can degrade severely with small changes in their operating environment. Generally, collecting massive-scale training datasets is costly or even infeasible, as for certain fields, only very limited or no examples at all can be gathered. Nevertheless, collecting, labeling, and vetting massive amounts of practical training data is certainly difficult and expensive, as it requires the painstaking efforts of experienced human annotators or experts, and in many cases, prohibitively costly or impossible due to some reason, such as privacy, safety or ethic issues. Li Liu 0002, Timothy M. Hospedales, Yann LeCun, Mingsheng Long, Jiebo Luo 0001, Wanli Ouyang, Matti Pietikäinen, Tinne Tuytelaars |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Deep Learning for Visual Speech Analysis: A SurveyabstractVisual speech, referring to the visual domain of speech, has attracted increasing attention due to its wide applications, such as public security, medical treatment, military defense, and film entertainment. As a powerful AI strategy, deep learning techniques have extensively promoted the development of visual speech learning. Over the past five years, numerous deep learning based methods have been proposed to address various problems in this area, especially automatic visual speech recognition and generation. To push forward future research on visual speech, this paper will present a comprehensive review of recent progress in deep learning methods on visual speech analysis. We cover different aspects of visual speech, including fundamental problems, challenges, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. Besides, we also identify gaps in current research and discuss inspiring future research directions. Changchong Sheng, Gangyao Kuang, Liang Bai 0003, Chenping Hou, Yulan Guo, Xin Xu 0001, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Highly Efficient and Unsupervised Framework for Moving Object Detection in Satellite VideosabstractMoving object detection in satellite videos (SVMOD) is a challenging task due to the extremely dim and small target characteristics. Current learning-based methods extract spatio-temporal information from multi-frame dense representation with labor-intensive manual labels to tackle SVMOD, which needs high annotation costs and contains tremendous computational redundancy due to the severe imbalance between foreground and background regions. In this paper, we propose a highly efficient unsupervised framework for SVMOD. Specifically, we propose a generic unsupervised framework for SVMOD, in which pseudo labels generated by a traditional method can evolve with the training process to promote detection performance. Furthermore, we propose a highly efficient and effective sparse convolutional anchor-free detection network by sampling the dense multi-frame image form into a sparse spatio-temporal point cloud representation and skipping the redundant computation on background regions. Coping these two designs, we can achieve both high efficiency (label and computation efficiency) and effectiveness. Extensive experiments demonstrate that our method can not only process 98.8 frames per second on 1024 ×1024 images but also achieve state-of-the-art performance. Wei An 0003, Yifan Zhang 0030, Zhuo Su 0002, Weidong Sheng, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Lightweight Pixel Difference Networks for Efficient Visual Representation LearningabstractRecently, there have been tremendous efforts in developing lightweight Deep Neural Networks (DNNs) with satisfactory accuracy, which can enable the ubiquitous deployment of DNNs in edge devices. The core challenge of developing compact and efficient DNNs lies in how to balance the competing goals of achieving high accuracy and high efficiency. In this paper we propose two novel types of convolutions, dubbed Pixel Difference Convolution (PDC) and Binary PDC (Bi-PDC) which enjoy the following benefits: capturing higher-order local differential information, computationally efficient, and able to be integrated with existing DNNs. With PDC and Bi-PDC, we further present two lightweight deep networks named Pixel Difference Networks (PiDiNet) and Binary PiDiNet (Bi-PiDiNet) respectively to learn highly efficient yet more accurate representations for visual tasks including edge detection and object recognition. Extensive experiments on popular datasets (BSDS500, ImageNet, LFW, YTF, etc.) show that PiDiNet and Bi-PiDiNet achieve the best accuracy-efficiency trade-off. For edge detection, PiDiNet is the first network that can be trained without ImageNet, and can achieve the human-level performance on BSDS500 at 100 FPS and with 1 M parameters. For object recognition, among existing Binary DNNs, Bi-PiDiNet achieves the best accuracy and a nearly 2× reduction of computational cost on ResNet18. Zhuo Su 0002, Longguang Wang, Hua Zhang 0008, Zhen Liu 0004, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Affective Computing [Scanning the Issue]abstractThe articles in this special issue cover four major subfields in affective computing, namely affect analysis, affect synthesis, applications, and ethics. Björn W. Schuller, Matti Pietikäinen |
Proc. IEEE | 2 |
| 2023 | Facial Micro-Expressions: An OverviewabstractMicro-expression (ME) is an involuntary, fleeting, and subtle facial expression. It may occur in high-stake situations when people attempt to conceal or suppress their true feelings. Therefore, MEs can provide essential clues to people’s true feelings and have plenty of potential applications, such as national security, clinical diagnosis, and interrogations. In recent years, ME analysis has gained much attention in various fields due to its practical importance, especially automatic ME analysis in computer vision as MEs are difficult to process by naked eyes. In this survey, we provide a comprehensive review of ME development in the field of computer vision, from the ME studies in psychology and early attempts in computer vision to various computational ME analysis methods and future directions. Four main tasks in ME analysis are specifically discussed, including ME spotting, ME recognition, ME action unit detection, and ME generation in terms of the approaches, advance developments, and challenges. Through this survey, readers can understand MEs in both aspects of psychology and computer vision, and apprehend the future research direction in ME analysis. Guoying Zhao 0001, Yante Li, Matti Pietikäinen |
Proc. IEEE | 4 |
| 2023 | Uncertainty-Guided Semi-Supervised Few-Shot Class-Incremental Learning With Knowledge DistillationabstractClass-Incremental Learning (CIL) aims at incrementally learning novel classes without forgetting old ones. This capability becomes more challenging when novel tasks contain one or a few labeled training samples, which leads to a more practical learning scenario,i.e., Few-Shot Class- Incremental Learning (FSCIL). The dilemma on FSCIL lies in serious overfitting and exacerbated catastrophic forgetting caused by the limited training data from novel classes. In this paper, excited by the easy accessibility of unlabeled data, we conduct a pioneering work and focus on a Semi-Supervised Few-Shot Class-Incremental Learning (Semi-FSCIL) problem, which requires the model incrementally to learn new classes from extremely limited labeled samples and a large number of unlabeled samples. To address this problem, a simple but efficient framework is first constructed based on the knowledge distillation technique to alleviate catastrophic forgetting. To efficiently mitigate the overfitting problem on novel categories with unlabeled data, uncertainty-guided semi-supervised learning is incorporated into this framework to select unlabeled samples into incremental learning sessions considering the model uncertainty. This process provides extra reliable supervision for the distillation process and contributes to better formulating the class means. Our extensive experiments on CIFAR100, miniImageNet and CUB200 datasets demonstrate the promising performance of our proposed method, and define baselines in this new research direction. Yawen Cui, Wanxia Deng, Xin Xu 0001, Zhen Liu 0004, Zhong Liu 0002, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Multim. | 6 |
| 2023 | Importance-Aware Information Bottleneck Learning Paradigm for Lip ReadingabstractLip reading is the task of decoding text from speakers' mouth movements. Numerous deep learning-based methods have been proposed to address this task. However, these existing deep lip reading models suffer from poor generalization due to overfitting the training data. To resolve this issue, we present a novel learning paradigm that aims to improve the interpretability and generalization of lip reading models. In specific, a Variational Temporal Mask (VTM) module is customized to automatically analyze the importance of frame-level features. Furthermore, the prediction consistency constraints of global information and local temporal important features are introduced to strengthen the model generalization. We evaluate the novel learning paradigm with multiple lip reading baseline models on the LRW and LRW-1000 datasets. Experiments show that the proposed framework significantly improves the generalization performance and interpretability of lip reading models. Changchong Sheng, Li Liu 0002, Wanxia Deng, Liang Bai 0003, Zhong Liu 0002, Songyang Lao, Gangyao Kuang, Matti Pietikäinen |
IEEE Trans. Multim. | 8 |
| 2022 | SVNet: Where SO(3) Equivariance Meets Binarization on Point Cloud RepresentationabstractEfficiency and robustness are increasingly needed for applications on 3D point clouds, with the ubiquitous use of edge devices in scenarios like autonomous driving and robotics, which often demand real-time and reliable responses. The paper tackles the challenge by designing a general framework to construct 3D learning architectures with SO(3) equivariance and network binarization. However, a naive combination of equivariant networks and binarization either causes sub-optimal computational efficiency or geometric ambiguity. We propose to locate both scalar and vector features in our networks to avoid both cases. Precisely, the presence of scalar features makes the major part of the network binarizable, while vector features serve to retain rich structural information and ensure SO(3) equivariance. The proposed approach can be applied to general backbones like PointNet and DGCNN. Meanwhile, experiments on ModelNet40, ShapeNet, and the real-world dataset ScanObjectNN, demonstrated that the method achieves a great trade-off between efficiency, rotation robustness, and accuracy. The codes are available at https://github.com/zhuoinoulu/svnet. Zhuo Su 0002, Max Welling, Matti Pietikäinen, Li Liu 0002 |
3DV | 3 |
| 2022 | Dynamic Binary Neural Network by Learning Channel-Wise ThresholdsabstractBinary neural networks (BNNs) constrain weights and activations to +1 or -1 with limited storage and computational cost, which is hardware-friendly for portable devices. Recently, BNNs have achieved remarkable progress and been adopted into various fields. However, the performance of BNNs is sensitive to activation distribution. The existing BNNs utilized the Sign function with predefined or learned static thresholds to binarize activations. This process limits representation capacity of BNNs since different samples may adapt to unequal thresholds. To address this problem, we propose a dynamic BNN (DyBNN) incorporating dynamic learnable channel-wise thresholds of Sign function and shift parameters of PReLU. The method aggregates the global information into the hyper function and effectively increases the feature expression ability. The experimental results prove that our method is an effective and straightforward way to reduce information loss and enhance performance of BNNs. The DyBNN based on two backbones of ReActNet (MobileNetV1 and ResNet18) achieve 71.2% and 67.4% top1-accuracy on ImageNet dataset, outperforming baselines by a large margin (i.e., 1.8% and 1.5% respectively). Zhuo Su 0002, Yang-He Feng, Xin Lu 0002, Matti Pietikäinen, Li Liu 0002 |
ICASSP | 5 |
| 2022 | Analyzing Group-Level Emotion with Global Alignment Kernel based ApproachabstractFrom the perspective of social science, understanding group emotion has become increasingly important for teams to considerably accomplish organizational work. Currently, automatically analyzing the perceived affect of a group of people has been received increasingly interest in affective computing community. The variability in group size makes difficulty for group-level emotion recognition to straightforwardly measure the feature distance of two group-level images. Recent works attempted to resolve the preceding problem by using feature encoding. However, the early works lack of efficiency. To alleviate this problem, this article aims to design a new method to effectively analyze the group behavior from a group-level image. Motivated by time-series kernel approaches explored in dynamic facial expression classification, this article mainly concentrates on global alignment kernel and design support vector machine with the combined global alignment kernels (SVM-CGAK) to better recognize group-level emotion. Specifically, we first propose to use global alignment kernel to explicitly measure the distance of two group-level images. For improving the performance of global alignment kernel, we use the global weight sort scheme based on their spatial relation information to sort the faces from group-level image, making an efficient data structure to the global alignment kernel. With this new global alignment kernel, we construct the backbone of SVM-CGAK, namely, support vector machine with global alignment kernel. Furthermore, considering the challenging environment, we construct two global alignment kernels based on Reisz-based Volume Local Binary Pattern and deep convolutional neural network features, respectively. Lastly, to make the robustness of group-level emotion recognition, we propose SVM-CGAK combining both global alignment kernels with multiple kernel learning approach. It can enhance the discriminative ability of each global alignment kernel. Intensive experiments are conducted on three challenging group-level emotion databases. The experimental results demonstrate that the proposed approach achieves promising performance for group-level emotion recognition compared with the recent state-of-the-art methods. Xiaohua Huang 0003, Abhinav Dhall, Roland Göcke, Matti Pietikäinen, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | Deep Ladder-Suppression Network for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims at learning a classifier for an unlabeled target domain by transferring knowledge from a labeled source domain with a related but different distribution. Most existing approaches learn domain-invariant features by adapting the entire information of the images. However, forcing adaptation of domain-specific variations undermines the effectiveness of the learned features. To address this problem, we propose a novel, yet elegant module, called the deep ladder-suppression network (DLSN), which is designed to better learn the cross-domain shared content by suppressing domain-specific variations. Our proposed DLSN is an autoencoder with lateral connections from the encoder to the decoder. By this design, the domain-specific details, which are only necessary for reconstructing the unlabeled target data, are directly fed to the decoder to complete the reconstruction task, relieving the pressure of learning domain-specific variations at the later layers of the shared encoder. As a result, DLSN allows the shared encoder to focus on learning cross-domain shared content and ignores the domain-specific variations. Notably, the proposed DLSN can be used as a standard module to be integrated with various existing UDA frameworks to further boost performance. Without whistles and bells, extensive experimental results on four gold-standard domain adaptation datasets, for example: 1) Digits; 2) Office31; 3) Office-Home; and 4) VisDA-C, demonstrate that the proposed DLSN can consistently and significantly improve the performance of various popular UDA frameworks. Wanxia Deng, Lingjun Zhao, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Cybern. | 5 |
| 2022 | Hyperspectral Estimation of Soil Copper Concentration Based on Improved TabNet Model in the Eastern Junggar CoalfieldabstractChina is the largest coal consumer in the world. The massive exploitation and utilization of coal resources has resulted in serious problems of heavy metal pollution and environmental contamination, such as soil degradation, water pollution, crop damage, and even threatening human lives. Therefore, monitoring soil heavy metal pollution quickly and in real time is an urgent task at present. This research not only formulated a new preprocessing method enlightened by few-shot learning for soil hyperspectral data, but also combined it with other soil-related auxiliary information to extract effective information from the soil hyperspectrum, at the end of which different regression methods were adopted to predict soil heavy metal contamination. This test used 168 actual soil samples from the Eastern Junggar coalfield in Xinjiang for verification. Since copper in the soil is a trace element and the corresponding spectral characteristics are affected by other impurities, improper use of hyperspectral preprocessing methods may introduce interference information or may delete useful information, which makes the model effect unsatisfied. To effectively address the above problems, the preprocessing method of this experiment combined the second-order differential derivation, data enhancement method together with the addition of auxiliary information to allow more effective features to be entered into the model. Next, the Attentive Interpretable Tabular Learning (TabNet) model was improved in three different ways using the original TabNet model and three improved TabNet models to create regression models. One of the improved TabNet models had the best effect, with a list of the top 30 features according to the degree of importance. Meanwhile, the regression prediction of Cu content using four different convolutional neural networks (CNN) revealed that the model with the residual block was the strongest and slightly outperformed the improved TabNet model, but lacked interpretation of the input data. Besides, this experiment also employed different pre-processing methods for regression prediction on various models, and found that the traditional pre-processing methods performed best in traditional regression models (e.g., PLSR) and underperformed in deep learning models. The selected optimal model was compared with partial least square regression (PLSR), and convolutional neural network (CNN) models. The results indicated that both the improved TabNet model and improved CNN model had better performance using the new preprocessing approach proposed in this paper, with improved TabNet yielding a coefficient of determination (R2), root mean square error (RMSE) and ratio of performance to interquartile range (RPIQ) of 0.94, 1.341 and 4.474, respectively. The improved CNN model had a coefficient of determination of 0.942, a root mean square error of 1.324 and an interquartile range of 4.531 in the test dataset. Yuan Wang 0034, Abdugheni Abliz, Hongbing Ma, Li Liu 0002, Alishir Kurban, Ümüt Halik, Matti Pietikäinen |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Informative Feature Disentanglement for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims at learning a classifier for an unlabeled target domain by transferring knowledge from a labeled source domain with a related but different distribution. The strategy of aligning the two domains in latent feature space via metric discrepancy or adversarial learning has achieved considerable progress. However, these existing approaches mainly focus on adapting the entire image and ignore the bottleneck that occurs when forced adaptation of uninformative domain-specific variations undermines the effectiveness of learned features. To address this problem, we propose a novel component called Informative Feature Disentanglement (IFD), which is equipped with the adversarial network or the metric discrepancy model, respectively. Accordingly, the new network architectures, named IFDAN and IFDMN, enable informative feature refinement before the adaptation. The proposed IFD is designed to disentangle informative features from the uninformative domain-specific variations, which are produced by a Variational Autoencoder (VAE) with lateral connections from the encoder to the decoder. We cooperatively apply the IFD to conduct supervised disentanglement for the source domain and unsupervised disentanglement for the target domain. In this way, informative features are disentangled from the domain-specific details before the adaptation. Extensive experimental results on three gold-standard domain adaptation datasets, e.g., Office31, Office-Home and VisDA-C, demonstrate the effectiveness of the proposed IFDAN and IFDMN models for UDA. Wanxia Deng, Lingjun Zhao, Qing Liao 0001, Deke Guo, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Multim. | 7 |
| 2022 | Adaptive Semantic-Spatio-Temporal Graph Convolutional Network for Lip ReadingabstractThe goal of this work is to recognize words, phrases, and sentences being spoken by a talking face without given the audio. Current deep learning approaches for lip reading focus on exploring the appearance and optical flow information of videos. However, these methods do not fully exploit the characteristics of lip motion. In addition to appearance and optical flow, the mouth contour deformation usually conveys significant information that is complementary to others. However, the modeling of dynamic mouth contour has received little attention than that of appearance and optical flow. In this work, we propose a novel model of dynamic mouth contours called Adaptive Semantic-Spatio-Temporal Graph Convolution Network (ASST-GCN), to go beyond previous methods by automatically learning both the spatial and temporal information from videos. To combine the complementary information from appearance and mouth contour, a two-stream visual front-end network is proposed. Experimental results demonstrate that the proposed method significantly outperforms the state-of-the-art lip reading methods on several large-scale lip reading benchmarks. Changchong Sheng, Xinzhong Zhu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Multim. | 4 |
| 2021 | Pixel Difference Networks for Efficient Edge DetectionabstractRecently, deep Convolutional Neural Networks (CNNs) can achieve human-level performance in edge detection with the rich and abstract edge representation capacities. However, the high performance of CNN based edge detection is achieved with a large pretrained CNN backbone, which is memory and energy consuming. In addition, it is surprising that the previous wisdom from the traditional edge detectors, such as Canny, Sobel, and LBP are rarely investigated in the rapid-developing deep learning era. To address these issues, we propose a simple, lightweight yet effective architecture named Pixel Difference Network (PiDiNet) for efficient edge detection. PiDiNet adopts novel pixel difference convolutions that integrate the traditional edge detection operators into the popular convolutional operations in modern CNNs for enhanced performance on the task, which enjoys the best of both worlds. Extensive experiments on BSDS500, NYUD, and Multicue are provided to demonstrate its effectiveness, and its high training and inference efficiency. Surprisingly, when training from scratch with only the BSDS500 and VOC datasets, PiDiNet can surpass the recorded result of human perception (0.807 vs. 0.803 in ODS F-measure) on the BSDS500 dataset with 100 FPS and less than 1M parameters. A faster version of PiDiNet with less than 0.1M parameters can still achieve comparable performance among state of the arts with 200 FPS. Results on the NYUD and Multicue datasets show similar observations. The codes are available at https://github.com/zhuoinoulu/pidinet. Zhuo Su 0002, Zitong Yu, Dewen Hu, Qing Liao 0001, Qi Tian 0001, Matti Pietikäinen, Li Liu 0002 |
ICCV | 7 |
| 2021 | Informative Class-Conditioned Feature Alignment for Unsupervised Domain AdaptationabstractThe goal of unsupervised domain adaptation is to learn a task classifier that performs well for the unlabeled target domain by borrowing rich knowledge from a well-labeled source domain. Although remarkable breakthroughs have been achieved in learning transferable representation across domains, two bottlenecks remain to be further explored. First, many existing approaches focus primarily on the adaptation of the entire image, ignoring the limitation that not all features are transferable and informative for the object classification task. Second, the features of the two domains are typically aligned without considering the class labels; this can lead the resulting representations to be domain-invariant but non-discriminative to the category. To overcome the two issues, we present a novel Informative Class-Conditioned Feature Alignment (IC2FA) approach for UDA, which utilizes a twofold method: informative feature disentanglement and class-conditioned feature alignment, designed to address the above two challenges, respectively. More specifically, to surmount the first drawback, we cooperatively disentangle the two domains to obtain informative transferable features; here, Variational Information Bottleneck (VIB) is employed to encourage the learning of task-related semantic representations and suppress task-unrelated information. With regard to the second bottleneck, we optimize a new metric, termed Conditional Sliced Wasserstein Distance (CSWD), which explicitly estimates the intra-class discrepancy and the inter-class margin. The intra-class and inter-class CSWDs are minimized and maximized, respectively, to yield the domain-invariant discriminative features. IC2FA equips class-conditioned feature alignment with informative feature disentanglement and causes the two procedures to work cooperatively, which facilitates informative discriminative features adaptation. Extensive experimental results on three domain adaptation datasets confirm the superiority of IC2FA. Wanxia Deng, Yawen Cui, Zhen Liu 0004, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
ACM Multimedia | 6 |
| 2021 | Cross-modal Self-Supervised Learning for Lip Reading: When Contrastive Learning meets Adversarial TrainingabstractThe goal of this work is to learn discriminative visual representations for lip reading without access to manual text annotation. Recent advances in cross-modal self-supervised learning have shown that the corresponding audio can serve as a supervisory signal to learn effective visual representations for lip reading. However, existing methods only exploit the natural synchronization of the video and the corresponding audio. We find that both video and audio are actually composed of speech-related information, identity-related information, and modal information. To make the visual representations (i) more discriminative for lip reading and (ii) indiscriminate with respect to the identities and modals, we propose a novel self-supervised learning framework called Adversarial Dual-Contrast Self-Supervised Learning (ADC-SSL), to go beyond previous methods by explicitly forcing the visual representations disentangled from speech-unrelated information. Experimental results clearly show that the proposed method outperforms state-of-the-art cross-modal self-supervised baselines by a large margin. Besides, ADC-SSL can outperform its supervised counterpart without any finetune. Changchong Sheng, Matti Pietikäinen, Qi Tian 0001, Li Liu 0002 |
ACM Multimedia | 2 |
| 2021 | Editorial for the special issue of IMAVIS on automatic face analytics for human behavior understanding
Xiaohua Huang 0003, Abhinav Dhall, Guoying Zhao 0001, Wenming Zheng, Matti Pietikäinen |
Image Vis. Comput. | 5 |
| 2021 | Deep ladder reconstruction-classification network for unsupervised domain adaptation
Wanxia Deng, Zhuo Su 0002, Qiang Qiu 0001, Lingjun Zhao, Gangyao Kuang, Matti Pietikäinen, Huaxin Xiao, Li Liu 0002 |
Pattern Recognit. Lett. | 6 |
| 2020 | Dynamic Group Convolution for Accelerating Convolutional Neural Networks
Zhuo Su 0002, Linpu Fang, Wenxiong Kang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
ECCV (6) | 5 |
| 2020 | Deep Learning for Generic Object Detection: A SurveyabstractAbstract Object detection, one of the most fundamental and challenging problems in computer vision, seeks to locate object instances from a large number of predefined categories in natural images. Deep learning techniques have emerged as a powerful strategy for learning feature representations directly from data and have led to remarkable breakthroughs in the field of generic object detection. Given this period of rapid evolution, the goal of this paper is to provide a comprehensive survey of the recent achievements in this field brought about by deep learning techniques. More than 300 research contributions are included in this survey, covering many aspects of generic object detection: detection frameworks, object feature representation, object proposal generation, context modeling, training strategies, and evaluation metrics. We finish the survey by identifying promising directions for future research. Li Liu 0002, Wanli Ouyang, Xiaogang Wang 0001, Paul W. Fieguth, Jie Chen 0001, Xinwang Liu 0002, Matti Pietikäinen |
Int. J. Comput. Vis. | 7 |
| 2020 | Efficient Visual Recognition
Li Liu 0002, Matti Pietikäinen, Jie Qin 0004, Wanli Ouyang, Luc Van Gool |
Int. J. Comput. Vis. | 2 |
| 2019 | BIRD: Learning Binary and Illumination Robust Descriptor for Face Recognition
Zhuo Su 0002, Matti Pietikäinen, Li Liu 0002 |
BMVC | 2 |
| 2019 | From BoW to CNN: Two Decades of Texture Representation for Texture ClassificationabstractTexture is a fundamental characteristic of many types of images, and texture representation is one of the essential and challenging problems in computer vision and pattern recognition which has attracted extensive research attention over several decades. Since 2000, texture representations based on Bag of Words and on Convolutional Neural Networks have been extensively studied with impressive performance. Given this period of remarkable evolution, this paper aims to present a comprehensive survey of advances in texture representation over the last two decades. More than 250 major publications are cited in this survey covering different aspects of the research, including benchmark datasets and state of the art results. In retrospect of what has been achieved so far, the survey discusses open challenges and directions for future research. Li Liu 0002, Jie Chen 0001, Paul W. Fieguth, Guoying Zhao 0001, Rama Chellappa, Matti Pietikäinen |
Int. J. Comput. Vis. | 6 |
| 2019 | Guest Editors' Introduction to the Special Section on Compact and Efficient Feature Representation and Learning in Computer VisionabstractThe papers in this special section examine compact and efficient feature representation and learning in computer vision. Li Liu 0002, Matti Pietikäinen, Jie Chen 0001, Guoying Zhao 0001, Xiaogang Wang 0001, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Discriminative Spatiotemporal Local Binary Pattern with Revisited Integral Projection for Spontaneous Facial Micro-Expression RecognitionabstractRecently, there have been increasing interests in inferring mirco-expression from facial image sequences. Due to subtle facial movement of micro-expressions, feature extraction has become an important and critical issue for spontaneous facial micro-expression recognition. Recent works used spatiotemporal local binary pattern (STLBP) for micro-expression recognition and considered dynamic texture information to represent face images. However, they miss the shape attribute of face images. On the other hand, they extract the spatiotemporal features from the global face regions while ignore the discriminative information between two micro-expression classes. The above-mentioned problems seriously limit the application of STLBP to micro-expression recognition. In this paper, we propose a discriminative spatiotemporal local binary pattern based on an integral projection to resolve the problems of STLBP for micro-expression recognition. First, we revisit an integral projection for preserving the shape attribute of micro-expressions by using robust principal component analysis. Furthermore, a revisited integral projection is incorporated with local binary pattern across spatial and temporal domains. Specifically, we extract the novel spatiotemporal features incorporating shape attributes into spatiotemporal texture features. For increasing the discrimination of micro-expressions, we propose a new feature selection based on Laplacian method to extract the discriminative information for facial micro-expression recognition. Intensive experiments are conducted on three availably published micro-expression databases including CASME, CASME2 and SMIC databases. We compare our method with the state-of-the-art algorithms. Experimental results demonstrate that our proposed method achieves promising performance for micro-expression recognition. Xiaohua Huang 0003, Xin Liu 0012, Guoying Zhao 0001, Xiaoyi Feng, Matti Pietikäinen |
IEEE Trans. Affect. Comput. | 6 |
| 2019 | Texture Classification in Extreme Scale Variations Using GANetabstractResearch in texture recognition often concentrates on recognizing textures with intraclass variations, such as illumination, rotation, viewpoint, and small-scale changes. In contrast, in real-world applications, a change in scale can have a dramatic impact on texture appearance to the point of changing completely from one texture category to another. As a result, texture variations due to changes in scale are among the hardest to handle. In this paper, we conduct the first study of classifying textures with extreme variations in scale. To address this issue, we first propose and then reduce scale proposals on the basis of dominant texture patterns. Motivated by the challenges posed by this problem, we propose a new GANet network where we use a genetic algorithm to change the filters in the hidden layers during network training in order to promote the learning of more informative semantic texture patterns. Finally, we adopt a Fisher vector pooling of a convolutional neural network filter bank feature encoder for global texture representation. Because extreme scale variations are not necessarily present in most standard texture databases, to support the proposed extreme-scale aspects of texture understanding, we are developing a new dataset, the extreme scale variation textures (ESVaT), to test the performance of our framework. It is demonstrated that the proposed framework significantly outperforms the gold-standard texture features by more than 10% on ESVaT. We also test the performance of our proposed approach on the KTHTIPS2b and OS datasets and a further dataset synthetically derived from Forrest, showing the superior performance compared with the state-of-the-art. Li Liu 0002, Jie Chen 0001, Guoying Zhao 0001, Paul W. Fieguth, Xilin Chen 0001, Matti Pietikäinen |
IEEE Trans. Image Process. | 6 |
| 2018 | Sparse Tikhonov-Regularized Hashing for Multi-Modal LearningabstractThis paper mainly focuses on the role of regularization in Multi-Modal Learning (MML). Existing MML studies devote most of the efforts in maximizing the consensus of models from cues of different modalities. However, regularization methods are still far from fully explored. To fill in this gap, we propose a compact and efficient coding solution, termed by sparse Tikhonov-Regularized Hashing (STRH). The STRH enforces both the ℓ0-norm induced sparsity constraints and the Tikhonov regularization on the binary solution vectors which maximize cross-modal correlation. In addition, we raise the concerns on the challenging testing scenario of `Multi-modal Learning and Single-modal Prediction' (MLSP). Finally, we demonstrate that the STRH is an efficient hashing solutions by showing its superiority under the MLSP scenario. Lei Tian 0002, Xiaopeng Hong, Chunxiao Fan 0001, Yue Ming 0001, Matti Pietikäinen, Guoying Zhao 0001 |
ICIP | 5 |
| 2018 | Sparse projections matrix binary descriptors for face recognition
Chunxiao Fan 0001, Lei Tian 0002, Yue Ming 0001, Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen |
Neurocomputing | 6 |
| 2018 | Towards Reading Hidden Emotions: A Comparative Study of Spontaneous Micro-Expression Spotting and Recognition MethodsabstractMicro-expressions (MEs) are rapid, involuntary facial expressions which reveal emotions that people do not intend to show. Studying MEs is valuable as recognizing them has many important applications, particularly in forensic science and psychotherapy. However, analyzing spontaneous MEs is very challenging due to their short duration and low intensity. Automatic ME analysis includes two tasks: ME spotting and ME recognition. For ME spotting, previous studies have focused on posed rather than spontaneous videos. For ME recognition, the performance of previous studies is low. To address these challenges, we make the following contributions: (i) We propose the first method for spotting spontaneous MEs in long videos (by exploiting feature difference contrast). This method is training free and works on arbitrary unseen videos. (ii) We present an advanced ME recognition framework, which outperforms previous work by a large margin on two challenging spontaneous ME databases (SMIC and CASMEII). (iii) We propose the first automatic ME analysis system (MESR), which can spot and recognize MEs from spontaneous video data. Finally, we show our method outperforms humans in the ME recognition task by a large margin, and achieves comparable performance to humans at the very challenging task of spotting and then recognizing spontaneous MEs. Xiaopeng Hong, Antti Moilanen, Xiaohua Huang 0003, Tomas Pfister, Guoying Zhao 0001, Matti Pietikäinen |
IEEE Trans. Affect. Comput. | 7 |
| 2018 | Multimodal Framework for Analyzing the Affect of a Group of PeopleabstractWith the advances in multimedia and the world wide web, users upload millions of images and videos everyone on social networking platforms on the Internet. From the perspective of automatic human behavior understanding, it is of interest to analyze and model the affects that are exhibited by groups of people who are participating in social events in these images. However, the analysis of the affect that is expressed by multiple people is challenging due to the varied indoor and outdoor settings. Recently, a few interesting works have investigated face-based group-level emotion recognition (GER). In this paper, we propose a multimodal framework for enhancing the affective analysis ability of GER in challenging environments. Specifically, for encoding a person's information in a group-level image, we first propose an information aggregation method for generating feature descriptions of face, upper body, and scene. Later, we revisit localized multiple kernel learning for fusing face, upper body, and scene information for GER against challenging environments. Intensive experiments are performed on two challenging group-level emotion databases (HAPPEI and GAFF) to investigate the roles of the face, upper body, scene information, and the multimodal framework. Experimental results demonstrate that the multimodal framework achieves promising performance for GER. Xiaohua Huang 0003, Abhinav Dhall, Roland Göcke, Matti Pietikäinen, Guoying Zhao 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | Robust local features for remote face recognition
Jie Chen 0001, Vishal M. Patel, Li Liu 0002, Vili Kellokumpu, Guoying Zhao 0001, Matti Pietikäinen, Rama Chellappa |
Image Vis. Comput. | 6 |
| 2017 | Local binary features for texture classification: Taxonomy and experimental study
Li Liu 0002, Paul W. Fieguth, Yulan Guo, Xiaogang Wang 0001, Matti Pietikäinen |
Pattern Recognit. | 5 |
| 2017 | HEp-2 Cell Classification via Combining Multiresolution Co-Occurrence Texture and Large Region Shape InformationabstractIndirect immunofluorescence imaging of human epithelial type 2 (HEp-2) cell image is an effective evidence to diagnose autoimmune diseases. Recently, computer-aided diagnosis of autoimmune diseases by the HEp-2 cell classification has attracted great attention. However, the HEp-2 cell classification task is quite challenging due to large intraclass and small interclass variations. In this paper, we propose an effective approach for the automatic HEp-2 cell classification by combining multiresolution co-occurrence texture and large regional shape information. To be more specific, we propose to: 1) capture multiresolution co-occurrence texture information by a novel pairwise rotation-invariant co-occurrence of local Gabor binary pattern descriptor; 2) depict large regional shape information by using an improved Fisher vector model with RootSIFT features, which are sampled from large image patches in multiple scales; and 3) combine both features. We evaluate systematically the proposed approach on the IEEE International Conference on Pattern Recognition (ICPR) 2012, the IEEE International Conference on Image Processing (ICIP) 2013, and the ICPR 2014 contest datasets. The proposed method based on the combination of the introduced two features outperforms the winners of the ICPR 2012 contest using the same experimental protocol. Our method also greatly improves the winner of the ICIP 2013 contest under four different experimental setups. Using the leave-one-specimen-out evaluation strategy, our method achieves comparable performance with the winner of the ICPR 2014 contest that combined four features. Xianbiao Qi, Guoying Zhao 0001, Chun-Guang Li, Jun Guo 0002, Matti Pietikäinen |
IEEE J. Biomed. Health Informatics | 5 |
| 2016 | Evaluation of LBP and Deep Texture Descriptors with a New Robustness Benchmark
Li Liu 0002, Paul W. Fieguth, Xiaogang Wang 0001, Matti Pietikäinen, Dewen Hu |
ECCV (3) | 4 |
| 2016 | Generalized face anti-spoofing by detecting pulse from face videosabstractFace biometric systems are vulnerable to spoofing attacks. Such attacks can be performed in many ways, including presenting a falsified image, video or 3D mask of a valid user. A widely used approach for differentiating genuine faces from fake ones has been to capture their inherent differences in (2D or 3D) texture using local descriptors. One limitation of these methods is that they may fail if an unseen attack type, e.g. a highly realistic 3D mask which resembles real skin texture, is used in spoofing. Here we propose a robust anti-spoofing method by detecting pulse from face videos. Based on the fact that a pulse signal exists in a real living face but not in any mask or print material, the method could be a generalized solution for face liveness detection. The proposed method is evaluated first on a 3D mask spoofing database 3DMAD to demonstrate its effectiveness in detecting 3D mask attacks. More importantly, our cross-database experiment with high quality REAL-F masks shows that the pulse based method is able to detect even the previously unseen mask type whereas texture based methods fail to generalize beyond the development data. Finally, we propose a robust cascade system combining two complementary attack-specific spoof detectors, i.e. utilize pulse detection against print attacks and color texture analysis against video attacks. Jukka Komulainen, Guoying Zhao 0001, Pong C. Yuen, Matti Pietikäinen |
ICPR | 5 |
| 2016 | Multi-modal emotion analysis from facial expressions and electroencephalogram
Xiaohua Huang 0003, Jukka Kortelainen, Guoying Zhao 0001, Antti Moilanen, Tapio Seppänen, Matti Pietikäinen |
Comput. Vis. Image Underst. | 7 |
| 2016 | Editorial of special issue on spontaneous facial behaviour analysis
Stefanos Zafeiriou, Guoying Zhao 0001, Matti Pietikäinen, Rama Chellappa, Irene Kotsia, Jeffrey F. Cohn |
Comput. Vis. Image Underst. | 3 |
| 2016 | RoLoD: Robust local descriptors for computer vision
Jie Chen 0001, Zhen Lei 0001, Li Liu 0002, Guoying Zhao 0001, Matti Pietikäinen |
Neurocomputing | 5 |
| 2016 | Capturing correlations of local features for image representation
Xiaopeng Hong, Guoying Zhao 0001, Stefanos Zafeiriou, Maja Pantic, Matti Pietikäinen |
Neurocomputing | 5 |
| 2016 | Spontaneous facial micro-expression analysis using Spatiotemporal Completed Local Quantized Patterns
Xiaohua Huang 0003, Guoying Zhao 0001, Xiaopeng Hong, Wenming Zheng, Matti Pietikäinen |
Neurocomputing | 5 |
| 2016 | Dynamic texture and scene classification by transferring deep image features
Xianbiao Qi, Chun-Guang Li, Guoying Zhao 0001, Xiaopeng Hong, Matti Pietikäinen |
Neurocomputing | 5 |
| 2016 | LOAD: Local orientation adaptive descriptor for texture and material classification
Xianbiao Qi, Guoying Zhao 0001, LinLin Shen, Qingquan Li 0001, Matti Pietikäinen |
Neurocomputing | 5 |
| 2016 | Extended local binary patterns for face recognition
Li Liu 0002, Paul W. Fieguth, Guoying Zhao 0001, Matti Pietikäinen, Dewen Hu |
Inf. Sci. | 4 |
| 2016 | Exploring illumination robust descriptors for human epithelial type 2 cell classification
Xianbiao Qi, Guoying Zhao 0001, Jie Chen 0001, Matti Pietikäinen |
Pattern Recognit. | 4 |
| 2016 | Data-driven techniques for smoothing histograms of local binary patterns
Juha Ylioinas, Norman Poh, Jukka Holappa, Matti Pietikäinen |
Pattern Recognit. | 4 |
| 2016 | HEp-2 cell classification: The role of Gaussian Scale Space Theory as a pre-processing approach
Xianbiao Qi, Guoying Zhao 0001, Jie Chen 0001, Matti Pietikäinen |
Pattern Recognit. Lett. | 4 |
| 2016 | Dynamic Facial Expression Recognition With Atlas Construction and Sparse RepresentationabstractIn this paper, a new dynamic facial expression recognition method is proposed. Dynamic facial expression recognition is formulated as a longitudinal groupwise registration problem. The main contributions of this method lie in the following aspects: 1) subject-specific facial feature movements of different expressions are described by a diffeomorphic growth model; 2) salient longitudinal facial expression atlas is built for each expression by a sparse groupwise image registration method, which can describe the overall facial feature changes among the whole population and can suppress the bias due to large intersubject facial variations; and 3) both the image appearance information in spatial domain and topological evolution information in temporal domain are used to guide recognition by a sparse representation method. The proposed framework has been extensively evaluated on five databases for different applications: the extended Cohn-Kanade, MMI, FERA, and AFEW databases for dynamic facial expression recognition, and UNBC-McMaster database for spontaneous pain expression monitoring. This framework is also compared with several state-of-the-art dynamic facial expression recognition methods. The experimental results demonstrate that the recognition rates of the new method are consistently higher than other methods under comparison. Yimo Guo, Guoying Zhao 0001, Matti Pietikäinen |
IEEE Trans. Image Process. | 3 |
| 2016 | Median Robust Extended Local Binary Pattern for Texture ClassificationabstractLocal binary patterns (LBP) are considered among the most computationally efficient high-performance texture features. However, the LBP method is very sensitive to image noise and is unable to capture macrostructure information. To best address these disadvantages, in this paper, we introduce a novel descriptor for texture classification, the median robust extended LBP (MRELBP). Different from the traditional LBP and many LBP variants, MRELBP compares regional image medians rather than raw image intensities. A multiscale LBP type descriptor is computed by efficiently comparing image medians over a novel sampling scheme, which can capture both microstructure and macrostructure texture information. A comprehensive evaluation on benchmark data sets reveals MRELBP's high performance-robust to gray scale variations, rotation changes and noise-but at a low computational cost. MRELBP produces the best classification scores of 99.82%, 99.38%, and 99.77% on three popular Outex test suites. More importantly, MRELBP is shown to be highly robust to image noise, including Gaussian noise, Gaussian blur, salt-and-pepper noise, and random pixel corruption. Li Liu 0002, Songyang Lao, Paul W. Fieguth, Yulan Guo, Xiaogang Wang 0001, Matti Pietikäinen |
IEEE Trans. Image Process. | 6 |
| 2015 | Spatiotemporal Integration of Optical Flow Vectors for Micro-expression Detection
Devangini Patel, Guoying Zhao 0001, Matti Pietikäinen |
ACIVS | 3 |
| 2015 | A Task-Driven Eye Tracking Dataset for Visual Attention Analysis
Yingyue Xu, Xiaopeng Hong, Qiuhai He, Guoying Zhao 0001, Matti Pietikäinen |
ACIVS | 5 |
| 2015 | Riesz-based Volume Local Binary Pattern and A Novel Group Expression Model for Group Happiness Intensity AnalysisabstractAutomatic emotion analysis and understanding has received much attention over the years in affective computing. Recently, there are increasing interests in inferring the emotional intensity of a group of people. For group emotional intensity analysis, feature extraction and group expression model are two critical issues. In this paper, we propose a new method to estimate the happiness intensity of a group of people in an image. Firstly, we combine the Riesz transform and the local binary pattern descriptor, named Riesz-based volume local binary pattern, which considers neighbouring changes not only in the spatial domain of a face but also along the different Riesz faces. Secondly, we exploit the continuous conditional random fields for constructing a new group expression model, which considers global and local attributes. Intensive experiments are performed on three challenging facial expression databases to evaluate the novel feature. Furthermore, experiments are conducted on the HAPPEI database to evaluate the new group expression model with the new feature. Our experimental results demonstrate the promising performance for group happiness intensity analysis. Xiaohua Huang 0003, Abhinav Dhall, Guoying Zhao 0001, Roland Göcke, Matti Pietikäinen |
BMVC | 5 |
| 2015 | Median robust extended local binary pattern for texture classificationabstractLocal Binary Patterns (LBP) are among the most computationally efficient amongst high-performance texture features. However, LBP is very sensitive to image noise and is unable to capture macrostructure information. To best address these disadvantages, in this paper we introduce a novel descriptor for texture classification, the Median Robust Extended Local Binary Pattern (MRELBP). In contrast to traditional LBP and many LBP variants, MRELBP compares local image medians instead of raw image intensities. We develop a multiscale LBP-type descriptor by efficiently comparing image medians over a novel sampling scheme, which can capture both microstructure and macrostructure. A comprehensive evaluation on benchmark datasets reveals MRELBP's remarkable performance (robust to gray scale variations, rotation changes and noise) relative to state-of-the-art algorithms, but nevertheless at a low computational cost, producing the best classification scores of 99.82%, 99.38% and 99.77% on three popular Outex test suites. Furthermore, MRELBP is also shown to be highly robust to image noise including Gaussian noise, Gaussian blur, Salt-and-Pepper noise and random pixel corruption. Li Liu 0002, Paul W. Fieguth, Matti Pietikäinen, Songyang Lao |
ICIP | 3 |
| 2015 | Globally rotation invariant multi-scale co-occurrence local binary pattern
Xianbiao Qi, LinLin Shen, Guoying Zhao 0001, Qingquan Li 0001, Matti Pietikäinen |
Image Vis. Comput. | 5 |
| 2014 | Remote Heart Rate Measurement from Face Videos under Realistic SituationsabstractHeart rate is an important indicator of people's physiological state. Recently, several papers reported methods to measure heart rate remotely from face videos. Those methods work well on stationary subjects under well controlled conditions, but their performance significantly degrades if the videos are recorded under more challenging conditions, specifically when subjects' motions and illumination variations are involved. We propose a framework which utilizes face tracking and Normalized Least Mean Square adaptive filtering methods to counter their influences. We test our framework on a large difficult and public database MAHNOB-HCI and demonstrate that our method substantially outperforms all previous methods. We also use our method for long term heart rate monitoring in a game evaluation scenario and achieve promising results. Jie Chen 0001, Guoying Zhao 0001, Matti Pietikäinen |
CVPR | 4 |
| 2014 | Generalized textured contact lens detection by extracting BSIF description from Cartesian iris imagesabstractTextured contact lenses cause severe problems for iris biometric systems because they can be used to alter the appearance of iris texture in order to deliberately increase the false positive and, especially, false negative match rates. Many texture analysis based techniques have been proposed for detecting the presence of cosmetic contact lenses. However, it has been shown recently that the generalization capability of the existing approaches is not sufficient because they have been developed for detecting specific lens texture patterns and evaluated only on those same lens types seen during development phase. This scenario does not apply in unpredictable practical applications because unseen lens patterns will be definitely experienced in operation. In this paper, we address this issue by studying the effect of different iris image preprocessing techniques and introducing a novel approach formore generalized cosmetic contact lens detection using binarized statistical image features (BSIF).Our extensive experimental analysis on benchmark datasets shows that the BSIF description extracted from preprocessed Cartesian iris texture images yields to promising generalization capabilities across unseen texture patterns and different iris sensors with mean equal error rate of 0.14%and 0.88%, respectively. The findings support the intuition that the textural differences between genuine iris texture and fake ones are best described by preserving the regular structure of different printing signatures without transforming the iris images into polar coordinate system. Jukka Komulainen, Abdenour Hadid, Matti Pietikäinen |
IJCB | 3 |
| 2014 | Extended local binary pattern fusion for face recognitionabstractThis paper presents a simple, novel, yet highly effective approach for robust face recognition. Given LBP-like descriptors based on local accumulated pixel differences, Angular Differences (AD) and Radial Differences (RD), the local differences are decomposed into complementary components of signs and magnitudes. The proposed descriptors have desirable features: (1) robustness to lighting, pose, and expression; (2) computation efficiency; (3) encoding of both microstructures and macrostructures; (4) consistent in form with traditional LBP, thus inheriting the merits of LBP; and (5) no required training, improving generalizability. From a given face image, we obtain six histogram features, each of which is obtained by concatenating spatial histograms extracted from nonoverlapping subregions. The Whitened PCA technique is used for dimensionality reduction, followed by Nearest Neighbor classification. We have evaluated the effectiveness of the proposed method on the Extended Yale B and CAS-PEAL-R1 databases. The proposed method impressively outperforms other well known systems, including what we believe to be the best reported performance for the the CAS-PEAL-R1 lighting probe set with a recognition rate of 72.3%. Li Liu 0002, Paul W. Fieguth, Guoying Zhao 0001, Matti Pietikäinen |
ICIP | 4 |
| 2014 | Improved Spatiotemporal Local Monogenic Binary Pattern for Emotion Recognition in The WildabstractLocal binary pattern from three orthogonal planes (LBP-TOP) has been widely used in emotion recognition in the wild. However, it suffers from illumination and pose changes. This paper mainly focuses on the robustness of LBP-TOP to unconstrained environment. Recent proposed method, spatiotemporal local monogenic binary pattern (STLMBP), was verified to work promisingly in different illumination conditions. Thus this paper proposes an improved spatiotemporal feature descriptor based on STLMBP. The improved descriptor uses not only magnitude and orientation, but also the phase information, which provide complementary information. In detail, the magnitude, orientation and phase images are obtained by using an effective monogenic filter, and multiple feature vectors are finally fused by multiple kernel learning. STLMBP and the proposed method are evaluated in the Acted Facial Expression in the Wild as part of the 2014 Emotion Recognition in the Wild Challenge. They achieve competitive results, with an accuracy gain of 6.35% and 7.65% above the challenge baseline (LBP-TOP) over video. Xiaohua Huang 0003, Qiuhai He, Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen |
ICMI | 5 |
| 2014 | Pose Estimation via Complex-Frequency Domain Analysis of Image Gradient OrientationsabstractHead Pose Estimation (HPE) has recently attracted a lot of interests in various computer vision applications. One challenging problem for accurate HPE is to model the intrinsic variations among poses, and suppress the extraneous variations derived from other factors, such as the illumination changes, outliers, and noise. To this end, this paper proposes a simple and efficient facial description for head pose estimation from images. To handle the illumination changes, we characterize each image pixel by its image gradient orientation (IGO), rather than the intensity, which is sensitive to illumination changes. We then carry out complex-frequency domain analysis of the IGO image via the two-dimensional image transform, such as the 2D Discrete Cosine Transform (DCT2), to encode the spatial configuration of image gradient orientations. The proposed facial description is called IGO-DCT2. It is robust to illumination changes, outliers, and noise. In addition, it is learning free and computationally efficient. Finally, the fine-grain head pose estimation is regarded as a regression problem and off-the-shelf non-linear regression models are used to learn the mapping from the feature space to the continuous pose labels. Experimental results show the proposed facial description achieves highly competitive results on the publicly available FacePix dataset. Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen |
ICPR | 3 |
| 2014 | Robust Facial Expression Recognition Using Revised Canonical CorrelationabstractThe poor alignment and large variations in the temporal sale of facial expressions are two crucial problems for facial expression recognition (FER). Canonical correlation (CC) has recently received increasing attention in surveillance face recognition because of its robustness to variations of alignment. But it could not suit well to FER, because the facial expression variations and temporal information are ignored for canonical subspace. This paper proposes the revised canonical correlation method to address the two above described issues for making FER be robust to false detection or mis-alignment. Firstly, this paper presents the local binary pattern to describe the appearance features for enhancing the spatial variations of facial expression. Secondly, this paper proposes the temporal orthogonal locality preserved projection for building a canonical subspace of a video clip, where it mostly captures the motion changes of facial expressions. Then this paper presents the discriminative CC to model the low-dimensional feature space, which increases robustness to imprecise alignment and strengthens discrimination for facial expressions. Extensive experimental results on Extended Cohn-Kanade and MAHNOB-HCI databases demonstrate that the proposed method achieves the best results in recognizing facial expressions and performs robustly with ordinary on general face detection and eye detection. Xiaohua Huang 0003, Guoying Zhao 0001, Matti Pietikäinen, Wenming Zheng |
ICPR | 3 |
| 2014 | Spotting Rapid Facial Movements from Videos Using Appearance-Based Feature Difference AnalysisabstractSpotting micro-expressions is a primary step for continuous emotion recognition from videos. Spotting in this context refers to automatically finding the temporal locations of the face-related events from a video sequence. Rapid facial movements mainly include micro-expressions and eye blinks. However, the role of eye blinks in expressing emotions is still controversial, and often they are considered as micro-expressions as well. In this paper a simple method for automatically spotting rapid facial movements from videos is proposed. The method relies on analyzing differences in appearance-based features of sequential frames. In addition to finding the temporal locations, the system is able to provide spatial information about the movements in the face. Micro-expression spotting experiments are carried out on three datasets consisting only of spontaneous micro-expressions. Baseline micro-expression spotting results are provided for these three datasets including the publicly available CASME database. Also an example of spatial localization of the spotted rapid movements is presented. Antti Moilanen, Guoying Zhao 0001, Matti Pietikäinen |
ICPR | 3 |
| 2014 | Facial 3D Shape Estimation from Images for Visual Speech AnimationabstractIn this paper we describe the first version of our system for estimating 3D shape sequences from images of the frontal face. This approach is developed with 3D Visual Speech Animation (VSA) as the target application. In particular, the focus is on the usability of an existing state-of-the-art image-based VSA system and subsequent on-line estimation of the corresponding 3D facial shape sequence from its output. This has the added advantage of a 3D visual speech, which is mainly render ability of the face in different poses and illumination conditions. The idea is based on the detection of landmarks from the facial image which are then used to determine the pose and shape. The method belongs to the category of methods which use a prior 3D Morph able Models (3D-MM) trained using 3D facial data. For the time being it is developed for a person-specific domain, i.e. the 3D-MM and the 2D facial landmark detector are trained using the data of a single person and tested with the same person-specific data. Utpala Musti, Ziheng Zhou 0003, Matti Pietikäinen |
ICPR | 3 |
| 2014 | An In-depth Examination of Local Binary Descriptors in Unconstrained Face RecognitionabstractAutomatic face recognition in unconstrained conditions is a difficult task which has recently attained increasing attention. In this domain, face verification methods have significantly improved since the release of the Labeled Faces in the Wild database, but the related problem of face identification, is still lacking considerations, which is partly because of the shortage of representative databases. Only recently, two new datasets called Remote Face and Point-and-Shoot Challenge were published providing appropriate benchmarks for the research community to investigate the problem of face recognition in challenging imaging conditions, in both, verification and identification modes. In this paper we provide an in-depth examination of three local binary description methods in unconstrained face recognition evaluating them on these two recently published datasets. In detail, we investigate three well established methods separately and fusing them at rank- and score-levels. We are using a well-defined evaluation protocol allowing a fair comparison of our results for future examinations. Juha Ylioinas, Abdenour Hadid, Juho Kannala, Matti Pietikäinen |
ICPR | 4 |
| 2014 | Learning local image descriptors using binary decision treesabstractIn this paper we propose a unified framework for learning such local image descriptors that describe pixel neighborhoods using binary codes. The descriptors are constructed using binary decision trees which are learnt from a set of training image patches. Our framework generalizes several previously proposed binary descriptors, such as BRIEF, LBP and their variants, and provides a principled way to learn new constructions which have not been previously studied. Further, the proposed framework can utilize both labeled or unlabeled training data, and hence fits to both supervised and unsupervised learning scenarios. We evaluate our framework using varying levels of supervision in the learning phase. The experiments show that our descriptor constructions perform comparably to benchmark descriptors in two different applications, namely texture categorization and age group classification from facial images. Juha Ylioinas, Juho Kannala, Abdenour Hadid, Matti Pietikäinen |
WACV | 4 |
| 2014 | A review of recent advances in visual speech decoding
Ziheng Zhou 0003, Guoying Zhao 0001, Xiaopeng Hong, Matti Pietikäinen |
Image Vis. Comput. | 4 |
| 2014 | Learning Discriminant Face DescriptorabstractLocal feature descriptor is an important module for face recognition and those like Gabor and local binary patterns (LBP) have proven effective face descriptors. Traditionally, the form of such local descriptors is predefined in a handcrafted way. In this paper, we propose a method to learn a discriminant face descriptor (DFD) in a data-driven way. The idea is to learn the most discriminant local features that minimize the difference of the features between images of the same person and maximize that between images from different people. In particular, we propose to enhance the discriminative ability of face representation in three aspects. First, the discriminant image filters are learned. Second, the optimal neighborhood sampling strategy is soft determined. Third, the dominant patterns are statistically constructed. Discriminative learning is incorporated to extract effective and robust features. We further apply the proposed method to the heterogeneous (cross-modality) face recognition problem and learn DFD in a coupled way (coupled DFD or C-DFD) to reduce the gap between features of heterogeneous face images to improve the performance of this challenging problem. Extensive experiments on FERET, CAS-PEAL-R1, LFW, and HFB face databases validate the effectiveness of the proposed DFD learning on both homogeneous and heterogeneous face recognition problems. The DFD improves POEM and LQP by about 4.5 percent on LFW database and the C-DFD enhances the heterogeneous face recognition performance of LBP by over 25 percent. Zhen Lei 0001, Matti Pietikäinen, Stan Z. Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | A Compact Representation of Visual Speech Data Using Latent VariablesabstractThe problem of visual speech recognition involves the decoding of the video dynamics of a talking mouth in a high-dimensional visual space. In this paper, we propose a generative latent variable model to provide a compact representation of visual speech data. The model uses latent variables to separately represent the interspeaker variations of visual appearances and those caused by uttering within images, and incorporates the structural information of the visual data through placing priors of the latent variables along a curve embedded within a path graph. Ziheng Zhou 0003, Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | Combining LBP Difference and Feature Correlation for Texture DescriptionabstractEffective characterization of texture images requires exploiting multiple visual cues from the image appearance. The local binary pattern (LBP) and its variants achieve great success in texture description. However, because the LBP(-like) feature is an index of discrete patterns rather than a numerical feature, it is difficult to combine the LBP(-like) feature with other discriminative ones by a compact descriptor. To overcome the problem derived from the nonnumerical constraint of the LBP, this paper proposes a numerical variant accordingly, named the LBP difference (LBPD). The LBPD characterizes the extent to which one LBP varies from the average local structure of an image region of interest. It is simple, rotation invariant, and computationally efficient. To achieve enhanced performance, we combine the LBPD with other discriminative cues by a covariance matrix. The proposed descriptor, termed the covariance and LBPD descriptor (COV-LBPD), is able to capture the intrinsic correlation between the LBPD and other features in a compact manner. Experimental results show that the COV-LBPD achieves promising results on publicly available data sets. Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen, Xilin Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2013 | RLBP: Robust Local Binary PatternabstractIn this paper, we propose a simple and robust local descriptor, called the robust local binary pattern (RLBP). The local binary pattern (LBP) works very successfully in many domains, such as texture classification, human detection and face recognition. However, an issue of LBP is that it is not so robust to the noise present in the image. We improve the robustness of LBP by changing the coding bit of LBP. Experimental results on the Brodatz and UIUC texture databases show that RLBP impressively outperforms the other widely used descriptors (e.g., SIFT, Gabor, MR8 and LBP) and other variants of LBP (e.g., completed LBP), especially when we add noise in the images. In addition, experimental results on human face recognition also show a promising performance comparable to the best known results on the Face Recognition Grand Challenge (FRGC) face dataset. Jie Chen 0001, Vili Kellokumpu, Guoying Zhao 0001, Matti Pietikäinen |
BMVC | 4 |
| 2013 | Emotion recognition from facial images with arbitrary viewsabstractFacial expression recognition has been predominantly utilized to analyze the emotional status of human beings. In practice nearly frontal-view facial images may not be available. Therefore, a desirable property of facial expression recognition would allow the user to have any head pose. Some methods on non-frontal-view facial images were recently proposed to recognize the facial expressions by building discriminative subspace in specific views. We argue that this kind of approach ignores (1) the discrimination of inter-class samples with the same view label and (2) the closeness of intra-class samples with all view labels. This paper proposes a new method to recognize arbitrary-view facial expressions by using discriminative neighborhood preserving embedding and multi-view concepts. It first captures the discriminative property of inter-class samples. In addition, it explores the closeness of intra-class samples with arbitrary view in a low-dimensional subspace. Experimental results on BU-3DFE and Multi-PIE databases show that our approach achieves promising results for recognizing facial expressions with arbitrary views. Xiaohua Huang 0003, Guoying Zhao 0001, Matti Pietikäinen |
BMVC | 3 |
| 2013 | Demographic classification from face videos using manifold learning
Abdenour Hadid, Matti Pietikäinen |
Neurocomputing | 2 |
| 2013 | Machine learning in motion analysis: New advances
Matti Pietikäinen, Matthew Turk 0001, Liang Wang 0001, Guoying Zhao 0001, Li Cheng 0001 |
Image Vis. Comput. | 1 |
| 2013 | Multiscale Local Phase Quantization for Robust Component-Based Face Recognition Using Kernel Fusion of Multiple DescriptorsabstractFace recognition subject to uncontrolled illumination and blur is challenging. Interestingly, image degradation caused by blurring, often present in real-world imagery, has mostly been overlooked by the face recognition community. Such degradation corrupts face information and affects image alignment, which together negatively impact recognition accuracy. We propose a number of countermeasures designed to achieve system robustness to blurring. First, we propose a novel blur-robust face image descriptor based on Local Phase Quantization (LPQ) and extend it to a multiscale framework (MLPQ) to increase its effectiveness. To maximize the insensitivity to misalignment, the MLPQ descriptor is computed regionally by adopting a component-based framework. Second, the regional features are combined using kernel fusion. Third, the proposed MLPQ representation is combined with the Multiscale Local Binary Pattern (MLBP) descriptor using kernel fusion to increase insensitivity to illumination. Kernel Discriminant Analysis (KDA) of the combined features extracts discriminative information for face recognition. Last, two geometric normalizations are used to generate and combine multiple scores from different face image scales to further enhance the accuracy. The proposed approach has been comprehensively evaluated using the combined Yale and Extended Yale database B (degraded by artificially induced linear motion blur) as well as the FERET, FRGC 2.0, and LFW databases. The combined system is comparable to state-of-the-art approaches using similar system configurations. The reported work provides a new insight into the merits of various face representation and fusion methods, as well as their role in dealing with variable lighting and blur degradation. Chi-Ho Chan, Muhammad Atif Tahir, Josef Kittler, Matti Pietikäinen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2013 | Automatic Dynamic Texture Segmentation Using Local Descriptors and Optical FlowabstractA dynamic texture (DT) is an extension of the texture to the temporal domain. How to segment a DT is a challenging problem. In this paper, we address the problem of segmenting a DT into disjoint regions. A DT might be different from its spatial mode (i.e., appearance) and/or temporal mode (i.e., motion field). To this end, we develop a framework based on the appearance and motion modes. For the appearance mode, we use a new local spatial texture descriptor to describe the spatial mode of the DT; for the motion mode, we use the optical flow and the local temporal texture descriptor to represent the temporal variations of the DT. In addition, for the optical flow, we use the histogram of oriented optical flow (HOOF) to organize them. To compute the distance between two HOOFs, we develop a simple effective and efficient distance measure based on Weber's law. Furthermore, we also address the problem of threshold selection by proposing a method for determining thresholds for the segmentation method by an offline supervised statistical learning. The experimental results show that our method provides very good segmentation results compared to the state-of-the-art methods in segmenting regions that differ in their dynamics. Jie Chen 0001, Guoying Zhao 0001, Mikko Salo, Esa Rahtu, Matti Pietikäinen |
IEEE Trans. Image Process. | 5 |
| 2013 | Video Texture Synthesis With Multi-Frame LBP-TOP and Diffeomorphic Growth ModelabstractVideo texture synthesis is the process of providing a continuous and infinitely varying stream of frames, which plays an important role in computer vision and graphics. However, it still remains a challenging problem to generate high-quality synthesis results. Considering the two key factors that affect the synthesis performance, frame representation and blending artifacts, we improve the synthesis performance from two aspects: 1) Effective frame representation is designed to capture both the image appearance information in spatial domain and the longitudinal information in temporal domain. 2) Artifacts that degrade the synthesis quality are significantly suppressed on the basis of a diffeomorphic growth model. The proposed video texture synthesis approach has two major stages: video stitching stage and transition smoothing stage. In the first stage, a video texture synthesis model is proposed to generate an infinite video flow. To find similar frames for stitching video clips, we present a new spatial-temporal descriptor to provide an effective representation for different types of dynamic textures. In the second stage, a smoothing method is proposed to improve synthesis quality, especially in the aspect of temporal continuity. It aims to establish a diffeomorphic growth model to emulate local dynamics around stitched frames. The proposed approach is thoroughly tested on public databases and videos from the Internet, and is evaluated in both qualitative and quantitative ways. Yimo Guo, Guoying Zhao 0001, Ziheng Zhou 0003, Matti Pietikäinen |
IEEE Trans. Image Process. | 4 |
| 2012 | Efficient Image Appearance Description Using Dense Sampling Based Local Binary Patterns
Juha Ylioinas, Abdenour Hadid, Yimo Guo, Matti Pietikäinen |
ACCV (3) | 4 |
| 2012 | Dynamic Facial Expression Recognition Using Longitudinal Facial Expression Atlases
Yimo Guo, Guoying Zhao 0001, Matti Pietikäinen |
ECCV (2) | 3 |
| 2012 | Unsupervised dynamic texture segmentation using local descriptors in volumes
Jie Chen 0001, Guoying Zhao 0001, Matti Pietikäinen |
ICPR | 3 |
| 2012 | Can gait biometrics be Spoofed?
Abdenour Hadid, Mohammad Ghahramani, Vili Kellokumpu, Matti Pietikäinen, John D. Bustard, Mark S. Nixon |
ICPR | 4 |
| 2012 | Combining local and global correlation for texture description
Xiaopeng Hong, Guoying Zhao 0001, Matti Pietikäinen, Xilin Chen 0001 |
ICPR | 3 |
| 2012 | CMV100: A dataset for people tracking and re-identification in sparse camera networks
Valtteri Takala, Matti Pietikäinen |
ICPR | 2 |
| 2012 | Age Classification in Unconstrained Conditions Using LBP Variants
Juha Ylioinas, Abdenour Hadid, Matti Pietikäinen |
ICPR | 3 |
| 2012 | Discriminative features for texture description
Yimo Guo, Guoying Zhao 0001, Matti Pietikäinen |
Pattern Recognit. | 3 |
| 2012 | Towards a dynamic expression recognition system under facial occlusion
Xiaohua Huang 0003, Guoying Zhao 0001, Wenming Zheng, Matti Pietikäinen |
Pattern Recognit. Lett. | 4 |
| 2012 | Spatiotemporal Local Monogenic Binary Patterns for Facial Expression RecognitionabstractFeature representation is an important research topic in facial expression recognition from video sequences. In this letter, we propose to use spatiotemporal monogenic binary patterns to describe both appearance and motion information of the dynamic sequences. Firstly, we use monogenic signals analysis to extract the magnitude, the real picture and the imaginary picture of the orientation of each frame, since the magnitude can provide much appearance information and the orientation can provide complementary information. Secondly, the phase-quadrant encoding method and the local bit exclusive operator are utilized to encode the real and imaginary pictures from orientation in three orthogonal planes, and the local binary pattern operator is used to capture the texture and motion information from the magnitude through three orthogonal planes. Finally, both concatenation method and multiple kernel learning method are respectively exploited to handle the feature fusion. The experimental results on the Extended Cohn-Kanade and Oulu-CASIA facial expression databases demonstrate that the proposed methods perform better than the state-of-the-art methods, and are robust to illumination variations. Xiaohua Huang 0003, Guoying Zhao 0001, Wenming Zheng, Matti Pietikäinen |
IEEE Signal Process. Lett. | 4 |
| 2012 | An Image-Based Visual Speech Animation SystemabstractAn image-based visual speech animation system is presented in this paper. A video model is proposed to preserve the video dynamics of a talking face. The model represents a video sequence by a low-dimensional continuous curve embedded in a path graph and establishes a map from the curve to the image domain. When selecting video segments for synthesis, we loosen the traditional requirement of using triphone as the unit to allow segments to contain longer natural talking motion. Dense videos are sampled from the segments, concatenated, and downsampled to train a video model that enables efficient time alignment and motion smoothing for the final video synthesis. Different viseme definitions are used to investigate the impact of visemes on the video realism of the animated talking face. The system is built on a public database and tested both objectively and subjectively. Ziheng Zhou 0003, Guoying Zhao 0001, Yimo Guo, Matti Pietikäinen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | Rotation-Invariant Image and Video Description With Local Binary Pattern FeaturesabstractIn this paper, we propose a novel approach to compute rotation-invariant features from histograms of local noninvariant patterns. We apply this approach to both static and dynamic local binary pattern (LBP) descriptors. For static-texture description, we present LBP histogram Fourier (LBP-HF) features, and for dynamic-texture recognition, we present two rotation-invariant descriptors computed from the LBPs from three orthogonal planes (LBP-TOP) features in the spatiotemporal domain. LBP-HF is a novel rotation-invariant image descriptor computed from discrete Fourier transforms of LBP histograms. The approach can be also generalized to embed any uniform features into this framework, and combining the supplementary information, e.g., sign and magnitude components of the LBP, together can improve the description ability. Moreover, two variants of rotation-invariant descriptors are proposed to the LBP-TOP, which is an effective descriptor for dynamic-texture recognition, as shown by its recent success in different application problems, but it is not rotation invariant. In the experiments, it is shown that the LBP-HF and its extensions outperform noninvariant and earlier versions of the rotation-invariant LBP in the rotation-invariant texture classification. In experiments on two dynamic-texture databases with rotations or view variations, the proposed video features can effectively deal with rotation variations of dynamic textures (DTs). They also are robust with respect to changes in viewpoint, outperforming recent methods proposed for view-invariant recognition of DTs. Guoying Zhao 0001, Timo Ahonen, Jiri Matas, Matti Pietikäinen |
IEEE Trans. Image Process. | 4 |
| 2011 | Texture Classification using a Linear Configuration Model based DescriptorabstractWe investigate rotation invariant image description and develop a linear model based descriptor namely MiC, which is suited to modeling microscopic configuration of images. To explore multi-channel discriminative information of both the microscopic configuration and local structures, the feature extraction process is formulated as an unsupervised framework that consists of: 1) the configuration model to encode image microscopic configuration; and 2) local patterns to describe local structural information. In this way, images are represented by a novel feature: local configuration pattern (LCP). We evaluate the performance of the proposed method by classifying textures present in three challenging texture databases: Outex_TC_00012, KTH-TIPS2 and Columbia-Utrecht (CUReT). The encouraging results show that LCPs is highly discriminative. 1 Yimo Guo, Guoying Zhao 0001, Matti Pietikäinen |
BMVC | 3 |
| 2011 | Towards a practical lipreading systemabstractA practical lipreading system can be considered either as subject dependent (SD) or subject-independent (SI). An SD system is user-specific, i.e., customized for some particular user while an SI system has to cope with a large number of users. These two types of systems pose variant challenges and have to be treated differently. In this paper, we propose a simple deterministic model to tackle the problem. The model first seeks a low-dimensional manifold where visual features extracted from the frames of a video can be projected onto a continuous deterministic curve embedded in a path graph. Moreover, it can map arbitrary points on the curve back into the image space, making it suitable for temporal interpolation. Based on the model, we develop two separate strategies for SD and SI lipreading. The former is turned into a simple curve-matching problem while for the latter, we propose a video-normalization scheme to improve the system developed by Zhao et al. We evaluated our system on the OuluVS database and achieved recognition rates more than 20% higher than the ones reported by Zhao et al. in both SD and SI testing scenarios. Ziheng Zhou 0003, Guoying Zhao 0001, Matti Pietikäinen |
CVPR | 3 |
| 2011 | Local frequency descriptor for low-resolution face recognitionabstractFace recognition from low-resolution images is a common yet challenging case in real applications. Since the high-frequency information is lost in low-resolution images, it is necessary to explore robust information in the low frequency domain. In this paper, we propose an effective local frequency descriptor (LFD) for low resolution face recognition, by building upon the ideas behind local phase quantization (LPQ) and exploring both blur-invariant magnitude and phase information in the low frequency domain. The proposed descriptor is more descriptive than LPQ with more comprehensive information. In addition, a statistical uniform pattern definition method is introduced to improve the efficiency of the proposed descriptor. Experimental results on FERET and a real video database show that LFD is effective and robust for low-resolution face recognition. Zhen Lei 0001, Timo Ahonen, Matti Pietikäinen, Stan Z. Li |
FG | 3 |
| 2011 | Competition on counter measures to 2-D facial spoofing attacksabstractSpoofing identities using photographs is one of the most common techniques to attack 2-D face recognition systems. There seems to exist no comparative studies of different techniques using the same protocols and data. The motivation behind this competition is to compare the performance of different state-of-the-art algorithms on the same database using a unique evaluation method. Six different teams from universities around the world have participated in the contest. Use of one or multiple techniques from motion, texture analysis and liveness detection appears to be the common trend in this competition. Most of the algorithms are able to clearly separate spoof attempts from real accesses. The results suggest the investigation of more complex attacks. Murali Mohan Chakka, André Anjos, Sébastien Marcel, Roberto Tronci, Daniele Muntoni, Gianluca Fadda, Maurizio Pili, Nicola Sirena, Gabriele Murgia, Marco Ristori, Fabio Roli, Dong Yi, Zhen Lei 0001, Stan Z. Li, William Robson Schwartz, Anderson Rocha 0001, Hélio Pedrini, Javier Lorenzo-Navarro, Modesto Castrillón-Santana, Jukka Komulainen, Abdenour Hadid, Matti Pietikäinen |
IJCB | 24 |
| 2011 | Face spoofing detection from single images using micro-texture analysisabstractCurrent face biometric systems are vulnerable to spoo ing attacks. A spoofing attack occurs when a person tries to masquerade as someone else by falsifying data and thereby gaining illegitimate access. Inspired by image quality assessment, characterization of printing artifacts, and differences in light reflection, we propose to approach the problem of spoofing detection from texture analysis point of view. Indeed, face prints usually contain printing quality defects that can be well detected using texture features. Hence, we present a novel approach based on analyzing facial image textures for detecting whether there is a live person in front of the camera or a face print. The proposed approach analyzes the texture of the facial images using multi-scale local binary patterns (LBP). Compared to many previous works, our proposed approach is robust, computationally fast and does not require user-cooperation. In addition, the texture features that are used for spoofing detection can also be used for face recognition. This provides a unique feature space for coupling spoofing detection and face recognition. Extensive experimental analysis on a publicly avail able database showed excellent results compared to existing works. Jukka Komulainen, Abdenour Hadid, Matti Pietikäinen |
IJCB | 3 |
| 2011 | Recognising spontaneous facial micro-expressionsabstractFacial micro-expressions are rapid involuntary facial expressions which reveal suppressed affect. To the best knowledge of the authors, there is no previous work that successfully recognises spontaneous facial micro-expressions. In this paper we show how a temporal interpolation model together with the first comprehensive spontaneous micro-expression corpus enable us to accurately recognise these very short expressions. We designed an induced emotion suppression experiment to collect the new corpus using a high-speed camera. The system is the first to recognise spontaneous facial micro-expressions and achieves very promising results that compare favourably with the human micro-expression detection accuracy. Tomas Pfister, Guoying Zhao 0001, Matti Pietikäinen |
ICCV | 4 |
| 2011 | Facial expression recognition from near-infrared videos
Guoying Zhao 0001, Xiaohua Huang 0003, Matti Taini, Stan Z. Li, Matti Pietikäinen |
Image Vis. Comput. | 5 |
| 2011 | Recognition of human actions using texture descriptors
Vili Kellokumpu, Guoying Zhao 0001, Matti Pietikäinen |
Mach. Vis. Appl. | 3 |
| 2011 | Face Recognition by Exploring Information Jointly in Space, Scale and OrientationabstractInformation jointly contained in image space, scale and orientation domains can provide rich important clues not seen in either individual of these domains. The position, spatial frequency and orientation selectivity properties are believed to have an important role in visual perception. This paper proposes a novel face representation and recognition approach by exploring information jointly in image space, scale and orientation domains. Specifically, the face image is first decomposed into different scale and orientation responses by convolving multiscale and multiorientation Gabor filters. Second, local binary pattern analysis is used to describe the neighboring relationship not only in image space, but also in different scale and orientation responses. This way, information from different domains is explored to give a good face representation for recognition. Discriminant classification is then performed based upon weighted histogram intersection or conditional mutual information with linear discriminant analysis techniques. Extensive experimental results on FERET, AR, and FRGC ver 2.0 databases show the significant advantages of the proposed method over the existing ones. Zhen Lei 0001, Shengcai Liao, Matti Pietikäinen, Stan Z. Li |
IEEE Trans. Image Process. | 3 |
| 2010 | Descriptor Learning Based on Fisher Separation Criterion for Texture Classification
Yimo Guo, Guoying Zhao 0001, Matti Pietikäinen, Zhengguang Xu |
ACCV (3) | 3 |
| 2010 | Dynamic Facial Expression Recognition Using Boosted Component-Based Spatiotemporal Features and Multi-classifier Fusion
Xiaohua Huang 0003, Guoying Zhao 0001, Matti Pietikäinen, Wenming Zheng |
ACIVS (2) | 3 |
| 2010 | Modeling pixel process with scale invariant local patterns for background subtraction in complex scenesabstractBackground modeling plays an important role in video surveillance, yet in complex scenes it is still a challenging problem. Among many difficulties, problems caused by illumination variations and dynamic backgrounds are the key aspects. In this work, we develop an efficient background subtraction framework to tackle these problems. First, we propose a scale invariant local ternary pattern operator, and show that it is effective for handling illumination variations, especially for moving soft shadows. Second, we propose a pattern kernel density estimation technique to effectively model the probability distribution of local patterns in the pixel process, which utilizes only one single LBP-like pattern instead of histogram as feature. Third, we develop multimodal background models with the above techniques and a multiscale fusion scheme for handling complex dynamic backgrounds. Exhaustive experimental evaluations on complex scenes show that the proposed method is fast and effective, achieving more than 10% improvement in accuracy compared over existing state-of-the-art algorithms. Shengcai Liao, Guoying Zhao 0001, Vili Kellokumpu, Matti Pietikäinen, Stan Z. Li |
CVPR | 4 |
| 2010 | Spatial-Temporal Granularity-Tunable Gradients Partition (STGGP) Descriptors for Human Detection
Yazhou Liu, Shiguang Shan, Xilin Chen 0001, Janne Heikkilä, Wen Gao 0001, Matti Pietikäinen |
ECCV (1) | 6 |
| 2010 | Recovering the Topology of Multiple Cameras by Finding Continuous Paths in a TrellisabstractIn this paper, we propose an unsupervised method for recovering the topology of multiple cameras with non-overlapping fields of view. The nodes in the topology graph are defined as entry/exit zones in each camera while the connectivity between nodes is inferred through finding continuous paths in a trellis where appearance information and temporal information of moving objects are encoded. Unlike previous methods which assume a single mode transition distribution between nodes, our method is capable of dealing with multi-modal transition situations when both cars and pedestrians are in the scene. Results on simulated and real-life datasets demonstrate the effectiveness of the proposed method. Yinghao Cai, Kaiqi Huang, Tieniu Tan, Matti Pietikäinen |
ICPR | 4 |
| 2010 | Matching Groups of People by Covariance DescriptorabstractIn this paper, we present a new solution to the problem of matching groups of people across multiple non-overlapping cameras. Similar to the problem of matching individuals across cameras, matching groups of people also faces challenges such as variations of illumination conditions, poses and camera parameters. Moreover, people often swap their positions while walking in a group. In this paper, we propose to use covariance descriptor in appearance matching of group images. Covariance descriptor is shown to be a discriminative descriptor which captures both appearance and statistical properties of image regions. Furthermore, it presents a natural way of combining multiple heterogeneous features together with a relatively low dimensionality. Experimental results on two different datasets demonstrate the effectiveness of the proposed method. Yinghao Cai, Valtteri Takala, Matti Pietikäinen |
ICPR | 3 |
| 2010 | Boosting Clusters of Samples for Sequence Matching in Camera NetworksabstractThis study introduces a novel classification algorithm for learning and matching sequences in view independent object tracking. The proposed learning method uses adaptive boosting and classification trees on a wide collection (shape, pose, color, texture, etc.) of image features that constitute a model for tracked objects. The temporal dimension is taken into account by using k-mean clusters of sequence samples. Most of the utilized object descriptors have a temporal quality also. We argue that with a proper boosting approach and decent number of reasonably descriptive image features it is feasible to do view-independent sequence matching in sparse camera networks. The experiments on real-life surveillance data support this statement. Valtteri Takala, Yinghao Cai, Matti Pietikäinen |
ICPR | 3 |
| 2010 | Lipreading: A Graph Embedding ApproachabstractIn this paper, we propose a novel graph embedding method for the problem of lipreading. To characterize the temporal connections among video frames of the same utterance, a new distance metric is defined on a pair of frames and graphs are constructed to represent the video dynamics based on the distances between frames. Audio information is used to assist in calculating such distances. For each utterance, a subspace of the visual feature space is learned from a well-defined intrinsic and penalty graph within a graph-embedding framework. Video dynamics are found to be well preserved along some dimensions of the subspace. Discriminatory cues are then decoded from curves of the projected visual features to classify different utterances. Ziheng Zhou 0003, Guoying Zhao 0001, Matti Pietikäinen |
ICPR | 3 |
| 2010 | WLD: A Robust Local Image DescriptorabstractInspired by Weber's Law, this paper proposes a simple, yet very powerful and robust local descriptor, called the Weber Local Descriptor (WLD). It is based on the fact that human perception of a pattern depends not only on the change of a stimulus (such as sound, lighting) but also on the original intensity of the stimulus. Specifically, WLD consists of two components: differential excitation and orientation. The differential excitation component is a function of the ratio between two terms: One is the relative intensity differences of a current pixel against its neighbors, the other is the intensity of the current pixel. The orientation component is the gradient orientation of the current pixel. For a given image, we use the two components to construct a concatenated WLD histogram. Experimental results on the Brodatz and KTH-TIPS2-a texture databases show that WLD impressively outperforms the other widely used descriptors (e.g., Gabor and SIFT). In addition, experimental results on human face detection also show a promising performance comparable to the best known results on the MIT+CMU frontal face test set, the AR face data set, and the CMU profile test set. Jie Chen 0001, Shiguang Shan, Chu He, Guoying Zhao 0001, Matti Pietikäinen, Xilin Chen 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2009 | A New Gabor Phase Difference Pattern for Face and Ear Recognition
Yimo Guo, Guoying Zhao 0001, Jie Chen 0001, Matti Pietikäinen, Zhengguang Xu |
CAIP | 4 |
| 2009 | Learning mappings for face synthesis from near infrared to visual light imagesabstractThis paper deals with a new problem in face recognition research, in which the enrollment and query face samples are captured under different lighting conditions. In our case, the enrollment samples are visual light (VIS) images, whereas the query samples are taken under near infrared (NIR) condition. It is very difficult to directly match the face samples captured under these two lighting conditions due to their different visual appearances. In this paper, we propose a novel method for synthesizing VIS images from NIR images based on learning the mappings between images of different spectra (i.e., NIR and VIS). In our approach, we reduce the inter-spectral differences significantly, thus allowing effective matching between faces taken under different imaging conditions. Face recognition experiments clearly show the efficacy of the proposed approach. Jie Chen 0001, Dong Yi, Jimei Yang, Guoying Zhao 0001, Stan Z. Li, Matti Pietikäinen |
CVPR | 6 |
| 2009 | Dynamic texture synthesis using a spatial temporal descriptorabstractDynamic textures are image sequences with visual pattern repetition in time and space, such as smoke, flames, moving objects and so on. Dynamic texture synthesis is to provide a continuous and infinitely varying stream of images by doing operations on dynamic textures. Considering that the previous video texture method provides high-quality visual results, but its representation does not well explore the temporal correlation among frames, we develop a novel spatial temporal descriptor for frame description accompanied with a similarity measure on the basis of the video texture method. Compared with the previous one, our method considers both the spatial and temporal domains of video sequences in representation; moreover, combines the local and global description on each spatial-temporal plane. From experimental results, the proposed method achieves better performance in both the syntheses of natural scene and human motion. Especially, it has the characteristic to be robust to noise in remodeling videos into infinite time domain. Yimo Guo, Guoying Zhao 0001, Jie Chen 0001, Matti Pietikäinen, Zhengguang Xu |
ICIP | 4 |
| 2009 | Combining appearance and motion for face and gender recognition from videos
Abdenour Hadid, Matti Pietikäinen |
Pattern Recognit. | 2 |
| 2009 | Description of interest regions with local binary patterns
Marko Heikkilä, Matti Pietikäinen, Cordelia Schmid |
Pattern Recognit. | 2 |
| 2009 | Image description using joint distribution of filter bank responses
Timo Ahonen, Matti Pietikäinen |
Pattern Recognit. Lett. | 2 |
| 2009 | Boosted multi-resolution spatiotemporal descriptors for facial expression recognition
Guoying Zhao 0001, Matti Pietikäinen |
Pattern Recognit. Lett. | 2 |
| 2009 | Lipreading With Local Spatiotemporal DescriptorsabstractVisual speech information plays an important role in lipreading under noisy conditions or for listeners with a hearing impairment. In this paper, we present local spatiotemporal descriptors to represent and recognize spoken isolated phrases based solely on visual input. Spatiotemporal local binary patterns extracted from mouth regions are used for describing isolated phrase sequences. In our experiments with 817 sequences from ten phrases and 20 speakers, promising accuracies of 62% and 70% were obtained in speaker-independent and speaker-dependent recognition, respectively. In comparison with other methods on AVLetters database, the accuracy, 62.8%, of our method clearly outperforms the others. Analysis of the confusion matrix for 26 English letters shows the good clustering characteristics of visemes for the proposed descriptors. The advantages of our approach include local processing and robustness to monotonic gray-scale changes. Moreover, no error prone segmentation of moving lips is needed. Guoying Zhao 0001, Mark Barnard, Matti Pietikäinen |
IEEE Trans. Multim. | 3 |
| 2008 | Human Activity Recognition Using a Dynamic Texture Based MethodabstractWe present a novel approach for human activity reco gnition. The method uses dynamic texture descriptors to describe human movements in a spatiotemporal way. The same features are also use d for human detection, which makes our whole approach computationally simple. Following recent trends in computer vision research , our method works on image data rather than silhouettes. We test our met hod on a publicly available dataset and compare our result to the sta te of the art methods. Vili Kellokumpu, Guoying Zhao 0001, Matti Pietikäinen |
BMVC | 3 |
| 2008 | A robust descriptor based on Weber's LawabstractInspired by Weber’s Law, this paper proposes a simple, yet very powerful and robust local descriptor, Weber Local Descriptor (WLD). It is based on the fact that human perception of a pattern depends on not only the change of a stimulus (such as sound, lighting, et al.) but also the original intensity of the stimulus. Specifically, WLD consists of two components: its differential excitation and orientation. A differential excitation is a function of the ratio between two terms: One is the relative intensity differences of its neighbors against a current pixel; the other is the intensity of the current pixel. An orientation is the gradient orientation of the current pixel. For a given image, we use the differential excitation and the orientation components to construct a concatenated WLD histogram feature. Experimental results on Brodatz textures show that WLD impressively outperforms the other classical descriptors (e.g., Gabor). Especially, experimental results on face detection show a promising performance. Although we train only one classifier based on WLD features, the classifier obtains a comparable performance to state-of-the-art methods on MIT+CMU frontal face test set, AR face dataset and CMU profile test set. Jie Chen 0001, Shiguang Shan, Guoying Zhao 0001, Xilin Chen 0001, Wen Gao 0001, Matti Pietikäinen |
CVPR | 6 |
| 2008 | Gabor volume based local binary pattern for face representation and recognitionabstractThis paper presents a novel face representation and recognition approach. The face image is first decomposed by multi-scale and multi-orientation Gabor filters and local binary pattern (LBP) analysis is then applied on the derived Gabor magnitude responses. Different from (W.C. Zhang et al., 2005), the present method not only describes the neighboring relationship in spatial domain, but also exploit those between different scales (frequency) and orientations. Specifically, we first reformulate the Gabor magnitude responses as a 3rd-order volume and then apply LBP analysis on three orthogonal planes of the Gabor volume, named GV-LBP-TOP in short, in a hope to encode sufficient information for face representation. Further, a computationally effective version, E-GV-LBP, is proposed to depict the neighboring changes in spatial, frequency and orientation domains simultaneously. Finally, the weighted histogram intersection metric is utilized to measure the dissimilarity of faces. Experimental results on FERET and FRGC ver 2.0 databases show the significant advantages of the proposed method. Zhen Lei 0001, Shengcai Liao, Ran He 0001, Matti Pietikäinen, Stan Z. Li |
FG | 4 |
| 2008 | Unsupervised dynamic texture segmentation using local spatiotemporal descriptorsabstractDynamic texture (DT) is an extension of texture to the temporal domain. In this paper, we address the problem of segmenting DT into disjoint regions in an unsupervised way. Each region is characterized by histograms of local binary patterns and contrast in a spatiotemporal mode. It combines the motion and appearance of DT together. Experimental results show that our method is effective in segmenting regions that differ in their dynamics. Jie Chen 0001, Guoying Zhao 0001, Matti Pietikäinen |
ICPR | 3 |
| 2008 | Combining motion and appearance for gender classification from video sequencesabstractWe investigate whether combining appearance (face structure) and motion (the way a person is talking and moving his/her facial features) boosts gender classification from face sequences. We propose and compare different schemes based on appearance only, motion only, and combination of appearance and motion. Experiments on various face video datasets of persons uttering phrases or expressing emotions show that combination of motion and appearance is useful for gender analysis of familiar faces, yielding in classification accuracy of 100%. However, for unfamiliar faces, motion seems to not provide additional discriminative information as the best performance (96.3%) is obtained using an appearance based approach with Local Binary Pattern (LBP) features and Support Vector Machines (SVMs). Abdenour Hadid, Matti Pietikäinen |
ICPR | 2 |
| 2008 | A Bayesian Local Binary Pattern texture descriptorabstractIn this paper, a Bayesian LBP operator is proposed. This operator is formulated in a novel Filtering, Labeling and Statistic (FLS) framework for texture descriptors. In the framework, the local labeling procedure, which is a part of many popular descriptors such as LBP, SIFT and VZ, can be modeled as a probability and optimization process. This enables the use of more reliable prior and likelihood information and reduces the sensitivity to noise. The BLBP operator pursues a label image, when given the filtered vector image, by maximizing the joint probability of two images under the criterion of MAP. The proposed approach is evaluated on texture retrieval schemes using entire Brodatz database. The result reveals BLBP operator’s efficient performance and FLS framework’s capability to in-depth analysis of the texture descriptors on a common background. Chu He, Timo Ahonen, Matti Pietikäinen |
ICPR | 3 |
| 2008 | Facial expression recognition from near-infrared video sequencesabstractFacial expressions can be thought as specific dynamic textures where local appearance and motion information need to be taken into account. We utilize local spatiotemporal operators to describe facial expressions. All current facial expression recognition databases are captured in visible light spectrum. Visual light usually changes with locations, and can also vary with time, which can cause significant variations in image appearance and texture. In this paper, we present a novel research on a dynamic facial expression recognition from near-infrared (NIR) video sequences. NIR imaging is robust with respect to illumination changes. Experiments on a new NIR database show promising and robust results against illumination variations. Matti Taini, Guoying Zhao 0001, Stan Z. Li, Matti Pietikäinen |
ICPR | 4 |
| 2007 | Multi-Object Tracking Using Color, Texture and MotionabstractIn this paper, we introduce a novel real-time tracker based on color, texture and motion information. RGB color histogram and correlogram (autocorrelogram) are exploited as color cues and texture properties are represented by local binary patterns (LBP). Object's motion is taken into account through location and trajectory. After extraction, these features are used to build a unifying distance measure. The measure is utilized in tracking and in the classification event, in which an object is leaving a group. The initial object detection is done by a texture-based background subtraction algorithm. The experiments on indoor and outdoor surveillance videos show that a unified system works better than the versions based on single features. It also copes well with low illumination conditions and low frame rates which are common in large scale surveillance systems. Valtteri Takala, Matti Pietikäinen |
CVPR | 2 |
| 2007 | Experiments with Facial Expression Recognition using Spatiotemporal Local Binary PatternsabstractIn this paper, the recently introduced method for facial expression recognition using spatiotemporal local binary patterns is reviewed and experiments are carried out to investigate the robustness of the approach. In experiments with the Cohn-Kanade facial expression database, our results from the cross-validation with low resolutions and low frame rates are promising. Advantages of our approach include local processing, robustness to low quality of videos and simple computation. Guoying Zhao 0001, Matti Pietikäinen |
ICME | 2 |
| 2007 | Dynamic Texture Recognition Using Local Binary Patterns with an Application to Facial ExpressionsabstractDynamic texture (DT) is an extension of texture to the temporal domain. Description and recognition of DTs have attracted growing attention. In this paper, a novel approach for recognizing DTs is proposed and its simplifications and extensions to facial image analysis are also considered. First, the textures are modeled with volume local binary patterns (VLBP), which are an extension of the LBP operator widely used in ordinary texture analysis, combining motion and appearance. To make the approach computationally simple and easy to extend, only the co-occurrences of the local binary patterns on three orthogonal planes (LBP-TOP) are then considered. A block-based method is also proposed to deal with specific dynamic events such as facial expressions in which local information and its spatial locations should also be taken into account. In experiments with two DT databases, DynTex and Massachusetts Institute of Technology (MIT), both the VLBP and LBP-TOP clearly outperformed the earlier approaches. The proposed block-based method was evaluated with the Cohn-Kanade facial expression database with excellent results. The advantages of our approach include local processing, robustness to monotonic gray-scale changes, and simple computation. Guoying Zhao 0001, Matti Pietikäinen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Contextual Analysis of Textured Scene ImagesabstractClassifying image regions into one of several pre-defined semantic categories is a typical image understanding problem. Different image regions and object types might have very similar color or texture characteristics making it difficult to categorize them. Without contextual information it is often impossible to find reasonable semantic labeling for outdoor images. In this paper, we combine an efficient SVM-based local classifier with the conditional random field framework to incorporate spatial contex information to the classification. The images are represented with powerful local texture features. Then a discriminative multiclass model for finding good labeling for the image is learned. The performance of the method was evaluated with two different datasets. The approach was also shown to be useful in more general image retrieval and annotation tasks based on classification. 1 Markus Turtinen, Matti Pietikäinen |
BMVC | 2 |
| 2006 | Eyes Location Using a Neural Network
Xiaoyi Feng, Li-ping Yang, Zhi Dang, Matti Pietikäinen |
ISNN (2) | 4 |
| 2006 | Face Description with Local Binary Patterns: Application to Face RecognitionabstractThis paper presents a novel and efficient facial image representation based on local binary pattern (LBP) texture features. The face image is divided into several regions from which the LBP feature distributions are extracted and concatenated into an enhanced feature vector to be used as a face descriptor. The performance of the proposed method is assessed in the face recognition problem under different challenges. Other applications and several extensions are also discussed. Timo Ahonen, Abdenour Hadid, Matti Pietikäinen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2006 | A Texture-Based Method for Modeling the Background and Detecting Moving ObjectsabstractThis paper presents a novel and efficient texture-based method for modeling the background and detecting moving objects from a video sequence. Each pixel is modeled as a group of adaptive local binary pattern histograms that are calculated over a circular region around the pixel. The approach provides us with many advantages compared to the state-of-the-art. Experimental results clearly justify our model. Marko Heikkilä, Matti Pietikäinen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | A Novel Real Time System for Facial Expression Recognition
Xiaoyi Feng, Matti Pietikäinen, Abdenour Hadid, Hongmei Xie |
ACII | 2 |
| 2005 | Incremental locally linear embedding
Olga Kouropteva, Oleg Okun, Matti Pietikäinen |
Pattern Recognit. | 3 |
| 2004 | A Texture-based Method for Detecting Moving ObjectsabstractThe detection of moving objects from video frames plays an important and often very critical role in different kinds of machine vision applications including human detection and tracking, traffic monitoring, humanmachine interfaces and military applications, since it usually is one of the first phases in a system architecture. A common way to detect moving objects is background subtraction. In background subtraction, moving objects are detected by comparing each video frame against an existing model of the scene background. In this paper, we propose a novel block-based algorithm for background subtraction. The algorithm is based on the Local Binary Pattern (LBP) texture measure. Each image block is modelled as a group of weighted adaptive LBP histograms. The algorithm operates in real-time under the assumption of a stationary camera with fixed focal length. It can adapt to inherent changes in scene background and can also handle multimodal backgrounds. Marko Heikkilä, Matti Pietikäinen, Janne Heikkilä |
BMVC | 2 |
| 2004 | A Discriminative Feature Space for Detecting and Recognizing Faces
Abdenour Hadid, Matti Pietikäinen, Timo Ahonen |
CVPR (2) | 2 |
| 2004 | Face Recognition with Local Binary Patterns
Timo Ahonen, Abdenour Hadid, Matti Pietikäinen |
ECCV (1) | 3 |
| 2004 | Classification with color and texture: jointly or separately?
Topi Mäenpää, Matti Pietikäinen |
Pattern Recognit. | 2 |
| 2004 | View-based recognition of real-world textures
Matti Pietikäinen, Tomi Nurmela, Topi Mäenpää, Markus Turtinen |
Pattern Recognit. | 1 |
| 2003 | Classification of handwritten digits using supervised locally linear embedding algorithm and support vector machine
Olga Kouropteva, Oleg Okun, Matti Pietikäinen |
ESANN | 3 |
| 2003 | Supervised Locally Linear Embedding
Dick de Ridder, Olga Kouropteva, Oleg Okun, Matti Pietikäinen, Robert P. W. Duin |
ICANN | 4 |
| 2003 | Optimising Colour and Texture Features for Real-time Visual Inspection
Topi Mäenpää, Jaakko Viertola, Matti Pietikäinen |
Pattern Anal. Appl. | 3 |
| 2002 | Multiresolution Gray-Scale and Rotation Invariant Texture Classification with Local Binary PatternsabstractPresents a theoretically very simple, yet efficient, multiresolution approach to gray-scale and rotation invariant texture classification based on local binary patterns and nonparametric discrimination of sample and prototype distributions. The method is based on recognizing that certain local binary patterns, termed "uniform," are fundamental properties of local image texture and their occurrence histogram is proven to be a very powerful texture feature. We derive a generalized gray-scale and rotation invariant operator presentation that allows for detecting the "uniform" patterns for any quantization of the angular space and for any spatial resolution and presents a method for combining multiple operators for multiresolution analysis. The proposed approach is very robust in terms of gray-scale variations since the operator is, by definition, invariant against any monotonic transformation of the gray scale. Another advantage is computational simplicity as the operator can be realized with a few operations in a small neighborhood and a lookup table. Experimental results demonstrate that good discrimination can be achieved with the occurrence statistics of simple rotation invariant local binary patterns. Timo Ojala, Matti Pietikäinen, Topi Mäenpää |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Edge-Based Method for Text Detection from Complex Document ImagesabstractDetection of text from documents in which text is embedded in complex colored and textured backgrounds is a very challenging problem. In this paper, we propose a simple texture-based approach based on edge information for this task. The performance of our method is compared to that obtained by a method based on the discrete cosine transform which was recently proposed by Y. Zhong et al. (2000) for text localization in compressed digital video. In our experiments, both methods performed about equally well for small-sized text, but our method was better in the case of large-sized text. The principal advantage of our approach is that in addition to the text detection problem, the same edge representation can also be used for other image interpretation tasks. Matti Pietikäinen, Oleg Okun |
ICDAR | 1 |
| 2001 | Texture discrimination with multidimensional distributions of signed gray-level differences
Timo Ojala, Kimmo Valkealahti, Erkki Oja, Matti Pietikäinen |
Pattern Recognit. | 4 |
| 2000 | Gray Scale and Rotation Invariant Texture Classification with Local Binary Patterns
Timo Ojala, Matti Pietikäinen, Topi Mäenpää |
ECCV (1) | 2 |
| 2000 | Robust Texture Classification by Subsets of Local Binary PatternsabstractRecently, a nonparametric approach to texture analysis has been developed, in which the distributions of simple texture measures based on local binary patterns (LBP) are used for texture description. The basic LBP encodes 256 simple feature detectors in a single 3/spl times/3 operator. This paper shows that a properly selected subset of patterns encoded in LBP forms an efficient and robust texture description which can achieve better classification rates in comparison with the whole LBP histogram. Experiments on classification of textures from the Columbia-Utrecht (CURET) database demonstrate the robustness of the approach. Topi Mäenpää, Timo Ojala, Matti Pietikäinen, Maricor Soriano |
ICPR | 3 |
| 2000 | Texture Classification by Multi-Predicate Local Binary Pattern OperatorsabstractThis paper benchmarks the simple local binary pattern (LBP) approach in the supervised texture segmentation problems of the recent comparative study of Randen and Husoy (1999). A multi-predicate version of LBP is proposed, making our approach even more powerful for images containing textures at multiple scales. Topi Mäenpää, Matti Pietikäinen, Timo Ojala |
ICPR | 2 |
| 2000 | Automatic Ground-Truth Generation for Skew-Tolerance Evaluation of Document Layout Analysis MethodsabstractGeneration of ground-truths is of great importance for unbiased performance evaluation of document layout analysis methods. This is especially necessary because many methods are claimed to be skew-tolerant. However, experimental evaluation of this fact is often based only on human subjective judgement and restricted to a few experiments. The main obstacle for obtaining human-independent and more automated performance evaluation is that usually there are only ground-truths for upright images, i.e., images with no skew of text lines, because currently available ground-truthing techniques are too time-consuming. We propose a methodology of automatic generation of ground-truths for skewed images by using the ground-truths available for upright images. This methodology is simple and quite fast because processing is done at the level of small square blocks, but not at pixel level. Oleg Okun, Matti Pietikäinen |
ICPR | 2 |
| 2000 | Rotation-invariant texture classification using feature distributions
Matti Pietikäinen, Timo Ojala |
Pattern Recognit. | 1 |
| 2000 | Adaptive document image binarization
Jaakko J. Sauvola, Matti Pietikäinen |
Pattern Recognit. | 2 |
| 1999 | Robust Skew Estimation on Low-Resolution Document ImagesabstractPresents a new technique for estimating an arbitrary skew angle using a simple and efficient text row accumulation based on first- and second-order statistics. In contrast to other works, where an image resolution of 100-300 dpi is necessary, a lower resolution (50 dpi) is enough for our method to obtain a correct result. The performance of the method was evaluated on the database UW-1 of the University of Washington and compared with an advanced Hough transform-based method. A significantly lower error rate for our method confirms its advantages and benefits for automated document processing. Oleg Okun, Matti Pietikäinen, Jaakko J. Sauvola |
ICDAR | 2 |
| 1999 | Document skew estimation without angle range restriction
Oleg Okun, Matti Pietikäinen, Jaakko J. Sauvola |
Int. J. Document Anal. Recognit. | 2 |
| 1999 | Unsupervised texture segmentation using feature distributions
Timo Ojala, Matti Pietikäinen |
Pattern Recognit. | 2 |
| 1998 | Nonparametric multichannel texture description with simple spatial operatorsabstractA multichannel approach to texture description is proposed by approximating joint occurrences of multiple features with marginal distributions, as 1D histograms, and combining similarity scores for 1D histograms into an aggregate similarity score. A stepwise feature selection algorithm is used to choose the best feature combination in a particular dimension. In classification experiments with Brodatz textures and MeasTex test suites the proposed method performs favorably compared to GLCM, Gabor and GMRF features. Timo Ojala, Matti Pietikäinen |
ICPR | 2 |
| 1997 | A distributed management system for testing document image analysis algorithmsabstractWe describe a new approach to manage the testing of document analysis and understanding applications. We propose and present a collection of document images, a set of techniques to prepare the test cases interactively and means to control the testing process. The systems architecture is designed to be distributed, scalable and platform independent utilizing Java, C++ and object-oriented databases. The main features of this system are a basic document categorization and ground truth, degradation models, custom test case creation facilities, a test management module (pipelining, test history), the ability to embed document analysis algorithms into the system, remote usage facilities and robust graphical user interfaces. Jaakko J. Sauvola, Sami Haapakoski, Hannu Kauniskangas, Tapio Seppänen, Matti Pietikäinen, David S. Doermann |
ICDAR | 5 |
| 1997 | Adaptive Document BinarizationabstractA new method is presented for adaptive document image binarization, where the page is considered as a collection of subcomponents such as text, background and picture. The problems caused by noise, illumination and many source type related degradations are addressed. The algorithm uses document characteristics to determine (surface) attributes, often used in document segmentation. Using characteristic analysis, two new algorithms are applied to determine a local threshold for each pixel. An algorithm based on soft decision control is used for thresholding the background and picture regions. An approach utilizing local mean and variance of gray values is applied to textual regions. Tests were performed with images including different types of document components and degradations. The results show that the method adapts and performs well in each case. Jaakko J. Sauvola, Tapio Seppänen, Sami Haapakoski, Matti Pietikäinen |
ICDAR | 4 |
| 1996 | The Development of a General Framework for Intelligent Document Image RetrievalabstractWork has recently begun on a joint project between the Universities of Maryland and Oulu on the development of a system for Intelligent Document Image Retrieval (IDIR). The IDIR system will provide close connections with and utilization of document analysis and image processing techniques, advanced computing and networking, and modern approaches to database management. The system design consists of aggressively modularized components to enhance the development of individual parts which are used in the complete solution, including: Interface specifications, multipurpose feature extraction, an integrated efficient query language, physical retrieval from an object-oriented database, and delivery of retrieved objects. In this paper, we introduce the general framework, feature extraction modules, query capabilities, a graphical query interface, and the application interface. We demonstrate each component of the system and how the query mechanisms can be used to handle both content and struc... David S. Doermann, Jaakko J. Sauvola, Hannu Kauniskangas, Christian K. Shin, Matti Pietikäinen, Azriel Rosenfeld |
DAS | 5 |
| 1996 | Accurate color discrimination with classification based on feature distributionsabstractThis paper considers a pattern recognition approach to accurate camera-based color measurements. The performances of Swain and Ballard's color indexing method based on 3-dimensional histograms and of a simplified version based on three 1-dimensional histograms are evaluated using three common color spaces and two sets of test images. Matti Pietikäinen, Sami Nieminen, Elzbieta Marszalec, Timo Ojala |
ICPR | 1 |
| 1996 | A document management interface utilizing page decomposition and content-based compressionabstractPage decomposition is very important in order to achieve an efficient interface for electronic (re)usability of documents. This paper proposes an approach in which the page is first segmented and classified into background text and image areas. These areas are then manipulated separately with a dedicated compression method applied to each class and represented in a way which provides an easy and open interface; for example, to send, manage or archive selected or all parts of the document. We call this a document management interface (DMI). Jaakko J. Sauvola, Matti Pietikäinen |
ICPR | 2 |
| 1996 | Some Aspects of RGB Vision and Its Applications in IndustryabstractRGB machine vision overlaps both colorimetry instrumentation and gray scale machine vision, and offers advantages over both. Proper calibration of the color camera which makes measurements independent of changes in illumination is essential for the reliability of color machine vision. A practical approach to on-line color camera calibration under unstable illumination conditions is presented and the performance of the procedure is evaluated. New potential applications of RGB machine vision are discussed from the perspective of physics-based vision and the calibration procedure developed here. Elzbieta Marsylaec, Matti Pietikäinen |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1996 | Determining Composition of Grain Mixtures by Texture Classification Based on Feature DistributionsabstractTexture analysis has many areas of potential application in industry. The problem of determining composition of grain mixtures by texture analysis was recently studied by Kjell. He obtained promising results when using all nine Laws' 3 × 3 features simultaneously and an ordinary feature vector classifier. In this paper the performance of texture classification based on feature distributions in this problem is evaluated. The results obtained are compared to those obtained with a feature vector classifier. The use of distributions of gray level differences as texture measures is also considered. Timo Ojala, Matti Pietikäinen, Jarkko Nisula |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1996 | A comparative study of texture measures with classification based on featured distributions
Timo Ojala, Matti Pietikäinen, David Harwood |
Pattern Recognit. | 2 |
| 1995 | Page segmentation and classification using fast feature extraction and connectivity analysisabstractPage segmentation and classification are important parts of the document analysis process. The aim is to extract and classify different parts of the page. This paper proposes an approach in which these two phases are combined. The integration process includes fast feature extraction with rule-based classification and label propagation using connectivity analysis providing classified areas in three categories: background, text and picture. Jaakko J. Sauvola, Matti Pietikäinen |
ICDAR | 2 |
| 1995 | An Experimental Comparison of Autoregressive and Fourier-Based Descriptors in 2D Shape ClassificationabstractAn experimental comparison of shape classification methods based on autoregressive modeling and Fourier descriptors of closed contours is carried out. The performance is evaluated using two independent sets of data: images of letters and airplanes. Silhouette contours are extracted from non-occluded 2D objects rotated, scaled, and translated in 3D space. Several versions of both types of methods are implemented and tested systematically. The comparison clearly shows better performance of Fourier-based methods, especially for images containing noise.> Hannu Kauppinen, Tapio Seppänen, Matti Pietikäinen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1995 | Texture classification by center-symmetric auto-correlation, using Kullback discrimination of distributions
David Harwood, Timo Ojala, Matti Pietikäinen, Shalom Kelman, Larry Davis 0001 |
Pattern Recognit. Lett. | 3 |
| 1994 | Online color camera calibrationabstractThis paper presents a practical approach to online color camera calibration under changing illumination conditions. Calibration is based on a color camera model and color constancy. In the first stage of the calibration a reconstruction of the current illuminant is performed based on an image taken from a scene with a calibrated color set. Using the updated power spectrum distribution of the current illumination the output camera values are corrected. Based on the corrected camera output values reconstruction of the reflectance spectrum of a random color object is performed and color features of the object can be determined under a required illuminant color temperature. The simulation experiments of the algorithm and results of tests on real images are discussed and the performance of the calibration is evaluated. The approach provides a real time color camera calibration procedure for many applications of machine vision. Elzbieta Marszalec, Matti Pietikäinen |
ICPR (1) | 2 |
| 1994 | Performance evaluation of texture measures with classification based on Kullback discrimination of distributionsabstractThis paper evaluates the performance both of some texture measures which have been successfully used in various applications and of some new promising approaches. For classification a method based on Kullback discrimination of sample and prototype distributions is used. The classification results for single features with one-dimensional feature value distributions and for pairs of complementary features with two-dimensional distributions are presented. Timo Ojala, Matti Pietikäinen, David Harwood |
ICPR (1) | 2 |
| 1992 | Evaluating quality of surface description using robust methodsabstractIn previous work (Koivunen and Pietikainen, 1991) the authors presented a segmentation method that combines useful properties of edge and region-based segmentation. The least squares estimation used gives good results when pixels in the neighborhood are from one statistical population, and the noise is Gaussian distributed. To be able to deal with very deviant pixel values, the authors applied an iterative reweighting least squares method and a least trimmed squares method for surface description. This paper presents the improvements on the robustness of the surface description, and a quantitative analysis of the quality of the description. The validity of the assumptions used is also evaluated quantitatively. Both synthetic and real range images are used for test images.> Visa Koivunen, Matti Pietikäinen |
ICPR (3) | 2 |
| 1992 | Edge-based texture measures for surface inspectionabstractPietikainen and Rosenfeld (1982) introduced a class of texture measures based on first-order statistics derived from edges in an image. the objective of this paper is to evaluate the performance of these measures and some new edge-based texture measures using two different types of data sets: images taken from the Brodatz album and images from a practical wood surface inspection problem. The results obtained for edge-based measures are compared to those obtained by popular second-order texture measures and tonal features. The role of the classifier on the performance is also studied by comparing the results obtained for three parametric classifiers and for a nonparametric k-nearest neighbor classifier. The results indicate that edge-based approaches are very promising for surface inspection problems, because they are relatively simple to compute and have performed very well in experiments.> Timo Ojala, Matti Pietikäinen, Olli Silvén |
ICPR (2) | 2 |
| 1990 | A hybrid computer architecture for machine visionabstractA hybrid computer architecture for machine vision which combines the useful properties of different types of architectures is introduced. HYBRID, an experimental hybrid system consisting of specialized Datacube-compatible processors and a transputer network, has been developed in a Sun-3 environment. The VLSI implementation of an edge-preserving smoothing operator for the low-level vision system is described, and the performance of transputer-based systems for higher-level vision is evaluated. Methods for analyzing and optimizing the performance of a hybrid architecture are discussed.> Matti Pietikäinen, Tapio Seppänen, Pertti Alapuranen |
ICPR (2) | 1 |
| 1990 | Color segmentation by hierarchical connected components analysis with image enhancement by symmetric neighborhood filtersabstractThe authors introduce a general-purpose procedure for image segmentation which combines iterative image enhancement by a symmetric neighborhood filter (SNF) with an iterative, hierarchical connected component (HCC) analysis. Color vector versions of both SNF and HCC are developed, using two common kinds of metrics applied to color vector intensity differences-a simple Euclidean-type metric and a coordinate maximum-type metric. The segmentation procedures are illustrated with three-band color images of indoor and outdoor scenes.> Tapani Westman, David Harwood, Toni Laitinen, Matti Pietikäinen |
ICPR (1) | 4 |
| 1990 | Automated visual inspection of rolled metal surfaces
Timo Piironen, Olli Silvén, Matti Pietikäinen, Toni Laitinen, Esko Strömmer |
Mach. Vis. Appl. | 3 |
| 1987 | Ferguson patch based method for representation of 3-D scenesabstractThe need for 3-D computer vision in robot control, measurement of object geometry and autonomous vehicle navigation is clear. Many current 3-D schemes rely on algebraic surface representation methods. This paper presents a parametric surface representation method based on the Ferguson surface patch. A parametric representation method appears to be better for describing free form surfaces than algebraic methods. This kind of method may also advance the integration of CAD/CAM and computer vision. Heikki Ailisto, Matti Pietikäinen |
ICRA | 2 |
| 1983 | Multispectral image smoothing guided by global distribution of pixel valuesabstractMultispectral images can be effectively smoothed by using the global distribution of pixel values to guide a local selective averaging process. After several iterations of this process, a typical image is virtually segmented into regions of constant value while significant edges in the image are preserved. Leslie J. Kitchen, Matti Pietikäinen, Azriel Rosenfeld, Chengye Wang |
IEEE Trans. Syst. Man Cybern. | 2 |
| 1983 | Experiments with texture classification using averages of local pattern matchesabstractLaws has introduced a class of texture features based on average degrees of match of the pixel neighbourhoods with a set of standard masks. These features yield better texture classification than standard features based on pairs of pixels. Simplifications of these features are investigated. Their performance is not greatly affected by their exact form and also appears to remain the same if only local match maxima are used. An alternative definition of such features is also presented, based on sums and differences of Gaussian convolutions. Matti Pietikäinen, Azriel Rosenfeld, Larry Davis 0001 |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1982 | Split-and-link algorithms for image segmentation
Matti Pietikäinen, Azriel Rosenfeld, Ingrid Walter |
Pattern Recognit. | 1 |