Xuefeng Liang

dblp:67/4179 · DBLP profile ↗
← Back
49ranked-venue papers
9as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 19 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Rethinking multi-pattern mining from a perspective of pattern prototype learning
Guanghui Shi, Xuefeng Liang
Pattern Recognit.2
2025 Large Multimodal Model is a Better Comparator on Facial Beauty Prediction
abstract
Order learning has been proven to improve the generalization of facial beauty prediction (FBP) models. However, the scope for advancement remains due to the constraints of current FBP datasets and model scales. In this study, we propose a Large Multimodal Model based Order Learning (LMOL), pioneering the use of a Large Multimodal Model (LMM) as a comparator in order learning. Meanwhile, we develop an FB instruction dataset to fine-tune the LMM, thereby enabling LMOL to discern the FB order. Experiments on three datasets showcase substantial performance improvements in FBP tasks, especially with a notable 18% increase in the Pearson Correlation Coefficient on zero-shot tests (two unseen datasets). This study demonstrate that LMM is a better comparator for FBP tasks.
Zhenyou Liu, Xuefeng Liang
ICASSP2
2025 Sample-level Self-paced Learning to Tackle Multimodal Imbalance Problem
abstract
The issue of multimodal imbalance has attracted widespread attention recently, and spurred the proposal of numerous dataset-level modulation strategies. However, we observe that the degree of modality imbalance may vary significantly across different samples, suggesting that dataset-level strategies may fail to learn the full spectrum of multimodal information present in certain samples. In this paper, we propose a sample-level multimodal self-paced learning strategy (SMSL). It first assesses the degree of modality imbalance in each sample and progressively learns the weaker modalities in these samples through self-paced learning with a two-phase modulation. This approach not only helps to fully leverage the multimodal information of each sample but also ensures the model's stability. Experiments on two benchmark datasets, CREMA-D and IEMOCAP, have demonstrated the robustness and effectiveness of SMSL.
Xuefeng Liang
ICASSP2
2025 LipReading for Low-resource Languages by Language Dynamic LoRA
abstract
Lipreading in low-resource languages is a highly demanded yet extremely challenging task. Recent research has shown that transferring the lipreading knowledge from data-rich languages to low-resource ones can improve the performance. However, unique lip movements in low-resource languages and heterogeneous syntactic structures and lexical conventions across languages render such knowledge transfer ineffective. To address these challenges, we introduce a Dynamic Low-Rank Fine-tuning to adaptively adjust model parameters for learning a set of basic lip shapes shared across languages, thereby better recognizing lip movements, especially those unique to low-resource languages. Moreover, we design an instruction tuning on multilingual LLMs to enhance the model’s cross-lingual text mapping capability. Experiments conducted on low-resource language datasets, Italian and Portuguese, demonstrate the effectiveness of our proposed method.
Shuai Zou, Xuefeng Liang
ICASSP2
2025 Learning Separable Fine-Grained Representation via Dendrogram Construction from Coarse Labels for Fine-grained Visual Recognition
Guanghui Shi, Xuefeng Liang
ICCV2
2025 Learning subjective time-series data via Utopia Label Distribution Approximation
Xuefeng Liang, Hexin Jiang, Yin Zhao
Pattern Recognit.1
2024 Quadripartite: Tackle Noisy Labels by Better Utilizing Uncertain Samples
abstract
To tackle the noisy label problem, sample selection methods have demonstrated particular advantages. An emerging method, named Tripartite, was proposed to divide the data into clean, uncertain, and noisy subsets and applies a low-weight training strategy to the uncertain subset to alleviate the harm of noisy labels. However, we find the uncertain subset selected by the Tripartite contains many noisy samples as well, which considerably lowers the performance of models even when using low-weight training. Therefore, we propose the Quadripartite, which further divides the uncertain subset into an uncertain clean and an uncertain noisy subset. The data visualization shows that these two new subsets have a high sample selection quality. It addresses a tough challenge that most sample selection methods face: how to reliably select the hard clean samples. The Quadripartite employs a weighted learning on the uncertain clean subset to maximize the utilization of clean samples, while dropping the uncertain noisy subset during the training stage to mitigate the negative effects caused by noisy samples. Extensive experiments show Quadripartite outperforms the state-of-the-art approaches on four benchmark datasets.
Xiaoxuan Cui, Xuefeng Liang, Guanghui Shi, Nesma El-haddar
IJCNN2
2024 Leveraging Knowledge of Modality Experts for Incomplete Multimodal Learning
Hexin Jiang, Xuefeng Liang
ACM Multimedia3
2024 Multipattern Mining Using Pattern-Level Contrastive Learning and Multipattern Activation Map
abstract
Visual patterns are basic elements in images and represent the discernible regularity in the visual world. Thus, mining visual patterns is a fundamental task in computer vision. Most previous studies consider that only one visual pattern exists in a category, and then builds up a one-to-one mapping using category label. In reality, however, many categories include multiple patterns, which are many-to-one mappings. Without knowing the information of patterns, few existing pattern mining methods can discover and distinguish varied patterns in a category. To tackle this problem, we propose a novel framework, PaclMap, which learns medium-grained features to represent patterns. It includes an unsupervised pattern-level contrastive learning and a multipattern activation map. Their joint optimization encourages the network to mine both discriminative and frequent patterns in a category. Extensive experiments conducted on four benchmark datasets (Place-20, imagenet large scale visual recognition challenge (ILSVRC)-20, visual object classes (VOC), and Travel) demonstrate that PaclMap outperforms six state-of-the-art methods with average improvements of 2.9% on accuracy and 12.3% on frequency, respectively.
Xuefeng Liang, Zhihui Liang, Huiwen Shi, Ying Zhou 0021
IEEE Trans. Neural Networks Learn. Syst.1
2023 Interpreting Positional Information in Perspective of Word Order
abstract
The attention mechanism is a powerful and effective method utilized in natural language processing.However, it has been observed that this method is insensitive to positional information.Although several studies have attempted to improve positional encoding and investigate the influence of word order perturbation, it remains unclear how positional encoding impacts NLP models from the perspective of word order.In this paper, we aim to shed light on this problem by analyzing the working mechanism of the attention module and investigating the root cause of its inability to encode positional information.Our hypothesis is that the insensitivity can be attributed to the weight sum operation utilized in the attention module.To verify this hypothesis, we propose a novel weight concatenation operation and evaluate its efficacy in neural machine translation tasks.Our enhanced experimental results not only reveal that the proposed operation can effectively encode positional information but also confirm our hypothesis.
Xilong Zhang, Xuefeng Liang
ACL (1)4
2023 Adaptive Mask Co-Optimization for Modal Dependence in Multimodal Learning
abstract
Multimodal learning has demonstrated a great advantage in emotion recognition tasks due to the richer information from different modalities. However, multimodal models may incline to rely on some modalities that are easier to be learned, while under-fit the other modalities and lead to sub-optimal results. To address this problem, we propose a novel plug-in module, Adaptive Mask Co-optimization (AMCo), which could be inserted into advanced models. The adaptive mask can encourage the model to fit other modalities better by making dependent modalities harder to be learned. The cooptimization can preserve the performance of models on dependent modalities without degradation. The extensive experiments on the IEMOCAP dataset show AMCo can improve four state-of-the-art models by 1.14% ~ 3.03% in terms of accuracy.
Xuefeng Liang, Shiquan Zheng, Huijun Xuan, Takatsune Kumada
ICASSP2
2023 Pairwise-Emotion Data Distribution Smoothing for Emotion Recognition
Hexin Jiang, Xuefeng Liang
PRCV (3)2
2023 MetaSelection: A Learnable Masked AutoEncoder for Multimodal Sentiment Feature Selection
Xuefeng Liang, Huijun Xuan
PRCV (7)1
2023 Stable Visual Pattern Mining via Pattern Probability Distribution
Xuefeng Liang, Guanghui Shi
PRCV (8)2
2022 A Single-Pathway Biomimetic Model for Potential Collision Prediction
Guodong Lei, Xuefeng Liang
PRCV (3)3
2022 A joint framework for mining discriminative and frequent visual representation
Ying Zhou 0021, Xuefeng Liang, Zhihui Liang, Yu Gu 0015, Yifei Yin
Neurocomputing2
2022 Multi-Classifier Interactive Learning for Ambiguous Speech Emotion Recognition
abstract
In recent years, speech emotion recognition technology is of great significance in widespread applications such as call centers, social robots and health care. Thus, the speech emotion recognition has been attracted much attention in both industry and academic. Since emotions existing in an entire utterance may have varied probabilities, speech emotion is likely to be ambiguous, which poses great challenges to recognition tasks. However, previous studies commonly assigned a single-label or multi-label to each utterance in certain. Therefore, their algorithms result in low accuracies because of the inappropriate representation. Inspired by the optimally interacting theory, we address the ambiguous speech emotions by proposing a novel multi-classifier interactive learning (MCIL) method. In MCIL, multiple different classifiers first mimic several individuals, who have inconsistent cognitions of ambiguous emotions, and construct new ambiguous labels (the emotion probability distribution). Then, they are retrained with the new labels to interact with their cognitions. This procedure enables each classifier to learn better representations of ambiguous data from others, and further improves the recognition ability. The experiments on three benchmark corpora (MAS, IEMOCAP, and FAU-AIBO) demonstrate that MCIL does not only improve each classifier’s performance, but also raises their recognition consistency from moderate to substantial.
Ying Zhou 0021, Xuefeng Liang, Yu Gu 0015, Yifei Yin, Longshan Yao
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Polarimetric Multipath Convolutional Neural Network for PolSAR Image Classification
abstract
Scatter targets of complex land covers in polarimetric synthetic aperture radar (PolSAR) images are often randomly oriented and cause randomly fluctuating echoes, which brings a challenge to PolSAR image classification. Therefore, many existing methods have alleviated this problem through orientation compensation. However, there are still two obstacles that limit the improvement of classification accuracy. On the one hand, generally, these methods process PolSAR images with fixed polarization rotation angles, which is experience-dependent and inflexible. On the other hand, for the different land covers of a PolSAR image, the existing methods do not consider these rotation angles separately. For the first obstacle, we design a group of convolution kernels called polarization rotation kernels (PRKs) and utilize them to build the polarimetric convolutional neural network (CNN) (PolCNN). The PolCNN is the base network of our final model, and it can learn polarization rotation angles adaptively. For the second obstacle, we extend the PolCNN into a multipath structure, the final model polarimetric multipath CNN (PolMPCNN). The polarization rotation angles of different land covers are directly related to the networks of different paths within the PolMPCNN. Furthermore, we also put forward the two-scale sampling and the stagewise training algorithm in order that our PolMPCNN can fit different scales of PolSAR targets and pays more attention to difficult training samples. Experiments on real PolSAR images show that the proposed model achieves the best classification results with an extremely low sampling rate of 0.1%.
Yuanhao Cui, Fang Liu 0001, Licheng Jiao, Yuwei Guo 0001, Xuefeng Liang, Lingling Li 0002, Shuyuan Yang 0001, Xiaoxue Qian
IEEE Trans. Geosci. Remote. Sens.5
2022 Unsupervised Outlier Detection Using Memory and Contrastive Learning
abstract
Outlier detection is to separate anomalous data from inliers in the dataset. Recently, the most deep learning methods of outlier detection leverage an auxiliary reconstruction task by assuming that outliers are more difficult to recover than normal samples (inliers). However, it is not always true in deep auto-encoder (AE) based models. The auto-encoder based detectors may recover certain outliers even if outliers are not in the training data, because they do not constrain the feature learning. Instead, we think outlier detection can be done in the feature space by measuring the distance between outliers' features and the consistency feature of inliers. To achieve this, we propose an unsupervised outlier detection method using a memory module and a contrastive learning module (MCOD). The memory module constrains the consistency of features, which merely represent the normal data. The contrastive learning module learns more discriminative features, which boosts the distinction between outliers and inliers. Extensive experiments on four benchmark datasets show that our proposed MCOD performs well and outperforms eleven state-of-the-art methods.
Ning Huyan, Dou Quan, Xiangrong Zhang, Xuefeng Liang, Jocelyn Chanussot, Licheng Jiao
IEEE Trans. Image Process.4
2022 Element-Wise Feature Relation Learning Network for Cross-Spectral Image Patch Matching
abstract
Recently, the majority of successful matching approaches are based on convolutional neural networks, which focus on learning the invariant and discriminative features for individual image patches based on image content. However, the image patch matching task is essentially to predict the matching relationship of patch pairs, that is, matching (similar) or non-matching (dissimilar). Therefore, we consider that the feature relation (FR) learning is more important than individual feature learning for image patch matching problem. Motivated by this, we propose an element-wise FR learning network for image patch matching, which transforms the image patch matching task into an image relationship-based pattern classification problem and dramatically improves generalization performances on image matching. Meanwhile, the proposed element-wise learning methods encourage full interaction between feature information and can naturally learn FR. Moreover, we propose to aggregate FR from multilevels, which integrates the multiscale FR for more precise matching. Experimental results demonstrate that our proposal achieves superior performances on cross-spectral image patch matching and single spectral image patch matching, and good generalization on image patch retrieval.
Dou Quan, Shuang Wang 0001, Ning Huyan, Jocelyn Chanussot, Ruojing Wang, Xuefeng Liang, Biao Hou, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.6
2021 Progressive Co-Teaching for Ambiguous Speech Emotion Recognition
abstract
Speech emotion recognition is a challenging task due to the ambiguity of emotion, which makes it difficult to learn the features of emotion data using machine learning algorithms. However, previous studies conventionally ignore the ambiguity of emotion and treat the emotion data as the same difficulty level, which results in low recognition accuracy. Motivated by human and animal learning studies, we propose a novel method named Progressive Co-teaching (PCT) to learn speech emotion features from simple to difficult. PCT method automatically identifies the difficulty level of data by itself using loss values, and then each network exchanges easy instances with small loss to peer network for early training. The rest instances with large loss are added gradually for later training. The experiment results demonstrate that our method achieves an improvement of 3.8% and 1.27% on MAS and IEMOCAP database than the state-of-the-arts, respectively.
Yifei Yin, Yu Gu 0015, Longshan Yao, Ying Zhou 0021, Xuefeng Liang
ICASSP5
2021 CALLip: Lipreading using Contrastive and Attribute Learning
abstract
Lipreading, aiming at interpreting speech by watching the lip movements of the speaker, has great significance in human communication and speech understanding. Despite having reached a feasible performance, lipreading still faces two crucial challenges: 1) the considerable lip movement variations cross different persons when they utter the same words; 2) the similar lip movements of people when they utter some confused phonemes. To tackle these two problems, we propose a novel lipreading framework, CALLip, which employs attribute learning and contrastive learning. The attribute learning extracts the speaker identity-aware features through a speaker recognition branch, which are able to normalize the lip shapes to eliminate cross-speaker variations. Considering that audio signals are intrinsically more distinguishable than visual signals, the contrastive learning is devised between visual and audio signals to enhance the discrimination of visual features and alleviate the viseme confusion problem. Experimental results show that CALLip does learn better features of lip movements. The comparisons on both English and Chinese benchmark datasets, GRID and CMLR, demonstrate that CALLip outperforms six state-of-the-art lipreading methods without using any additional data.
Xuefeng Liang, Chaowei Fang
ACM Multimedia2
2020 Jointly Discriminating and Frequent Visual Representation Mining
Qiannan Wang, Ying Zhou 0021, Zhaoyan Zhu, Xuefeng Liang, Yu Gu 0015
ACCV (3)4
2020 Deep Relevance Feature Clustering for Discovering Visual Representation of Tourism Destination
Qiannan Wang, Zhaoyan Zhu, Xuefeng Liang, Huiwen Shi, Pu Cao
PRCV (3)3
2019 AFD-Net: Aggregated Feature Difference Learning for Cross-Spectral Image Patch Matching
abstract
Image patch matching across different spectral domains is more challenging than in a single spectral domain. We consider the reason is twofold: 1. the weaker discriminative feature learned by conventional methods; 2. the significant appearance difference between two images domains. To tackle these problems, we propose an aggregated feature difference learning network (AFD-Net). Unlike other methods that merely rely on the high-level features, we find the feature differences in other levels also provide useful learning information. Thus, the multi-level feature differences are aggregated to enhance the discrimination. To make features invariant across different domains, we introduce a domain invariant feature extraction network based on instance normalization (IN). In order to optimize the AFD-Net, we borrow the large margin cosine loss which can minimize intra-class distance and maximize inter-class distance between matching and non-matching samples. Extensive experiments show that AFD-Net largely outperforms the state-of-the-arts on the cross-spectral dataset, meanwhile, demonstrates a considerable generalizability on a single spectral dataset.
Dou Quan, Xuefeng Liang, Shuang Wang 0001, Shaowei Wei, Ning Huyan, Licheng Jiao
ICCV2
2019 Better and Faster: Exponential Loss for Image Patch Matching
abstract
Recent studies on image patch matching are paying more attention on hard sample learning, because easy samples do not contribute much to the network optimization. They have proposed various hard negative sample mining strategies, but very few addressed this problem from the perspective of loss functions. Our research shows that the conventional Siamese and triplet losses treat all samples linearly, thus make the training time consuming. Instead, we propose the exponential Siamese and triplet losses, which can naturally focus more on hard samples and put less emphasis on easy ones, meanwhile, speed up the optimization. To assist the exponential losses, we introduce the hard positive sample mining to further enhance the effectiveness. The extensive experiments demonstrate our proposal improves both metric and descriptor learning on several well accepted benchmarks, and outperforms the state-of-the-arts on the UBC dataset. Moreover, it also shows a better generalizability on cross-spectral image matching and image retrieval tasks.
Shuang Wang 0001, Xuefeng Liang, Dou Quan, Bowu Yang, Shaowei Wei, Licheng Jiao
ICCV3
2019 Golden Ratio: The Attributes of Facial Attractiveness Learned By CNN
abstract
The recent success of deep learning has promoted the applications in facial attractiveness prediction and enhancement. However, what attributes have been learned to represent facial attractiveness is not well discovered yet. In this work, we find that DNN can learn both local and global shape-cues of face (Golden Ratio) that associate with facial attractiveness. This finding is concluded from a newly trained CNN model and an interpretation of visualizing activation of the category-specific neurons. The CNN model is trained on thousands extremely attractive/unattractive face images, and achieves an accuracy of 98.05%. The deconvolutional neural network generates face-like representations that depict the intuitive attributes of four face attractive categories. The results are consistent with the beauty ratios of facial attractiveness in psychological research.
Xuefeng Liang, Song Tong, Takatsune Kumada, Sunao Iwaki
ICIP1
2019 Photo Shot-Type Disambiguation by Multi-Classifier Semi-Supervised Learning
abstract
The previous studies showed that classifying photo shot-types into either narrow-type or wide-type was non-trivial because of ambiguous photos. To disambiguate such data, we define this problem as learning an ambiguity distribution instead of the conventional binary label, and then propose a multi-classifier semi-supervised learning framework (MCSSL) to tackle two issues: (1) Commonly, the ambiguity distribution of such photos are unavailable in practice. By introducing semi-supervised learning to multiple different classifiers (DNNs), MCSSL can mimic the inconsistent recognitions among multiple sources to determine the ambiguous degrees automatically for a photo. (2) The widely accepted Cross-Entropy loss is unfeasible for ambiguity learning. Instead, MCSSL applies the Kullback-Leibler divergence for DNNs to measure the similarity between the predicted distribution and the one given by multiple sources. The experiments on 122K photos demonstrate that MCSSL raises the classification accuracy of all DNNs, and improve the recognition consistency among them from substantial to almost perfect, which is our goal of data disambiguation.
Xuefeng Liang, Song Tong, Takatsune Kumada
ICIP2
2019 An Improved Fully Convolutional Network for Learning Rich Building Features
abstract
Many efficient approaches are proposed to detect building in remote sensing images. In this paper, in order to learning rich building features better, we propose a full convolutional network with dense connection. There contributions are made: 1) To strengthen feature propagation, an improved dense network is introduced to the full convolution network. 2) We have designed top-down short connections to facilitate the fusion of high and low feature information. 3) In addition, we add the weighted cross entropy edge loss function to make the network pay more attention to building edge in detail. Experiments show that the proposed method achieves excellent performance on the remote sensing image data taken by the QuickBird satellite.
Shuang Wang 0001, Pei He, Dou Quan, Xuefeng Liang, Biao Hou
IGARSS6
2019 Low-light image enhancement using Gaussian Process for features retrieval
Yuen Peng Loh, Xuefeng Liang, Chee Seng Chan
Signal Process. Image Commun.2
2019 Face super-resolution via bilayer contextual representation
Kangli Zeng, Tao Lu 0001, Xuefeng Liang, Kai Li 0005, Yanduo Zhang
Signal Process. Image Commun.3
2018 Cross-Spectral Image Patch Matching by Learning Features of the Spatially Connected Patches in a Shared Space
Dou Quan, Shuai Fang, Xuefeng Liang, Shuang Wang 0001, Licheng Jiao
ACCV (2)3
2018 Deep Generative Matching Network for Optical and SAR Image Registration
abstract
Multimodal remote sensing images contain complementary information, thus, could potentially benefit many remote sensing applications. To this end, the image registration is a common requirement for utilizing the multimodal images. However, due to the rather different imaging mechanisms, multimodal image registration becomes much more challenging than ordinary registration, particular for optical and synthetic aperture radar (SAR) images. In this work, we design a deep matching network to exploit the latent and coherent features between multimodal patch pairs for inferring their matching labels. But, the network requires immense data for training, which is not usually met. To address this issue, we propose a generative matching network (GMN) to generate the coupled optical and SAR images, hence, improve the quantity and diversity of the training data. The experimental results show that our proposal significantly improves the registration performance of optical and SAR image registration, and achieves subpixel or close to subpixel error.
Dou Quan, Shuang Wang 0001, Xuefeng Liang, Ruojing Wang, Shuai Fang, Biao Hou, Licheng Jiao
IGARSS3
2018 An external learning assisted self-examples learning for image super-resolution
Bo Yue, Shuang Wang 0001, Xuefeng Liang, Licheng Jiao
Neurocomputing3
2018 How Does the Low-Rank Matrix Decomposition Help Internal and External Learnings for Super-Resolution
abstract
Wisely utilizing the internal and external learning methods is a new challenge in super-resolution problem. To address this issue, we analyze the attributes of two methodologies and find two observations of their recovered details: 1) they are complementary in both feature space and image plane and 2) they distribute sparsely in the spatial space. These inspire us to propose a low-rank solution which effectively integrates two learning methods and then achieves a superior result. To fit this solution, the internal learning method and the external learning method are tailored to produce multiple preliminary results. Our theoretical analysis and experiment prove that the proposed low-rank solution does not require massive inputs to guarantee the performance, and thereby simplifying the design of two learning methods for the solution. Intensive experiments show the proposed solution improves the single learning method in both qualitative and quantitative assessments. Surprisingly, it shows more superior capability on noisy images and outperforms state-of-the-art methods.
Shuang Wang 0001, Bo Yue, Xuefeng Liang, Licheng Jiao
IEEE Trans. Image Process.3
2017 SAR images super-resolution via cartoon-texture image decomposition and jointly optimized regressors
abstract
This paper presents a novel approach to enhance the spatial resolution of Synthetic Aperture Radar(SAR) images. SAR images super-resolution(SR) reconstruction is challenging since SAR images has more complex structures. Inspired by the recent advance on natural image SR techniques, we propose a joint learning based strategy[1], combined with the characteristics of SAR image, to reconstruct HR SAR images from LR SAR images. Our method has ability to handle the complicated structures of SAR images. Besides, SAR images are decomposed into cartoon components and texture components and processed respectively. The purpose of decomposing strategy is to reduce the influence of speckle noise of SAR images. The experimental results and comparative analyses verify the effectiveness of this algorithm.
Shuang Wang 0001, Caijin Xu, Bo Yue, Xuefeng Liang
IGARSS6
2017 Robust coupled dictionary learning with ℓ1-norm coefficients transition constraint for noisy image super-resolution
Bo Yue, Shuang Wang 0001, Xuefeng Liang, Licheng Jiao
Signal Process.3
2016 Visual attention inspired distant view and close-up view classification
abstract
The images of distant view and close-up view indicate a photographers' attention which can be further utilized for user behavior analysis and scene evaluation. As images may compose arbitrary contexts, distant view and close-up view classification becomes non-trivial. In this work, we found two cues can represent human visual attention, i.e. focus cue and scale cue. We model the focus cue in frequency domain using the Discrete Wavelet Transform, and employ signal distribution as the focus feature. For the scale cue, we model it by defining a spatial size and a conceptual size in the image using the Edge Box and Convolution Neural Network. By integrating these two models, a robust scheme is proposed for this non-trivial task. Experiments on a newly retrieved dataset, which has 2137 natural images, show the classification accuracy achieves up to 97.3%.
Song Tong, Yuen Peng Loh, Xuefeng Liang, Takatsune Kumada
ICIP3
2015 Discovering Obscure Sightseeing Spots by Analysis of Geo-tagged Social Images
abstract
In contrast to conventional studies of discovering hot spots, by analyzing geo-tagged images on Flickr, we introduce novel methods to discover obscure sightseeing spots that are less well-known while still worth visiting. To this end, we face two new challenges that the classical authority analysis based methods do not encounter: how to discover and rank spots on the basis of 1) popularity (obscurity level) and 2) scenery quality. For the first challenge, we estimate the obscurity level of a spot in accordance with the visiting asymmetry between photographers who are familiar with a target city and those who are not. For the second challenge, the behavior of both viewers who browsed the images and photographers are analyzed per each spot. We also develop an application system to help users to explore sightseeing spots with different geographical granularities. Experimental evaluations and analysis on a real dataset well demonstrate the effectiveness of the proposed methods.
Chenyi Zhuang, Qiang Ma 0001, Xuefeng Liang, Masatoshi Yoshikawa
ASONAM3
2015 Decision-Tree Based Hybrid Filter-Wrapping Method for the Fusion of Multiple Feature Sets
Cuicui Zhang, Xuefeng Liang, Naixue Xiong
ICIG (2)2
2015 External and internal learning for single-image super-resolution
abstract
Super-resolution (SR) problem still faces a challenge of wisely utilizing diverse learned priors to recover the lost details in low resolution images. In this work, we propose a novel method using low rank decomposition which integrates diverse priors learned from external and internal learning to construct SR image. The proposed method first applies an external dictionary learning to get the meta-detail that is commonly shared among images, and then introduces an internal prior learning to learn the local self-similarity (local structure) that is shared in the image. Both are essential but different priors for SR image construction. With these priors, a bank of preliminary HR images are obtained but with estimation errors and noise. To restrain the errors and noise, we consider these HR images as a high dimension data in dimension reduction problem, and solve it using a low rank decomposition. Experimental results show the proposed method preserves image details effectively, also outperforms state-of-the-arts in both visual and quantitative assessments, especially in dealing with the noise.
Shuang Wang 0001, Chris S. Lin, Xuefeng Liang, Bo Yue, Licheng Jiao
ICIP3
2014 Inlier Estimation for Moving Camera Motion Segmentation
Xuefeng Liang, Cuicui Zhang, Takashi Matsuyama
ACCV (4)1
2014 Anaba: An obscure sightseeing spots discovering system
abstract
Discovering and recommending points of interest (POI) are drawing more attention to meet the increasing demand from personalized tours. Unlike conventional systems focusing on popular sightseeing locations, we develop a system, Anaba, to discover the obscure sightseeing spots that are less well-known while still worth visiting. By analyzing geo-tagged images on image hosting websites (Flickr, etc.), Anaba discovers and ranks sightseeing spots based on their obscurity levels and scenery quality. Anaba first selects obscure candidates in accordance with the asymmetry between visitors who are familiar with a target area and those who are not. Then, it evaluates the scenery quality of each candidate by considering both social appreciation and the content of images shot around there. The experiments on a newly retrieved dataset demonstrate the effectiveness of the proposed system.
Chenyi Zhuang, Qiang Ma 0001, Xuefeng Liang, Masatoshi Yoshikawa
ICME3
2012 Multi-subregion face recognition using coarse-to-fine Quad-tree decomposition
Cuicui Zhang, Xuefeng Liang, Takashi Matsuyama
ICPR2
2008 A compensation scheme of fingerprint distortion using combined radial basis function model for ubiquitous services
Xuefeng Liang, Naixue Xiong, Laurence T. Yang, Hui Zhang 0001, Jong Hyuk Park 0001
Comput. Commun.1
2007 A Combinatorial Approach to Fingerprint Binarization and Minutiae Extraction Using Euclidean Distance Transform
abstract
Most of the fingerprint matching techniques require extraction of minutiae that are ridge endings or bifurcations of ridge lines in a fingerprint image. Crucial to this step is either detecting ridges from the gray-level image or binarizing the image and then extracting the minutiae. In this work, we firstly exploit the property of almost equal width of ridges and valleys for binarization. Computing the width of arbitrary shapes is a nontrivial task. So, we estimate the width using Euclidean distance transform (EDT) and provide a near-linear time algorithm for binarization. Secondly, instead of using thinned binary images for minutiae extraction, we detect minutiae straightaway from the binarized fingerprint images using EDT. We also use EDT values to get rid of spurs and bridges in the fingerprint image. Unlike many other previous methods, our work depends minimally on arbitrary selection of parameters.
Xuefeng Liang, Arijit Bishnu, Tetsuo Asano
Int. J. Pattern Recognit. Artif. Intell.1
2007 A Robust Fingerprint Indexing Scheme Using Minutia Neighborhood Structure and Low-Order Delaunay Triangles
abstract
Fingerprint indexing is a key technique in automatic fingerprint identification systems (AFIS). However, handling fingerprint distortion is still a problem. This paper concentrates on a more accurate fingerprint indexing algorithm that efficiently retrieves the topNpossible matching candidates from a huge database. To this end, we design a novel feature based on minutia neighborhood structure (we call this minutia detail and it contains richer minutia information) and a more stable triangulation algorithm (low-order Delaunay triangles, consisting of order 0 and 1 Delaunay triangles), which are both insensitive to fingerprint distortion. The indexing features include minutia detail and attributes of low-order Delaunay triangle (its handedness, angles, maximum edge, and related angles between orientation field and edges). Experiments on databases FVC2002 and FVC2004 show that the proposed algorithm considerably narrows down the search space in fingerprint databases and is stable for various fingerprints. We also compared it with other indexing approaches, and the results show our algorithm has better performance, especially on fingerprints with distortion.
Xuefeng Liang, Arijit Bishnu, Tetsuo Asano
IEEE Trans. Inf. Forensics Secur.1
2006 Feature Extraction for Time Series Classification Using Discriminating Wavelet Coefficients
Mao Song Lin, Xuefeng Liang
ISNN (1)4
2004 A Near-Linear Time Algorithm for Binarization of Fingerprint Images Using Distance Transform
Xuefeng Liang, Arijit Bishnu, Tetsuo Asano
IWCIA1