Jun Guo 0002

dblp:204/6231 · DBLP profile ↗
← Back
194ranked-venue papers
0as first author
60since 2021 · last 2026
0000-0001-9045-1339ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 121 · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 87 · 25 since 2021Databases, data management, data science and information retrieval · 15 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021Computer networks · 4Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing
abstract
Semantic segmentation takes a pivotal role in various applications such as autonomous driving and medical image analysis. When deploying segmentation models in practice, it is critical to test their behaviors in varied and complex scenes in advance. In this paper, we construct an automatic data generation pipeline Gen4Seg to stress-test semantic segmentation models by generating various challenging samples with different attribute changes. Beyond previous evaluation paradigms focusing solely on global weather and style transfer, we investigate variations in both appearance and geometry attributes at the object and image level. These include object color, material, size, and position, as well as image-level variations such as weather and style. To achieve this, we propose to edit visual attributes of existing real images with precise control of structural information, empowered by diffusion models. In this way, the existing segmentation labels can be reused for the edited images, which greatly reduces the labor costs of constructing datasets. Using our pipeline, we construct two new benchmarks, Pascal-EA and COCO-EA. We benchmark a broad variety of semantic segmentation models, spanning from conventional close-set models to recent open-vocabulary large models. We have several key findings: 1) advanced open-vocabulary models do not exhibit greater robustness compared to closed-set methods under geometric variations; 2) traditional data augmentation techniques, such as CutOut and CutMix, are limited in enhancing robustness against appearance variations; 3) our generation pipeline can also be employed as a data augmentation tool and improve both in-distribution and out-of-distribution performances. Our work suggests the potential of generative models as effective tools for automatically analyzing segmentation models, and we hope our findings will assist practitioners and researchers in developing more robust and reliable segmentation models.
Zijin Yin, Bing Li 0015, Kongming Liang, Hao Sun 0015, Zhongjiang He, Zhanyu Ma, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 HumanRecon: Neural reconstruction of dynamic human using geometric cues and physical priors
Junhui Yin, Wei Yin 0006, Hao Chen 0041, Xuqian Ren, Zhanyu Ma, Jun Guo 0002, Yifan Liu 0001
Pattern Recognit.6
2026 From Sight to Insight: Enhancing Confusable Structure Segmentation via Vision-Language Mutual Prompting
abstract
Confusable structure segmentation (CSS) is a type of semantic segmentation applied in remote sensing sea fog detection, medical image segmentation, camouflaged object detection, etc. Structural similarity and visual ambiguity are two critical issues in CSS that pose difficulties in distinguishing foreground objects from the background. Current methods focus primarily on enhancing visual representations and do not often incorporate multimodal information, which leads to performance bottlenecks. Inspired by recent achievements in vision-language models, we proposeVision-LanguageMutualPrompting (VLMP), a novel and unified language-guided framework that leverages text prompts to enhance CSS. Specifically, VLMP consists of vision-to-language prompting and language-to-vision prompting, which bidirectionally model the interactions between visual and linguistic features, thereby facilitating cross-modal complementary information flow. To prevent the predominance of one modality over another, we design a feature integration modulator that modulates and balances feature weights for adaptive multimodal fusion. Our framework is designed to be modular and flexible, allowing for integration with any backbone, including CNNs and transformers. We evaluate VLMP with three diverse datasets: SFDD-H8, QaTa-COV19, and CAMO-COD10K. Extensive experiments demonstrate the effectiveness and superiority of the proposed framework over those of state-of-the-art methods across these datasets. This shift from basicsightto deeperinsightin CSS through vision-language integration represents a significant advancement in the field.
Yihao Zuo, Mengqiu Xu, Kaixin Chen 0001, Ming Wu 0001, Zhanyu Ma, Jun Guo 0002
IEEE Trans. Multim.8
2025 ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer
abstract
The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos, failing to transfer specific subject motion; 2) struggle to preserve the diversity and accuracy of motion as transferring to subjects with varying shapes. To overcome these, we introduce ConMo, a zero-shot framework that disentangle and recompose the motions of subjects and camera movements. ConMo isolates individual subject and background motion cues from complex trajectories in source videos using only subject masks, and reassembles them for target video generation. This approach enables more accurate motion control across diverse subjects and improves performance in multi-subject scenarios. Additionally, we propose soft guidance in the recomposition stage which controls the retention of original motion to adjust shape constraints, aiding subject shape adaptation and semantic transformation. Unlike previous methods, ConMo unlocks a wide range of applications, including subject size and position editing, subject removal, semantic modifications, and camera motion simulation. Extensive experiments demonstrate that ConMo significantly outperforms state-of-the-art methods in motion fidelity and semantic consistency. The code is available at https://github.com/Andyplus1/ConMo.
Jiayi Gao, Zijin Yin, Changcheng Hua, Yuxin Peng 0001, Kongming Liang, Zhanyu Ma, Jun Guo 0002, Yang Liu 0105
CVPR7
2025 MFogHub: Bridging Multi-Regional and Multi-Satellite Data for Global Marine Fog Detection and Forecasting
abstract
Deep learning approaches for marine fog detection and forecasting have outperformed traditional methods, demonstrating significant scientific and practical importance. However, the limited availability of open-source datasets remains a major challenge. Existing datasets, often focused on a single region or satellite, restrict the ability to evaluate model performance across diverse conditions and hinder the exploration of intrinsic marine fog characteristics. To address these limitations, we introduce MFogHub, the first multi-regional and multi-satellite dataset to integrate annotated marine fog observations from 15 coastal fog-prone regions and six geostationary satellites, comprising over 68,000 high-resolution samples. By encompassing diverse regions and satellite perspectives, MFogHub facilitates rigorous evaluation of both detection and forecasting methods under varying conditions. Extensive experiments with 16 baseline models demonstrate that MFogHub can reveal generalization fluctuations due to regional and satellite discrepancy, while also serving as a valuable resource for the development of targeted and scalable fog prediction techniques. Through MFogHub, we aim to advance both the practical monitoring and scientific understanding of marine fog dynamics on a global scale. The dataset and code are at https://github.com/kaka0910/MFogHub.
Mengqiu Xu, Kaixin Chen 0001, Heng Guo 0003, Ming Wu 0001, Jun Guo 0002
CVPR8
2025 Large Language Model-Empowered Adversarial Fusion for Typhoon Track Prediction
abstract
Accurate prediction of typhoon tracks is essential for effective disaster prevention strategies. Given that typhoon tracks can be conceptualized as a special class of time series, promising results have been achieved via the learning transferability of large language models (LLMs) in time series forecasting. However, limitations primarily lie in the incomplete utilization of the temporal and channel features of typhoon tracks. To address this, we propose LAF, an LLM-empowered framework for typhoon track prediction with adversarial fusion. We begin by transferring knowledge from pretrained LLMs to typhoon tracks to capture temporal features. To exploit purer channel features, we design a channel feature extraction strategy to limit the noise introduced. Then, we implement adversarial fusion between temporal and channel features to effectively narrow their discrepancy, facilitating a more comprehensive fusion for typhoon track prediction. Evaluations on Northwest Pacific typhoon track data demonstrate the effectiveness of the proposed model.
Lei Luo 0008, Jiahao Luan, Anudeep Vurity, Sumanth Sriram, Jun Guo 0002
ICASSP6
2025 Bi-directional generative retrieval-augmented diffusion models for document-level informative argument extraction
Lei Luo 0008, Xuanzhi Chen, Liming Mao, Jun Guo 0002
Neurocomputing6
2025 Neural Normalized Cut: A differential and generalizable approach for spectral clustering
Shangzhi Zhang, Chun-Guang Li, Xianbiao Qi, Rong Xiao 0003, Jun Guo 0002
Pattern Recognit.6
2025 Detailed Object Description With Controllable Dimensions
abstract
Object description plays an important role for visually impaired individuals to understand and compare the differences between objects. Recent multimodal large language models (MLLMs) exhibit powerful perceptual abilities and demonstrate impressive potential for generating object-centric descriptions. However, the descriptions generated by such models may still usually contain a lot of content that is not relevant to the user intent or miss some important object dimension details. Under special scenarios, users may only need the details of certain dimensions of an object. In this paper, we propose a training-free object description refinement pipeline,Dimension Tailor, designed to enhance user-specified details in object descriptions. This pipeline includes three steps: dimension extracting, erasing, and supplementing, which decompose the description into user-specified dimensions. Dimension Tailor can not only improve the quality of object details but also offer flexibility in including or excluding specific dimensions based on user preferences. We conducted extensive experiments to demonstrate the effectiveness of Dimension Tailor on controllable object descriptions. Notably, the proposed pipeline can consistently improve the performance of the recent MLLMs. The code is currently accessible athttps://github.com/PRIS-CV/ControllableObjectDescription.
Haiwen Zhang, Baoteng Li, Kongming Liang, Hao Sun 0015, Zhongjiang He, Zhanyu Ma, Jun Guo 0002
IEEE Trans. Multim.8
2024 Dual-Prior Augmented Decoding Network for Long Tail Distribution in HOI Detection
abstract
Human object interaction detection aims at localizing human-object pairs and recognizing their interactions. Trapped by the long-tailed distribution of the data, existing HOI detection methods often have difficulty recognizing the tail categories. Many approaches try to improve the recognition of HOI tasks by utilizing external knowledge (e.g. pre-trained visual-language models). However, these approaches mainly utilize external knowledge at the HOI combination level and achieve limited improvement in the tail categories. In this paper, we propose a dual-prior augmented decoding network by decomposing the HOI task into two sub-tasks: human-object pair detection and interaction recognition. For each subtask, we leverage external knowledge to enhance the model's ability at a finer granularity. Specifically, we acquire the prior candidates from an external classifier and embed them to assist the subsequent decoding process. Thus, the long-tail problem is mitigated from a coarse-to-fine level with the corresponding external knowledge. Our approach outperforms existing state-of-the-art models in various settings and significantly boosts the performance on the tail HOI categories. The source code is available at https://github.com/PRIS-CV/DP-ADN.
Jiayi Gao, Kongming Liang, Wei Chen 0071, Zhanyu Ma, Jun Guo 0002
AAAI6
2024 Benchmarking Segmentation Models with Mask-Preserved Attribute Editing
abstract
When deploying segmentation models in practice, it is critical to evaluate their behaviors in varied and complex scenes. Different from the previous evaluation paradigms only in consideration of global attribute variations (e.g. adverse weather), we investigate both local and global attribute variations for robustness evaluation. To achieve this, we construct a mask-preserved attribute editing pipeline to edit visual attributes of real images with precise control of structural information. Therefore, the original segmentation labels can be reused for the edited images. Using our pipeline, we construct a benchmark covering both object and image attributes (e.g. color, material, pattern, style). We evaluate a broad variety of semantic segmentation models, spanning from conventional close-set models to recent open-vocabulary large models on their robustness to different types of variations. We find that both local and global attribute variations affect segmentation performances, and the sensitivity of models diverges across different variation types. We argue that local attributes have the same importance as global attributes, and should be considered in the robustness evaluation of segmentation models. Code: https://github.com/PRIS-CV/Pascal-EA.
Zijin Yin, Kongming Liang, Bing Li 0015, Zhanyu Ma, Jun Guo 0002
CVPR5
2024 Efficient Face Super-Resolution via Wavelet-based Feature Enhancement Network
abstract
Face super-resolution aims to reconstruct a high-resolution face image from a low-resolution face image. Previous methods typically employ an encoder-decoder structure to extract facial structural features, where the direct downsampling inevitably introduces distortions, especially to high-frequency features such as edges. To address this issue, we propose a wavelet-based feature enhancement network, which mitigates feature distortion by losslessly decomposing the input feature into high and low-frequency components using the wavelet transform and processing them separately. To improve the efficiency of facial feature extraction, a full domain Transformer is further proposed to enhance local, regional, and global facial features. Such designs allow our method to perform better without stacking many modules as previous methods did. Experiments show that our method effectively balances performance, model size, and speed. Code link: https://github.com/PRIS-CV/WFEN.
Heng Guo 0003, Xuannan Liu, Kongming Liang, Jiani Hu, Zhanyu Ma, Jun Guo 0002
ACM Multimedia7
2024 Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video Understanding
abstract
With the emergence of large pre-trained multimodal video models, multiple benchmarks have been proposed to evaluate model capabilities. However, most of the benchmarks are human-centric, with evaluation data and tasks centered around human applications. Animals are an integral part of the natural world, and animal-centric video understanding is crucial for animal welfare and conservation efforts. Yet, existing benchmarks overlook evaluations focused on animals, limiting the application of the models. To address this limitation, our work established an animal-centric benchmark, namely Animal-Bench, to allow for a comprehensive evaluation of model capabilities in real-world contexts, overcoming agent-bias in previous benchmarks. Animal-Bench includes 13 tasks encompassing both common tasks shared with humans and special tasks relevant to animal conservation, spanning 7 major animal categories and 819 species, comprising a total of 41,839 data entries. To generate this benchmark, we defined a task system centered on animals and proposed an automated pipeline for animal-centric data processing. To further validate the robustness of models against real-world challenges, we utilized a video editing approach to simulate realistic scenarios like weather changes and shooting parameters due to animal movements. We evaluated 8 current multimodal video models on our benchmark and found considerable room for improvement. We hope our work provides insights for the community and opens up new avenues for research in multimodal video models. Our data and code will be released at https://github.com/PRIS-CV/Animal-Bench.
Yinuo Jing, Ruxu Zhang, Kongming Liang, Zhongjiang He, Zhanyu Ma, Jun Guo 0002
NeurIPS7
2024 Learning Dynamic Prototypes for Visual Pattern Debiasing
abstract
Abstract Deep learning has achieved great success in academic benchmarks but fails to work effectively in the real world due to the potential dataset bias. The current learning methods are prone to inheriting or even amplifying the bias present in a training dataset and under-represent specific demographic groups. More recently, some dataset debiasing methods have been developed to address the above challenges based on the awareness of protected or sensitive attribute labels. However, the number of protected or sensitive attributes may be considerably large, making it laborious and costly to acquire sufficient manual annotation. To this end, we propose a prototype-based network to dynamically balance the learning of different subgroups for a given dataset. First, an object pattern embedding mechanism is presented to make the network focus on the foreground region. Then we design a prototype learning method to discover and extract the visual patterns from the training data in an unsupervised way. The number of prototypes is dynamic depending on the pattern structure of the feature space. We evaluate the proposed prototype-based network on three widely used polyp segmentation datasets with abundant qualitative and quantitative experiments. Experimental results show that our proposed method outperforms the CNN-based and transformer-based state-of-the-art methods in terms of both effectiveness and fairness metrics. Moreover, extensive ablation studies are conducted to show the effectiveness of each proposed component and various parameter values. Lastly, we analyze how the number of prototypes grows during the training process and visualize the associated subgroups for each learned prototype. The code and data will be released at https://github.com/zijinY/dynamic-prototype-debiasing .
Kongming Liang, Zijin Yin, Min Min, Zhanyu Ma, Jun Guo 0002
Int. J. Comput. Vis.6
2024 Scalable, explainable, adaptive information extraction from structure-aware nearest neighbor
Shudong Lu, Si Li 0001, Jun Guo 0002
Neurocomputing3
2024 Semi-Supervised Learning for FGVC With Out-of-Category Data
abstract
Despite great strides made on fine-grained visual classification (FGVC), current methods are still heavily reliant on fully-supervised paradigms where ample expert labels are called for. Semi-supervised learning (SSL) techniques, acquiring knowledge from unlabeled data, provide a considerable means forward and have shown great promise for coarse-grained problems. However, exiting SSL paradigms mostly assume in-category (i.e., category-aligned) unlabeled data, which hinders their effectiveness when re-proposed on FGVC. In this paper, we put forward a novel design specifically aimed at making out-of-category data work for semi-supervised FGVC. We work off an important assumption that all fine-grained categories naturally follow a hierarchical structure (e.g., the phylogenetic tree of "Aves" that covers all bird species). It follows that, instead of operating on individual samples, we can instead predict sample relations within this tree structure as the optimization goal of SSL. Beyond this, we further introduced two strategies uniquely brought by these tree structures to achieve inter-sample consistency regularization and reliable pseudo-relation. Our experimental results reveal that (i) the proposed method yields good robustness against out-of-category data, and (ii) it can be equipped with prior arts, boosting their performance thus yielding state-of-the-art results.
Ruoyi Du, Dongliang Chang, Zhanyu Ma, Kongming Liang, Yi-Zhe Song, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Bi-Directional Ensemble Feature Reconstruction Network for Few-Shot Fine-Grained Classification
abstract
The main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting - a quick pilot study reveals that they in fact push for the opposite (i.e., lower inter-class variations and higher intra-class variations). To alleviate this problem, prior works predominately use a support set to reconstruct the query image and then utilize metric learning to determine its category. Upon careful inspection, we further reveal that such unidirectional reconstruction methods only help to increase inter-class variations and are not effective in tackling intra-class variations. In this paper, we introduce a bi-reconstruction mechanism that can simultaneously accommodate for inter-class and intra-class variations. In addition to using the support set to reconstruct the query set for increasing inter-class variations, we further use the query set to reconstruct the support set for reducing intra-class variations. This design effectively helps the model to explore more subtle and discriminative features which is key for the fine-grained problem in hand. Furthermore, we also construct a self-reconstruction module to work alongside the bi-directional module to make the features even more discriminative. We introduce the snapshot ensemble method in the episodic learning strategy - a simple trick to further improve model performance without increasing training costs. Experimental results on three widely used fine-grained image classification datasets, as well as general and cross-domain few-shot image datasets, consistently show considerable improvements compared with other methods.
Jijie Wu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jie Cao 0014, Jun Guo 0002, Yi-Zhe Song
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Relation fusion propagation network for transductive few-shot learning
Hongyu Hao, Weichao Ge, Ming Wu 0001, Jun Guo 0002
Pattern Recognit.7
2024 A simple scheme to amplify inter-class discrepancy for improving few-shot fine-grained image classification
abstract
Few-shot image classification is a challenging topic in pattern recognition and computer vision. Few-shot fine-grained image classification is even more challenging, due to not only the few shots of labelled samples but also the subtle differences to distinguish subcategories in fine-grained images. A recent method called task discrepancy maximisation (TDM) can be embedded into the feature map reconstruction network (FRN) to generate discriminative features, by preserving the appearance details through reconstructing the query image and then assigning higher weights to more discriminative channels, producing the state-of-the-art performance for few-shot fine-grained image classification. However, due to the small inter-class discrepancy in fine-grained images and the small training set in few-shot learning, the training of FRN+TDM can result in excessively flexible boundaries between subcategories and hence overfitting. To resolve this problem, we propose a simple scheme to amplify inter-class discrepancy and thus improve FRN+TDM. To achieve this aim, instead of developing new modules, our scheme only involves two simple amendments to FRN+TDM: relaxing the inter-class score in TDM, and adding a centre loss to FRN. Extensive experiments on five benchmark datasets showcase that, although embarrassingly simple, our scheme is quite effective to improve the performance of few-shot fine-grained image classification. The code is available at https://github.com/Airgods/AFRN.git.
Zijie Guo, Rui Zhu 0006, Zhanyu Ma, Jun Guo 0002, Jing-Hao Xue
Pattern Recognit.5
2024 Mind the Gap: Open Set Domain Adaptation via Mutual-to-Separate Framework
abstract
Unsupervised domain adaptation aims to leverage labeled data from a source domain to learn a classifier for an unlabeled target domain. Amongst its many variants, open set domain adaptation (OSDA) is perhaps the most challenging one, as it further assumes the presence of unknown classes in the target domain. In this paper, we study OSDA with a particular focus on enriching its ability to traverse across larger domain gaps, and we show that existing state-of-the-art methods suffer a considerable performance drop in the presence of larger domain gaps, especially on a new dataset (PACS) that we re-purposed for OSDA. Exploring this is pivotal for OSDA as with increasing domain shift, identifying unknown samples in the target domain becomes harder for the model, thus making negative transfer between source and target domains more challenging. Accordingly, we propose a Mutual-to-Separate (MTS) framework to address the larger domain gaps. Essentially we design two networks – (a) Sample Separation Network (SSN): which is trained to learn a hyperplane for separating unknown samples from known ones, and (b) Distribution Matching Network (DMN): which is trained to maximise domain confusion between source and target domains without unknown samples under the guidance of the SSN. The key insight lies in how we exploit the mutually beneficial information between these two networks. On closer observation, we see that SSN can reveal which samples in the target domain belong to the unknown class by instance weighting whereas, DMN pushes apart the samples that most likely belong to the unknown class in the target domain, which in turn reduces the difficulty of SSN in identifying unknown samples. It follows that (a) and (b) will mutually supervise each other and alternate until convergence, which can better align the source and target domains in the shared label space. Extensive experiments on five datasets (Office-31, Office-Home, PACS, VisDA, andmini_DomainNet) demonstrate the efficiency of the proposed method. Detailed ablation experiments also validate the effectiveness of each component and the generality of the proposed framework. Codes are available at: https://github.com/PRIS-CV/Mutual-to-Separate.
Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Yi-Zhe Song, Ruiping Wang 0001, Jun Guo 0002
IEEE Trans. Circuits Syst. Video Technol.6
2023 Bi-directional Feature Reconstruction Network for Fine-Grained Few-Shot Image Classification
abstract
The main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting -- a quick pilot study reveals that they in fact push for the opposite (i.e., lower inter-class variations and higher intra-class variations). To alleviate this problem, prior works predominately use a support set to reconstruct the query image and then utilize metric learning to determine its category. Upon careful inspection, we further reveal that such unidirectional reconstruction methods only help to increase inter-class variations and are not effective in tackling intra-class variations. In this paper, we for the first time introduce a bi-reconstruction mechanism that can simultaneously accommodate for inter-class and intra-class variations. In addition to using the support set to reconstruct the query set for increasing inter-class variations, we further use the query set to reconstruct the support set for reducing intra-class variations. This design effectively helps the model to explore more subtle and discriminative features which is key for the fine-grained problem in hand. Furthermore, we also construct a self-reconstruction module to work alongside the bi-directional module to make the features even more discriminative. Experimental results on three widely used fine-grained image classification datasets consistently show considerable improvements compared with other methods. Codes are available at: https://github.com/PRIS-CV/Bi-FRN.
Jijie Wu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jie Cao 0014, Jun Guo 0002, Yi-Zhe Song
AAAI7
2023 Semantic Memory Guided Image Representation for Polyp Segmentation
abstract
Polyp segmentation is important in the early diagnosis and treatment of colorectal cancer. Since polyps vary in shape, size, color, and texture, accurate polyp segmentation is very challenging. One promising solution is to model the contextual relation for each pixel. However, previous methods only focus on learning the dependencies between the position within an individual image and ignore the contextual relation across different images. In this paper, we propose a memory-based feature enhancement module to capture the cross-image contextual relations. Specifically, we first present a polyp-centric representation. Then a semantic memory is designed to extract the polyp prototypes across different images. The feature at one position can be further enhanced by the contextual embeddings stored in the semantic memory. The enhanced feature is propagated into the features of the previous levels as the multi-scale guidance. The experimental results show that our method achieves better performance than other state-of-the-art methods.
Zijin Yin, Runpu Wei, Kongming Liang, Yiyang Lin, Zhanyu Ma, Min Min, Jun Guo 0002
ICASSP8
2023 Attribute Learning with Knowledge Enhanced Partial Annotations
abstract
Under limited annotation cost, large-scale attribute learning datasets only contain partial labels for each image. The conventional methods treat the un-annotated attributes as negative or ignore their loss without considering the associated knowledge. In this paper, we present a knowledge enhanced selective loss for partially labeled attribute learning. Given a visual instance, we investigate the object-attribute co-occurrence as internal knowledge to subdivide the unannotated attributes into feasible and infeasible sets. Based on that, we can enhance the model to focus on the learning of feasible un-annotated attributes and remove the distraction from the infeasible ones. Besides the internal knowledge, we adopt external knowledge to excavate the unseen object-attribute pairs. Experimental results show that our proposed loss can achieve state-of-the-art performance on the newly cleaned VAW2 dataset that contains 170,407 instances, 1763 objects, and 591 attributes. The code and VAW2 dataset are available at https://github.com/GriffinLiang/seal.
Kongming Liang, Wei Chen 0071, Zhanyu Ma, Jun Guo 0002
ICIP6
2023 Hierarchical Visual Attribute Learning in the Wild
abstract
Observing objects' attributes at different levels of detail is a fundamental aspect of how humans perceive and understand the world around them. Existing studies focused on attribute prediction in a flat way, but they overlook the underlying attribute hierarchy, e.g., navy blue is a subcategory of blue. In recent years, large language models, e.g., ChatGPT, have emerged with the ability to perform an extensive range of natural language processing tasks like text generation and classification. The factual knowledge learned by LLM can assist us build the hierarchical relations of visual attributes in the wild. Based on that, we propose a model called the object-specific attribute relation net, which takes advantage of three types of relations among attributes - positive, negative, and hierarchical - to better facilitate attribute recognition in images. Guided by the extracted hierarchical relations, our model can predict attributes from coarse to fine. Additionally, we introduce several evaluation metrics for attribute hierarchy to comprehensively assess the model's ability to comprehend hierarchical relations. Our extensive experiments demonstrate that our proposed hierarchical annotation brings improvements to the model's understanding of hierarchical relations of attributes, and the object-specific attribute relation net can recognize visual attributes more accurately.
Kongming Liang, Haiwen Zhang, Zhanyu Ma, Jun Guo 0002
ACM Multimedia5
2023 Making a Bird AI Expert Work for You and Me
abstract
As powerful as fine-grained visual classification (FGVC) is, responding your query with a bird name of “Whip-poor-will” or “Mallard” probably does not make much sense. This however commonly accepted in the literature, underlines a fundamental question interfacing AI and human – what constitutes transferable knowledge for human to learn from AI? This paper sets out to answer this very question using FGVC as a test bed. Specifically, we envisage a scenario where a trained FGVC model (the AI expert) functions as a knowledge provider in enabling average people (you and me) to become better domain experts ourselves,i.e.,those capable in distinguishing between “Whip-poor-will” and “Mallard”. Fig. 1 lays out our approach in answering this question. Assuming an AI expert trained using expert human labels, we ask (i) what is the best transferable knowledge we can extract from AI, and (ii) what is the most practical means to measure the gains in expertise given that knowledge? On the former, we propose to represent knowledge as highly discriminative visual regions that are expert-exclusive. For that, we devise a multi-stage learning framework, which starts with modelling visual attention of domain experts and novices separately, before discriminatively distilling their differences to acquire those exclusive to experts. For the latter, we simulate the evaluation process as a book guide to best accommodate the learning practice of that is accustomed to humans. A comprehensive human study of 15,000 trials shows our method is able to consistently improve people of divergent bird expertise to recognise once unrecognisable birds. To counter the lack of reproducibility of perceptual studies, and in turn to make a sustainable direction out of our “AI for Human” effort, we further propose a quantitative metric, namely Transferable Effective Model Attention (TEMI). TEMI acts as a crude but benchmarkable metric to replace large-scale human studies, and therefore allows future efforts in this direction to be comparable to ours. We attest to the integrity of TEMI by (i) empirically showing a strong correlation between TEMI scores and raw human study data, and (ii) its expected behaviour holds for a large body of attention models. Last but not least, our approach also leads to improved FGVC performance in the conventional benchmarking sense, when the extracted knowledge defined is utilised as means to achieve discriminative localisation. Codes and all details on the human study are available at:https://github.com/PRIS-CV/Making-a-Bird-AI-Expert-Work-for-You-and-Me.
Dongliang Chang, Kaiyue Pang, Ruoyi Du, Yujun Tong, Yi-Zhe Song, Zhanyu Ma, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 M²NAS: Joint Neural Architecture Optimization System With Network Transmission
abstract
Differentiable neural architecture search (NAS) methods have achieved comparable results for low search costs and high performance. Existing differentiable methods focus on searching microstructures in micro space, lacking in exploring macrostructures. However, different networks should have different macrostructures rather than all evenly distributed. This article proposes an M2NAS optimization system to optimize macrostructure and microstructure jointly. Specifically, we initialize a network with maximum complexity, where down-samplings occur at the beginning. Then, iteratively optimize the macrostructure and microstructure. For macro search, we establish a macro space and explore this space by generating candidates and initializing weights and architectural parameters for candidate networks, i.e., network transmission. For micro search, we propose a progressive pruning method to eliminate the vast quantization error caused by one-time pruning. With the number of parameters decreasing during the search, M2NAS obtains a series of networks with different complexity through network selection, forming a Pareto-optimal set. Results show that networks obtained by M2NAS have a better tradeoff between accuracy and complexity. A network searched on ImageNet achieves 77.4% and 76.3% top-1 accuracy, with and without data augment during retraining. When taking different pretrained models as backbones and combining them with Faster RCNN on COCO, our model can get 35.6%$(AP)$and 55.7%$(AP_{50})$, higher than typical NAS models, second only to ResNet-50 results.
Lanfei Wang, Lingxi Xie, Kaifeng Bi, Kaili Zhao, Jun Guo 0002, Qi Tian 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Weakly Supervised Sea Fog Detection in Remote Sensing Images via Prototype Learning
abstract
Sea fog detection is a challenging and significant task in the field of remote sensing. Deep learning-based methods have shown promising potential, but require a large amount of pixel-level labeled data that are time-consuming and labor-intensive to acquire. To scale up the dataset and overcome the limitations of pixel-level annotation, we attempt to explore the existing knowledge from historical statistics for label efficient sea fog detection. In this paper, we propose an image-level Weakly Supervised Sea Fog Detection Dataset (WS-SFDD) and a novel weakly supervised sea fog detection framework via prototype learning, named ProCAM. According to the sea fog events recorded by the Marine Weather Review published quarterly by the National Meteorological Center of China, we collect the sea fog images from Himawari-8 satellite data and obtain free image-level labels to construct the dataset. However, with image-level annotations, existing weakly supervised semantic segmentation methods mainly rely on class activation maps (CAMs) and have limitations when applied to such a specific scenario: 1) the pseudo labels mainly cover the most discriminative part of object regions that are incomplete; 2) the background is complex with varying atmospheric conditions and it is difficult to distinguish sea fog from low clouds due to their high similarity in spectral characteristics; 3) the co-occurring context like ‘sea’ distracts the model and thus degrades the performance. To address the above issues, in our proposed ProCAM, we first design a prototype re-activation (PRA) module that reactivates self-similar sea fog regions by pixel-to-prototype feature matching to improve the robustness and completeness of CAMs. Then, we develop a pixel-to-prototype contrastive (PPC) learning method to increase the distance between sea fog and background in the embedding space for learning more discriminative dense features. Finally, a self-augmented regularization (SAR) strategy is presented to decouple sea fog from its co-occurring context and thus avoid background interference. Extensive experiments on the WS-SFDD dataset demonstrate our proposed method ProCAM achieves superior performance with an F1-score of 77.59% and a critical success index of 63.39%. To the best of our knowledge, this is the first work to perform image-level weakly supervised sea fog detection in remote sensing images. The dataset and code are available at https://github.com/yixianghuang/ProCAM.
Ming Wu 0001, Xin Jiang 0036, Jiaao Li, Mengqiu Xu, Jun Guo 0002
IEEE Trans. Geosci. Remote. Sens.7
2023 A Real-Time Memory Updating Strategy for Unsupervised Person Re-Identification
abstract
Recently, clustering-based methods have been the dominant solution for unsupervised person re-identification (ReID). Memory-based contrastive learning is widely used for its effectiveness in unsupervised representation learning. However, we find that the inaccurate cluster proxies and the momentum updating strategy do harm to the contrastive learning system. In this paper, we propose a real-time memory updating strategy (RTMem) to update the cluster centroid with a randomly sampled instance feature in the current mini-batch without momentum. Compared to the method that calculates the mean feature vectors as the cluster centroid and updating it with momentum, RTMem enables the features to be up-to-date for each cluster. Based on RTMem, we propose two contrastive losses, i.e., sample-to-instance and sample-to-cluster, to align the relationships between samples to each cluster and to all outliers not belonging to any other clusters. On the one hand, sample-to-instance loss explores the sample relationships of the whole dataset to enhance the capability of density-based clustering algorithm, which relies on similarity measurement for the instance-level images. On the other hand, with pseudo-labels generated by the density-based clustering algorithm, sample-to-cluster loss enforces the sample to be close to its cluster proxy while being far from other proxies. With the simple RTMem contrastive learning strategy, the performance of the corresponding baseline is improved by 9.3% on Market-1501 dataset. Our method consistently outperforms state-of-the-art unsupervised learning person ReID methods on three benchmark datasets. Code is made available at:https://github.com/PRIS-CV/RTMem.
Junhui Yin, Xinyu Zhang 0015, Zhanyu Ma, Jun Guo 0002, Yifan Liu 0001
IEEE Trans. Image Process.4
2023 On the Comparisons of Decorrelation Approaches for Non-Gaussian Neutral Vector Variables
abstract
-norm equals one. In addition, its neutral properties make it significantly different from the commonly studied vector variables (e.g., the Gaussian vector variables). Due to the aforementioned properties, the conventionally applied linear transformation approaches [e.g., principal component analysis (PCA) and independent component analysis (ICA)] are not suitable for neutral vector variables, as PCA cannot transform a neutral vector variable, which is highly negatively correlated, into a set of mutually independent scalar variables and ICA cannot preserve the bounded property after transformation. In recent work, we proposed an efficient nonlinear transformation approach, i.e., the parallel nonlinear transformation (PNT), for decorrelating neutral vector variables. In this article, we extensively compare PNT with PCA and ICA through both theoretical analysis and experimental evaluations. The results of our investigations demonstrate the superiority of PNT for decorrelating the neutral vector variables.
Zhanyu Ma, Xiaoou Lu, Jiyang Xie 0001, Zhen Yang 0004, Jing-Hao Xue, Zheng-Hua Tan, Bo Xiao 0006, Jun Guo 0002
IEEE Trans. Neural Networks Learn. Syst.8
2022 Learning Invariant Visual Representations for Compositional Zero-Shot Learning
Tian Zhang 0029, Kongming Liang, Ruoyi Du, Zhanyu Ma, Jun Guo 0002
ECCV (24)6
2022 Pointshift: Point-Wise Shift MLP for Pixel-Level Cloud Type Classification in Meteorological Satellite Imagery
abstract
The deep neural network has recently achieved promising results on cloud type classification, which gets rid of the hand-crafted features and plays an essential role in climate change analysis. Previous methods perform context reasoning with single-scale representation at the centre of the network, which is challenging to capture sufficient contextual information. In this paper, we propose a point-wise shift multi-layer perceptron (MLP) for pixel-level cloud type classification, termed PointShift, which effectively models point-wise and multi-scale neighbour information. We design a shift operation to compose a multi-scale receptive field in a non-parametric manner. Besides, we introduce split attention to improve the interaction of feature channels. Extensive experiments on the Himawari-8 image dataset demonstrate that our proposed architecture achieves the best mIoU of 71.06% and a competitive trade-off between efficiency and performance.
Zhaoqing Wang, Xin Jiang 0036, Ming Wu 0001, Jun Guo 0002
IGARSS6
2022 ENDE-GNN: An Encoder-decoder GNN Framework for Sketch Semantic Segmentation
abstract
Sketch semantic segmentation serves as an important part of sketch interpretation. Recently, some researchers have obtained significant results using graph neural networks (GNN) for this task. However, existing GNN-based methods usually neglect the drawing order of sketches thus missing out the sequence information inherent to sketches. Towards solving this problem to achieve better performance on sketch semantic segmentation, we propose an encoder-decoder GNN framework named ENDE-GNN. Working with an auxiliary decoder, our ENDE-GNN guides the GNN backbone network to not only extract the inter-stroke and intra-stroke features, but also pays attention to the drawing order of sketches. This decoder acts during training only, preventing any additional overhead during testing. The proposed ENDE-GNN obtains state-of-the-art per-formances on three public sketch semantic segmentation datasets, namely SPG, SketchSeg-150K, and CreativeSketch. We further evaluate the effectiveness of ENDE-GNN via ablation studies and visualizations. Codes are available at https://github.com/PRIS-CV/ENDE_For_SSS.
Yixiao Zheng, Jiyang Xie 0001, Aneeshan Sain, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002
VCIP6
2022 Event detection from text using path-aware graph convolutional network
Shudong Lu, Si Li 0001, Haibo Lan, Jun Guo 0002
Appl. Intell.6
2022 Scalable NAS with factorizable architectural parameters
Lanfei Wang, Lingxi Xie, Kaili Zhao, Jun Guo 0002, Qi Tian 0001
Neurocomputing4
2022 Dual-granularity feature alignment for cross-modality person re-identification
Junhui Yin, Zhanyu Ma, Jiyang Xie 0001, Shibo Nie, Kongming Liang, Jun Guo 0002
Neurocomputing6
2022 Explainable document-level event extraction via back-tracing to sentence-level event clues
Shudong Lu, Si Li 0001, Jun Guo 0002
Knowl. Based Syst.4
2022 Leveraging speaker-aware structure and factual knowledge for faithful dialogue summarization
Weiran Xu, Chunyun Zhang, Jun Guo 0002
Knowl. Based Syst.4
2022 A Correlation Context-Driven Method for Sea Fog Detection in Meteorological Satellite Imagery
abstract
Sea fog detection is a challenging and essential issue in satellite remote sensing. Although conventional threshold methods and deep learning methods can achieve pixel-level classification, it is difficult to distinguish ambiguous boundaries and thin structures from the background. Considering the correlations between neighbor pixels and the affinities between superpixels, a correlation context-driven method for sea fog detection is proposed in this letter, which mainly consists of a two-stage superpixel-based fully convolutional network (SFCNet), named SFCNet. A fully connected Conditional Random Field (CRF) is utilized to model the dependencies between pixels. To alleviate the problem of high cloud occlusion, an attentive Generative Adversarial Network (GAN) is implemented for image enhancement by exploiting contextual information. Experimental results demonstrate that our proposed method achieves 91.65% mIoU and obtains more refined segmentation results, performing well in detecting fogs in small, broken bits and weak contrast thin structures, as well as detects more obscured parts.
Ming Wu 0001, Jun Guo 0002, Mengqiu Xu
IEEE Geosci. Remote. Sens. Lett.3
2022 Progressive Learning of Category-Consistent Multi-Granularity Features for Fine-Grained Visual Classification
abstract
Fine-grained visual classification (FGVC) is much more challenging than traditional classification tasks due to the inherently subtle intra-class object variations. Recent works are mainly part-driven (either explicitly or implicitly), with the assumption that fine-grained information naturally rests within the parts. In this paper, we take a different stance, and show that part operations are not strictly necessary - the key lies with encouraging the network to learn at different granularities and progressively fusing multi-granularity features together. In particular, we propose: (i) a progressive training strategy that effectively fuses features from different granularities, and (ii) a consistent block convolution that encourages the network to learn the category-consistent features at specific granularities. We evaluate on several standard FGVC benchmark datasets, and demonstrate the proposed method consistently outperforms existing alternatives or delivers competitive results. Codes are available at https://github.com/PRIS-CV/PMG-V2.
Ruoyi Du, Jiyang Xie 0001, Zhanyu Ma, Dongliang Chang, Yi-Zhe Song, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 GPCA: A Probabilistic Framework for Gaussian Process Embedded Channel Attention
abstract
Channel attention mechanisms have been commonly applied in many visual tasks for effective performance improvement. It is able to reinforce the informative channels as well as to suppress the useless channels. Recently, different channel attention modules have been proposed and implemented in various ways. Generally speaking, they are mainly based on convolution and pooling operations. In this paper, we propose Gaussian process embedded channel attention (GPCA) module and further interpret the channel attention schemes in a probabilistic way. The GPCA module intends to model the correlations among the channels, which are assumed to be captured by beta distributed variables. As the beta distribution cannot be integrated into the end-to-end training of convolutional neural networks (CNNs) with a mathematically tractable solution, we utilize an approximation of the beta distribution to solve this problem. To specify, we adapt a Sigmoid-Gaussian approximation, in which the Gaussian distributed variables are transferred into the interval [0,1]. The Gaussian process is then utilized to model the correlations among different channels. In this case, a mathematically tractable solution is derived. The GPCA module can be efficiently implemented and integrated into the end-to-end training of the CNNs. Experimental results demonstrate the promising performance of the proposed GPCA module. Codes are available at https://github.com/PRIS-CV/GPCA.
Jiyang Xie 0001, Zhanyu Ma, Dongliang Chang, Guoqiang Zhang 0003, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Advanced Dropout: A Model-Free Methodology for Bayesian Dropout Optimization
abstract
Due to lack of data, overfitting ubiquitously exists in real-world applications of deep neural networks (DNNs). We propose advanced dropout, a model-free methodology, to mitigate overfitting and improve the performance of DNNs. The advanced dropout technique applies a model-free and easily implemented distribution with parametric prior, and adaptively adjusts dropout rate. Specifically, the distribution parameters are optimized by stochastic gradient variational Bayes in order to carry out an end-to-end training. We evaluate the effectiveness of the advanced dropout against nine dropout techniques on seven computer vision datasets (five small-scale datasets and two large-scale datasets) with various base models. The advanced dropout outperforms all the referred techniques on all the datasets. We further compare the effectiveness ratios and find that advanced dropout achieves the highest one on most cases. Next, we conduct a set of analysis of dropout rate characteristics, including convergence of the adaptive dropout rate, the learned distributions of dropout masks, and a comparison with dropout rate generation without an explicit distribution. In addition, the ability of overfitting prevention is evaluated and confirmed. Finally, we extend the application of the advanced dropout to uncertainty inference, network pruning, text classification, and regression. The proposed advanced dropout is also superior to the corresponding referred methods. Codes are available at https://github.com/PRIS-CV/AdvancedDropout.
Jiyang Xie 0001, Zhanyu Ma, Jianjun Lei 0001, Guoqiang Zhang 0003, Jing-Hao Xue, Zheng-Hua Tan, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2022 The devil in the tail: Cluster consolidation plus cluster adaptive balancing loss for unsupervised person re-identification
Chaoqun Lin, Chun-Guang Li, Jun Guo 0002
Pattern Recognit.5
2022 MPCCL: Multiview predictive coding with contrastive learning for person re-identification
Junhui Yin, Jiyang Xie 0001, Zhanyu Ma, Jun Guo 0002
Pattern Recognit.4
2022 Unsupervised person re-identification via simultaneous clustering and mask prediction
Junhui Yin, Siqing Zhang 0001, Jiyang Xie 0001, Zhanyu Ma, Jun Guo 0002
Pattern Recognit.5
2022 CorefDPR: A Joint Model for Coreference Resolution and Dropped Pronoun Recovery in Chinese Conversations
abstract
In this work, we present that coreference resolution and dropped pronoun recovery are two strongly related tasks in Chinese conversations, as recovering the dropped pronoun needs to explore the referent of the pronoun at first. Meanwhile, the omitted entity mention should be recovered before its coreferences are resolved. This motivates us to propose CorefDPR, a novel model to jointly resolve these two tasks and make them enhance each other. CorefDPR firstly utilizes a pre-trained language model to encode tokens in the conversation snippet. Then, the coreference resolution layer detects all entity mentions from the candidate text spans and groups them as coreferent mention clusters based on the contextualized token states. Furthermore, the pronoun recovery layer explores the referent of each dropped pronoun from the coreferent mention clusters and predicts the probability distribution over pronoun category for each token. Finally, a general conditional random fields (GCRF) is employed to globally optimize the pronoun recovery sequence of the snippet by modeling both intra-utterance and cross-utterance pronoun dependencies, and the recovered pronouns are further linked back to corresponding mention clusters to complete them. Experimental results on the benchmark demonstrate that our proposed model outperformed the state-of-the-art baselines of both these two tasks, and the exploratory experiments also demonstrate that these two tasks mutually benefit each other.
Si Li 0001, Sheng Gao 0001, Jun Guo 0002
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Cluster-Guided Asymmetric Contrastive Learning for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) aims to match pedestrian images from different camera views in an unsupervised setting. Existing methods for unsupervised person Re-ID are usually built upon the pseudo labels from clustering. However, the result of clustering depends heavily on the quality of the learned features, which are overwhelmingly dominated by colors in images. In this paper, we attempt to suppress the negative dominating influence of colors to learn more effective features for unsupervised person Re-ID. Specifically, we propose a Cluster-guided Asymmetric Contrastive Learning (CACL) approach for unsupervised person Re-ID, in which clustering result is leveraged to guide the feature learning in a properly designed asymmetric contrastive learning framework. In CACL, both instance-level and cluster-level contrastive learning are employed to help the siamese network learn discriminant features with respect to the clustering result within and between different data augmentation views, respectively. In addition, we also present a cluster refinement method, and validate that the cluster refinement step helps CACL significantly. Extensive experiments conducted on three benchmark datasets demonstrate the superior performance of our proposal.
Chun-Guang Li, Jun Guo 0002
IEEE Trans. Image Process.3
2022 Dirichlet Process Mixture of Generalized Inverted Dirichlet Distributions for Positive Vector Data With Extended Variational Inference
abstract
A Bayesian nonparametric approach for estimation of a Dirichlet process (DP) mixture of generalized inverted Dirichlet distributions [i.e., an infinite generalized inverted Dirichlet mixture model (InGIDMM)] has been proposed. The generalized inverted Dirichlet distribution has been proven to be efficient in modeling the vectors that contain only positive elements. Under the classical variational inference (VI) framework, the key challenge in the Bayesian estimation of InGIDMM is that the expectation of the joint distribution of data and variables cannot be explicitly calculated. Therefore, numerical methods are usually applied to simulate the optimal posterior distributions. With the recently proposed extended VI (EVI) framework, we introduce lower bound approximations to the original variational objective function in the VI framework such that an analytically tractable solution can be derived. Hence, the problem in numerical simulation has been overcome. By applying the DP mixture technique, an InGIDMM can automatically determine the number of mixture components from the observed data. Moreover, the DP mixture model with an infinite number of mixture components also avoids the problems of underfitting and overfitting. The performance of the proposed approach is demonstrated with both synthesized data and real-life data applications.
Zhanyu Ma, Yuping Lai, Jiyang Xie 0001, Deyu Meng, W. Bastiaan Kleijn, Jun Guo 0002, Jingyi Yu 0001
IEEE Trans. Neural Networks Learn. Syst.6
2021 A Joint Model for Dropped Pronoun Recovery and Conversational Discourse Parsing in Chinese Conversational Speech
abstract
Jingxuan Yang, Kerui Xu, Jun Xu, Si Li, Sheng Gao, Jun Guo, Nianwen Xue, Ji-Rong Wen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Kerui Xu, Jun Xu 0001, Si Li 0001, Sheng Gao 0001, Jun Guo 0002, Nianwen Xue, Ji-Rong Wen
ACL/IJCNLP (1)6
2021 Your "Flamingo" is My "Bird": Fine-Grained, or Not
abstract
Whether what you see in Figure 1 is a "flamingo" or a "bird", is the question we ask in this paper. While fine-grained visual classification (FGVC) strives to arrive at the former, for the majority of us non-experts just "bird" would probably suffice. The real question is therefore – how can we tailor for different fine-grained definitions under divergent levels of expertise. For that, we re-envisage the traditional setting of FGVC, from single-label classification, to that of top-down traversal of a pre-defined coarse-to-fine label hierarchy – so that our answer becomes "bird" ⇒ "Phoenicopteriformes" ⇒ "Phoenicopteridae" ⇒ "flamingo".To approach this new problem, we first conduct a comprehensive human study where we confirm that most participants prefer multi-granularity labels, regardless whether they consider themselves experts. We then discover the key intuition that: coarse-level label prediction exacerbates fine-grained feature learning, yet fine-level feature betters the learning of coarse-level classifier. This discovery enables us to design a very simple albeit surprisingly effective solution to our new problem, where we (i) leverage level-specific classification heads to disentangle coarse-level features with fine-grained ones, and (ii) allow finer-grained features to participate in coarser-grained label predictions, which in turn helps with better disentanglement. Experiments show that our method achieves superior performance in the new FGVC setting, and performs better than state-of-the-art on the traditional single-label FGVC problem as well. Thanks to its simplicity, our method can be easily implemented on top of any existing FGVC frameworks and is parameter-free.
Dongliang Chang, Kaiyue Pang, Yixiao Zheng, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002
CVPR6
2021 Joint Topology-preserving and Feature-refinement Network for Curvilinear Structure Segmentation
abstract
Curvilinear structure segmentation (CSS) is under semantic segmentation, whose applications include crack detection, aerial road extraction, and biomedical image segmentation. In general, geometric topology and pixel-wise features are two critical aspects of CSS. However, most semantic segmentation methods only focus on enhancing feature representations while existing CSS techniques emphasize preserving topology alone. In this paper, we present a Joint Topology-preserving and Feature-refinement Network (JTFN) that jointly models global topology and refined features based on an iterative feedback learning strategy. Specifically, we explore the structure of objects to help preserve corresponding topologies of predicted masks, thus design a reciprocative two-stream module for CSS and boundary detection. In addition, we introduce such topology-aware predictions as feedback guidance that refines attentive features by supplementing and enhancing saliencies. To the best of our knowledge, this is the first work that jointly addresses topology preserving and feature refinement for CSS. We evaluate JTFN on four datasets of diverse applications: Crack500, CrackTree200, Roads, and DRIVE. Results show that JTFN performs best in comparison with alternative methods. Code is available.1
Mingfei Cheng, Kaili Zhao, Xuhong Guo, Jun Guo 0002
ICCV5
2021 Knowledge Transfer Based Fine-Grained Visual Classification
abstract
Fine-grained visual classification (FGVC) aims to distinguish the sub-classes of the same category and its essential solution is to mine the subtle and discriminative regions. Convolution neural networks (CNNs), which employ the cross entropy loss (CE-loss) as the loss function, show poor performance since the model can only learn the most discriminative part and ignore other meaningful regions. Some existing works try to solve this problem by mining more discriminative regions by some detection techniques or attention mechanisms. However, most of them will meet the background noise problem when trying to find more discriminative regions. In this paper, we address it in a knowledge transfer learning manner. Multiple models are trained one by one, and all previously trained models are regarded as teacher models to supervise the training of the current one. Specifically, a orthogonal loss (OR-loss) is proposed to encourage the network to find diverse and meaningful regions. In addition, the first model is trained with only CE-Loss. Finally, all models’ outputs with complementary knowledge are combined together for the final prediction result. We demonstrate the superiority of the proposed method and obtain state-of-the-art (SOTA) performances on three popular FGVC datasets.
Siqing Zhang 0001, Ruoyi Du, Dongliang Chang, Zhanyu Ma, Jun Guo 0002
ICME5
2021 Exploring Category-Shared and Category-Specific Features for Fine-Grained Image Classification
Dongliang Chang, Bo Xiao 0002, Zhanyu Ma, Jun Guo 0002, Yaning Chang
PRCV (1)6
2021 Meta-Learned Specific Scenario Interest Network for User Preference Prediction
abstract
User preference prediction is a task of learning user interests through user-item interactions. Most existing studies capture user interests based on historical behaviors without considering specific scenario information. However, the users may have special interests in these specific scenarios and sometimes user historical behaviors are limited. In this paper, we propose a Meta-Learned Specific Scenario Interest Network (Meta-SSIN) to predict user preference of target item by capturing specific scenario interests. Meta-SSIN uses multiple independent meta-learning modules to model historical behaviors in each scenario. The independent module can capture special interests based on limited behaviors. Experimental results on three datasets show that Meta-SSIN outperforms compared state-of-the-art methods.
Hehuan Liu, Si Li 0001, Jun Guo 0002
SIGIR6
2021 Progressive Co-Attention Network for Fine-Grained Visual Classification
abstract
Fine-grained visual classification aims to recognize images belonging to multiple sub-categories within a same category. It is a challenging task due to the inherently subtle variations among highly-confused categories. Most existing methods only take an individual image as input, which may limit the ability of models to recognize contrastive clues from different images. In this paper, we propose an effective method called progressive co-attention network (PCA-Net) to tackle this problem. Specifically, we calculate the channel-wise similarity by encouraging interaction between the feature channels within same-category image pairs to capture the common discriminative features. Considering that complementary information is also crucial for recognition, we erase the prominent areas enhanced by the channel interaction to force the network to focus on other discriminative regions. The proposed model has achieved competitive results on three fine-grained visual classification benchmark datasets: CUB-200-2011, Stanford Cars, and FGVC Aircraft.
Tian Zhang 0029, Dongliang Chang, Zhanyu Ma, Jun Guo 0002
VCIP4
2021 Adapting User Preference to Online Feedback in Multi-round Conversational Recommendation
abstract
This paper concerns user preference estimation in multi-round conversational recommender systems (CRS), which interacts with users by asking questions about attributes and recommending items multiple times in one conversation. Multi-round CRS such as EAR have been proposed in which the user's online feedback at both attribute level and item level can be utilized to estimate user preference and make recommendations. Though preliminary success has been shown, existing user preference models in CRS usually use the online feedback information as independent features or training instances, overlooking the relation between attribute-level and item-level feedback signals. The relation can be used to more precisely identify the reasons (e.g., some certain attributes) that trigger the rejection of an item, leading to more fine-grained utilization of the feedback information. To address aforementioned issue, this paper proposes a novel preference estimation model tailored for multi-round CRS, called Feedback-guided Preference Adaptation Network (FPAN). In FPAN, two gating modules are designed to respectively adapt the original user embedding and item-level feedback, both according to the online attribute-level feedback. The gating modules utilize the fine-grained attribute-level feedback to revise the user embedding and coarse-grained item-level feedback, achieving more accurate user preference estimation by considering the relation between feedback. Experimental results on two benchmarks showed that FPAN outperformed the state-of-the-art user preference models in CRS, and the multi-round CRS can also be enhanced by using FPAN as its recommender component.
Kerui Xu, Jun Xu 0001, Sheng Gao 0001, Jun Guo 0002, Ji-Rong Wen
WSDM5
2021 Deep InterBoost networks for small-sample image classification
Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jun Guo 0002
Neurocomputing7
2021 Utilizing graph neural networks to improving dialogue-based relation extraction
Weiran Xu, Sheng Gao 0001, Jun Guo 0002
Neurocomputing4
2021 ReMarNet: Conjoint Relation and Margin Learning for Small-Sample Image Classification
abstract
Despite achieving state-of-the-art performance, deep learning methods generally require a large amount of labeled data during training and may suffer from overfitting when the sample size is small. To ensure good generalizability of deep networks under small sample sizes, learning discriminative features is crucial. To this end, several loss functions have been proposed to encourage large intra-class compactness and inter-class separability. In this paper, we propose to enhance the discriminative power of features from a new perspective by introducing a novel neural network termed Relation-and-Margin learning Network (ReMarNet). Our method assembles two networks of different backbones so as to learn the features that can perform excellently in both of the aforementioned two classification mechanisms. Specifically, a relation network is used to learn the features that can support classification based on the similarity between a sample and a class prototype; at the meantime, a fully connected network with the cross entropy loss is used for classification via the decision boundary. Experiments on four image datasets demonstrate that our approach is effective in learning discriminative features from a small set of labeled samples and achieves competitive performance against state-of-the-art methods. Code is available at https://github.com/liyunyu08/ReMarNet.
Liyun Yu, Zhanyu Ma, Jing-Hao Xue, Jie Cao 0014, Jun Guo 0002
IEEE Trans. Circuits Syst. Video Technol.7
2021 Fine-Grained Instance-Level Sketch-Based Video Retrieval
abstract
Existing sketch-analysis work studies sketches depicting static objects or scenes. In this work, we propose a novel cross-modal retrieval problem of fine-grained instance-level sketch-based video retrieval (FG-SBVR), where a sketch sequence is used as a query to retrieve a specific target video instance. Compared with sketch-based still image retrieval, and coarse-grained category-level video retrieval, this is more challenging as both visual appearance and motion need to be simultaneously matched at a fine-grained level. We contribute the first FG-SBVR dataset with rich annotations. We then introduce a novel multi-stream multi-modality deep network to perform FG-SBVR under both strong and weakly supervised settings. The key component of the network is a relation module, designed to prevent model overfitting given scarce training data. We show that this model significantly outperforms a number of existing state-of-the-art models designed for video analysis.
Peng Xu 0005, Kun Liu 0016, Tao Xiang 0002, Timothy M. Hospedales, Zhanyu Ma, Jun Guo 0002, Yi-Zhe Song
IEEE Trans. Circuits Syst. Video Technol.6
2021 DS-UI: Dual-Supervised Mixture of Gaussian Mixture Models for Uncertainty Inference in Image Recognition
abstract
This paper proposes a dual-supervised uncertainty inference (DS-UI) framework for improving Bayesian estimation-based UI in DNN-based image recognition. In the DS-UI, we combine the classifier of a DNN, i.e., the last fully-connected (FC) layer, with a mixture of Gaussian mixture models (MoGMM) to obtain an MoGMM-FC layer. Unlike existing UI methods for DNNs, which only calculate the means or modes of the DNN outputs' distributions, the proposed MoGMM-FC layer acts as a probabilistic interpreter for the features that are inputs of the classifier to directly calculate the probabilities of them for the DS-UI. In addition, we propose a dual-supervised stochastic gradient-based variational Bayes (DS-SGVB) algorithm for the MoGMM-FC layer optimization. Unlike conventional SGVB and optimization algorithms in other UI methods, the DS-SGVB not only models the samples in the specific class for each Gaussian mixture model (GMM) in the MoGMM, but also considers the negative samples from other classes for the GMM to reduce the intra-class distances and enlarge the inter-class margins simultaneously for enhancing the learning ability of the MoGMM-FC layer in the DS-UI. Experimental results show the DS-UI outperforms the state-of-the-art UI methods in misclassification detection. We further evaluate the DS-UI in open-set out-of-domain/-distribution detection and find statistically significant improvements. Visualizations of the feature spaces demonstrate the superiority of the DS-UI. Codes are available at https://github.com/PRIS-CV/DS-UI.
Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue, Guoqiang Zhang 0003, Yinhe Zheng, Jun Guo 0002
IEEE Trans. Image Process.7
2020 CGTR: Convolution Graph Topology Representation for Document Ranking
abstract
Contextualized neural language models have gained much attention in Information Retrieval (IR) with its ability to achieve better text understanding by capturing contextual structure. However, to achieve better document understanding, it is necessary to involve global structure of a document. In this paper, we take the advantage of Graph Convolutional Networks (GCN) to model global word-relation structure of a document to improve context-aware document ranking. We propose to build a graph for a document to model the global structure. The nodes and edges of the graph are constructed from contextual embeddings. Then we apply graph convolution on the graph to learning a new representation, and this representation covers both contextual and global structure information. The experimental results show that our method outperforms the state-of-the-art contextual language models, which demonstrate that incorporating global structure is useful for improving document ranking and GCN is an effective way to achieve it.
Jiayue Zhang, Weiran Xu, Jun Guo 0002
CIKM5
2020 Improving Abstractive Dialogue Summarization with Graph Structures and Topic Words
abstract
Recently, people have been beginning paying more attention to the abstractive dialogue summarization task. Since the information flows are exchanged between at least two interlocutors and key elements about a certain event are often spanned across multiple utterances, it is necessary for researchers to explore the inherent relations and structures of dialogue contents. However, the existing approaches often process the dialogue with sequence-based models, which are hard to capture long-distance inter-sentence relations. In this paper, we propose a Topic-word Guided Dialogue Graph Attention (TGDGA) network to model the dialogue as an interaction graph according to the topic word information. A masked graph self-attention mechanism is used to integrate cross-sentence information flows and focus more on the related utterances, which makes it better to understand the dialogue. Moreover, the topic word features are introduced to assist the decoding process. We evaluate our model on the SAMSum Corpus and Automobile Master Corpus. The experimental results show that our method outperforms most of the baselines.
Weiran Xu, Jun Guo 0002
COLING3
2020 Fine-Grained Visual Classification via Progressive Multi-granularity Training of Jigsaw Patches
Ruoyi Du, Dongliang Chang, Ayan Kumar Bhunia, Jiyang Xie 0001, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002
ECCV (20)7
2020 Progressive Refinement Network for Occluded Pedestrian Detection
Kaili Zhao, Wen-Sheng Chu, Honggang Zhang 0002, Jun Guo 0002
ECCV (23)5
2020 Self-paced Bottom-up Clustering Network with Side Information for Person Re-Identification
abstract
Person re-identification (Re-ID) has attracted a lot of research attention in recent years. However, supervised methods demand enormous amount of manually annotated data. In this paper, we propose a Self-Paced bottom-up Clustering Network with Side Information (SPCNet-SI) for unsupervised person Re-ID, where the side information comes from the serial number of the camera associated with each image. Specifically, our proposed SPCNet-SI exploits the camera side information to guide the feature learning and uses soft label in bottom-up clustering process, in which the camera association information is used in the repelled loss and the soft label based cluster information is used to select the candidate cluster pairs to merge. Moreover, a self-paced dynamic mechanism is developed to regularize the merging process such that the clustering is implemented in an easy-to-hard way with a slow-to-fast merging process. Experiments on two benchmark datasets Market-1501 and DukeMTMC-ReID demonstrate promising performance.
Chun-Guang Li, Ruo-Pei Guo, Jun Guo 0002
ICPR4
2020 Dual-attention Guided Dropblock Module for Weakly Supervised Object Localization
abstract
Attention mechanisms is frequently used to learn the discriminative features for better feature representations. In this paper, we extend the attention mechanism to the task of weakly supervised object localization (WSOL) and propose the dual-attention guided dropblock module (DGDM), which aims at learning the informative and complementary visual patterns for WSOL. This module contains two key components, the channel attention guided dropout (CAGD) and the spatial attention guided dropblock (SAGD). To model channel interdependencies, the CAGD ranks the channel attentions and treats the top-k attentions with the largest magnitudes as the important ones. It also keeps some low-valued elements to increase their value if they become important during training. The SAGD can efficiently remove the most discriminative information by erasing the contiguous regions of feature maps rather than individual pixels. This guides the model to capture the less discriminative parts for classification. Furthermore, it can also distinguish the foreground objects from the background regions to alleviate the attention misdirection. Experimental results demonstrate that the proposed method achieves new state-of-the-art localization performance.
Junhui Yin, Siqing Zhang 0001, Dongliang Chang, Zhanyu Ma, Jun Guo 0002
ICPR5
2020 Learning Graph Topology Representation with Attention Networks
abstract
Contextualized neural language models have gained much attention in Information Retrieval (IR) with its ability to achieve better word understanding by capturing contextual structure on sentence level. However, to understand a document better, it is necessary to involve contextual structure from document level. Moreover, some words contributes more information to delivering the meaning of a document. Motivated by this, in this paper, we take the advantages of Graph Convolutional Networks (GCN) and Graph Attention Networks (GAN) to model global word-relation structure of a document with attention mechanism to improve context-aware document ranking. We propose to build a graph for a document to model the global contextual structure. The nodes and edges of the graph are constructed from contextual embeddings. We first apply graph convolution on the graph and then use attention networks to explore the influence of more informative words to obtain a new representation. This representation covers both local contextual and global structure information. The experimental results show that our method outperforms the state-of-the-art contextual language models, which demonstrate that incorporating contextual structure is useful for improving document ranking.
Jiayue Zhang, Weiran Xu, Jun Guo 0002, Honggang Zhang 0002
VCIP4
2020 Density-adaptive kernel based efficient reranking approaches for person reidentification
Ruo-Pei Guo, Chun-Guang Li, Jiaru Lin, Jun Guo 0002
Neurocomputing5
2020 Cross-sentence N-ary relation classification using LSTMs on graph and sequence structures
Weiran Xu, Sheng Gao 0001, Jun Guo 0002
Knowl. Based Syst.4
2020 Outline Extraction with Question-Specific Memory Cells
abstract
Outline extraction has been widely applied in online consultation to help experts quickly understand individual cases. Given a specific case described as unstructured plain text, outline extraction aims to make a summary for this case by answering a set of questions, which in fact is a new type of machine reading comprehension task. Inspired by a recently popular memory network, we propose a novel question-specific memory cell network (QSMCN) to extract information related to multiple questions on-the-fly as it reads texts. QSMCN constructs a specific memory cell for each question, which is sequentially expanded in recurrent neural network style. Each cell contains three specific vectors to first identify whether current input is related to corresponding question and then update question-specific case representation. We add a penalization term in the loss function to make extracted knowledge more reasonable and interpretable. To support this study, we construct a new outline extraction corpus, InjuryCase, 1 which is composed of 3,995 real Chinese occupational injury cases. Experimental results show that our method makes a significant improvement. We further apply the proposed framework on two multi-aspect extraction tasks and find that the proposed model also remarkably outperforms existing state-of-the-art methods of the aspect extraction task.
Haotian Cui, Si Li 0001, Sheng Gao 0001, Jun Guo 0002, Zhengdong Lu
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2020 Deep Neural Network-Based Impacts Analysis of Multimodal Factors on Heat Demand Prediction
abstract
Prediction of heat demand using artificial neural networks has attracted enormous research attention. Weather conditions, such as direct solar irradiance and wind speed, have been identified as key parameters affecting heat demand. This paper employs an Elman neural network to investigate the impacts of direct solar irradiance and wind speed on the heat demand from the perspective of the entire district heating network. Results of the overall mean absolute percentage error (MAPE) show that direct solar irradiance and wind speed have quite similar impacts. However, the involvement of direct solar irradiance can clearly reduce the maximum absolute deviation when only involving direct solar irradiance and wind speed, respectively. In addition, the simultaneous involvement of both wind speed and direct solar irradiance does not show an obvious improvement of MAPE. Moreover, the prediction accuracy can also be affected by other factors like data discontinuity and outliers.
Zhanyu Ma, Jiyang Xie 0001, Qie Sun, Fredrik Wallin, Zhongwei Si, Jun Guo 0002
IEEE Trans. Big Data7
2020 The Devil is in the Channels: Mutual-Channel Loss for Fine-Grained Image Classification
abstract
The key to solving fine-grained image categorization is finding discriminate and local regions that correspond to subtle visual traits. Great strides have been made, with complex networks designed specifically to learn part-level discriminate feature representations. In this paper, we show that it is possible to cultivate subtle details without the need for overly complicated network designs or training mechanisms - a single loss is all it takes. The main trick lies with how we delve into individual feature channels early on, as opposed to the convention of starting from a consolidated feature map. The proposed loss function, termed as mutual-channel loss (MC-Loss), consists of two channel-specific components: a discriminality component and a diversity component. The discriminality component forces all feature channels belonging to the same class to be discriminative, through a novel channel-wise attention mechanism. The diversity component additionally constraints channels so that they become mutually exclusive across the spatial dimension. The end result is therefore a set of feature channels, each of which reflects different locally discriminative regions for a specific class. The MC-Loss can be trained end-to-end, without the need for any bounding-box/part annotations, and yields highly discriminative regions during inference. Experimental results show our MC-Loss when implemented on top of common base networks can achieve state-of-the-art performance on all four fine-grained categorization datasets (CUB-Birds, FGVC-Aircraft, Flowers-102, and Stanford Cars). Ablative studies further demonstrate the superiority of the MC-Loss when compared with other recently proposed general-purpose losses for visual classification, on two different base networks.
Dongliang Chang, Jiyang Xie 0001, Ayan Kumar Bhunia, Zhanyu Ma, Ming Wu 0001, Jun Guo 0002, Yi-Zhe Song
IEEE Trans. Image Process.8
2020 OSLNet: Deep Small-Sample Classification With an Orthogonal Softmax Layer
abstract
A deep neural network of multiple nonlinear layers forms a large function space, which can easily lead to overfitting when it encounters small-sample data. To mitigate overfitting in small-sample classification, learning more discriminative features from small-sample data is becoming a new trend. To this end, this paper aims to find a subspace of neural networks that can facilitate a large decision margin. Specifically, we propose the Orthogonal Softmax Layer (OSL), which makes the weight vectors in the classification layer remain orthogonal during both the training and test processes. The Rademacher complexity of a network using the OSL is only 1/K, where K is the number of classes, of that of a network using the fully connected classification layer, leading to a tighter generalization error bound. Experimental results demonstrate that the proposed OSL has better performance than the methods used for comparison on four small-sample benchmark datasets, as well as its applicability to large-sample datasets. Codes are available at: https://github.com/dongliangchang/OSLNet.
Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jingyi Yu 0001, Jun Guo 0002
IEEE Trans. Image Process.8
2020 Insights Into Multiple/Single Lower Bound Approximation for Extended Variational Inference in Non-Gaussian Structured Data Modeling
abstract
For most of the non-Gaussian statistical models, the data being modeled represent strongly structured properties, such as scalar data with bounded support (e.g., beta distribution), vector data with unit length (e.g., Dirichlet distribution), and vector data with positive elements (e.g., generalized inverted Dirichlet distribution). In practical implementations of non-Gaussian statistical models, it is infeasible to find an analytically tractable solution to estimating the posterior distributions of the parameters. Variational inference (VI) is a widely used framework in Bayesian estimation. Recently, an improved framework, namely, the extended VI (EVI), has been introduced and applied successfully to a number of non-Gaussian statistical models. EVI derives analytically tractable solutions by introducing lower bound approximations to the variational objective function. In this paper, we compare two approximation strategies, namely, the multiple lower bounds (MLBs) approximation and the single lower bound (SLB) approximation, which can be applied to carry out the EVI. For implementation, two different conditions, the weak and the strong conditions, are discussed. Convergence of the EVI depends on the selection of the lower bound, regardless of the choice of weak or strong condition. We also discuss the convergence properties to clarify the differences between MLB and SLB. Extensive comparisons are made based on some EVI-based non-Gaussian statistical models. Theoretical analysis is conducted to demonstrate the differences between the weak and strong conditions. Experimental results based on real data show advantages of the SLB approximation over the MLB approximation.
Zhanyu Ma, Jiyang Xie 0001, Yuping Lai, Jalil Taghia, Jing-Hao Xue, Jun Guo 0002
IEEE Trans. Neural Networks Learn. Syst.6
2019 Self-Supervised Convolutional Subspace Clustering Network
abstract
Subspace clustering methods based on data self-expression have become very popular for learning from data that lie in a union of low-dimensional linear subspaces. However, the applicability of subspace clustering has been limited because practical visual data in raw form do not necessarily lie in such linear subspaces. On the other hand, while Convolutional Neural Network (ConvNet) has been demonstrated to be a powerful tool for extracting discriminative features from visual data, training such a ConvNet usually requires a large amount of labeled data, which are unavailable in subspace clustering applications. To achieve simultaneous feature learning and subspace clustering, we propose an end-to-end trainable framework, called Self-Supervised Convolutional Subspace Clustering Network (S$^2$ConvSCN), that combines a ConvNet module (for feature learning), a self-expression module (for subspace clustering) and a spectral clustering module (for self-supervision) into a joint optimization framework. Particularly, we introduce a dual self-supervision that exploits the output of spectral clustering to supervise the training of the feature learning module (via a classification loss) and the self-expression module (via a spectral clustering loss). Our experiments on four benchmark datasets show the effectiveness of the dual self-supervision and demonstrate superior performance of our proposed approach.
Chun-Guang Li, Chong You, Xianbiao Qi, Honggang Zhang 0002, Jun Guo 0002, Zhouchen Lin
CVPR6
2019 Name Entity Recognition with Policy-Value Networks
abstract
In this paper we propose a novel reinforcement learning based model for named entity recognition (NER), referred to as MM-NER. Inspired by the methodology of the AlphaGo Zero, MM-NER formalizes the problem of named entity recognition with a Monte-Carlo tree search (MCTS) enhanced Markov decision process (MDP) model, in which the time steps correspond to the positions of words in a sentence from left to right, and each action corresponds to assign an NER tag to a word. Two Gated Recurrent Units (GRU) are used to summarize the past tag assignments and words in the sentence. Based on the outputs of GRUs, the policy for guiding the tag assignment and the value for predicting the whole tagging accuracy of the whole sentence are produced. The policy and value are then strengthened with MCTS, which takes the produced raw policy and value as inputs, simulates and evaluates the possible tag assignments at the subsequent positions, and outputs a better search policy for assigning tags. A reinforcement learning algorithm is proposed to train the model parameters. Empirically, we show that MM-NER can accurately predict the tags thanks to the exploratory decision making mechanism introduced by MCTS. It outperformed the conventional sequence tagging baselines and performed equally well with the state-of-the-art baseline BLSTM-CRF.
Yadi Lao, Jun Xu 0001, Sheng Gao 0001, Jun Guo 0002, Ji-Rong Wen
SIGIR4
2019 Finding Salient Context based on Semantic Matching for Relevance Ranking
abstract
We propose a salient-context based semantic matching method to improve relevance ranking in information retrieval. We first propose a new notion of salient context and then define how to measure it. Then we show how the most salient context can be located with a sliding window technique. Finally, we use the semantic similarity between a query term and the most salient context terms in a corpus of documents to rank those documents. Experiments on various TREC collections show the effectiveness of our model compared to the state-of-the-art methods.
Jiayue Zhang, Weiran Xu, Jun Guo 0002, Yan Li 0035
VCIP4
2019 Image-text dual neural network with decision strategy for small-sample image classification
Fangyi Zhu, Zhanyu Ma, Guang Chen 0003, Jen-Tzung Chien, Jing-Hao Xue, Jun Guo 0002
Neurocomputing7
2019 Compressive Binary Patterns: Designing a Robust Binary Face Descriptor with Random-Field Eigenfilters
abstract
A binary descriptor typically consists of three stages: image filtering, binarization, and spatial histogram. This paper first demonstrates that the binary code of the maximum-variance filtering responses leads to the lowest bit error rate under Gaussian noise. Then, an optimal eigenfilter bank is derived from a universal assumption on the local stationary random field. Finally, compressive binary patterns (CBP) is designed by replacing the local derivative filters of local binary patterns (LBP) with these novel random-field eigenfilters, which leads to a compact and robust binary descriptor that characterizes the most stable local structures that are resistant to image noise and degradation. A scattering-like operator is subsequently applied to enhance the distinctiveness of the descriptor. Surprisingly, the results obtained from experiments on the FERET, LFW, and PaSC databases show that the scattering CBP (SCBP) descriptor, which is handcrafted by only 6 optimal eigenfilters under restrictive assumptions, outperforms the state-of-the-art learning-based face descriptors in terms of both matching accuracy and robustness. In particular, on probe images degraded with noise, blur, JPEG compression, and reduced resolution, SCBP outperforms other descriptors by a greater than 10 percent accuracy margin.
Weihong Deng, Jiani Hu, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Variational Bayesian Learning for Dirichlet Process Mixture of Inverted Dirichlet Distributions in Non-Gaussian Image Feature Modeling
abstract
In this paper, we develop a novel variational Bayesian learning method for the Dirichlet process (DP) mixture of the inverted Dirichlet distributions, which has been shown to be very flexible for modeling vectors with positive elements. The recently proposed extended variational inference (EVI) framework is adopted to derive an analytically tractable solution. The convergency of the proposed algorithm is theoretically guaranteed by introducing single lower bound approximation to the original objective function in the EVI framework. In principle, the proposed model can be viewed as an infinite inverted Dirichlet mixture model that allows the automatic determination of the number of mixture components from data. Therefore, the problem of predetermining the optimal number of mixing components has been overcome. Moreover, the problems of overfitting and underfitting are avoided by the Bayesian estimation approach. Compared with several recently proposed DP-related methods and conventional applied methods, the good performance and effectiveness of the proposed method have been demonstrated with both synthesized data and real data evaluations.
Zhanyu Ma, Yuping Lai, W. Bastiaan Kleijn, Yi-Zhe Song, Liang Wang 0001, Jun Guo 0002
IEEE Trans. Neural Networks Learn. Syst.6
2018 SketchMate: Deep Hashing for Million-Scale Human Sketch Retrieval
abstract
We propose a deep hashing framework for sketch retrieval that, for the first time, works on a multi-million scale human sketch dataset. Leveraging on this large dataset, we explore a few sketch-specific traits that were otherwise under-studied in prior literature. Instead of following the conventional sketch recognition task, we introduce the novel problem of sketch hashing retrieval which is not only more challenging, but also offers a better testbed for large-scale sketch analysis, since: (i) more fine-grained sketch feature learning is required to accommodate the large variations in style and Abstraction, and (ii) a compact binary code needs to be learned at the same time to enable efficient retrieval. Key to our network design is the embedding of unique characteristics of human sketch, where (i) a two-branch CNN-RNN architecture is adapted to explore the temporal ordering of strokes, and (ii) a novel hashing loss is specifically designed to accommodate both the temporal and Abstract traits of sketches. By working with a 3.8M sketch dataset, we show that state-of-the-art hashing models specifically engineered for static images fail to perform well on temporal sketch data. Our network on the other hand not only offers the best retrieval performance on various code sizes, but also yields the best generalization performance under a zero-shot setting and when re-purposed for sketch recognition. Such superior performances effectively demonstrate the benefit of our sketch-specific design.
Peng Xu 0005, Yongye Huang, Tongtong Yuan, Kaiyue Pang, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Zhanyu Ma, Jun Guo 0002
CVPR9
2018 Constrained Sparse Subspace Clustering with Side-Information
abstract
Subspace clustering refers to the problem of segmenting high dimensional data drawn from a union of subspaces into the respective subspaces. In some applications, partial side-information to indicate “must-link” or “cannot-link” in clustering is available. This leads to the task of subspace clustering with side-information. However, in prior work the supervision value of the side-information for subspace clustering has not been fully exploited. To this end, in this paper, we present an enhanced approach for constrained subspace clustering with side-information, termed Constrained Sparse Subspace Clustering plus (CSSC+), in which the side-information is used not only in the stage of learning an affinity matrix but also in the stage of spectral clustering. Moreover, we propose to estimate clustering accuracy based on the partial side-information and theoretically justify the connection to the ground-truth clustering accuracy in terms of the Rand index. We conduct experiments on three cancer gene expression datasets to validate the effectiveness of our proposals.
Chun-Guang Li, Jun Guo 0002
ICPR3
2018 Supervised latent Dirichlet allocation with a mixture of sparse softmax
Zhanyu Ma, Feiyue Huang, Xiaojie Wang 0006, Jun Guo 0002
Neurocomputing7
2018 Corrigendum to "Supervised latent Dirichlet allocation with a mixture of sparse softmax" [Neurocomputing, volume 312, 27 October 2018, Pages 324-335]
Zhanyu Ma, Feiyue Huang, Xiaojie Wang 0006, Jun Guo 0002
Neurocomputing7
2018 A novel non-Gaussian embedding based model for recommender systems
Sheng Gao 0001, Qinjie Lyu, Jun Guo 0002, Patrick Gallinari
Neurocomputing4
2018 Cross-modal subspace learning for fine-grained sketch-based image retrieval
Peng Xu 0005, Qiyue Yin, Yongye Huang, Yi-Zhe Song, Zhanyu Ma, Liang Wang 0001, Tao Xiang 0002, W. Bastiaan Kleijn, Jun Guo 0002
Neurocomputing9
2018 Face Recognition via Collaborative Representation: Its Discriminant Nature and Superposed Representation
abstract
Collaborative representation methods, such as sparse subspace clustering (SSC) and sparse representation-based classification (SRC), have achieved great success in face clustering and classification by directly utilizing the training images as the dictionary bases. In this paper, we reveal that the superior performance of collaborative representation relies heavily on the sufficiently large class separability of the controlled face datasets such as Extended Yale B. On the uncontrolled or undersampled dataset, however, collaborative representation suffers from the misleading coefficients of the incorrect classes. To address this limitation, inspired by the success of linear discriminant analysis (LDA), we develop a superposed linear representation classifier (SLRC) to cast the recognition problem by representing the test image in term of a superposition of the class centroids and the shared intra-class differences. In spite of its simplicity and approximation, the SLRC largely improves the generalization ability of collaborative representation, and competes well with more sophisticated dictionary learning techniques, on the experiments of AR and FRGC databases. Enforced with the sparsity constraint, SLRC achieves the state-of-the-art performance on FERET database using single sample per person.
Weihong Deng, Jiani Hu, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 From one to many: Pose-Aware Metric Learning for single-sample face recognition
Weihong Deng, Jiani Hu, Zhongjun Wu, Jun Guo 0002
Pattern Recognit.4
2018 Decorrelation of Neutral Vector Variables: Theory and Applications
abstract
In this paper, we propose novel strategies for neutral vector variable decorrelation. Two fundamental invertible transformations, namely, serial nonlinear transformation and parallel nonlinear transformation, are proposed to carry out the decorrelation. For a neutral vector variable, which is not multivariate-Gaussian distributed, the conventional principal component analysis cannot yield mutually independent scalar variables. With the two proposed transformations, a highly negatively correlated neutral vector can be transformed to a set of mutually independent scalar variables with the same degrees of freedom. We also evaluate the decorrelation performances for the vectors generated from a single Dirichlet distribution and a mixture of Dirichlet distributions. The mutual independence is verified with the distance correlation measurement. The advantages of the proposed decorrelation strategies are intensively studied and demonstrated with synthesized data and practical application evaluations.
Zhanyu Ma, Jing-Hao Xue, Arne Leijon, Zheng-Hua Tan, Zhen Yang 0004, Jun Guo 0002
IEEE Trans. Neural Networks Learn. Syst.6
2018 Spoofing Detection in Automatic Speaker Verification Systems Using DNN Classifiers and Dynamic Acoustic Features
abstract
With the development of speech synthesis technology, automatic speaker verification (ASV) systems have encountered the serious challenge of spoofing attacks. In order to improve the security of ASV systems, many antispoofing countermeasures have been developed. In the front-end domain, much research has been conducted on finding effective features which can distinguish spoofed speech from genuine speech and the published results show that dynamic acoustic features work more effectively than static ones. In the back-end domain, Gaussian mixture model (GMM) and deep neural networks (DNNs) are the two most popular types of classifiers used for spoofing detection. The log-likelihood ratios (LLRs) generated by the difference of human and spoofing log-likelihoods are used as spoofing detection scores. In this paper, we train a five-layer DNN spoofing detection classifier using dynamic acoustic features and propose a novel, simple scoring method only using human log-likelihoods (HLLs) for spoofing detection. We mathematically prove that the new HLL scoring method is more suitable for the spoofing detection task than the classical LLR scoring method, especially when the spoofing speech is very similar to the human speech. We extensively investigate the performance of five different dynamic filter bank-based cepstral features and constant Q cepstral coefficients (CQCC) in conjunction with the DNN-HLL method. The experimental results show that, compared to the GMM-LLR method, the DNN-HLL method is able to significantly improve the spoofing detection accuracy. Compared with the CQCC-based GMM-LLR baseline, the proposed DNN-HLL model reduces the average equal error rate of all attack types to 0.045%, thus exceeding the performance of previously published approaches for the ASVspoof 2015 Challenge task. Fusing the CQCC-based DNN-HLL spoofing detection system with ASV systems, the false acceptance rate on spoofing attacks can be reduced significantly.
Hong Yu 0006, Zheng-Hua Tan, Zhanyu Ma, Rainer Martin 0001, Jun Guo 0002
IEEE Trans. Neural Networks Learn. Syst.5
2018 Spatial Pyramid-Based Statistical Features for Person Re-Identification: A Comprehensive Evaluation
abstract
Person re-identification (Re-Id) across nonoverlapping camera views is one of challenging problems in surveillance video analysis. The difficulties in person Re-Id mainly come from the large appearance variations caused by camera view angle, human pose, illumination, and occlusion. Recently, extensive efforts have been cast into addressing this problem by developing invariant features or discriminative distance metrics. However, there is still a lack of systematic evaluations on the pipeline for feature extraction and combination. In this paper, we propose a spatial pyramid-based statistical feature extraction framework as a unified pipeline of feature extraction and combination for person Re-Id, and systematically evaluate the configuration details in feature extraction and the fusion strategies in feature combination. Extensive experiments on benchmark datasets demonstrate the critical components in feature extraction. Moreover, by combining multiple features, our proposed approach can yield state-of-the-art performance. It should be mentioned that our approach achieves rank 1 matching rate of 45.8% on dataset VIPeR and 61.5% on dataset CUHK01, respectively.
Jianlou Si, Honggang Zhang 0002, Chun-Guang Li, Jun Guo 0002
IEEE Trans. Syst. Man Cybern. Syst.4
2018 A Survey on Machine Learning-Based Mobile Big Data Analysis: Challenges and Applications
abstract
This paper attempts to identify the requirement and the development of machine learning‐based mobile big data (MBD) analysis through discussing the insights of challenges in the mobile big data. Furthermore, it reviews the state‐of‐the‐art applications of data analysis in the area of MBD. Firstly, we introduce the development of MBD. Secondly, the frequently applied data analysis methods are reviewed. Three typical applications of MBD analysis, namely, wireless channel modeling, human online and offline behavior analysis, and speech recognition in the Internet of Vehicles, are introduced, respectively. Finally, we summarize the main challenges and future development directions of mobile big data analysis.
Jiyang Xie 0001, Yanting Zhang 0001, Hong Yu 0006, Jinnan Zhan, Zhanyu Ma, Yuanyuan Qiao 0002, Jianhua Zhang 0001, Jun Guo 0002
Wirel. Commun. Mob. Comput.10
2017 Designing an adaptive attention mechanism for relation classification
abstract
Entity pair provide essential information for identifying relation type. Aiming at this characteristic, Position Feature is widely used in current relation classification systems to highlight the words close to them. However, semantic knowledge involved in entity pair has not been fully utilized. To overcome this issue, we propose an Entity-pair-based Attention Mechanism, which is specially designed for relation classification. Recently, attention mechanism significantly promotes the development of deep learning in NLP. Inspired by this, for specific instance(entity pair, sentence), the corresponding entity pair information is incorporated as prior knowledge to adaptively compute attention weights for generating sentence representation. Experimental results on SemEval-2010 Task 8 dataset show that our method outperforms most of the state-of-the-art models, without external linguistic features.
Pengda Qin, Weiran Xu, Jun Guo 0002
IJCNN3
2017 Adversarial Network Bottleneck Features for Noise Robust Speaker Verification
abstract
In this paper, we propose a noise robust bottleneck feature representation which is generated by an adversarial network (AN).The AN includes two cascade connected networks, an encoding network (EN) and a discriminative network (DN).Melfrequency cepstral coefficients (MFCCs) of clean and noisy speech are used as input to the EN and the output of the EN is used as the noise robust feature.The EN and DN are trained in turn, namely, when training the DN, noise types are selected as the training labels and when training the EN, all labels are set as the same, i.e., the clean speech label, which aims to make the AN features invariant to noise and thus achieve noise robustness.We evaluate the performance of the proposed feature on a Gaussian Mixture Model-Universal Background Model based speaker verification system, and make comparison to MFCC features of speech enhanced by short-time spectral amplitude minimum mean square error (STSA-MMSE) and deep neural network-based speech enhancement (DNN-SE) methods.Experimental results on the RSR2015 database show that the proposed AN bottleneck feature (AN-BN) dramatically outperforms the STSA-MMSE and DNN-SE based MFCCs for different noise types and signal-to-noise ratios.Furthermore, the AN-BN feature is able to improve the speaker verification performance under the clean condition.
Hong Yu 0006, Zheng-Hua Tan, Zhanyu Ma, Jun Guo 0002
INTERSPEECH4
2017 A Targeted Retraining Scheme of Unsupervised Word Embeddings for Specific Supervised Tasks
Pengda Qin, Weiran Xu, Jun Guo 0002
PAKDD (2)3
2017 Product ranking using hierarchical aspect structures
Si Li 0001, Zhaoyan Ming, Yan Leng, Jun Guo 0002
J. Intell. Inf. Syst.4
2017 A temporal model in Electronic Health Record search
Jiayue Zhang, Weiran Xu, Jun Guo 0002, Sheng Gao 0001
Knowl. Based Syst.3
2017 Lighting-aware face frontalization for unconstrained face recognition
Weihong Deng, Jiani Hu, Zhongjun Wu, Jun Guo 0002
Pattern Recognit.4
2017 Fine-grained face verification: FGLFW database, baselines, and human-DCMN partnership
Weihong Deng, Jiani Hu, Nanhai Zhang, Binghui Chen, Jun Guo 0002
Pattern Recognit.5
2017 A Novel Embedding Method for Information Diffusion Prediction in Social Network Big Data
abstract
With the increase of social networking websites and the interaction frequency among users, the prediction of information diffusion is required to support effective generalization and efficient inference in the context of social big data era. However, the existing models either rely on expensive probabilistic modeling of information diffusion based on partially known network structures, or discover the implicit structures of diffusion from users' behaviors without considering the impacts of different diffused contents. To address the issues, in this paper, we propose a novel information-dependent embedding-based diffusion prediction (IEDP) model to map the users in observed diffusion process into a latent embedding space, then the temporal order of users with the timestamps in the cascade can be preserved by the embedding distance of users. Our proposed model further learns the propagation probability of information in the cascade as a function of the relative positions of information-specific user embeddings in the information-dependent subspace. Then, the problem of temporal propagation prediction can be converted into the task of spatial probability learning in the embedding space. Moreover, we present an efficient margin-based optimization algorithm with a fast computation to make the inference of the information diffusion in the latent embedding space. When applying our proposed method to several social network datasets, the experimental results show the effectiveness of our proposed approach for the information diffusion prediction and the efficiency with respect to the inference speed compared with the state-of-the-art methods.
Sheng Gao 0001, Huacan Pang, Patrick Gallinari, Jun Guo 0002, Nei Kato
IEEE Trans. Ind. Informatics4
2017 HEp-2 Cell Classification via Combining Multiresolution Co-Occurrence Texture and Large Region Shape Information
abstract
Indirect immunofluorescence imaging of human epithelial type 2 (HEp-2) cell image is an effective evidence to diagnose autoimmune diseases. Recently, computer-aided diagnosis of autoimmune diseases by the HEp-2 cell classification has attracted great attention. However, the HEp-2 cell classification task is quite challenging due to large intraclass and small interclass variations. In this paper, we propose an effective approach for the automatic HEp-2 cell classification by combining multiresolution co-occurrence texture and large regional shape information. To be more specific, we propose to: 1) capture multiresolution co-occurrence texture information by a novel pairwise rotation-invariant co-occurrence of local Gabor binary pattern descriptor; 2) depict large regional shape information by using an improved Fisher vector model with RootSIFT features, which are sampled from large image patches in multiple scales; and 3) combine both features. We evaluate systematically the proposed approach on the IEEE International Conference on Pattern Recognition (ICPR) 2012, the IEEE International Conference on Image Processing (ICIP) 2013, and the ICPR 2014 contest datasets. The proposed method based on the combination of the introduced two features outperforms the winners of the ICPR 2012 contest using the same experimental protocol. Our method also greatly improves the winner of the ICIP 2013 contest under four different experimental setups. Using the leave-one-specimen-out evaluation strategy, our method achieves comparable performance with the winner of the ICPR 2014 contest that combined four features.
Xianbiao Qi, Guoying Zhao 0001, Chun-Guang Li, Jun Guo 0002, Matti Pietikäinen
IEEE J. Biomed. Health Informatics4
2016 Low-rank and structured sparse subspace clustering
abstract
High dimensional data often lie approximately in low dimensional subspaces corresponding to multiple classes or categories. Segmenting the high dimensional data into their corresponding low dimensional subspaces is referred as subspace clustering. State of the art methods solve this problem in two steps. First, an affinity matrix is built from data based on self-expressiveness model, in which each data point is expressed as a linear combination of other data points. Second, the segmentation is obtained by spectral clustering. However, solving two dependent steps separately is still suboptimal. In this paper, we propose a joint affinity learning and spectral clustering approach for low-rank representation based subspace clustering, termed Low-Rank and Structured Sparse Subspace Clustering (LRS3C), where a subspace structured norm that depends on subspace clustering result is introduced into the objective of low-rank representation problem. We solve it efficiently via a combination of Linearized Alternation Direction Method (LADM) with spectral clustering. Experiments on Hopkins 155 motion segmentation database and Extended Yale B data set demonstrated the effectiveness of our method.
Chun-Guang Li, Honggang Zhang 0002, Jun Guo 0002
VCIP4
2016 Feature selection for neutral vector in EEG signal classification
Zhanyu Ma, Zheng-Hua Tan, Jun Guo 0002
Neurocomputing3
2016 A novel negative sampling based on TFIDF for learning word representation
Pengda Qin, Weiran Xu, Jun Guo 0002
Neurocomputing3
2016 An empirical convolutional neural network approach for semantic relation classification
Pengda Qin, Weiran Xu, Jun Guo 0002
Neurocomputing3
2016 A Comprehensive Review of Smart Energy Meters in Intelligent Energy Networks
abstract
The significant increase in energy consumption and the rapid development of renewable energy, such as solar power and wind power, have brought huge challenges to energy security and the environment, which, in the meantime, stimulate the development of energy networks toward a more intelligent direction. Smart meters are the most fundamental components in the intelligent energy networks (IENs). In addition to measuring energy flows, smart energy meters can exchange the information on energy consumption and the status of energy networks between utility companies and consumers. Furthermore, smart energy meters can also be used to monitor and control home appliances and other devices according to the individual consumer's instruction. This paper systematically reviews the development and deployment of smart energy meters, including smart electricity meters, smart heat meters, and smart gas meters. By examining various functions and applications of smart energy meters, as well as associated benefits and costs, this paper provides insights and guidelines regarding the future development of smart meters.
Qie Sun, Zhanyu Ma, Chao Wang 0015, Javier Campillo, Qi Zhang 0019, Fredrik Wallin, Jun Guo 0002
IEEE Internet Things J.8
2016 On the Outage Probability of Device-to-Device-Communication-Enabled Multichannel Cellular Networks: An RSS-Threshold-Based Perspective
abstract
In this paper, we study the outage probability of device-to-device (D2D)-communication-enabled cellular networks from a general threshold-based perspective. Specifically, a mobile user equipment (TIE) transmits in D2D mode if the received signal strength (RSS) from the nearest base station (BS) is less than a specified threshold β ≥ 0; otherwise, it connects to the nearest BS and transmits in cellular mode. The RSS-threshold-based setting is general in the sense that by varying β from β = 0 to β = ∞, the network accordingly evolves from a traditional cellular network (including only cellular mode) toward a wireless ad hoc network (including only D2D mode). We provide a unified framework to analyze the downlink outage probability in a multichannel environment with Rayleigh fading, where the spatial distributions of BSs and TIEs are well explicitly accounted for by utilizing stochastic geometry. We derive closed-form expressions for the outage probability of a generic TIE and that in both cellular mode and D2D mode and quantify the performance gains in outage probability that can be obtained by allowing such RSS-thresholdbased D2D communications. We show that increasing the number of channels, although able to support more cellular TIEs, may result in an increase of outage probability in the D2D-enabled cellular network. The corresponding condition and reason are also identified by applying our framework.
Jiajia Liu 0001, Hiroki Nishiyama 0001, Nei Kato, Jun Guo 0002
IEEE J. Sel. Areas Commun.4
2016 Aspect-based latent factor model by integrating ratings and reviews for recommender system
Sheng Gao 0001, Jun Guo 0002
Knowl. Based Syst.4
2015 VecLP: A Realtime Video Recommendation System for Live TV Programs
abstract
We propose VecLP, a novel Internet Video recommendation system working for Live TV Programs in this paper. Given little information on the live TV programs, our proposed VecLP system can effectively collect necessary information on both the programs and the subscribers as well as a large volume of related online videos, and then recommend the relevant Internet videos to the subscribers. For that, the key frames are firstly detected from the live TV programs, and then visual and textual features are extracted from these frames to enhance the understanding of the TV broadcasts. Furthermore, by utilizing the subscribers' profiles and their social relationships, a user preference model is constructed, which greatly improves the diversity of the recommendations in our system. The subscriber's browsing history is also recorded and used to make a further personalized recommendation. This work also illustrates how our proposed VecLP system makes it happen. Finally, we dispose some sort of new recommendation strategies in use at the system to meet special needs from diverse live TV programs and throw light upon how to fuse these strategies.
Sheng Gao 0001, Honggang Zhang 0002, Jianxin Liao, Jun Guo 0002
AAAI7
2015 Improving Cross-Domain Recommendation through Probabilistic Cluster-Level Latent Factor Model
abstract
Cross-domain recommendation has been proposed to transfer user behavior pattern by pooling together the rating data from multiple domains to alleviate the sparsity problem appearing in single rating domains. However, previous models only assume that multiple domains share a latent common rating pattern based on the user-item co-clustering. To capture diversities among different domains, we propose a novel Probabilistic Cluster-level Latent Factor (PCLF) model to improve the cross-domain recommendation performance. Experiments on several real world datasets demonstrate that our proposed model outperforms the state-of-the-art methods for the cross-domain recommendation task.
Siting Ren, Sheng Gao 0001, Jianxin Liao, Jun Guo 0002
AAAI4
2015 Making better use of edges via perceptual grouping
abstract
We propose a perceptual grouping framework that organizes image edges into meaningful structures and demonstrate its usefulness on various computer vision tasks. Our grouper formulates edge grouping as a graph partition problem, where a learning to rank method is developed to encode probabilities of candidate edge pairs. In particular, RankSVM is employed for the first time to combine multiple Gestalt principles as cue for edge grouping. Afterwards, an edge grouping based object proposal measure is introduced that yields proposals comparable to state-of-the-art alternatives. We further show how human-like sketches can be generated from edge groupings and consequently used to deliver state-of-the-art sketch-based image retrieval performance. Last but not least, we tackle the problem of freehand human sketch segmentation by utilizing the proposed grouper to cluster strokes into semantic object parts.
Yonggang Qi, Yi-Zhe Song, Tao Xiang 0002, Honggang Zhang 0002, Timothy M. Hospedales, Yi Li 0004, Jun Guo 0002
CVPR7
2015 Learning Semi-Supervised Representation Towards a Unified Optimization Framework for Semi-Supervised Learning
abstract
State of the art approaches for Semi-Supervised Learning (SSL) usually follow a two-stage framework -- constructing an affinity matrix from the data and then propagating the partial labels on this affinity matrix to infer those unknown labels. While such a two-stage framework has been successful in many applications, solving two subproblems separately only once is still suboptimal because it does not fully exploit the correlation between the affinity and the labels. In this paper, we formulate the two stages of SSL into a unified optimization framework, which learns both the affinity matrix and the unknown labels simultaneously. In the unified framework, both the given labels and the estimated labels are used to learn the affinity matrix and to infer the unknown labels. We solve the unified optimization problem via an alternating direction method of multipliers combined with label propagation. Extensive experiments on a synthetic data set and several benchmark data sets demonstrate the effectiveness of our approach.
Chun-Guang Li, Zhouchen Lin, Honggang Zhang 0002, Jun Guo 0002
ICCV4
2015 A stochastic geometry analysis of D2D overlaying multi-channel downlink cellular networks
abstract
Based on the tool of stochastic geometry, we present in this paper a framework for analyzing the coverage probability and ergodic rate in a D2D overlaying multi-channel downlink cellular network. Different from previous works, 1) we consider a flexible new scheme for mobile UEs to select operation mode individually, under which a mobile UE decides to establish a cellular link (with a BS) or a D2D link (with a neighboring UE) based on the pilot signal strength received from its nearest BS; 2) we allow a mobile UE which is located far from BSs to connect to a nearby BS via another intermediate UE in a two-hop manner. Our results indicate that the developed framework is very helpful for network designers to efficiently determine the optimal network parameters at which the optimum system performance can be achieved. Furthermore, as corroborated by extensive numerical results, enabling the D2D link based two-hop connection can significantly improve the network coverage performance, especially for the low SIR regime.
Jiajia Liu 0001, Shangwei Zhang, Hiroki Nishiyama 0001, Nei Kato, Jun Guo 0002
INFOCOM5
2015 Activation force-based air pollution observation station clustering
Di Huang 0006, Hong Yu 0006, Huanyu Zhou, Zhanyu Ma, Weisong Hu, Jun Guo 0002
QSHINE7
2015 A multi-level system for sequential update summarization
Chunyun Zhang, Zhanyu Ma, Jiayue Zhang, Weiran Xu, Jun Guo 0002
QSHINE5
2015 DeepEmo: Real-world facial expression analysis via deep learning
abstract
Recent automatic facial expression recognition research has focused on optimizing performance on a few databases that were collected under controlled pose and lighting conditions, and has produced nearly perfect accuracy. This paper explores the necessary characteristics of the training dataset, feature representations and machine learning algorithms for a system that operates reliably in more realistic conditions. A new database, Real-world Affective Face Database (RAF-DB), is presented which contains about 30,000 greatly-diverse facial images from social networks. Crowdsourcing results suggest that real-world expression recognition problem is a typical imbalanced multi-label classification problem, and the balanced, single-label datasets currently used in the literature could potentially lead research into misleading algorithmic solutions. A deep learning architecture, DeepEmo, is proposed to address the real-world challenge of emotion recognition by learning the highlevel feature representations which are highly effective for discriminating realistic facial expressions. Extensive experimental results show that the deep learning method is significantly superior to handcrafted features, and with the near-frontal pose constraint, human-level recognition accuracy is achievable.
Weihong Deng, Jiani Hu, Jun Guo 0002
VCIP4
2015 Improving tag matrix completion for image annotation and retrieval
abstract
Image annotation is a fundamental and challenging task in the field of semantic image retrieval. In this paper, we deal with image annotation via matrix completion. Concretely, we formulate the problem of annotating the tags of an image into a constrained optimization problem, in which the constraint is to keep the consistency with the given initial labels and the objective is to minimize the discrepancy between the correlation in visual content and the correlation in semantic tags. We solve the optimization problem with the linearized alternating direction method. Experimental results on benchmark data demonstrate the effectiveness of our proposals.
Zhen Qin 0001, Chun-Guang Li, Honggang Zhang 0002, Jun Guo 0002
VCIP4
2015 Knowledge base completion by learning pairwise-interaction differentiated embeddings
Yu Zhao 0019, Sheng Gao 0001, Patrick Gallinari, Jun Guo 0002
Data Min. Knowl. Discov.4
2015 Im2Sketch: Sketch generation by unconflicted perceptual grouping
Yonggang Qi, Jun Guo 0002, Yi-Zhe Song, Tao Xiang 0002, Honggang Zhang 0002, Zheng-Hua Tan
Neurocomputing2
2015 Multi-label learning with prior knowledge for facial expression analysis
Kaili Zhao, Honggang Zhang 0002, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002
Neurocomputing5
2015 Construction of semantic bootstrapping models for relation extraction
Chunyun Zhang, Weiran Xu, Zhanyu Ma, Sheng Gao 0001, Qun Li 0002, Jun Guo 0002
Knowl. Based Syst.6
2015 Mining activation force defined dependency patterns for relation extraction
Chunyun Zhang, Yichang Zhang, Weiran Xu, Zhanyu Ma, Yan Leng, Jun Guo 0002
Knowl. Based Syst.6
2015 A holistic model of mining product aspects and associated sentiments from online reviews
Yan Li 0035, Zhen Qin 0001, Weiran Xu, Jun Guo 0002
Multim. Tools Appl.4
2015 Variational Bayesian Matrix Factorization for Bounded Support Data
abstract
A novel Bayesian matrix factorization method for bounded support data is presented. Each entry in the observation matrix is assumed to be beta distributed. As the beta distribution has two parameters, two parameter matrices can be obtained, which matrices contain only nonnegative values. In order to provide low-rank matrix factorization, the nonnegative matrix factorization (NMF) technique is applied. Furthermore, each entry in the factorized matrices, i.e., the basis and excitation matrices, is assigned with gamma prior. Therefore, we name this method as beta-gamma NMF (BG-NMF). Due to the integral expression of the gamma function, estimation of the posterior distribution in the BG-NMF model can not be presented by an analytically tractable solution. With the variational inference framework and the relative convexity property of the log-inverse-beta function, we propose a new lower-bound to approximate the objective function. With this new lower-bound, we derive an analytically tractable solution to approximately calculate the posterior distributions. Each of the approximated posterior distributions is also gamma distributed, which retains the conjugacy of the Bayesian estimation. In addition, a sparse BG-NMF can be obtained by including a sparseness constraint to the gamma prior. Evaluations with synthetic data and real life data demonstrate the good performance of the proposed method.
Zhanyu Ma, Andrew E. Teschendorff, Arne Leijon, Yuanyuan Qiao 0002, Honggang Zhang 0002, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2015 Line spectral frequencies modeling by a mixture of von Mises-Fisher distributions
Zhanyu Ma, Jalil Taghia, W. Bastiaan Kleijn, Arne Leijon, Jun Guo 0002
Signal Process.5
2014 Transformed Principal Gradient Orientation for Robust and Precise Batch Face Alignment
Weihong Deng, Jiani Hu, Jun Guo 0002
ACCV (4)4
2014 Linear Ranking Analysis
abstract
We extend the classical linear discriminant analysis (LDA) technique to linear ranking analysis (LRA), by considering the ranking order of classes centroids on the projected subspace. Under the constrain on the ranking order of the classes, two criteria are proposed: 1) minimization of the classification error with the assumption that each class is homogenous Guassian distributed, 2) maximization of the sum (average) of the K minimum distances of all neighboring-class (centroid) pairs. Both criteria can be efficiently solved by the convex optimization for one-dimensional subspace. Greedy algorithm is applied to extend the results to the multi-dimensional subspace. Experimental results show that 1) LRA with both criteria achieve state-of-the-art performance on the tasks of ranking learning and zero-shot learning, and 2) the maximum margin criterion provides a discriminative subspace selection method, which can significantly remedy the class separation problem in comparing with several representative extensions of LDA.
Weihong Deng, Jiani Hu, Jun Guo 0002
CVPR3
2014 Nonlinear estimation of missing ΔLSF parameters by a mixture of Dirichlet distributions
abstract
In packet networks, a reliable scheme to handle packet loss during speech transmission is of great importance. As a common representation of the linear predictive coding (LPC) model, the line spectral frequency (LSF) parameters are widely used in speech quantization and transmission. In this paper, we propose a novel scheme to estimate the missing values occurring during LPC model transmission. In order to exploit the boundary and ordering properties of the LSF parameters, we utilize the ΔLSF representation and apply the Dirichlet mixture model (DMM) to capture the correlations among the elements in the ΔLSF vector. With the conditional distribution of the missing part given the received part, an optimal nonlinear minimum mean square error estimator for the missing values is proposed. Compared to the previously presented Gaussian mixture model based method, the proposed DMM based nonlinear estimator shows a convincing improvement.
Zhanyu Ma, Rainer Martin 0001, Jun Guo 0002, Honggang Zhang 0002
ICASSP3
2014 An adaptive group lasso based multi-label regression approach for facial expression analysis
abstract
In the realm of facial expression analysis, numerous attempts have been made to link each facial picture to one affective category. Nevertheless, in our daily life, few of the facial expressions are exactly one of the predefined affective states. Therefore, to analyze the facial expressions more effectively, this paper proposes an Adaptive Group Lasso based Multilabel Regression approach, which depicts each facial expression with multiple continuous values of predefined affective states. Adaptive Group Lasso is adopted to depict the relationship between different labels which different facial expressions share some same affective facial areas (patches). Moreover, to solve the multi-label regression problem, a convex optimization formulation is presented, which would guarantee a global optimal solution. The experiment results based on JAFFE dataset have verified the superior performance of our approach.
Kaili Zhao, Honggang Zhang 0002, Jun Guo 0002
ICIP3
2014 Online Regression of Grandmother-Cell Responses with Visual Experience Learning for Face Recognition
abstract
Grandmother cell is a term in neuroscience to imitate the simplistic notion that the brain has a separate neuron to represent every familiar face, with important properties of sparseness and invariance. This paper proposes a linear regression based classification model for face recognition, which learn a mapping from the training feature vectors to the grandmother-cell-like codes, with one unit corresponding to an individual. Two kinds of visual experiences are incorporated to enhance the generalization capability of the regression mapping. First, the regression model maps the intra-personal facial differences of the unknown faces to the zeros vectors, so that any similar variation on the familiar face would not affect the regression result. Second, to adapt to the evolution of facial appearance, the model feeds the selected testing images back to incrementally retrain the regression mapping, and decrement ally remove the influence of outdated training images, all in an unsupervised manner. Experiments results on Extended Yale B, FERET, and AR databases demonstrate the efficacy of the proposed regression based face recognition algorithms.
Jiani Hu, Weihong Deng, Jun Guo 0002
ICPR3
2014 Max-K-Min Distance Analysis for Dimension Reduction
abstract
We propose a new criterion for discriminative dimension reduction, Max-K-Min Distance Analysis (MKMDA). Given a data set with C classes, MKMDA maximizes the sum of the K minimum pair wise distance of these C classes on the selected one-dimensional subspace. The set of the possible one-dimensional subspace, for which the order of the projected class centroids is identical, define a convex region with associated convex sum of K smallest margin functions. This allows for the maximization of the margin function using standard convex optimization algorithms. This result is further extended to obtain the d-dimensional subspace for any given d by iterative applying our algorithm to the null space of the (d -- 1)-dimensional subspace. The effectiveness of the proposed criterion and corresponding algorithm is shown by the visualization and classification experiments on both synthetic data and real data sets.
Jiani Hu, Weihong Deng, Jun Guo 0002
ICPR3
2014 Confidence Estimation and Reputation Analysis in Aspect Extraction
abstract
Extracting product aspects and their associated sentiments is one of the key tasks in sentiment analysis. Estimating the confidences of extracted aspects is important to ensure the performance. To tackle the issue, this paper proposes a two-step estimation method. Collocations of product features and opinion words are initially extracted through pattern bootstrapping. A criterion synthesizing two measurements, Popularity and Reliability, is novelly exploited to assess both patterns and features. Then the features are further clustered into aspects based on path similarities in the Word Net. Each cluster is assigned a weight based on its Compactness and Texture, and the light ones are filtered out. In addition, this paper also captures global aspect reputations by aggregating sentiment strengths through opinion collocations. Experimental results on a benchmark data set with 5 products demonstrate the effectiveness and reliability of our proposed method.
Yan Li 0035, Zhen Qin 0001, Weiran Xu, Jun Guo 0002
ICPR5
2014 A Feature Extraction Method Based on Word Embedding for Word Similarity Computing
Weitai Zhang, Weiran Xu, Guang Chen 0003, Jun Guo 0002
NLPCC4
2014 A multimedia information fusion framework for web image categorization
Wenting Lu, Lei Li 0001, Tao Li 0001, Honggang Zhang 0002, Jun Guo 0002
Multim. Tools Appl.6
2014 A high-performance training-free approach for hand gesture recognition with accelerometer
Mingzhi Dong, Ying Duan, Weihong Deng, Kaili Zhao, Jun Guo 0002
Multim. Tools Appl.6
2014 Transform-Invariant PCA: A Unified Approach to Fully Automatic FaceAlignment, Representation, and Recognition
abstract
We develop a transform-invariant PCA (TIPCA) technique which aims to accurately characterize the intrinsic structures of the human face that are invariant to the in-plane transformations of the training images. Specially, TIPCA alternately aligns the image ensemble and creates the optimal eigenspace, with the objective to minimize the mean square error between the aligned images and their reconstructions. The learning from the FERET facial image ensemble of 1,196 subjects validates the mutual promotion between image alignment and eigenspace representation, which eventually leads to the optimized coding and recognition performance that surpasses the handcrafted alignment based on facial landmarks. Experimental results also suggest that state-of-the-art invariant descriptors, such as local binary pattern (LBP), histogram of oriented gradient (HOG), and Gabor energy filter (GEF), and classification methods, such as sparse representation based classification (SRC) and support vector machine (SVM), can benefit from using the TIPCA-aligned faces, instead of the manually eye-aligned faces that are widely regarded as the ground-truth alignment. Favorable accuracies against the state-of-the-art results on face coding and face recognition are reported.
Weihong Deng, Jiani Hu, Jiwen Lu, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 Pairwise Rotation Invariant Co-Occurrence Local Binary Pattern
abstract
Designing effective features is a fundamental problem in computer vision. However, it is usually difficult to achieve a great tradeoff between discriminative power and robustness. Previous works shown that spatial co-occurrence can boost the discriminative power of features. However the current existing co-occurrence features are taking few considerations to the robustness and hence suffering from sensitivity to geometric and photometric variations. In this work, we study the Transform Invariance (TI) of co-occurrence features. Concretely we formally introduce a Pairwise Transform Invariance (PTI) principle, and then propose a novel Pairwise Rotation Invariant Co-occurrence Local Binary Pattern (PRICoLBP) feature, and further extend it to incorporate multi-scale, multi-orientation, and multi-channel information. Different from other LBP variants, PRICoLBP can not only capture the spatial context co-occurrence information effectively, but also possess rotation invariance. We evaluate PRICoLBP comprehensively on nine benchmark data sets from five different perspectives, e.g., encoding strategy, rotation invariance, the number of templates, speed, and discriminative power compared to other LBP variants. Furthermore we apply PRICoLBP to six different but related applications-texture, material, flower, leaf, food, and scene classification, and demonstrate that PRICoLBP is efficient, effective, and of a well-balanced tradeoff between the discriminative power and robustness.
Xianbiao Qi, Rong Xiao 0003, Chun-Guang Li, Yu Qiao 0001, Jun Guo 0002, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.5
2014 Equidistant prototypes embedding for single sample based face recognition with generic learning and incremental learning
Weihong Deng, Jiani Hu, Xiuzhuang Zhou, Jun Guo 0002
Pattern Recognit.4
2014 Dirichlet mixture modeling to estimate an empirical lower bound for LSF quantization
Zhanyu Ma, Saikat Chatterjee, W. Bastiaan Kleijn, Jun Guo 0002
Signal Process.4
2013 A Maximum K-Min Approach for Classification
Mingzhi Dong, Weihong Deng, Jun Guo 0002, Honggang Zhang 0002
AAAI5
2013 Promoting electronic health record search through a time-aware approach
abstract
In this paper, we propose a time-aware approach to promoting textual retrieval performance for Electronic Health Record (EHR) search. The proposed approach focuses on identifying patients cohorts from the perspective of EHR temporal correlation. First, an EHR temporal profile is created according to EHR distribution on time interval for each patient. Second, the temporal similarity is computed and used as a feature for discovering temporal cohorts. In each cohort, the highest-ranked profile in textual retrieval is considered as the centroid, and a temporal relevance score is computed by multiplying temporal similarity with the textual relevance of the centroid. Finally, the temporal relevance is combined linearly with the textual relevance for re-ranking. Extensive experiments are conducted to demonstrate the effectiveness of the proposed approach in promoting retrieval performance for EHR search.
Jiayue Zhang, Jimmy Huang 0001, Jun Guo 0002, Weiran Xu
BIBM3
2013 Multi-scale Joint Encoding of Local Binary Patterns for Texture and Material Classification
Xianbiao Qi, Yu Qiao 0001, Chun-Guang Li, Jun Guo 0002
BMVC4
2013 Exploring Cross-Channel Texture Correlation for Color Texture Classification
abstract
This paper proposes a novel approach to encode cross-channel texture correlation for color texture classification task. Firstly, we quantitatively study the correlation between different color channels using Local Binary Pattern (LBP) as the texture descriptor and using Shannon’s information theory to measure the correlation. We find that (R, G) channel pair exhibits stronger correlation than (R, B) and (G, B) channel pairs. Secondly, we propose a novel descriptor to encode the cross-channel texture correlation. The proposed descriptor can capture well the relative variance of texture patterns between different channels. Meanwhile, our descriptor is computationally efficient and robust to image rotation. We conduct extensive experiments on four challenging color texture databases to validate the effectiveness of the proposed approach. The experimental results show that the proposed approach significantly outperforms its mostly relevant counterpart (Multichannel color LBP), and achieves the state-of-the-art performance.
Xianbiao Qi, Yu Qiao 0001, Chun-Guang Li, Jun Guo 0002
BMVC4
2013 In Defense of Sparsity Based Face Recognition
abstract
The success of sparse representation based classification (SRC) has largely boosted the research of sparsity based face recognition in recent years. A prevailing view is that the sparsity based face recognition performs well only when the training images have been carefully controlled and the number of samples per class is sufficiently large. This paper challenges the prevailing view by proposing a ``prototype plus variation'' representation model for sparsity based face recognition. Based on the new model, a Superposed SRC (SSRC), in which the dictionary is assembled by the class centroids and the sample-to-centroid differences, leads to a substantial improvement on SRC. The experiments results on AR, FERET and FRGC databases validate that, if the proposed prototype plus variation representation model is applied, sparse coding plays a crucial role in face recognition, and performs well even when the dictionary bases are collected under uncontrolled conditions and only a single sample per classes is available.
Weihong Deng, Jiani Hu, Jun Guo 0002
CVPR3
2013 Latent Factor BlockModel for Modelling Relational Data
Sheng Gao 0001, Ludovic Denoyer, Patrick Gallinari, Jun Guo 0002
ECIR4
2013 Local alignment for query by humming
abstract
Query by humming (QBH) allows users to retrieve songs by humming a clip. In the previous work, the query has been regarded as a fragment of the music, so the task of QBH is considered to find a subsequence, which is most similar to the whole query, from the database. Taking into account humming errors, especially at the beginning or ending of the query, we assume that only part of the query is a subsequence of the music. Based on this assumption, we propose a local alignment framework which searches for the best match common subsequence between the query and database music. To verify the effectiveness of local alignment, two popular match algorithms, i.e. Linear Scaling and Dynamic Time Warping, are extended to identify the common subsequence. Experimental results on the 2010 MIREX-QBH corpus show that the new algorithms improve the retrieval accuracy significantly.
Qiang Wang 0048, Gang Liu 0008, Chun-Guang Li, Jun Guo 0002
ICASSP5
2013 Ordered histogram of shapemes: An ordered bag-of-features based shape descriptor for efficient shape matching
abstract
In this paper, we enhance the Shape Context-based descriptor, shapemes, by introducing an ordered bag-of-features model and dynamic programming. The proposed descriptor consists of a series of sub-histograms of shapemes, each of which represents a subset of sampled points. The division of the sampled points is based on their sequential positions on the contour of the shape, so the representation has intrinsic order and is therefore named ordered histogram of shapemes. Then dynamic programming is utilized for descriptor matching. The framework is effective and efficient owing to the following properties: 1) points division approach together with dynamic programming for invariance under the change of starting point, 2) Earth Mover's Distance for discriminative power, and 3) pre-caculated shapemes dissimilarity matrix for fast descriptor distance calculation. Experiments on standard shape database and real world application scenario demonstrate the effectiveness and efficiency of the descriptor and the matching framework. We make our code and experimental data publicly available for future reference.
Lunshao Chai, Zhen Qin 0001, Qun Li 0002, Honggang Zhang 0002, Jun Guo 0002
ICIP5
2013 Representative reference-set and betweenness centrality for scene image categorization
abstract
Reference-based image classification approach introduces a reference-set for both image representation and dictionary learning. It significantly reduces the dimensionality of represented images and shows outstanding performance even with randomly selected reference images and simple distance measure. In this paper, we improve upon existing work with two major contributions. First, we show that a more representative reference-set contributes to better classification accuracy. To this end, we carefully adapt the K-means clustering algorithm in the feature space to select a distinguished reference-set. Second, in the image classification process, we propose to represent each image by measuring its betweenness centrality in a social network composed of the representative reference-set in each class, leading to a more coherent distance measure that considers the overall connectivity between the probe image and the reference-set. Extensive experiment results demonstrate that our proposed scheme achieves better performance than existing methods.
Qun Li 0002, Zhen Qin 0001, Lunshao Chai, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu
ICIP5
2013 Sketching by perceptual grouping
abstract
Sketch is used for rendering the visual world since prehistoric times, and has become ubiquitous nowadays with the increasing availability of touchscreens on portable devices. However, how to automatically map images to sketches, a problem that has profound implications on applications such as sketch-based image retrieval, still remains open. In this paper, we propose a novel method that draws a sketch automatically from a single natural image. Sketch extraction is posed within an unified contour grouping framework, where perceptual grouping is first used to form contour segment groups, followed by a group-based contour simplification method that generate the final sketches. In our experiment, for the first time we pose sketch evaluation as a sketch-based object recognition problem and the results validate the effectiveness of our system over the state-of-the-arts alternatives.
Yonggang Qi, Jun Guo 0002, Yi Li 0004, Honggang Zhang 0002, Tao Xiang 0002, Yi-Zhe Song
ICIP2
2013 Cross-Domain Recommendation via Cluster-Level Latent Factor Model
Sheng Gao 0001, Shantao Li, Patrick Gallinari, Jun Guo 0002
ECML/PKDD (2)6
2013 Perceptual grouping via untangling Gestalt principles
abstract
Gestalt principles, a set of conjoining rules derived from human visual studies, have been known to play an important role in computer vision. Many applications such as image segmentation, contour grouping and scene understanding often rely on such rules to work. However, the problem of Gestalt confliction, i.e., the relative importance of each rule compared with another, remains unsolved. In this paper, we investigate the problem of perceptual grouping by quantifying the confliction among three commonly used rules: similarity, continuity and proximity. More specifically, we propose to quantify the importance of Gestalt rules by solving a learning to rank problem, and formulate a multi-label graph-cuts algorithm to group image primitives while taking into account the learned Gestalt confliction. Our experiment results confirm the existence of Gestalt confliction in perceptual grouping and demonstrate an improved performance when such a confliction is accounted for via the proposed grouping algorithm. Finally, a novel cross domain image classification method is proposed by exploiting perceptual grouping as representation.
Yonggang Qi, Jun Guo 0002, Yi Li 0004, Honggang Zhang 0002, Tao Xiang 0002, Yi-Zhe Song, Zheng-Hua Tan
VCIP2
2013 A multi-label classification approach for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) techniques have already been adopted in numerous multimedia systems. Plenty of previous research assumes that each facial picture should be linked to only one of the predefined affective labels. Nevertheless, in practical applications, few of the expressions are exactly one of the predefined affective states. Therefore, to depict the facial expressions more accurately, this paper proposes a multi-label classification approach for FER and each facial expression would be labeled with one or multiple affective states. Meanwhile, by modeling the relationship between labels via Group Lasso regularization term, a maximum margin multi-label classifier is presented and the convex optimization formulation guarantees a global optimal solution. To evaluate the performance of our classifier, the JAFFE dataset is extended into a multi-label facial expression dataset by setting threshold to its continuous labels marked in the original dataset and the labeling results have shown that multiple labels can output a far more accurate description of facial expression. At the same time, the classification results have verified the superior performance of our algorithm.
Kaili Zhao, Honggang Zhang 0002, Mingzhi Dong, Jun Guo 0002, Yonggang Qi, Yi-Zhe Song
VCIP4
2013 Bases sorting: Generalizing the concept of frequency for over-complete dictionaries
Chun-Guang Li, Zhouchen Lin, Jun Guo 0002
Neurocomputing3
2013 Text extraction from natural scene image: A survey
Honggang Zhang 0002, Kaili Zhao, Yi-Zhe Song, Jun Guo 0002
Neurocomputing4
2013 A query by humming system based on locality sensitive hashing indexes
Qiang Wang 0048, Gang Liu 0008, Jun Guo 0002
Signal Process.4
2013 Reference-Based Scheme Combined With K-SVD for Scene Image Categorization
abstract
A reference-based algorithm for scene image categorization is presented in this letter. In addition to using a reference-set for images representation, we also associate the reference-set with training data in sparse codes during the dictionary learning process. The reference-set is combined with the reconstruction error to form a unified objective function. The optimal solution is efficiently obtained using the K-SVD algorithm. After dictionaries are constructed, Locality-constrained Linear Coding (LLC) features of images are extracted. Then, we represent each image feature vector using the similarities between the image and the reference-set, leading to a significant reduction of the dimensionality in the feature space. Experimental results demonstrate that our method achieves outstanding performance.
Qun Li 0002, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu
IEEE Signal Process. Lett.3
2013 Web Multimedia Object Classification Using Cross-Domain Correlation Knowledge
abstract
Given a collection of web images with the corresponding textual descriptions, in this paper, we propose a novel cross-domain learning method to classify these web multimedia objects by transferring the correlation knowledge among different information sources. Here, the knowledge is extracted from unlabeled objects through unsupervised learning and applied to perform supervised classification tasks. To mine more meaningful correlation knowledge, instead of using commonly used visual words in the traditional bag-of-visual-words (BoW) model, we discover higher level visual components (words and phrases) to incorporate the spatial and semantic information into our image representation model, i.e., bag-of-visual-phrases (BoP). By combining the enriched visual components with the textual words, we calculate the frequently co-occurring pairs among them to construct a cross-domain correlated graph in which the correlation knowledge is mined. After that, we investigate two different strategies to apply such knowledge to enrich the feature space where the supervised classification is performed. By transferring such knowledge, our cross-domain transfer learning method can not only handle large scale web multimedia objects, but also deal with the situation that the textual descriptions of a small portion of web images are missing. Empirical experiments on two different datasets of web multimedia objects are conducted to demonstrate the efficacy and effectiveness of our proposed cross-domain transfer learning method.
Wenting Lu, Tao Li 0001, Weidong Guo, Honggang Zhang 0002, Jun Guo 0002
IEEE Trans. Multim.6
2012 Pairwise Rotation Invariant Co-occurrence Local Binary Pattern
Xianbiao Qi, Rong Xiao 0003, Jun Guo 0002, Lei Zhang 0001
ECCV (6)3
2012 Entropy based locality sensitive hashing
abstract
Nearest neighbor problem has recently been a research focus, especially on large amounts of data. Locality sensitive hashing (LSH) scheme based on p-stable distributions is a good solution to the approximate nearest neighbor (ANN) problem, but points are always mapped to a poor distribution. This paper proposes a set of new hash mapping functions based on entropy for LSH. Using our new hash functions the distribution of mapped values will be approximately uniform, which is the maximum entropy distribution. This paper also provides a method on how these parameters should be adjusted to get better performance. Experimental results show that the proposed method will be more accurate with the same time consuming.
Qiang Wang 0048, Gang Liu 0008, Jun Guo 0002
ICASSP4
2012 Re-ranking using compression-based distance measure for Content-based Commercial Product Image Retrieval
abstract
With the prevalence of E-Commerce sites such as eBay, Content-based Commercial Product Image Retrieval (CBCPIR) has become an emerging application-oriented field of Content-based Image Retrieval (CBIR). Though a number of traditional CBIR techniques and evaluation criterions have been applied directly or with minor modifications, they tend to neglect one critical factor that greatly affects user experience: users usually care about the exact ranks of the results, especially few top ones, which should share very high similarity with the query image. In this work, we propose a novel two-stage retrieval framework that uses a compression-based re-ranking method and a new subjective retrieval evaluation criterion to address such a problem. More specifically, we extend the state-of-art texture descriptor Campana-Keogh (CK) method from data mining in several aspects and validate the superiority of our framework via extensive experiments and real-world user feedback. We also make our code and CBCPIR dataset publicly available. The number of images of the latter is much larger than current freely accessible ones and better represents real-world commercial product images.
Lunshao Chai, Zhen Qin 0001, Honggang Zhang 0002, Jun Guo 0002, Christian R. Shelton
ICIP4
2012 Integrative labeling based statistical color models with application to skin detection
abstract
To alleviate the workload of labeling before estimating certain color distributions, integrative labeling is introduced, which merely needs to figure out whether a picture contains positive-class regions or not and then all pixels of the picture are treated as positive or negative class training samples. Integrative labeling, however, results in heavy mixture of training samples. Thus traditional generative density estimation methods can't be used directly in that they perform poorly with heavily polluted training samples. In this paper, by utilizing the prior knowledge of high separability between positive and negative class color distributions, a discriminative learning based GMM(DiscGMM) is proposed for integrative labeling. Besides generating the polluted positive-class samples with comparatively high probability, optimal parameters found by DiscGMM also enjoy a comparatively low probability of generating negative-class samples. The parameter learning problem is solved by a modified Expectation Maximization (EM) algorithm. In an integrative labeling experiment of skin detection, DiscGMM is testified to enjoy much better performance than generative density estimation methods and shows qualified results.
Mingzhi Dong, Jun Guo 0002, Weihong Deng, Weiran Xu
ICIP3
2012 Codebook optimization using word activation forces for scene categorization
abstract
Visual codebook based quantization of robust appearance descriptors extracted from local image patches is an effective means of capturing image statistics for texture analysis and natural scene classification. In this paper, based on the newly proposed statistics of word activation forces (WAFs), we optimize the codebook. Currently, codebooks are typically created from a set of training images using a clustering algorithm. However, these codebooks are often functionally limited due to redundancy. We show that WAFs can remove the redundancy efficiently. In the experiment, the proposed method achieved the state-of-the-art performance on the Caltech-101, fifteen natural scene categories and VOC2007 databases. The optimization method also offers insights into the success of several recently proposed images classification approaches, including vector quantization (VQ) coding in the Spatial Pyramid Matching (SPM), sparse coding SPM (ScSPM), and Locality-constrained Linear Coding (LLC).
Qun Li 0002, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu
ICIP3
2012 A Linear Max K-min classifier
Mingzhi Dong, Weihong Deng, Qiang Wang 0048, Caixia Yuan, Jun Guo 0002, Liwei Ma
ICPR6
2012 Query by humming via hierarchical filters
Qiang Wang 0048, Gang Liu 0008, Jun Guo 0002
ICPR5
2012 Robust lyric search based on weighted syllable confusion matrix
Baoxiang Li, Fengxiang Chang, Qiang Wang 0048, Gang Liu 0008, Jun Guo 0002
ICPR5
2012 Tempo variation based multilayer filters for query by humming
Qiang Wang 0048, Baoxiang Li, Gang Liu 0008, Jun Guo 0002
ICPR5
2012 A rapid flower/leaf recognition system
abstract
In this work, we introduce a rapid and accurate flower/leaf recognition system. The system could process one query in less than 0.35s with users' simple interaction. Meanwhile, high accuracy and recall is achieved. Furthermore, low computational resource and memory cost are required by the system. Now, the system is demonstrated on 172 categories of flowers, the largest flower dataset until now, and 220 categories of leaves.
Xianbiao Qi, Rong Xiao 0003, Lei Zhang 0001, Chun-Guang Li, Jun Guo 0002
ACM Multimedia5
2012 Extended SRC: Undersampled Face Recognition via Intraclass Variant Dictionary
abstract
Sparse Representation-Based Classification (SRC) is a face recognition breakthrough in recent years which has successfully addressed the recognition problem with sufficient training images of each gallery subject. In this paper, we extend SRC to applications where there are very few, or even a single, training images per subject. Assuming that the intraclass variations of one subject can be approximated by a sparse linear combination of those of other subjects, Extended Sparse Representation-Based Classifier (ESRC) applies an auxiliary intraclass variant dictionary to represent the possible variation between the training and testing images. The dictionary atoms typically represent intraclass sample differences computed from either the gallery faces themselves or the generic faces that are outside the gallery. Experimental results on the AR and FERET databases show that ESRC has better generalization ability than SRC for undersampled face recognition under variable expressions, illuminations, disguises, and ages. The superior results of ESRC suggest that if the dictionary is properly constructed, SRC algorithms can generalize well to the large-scale face recognition problem, even with a single training image per class.
Weihong Deng, Jiani Hu, Jun Guo 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 The small sample size problem of ICA: A comparative study and analysis
Weihong Deng, Yebin Liu, Jiani Hu, Jun Guo 0002
Pattern Recognit.4
2011 Web Multimedia Object Clustering via Information Fusion
abstract
Multimedia information plays an increasingly important role in humans daily activities. Given a set of web multimedia objects (images with corresponding texts), a challenging problem is how to group these images into several clusters using the available information. Previous researches focus on either adopting individual information, or simply combining image and text information together for clustering. In this paper, we propose a novel approach (Dynamic Weighted Clustering) to separate images under the "supervision" of text descriptions, Also, we provide a comparative experimental investigation on utilizing text and image information to tackle web image clustering. Empirical experiments on a manually collected web multimedia object (related to the events after disasters) dataset are conducted to demonstrate the efficacy of our proposed method.
Wenting Lu, Lei Li 0001, Tao Li 0001, Honggang Zhang 0002, Jun Guo 0002
ICDAR5
2011 Weakly supervised locality sensitive hashing for duplicate image retrieval
abstract
Locality sensitive hashing (LSH) is quite popular in high dimensional data indexing. However, most of existing methods perform hashing in an unsupervised way, that is to say, hash functions are randomly generated without the prior information of the data. In this paper, we propose two improved LSH algorithms based on weakly supervised learning technique, which need only small quantities of labeled sample pairs. One is to select the most appropriate hash functions from a pool of functions using sample pairs labeled with “similar” or “dissimilar”. The other is to generate hash functions with positive sample pairs. The experiments show that the proposed algorithms reduce the search complexity compared with original LSH.
Honggang Zhang 0002, Jun Guo 0002
ICIP3
2011 Structural fingerprint based hierarchical filtering in song identification
abstract
Automatic song identification has long been a research focus. In this paper, a novel structural fingerprint based hierarchical filtering method is proposed and it consists of two parts: one is the generation of fingerprint with both long structural information and low collision, and the other is an efficient searching algorithm based on a set of selective 2-level filters. Experiments conducted on a database of 10,000 songs show that our approach is fast enough and can achieve the accuracy of 99.7% on 5 second clips with the SNR at 0db comparable to the state-of-the-art.
Qiang Wang 0048, Gang Liu 0008, Jun Guo 0002
ICME4
2011 Product comparison using comparative relations
abstract
This paper proposes a novel Product Comparison approach. The comparative relations between products are first mined from both user reviews on multiple review websites and community-based question answering pairs containing product comparison information. A unified graph model is then developed to integrate the resultant comparative relations for product comparison. Experiments on popular electronic products show that the proposed approach outperforms the state-of-the-art methods.
Si Li 0001, Zhengjun Zha, Zhaoyan Ming, Meng Wang 0001, Tat-Seng Chua, Jun Guo 0002, Weiran Xu
SIGIR6
2010 Matching Image with Multiple Local Features
abstract
In this paper, we present the fusional feature composed of Affine-SIFT, MSER and color moment invariants. The fusional feature is more robust and distinctive than a single local feature. Instead of adding three local features together simply, an efficient two-level matching strategy is devised with the fusional feature, which speeds up the establishment of the local correspondences. To remove partial false positives, an affine transformation is estimated with the weighted RANSAC which decreases iteration times. The experimental results show that our approach can achieve more accurate correspondence. We prospect to apply the fusional feature and match strategy to image retrieval in the end.
Honggang Zhang 0002, Jun Guo 0002
ICPR5
2010 Local Sparse Representation Based Classification
abstract
In this paper, we address the computational complexity issue in Sparse Representation based Classification (SRC). In SRC, it is time consuming to find a global sparse representation. To remedy this deficiency, we propose a Local Sparse Representation based Classification (LSRC) scheme, which performs sparse decomposition in local neighborhood. In LSRC, instead of solving the l1-norm constrained least square problem for all of training samples we solve a similar problem in a local neighborhood for each test sample. Experiments on face recognition data sets ORL and Extended Yale B demonstrated that the proposed LSRC algorithm can reduce the computational complexity and remain the comparative classification accuracy and robustness.
Chun-Guang Li, Jun Guo 0002, Honggang Zhang 0002
ICPR2
2010 Exploiting Combined Multi-level Model for Document Sentiment Analysis
abstract
This paper focuses on the task of text sentiment analysis in hybrid online articles and web pages. Traditional approaches of text sentiment analysis typically work at a particular level, such as phrase, sentence or document level, which might not be suitable for the documents with too few or too many words. Considering every level analysis has its own advantages, we expect that a combination model may achieve better performance. In this paper, a novel combined model based on phrase and sentence level's analyses and a discussion on the complementation of different levels' analyses are presented. For the phrase-level sentiment analysis, a newly defined Left-Middle-Right template and the Conditional Random Fields are used to extract the sentiment words. The Maximum Entropy model is used in the sentence-level sentiment analysis. The experiment results verify that the combination model with specific combination of features is better than single level model.
Si Li 0001, Hao Zhang 0022, Weiran Xu, Guang Chen 0003, Jun Guo 0002
ICPR5
2010 Locality preserving and global discriminant projection with prior information
Honggang Zhang 0002, Weihong Deng, Jun Guo 0002, Jie Yang 0001
Mach. Vis. Appl.3
2010 Robust, accurate and efficient face recognition from a single training image: A uniform pursuit approach
Weihong Deng, Jiani Hu, Jun Guo 0002, Tom Weidong Cai, David Dagan Feng
Pattern Recognit.3
2010 Emulating biological strategies for uncontrolled face recognition
Weihong Deng, Jiani Hu, Jun Guo 0002, Tom Weidong Cai, David Dagan Feng
Pattern Recognit.3
2009 Learning Bundle Manifold by Double Neighborhood Graphs
Chun-Guang Li, Jun Guo 0002, Honggang Zhang 0002
ACCV (3)2
2009 Detection of Vehicle Manufacture Logos Using Contextual Information
Wenting Lu, Honggang Zhang 0002, Kunyan Lan, Jun Guo 0002
ACCV (2)4
2009 HCL2000 - A Large-scale Handwritten Chinese Character Database for Handwritten Character Recognition
abstract
In this paper, we present a large scale offline handwritten Chinese character database-HCL2000 which will be made public available for the research community. The database contains 3,755 frequently used simplified Chinese-characters written by 1,000 different subjects. The writerspsila information is incorporated in the database to facilitate testing on grouping writers with different background such as age, occupation, gender, and education etc. We investigate some characteristics of writing styles from different groups of writers. We evaluate HCL2000 database using three different algorithms as a baseline. We decide to publish the database along with this paper and make it free for a research purpose.
Honggang Zhang 0002, Jun Guo 0002, Guang Chen 0003, Chun-Guang Li
ICDAR2
2009 Semi-supervised Learning Based on Label Propagation through Submanifold
Jiani Hu, Weihong Deng, Jun Guo 0002
ISNN (1)3
2009 Emotion Recognition of Pop Music Based on Maximum Entropy with Priors
Jun Guo 0002
PAKDD3
2009 Intrinsic Dimensionality Estimation within Neighborhood Convex Hull
abstract
In this paper, a novel method to estimate the intrinsic dimensionality of high-dimensional data set is proposed. Based on neighborhood information, our method calculates the non-negative locally linear reconstruction coefficients from its neighbors for each data point, and the numbers of those dominant positive reconstruction coefficients are regarded as a faithful guide to the intrinsic dimensionality of data set. The proposed method requires no parametric assumption on data distribution and is easy to implement in the general framework of manifold learning. Experimental results on several synthesized data sets and real data sets have shown the benefits of the proposed method.
Chun-Guang Li, Jun Guo 0002
Int. J. Pattern Recognit. Artif. Intell.2
2009 Learning a locality discriminating projection for classification
Jiani Hu, Weihong Deng, Jun Guo 0002, Weiran Xu
Knowl. Based Syst.3
2008 Handwritten Chinese character recognition using Local Discriminant Projection with Prior Information
abstract
In this paper, we propose a new method to model the manifold of handwritten Chinese characters using the local discriminant projection. We utilize a cascade framework that combines global similarity with local discriminative cues to recognize Chinese characters. We find the similarity of different characters using a nearest-neighbor (NN) classifier, and followed by the Local Discriminant Projection with Prior Information (LDPPI) to map similar characters within a cluster to a low-dimensional space. We evaluate the proposed method on two large public datasets, ETL9B which contains 607,200 handwritten characters from 200 people, and HCL2000 which contains 3,755,000 characters written by 1,000 people. The experimental results demonstrate that the proposed method achieves 0.74% error rate on ETL9B database and 1.88% on HCL2000 database.
Honggang Zhang 0002, Jie Yang 0001, Weihong Deng, Jun Guo 0002
ICPR4
2008 Comments on "Globally Maximizing, Locally Minimizing: Unsupervised Discriminant Projection with Application to Face and Palm Biometrics"
abstract
In [1], UDP is proposed to address the limitation of LPP for the clustering and classification tasks. In this communication, we show that the basic ideas of UDP and LPP are identical. In particular, UDP is just a simplified version of LPP on the assumption that the local density is uniform.
Weihong Deng, Jiani Hu, Jun Guo 0002, Honggang Zhang 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 Locality discriminating indexing for document classification
abstract
This paper introduces a locality discriminating indexing (LDI) algorithm for document classification. Based on the hypothesis that samples from different classes reside in class-specific manifold structures, LDI seeks for a projection which best preserves the within-class local structures while suppresses the between-class overlap. Comparative experiments show that the proposed method isable to derives compact discriminating document representations for classification.
Jiani Hu, Weihong Deng, Jun Guo 0002, Weiran Xu
SIGIR3
2006 A Boosting Approach for Utterance Verification
Chengyu Dong, Dezhi Huang, Jun Guo 0002, Haila Wang
ICIC (2)4
2006 Multi-scale Support Vector Machine for Regression Estimation
Zhen Yang 0004, Jun Guo 0002, Weiran Xu, Xiangfei Nie, Jianjun Lei 0001
ISNN (1)2
2006 Model-Based Feature Compensation for Robust Speech Recognition
Haifeng Shen, Qunxia Li, Jun Guo 0002, Gang Liu 0008
Fundam. Informaticae3
2005 Two-Domain Feature Compensation for Robust Speech Recognition
Haifeng Shen, Gang Liu 0008, Jun Guo 0002, Qunxia Li
ISNN (2)3
2005 Non-stationary Environment Compensation Using Sequential EM Algorithm for Robust Speech Recognition
Haifeng Shen, Jun Guo 0002, Gang Liu 0008, Qunxia Li
PKDD2