VLDB 2026 Research / reviewers in the wild / expert
Meng Yang 0001
dblp:44/2761-1
· DBLP profile ↗
107ranked-venue papers
26as first author
40since 2021 · last 2026
0000-0002-0795-3221ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 71 · 22 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 12 first-author · 18 since 2021Security and privacy · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AEA-FIRM: Adaptive Elastic Alignment With Fine-Grained Representation Mining for Text-Based Aerial Pedestrian RetrievalabstractUnmanned aerial vehicles (UAVs) have garnered significant attention due to their operational flexibility, enabling expanded application scenarios across diverse fields. The Text-Based Pedestrian Retrieval (TBPR) task aims to identify corresponding images from textual descriptions, yet existing research has primarily focused on ground-level views. To broaden the applicability of TBPR systems, we introduce aerial-view analysis and propose a novel Text-Based Aerial Pedestrian Retrieval (TBAPR) task. This task introduces unique challenges, particularly the dual gaps in cross-view (aerial vs. ground) and cross-modal (text vs. image) matching, which are more complex than traditional TBPR or aerial-ground pedestrian understanding tasks. To address these challenges, we propose an Adaptive Elastic Alignment Network with FIne-Grained Representation Mining (AEA-FIRM). Our framework tackles the cross-view gap through an AEA loss that adaptively prioritizes critical semantic features while dynamically aligning textual and aerial semantics under challenging conditions. Concurrently, the FIRM module refines visual-linguistic representations by mining fine-grained pedestrian attributes and explicitly textualizing them for cross-modal matching verification. Extensive experiments demonstrate that AEA-FIRM achieves state-of-the-art performance, outperforming existing TBPR methods by 4.87% in Rank-1 accuracy. Our code and dataset are available at https://github.com/xbdxwyh/AEA-FIRM-main.git. Yihao Wang 0010, Meng Yang 0001, Rui Cao 0003, Guangwei Gao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | SEC-Prompt: SEmantic Complementary Prompting for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) presents a significant challenge in machine learning, requiring models to integrate new classes from limited examples while preserving performance on previously learned classes. Recently, prompt-based CIL approaches leverage ample data to train prompts, effectively mitigating catastrophic forgetting. However, these methods do not account for the semantic features embedded in prompts, exacerbating the plasticity-stability dilemma in few-shot incremental learning. In this paper, we propose a novel and simple framework named SEmantic Complementary Prompt(SEC-Prompt), which learns two sets of semantically complementary prompts based on an adaptive query: discriminative prompts(D-Prompt) and non-discriminative prompts(ND-Prompt). D-Prompt enhances the separation of class-specific feature distributions by strengthening key discriminative features, while ND-Prompt balances non-discriminative information to promote generalization to novel classes. To efficiently learn high-quality knowledge from limited samples, we leverage ND-Prompt for data augmentation to increase sample diversity and introduce Prompt Clustering Loss to prevent noise contamination in D-Prompt, ensuring robust discriminative feature learning and improved generalization. Our experimental results showcase state-of-the-art performance across three benchmark datasets, including CIFAR100, ImageNet-R and CUB datasets. Meng Yang 0001 |
CVPR | 2 |
| 2025 | Language-based reasoning graph neural network for commonsense question answering
Meng Yang 0001, Yihao Wang 0010 |
Neural Networks | 1 |
| 2025 | MultiSpectral Transformer Fusion via exploiting similarity and complementarity for robust pedestrian detection
Song Hou, Meng Yang 0001, Wei-Shi Zheng 0001, Shibo Gao |
Pattern Recognit. | 2 |
| 2024 | Overcome Noise and Bias: Segmentation-Aided Multi-Granularity Denoising and Debiasing for Enhanced Quarduples Extraction in DialogueabstractDialogue Aspect-based Sentiment Quadruple analysis (DiaASQ) extends ABSA to more complex real-world scenarios (i.e., dialogues), which makes existing generation methods encounter heightened noise and order bias challenges, leading to decreased robustness and accuracy.To address these, we propose the Segmentation-Aided multi-grained Denoising and Debiasing (SADD) method.For noise, we propose the Multi-Granularity Denoising Generation model (MGDG), achieving word-level denoising via sequence labeling and utterancelevel denoising via topic-aware dialogue segmentation.Denoised Attention in MGDG integrates multi-grained denoising information to help generate denoised output.For order bias, we first theoretically analyze its direct cause as the gap between ideal and actual training objectives and propose a distribution-based solution.Since this solution introduces a one-to-many learning challenge, our proposed Segmentationaided Order Bias Mitigation (SOBM) method utilizes dialogue segmentation to supplement order diversity, concurrently mitigating this challenge and order bias.Experiments demonstrate SADD's effectiveness, achieving state-ofthe-art results with a 6.52% F1 improvement. Xianlong Luo, Meng Yang 0001, Yihao Wang 0010 |
EMNLP | 2 |
| 2024 | Implicit-Knowledge-Guided Align Before Understanding for KB-VQAabstractVisual Question Answering based on external Knowledge Bases(KB-VQA) have gained more attention in recent years. Due to the large scale of knowledge base data and the lack of manual annotation information in the process of retrieving knowledge for image problem data, the current model usually suffers performance due to the low accuracy of knowledge retrieval. In order to solve the current challenges, we propose an implicit-knowledge-guided alignment before understanding (IK-ALBUN) model, which improves the knowledge base retrieval performance by aligning multimodal information with the knowledge base in the first stage, and conduct multi-modal semantic fusion and understanding to complete the KB-VQA with the retrieved knowledge bases and the additionally designed contrastive loss. The model has completed experiments on several public datasets and proved the superior performance of the current method. Mao Feng, Meng Yang 0001 |
ICASSP | 3 |
| 2024 | Incomplete Observations Bias Suppression for Abductive Natural Language InferenceabstractAbductive natural language commonsense reasoning is a task aiming at inferring the most plausible explanation in narrative text for observed events. Previous works mostly concentrate on utilizing powerful pre-trained language models and making better use of excess training data to learn abundant event commonsense knowledge. However, the utilization of causal effect is hidden in the language reasoning process and the explicit constraint of the causal effect between events has not been explored, resulting in biased inference. The model may focus on one observed event and make the wrong prediction while ignoring the other helpful events. To reveal the problem we modify the original task by appending unrelated text to the context which won’t change the causal relation. And typical methods get worse in the new task as they are not good at utilizing the complementary between the two observations. Motivated by eliminating the shortcut from incomplete observation and utilizing the complementarity of the two observations, we propose an incomplete observation bias suppression method to guide the training process. Results show our approach can ease the problem revealed in the new task. Based on the proposed method and the new task, our method also get competitive result on the original task. Xianlong Luo, Meng Yang 0001 |
ICASSP | 3 |
| 2024 | Human Guided Cross-Modal Reasoning with Semantic Attention Learning for Visual Question AnsweringabstractOne of the major difficulties in the Visual Question Answering (VQA) task of real-world images is the long-tailed distribution of concepts which makes the model vulnerable to negative linguistic biases. To imitate human learning and reasoning, researchers have designed reasoning models, which, however, is still a black-box process and cannot guarantee the visual interpretability of the final answer. How to guide the direction of the model reasoning and improve the generalization ability is a challenge to be solved. We proposed a novel Human-Guided Cross-Modal Reasoning (HGCMR) with semantic attention learning to improve the reasoning ability. The cross-modal reasoning module of HGCMR imitates the reasoning steps via semantic attention learning to generate the contextural image and question representation. The supervision module of HGCMR automatically extracts the human-guided attention distribution over object regions from the provided reasoning patterns, so as to guide the reasoning process. With the attended image and question representation and human reasoning supervision, the proposed HGCMR finally complete the question-answering task with an output classifier. By evaluating models on the real-world dataset GQA, our HGCMR improves compositional and grounding performance. Mao Feng, Meng Yang 0001 |
ICASSP | 3 |
| 2024 | MCM-CSD: Multi-Granularity Context Modeling with Contrastive Speaker Detection for Emotion Recognition in Real-Time ConversationabstractEmotion recognition in conversation (ERC) has received extensive attention for its wide applications in recent years. Considering the actual situation, we focus on the real-time conversation scenarios, in which how to model the conversation emotion with only the historical contextual information and how to exploit the speaker information for emotion recognition have not been well studied. Therefore, we propose a novel multi-task learning model MCM-CSD, which combines the multi-granularity context modeling (MCM) based ERC with contrastive speaker detection (CSD). For the main task ERC, we design a bottom-up approach to fully extract the multi-granularity contextual information in both word and utterance levels. And for the auxiliary task CSD, we design a supervised contrastive learning loss that can easily distinguish previous speakers from the current speaker in multi-turn conversations. We conduct experiments on four benchmark datasets, the results show that our model can achieve state-of-the-art performance compared to previous methods. Furthermore, we perform ablation experiments and a case study to verify the effectiveness of each component and explain the significance of CSD in MCM-CSD. The code is available at https://github.com/WHOISJENNY/MCM-CSD. Yuan Xu 0028, Meng Yang 0001 |
ICASSP | 2 |
| 2024 | Fine-grained Semantic Alignment with Transferred Person-SAM for Text-based Person RetrievalabstractAddressing the disparity in description granularity and information gap between images and text has long been a formidable challenge in text-based person retrieval (TBPR) tasks. Recent researchers tried to solve this problem by random local alignment. However, they failed to capture the fine-grained relationships between images and text, so the information and modality gaps remain on the table. We align image regions and text phrases at the same semantic granularity to address the semantic atomicity gap. Our idea is first to extract and then exploit the relationships between fine-grained locals. We introduce a novel Fine-grained Semantic Alignment with Transferred Person-SAM (SAP-SAM) approach. By distilling and transferring knowledge, we propose a Person-SAM model to extract fine-grained semantic concepts at the same granularity from images and texts of TBPR and its relationships. With the extracted knowledge, we optimize the fine-grained matching via Explicit Local Concept Alignment and Attentive Cross-modal Decoding to discriminate fine-grained image and text features at the same granularity level and represent the important semantic concepts from both modalities, effectively alleviating the granularity and information gaps. We evaluate our proposed approach on three popular TBPR datasets, demonstrating that SAP-SAM achieves state-of-the-art results and underscores the effectiveness of end-to-end fine-grained local alignment in TBPR tasks. Yihao Wang 0010, Meng Yang 0001, Rui Cao 0003 |
ACM Multimedia | 2 |
| 2024 | A pathology-based diagnosis and prognosis intelligent system for oral squamous cell carcinoma using semi-supervised learningabstractPathological images are important for diagnosis and prognosis of oral squamous cell carcinoma (OSCC). However, it is difficult for pathologists to directly apply intuitive pathological image information to predict prognosis. Applying supervised learning (SL) to whole slide images (WSIs) analysis is labor-consuming and time-costing, and semi-supervised learning (SmSL) has provided a new opportunity to revisit classical approaches in digital pathology. In this study, we designed an intelligent SmSL system based on Self-supervised Pretraining (SP) and Adaptive Threshold (AT), named SPAT_SmSL, for the diagnosis and prognosis of OSCC on multi centers. Firstly, we used the SP technique and AT strategy to fully exploit the unlabeled data, both of which were integrated into the SPAT_SmSL algorithm to recognize tumor, stroma, and tumor-infiltrating lymphocytes (TILs) regions. Secondly, pathological variables including TIL-score and depth of invasion (DOI) were digitally quantified based on the results of image recognition. Finally, multivariable cox analysis was performed to identify independent prognostic factors affecting overall survival and establish a comprehensive predictive model for OSCC patients. The new SPAT_SmSL paradigm demonstrates superior performance in WSIs recognition and survival prediction, which potentially serves as a novel tool to build an expert digital pathological platform to meet the demand of intelligent diagnosis and prognosis, as well as facilitating clinicians with complementary information for individualized treatment in the future. Jiaying Zhou, Haoyuan Wu, Xiaojing Hong, Yunyi Huang, Bo Jia, Jiabin Lu, Meng Yang 0001 |
Expert Syst. Appl. | 9 |
| 2023 | Tagging-Assisted Generation Model with Encoder and Decoder Supervision for Aspect Sentiment Triplet ExtractionabstractASTE (Aspect Sentiment Triplet Extraction) has gained increasing attention.Recent advancements in the ASTE task have been primarily driven by Natural Language Generationbased (NLG) approaches.However, most NLG methods overlook the supervision of the encoder-decoder hidden representations and fail to fully utilize the semantic information provided by the labels to enhance supervision.These limitations can hinder the extraction of implicit aspects and opinions.To address these challenges, we propose a tagging-assisted generation model with encoder and decoder supervision (TAGS), which enhances the supervision of the encoder and decoder through multipleperspective tagging assistance and label semantic representations.Specifically, TAGS enhances the generation task by integrating an additional sequence tagging task, which improves the encoder's capability to distinguish the words of triplets.Moreover, it utilizes sequence tagging probabilities to guide the decoder, improving the generated content's quality.Furthermore, TAGS employs a selfdecoding process for labels to acquire the semantic representations of the labels and aligns the decoder's hidden states with these semantic representations, thereby achieving enhanced semantic supervision for the decoder's hidden states.Extensive experiments on various public benchmarks demonstrate that TAGS achieves state-of-the-art performance. Xianlong Luo, Meng Yang 0001, Yihao Wang 0010 |
EMNLP | 2 |
| 2023 | Self-Attention Prediction Correction with Channel Suppression for Weakly-Supervised Semantic SegmentationabstractSingle-stage weakly-supervised semantic segmentation (WSSS) with image-level labels has become a new research hotspot in the community for its lower cost and higher training efficiency. However, the pseudo label of WSSS generally suffers from somewhat noise, which limits the segmentation performance. In this paper, to explore the integral foreground activation, we propose the Channel Suppression (CS) module for preventing only activating the most discriminative regions, thereby improving the initial pseudo labels. To rectify the in-correct prediction, we explore the Self-Attention Prediction Correction (SAPC) module, which adaptively generates the category-wise prediction rectification weights. After extensive experiments, the proposed efficient single-stage framework achieves excellent performance with 67.6% mIoU and 39.9% mIoU on PASCAL VOC 2012 and MS COCO 2014 datasets, significantly exceeding several recent single-stage methods. Guoying Sun, Meng Yang 0001 |
ICME | 2 |
| 2023 | Multi-hop Attention GNN with Answer-Evidence Contrastive Loss for Multi-hop QAabstractMulti-hop question answering (QA) is a challenging task in natural language processing (NLP), which requires multi-step reasoning over the sentences from several passages and finding out the answer as well as the scattered evidence sentences. The existing QA models that are based on Graph Neural Network (GNN) have exhibited good performance, however, the advantages of GNN have not been brought into full play. In this paper, we incorporate an effective multi-hop attention mechanism into GNN to aggregate richer information from high-order nodes of the graph. In addition, when multiple tasks are jointly optimized, the performance of all tasks is usually unable to improve together. To address this problem, we design a novel answer-evidence contrastive learning loss, which encourages models to learn better shared representation and distinguish the evidence sentences from other confusing ones through answer-evidence similarity. Our experiments on HotpotQA dataset demonstrate that the proposed method achieves comparable results to the state-of-the-art models and helps the baseline model gain significant performance improvement. Meng Yang 0001 |
IJCNN | 2 |
| 2023 | Confidence-Guided Open-World Semi-supervised Learning
Jibang Li, Meng Yang 0001, Mao Feng |
PRCV (4) | 2 |
| 2023 | DictPrompt: Comprehensive dictionary-integrated prompt tuning for pre-trained language model
Rui Cao 0003, Yihao Wang 0010, Meng Yang 0001 |
Knowl. Based Syst. | 4 |
| 2023 | Discriminative semi-supervised learning via deep and dictionary representation for image classification
Meng Yang 0001, Jiaming Chen 0008, Mao Feng, Jian Yang 0003 |
Pattern Recognit. | 1 |
| 2023 | Deep Supervised Dual Cycle Adversarial Network for Cross-Modal RetrievalabstractCross-modal retrieval tasks, which are more natural and challenging than traditional retrieval tasks, have attracted increasing interest from researchers in recent years. Although different modalities with the same semantics have some potential relevance, the feature space heterogeneity still seriously weakens the performance of cross-modal retrieval models. To solve this problem, common space-based methods in which multimodal data is projected into a learned common space for similarity measurement have become the mainstream approach for cross-modal retrieval tasks. However, current methods entangle the modality style and semantic content in the common space and neglect to fully explore the semantic and discriminative representation/reconstruction of the semantic content. This often results in an unsatisfactory retrieval performance. To solve these issues, this paper proposes a new Deep Supervised Dual Cycle Adversarial Network (DSDCAN) model based on common space learning. It is composed of two cross-modal cycle GANs, one for the image and one for the text. The proposed cycle GAN model disentangles the semantic content and modality style features by making the data of one modality well reconstructed from the extracted modal style feature and the content feature of the other modality. Then, a discriminative semantic and label loss is proposed by fully considering the category, sample contrast, and label supervision to enhance the semantic discrimination of the common space representation. Besides this, to make the data distribution between two modalities similar, a second-order similarity is presented as a distance measurement of the cross-modal representation in the common space. Extensive experiments have been conducted on the Wikipedia, Pascal Sentence, NUS-WIDE-10k, PKU XMedia, MSCOCO, NUS-WIDE, Flickr30k and MIRFlickr datasets. The results demonstrate that the proposed method can achieve a higher performance than the state-of-the-art methods. Meng Yang 0001, Bob Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Pseudo-Label Noise Prevention, Suppression and Softening for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (ReID), including fully unsupervised ReID and unsupervised domain adaptive ReID, remains a challenge for the fields of biometrics and computer vision due to its difficulty in learning with unlabeled target domain data. Existing state-of-the-art methods, most of which generate pseudo-labels via unsupervised clustering for model optimization, are inevitably hampered by the under-explored problem of pseudo-label noise. Motivated by this, we propose a novel joint framework termed pseudo-label Noise Prevention, Suppression, and Softening (NPSS) for unsupervised person re-identification. Instead of refining generated label noise after clustering as many existing methods do, we start solving this issue from the source of pseudo-label noise by proposing a new Dynamic Camera-Adaptive Clustering (DCAC), which dynamically involves camera information to prevent noise caused by cross-camera variance, thus improving their quality during clustering. Moreover, we propose an Online Domain Union (ODU) mechanism for the classification model learning on the target domain via involving source domain data with their ground-truth labels, which effectively suppresses the indelible noisy pseudo-labels. Furthermore, we present the Self-Consistency Constraint (SCC) to soften the label noise in a single model with reduced computation and network parameter cost, which achieves intra-sample knowledge ensembling with our global-local SCC and cross-sample knowledge ensembling with our inter-instance SCC. Experiments demonstrate the effectiveness of our method as it surpasses state-of-the-art methods by a large margin on Market-1501, DukeMTMC-ReID, and MSMT17 benchmarks. The code is available at https://github.com/hjwang-824/NPSS. Haijian Wang, Meng Yang 0001, Wei-Shi Zheng 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Semi-Supervised Few-shot Learning via Multi-Factor ClusteringabstractThe scarcity of labeled data and the problem of model overfitting have been the challenges in few-shot learning. Recently, semi-supervised few-shot learning has been developed to obtain pseudo-labels of unlabeled samples for expanding the support set. However, the relationship between unlabeled and labeled data is not well exploited in generating pseudo labels, the noise of which will di-rectly harm the model learning. In this paper, we propose a Clustering-based semi-supervised Few-Shot Learning (cluster-FSL) method to solve the above problems in image classification. By using multi-factor collaborative representation, a novel Multi-Factor Clustering (MFC) is designed to fuse the information of few-shot data distribution, which can generate soft and hard pseudo-labels for unlabeled samples based on labeled data. And we exploit the pseudo labels of unlabeled samples by MFC to expand the support set for obtaining more distribution information. Furthermore, robust data augmentation is used for support set in the fine-tuning phase to increase the labeled samples' diversity. We verified the validity of the cluster-FSL by comparing it with other few-shot learning methods on three popular benchmark datasets, miniImageNet, tieredImageNet, and CUB-200-2011. The ablation experiments further demonstrate that our MFC can effectively fuse distribution information of labeled samples and provide high-quality pseudo-labels. Our code is available at: https://gitlab.com/smartllvlab/cluster-fsl Meng Yang 0001, Jia Shuai |
CVPR | 3 |
| 2022 | Exploring Pixel Alignment on Shallow Feature for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to cover the entire target object only under the image-level supervision. Most WSOL methods are stuck in mining the CAMs (class activation maps) of deep semantic features for they only focus on limited discriminative regions playing key role in classification. Recently, a new paradigm has emerged by localizing objects using the low-level feature through two stages. Existing two-stages methods usually train a classification network first to yield CAMs as pseudo labels to guide the learning of segment network, yet it does not consider the activations with more background noise or less discriminative area. In this paper, we propose a Pixel Alignment strategy to refine the object localization by improving the shallow-feature based CAMs generator with the joint supervision of pseudo-label mask, classification evaluation, and absolution size constraint on the activation map. More specifically, we utilize the class-specific pixel gradient to achieve a robust activation pseudo mask to background noise, which further supervises the activation generator with confident foreground and background regions. We also adapt a post-processing to excavate the target region in the conflict area (i.e., the non-overlap area of CAMs and the activations). Extensive experiments on CUB-2002011 and ILSVRC datasets indicate that our method outperforms the state-of-the-art among the two-stage works. Xinzi Cao, Meng Yang 0001, Guoying Sun |
IJCNN | 2 |
| 2022 | Consistency Learning based on Class-Aware Style Variation for Domain Generalizable Semantic SegmentationabstractDomain generalizable (DG) semantic segmentation, i.e., a semantic segmentation model pretrained from a source domain performs well in previously unseen target domains without any fine-tuning, remains an open question. A promising solution is learning style-agnostic and domain-invariant features with stylized augmented data. However, existing methods mainly focused on performing stylization on coarse-grained image-level features, while ignoring to explore fine-grained semantic style clues and high-order semantic context correlation, which are essential in enhancing the generalization. Motivated by this, we propose a novel framework termed Consistent Learning based on Class-Aware Style Variation (CL-CASV) for DG semantic segmentation. Specifically, with the guidance of class-level semantic information, our proposed Class-Aware Style Variation (CASV) module simulates imaging object and imaging condition style variation that can appear in complex real-world scenarios, thus generating fine-grained class-aware stylized images with rich style variation. Then the similarities between augmentations and original images are exploited via our Self-Correlation Consistency Learning (SCCL) that mines global context consistency from the views of channel correlation and spatial correlation in the feature and prediction spaces. Extensive experiments on mainstream benchmarks, including Cityscapes, GTAV, BDD100K, SYNTHIA, and Mapillary, demonstrate the effectiveness of our method as it surpasses the state-of-the-art methods. Siwei Su, Haijian Wang, Meng Yang 0001 |
ACM Multimedia | 3 |
| 2022 | Infrared and Near-Infrared Image Generation via Content Consistency and Style Adversarial Learning
Meng Yang 0001, Haijian Wang |
PRCV (1) | 2 |
| 2022 | Flexible entity marks and a fine-grained style control for knowledge based natural answer generation
Yongjie Huang, Meng Yang 0001 |
Knowl. Based Syst. | 2 |
| 2022 | Multi-feature sparse similar representation for person identification
Meng Yang 0001, Kangyin Ke, Guangwei Gao |
Pattern Recognit. | 1 |
| 2022 | Hierarchical Deep CNN Feature Set-Based Representation Learning for Robust Cross-Resolution Face RecognitionabstractCross-resolution face recognition (CRFR), which is important in intelligent surveillance and biometric forensics, refers to the problem of matching a low-resolution (LR) probe face image against high-resolution (HR) gallery face images. Existing shallow learning-based and deep learning-based methods focus on mapping the HR-LR face pairs into a joint feature space where the resolution discrepancy is mitigated. However, little works consider how to extract and utilize the intermediate discriminative features from the noisy LR query faces to further mitigate the resolution discrepancy due to the resolution limitations. In this study, we desire to fully exploit the multi-level deep convolutional neural network (CNN) feature set for robust CRFR. In particular, our contributions are threefold. (i) To learn more robust and discriminative features, we desire to adaptively fuse the contextual features from different layers. (ii) To fully exploit these contextual features, we design a feature set-based representation learning (FSRL) scheme to collaboratively represent the hierarchical features for more accurate recognition. Moreover, FSRL utilizes the primitive form of feature maps to keep the latent structural information, especially in noisy cases. (iii) To further promote the recognition performance, we desire to fuse the hierarchical recognition outputs from different stages. Meanwhile, the discriminability from different scales can also be fully integrated. By exploiting these advantages, the efficiency of the proposed method can be delivered. Experimental results on several face datasets have verified the superiority of the presented algorithm to the other competitive CRFR approaches. Guangwei Gao, Yi Yu 0001, Jian Yang 0003, Guo-Jun Qi, Meng Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Leaning compact and representative features for cross-modality person re-identification
Guangwei Gao, Hao Shao, Fei Wu 0004, Meng Yang 0001, Yi Yu 0001 |
World Wide Web | 4 |
| 2021 | Collaborative Feature Learning and Credible Soft Labeling for Unsupervised Domain Adaptive Person Re-Identification
Haijian Wang, Meng Yang 0001 |
IJCB | 2 |
| 2021 | Semi-Supervised Learning by Exploiting Unlabeled Data Correlations in a Dual-Branch NetworkabstractHow to utilize more abundant unlabeled data to effectively boost the performance of a model with limited labeled training data has attracted much attention. Recent semi-supervised learning works have achieved great success through assigning pseudo-labels to strongly-augmented versions of unlabeled data, and fitting their predictions from a model to the pseudo-labels. But the slow convergence of optimization and the neglect of unlabeled data with low confidence limit the performances of these methods. In this work, we proposed a novel framework of semi-supervised learning by Exploiting the Correlations of unlabeled data in a Dual-Branch Network (EC-DBN). In the proposed ECDBN, the augmented unlabeled data are explored to accelerate the convergence of the methods with data augmentation. Specifically, we apply strong data augmentation to one unlabeled image twice so that the model can obtain more information about unlabeled data from two strongly-augmented versions shared with the same pseudo-label. What’s more, we contract the distances between weakly-augmented version of unlabeled data and their strongly-augmented versions. To align the feature distributions of labeled and unlabeled data, a dual-branch network is also presented by introducing a consistency loss between these two branches. Our experiments demonstrate the effectiveness and better performance of our method on several standard benchmarks compared to other methods. Meng Yang 0001 |
ICME | 2 |
| 2021 | A Hierarchical Inter-Clause Interaction Network for Emotion Cause ExtractionabstractRecently, some methods with inter-clause interaction have achieved promising results on the task of emotion cause extraction. However, the inter-clause modeling modules are only applied on clause-level features rather than word-level features, thus weakening the ability to capture important word-level cues for identifying whether a clause is an emotion cause. In this paper, we propose a framework of Hierarchical Inter-Clause Interaction Network (HICIN), in which inter-clause interaction is applied on both word-level and clause-level features. Word-level interaction can capture the fine-grained semantic cues of each clause by the guidance of all clauses in the document and then generate more powerful clause-level features, while clause-level interaction makes the obtained clause-level features more discriminative. Experimental results show that our model can improve the performance effectively. Peiqin Lin, Meng Yang 0001 |
IJCNN | 2 |
| 2021 | Utilization of Question Categories in Multi-Document Machine Reading ComprehensionabstractMulti-document machine reading comprehension has become a hot topic in natural language processing due to its more realistic setting and wider applications. However, how to effectively exploit the information of multiple documents and the question is still a challenge. In this paper, we propose a new end-to-end reading comprehension model with the utilization of question categories. To compress the search space of the answer and pinpoint it more precisely, we make the best use of the question and its category to predict the length of the answer. To better evaluate the importance of each document and give a more suitable score, we integrate the question category into multi-step reasoning based document extraction. Besides, we propose a new question classification model based on keyword extraction to get the question categories. The experimental results show that our method outperforms the baselines on the English MS MARCO dataset and the Chinese DuReader dataset. Shaomin Zheng, Meng Yang 0001, Yongjie Huang, Peiqin Lin |
IJCNN | 2 |
| 2021 | Breadth First Reasoning Graph for Multi-hop Question AnsweringabstractRecently Graph Neural Network (GNN) has been used as a promising tool in multi-hop question answering task.However, the unnecessary updations and simple edge constructions prevent an accurate answer span extraction in a more direct and interpretable way.In this paper, we propose a novel model of Breadth First Reasoning Graph (BFR-Graph), which presents a new message passing way that better conforms to the reasoning process.In BFR-Graph, the reasoning message is required to start from the question node and pass to the next sentences node hop by hop until all the edges have been passed, which can effectively prevent each node from over-smoothing or being updated multiple times unnecessarily.To introduce more semantics, we also define the reasoning graph as a weighted graph with considering the number of co-occurrence entities and the distance between sentences.Then we present a more direct and interpretable way to aggregate scores from different levels of granularity based on the GNN.On Hot-potQA leaderboard, the proposed BFR-Graph achieves state-of-the-art on answer span prediction. Yongjie Huang, Meng Yang 0001 |
NAACL-HLT | 2 |
| 2021 | Generating Relevant, Correct and Fluent Answers in Natural Answer Generation
Yongjie Huang, Meng Yang 0001 |
NLPCC (2) | 2 |
| 2021 | Suppressing Style-Sensitive Features via Randomly Erasing for Domain Generalizable Semantic Segmentation
Siwei Su, Haijian Wang, Meng Yang 0001 |
PRCV (4) | 3 |
| 2021 | Adversarial Decoupling for Weakly Supervised Semantic Segmentation
Guoying Sun, Meng Yang 0001, Wenfeng Luo |
PRCV (4) | 2 |
| 2021 | Attention-based label consistency for semi-supervised deep learning based image classification
Jiaming Chen 0008, Meng Yang 0001 |
Neurocomputing | 2 |
| 2021 | Constructing multilayer locality-constrained matrix regression framework for noise robust face super-resolution
Guangwei Gao, Yi Yu 0001, Jin Xie 0001, Jian Yang 0003, Meng Yang 0001, Jian Zhang 0002 |
Pattern Recognit. | 5 |
| 2021 | Weakly-supervised semantic segmentation with saliency and incremental supervision updating
Wenfeng Luo, Meng Yang 0001, Wei-Shi Zheng 0001 |
Pattern Recognit. | 2 |
| 2021 | Surrogate network-based sparseness hyper-parameter optimization for deep expression recognition
Weicheng Xie 0001, Wenting Chen, LinLin Shen, Jinming Duan 0001, Meng Yang 0001 |
Pattern Recognit. | 5 |
| 2021 | Deep Selective Memory Network With Selective Attention and Inter-Aspect Modeling for Aspect Level Sentiment ClassificationabstractAspect level sentiment classification aims to recognize the sentiment polarity of each aspect term in a sentence. However, most of the existing methods usually applied the attention mechanism over position-weighted memory and did not consider inter-aspect information. To address these issues, we propose a novel framework for aspect level sentiment classification, Deep Selective Memory Network (DSMN), which selects the context memory dynamically for better guiding the multi-hop attention mechanism and integrates inter-aspect information with deep memory network. By designing a selective attention mechanism based on the distance information between an aspect and its context, DSMN focuses on different parts of the context memory in different memory network layers to capture abundant aspect-aware context information. Besides, to make full use of the inter-aspect information, we also design effective inter-aspect modeling modules to generate both semantic and relation information of the nearby aspects for the desired aspect. We evaluate the advantages of our framework on three benchmark datasets, and experiment results show that our framework achieves state-of-the-art performance. Peiqin Lin, Meng Yang 0001, Jian-Huang Lai |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Hierarchical Attention Network with Pairwise Loss for Chinese Zero Pronoun ResolutionabstractRecent neural network methods for Chinese zero pronoun resolution didn't take bidirectional attention between zero pronouns and candidate antecedents into consideration, and simply treated the task as a classification task, ignoring the relationship between different candidates of a zero pronoun. To solve these problems, we propose a Hierarchical Attention Network with Pairwise Loss (HAN-PL), for Chinese zero pronoun resolution. In the proposed HAN-PL, we design a two-layer attention model to generate more powerful representations for zero pronouns and candidate antecedents. Furthermore, we propose a novel pairwise loss by introducing the correct-antecedent similarity constraint and the pairwise-margin loss, making the learned model more discriminative. Extensive experiments have been conducted on OntoNotes 5.0 dataset, and our model achieves state-of-the-art performance in the task of Chinese zero pronoun resolution. Peiqin Lin, Meng Yang 0001 |
AAAI | 2 |
| 2020 | Learning Saliency-Free Model with Generic Features for Weakly-Supervised Semantic SegmentationabstractCurrent weakly-supervised semantic segmentation methods often estimate initial supervision from class activation maps (CAM), which produce sparse discriminative object seeds and rely on image saliency to provide background cues when only class labels are used. To eliminate the demand of extra data for training saliency detector, we propose to discover class pattern inherent in the lower layer convolution features, which are scarcely explored as in previous CAM methods. Specifically, we first project the convolution features into a low-dimension space and then decide on a decision boundary to generate class-agnostic maps for each semantic category that exists in the image. Features from Lower layer are more generic, thus capable of generating proxy ground-truth with more accurate and integral objects. Experiments on the PASCAL VOC 2012 dataset show that the proposed saliency-free method outperforms the previous approaches under the same weakly-supervised setting and achieves superior segmentation results, which are 64.5% on the validation set and 64.6% on the test set concerning mIoU metric. Wenfeng Luo, Meng Yang 0001 |
AAAI | 2 |
| 2020 | Erasing Integrated Learning: A Simple Yet Effective Approach for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to localize object with only weak supervision like image-level labels. However, a long-standing problem for available techniques based on the classification network is that they often result in highlighting the most discriminative parts rather than the entire extent of object. Nevertheless, trying to explore the integral extent of the object could degrade the performance of image classification on the contrary. To remedy this, we propose a simple yet powerful approach by introducing a novel adversarial erasing technique, erasing integrated learning (EIL). By integrating discriminative region mining and adversarial erasing in a single forward-backward propagation in a vanilla CNN, the proposed EIL explores the high response class-specific area and the less discriminative region simultaneously, thus could maintain high performance in classification and jointly discover the full extent of the object. Furthermore, we apply multiple EIL (MEIL) modules at different levels of the network in a sequential manner, which for the first time integrates semantic features of multiple levels and multiple scales through adversarial erasing learning. In particular, the proposed EIL and advanced MEIL both achieve a new state-of-the-art performance in CUB-200-2011 and ILSVRC 2016 benchmark, making significant improvement in localization while advancing high performance in image classification. Jinjie Mai, Meng Yang 0001, Wenfeng Luo |
CVPR | 2 |
| 2020 | Semi-supervised Semantic Segmentation via Strong-Weak Dual-Branch Network
Wenfeng Luo, Meng Yang 0001 |
ECCV (5) | 2 |
| 2020 | Lightweight Multiple Perspective Fusion with Information Enriching for BERT-Based Answer Selection
Meng Yang 0001, Peiqin Lin |
NLPCC (1) | 2 |
| 2020 | Cross-resolution face recognition with pose variations via multilayer locality-constrained structural orthogonal procrustes regression
Guangwei Gao, Yi Yu 0001, Meng Yang 0001, Heyou Chang, Dong Yue 0001 |
Inf. Sci. | 3 |
| 2020 | Semi-supervised Dual-Branch Network for image classification
Jiaming Chen 0008, Meng Yang 0001, Guangwei Gao |
Knowl. Based Syst. | 2 |
| 2020 | Face image super-resolution with pose via nuclear norm regularized structural orthogonal Procrustes regression
Guangwei Gao, Meng Yang 0001, Huimin Lu 0001, Wankou Yang, Hao Gao 0005 |
Neural Comput. Appl. | 3 |
| 2020 | Adaptive Convolution Local and Global Learning for Class-Level Joint Representation of Facial Recognition With a Single Sample Per Data SubjectabstractDue to the absence of training samples and intraclass variation, the extraction of discriminative facial features and construction of powerful classifiers have bottlenecks in improving the performance of facial recognition (FR) with a single sample per data subject (SSPDS). In this paper, we propose to learn regional adaptive convolution features that are locally and globally discriminative to facial identity and robust to facial variation. Then, a novel class-level joint representation framework is presented to exploit the distinctiveness and class-level commonality of different facial features. In the proposed class-level joint representation with regional adaptive convolution features (CJR-RACF), both discriminative facial features that are robust to facial variations and powerful representations for classification with generic facial variations have been fully exploited. Furthermore, the gallery discrimination is extracted by our proposed weight-embedded supervision in the training phase (denoted by CJR-RACFw), which is conducive to more specific features for FR with SSPDS. CJR-RACF and CJR-RACFw have been evaluated on several popular databases, including the large-scale CMU Multi-PIE, LFW, Megaface, and VGGFace datasets. Experimental results demonstrate the much higher robustness and effectiveness of the proposed methods compared to the state-of-the-art methods. Meng Yang 0001, Xing Wang 0012, LinLin Shen, Guangwei Gao |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2019 | Semantic GAN: Application for Cross-Domain Image Style TransferabstractImage style transfer has attracted much attention from many fields and received promising performance. However, style transfer in the cross-domain field, e.g., the transfer between near-infrared and visible light images, is rarely studied. In the cross-domain image style transfer, one key issue is mismatching problem existing in the generated semantic regions. In this paper, we propose a novel model of Semantic GAN, which integrates the semantic guidance and the recent CycleGAN. In particular, we present a semantic style loss with Gram matrix to well preserve the semantic information in the generated images. The proposed Semantic GAN can control the transfer in the right way with semantic masks and solve the mismatching problem. We apply our approach to two outdoor scene datasets to evaluate the performance of all competing methods. The experimental results show that our approach outperforms previous methods in addressing the mismatching problem and providing a good quality result. Meng Yang 0001 |
ICME | 2 |
| 2019 | Pedestrian re-Identification Based on Tree Branch Network with Local and Global LearningabstractDeep part-based methods in recent literature have revealed the great potential of learning local part-level representation for pedestrian image in the task of person re-identification. However, global features that capture discriminative holistic information of human body are usually ignored or not well exploited. This motivates us to investigate joint learning global and local features from pedestrian images. Specifically, in this work, we propose a novel framework termed tree branch network (TBN) for person re-identification. Given a pedestrain image, the feature maps generated by the backbone CNN, are partitioned recursively into several pieces, each of which is followed by a bottleneck structure that learns finer-grained features for each level in the hierarchical tree-like framework. In this way, representations are learned in a coarse-to-fine manner and finally assembled to produce more discriminative image descriptions. Experimental results demonstrate the effectiveness of the global and local feature learning method in the proposed TBN framework. We also show significant improvement in performance over state-of-the-art methods on three public benchmarks: Market-1501, CUHK-03 and DukeMTMC. Meng Yang 0001, Zhihui Lai 0001, Wei-Shi Zheng 0001, Zitong Yu |
ICME | 2 |
| 2019 | Deep Mask Memory Network with Semantic Dependency and Context Moment for Aspect Level Sentiment ClassificationabstractAspect level sentiment classification aims at identifying the sentiment of each aspect term in a sentence. Deep memory networks often use location information between context word and aspect to generate the memory. Although improved results are achieved, the relation information among aspects in the same sentence is ignored and the word location can't bring enough and accurate information for the analysis on the aspect sentiment. In this paper, we propose a novel framework for aspect level sentiment classification, deep mask memory network with semantic dependency and context moment (DMMN-SDCM), which integrates semantic parsing information of the aspect and the inter-aspect relation information into deep memory network. With the designed attention mechanism based on semantic dependency information, different parts of the context memory in different computational layers are selected and useful inter-aspect information in the same sentence is exploited for the desired aspect. To make full use of the inter-aspect relation information, we also jointly learn a context moment learning task, which aims to learn the sentiment distribution of the entire sentence for providing a background for the desired aspect. We examined the merit of our model on SemEval 2014 Datasets, and the experimental results show that our model achieves a state-of-the-art performance. Peiqin Lin, Meng Yang 0001, Jian-Huang Lai |
IJCAI | 2 |
| 2019 | Attention-Based Label Consistency for Semi-supervised Deep Learning
Jiaming Chen 0008, Meng Yang 0001 |
PRCV (1) | 2 |
| 2019 | Multi-branch Structure for Hierarchical Classification in Plant Disease Recognition
Zihao Mao, Jiaming Chen 0008, Meng Yang 0001 |
PRCV (3) | 3 |
| 2019 | Person ReID: Optimization of Domain Adaption Though Clothing Style Transfer Between Datasets
Haijian Wang, Meng Yang 0001, Linbin Ye |
PRCV (3) | 2 |
| 2019 | Triple-translation GAN with multi-layer sparse representation for face image synthesis
Linbin Ye, Bob Zhang 0001, Meng Yang 0001, Wei Lian |
Neurocomputing | 3 |
| 2019 | Robust joint representation with triple local feature for face recognition with single sample per person
Xing Wang 0012, Bob Zhang 0001, Meng Yang 0001, Kangyin Ke, Wei-Shi Zheng 0001 |
Knowl. Based Syst. | 3 |
| 2019 | Sparse deep feature learning for facial expression recognition
Weicheng Xie 0001, Xi Jia, LinLin Shen, Meng Yang 0001 |
Pattern Recognit. | 4 |
| 2018 | Toward Characteristic-Preserving Image-Based Virtual Try-On Network
Bochao Wang, Huabin Zheng, Xiaodan Liang, Meng Yang 0001 |
ECCV (13) | 6 |
| 2018 | Hierarchical Hybrid Code Networks for Task-Oriented Dialogue
Weiri Liang, Meng Yang 0001 |
ICIC (2) | 2 |
| 2018 | Semi-supervised convolutional neural networks with label propagation for image classificationabstractOver the past several years, deep learning has achieved promising performance in many visual tasks, e.g., face verification and object classification. However, a limited number of labeled training samples existing in practical applications is still a huge bottleneck for achieving a satisfactory performance. In this paper, we integrate class estimation of unlabeled training data with deep learning model which generates a novel semi-supervised convolutional neural network (SSCNN) trained by both the labeled training data and unlabeled data. In the framework of SSCNN, the deep convolution feature extraction and the class estimation of the unlabeled data are jointly learned. Specifically, deep convolution features are learned from the labeled training data and unlabeled data with confident class estimation. After the deep features are obtained, the label propagation algorithm is utilized to estimate the identities of unlabeled training samples. The alternative optimization of SSCNN makes the class estimation of unlabeled data more and more accurate due to the learned CNN feature more and more discriminative. We compared the proposed SSCNN with some representative semi-supervised learning approaches on MINIST and Cifar-10 databases. Extensive experiments on landmark databases show the effectiveness of our semi-supervised deep learning framework. Shiqi Yu 0001, Meng Yang 0001 |
ICPR | 3 |
| 2018 | Fast Skin Lesion Segmentation via Fully Convolutional Network with Residual Architecture and CRFabstractMelanoma is known to be the most fatal form of skin cancers. In order to achieve automated diagnosis of such disease, a system is needed to accurately locate suspicious skin lesions using images captured by standard digital cameras. Recently, there exists a trend for the use of Fully Convolutional Net-work(FCN) to perform image segmentation task. In this paper, we propose a FCN-based processing pipeline that incorporates a deep neural net and a graphical model, to attain a segmentation mask of lesion region from normal skin. Our method extends the residual network by adding a transposed convolution layer to yield a FCN architecture. We demonstrate that the noisy outcome from FCN can be refined by a fully connected Conditional Random Field(CRF). Our model enjoys three major advantages over existing algorithms: simpler process pipeline, state-of-art accuracy in terms of segmentation sensitivity(95.6%) and fast inference time. Wenfeng Luo, Meng Yang 0001 |
ICPR | 2 |
| 2018 | Adaptive convolution local and global learning for class-level joint representation of face recognition with single sample per personabstractDue to the absence of samples with intra-class variation, extracting discriminative facial features and building powerful classifiers are the bottlenecks of improving the performance of face recognition (FR) with single sample per person (SSPP). In this paper, we propose to learn regional adaptive convolution features which are locally and globally discriminative to face identity and robust to face variation. With collected generic facial variations, a novel class-level joint representation framework is presented to exploit the distinctiveness and class-level commonality of different facial features. In the proposed class-level joint representation with regional adaptive convolution feature (CJR-RACF), both discriminative facial features robust to various facial variations and powerful representation for classification with generic facial variations that can overcome the small-sample-size problem are fully exploited. CJR-RACF has been evaluated on several popular databases, including large-scale CMU Multi-PIE and LFW databases. Experimental results demonstrate the much higher robustness and effectiveness of CJR-RACF to complex facial variations compared to the state-of-the-art methods. Xing Wang 0012, LinLin Shen, Meng Yang 0001 |
ICPR | 4 |
| 2018 | Multi-feature Shared and Specific Representation for Pattern Classification
Kangyin Ke, Meng Yang 0001 |
PRCV (3) | 2 |
| 2018 | Semi-supervised Dictionary Active Learning for Pattern Classification
Qin Zhong, Meng Yang 0001, Tiancheng Zhang 0001 |
PRCV (3) | 2 |
| 2018 | Open snake model based on global guidance field for embryo vessel locationabstractThe development of vessels can provide important information about the growth status of animal embryos. It is, therefore, important to automatically locate the deformed vessel branches from the embryo images. However, very few vessel detectors can accurately locate all vessel branches when the captured images are low quality and the implied vessel shapes are complex. In this study, a new framework consisting of vessel region extraction and snake shape optimisation is proposed. The main contribution in this detector is a novel open snake model based on the global guidance field and deformation template initialisation. Experimental results on a specific application of an embryo vessel database [Database and source codes: https://github.com/wcxie/Egg‐embryro‐vessel‐location/ .] demonstrate that the proposed algorithm not only locates the vessel shape properly but also obtains the orientations of embryo vessel branches accurately. Comparison to traditional guidance fields and the active appearance model illustrates the effectiveness and competitiveness of the proposed model. Weicheng Xie 0001, Jinming Duan 0001, LinLin Shen, Yuexiang Li, Meng Yang 0001, Guojun Lin |
IET Comput. Vis. | 5 |
| 2018 | Facial expression synthesis with direction field preservation based mesh deformation and lighting fitting based wrinkle mapping
Weicheng Xie 0001, LinLin Shen, Meng Yang 0001, Jianmin Jiang |
Multim. Tools Appl. | 3 |
| 2018 | Robust, discriminative and comprehensive dictionary learning for face recognition
Guojun Lin, Meng Yang 0001, Jian Yang 0003, LinLin Shen, Weicheng Xie 0001 |
Pattern Recognit. | 2 |
| 2017 | Discriminative Semi-Supervised Dictionary Learning with Entropy Regularization for Pattern ClassificationabstractDictionary learning has played an important role in the success of sparse representation, which triggers the rapid developments of unsupervised and supervised dictionary learning methods. However, in most practical applications, there are usually quite limited labeled training samples while it is relatively easy to acquire abundant unlabeled training samples. Thus semi-supervised dictionary learning that aims to effectively explore the discrimination of unlabeled training data has attracted much attention of researchers. Although various regularizations have been introduced in the prevailing semi-supervised dictionary learning, how to design an effective unified model of dictionary learning and unlabeled-data class estimating and how to well explore the discrimination in the labeled and unlabeled data are still open. In this paper, we propose a novel discriminative semi-supervised dictionary learning model (DSSDL) by introducing discriminative representation, an identical coding of unlabeled data to the coding of testing data final classification, and an entropy regularization term. The coding strategy of unlabeled data can not only avoid the affect of its incorrect class estimation, but also make the learned discrimination be well exploited in the final classification. The introduced regularization of entropy can avoid overemphasizing on some uncertain estimated classes for unlabeled samples. Apart from the enhanced discrimination in the learned dictionary by the discriminative representation, an extended dictionary is used to mainly explore the discrimination embedded in the unlabeled data. Extensive experiments on face recognition, digit recognition and texture classification show the effectiveness of the proposed method. Meng Yang 0001 |
AAAI | 1 |
| 2017 | Semi-supervised dictionary learning with label propagation for image classificationabstractSparse coding and supervised dictionary learning have rapidly developed in recent years, and achieved impressive performance in image classification. However, there is usually a limited number of labeled training samples and a huge amount of unlabeled data in practical image classification, which degrades the discrimination of the learned dictionary. How to effectively utilize unlabeled training data and explore the information hidden in unlabeled data has drawn much attention of researchers. In this paper, we propose a novel discriminative semi-supervised dictionary learning method using label propagation (SSD-LP). Specifically, we utilize a label propagation algorithm based on class-specific reconstruction errors to accurately estimate the identities of unlabeled training samples, and develop an algorithm for optimizing the discriminative dictionary and discriminative coding vectors simultaneously. Extensive experiments on face recognition, digit recognition, and texture classification demonstrate the effectiveness of the proposed method. Meng Yang 0001 |
Comput. Vis. Media | 2 |
| 2017 | Discriminative analysis-synthesis dictionary learning for image classification
Meng Yang 0001, Heyou Chang, Weixin Luo |
Neurocomputing | 1 |
| 2017 | Fisher discrimination dictionary pair learning for image classification
Meng Yang 0001, Heyou Chang, Weixin Luo, Jian Yang 0003 |
Neurocomputing | 1 |
| 2017 | Image set classification based on synthetic examples and reverse training
Lin Zhang 0014, Qingjun Liang, Ying Shen 0005, Meng Yang 0001, Feng Liu 0013 |
Neurocomputing | 4 |
| 2017 | Joint regularized nearest points for image set based face recognition
Meng Yang 0001, Xing Wang 0012, Weiyang Liu, LinLin Shen |
Image Vis. Comput. | 1 |
| 2017 | Joint and collaborative representation with local adaptive convolution feature for face recognition with single sample per person
Meng Yang 0001, Xing Wang 0012, Guohang Zeng, LinLin Shen |
Pattern Recognit. | 1 |
| 2017 | Towards contactless palmprint recognition: A novel device, a new benchmark, and a collaborative representation based identification approach
Lin Zhang 0014, Lida Li, Anqi Yang, Ying Shen 0005, Meng Yang 0001 |
Pattern Recognit. | 5 |
| 2016 | Analysis-Synthesis Dictionary Learning for Universality-Particularity Representation Based ClassificationabstractDictionary learning has played an important role in the success of sparse representation. Although synthesis dictionary learning for sparse representation has been well studied for universality representation (i.e., the dictionary is universal to all classes) and particularity representation (i.e., the dictionary is class-particular), jointly learning an analysis dictionary and a synthesis dictionary is still in its infant stage. Universality-particularity representation can well match the intrinsic characteristics of data (i.e., different classes share commonality and distinctness), while analysis-synthesis dictionary can give a more complete view of data representation (i.e., analysis dictionary is a dual-viewpoint of synthesis dictionary). In this paper, we proposed a novel model of analysis-synthesis dictionary learning for universality-particularity (ASDL-UP) representation based classification. The discrimination of universality and particularity representation is jointly exploited by simultaneously learning a pair of analysis dictionary and synthesis dictionary. More specifically, we impose a label preserving term to analysis coding coefficients for universality representation. Fisher-like regularizations for analysis coding coefficients and the subsequent synthesis representation are introduced to particularity representation. Compared with other state-of-the-art dictionary learning methods, ASDL-UP has shown better or competitive performance in various classification tasks. Meng Yang 0001, Weiyang Liu, Weixin Luo, LinLin Shen |
AAAI | 1 |
| 2016 | Jointly Learning Non-negative Projection and Dictionary with Discriminative Graph Constraints for Classification
Weiyang Liu, Zhiding Yu, Yandong Wen, Rongmei Lin, Meng Yang 0001 |
BMVC | 5 |
| 2016 | Large-Margin Softmax Loss for Convolutional Neural NetworksabstractCross-entropy loss together with softmax is arguably one of the most common used supervision components in convolutional neural networks (CNNs). Despite its simplicity, popularity and excellent performance, the component does not explicitly encourage discriminative learning of features. In this paper, we propose a generalized large-margin softmax (L-Softmax) loss which explicitly encourages intra-class compactness and inter-class separability between learned features. Moreover, L-Softmax not only can adjust the desired margin but also can avoid overfitting. We also show that the L-Softmax loss can be optimized by typical stochastic gradient descent. Extensive experiments on four benchmark datasets demonstrate that the deeply-learned features with L-softmax loss become more discriminative, hence significantly boosting the performance on a variety of visual classification and verification tasks. Weiyang Liu, Yandong Wen, Zhiding Yu, Meng Yang 0001 |
ICML | 4 |
| 2016 | Schatten p-norm based principal component analysis
Heyou Chang, Lei Luo 0001, Jian Yang 0003, Meng Yang 0001 |
Neurocomputing | 4 |
| 2016 | Learning a structure adaptive dictionary for sparse representation based classification
Heyou Chang, Meng Yang 0001, Jian Yang 0003 |
Neurocomputing | 2 |
| 2016 | Structured regularized robust coding for face recognition
Xing Wang 0012, Meng Yang 0001, LinLin Shen |
Neurocomputing | 2 |
| 2016 | Structured occlusion coding for robust face recognition
Yandong Wen, Weiyang Liu, Meng Yang 0001, Yuli Fu 0001, Youjun Xiang, Rui Hu 0008 |
Neurocomputing | 3 |
| 2016 | 3D Ear Identification Using Block-Wise Statistics-Based Features and LC-KSVDabstractBiometrics authentication has been corroborated to be an effective method for recognizing a person's identity with high confidence. In this field, the use of three-dimensional (3D) ear shape is a recent trend. As a biometric identifier, the ear has several inherent merits. However, although a great deal of efforts have been devoted, there is still large room for improvement in developing a highly effective and efficient 3D ear identification approach. In this paper, we attempt to fill this gap to some extent by proposing a novel 3D ear classification scheme that makes use of the label consistent K-SVD (LC-KSVD) framework. As an effective supervised dictionary learning algorithm, LC-KSVD learns a single compact discriminative dictionary for sparse coding and a multi-class linear classifier simultaneously. To use the LC-KSVD framework, one key issue is how to extract feature vectors from 3D ear scans. To this end, we propose a blockwise statistics-based feature extraction scheme. Specifically, we divide a 3D ear region of interest into uniform blocks and extract a histogram of surface types from each block; histograms from all blocks are then concatenated to form the desired feature vector. Feature vectors extracted in this way are highly discriminative and are robust to mere misalignment between samples. Experiments demonstrate that our approach can achieve better recognition accuracy than the other state-of-the-art methods. More importantly, its computational complexity is extremely low, making it quite suitable for the large-scale identification applications. MATLAB source codes are publicly online available at http://sse.tongji.edu.cn/linzhang/LCKSVDEar/LCKSVDEar. htm. Lin Zhang 0014, Lida Li, Hongyu Li 0001, Meng Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2015 | Multi-kernel collaborative representation for image classificationabstractWe consider the image classification problem via multiple kernel collaborative representation (MKCR). We generalize the kernel collaborative representation based classification to a multi-kernel framework where multiple kernels are jointly learned with the representation coefficients. The intrinsic idea of multiple kernel learning is adopted in our MKCR model. Experimental results show MKCR converges within reasonable iterations and achieves state-of-the-art performance. Weiyang Liu, Zhiding Yu, Yandong Wen, Meng Yang 0001, Yuexian Zou |
ICIP | 4 |
| 2015 | Joint kernel dictionary and classifier learning for sparse coding via locality preserving K-SVDabstractWe present a locality preserving K-SVD (LP-KSVD) algorithm for joint dictionary and classifier learning, and further incorporate kernel into our framework. In LP-KSVD, we construct a locality preserving term based on the relations between input samples and dictionary atoms, and introduce the locality via nearest neighborhood to enforce the locality of representation. Motivated by the fact that locality-related methods works better in a more discriminative and separable space, we map the original feature space to the kernel space, where samples of different classes become more separable. Experimental results show the proposed approach has strong discrimination power and is comparable or outperforms some state-of-the-art approaches on public databases. Weiyang Liu, Zhiding Yu, Meng Yang 0001, Lijia Lu, Yuexian Zou |
ICME | 3 |
| 2015 | Joint representation and pattern learning for robust face recognition
Meng Yang 0001, Pengfei Zhu 0001, Feng Liu 0013, LinLin Shen |
Neurocomputing | 1 |
| 2014 | Local Generic Representation for Face Recognition with Single Sample per Person
Pengfei Zhu 0001, Meng Yang 0001, Lei Zhang 0006, Il-Yong Lee |
ACCV (3) | 2 |
| 2014 | Latent Dictionary Learning for Sparse Representation Based ClassificationabstractDictionary learning (DL) for sparse coding has shown promising results in classification tasks, while how to adaptively build the relationship between dictionary atoms and class labels is still an important open question. The existing dictionary learning approaches simply fix a dictionary atom to be either class-specific or shared by all classes beforehand, ignoring that the relationship needs to be updated during DL. To address this issue, in this paper we propose a novel latent dictionary learning (LDL) method to learn a discriminative dictionary and build its relationship to class labels adaptively. Each dictionary atom is jointly learned with a latent vector, which associates this atom to the representation of different classes. More specifically, we introduce a latent representation model, in which discrimination of the learned dictionary is exploited via minimizing the within-class scatter of coding coefficients and the latent-value weighted dictionary coherence. The optimal solution is efficiently obtained by the proposed solving algorithm. Correspondingly, a latent sparse representation based classifier is also presented. Experimental results demonstrate that our algorithm outperforms many recently proposed sparse representation and dictionary learning approaches for action, gender and face recognition. Meng Yang 0001, Dengxin Dai, Lilin Shen, Luc Van Gool |
CVPR | 1 |
| 2014 | Sparse Representation Based Fisher Discrimination Dictionary Learning for Image Classification
Meng Yang 0001, Lei Zhang 0006, Xiangchu Feng, David Zhang 0001 |
Int. J. Comput. Vis. | 1 |
| 2014 | Multi-granularity distance metric learning via neighborhood granule margin maximization
Pengfei Zhu 0001, Qinghua Hu, Wangmeng Zuo, Meng Yang 0001 |
Inf. Sci. | 4 |
| 2014 | Fast and robust face recognition via coding residual map learning based adaptive masking
Meng Yang 0001, Zhizhao Feng, Simon C. K. Shiu, Lei Zhang 0006 |
Pattern Recognit. | 1 |
| 2013 | Sparse Variation Dictionary Learning for Face Recognition with a Single Training Sample per PersonabstractFace recognition (FR) with a single training sample per person (STSPP) is a very challenging problem due to the lack of information to predict the variations in the query sample. Sparse representation based classification has shown interesting results in robust FR, however, its performance will deteriorate much for FR with STSPP. To address this issue, in this paper we learn a sparse variation dictionary from a generic training set to improve the query sample representation by STSPP. Instead of learning from the generic training set independently w.r.t. the gallery set, the proposed sparse variation dictionary learning (SVDL) method is adaptive to the gallery set by jointly learning a projection to connect the generic training set with the gallery set. The learnt sparse variation dictionary can be easily integrated into the framework of sparse representation based classification so that various variations in face images, including illumination, expression, occlusion, pose, etc., can be better handled. Experiments on the large-scale CMU Multi-PIE, FRGC and LFW databases demonstrate the promising performance of SVDL on FR with STSPP. Meng Yang 0001, Luc Van Gool, Lei Zhang 0006 |
ICCV | 1 |
| 2013 | Joint discriminative dimensionality reduction and dictionary learning for face recognition
Zhizhao Feng, Meng Yang 0001, Lei Zhang 0006, Yan Liu 0004, David Zhang 0001 |
Pattern Recognit. | 2 |
| 2013 | Gabor feature based robust representation and classification for face recognition with Gabor occlusion dictionary
Meng Yang 0001, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001 |
Pattern Recognit. | 1 |
| 2013 | Regularized Robust Coding for Face RecognitionabstractRecently the sparse representation based classification (SRC) has been proposed for robust face recognition (FR). In SRC, the testing image is coded as a sparse linear combination of the training samples, and the representation fidelity is measured by the l2-norm or l1 -norm of the coding residual. Such a sparse coding model assumes that the coding residual follows Gaussian or Laplacian distribution, which may not be effective enough to describe the coding residual in practical FR systems. Meanwhile, the sparsity constraint on the coding coefficients makes the computational cost of SRC very high. In this paper, we propose a new face coding model, namely regularized robust coding (RRC), which could robustly regress a given signal with regularized regression coefficients. By assuming that the coding residual and the coding coefficient are respectively independent and identically distributed, the RRC seeks for a maximum a posterior solution of the coding problem. An iteratively reweighted regularized robust coding (IR(3)C) algorithm is proposed to solve the RRC model efficiently. Extensive experiments on representative face databases demonstrate that the RRC is much more effective and efficient than state-of-the-art sparse representation based methods in dealing with face occlusion, corruption, lighting, and expression changes, etc. Meng Yang 0001, Lei Zhang 0006, Jian Yang 0003, David Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Robust Kernel Representation With Statistical Local Features for Face RecognitionabstractFactors such as misalignment, pose variation, and occlusion make robust face recognition a difficult problem. It is known that statistical features such as local binary pattern are effective for local feature extraction, whereas the recently proposed sparse or collaborative representation-based classification has shown interesting results in robust face recognition. In this paper, we propose a novel robust kernel representation model with statistical local features (SLF) for robust face recognition. Initially, multipartition max pooling is used to enhance the invariance of SLF to image registration error. Then, a kernel-based representation model is proposed to fully exploit the discrimination information embedded in the SLF, and robust regression is adopted to effectively handle the occlusion in face images. Extensive experiments are conducted on benchmark face databases, including extended Yale B, AR (A. Martinez and R. Benavente), multiple pose, illumination, and expression (multi-PIE), facial recognition technology (FERET), face recognition grand challenge (FRGC), and labeled faces in the wild (LFW), which have different variations of lighting, expression, pose, and occlusions, demonstrating the promising performance of the proposed method. Meng Yang 0001, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | Relaxed collaborative representation for pattern classificationabstractRegularized linear representation learning has led to interesting results in image classification, while how the object should be represented is a critical issue to be investigated. Considering the fact that the different features in a sample should contribute differently to the pattern representation and classification, in this paper we present a novel relaxed collaborative representation (RCR) model to effectively exploit the similarity and distinctiveness of features. In RCR, each feature vector is coded on its associated dictionary to allow flexibility of feature coding, while the variance of coding vectors is minimized to address the similarity among features. In addition, the distinctiveness of different features is exploited by weighting its distance to other features in the coding domain. The proposed RCR is simple, while our extensive experimental results on benchmark image databases (e.g., various face and flower databases) show that it is very competitive with state-of-the-art image classification methods. Meng Yang 0001, Lei Zhang 0006, David Zhang 0001, Shenlong Wang |
CVPR | 1 |
| 2012 | Efficient Misalignment-Robust Representation for Real-Time Face Recognition
Meng Yang 0001, Lei Zhang 0006, David Zhang 0001 |
ECCV (1) | 1 |
| 2012 | Monogenic Binary Coding: An Efficient Local Feature Extraction Approach to Face RecognitionabstractLocal-feature-based face recognition (FR) methods, such as Gabor features encoded by local binary pattern, could achieve state-of-the-art FR results in large-scale face databases such as FERET and FRGC. However, the time and space complexity of Gabor transformation are too high for many practical FR applications. In this paper, we propose a new and efficient local feature extraction scheme, namely monogenic binary coding (MBC), for face representation and recognition. Monogenic signal representation decomposes an original signal into three complementary components: amplitude, orientation, and phase. We encode the monogenic variation in each local region and monogenic feature in each pixel, and then calculate the statistical features (e.g., histogram) of the extracted local features. The local statistical features extracted from the complementary monogenic components (i.e., amplitude, orientation, and phase) are then fused for effective FR. It is shown that the proposed MBC scheme has significantly lower time and space complexity than the Gabor-transformation-based local feature methods. The extensive FR experiments on four large-scale databases demonstrated the effectiveness of MBC, whose performance is competitive with and even better than state-of-the-art local-feature-based FR methods. Meng Yang 0001, Lei Zhang 0006, Simon C. K. Shiu, David Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2011 | Robust sparse coding for face recognitionabstractRecently the sparse representation (or coding) based classification (SRC) has been successfully used in face recognition. In SRC, the testing image is represented as a sparse linear combination of the training samples, and the representation fidelity is measured by the l2-norm or l1-norm of coding residual. Such a sparse coding model actually assumes that the coding residual follows Gaussian or Laplacian distribution, which may not be accurate enough to describe the coding errors in practice. In this paper, we propose a new scheme, namely the robust sparse coding (RSC), by modeling the sparse coding as a sparsity-constrained robust regression problem. The RSC seeks for the MLE (maximum likelihood estimation) solution of the sparse coding problem, and it is much more robust to outliers (e.g., occlusions, corruptions, etc.) than SRC. An efficient iteratively reweighted sparse coding algorithm is proposed to solve the RSC model. Extensive experiments on representative face databases demonstrate that the RSC scheme is much more effective than state-of-the-art methods in dealing with face occlusion, corruption, lighting and expression changes, etc. Meng Yang 0001, Lei Zhang 0006, Jian Yang 0003, David Zhang 0001 |
CVPR | 1 |
| 2011 | Fisher Discrimination Dictionary Learning for sparse representationabstractSparse representation based classification has led to interesting image recognition results, while the dictionary used for sparse coding plays a key role in it. This paper presents a novel dictionary learning (DL) method to improve the pattern classification performance. Based on the Fisher discrimination criterion, a structured dictionary, whose dictionary atoms have correspondence to the class labels, is learned so that the reconstruction error after sparse coding can be used for pattern classification. Meanwhile, the Fisher discrimination criterion is imposed on the coding coefficients so that they have small within-class scatter but big between-class scatter. A new classification scheme associated with the proposed Fisher discrimination DL (FDDL) method is then presented by using both the discriminative information in the reconstruction error and sparse coding coefficients. The proposed FDDL is extensively evaluated on benchmark image databases in comparison with existing sparse representation and DL based classification methods. Meng Yang 0001, Lei Zhang 0006, Xiangchu Feng, David Zhang 0001 |
ICCV | 1 |
| 2011 | Sparse representation or collaborative representation: Which helps face recognition?abstractAs a recently proposed technique, sparse representation based classification (SRC) has been widely used for face recognition (FR). SRC first codes a testing sample as a sparse linear combination of all the training samples, and then classifies the testing sample by evaluating which class leads to the minimum representation error. While the importance of sparsity is much emphasized in SRC and many related works, the use of collaborative representation (CR) in SRC is ignored by most literature. However, is it really the l1-norm sparsity that improves the FR accuracy? This paper devotes to analyze the working mechanism of SRC, and indicates that it is the CR but not the l1-norm sparsity that makes SRC powerful for face classification. Consequently, we propose a very simple yet much more efficient face classification scheme, namely CR based classification with regularized least square (CRC_RLS). The extensive experiments clearly show that CRC_RLS has very competitive classification results, while it has significantly less complexity than SRC. Lei Zhang 0006, Meng Yang 0001, Xiangchu Feng |
ICCV | 2 |
| 2010 | Gabor Feature Based Sparse Representation for Face Recognition with Gabor Occlusion Dictionary
Meng Yang 0001, Lei Zhang 0006 |
ECCV (6) | 1 |
| 2010 | Metaface learning for sparse representation based face recognitionabstractFace recognition (FR) is an active yet challenging topic in computer vision applications. As a powerful tool to represent high dimensional data, recently sparse representation based classification (SRC) has been successfully used for FR. This paper discusses the metaface learning (MFL) of face images under the framework of SRC. Although directly using the training samples as dictionary bases can achieve good FR performance, a well learned dictionary matrix can lead to higher FR rate with less dictionary atoms. An SRC oriented unsupervised MFL algorithm is proposed in this paper and the experimental results on benchmark face databases demonstrated the improvements brought by the proposed MFL algorithm over original SRC. Meng Yang 0001, Lei Zhang 0006, Jian Yang 0003, David Zhang 0001 |
ICIP | 1 |
| 2010 | Monogenic Binary Pattern (MBP): A Novel Feature Extraction and Representation Model for Face RecognitionabstractA novel feature extraction method, namely monogenic binary pattern (MBP), is proposed in this paper based on the theory of monogenic signal analysis, and the histogram of MBP (HMBP) is subsequently presented for robust face representation and recognition. MBP consists of two parts: one is monogenic magnitude encoded via uniform LBP, and the other is monogenic orientation encoded as quadrant-bit codes. The HMBP is established by concatenating the histograms of MBP of all sub-regions. Compared with the well-known and powerful Gabor filtering based LBP schemes, one clear advantage of HMBP is its lower time and space complexity because monogenic signal analysis needs fewer convolutions and generates more compact feature vectors. The experimental results on the AR and FERET face databases validate that the proposed MBP algorithm has better performance than or comparable performance with state-of-the-art local feature based methods but with significantly lower time and space complexity. Meng Yang 0001, Lei Zhang 0006, Lin Zhang 0014, David Zhang 0001 |
ICPR | 1 |
| 2010 | On the Dimensionality Reduction for Sparse Representation Based Face RecognitionabstractFace recognition (FR) is an active yet challenging topic in computer vision applications. As a powerful tool to represent high dimensional data, recently sparse representation based classification (SRC) has been successfully used for FR. This paper discusses the dimensionality reduction (DR) of face images under the framework of SRC. Although one important merit of SRC is that it is insensitive to DR or feature extraction, a well trained projection matrix can lead to higher FR rate at a lower dimensionality. An SRC oriented unsupervised DR algorithm is proposed in this paper and the experimental results on benchmark face databases demonstrated the improvements brought by the proposed DR algorithm over PCA or random projection based DR under the SRC framework. Lei Zhang 0006, Meng Yang 0001, Zhizhao Feng, David Zhang 0001 |
ICPR | 2 |