Lan Yan

dblp:211/8085 · DBLP profile ↗
← Back
40ranked-venue papers
11as first author
26since 2021 · last 2025
0000-0001-6452-9649ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Open World Person Retrieval and Its Distributed Parallel Training Framework
Lan Yan, Wenbo Zheng 0001, Kenli Li 0001
ICA3PP (3)1
2025 Source-free domain adaptive person search
Lan Yan, Wenbo Zheng 0001, Kenli Li 0001
Pattern Recognit.1
2024 ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
abstract
Despite the advancements of open-source large language models (LLMs), e.g., LLaMA, they remain significantly limited in tool-use capabilities, i.e., using external tools (APIs) to fulfill human instructions. The reason is that current instruction tuning largely focuses on basic language tasks but ignores the tool-use domain. This is in contrast to the excellent tool-use capabilities of state-of-the-art (SOTA) closed-source LLMs, e.g., ChatGPT. To bridge this gap, we introduce ToolLLM, a general tool-use framework encompassing data construction, model training, and evaluation. We first present ToolBench, an instruction-tuning dataset for tool use, which is constructed automatically using ChatGPT. Specifically, the construction can be divided into three stages: (i) API collection: we collect 16,464 real-world RESTful APIs spanning 49 categories from RapidAPI Hub; (ii) instruction generation: we prompt ChatGPT to generate diverse instructions involving these APIs, covering both single-tool and multi-tool scenarios; (iii) solution path annotation: we use ChatGPT to search for a valid solution path (chain of API calls) for each instruction. To enhance the reasoning capabilities of LLMs, we develop a novel depth-first search-based decision tree algorithm. It enables LLMs to evaluate multiple reasoning traces and expand the search space. Moreover, to evaluate the tool-use capabilities of LLMs, we develop an automatic evaluator: ToolEval. Based on ToolBench, we fine-tune LLaMA to obtain an LLM ToolLLaMA, and equip it with a neural API retriever to recommend appropriate APIs for each instruction. Experiments show that ToolLLaMA demonstrates a remarkable ability to execute complex instructions and generalize to unseen APIs, and exhibits comparable performance to ChatGPT. Our ToolLLaMA also demonstrates strong zero-shot generalization ability in an out-of-distribution tool-use dataset: APIBench.
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin 0001, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou 0016, Mark Gerstein, Dahai Li, Zhiyuan Liu 0001, Maosong Sun 0001
ICLR5
2024 Unknown Instance Learning for Person Search
abstract
Although person search has achieved great progress with advanced deep neural networks, it remains an intricate challenge, necessitating a concurrent solution to pedestrian detection and person re-identification. Due to the difficulty of accurate labeling, the current widely used person search datasets contain numerous unknown identities in addition to the carefully labeled known ones. However, previous efforts usually simply utilize these unknown persons as negative samples, inadvertently disregarding the wealth of untapped sample information. To tackle this issue, we propose a novel Unknown Instance Learning (UIL) method that introduces an extra network branch to mine the potential identity information of unknown individuals during training, apart from the fundamental search branch. Specifically, a Clustering-Based Unknowns Labeling (CBUL) mechanism is introduced to effectively exploit the features extracted from unknown persons. In addition, recognizing the distinct roles of pseudo and real identity labels in discriminative feature learning, we design a Reformative Online Instance Matching (ROIM) loss and maintain lookup tables for real and pseudo labels separately. Extensive experiments on the CUHK-SYSU and PRW datasets demonstrate the superiority of our approach in comparison to other state-of-the-art person search models.
Lan Yan, Kenli Li 0001
ICME1
2024 Knowledge is power: Open-world knowledge representation learning for knowledge-based visual reasoning
Wenbo Zheng 0001, Lan Yan, Fei-Yue Wang 0001
Artif. Intell.2
2024 Knowledge-Embedded Mutual Guidance for Visual Reasoning
abstract
Visual reasoning between visual images and natural language is a long-standing challenge in computer vision. Most of the methods aim to look for answers to questions only on the basis of the analysis of the offered questions and images. Other approaches treat knowledge graphs as flattened tables to search for the answer. However, there are two major problems with these works: 1) the model disregards the fact that the world we surrounding us interlinks our hearing and speaking of natural language and 2) the model largely ignores the structure of the acrlong KG. To overcome these challenging deficiencies, a model should jointly consider two modalities of vision and language, as well as the rich structural and logical information embedded in knowledge graphs. To this end, we propose a general joint representation learning framework for visual reasoning, namely, knowledge-embedded mutual guidance. It realizes mutual guidance not only between visual data and natural language descriptions but also between knowledge graphs and reasoning models. In addition, it exploits the knowledge derived from the reasoning model to boost knowledge graphs when applying the visual relation detection task. The experimental results demonstrate that the proposed approach performs dramatically better than state-of-the-art methods on two benchmarks for visual reasoning.
Wenbo Zheng 0001, Lan Yan, Long Chen 0001, Qiang Li 0060, Fei-Yue Wang 0001
IEEE Trans. Cybern.2
2024 Webly Supervised Knowledge-Embedded Model for Visual Reasoning
abstract
Visual reasoning between visual images and natural language remains a long-standing challenge in computer vision. Conventional deep supervision methods target at finding answers to the questions relying on the datasets containing only a limited amount of images with textual ground-truth descriptions. Facing learning with limited labels, it is natural to expect to constitute a larger scale dataset consisting of several million visual data annotated with texts, but this approach is extremely time-intensive and laborious. Knowledge-based works usually treat knowledge graphs (KGs) as static flattened tables for searching the answer, but fail to take advantage of the dynamic update of KGs. To overcome these deficiencies, we propose a Webly supervised knowledge-embedded model for the task of visual reasoning. On the one hand, vitalized by the overwhelming successful Webly supervised learning, we make much use readily available images from the Web with their weakly annotated texts for an effective representation. On the other hand, we design a knowledge-embedded model, including the dynamically updated interaction mechanism between semantic representation models and KGs. Experimental results on two benchmark datasets demonstrate that our proposed model significantly achieves the most outstanding performance compared with other state-of-the-art approaches for the task of visual reasoning.
Wenbo Zheng 0001, Lan Yan, Fei-Yue Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 So Many Heads, So Many Wits: Multimodal Graph Reasoning for Text-Based Visual Question Answering
abstract
While texts related to images convey fundamental messages for scene understanding and reasoning, text-based visual question answering tasks concentrate on visual questions that require reading texts from images. However, most current methods add multimodal features that are independently extracted from a given image into a reasoning model without considering their inter- and intra-relationships according to three modalities (i.e., scene texts, questions, and images). To this end, we propose a novel text-based visual question answering model, multimodal graph reasoning. Our model first extracts intramodality relationships by taking the representations from identical modalities as semantic graphs. Then, we present graph multihead self-attention, which boosts each graph representation through graph-by-graph aggregation to capture the intermodality relationship. It is a case of “so many heads, so many wits” in the sense that as more semantic graphs are involved in this process, each graph representation becomes more effective. Finally, these representations are reprojected, and we perform answer prediction with their outputs. The experimental results demonstrate that our approach realizes substantially better performance compared with other state-of-the-art models.
Wenbo Zheng 0001, Lan Yan, Fei-Yue Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2023 WebCPM: Interactive Web Search for Chinese Long-form Question Answering
abstract
Yujia Qin, Zihan Cai, Dian Jin, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin, Xu Han, Ning Ding, Huadong Wang, Ruobing Xie, Fanchao Qi, Zhiyuan Liu, Maosong Sun, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yujia Qin, Zihan Cai, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin 0001, Xu Han 0007, Ning Ding 0002, Ruobing Xie, Fanchao Qi, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0024
ACL (1)4
2023 Balanced Supervised Contrastive Learning for Skin Lesion Classification
abstract
Deep neural networks have emerged as an important tool for computer-aided diagnosis. However, deep models for skin lesion classification still face the challenges of intra-class variation and inter-class similarity, as well as data imbalance. To address these challenges, in this paper, we propose a balanced supervised contrastive learning (BSCL) approach for the skin lesion classification task. Our model consists of two branches for supervised contrastive learning and classification, respectively. The introduced supervised contrastive learning branch helps the network to learn more discriminative representations. Moreover, we design both a category -averaging strategy which averages the instances of every class in a mini-batch, and a category-complement strategy which makes all categories to appear in each mini-batch, to balance the influence from different skin lesion categories. Besides, we introduce a multi-weighted classification loss to learn a balanced classifier. Extensive experiments on two benchmarks demonstrate that our approach is able to learn strong feature representations and achieve state-of-the-art skin lesion classification performance.
Lan Yan, Kenli Li 0001
SMC1
2023 Where Characteristics Effect Night-to-Day Translation Performance?
abstract
Inspired by the huge success of generative adversarial networks (GANs), GAN-based night-to-day translation methods have achieved excellent results. However, these methods have not been well visualized and understood, and thus cannot find the characteristics about their night-to-day translation performance. To this end, we present a simple clustering-based parsing approach to effectively understand the internal representations of the GAN-based night-to-day translator. In particular, we first cluster the internal representations of specific layers of the translator into a number of classes. Then, according to the proposed selection strategy, the class that has the most significant impact on the translation performance can be identified. Through experiments on three publicly available datasets, we find the answer of the question (title) is that the characteristics at the junction of bright and dark regions affect the performance of night-to-day translation.
Lan Yan, Wenbo Zheng 0001, Kenli Li 0001
SMC1
2023 Two Birds With One Stone: Knowledge-Embedded Temporal Convolutional Transformer for Depression Detection and Emotion Recognition
abstract
Depression is a critical problem in modern society that affects an estimated 350 million people worldwide, causing feelings of sadness and a lack of interest and pleasure. Emotional disorders are gaining interest and are closely entwined with depression, because one contributes to an understanding of the other. Despite the achievements in the two separate tasks of emotion recognition and depression detection, there has not been much prior effort to build a unified model that can connect these two tasks with different modalities, including multimedia (text, audio, and video) and unobtrusive physiological signals (e.g., electroencephalography). We propose a noveltemporal convolutional transformer with knowledge embeddingto address the joint task of depression detection and emotion recognition. This approach not only learns multimodal embeddings across domains via the temporal convolutional transformer but also exploits special-domain knowledge from medical knowledge graphs to improve the performance of detection and recognition. It is essential that the features learned by our method can be perceived as a priori and are suitable for increasing the performance of other related tasks. Our method illustrates the case of “two birds with one stone” in the sense that two or more tasks can be efficiently handled with our unique model, which captures effective features. Experimental results on ten real-world datasets show that the proposed approach significantly outperforms other state-of-the-art approaches. On the other hand, experiments in which our methodology is applied to other reasoning tasks show that our approach effectively supports model reasoning related to emotion and improves its performance.
Wenbo Zheng 0001, Lan Yan, Fei-Yue Wang 0001
IEEE Trans. Affect. Comput.2
2023 D→K→I: Data-Knowledge-Driven Group Intelligence Framework for Smart Service in Education Metaverse
abstract
Metaverse is the fusion of cyber–physical–social intelligence, and the fusion becomes the core and fundamental property of the metaverse. As an important part of social operationalization, the education domain leads to the birth of the education metaverse. This article answers three basic questions about smart services in the education metaverse: 1) learning scene; 2) technical framework; and 3) initial expansion. Specifically, four key elements constitute the learning scene in the education metaverse: 1) the learner; 2) its time; 3) space; and 4) learning event. In this learning scene, we propose a novel data-knowledge-driven group intelligence framework, aiming to transform data in the education metaverse into knowledge, and intersect and integrate intelligence with knowledge; based on this framework, we apply it to specific services, i.e., transaction and management services. We hope that our work opens the door to research on smart services in the education metaverse and more scholars will work for these challenges.
Wenbo Zheng 0001, Lan Yan, Liwei Ouyang, Ding Wen
IEEE Trans. Syst. Man Cybern. Syst.2
2022 IPGAN: Identity-Preservation Generative Adversarial Network for unsupervised photo-to-caricature translation
Lan Yan, Wenbo Zheng 0001, Chao Gou, Fei-Yue Wang 0001
Knowl. Based Syst.1
2022 An ACP-Based Parallel Approach for Color Image Encryption Using Redundant Blocks
abstract
Public concerns on image encryption grow significantly as the development and application of edge computing and the Internet of Things intensified recently. However, most existing image cryptosystems are not sophisticated enough to resist the two major attack strategies available currently, that is: 1) differential attacks and 2) chosen-plaintext attacks, which are famous for their destructive power, especially their capability of exploiting cryptosystems' features to recover the secret key. In this article, we propose an artificial image, computational experiment, and parallel execution (ACP)-based color image encryption approach using redundant blocks. First, a redundant blocks strategy with redundant spaces is proposed to prevent differential attacks and accelerate operating speed while guaranteeing the security of image cryptosystems. Second, real-world chaotic data (e.g., stock data) are obtained to generate artificial images and conduct computational experiments. Furthermore, artificial images are encrypted via real-world chaos, while the original images are encrypted via simulated chaos (such as Chen's hyperchaos). Finally, we design the process of parallel execution for image encryption and use DNA XOR to merge two groups of encrypted subimages to fuse the effect of the chaotic characteristics in both the real world and the simulation. The final encrypted image is realized through the recovery of redundant blocks. The ACP mechanism of color image encryption achieves the goal of improving the sophistication of chaos-based cryptosystems and resists both the differential and chosen-plaintext attacks. Experimental results and security analysis show that our approach provides not only excellent encryption but also security sufficient to prevent known attacks.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Fei-Yue Wang 0001
IEEE Trans. Cybern.2
2022 Fuzzy Deep Forest With Deep Contours Feature for Leaf Cultivar Classification
abstract
Deep learning is a compelling technique for feature extraction due to its adaptive capacity of processing and providing deeper image information. However, for the task of leaf cultivar classification, the deep learning-based classifier model is unable to extract contour features of leaf images deeply due to the lack of large specialized datasets and expert knowledge annotations. Also, the scale/size of the current leaf cultivar dataset does not meet the needs of deep neural networks (DNNs). In particular, the high model complexity of DNNs implies that deep-learning-based neural networks seem to must require a large dataset to achieve good performance, but facing the fact that the leaf cultivar dataset often is small, even some classes in this kind of datasets contain less than ten images/examples. To overcome these problems and inspired by the resounding success of fuzzy logic, we propose a novel fuzzy ensemble model for leaf cultivar classification. To extract the contours of leaves, we first propose generative adversarial networks-based methods. Second, to improve the ability of feature representation, we present a data augmentation method to transform our contour features. Third, to get the essential features of leaves, we design a novel generation of thefuzzy random forest. Finally, to achieve accurate classification, we design a novel deep learning strategy, namelydeep fuzzy representation learning, integrating and cascading a lot of our fuzzy random forests. Experimental results show that our model outperforms other existing state-of-the-arts on three real-world datasets, and performs much better than the original deep forest and DNN-based algorithms particularly.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Fei-Yue Wang 0001
IEEE Trans. Fuzzy Syst.2
2022 A Novel Vehicle Detection Framework Based on Parallel Vision
abstract
Autonomous driving has become a prevalent research topic in recent years, arousing the attention of many academic universities and commercial companies. As human drivers rely on visual information to discern road conditions and make driving decisions, autonomous driving calls for vision systems such as vehicle detection models. These vision models require a large amount of labeled data while collecting and annotating the real traffic data are time‐consuming and costly. Therefore, we present a novel vehicle detection framework based on the parallel vision to tackle the above issue, using the specially designed virtual data to help train the vehicle detection model. We also propose a method to construct large‐scale artificial scenes and generate the virtual data for the vision‐based autonomous driving schemes. Experimental results verify the effectiveness of our proposed framework, demonstrating that the combination of virtual and real data has better performance for training the vehicle detection model than the only use of real data.
Ying Zhuo, Lan Yan, Wenbo Zheng 0001, Yutian Zhang, Chao Gou
Wirel. Commun. Mob. Comput.2
2021 Two Heads are Better Than One: Hypergraph-Enhanced Graph Reasoning for Visual Event Ratiocination
abstract
Even with a still image, humans can ratiocinate various visual cause-and-effect descriptions before, at present, and after, as well as beyond the given image. However, it is challenging for models to achieve such task–the visual event ratiocination, owing to the limitations of time and space. To this end, we propose a novel multi-modal model, Hypergraph-Enhanced Graph Reasoning. First it represents the contents from the same modality as a semantic graph and mines the intra-modality relationship, therefore breaking the limitations in the spatial domain. Then, we introduce the Graph Self-Attention Enhancement. On the one hand, this enables semantic graph representations from different modalities to enhance each other and captures the inter-modality relationship along the line. On the other hand, it utilizes our built multi-modal hypergraphs in different moments to boost individual semantic graph representations, and breaks the limitations in the temporal domain. Our method illustrates the case of "two heads are better than one" in the sense that semantic graph representations with the help of the proposed enhancement mechanism are more robust than those without. Finally, we re-project these representations and leverage their outcomes to generate textual cause-and-effect descriptions. Experimental results show that our model achieves significantly higher performance in comparison with other state-of-the-arts.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Fei-Yue Wang 0001
ICML2
2021 Knowledge is Power: Hierarchical-Knowledge Embedded Meta-Learning for Visual Reasoning in Artistic Domains
abstract
This paper deals with the challenging problem of building visual reasoning models for answering questions related to artworks in artistic domains. The nature of abstract styles and cultural contexts within an artistic image makes the corresponding learning tasks extremely difficult. We propose a novel framework termed as Hierarchical-Knowledge Embedded Meta-Learning to address the critical issues of visual reasoning in artistic domains. In particular, we firstly present a deep relational model to capture and memorize the relations among different samples. Then, we provide the hierarchical-knowledge embedding that mines the implicit relationship between question-answer pairs for knowledge representation as the guidance of our meta-learner. This is a case of "knowledge is power" in the sense that the hierarchical knowledge representation is incorporated into our meta-learning based model. The final classification is derived from our model by learning to compare the features of samples. Experimental results show that our approach achieves significantly higher performance compared with other state-of-the-arts.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Fei-Yue Wang 0001
KDD2
2021 Weakly Supervised Sketch Based Person Search
abstract
Person search often requires a query photo of the target person. However, in many practical scenarios, there is no guarantee that such a photo is always available. In this paper, we define the problem of sketch based person search, which uses a sketch instead of a photo as the probe for retrieving. We tackle this problem in a weak supervision setting and propose a clustering and feature attention based weakly supervised learning framework, which contains two stages of pedestrian detection and sketch based person re-identification. Specially, we introduce multiple detectors, followed by fuzzy c-means clustering to achieve weakly supervised pedestrian detection. Moreover, we design an attention module to learn discriminative features in subsequent re-identification network. Extensive experiments show the superiority of our method.
Lan Yan, Wenbo Zheng 0001, Fei-Yue Wang 0001, Chao Gou
ICMR1
2021 Learning from the Negativity: Deep Negative Correlation Meta-Learning for Adversarial Image Classification
Wenbo Zheng 0001, Lan Yan, Fei-Yue Wang 0001, Chao Gou
MMM (1)2
2021 Fighting fire with fire: A spatial-frequency ensemble relation network with generative adversarial learning for adversarial image classification
abstract
Adversarial images generated by generative adversarial networks are not close to any existing benign images, and contain nonrobust features that have been identified as critical to the robustness of a machine learning model. Since adversarial images have an underlying distribution that differs from normal images, these kinds of images can offer valuable features for training a robust model. To deal with these special features, we focus on a novel machine learning task of adversarial images classification, where adversarial images can be used to investigate the problem of classifying adversarial images themselves. In the setting of this novel task, adversarial images are the ONLY kind of data used in training and testing, rather than not just a set of testing images as usual. To this end, we propose a novel spatial–frequency ensemble relation network with generative adversarial learning. First, we present a spatial–frequency ensemble representation learning to extract the feature of training images. Second, we design a meta-learning-based relation model to gain the relationship between images. Third, to achieve a robust model, we utilize generative adversarial learning and transform the relationship into a Jacobian matrix. Finally, we design a discriminator model that determines whether an adversarial image is from the matching category or not. Experimental results demonstrate that our approach achieves significantly higher performance compared with other state-of-the-arts.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Fei-Yue Wang 0001
Int. J. Intell. Syst.2
2021 Learning to learn by yourself: Unsupervised meta-learning with self-knowledge distillation for COVID-19 diagnosis from pneumonia cases
abstract
The goal of diagnosing the coronavirus disease 2019 (COVID-19) from suspected pneumonia cases, that is, recognizing COVID-19 from chest X-ray or computed tomography (CT) images, is to improve diagnostic accuracy, leading to faster intervention. The most important and challenging problem here is to design an effective and robust diagnosis model. To this end, there are three challenges to overcome: (1) The lack of training samples limits the success of existing deep-learning-based methods. (2) Many public COVID-19 data sets contain only a few images without fine-grained labels. (3) Due to the explosive growth of suspected cases, it is urgent and important to diagnose not only COVID-19 cases but also the cases of other types of pneumonia that are similar to the symptoms of COVID-19. To address these issues, we propose a novel framework called Unsupervised Meta-Learning with Self-Knowledge Distillation to address the problem of differentiating COVID-19 from pneumonia cases. During training, our model cannot use any true labels and aims to gain the ability of learning to learn by itself. In particular, we first present a deep diagnosis model based on a relation network to capture and memorize the relation among different images. Second, to enhance the performance of our model, we design a self-knowledge distillation mechanism that distills knowledge within our model itself. Our network is divided into several parts, and the knowledge in the deeper parts is squeezed into the shallow ones. The final results are derived from our model by learning to compare the features of images. Experimental results demonstrate that our approach achieves significantly higher performance than other state-of-the-art methods. Moreover, we construct a new COVID-19 pneumonia data set based on text mining, consisting of 2696 COVID-19 images (347 X-ray + 2349 CT), 10,155 images (9661 X-ray + 494 CT) about other types of pneumonia, and the fine-grained labels of all. Our data set considers not only a bacterial infection or viral infection which causes pneumonia but also a viral infection derived from the influenza virus or coronavirus.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Zhicheng Zhang 0004, Jun Jason Zhang, Fei-Yue Wang 0001
Int. J. Intell. Syst.2
2021 A loss-balanced multi-task model for simultaneous detection and segmentation
Kunfeng Wang, Yutong Wang 0001, Lan Yan, Fei-Yue Wang 0001
Neurocomputing4
2021 IsGAN: Identity-sensitive generative adversarial network for face photo-sketch synthesis
Lan Yan, Wenbo Zheng 0001, Chao Gou, Fei-Yue Wang 0001
Pattern Recognit.1
2021 Joint image-to-image translation with denoising using enhanced generative adversarial networks
Lan Yan, Wenbo Zheng 0001, Fei-Yue Wang 0001, Chao Gou
Signal Process. Image Commun.1
2020 Webly Supervised Knowledge Embedding Model for Visual Reasoning
abstract
Visual reasoning between visual image and natural language description is a long-standing challenge in computer vision. While recent approaches offer a great promise by compositionality or relational computing, most of them are oppressed by the challenge of training with datasets containing only a limited number of images with ground-truth texts. Besides, it is extremely time-consuming and difficult to build a larger dataset by annotating millions of images with text descriptions that may very likely lead to a biased model. Inspired by the majority success of webly supervised learning, we utilize readily-available web images with its noisy annotations for learning a robust representation. Our key idea is to presume on web images and corresponding tags along with fully annotated datasets in learning with knowledge embedding. We present a two-stage approach for the task that can augment knowledge through an effective embedding model with weakly supervised web data. This approach learns not only knowledge-based embeddings derived from key-value memory networks to make joint and full use of textual and visual information but also exploits the knowledge to improve the performance with knowledge-based representation learning for applying other general reasoning tasks. Experimental results on two benchmarks show that the proposed approach significantly improves performance compared with the state-of-the-art methods and guarantees the robustness of our model against visual reasoning tasks and other reasoning tasks.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Fei-Yue Wang 0001
CVPR2
2020 Weakly Supervised Person Search
abstract
While existing person search methods have achieved good performance, they require the images used for training contain labels about the identity and bounding box location of each person. However, it is expensive and difficult to manually annotate these labels in the large scale scenario. To overcome this issue, we consider weakly supervised person search. The weakly supervised setting means during training we only know which identities appear in the image set and how many individuals present in each image, without any identity or location information on the image. Facing this challenge, we propose a clustering and patch based weakly supervised learning (CPBWSL) framework, which separately addresses two sub-tasks including pedestrian detection and person re-identification. Particularly, we introduce multiple detectors to provide more detection results as well as fuzzy c-means clustering algorithm to cluster these results and remove low membership ones. Moreover, a patch based learning network is designed to generate different patches and learn discriminative patch features. Extensive experiments on two benchmarks indicate that the proposed weakly supervised setting is feasible and our method can achieve performance comparable to some fully supervised person search methods.
Lan Yan, Wenbo Zheng 0001, Fei-Yue Wang 0001, Chao Gou
DSAA1
2020 JND-GAN: Human-Vision-Systems Inspired Generative Adversarial Networks for Image-to-Image Translation
abstract
Image-to-image translation aims to learn the mapping between two visual domains. At the beginning of designing the existing image-to-image translation method, it was not considered whether the generated image is realistic or not. In this work, we present a novel approach to address the problem of generating fidelity in the area of image-to-image translation. In particular, humans judge whether an image is realistic or not with unique human vision's feeling rather than paying attention to the real-world semantics. Inspired by this, we propose an effective network loss to capture the pixel-level representations and human vision system information for verisimilar image-to-image translation. To enforce both structural and translation-model consistency during adaptation, we propose a novel Just-Noticeable-Difference loss based on a visual recognition task. The Just-Noticeable-Difference loss not only guides the overall representation to be discriminative, but also enforces our cycle loss before and after mapping between domains. Qualitative results show that our model can generate realistic images on a wide range of tasks without paired training data. For quantitative comparisons, we measure realism with user study and diversity with a perceptual distance metric. We apply the proposed model to domain adaptation and show competitive performance when compared to the state-of-the-art on many datasets.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Fei-Yue Wang 0001
ECAI2
2020 Feature Aggregation Attention Network for Single Image Dehazing
abstract
Due to its ill-posed nature, single image dehazing is a challenging problem. In this paper, we propose an end-to-end feature aggregation attention network (FAAN) for single image dehazing. It incorporates the idea of attention mechanism and residual learning and can adaptively aggregate different level features. In particular, in the proposed FANN, we design a novel block structure consisting of feature attention module, smoothed dilated convolution and local residual learning. The local residual learning allows the less useful information to be bypassed through multiple skip connections. The feature attention module is designed to assign more weight to important features. The smoothed dilated convolution is adopted to enlarge the receptive field without the negative influence of gridding artifacts. The experiments on the RESIDE dataset show that the proposed approach acquires state-of-the-art performance in both qualitative and quantitative measures.
Lan Yan, Wenbo Zheng 0001, Chao Gou, Fei-Yue Wang 0001
ICIP1
2020 Graph Attention Model Embedded With Multi-Modal Knowledge For Depression Detection
abstract
With more than 300 million people depressed worldwide annually, depression is a global problem. The goal of depression detection is to improve diagnostic accuracy and availability, leading to faster intervention. The most important and challenging problem here is to design an effective and robust depression detection model. To this end, there are two challenges to overcome: 1) Multi-modal (audio, image, text, etc.) information must be jointly considered to make accurate inferences. 2) Existing deep learning-based work suffers from multi-modal data sufficiency problem. To address these issues, we propose a graph attention model embedded with multi-modal knowledge for depression detection. This approach learns not only reasonable embeddings for nodes in the knowledge graph, but also exploits medical knowledge to improve the performance of classification and prediction with the knowledge attention mechanism. Experimental results on the real-world datasets show that the proposed approach significantly improves the classification and prediction performance compared with other major state-of-the-art approaches, with guaranteed the robustness with each modality of multi-modal data. Overall, this paper shows how multi-modal knowledge attention mechanism and deep-learning-based networks can be combined to assist mental health patients and practitioners.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Fei-Yue Wang 0001
ICME2
2020 Learning from the Guidance: Knowledge Embedded Meta-learning for Medical Visual Question Answering
Wenbo Zheng 0001, Lan Yan, Fei-Yue Wang 0001, Chao Gou
ICONIP (4)2
2020 Federated Meta-Learning for Fraudulent Credit Card Detection
abstract
Credit card transaction fraud costs billions of dollars to card issuers every year. Besides, the credit card transaction dataset is very skewed, there are much fewer samples of frauds than legitimate transactions. Due to the data security and privacy, different banks are usually not allowed to share their transaction datasets. These problems make traditional model difficult to learn the patterns of frauds and also difficult to detect them. In this paper, we introduce a novel framework termed as federated meta-learning for fraud detection. Different from the traditional technologies trained with data centralized in the cloud, our model enables banks to learn fraud detection model with the training data distributed on their own local database. A shared whole model is constructed by aggregating locallycomputed updates of fraud detection model. Banks can collectively reap the benefits of shared model without sharing the dataset and protect the sensitive information of cardholders. To achieve the good performance of classification, we further formulate an improved triplet-like metric learning, and design a novel meta-learning-based classifier, which allows joint comparison with K negative samples in each mini-batch. Experimental results demonstrate that the proposed approach achieves significantly higher performance compared with the other state-of-the-art approaches.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Fei-Yue Wang 0001
IJCAI2
2020 Learning from the Past: Meta-Continual Learning with Knowledge Embedding for Jointly Sketch, Cartoon, and Caricature Face Recognition
abstract
This paper deals with a challenging task of learning from different modalities by tackling the difficulty problem of jointly face recognition between abstract-like sketches, cartoons, caricatures and real-life photographs. Due to the significant variations in the abstract faces, building vision models for recognizing data from these modalities is an extremely challenging. We propose a novel framework termed as Meta-Continual Learning with Knowledge Embedding to address the task of jointly sketch, cartoon, and caricature face recognition. In particular, we firstly present a deep relational network to capture and memorize the relation among different samples. Secondly, we present the construction of our knowledge graph that relates image with the label as the guidance of our meta-learner. We then design a knowledge embedding mechanism to incorporate the knowledge representation into our network. Thirdly, to mitigate catastrophic forgetting, we use a meta-continual model that updates our ensemble model and improves its prediction accuracy. With this meta-continual model, our network can learn from its past. The final classification is derived from our network by learning to compare the features of samples. Experimental results demonstrate that our approach achieves significantly higher performance compared with other state-of-the-art approaches.
Wenbo Zheng 0001, Lan Yan, Fei-Yue Wang 0001, Chao Gou
ACM Multimedia2
2020 Learning to Classify: A Flow-Based Relation Network for Encrypted Traffic Classification
abstract
As the size and source of network traffic increase, so does the challenge of monitoring and analyzing network traffic. The challenging problems of classifying encrypted traffic are the imbalanced property of network data, the generalization on an unseen dataset, and overly dependent on data size. In this paper, we propose an application of a meta-learning approach to address these problems in encrypted traffic classification, named Flow-Based Relation Network (RBRN). The RBRN is an end-to-end classification model that learns representative features from the raw flows and then classifies them in a unified framework. Moreover, we design “hallucinator” to produce additional training samples for the imbalanced classification, and then focus on meta-learning to classify unseen categories from few labeled samples. We validate the effectiveness of the RBRN on the real-world network traffic dataset, and the experimental results demonstrate that the RBRN can achieve an excellent classification performance and outperform the state-of-the-art methods on encrypted traffic classification. What is more interesting, our model trained on the real-world dataset can generalize very well to unseen datasets, outperforming multiple state-of-art methods.
Wenbo Zheng 0001, Chao Gou, Lan Yan, Shaocong Mo
WWW3
2019 Forest Representation Learning with Multiscale Contour Feature Learning for Leaf Cultivar Classification
abstract
Automated plant species identification system could help botanists and layman in identifying plant species rapidly. Deep learning is robust for feature extraction as it is superior in providing deeper information of images. However, deep neural networks based leaf cultivar classification model can not deeply mine the contour features of leaf images. Besides, the deep neural networks are very data-hungry due to the large complexity of the models, which means that the model's performance can decrease significantly when the size of the training data decreases. In this paper, we propose the leaf cultivar classification model based on forest representation learning and multiscale contour feature learning. Firstly, we build the multiscale contour feature learning to extract the contour features of leaf images. Then, the original input features are transformed by the data augmentation method to enhance the ability of feature expression, and the autoencoder by forest performs the data of dimensionality reduction on the features. Finally, we use the improved deep forest algorithm and autoencoder by forest to build forest representation learning for classification. The experimental results show that the proposed algorithm has higher performance than the original deep forest (gcForest) algorithm and other existing start-of-art algorithms, and has higher performance and efficiency than other deep learning algorithms.
Wenbo Zheng 0001, Chao Gou, Lan Yan
BIBM3
2019 Guided Cyclegan Via Semi-Dual Optimal Transport for Photo-Realistic Face Super-Resolution
abstract
Face super-resolution has been studied for decades, and many approaches have been proposed to upsample low-resolution face images using information mined from paired low-resolution (LR) images and high-resolution (HR) images. However, most of this kind of works only simply sharpen the blurry edges in the upsampled face images and typically no photo-realistic face is reconstructed in the final result. In this paper, we present a GAN-based algorithm for face super-resolution which properly synthesizes photo-realistic super-recovered face. To this end, we introduce semi-dual optimal transport to optimize our model such that the distribution of its generated data can match the distribution of a target domain as much as possible. This way would endow our model with learning the mapping of distribution from unpaired LR images and HR images with desired properties. We demonstrate the robustness of our algorithm by testing it on Color FERET database and show that its performance is considerably superior to all state-of-the-art approaches.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Fei-Yue Wang 0001
ICIP2
2019 A Relation Network Embedded with Prior Features for Few-Shot Caricature Recognition
abstract
Caricature is a simple and abstract description of a person using her/his exaggerated characteristics. Due to amplified facial variations in the caricatures and significant differences among caricature and real face modalities, building vision models for recognizing each other between these modalities is an extremely challenging task. In addition, it is not easy to collect abundant samples of real faces and corresponding caricatures for training vision models, which makes the recognition more difficult. In this paper, we propose a novel relation network via meta learning to address the problem of few-shot caricature face recognition. In particular, we present a deep relation network to capture and memorize the relation among different samples. To employ the prior knowledge, we combine learned deep and handcrafted features to form the hybrid-prior representation via joint meta learning. Final recognition is derived from our relation network by learning to compare between the hybrid-prior features of samples. Experimental results on three caricature datasets of WebCaricature, IIIT-CFW, and Caricature-207 demonstrate that our method performs better than many existing ones for few-shot caricature recognition.
Wenbo Zheng 0001, Lan Yan, Chao Gou, Fei-Yue Wang 0001
ICME2
2019 The ParallelEye Dataset: A Large Collection of Virtual Images for Traffic Vision Research
abstract
Dataset plays an essential role in the training and testing of traffic vision algorithms. However, the collection and annotation of images from the real world is time-consuming, labor-intensive, and error-prone. Therefore, more and more researchers have begun to explore the virtual dataset, to overcome the disadvantages of real datasets. In this paper, we propose a systematic method to construct large-scale artificial scenes and collect a new virtual dataset (named “ParallelEye”) for the traffic vision research. The Unity3D rendering software is used to simulate environmental changes in the artificial scenes and generate ground-truth labels automatically, including semantic/instance segmentation, object bounding boxes, and so on. In addition, we utilize ParallelEye in combination with real datasets to conduct experiments. The experimental results show the inclusion of virtual data helps to enhance the per-class accuracy in object detection and semantic segmentation. Meanwhile, it is also illustrated that the virtual data with controllable imaging conditions can be used to design evaluation experiments flexibly.
Xuan Li 0006, Kunfeng Wang, Yonglin Tian, Lan Yan, Fang Deng, Fei-Yue Wang 0001
IEEE Trans. Intell. Transp. Syst.4
2018 The ParallelEye-CS Dataset: Constructing Artificial Scenes for Evaluating the Visual Intelligence of Intelligent Vehicles
abstract
Offline training and testing are playing an essential role in design and evaluation of intelligent vehicle vision algorithms. Nevertheless, long-term inconvenience concerning traditional image datasets is that manually collecting and annotating datasets from real scenes lack testing tasks and diverse environmental conditions. For that virtual datasets can make up for these regrets. In this paper, we propose to construct artificial scenes for evaluating the visual intelligence of intelligent vehicles and generate a new virtual dataset called “ParallelEye-CS”. First of all, the actual track map data is used to build 3D scene model of Chinese Flagship Intelligent Vehicle Proving Center Area, Changshu. Then, the computer graphics and virtual reality technologies are utilized to simulate the virtual testing tasks according to the Chinese Intelligent Vehicles Future Challenge (IVFC) tasks. Furthermore, the Unity3D platform is used to generate accurate ground-truth labels and change environmental conditions. As a result, we present a viable implementation method for constructing artificial scenes for traffic vision research. The experimental results show that our method is able to generate photorealistic virtual datasets with diverse testing tasks.
Xuan Li 0006, Yutong Wang 0001, Kunfeng Wang, Lan Yan, Fei-Yue Wang 0001
Intelligent Vehicles Symposium4