Shuhuan Mei

dblp:207/1938 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Image recognition and object detection · 43% Learning paradigms · 14% Representation and self-supervised learning · 14%
Computer graphics and multimedia
4 papers
Multimedia analysis and retrieval · 100%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 77% Data mining · 23%

Topics — the 12 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.922021
Plant Disease Recognition: A Large-Scale Benchmark Dataset and a Visual Region and Loss Reweighting Approach · IEEE Trans. Image Process. 2021
Multi-Task Deep Relative Attribute Learning for Visual Urban Perception · IEEE Trans. Image Process. 2020
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.412020
A Two-Stage Triplet Network Training Framework for Image Retrieval · IEEE Trans. Multim. 2020
Machine learning › Learning paradigms
multi-task learning
0.412020
Multi-Task Deep Relative Attribute Learning for Visual Urban Perception · IEEE Trans. Image Process. 2020
Computer vision › Image recognition and object detection › attribute recognition
relative attribute learning
0.412020
Multi-Task Deep Relative Attribute Learning for Visual Urban Perception · IEEE Trans. Image Process. 2020
Machine learning › Deep learning architectures and training
triplet network
0.412020
A Two-Stage Triplet Network Training Framework for Image Retrieval · IEEE Trans. Multim. 2020
Multimedia analysis and retrieval
image retrieval
0.412020
A Two-Stage Triplet Network Training Framework for Image Retrieval · IEEE Trans. Multim. 2020
Multimedia analysis and retrieval › image retrieval
instance-level image retrieval
0.412020
A Two-Stage Triplet Network Training Framework for Image Retrieval · IEEE Trans. Multim. 2020
Multimedia analysis and retrieval › social media analysis
venue category estimation
0.412019
Hierarchy-Dependent Cross-Platform Multi-View Feature Learning for Venue Category Prediction · IEEE Trans. Multim. 2019
Multimedia analysis and retrieval
food computing
0.312018
You Are What You Eat: Exploring Rich Recipe Information for Cross-Region Food Analysis · IEEE Trans. Multim. 2018
Multimedia analysis and retrieval
topic modeling
0.312018
You Are What You Eat: Exploring Rich Recipe Information for Cross-Region Food Analysis · IEEE Trans. Multim. 2018
Recommender systems › domain-specific recommendation
recipe recommendation
0.312017
A Delicious Recipe Analysis Framework for Exploring Multi-Modal Recipes with Various Attributes · ACM Multimedia 2017
Data mining › text mining
topic modeling
0.112017
A Delicious Recipe Analysis Framework for Exploring Multi-Modal Recipes with Various Attributes · ACM Multimedia 2017

Methods — techniques the papers use, named apart from their topics

weakly supervised learning · 1.0patch feature clustering · 1.0loss reweighting · 1.0LSTM · 1.0regional generalized-mean pooling · 0.9ranking loss · 0.9deep learning · 0.8multimodal embedding · 0.6deep visual features · 0.6structured sparsity · 0.4siamese network · 0.4multi-task learning · 0.4multi-view fusion · 0.4hierarchical structure embedding · 0.4probabilistic topic model · 0.3manifold ranking · 0.3
YearPublicationVenuePosition
2021 Plant Disease Recognition: A Large-Scale Benchmark Dataset and a Visual Region and Loss Reweighting Approach
abstract
Plant disease diagnosis is very critical for agriculture due to its importance for increasing crop production. Recent advances in image processing offer us a new way to solve this issue via visual plant disease analysis. However, there are few works in this area, not to mention systematic researches. In this paper, we systematically investigate the problem of visual plant disease recognition for plant disease diagnosis. Compared with other types of images, plant disease images generally exhibit randomly distributed lesions, diverse symptoms and complex backgrounds, and thus are hard to capture discriminative information. To facilitate the plant disease recognition research, we construct a new large-scale plant disease dataset with 271 plant disease categories and 220,592 images. Based on this dataset, we tackle plant disease recognition via reweighting both visual regions and loss to emphasize diseased parts. We first compute the weights of all the divided patches from each image based on the cluster distribution of these patches to indicate the discriminative level of each patch. Then we allocate the weight to each loss for each patch-label pair during weakly-supervised training to enable discriminative disease part learning. We finally extract patch features from the network trained with loss reweighting, and utilize the LSTM network to encode the weighed patch feature sequence into a comprehensive feature representation. Extensive evaluations on this dataset and another public dataset demonstrate the advantage of the proposed method. We expect this research will further the agenda of plant disease recognition in the community of image processing.
Xinda Liu, Weiqing Min, Shuhuan Mei, Lili Wang 0006, Shuqiang Jiang
IEEE Trans. Image Process.3
2020 Multi-Task Deep Relative Attribute Learning for Visual Urban Perception
abstract
Visual urban perception aims to quantify perceptual attributes (e.g., safe and depressing attributes) of physical urban environment from crowd-sourced street-view images and their pairwise comparisons. It has been receiving more and more attention in computer vision for various applications, such as perceptive attribute learning and urban scene understanding. Most existing methods adopt either (i) a regression model trained using image features and ranked scores converted from pairwise comparisons for perceptual attribute prediction or (ii) a pairwise ranking algorithm to independently learn each perceptual attribute. However, the former fails to directly exploit pairwise comparisons while the latter ignores the relationship among different attributes. To address them, we propose a Multi-Task Deep Relative Attribute Learning Network (MTDRALN) to learn all the relative attributes simultaneously via multi-task Siamese networks, where each Siamese network will predict one relative attribute. Combined with deep relative attribute learning, we utilize the structured sparsity to exploit the prior from natural attribute grouping, where all the attributes are divided into different groups based on semantic relatedness in advance. As a result, MTDRALN is capable of learning all the perceptual attributes simultaneously via multi-task learning. Besides the ranking sub-network, MTDRALN further introduces the classification sub-network, and these two types of losses from two sub-networks jointly constrain parameters of the deep network to make the network learn more discriminative visual features for relative attribute learning. In addition, our network can be trained in an end-to-end way to make deep feature learning and multi-task relative attribute learning reinforce each other. Extensive experiments on the large-scale Place Pulse 2.0 dataset validate the advantage of our proposed network. Our qualitative results along with visualization of saliency maps also show that the proposed network is able to learn effective features for perceptual attributes.
Weiqing Min, Shuhuan Mei, Linhu Liu, Shuqiang Jiang
IEEE Trans. Image Process.2
2020 A Two-Stage Triplet Network Training Framework for Image Retrieval
abstract
In this paper, we propose a novel framework for instance-level image retrieval. Recent methods focus on fine-tuning the Convolutional Neural Network (CNN) via a Siamese architecture to improve off-the-shelf CNN features. They generally use the ranking loss to train such networks, and do not take full use of supervised information for better network training, especially with more complex neural architectures. To solve this, we propose a two-stage triplet network training framework, which mainly consists of two stages. First, we propose a Double-Loss Regularized Triplet Network (DLRTN), which extends basic triplet network by attaching the classification sub-network, and is trained via simultaneously optimizing two different types of loss functions. Double-loss functions of DLRTN aim at specific retrieval task and can jointly boost the discriminative capability of DLRTN from different aspects via supervised learning. Second, considering feature maps of the last convolution layer extracted from DLRTN and regions detected from the region proposal network as the input, we then introduce the Regional Generalized-Mean Pooling (RGMP) layer for the triplet network, and re-train this network to learn pooling parameters. Through RGMP, we pool feature maps for each region and aggregate features of different regions from each image to Regional Generalized Activations of Convolutions (R-GAC) as final image representation. R-GAC is capable of generalizing existing Regional Maximum Activations of Convolutions (R-MAC) and is thus more robust to scale and translation. We conduct the experiment on six image retrieval datasets including standard benchmarks and recently introduced INSTRE dataset. Extensive experimental results demonstrate the effectiveness of the proposed framework.
Weiqing Min, Shuhuan Mei, Shuqiang Jiang
IEEE Trans. Multim.2
2019 Instance-level object retrieval via deep region CNN
Shuhuan Mei, Weiqing Min, Hua Duan, Shuqiang Jiang
Multim. Tools Appl.1
2019 Hierarchy-Dependent Cross-Platform Multi-View Feature Learning for Venue Category Prediction
abstract
In this paper, we focus on visual venue category prediction, which can facilitate various applications for location-based service and personalization. Considering the complementarity of different media platforms, it is reasonable to leverage venue-relevant media data from different platforms to boost the prediction performance. Intuitively, recognizing one venue category involves multiple semantic cues, especially objects and scenes and, thus, they should contribute together to venue category prediction. In addition, these venues can be organized in a natural hierarchical structure, which provides prior knowledge to guide venue category estimation. Taking these aspects into account, we propose a Hierarchy-dependent Cross-platform Multi-view Feature Learning (HCM-FL) framework for venue category prediction from videos by leveraging images from other platforms. HCM-FL includes two major components, namely Cross-Platform Transfer Deep Learning (CPTDL) and Multi-View Feature Learning with the Hierarchical Venue Structure (MVFL-HVS). CPTDL is capable of reinforcing the learned deep network from videos using images from other platforms. Specifically, CPTDL first trained a deep network using videos. These images from other platforms are filtered by the learnt network and these selected images are then fed into this learnt network to enhance it. Two kinds of pre-trained networks on the ImageNet and Places dataset are employed. Therefore, we can harness both object-oriented and scene-oriented deep features through these enhanced deep networks. MVFL-HVS is then developed to enable multi-view feature fusion. It is capable of embedding the hierarchical structure ontology to support more discriminative joint feature learning. We conduct the experiment on videos from Vine and images from Foursquare. These experimental results demonstrate the advantage of our proposed framework in jointly utilizing multi-platform data, multi-view deep features, and hierarchical venue structure knowledge.
Shuqiang Jiang, Weiqing Min, Shuhuan Mei
IEEE Trans. Multim.3
2018 You Are What You Eat: Exploring Rich Recipe Information for Cross-Region Food Analysis
abstract
Cuisine is a style of cooking and usually associated with a specific geographic region. Recipes from different cuisines shared on the web are an indicator of culinary cultures in different countries. Therefore, analysis of these recipes can lead to deep understanding of food from the cultural perspective. In this paper, we perform the first cross-region recipe analysis by jointly using the recipe ingredients, food images, and attributes such as the cuisine and course (e.g., main dish and dessert). For that solution, we propose a culinary culture analysis framework to discover the topics of ingredient bases and visualize them to enable various applications. We first propose a probabilistic topic model to discover cuisine-course specific topics. The manifold ranking method is then utilized to incorporate deep visual features to retrieve food images for topic visualization. At last, we applied the topic modeling and visualization method for three applications: 1) multimodal cuisine summarization with both recipe ingredients and images, 2) cuisine-course pattern analysis including topic-specific cuisine distribution and cuisine-specific course distribution of topics, and 3) cuisine recommendation for both cuisine-oriented and ingredient-oriented queries. Through these three applications, we can analyze the culinary cultures at both macro and micro levels. We conduct the experiment on a recipe database Yummly-66K with 66,615 recipes from 10 cuisines in Yummly. Qualitative and quantitative evaluation results have validated the effectiveness of topic modeling and visualization, and demonstrated the advantage of the framework in utilizing rich recipe information to analyze and interpret the culinary cultures from different regions.
Weiqing Min, Bing-Kun Bao, Shuhuan Mei, Yong Rui, Shuqiang Jiang
IEEE Trans. Multim.3
2017 A Delicious Recipe Analysis Framework for Exploring Multi-Modal Recipes with Various Attributes
abstract
Human beings have developed a diverse food culture. Many factors like ingredients, visual appearance, courses (e.g., breakfast and lunch), flavor and geographical regions affect our food perception and choice. In this work, we focus on multi-dimensional food analysis based on these food factors to benefit various applications like summary and recommendation. For that solution, we propose a delicious recipe analysis framework to incorporate various types of continuous and discrete attribute features and multi-modal information from recipes. First, we develop a Multi-Attribute Theme Modeling (MATM) method, which can incorporate arbitrary types of attribute features to jointly model them and the textual content. We then utilize a multi-modal embedding method to build the correlation between the learned textual theme features from MATM and visual features from the deep learning network. By learning attribute-theme relations and multi-modal correlation, we are able to fulfill different applications, including (1) flavor analysis and comparison for better understanding the flavor patterns from different dimensions, such as the region and course, (2) region-oriented multi-dimensional food summary with both multi-modal and multi-attribute information and (3) multi-attribute oriented recipe recommendation. Furthermore, our proposed framework is flexible and enables easy incorporation of arbitrary types of attributes and modalities. Qualitative and quantitative evaluation results have validated the effectiveness of the proposed method and framework on the collected Yummly dataset.
Weiqing Min, Shuqiang Jiang, Shuhui Wang, Shuhuan Mei
ACM Multimedia5