EDBT 2026 Demo / reviewers in the wild / expert
Weiqing Min
dblp:122/2626
· DBLP profile ↗
68ranked-venue papers
18as first author
38since 2021 · last 2026
0000-0001-6668-9208ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 17 first-author · 30 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 8 since 2021Computer networks · 10 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RFHNet: Relational and Frequency-Aware Hashing Network for Large-Scale Fine-Grained Food Image RetrievalabstractFine-grained food image retrieval is a key task in computational gastronomy, with applications in food traceability, dietary monitoring, and smart catering systems. Although hashing-based retrieval is attractive for large-scale search due to its storage efficiency and fast Hamming-distance computation, existing methods often perform poorly in fine-grained food scenarios, where subtle local semantics and frequency-sensitive visual cues are essential. To address this challenge, we propose RFHNet, a cascaded hierarchical hashing network that captures both global structure and fine-grained local details through multi-level representations. RFHNet includes three components: (1) Fine-grained Relation Modeling (FRM) to capture subtle visual differences among similar food components; (2) Multi-Frequency Modulated Fusion (MFMF) to extract informative multi-frequency features; and (3) Hierarchical Semantic Synergy (HSS) to adaptively integrate multi-level representations and generate discriminative hash codes. Experiments on six food-specific benchmarks show that RFHNet consistently outperforms state-of-the-art hashing methods, with mAP gains of 4.44% to 17.20% at 12 bits. These results validate the effectiveness of RFHNet for large-scale visual food retrieval and smart catering applications. The source code will be released upon publication. Weiqing Min, Guorui Sheng, Lili Wang 0009, Shuqiang Jiang |
ICMR | 2 |
| 2026 | ComAdPro: compositional learning with prototype adaptation for logo few-shot class-incremental recognition (ChinaMM 2025)
Jianxin Zhan, Wentai Chen, Sujuan Hou, Weiqing Min |
Multim. Syst. | 4 |
| 2026 | DPFA-net: a lightweight hybrid neural network with dual path feature aggregation for food image recognition
Xiangyi Zhu, Yingnan Sheng, Congrui Lv, Guorui Sheng, Weiqing Min, Shuqiang Jiang |
Multim. Syst. | 6 |
| 2026 | Large-Scale Logo DetectionabstractLogo detection is crucial for trademark compliance and media monitoring, enabling companies to monitor online trademark usage and evaluate brand visibility on social media and advertisements. The use of large datasets significantly improves accuracy and generalization, emphasizing the need for high-quality datasets to optimize performance and enhance reasoning abilities in visual detection models. This drove us to create Logo4500, an unparalleled dataset featuring 4,500 logo categories and over 293,000 meticulously labeled images. To ensure the dataset's quality, we meticulously designed the construction and annotation process, with detailed information provided in our paper. Compared to existing logo datasets, Logo4500 offers greater diversity and class imbalance, making it more reflective of real-world distribution. Leveraging this high-quality dataset, we introduce a benchmark called Frequency-Aware Learnable Dual Reweighting Network (FALDR-Net), which enhances the representation of ambiguous features and addresses class imbalance for large-scale logo detection. We conducted extensive experiments, evaluating various recent methods on this new dataset and several existing publicly available logo datasets, demonstrating its effectiveness. Additionally, we verified Logo4500's generalization ability in several tasks. We anticipate that Logo4500 and the benchmark will inspire further exploration in the logo-related research community, facilitating the advancement of visual foundation models. Sujuan Hou, Weiqing Min, Jianxin Zhan, Mengmeng Zhang 0008, Peng Li 0081, Shuqiang Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Structured-condensed prompt tuning in vision-language models for fine-grained image recognition
Xinda Liu, Weiqing Min, Guohua Geng, Shuqiang Jiang |
Pattern Recognit. | 3 |
| 2026 | LLM-informed global-local contextualization for zero-shot food detection
Weiqing Min, Guorui Sheng, Jingru Song, Yancun Yang, Shuqiang Jiang |
Pattern Recognit. | 2 |
| 2026 | SyMFood: Synergistic Multi-Modal Prompting for Fine-Grained Zero-Shot Food DetectionabstractFine-grained object detection in food computing is severely constrained by the vast diversity of food items and the high cost of data annotation. Existing Zero-Shot Food Detection (ZSFD) methods attempt to solve this by leveraging semantic information, but they suffer from two critical bottlenecks: (1) a "Semantic Dilemma" (SD) where textual descriptions are too ambiguous to distinguish visually similar food categories, and (2) an "Architectural Bottleneck" (AB) due to the granularity mismatch between high-level semantics and low-level visual features. In this article, we propose SyMFood (Synergistic Multi-modal Framework for Food Generalization), a novel ZSFD framework designed to systematically overcome these challenges. To resolve the SD, SyMFood employs a multi-modal prompt system, which combines rich descriptions from Large Language Models (LLMs) with unambiguous visual exemplars to provide precise semantic grounding. To break the AB, SyMFood introduces a “Refine-then-Fuse” architecture. This design first utilizes a Context-Aware Spatial-Channel Refinement (CaSC) block to enhance visual features independently. Subsequently, a Progressive Food Knowledge Fusion (ProFus) module performs bi-directional, iterative co-refinement between the enhanced visual features and multi-modal prompts across all scales. Extensive experiments across four challenging datasets, including food-specific (UEC FOOD 256, FOWA) and general-purpose (PASCAL VOC, MS COCO) benchmarks, validate our approach. The proposed method outperforms baselines, yielding a notable 8.5% improvement in Harmonic Mean on the FOWA dataset in the genearl ZSD (GZSD) setting. The source code will be available at https://github.com/Niko000202/SymFood0202. Weiqing Min, Shoulong Liu, Guorui Sheng, Shuqiang Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | CondFoodGen: A Conditional Two-Stream Network for Controllable Food Image GenerationabstractFood image generation is an important research direction in food computing, aiming to produce highly realistic images that accurately capture the visual characteristics of various dishes while adhering to specified input conditions. Existing methods that rely solely on textual descriptions struggle to handle the large intra-class variability of food, often resulting in limited diversity and accuracy. Although some approaches incorporate additional conditions, they generally lack optimizations for food-specific challenges, leading to inconsistencies in texture, shape, and color fidelity. To address these limitations, we propose CondFoodGen, a diffusion-based two-stream network for controllable food image generation. The architecture consists of a control stream and a generation stream, where the control stream provides conditional guidance to regulate the generation process. To optimize bidirectional interactions between the two streams, we introduce the Bidirectional Adaptive Gating (BAG) mechanism, which not only guides synthesis but also adaptively refines control representations through feedback from the generation stream. In addition, we propose the Wavelet-Guided Hierarchical Attention (WGHA) module, which combines wavelet-based multi-frequency analysis with hierarchical attention to enhance fine-grained texture fidelity and structural realism. A progressive multi-stage training strategy further stabilizes optimization and enables seamless integration of conditional guidance with bidirectional interaction. Extensive experiments on three food image datasets demonstrate that CondFoodGen consistently generates high-quality and diverse images. Compared with the best existing food image generation methods, our approach achieves an average improvement of about 11.0% across three evaluation metrics and compared to the leading conditional generation approaches, the average improvement reaches 16.2%. The source code, trained models, and supplementary materials are publicly available at https://github.com/housujuan123/CondFoodGen. Mengyao Zhao, Hao Xiong 0001, Weiqing Min, Sujuan Hou, Mengmeng Zhang 0008, Shuqiang Jiang |
IEEE Trans. Image Process. | 3 |
| 2026 | FoodHash: Context-Aware Proxy Interaction and Fusion for Food Image RetrievalabstractVision-based food image retrieval has garnered significant attention due to its potential for critical applications in dietary and health management. However, food images exhibit more complex feature distributions and lack the geometric regularity and structured patterns typically observed in general image retrieval tasks. This complexity poses a challenge for existing models to extract fine-grained features and semantic information, thereby compromising retrieval performance. To address this challenge, we propose FoodHash, a context-aware proxy interaction and fusion hashing method for food image retrieval. The method incorporates an Aggregation–Interaction–Propagation (AIP) module that facilitates contextual information exchange among patch tokens within the same feature map, guided by proxy tokens, thereby effectively capturing the intricate details of food images. Furthermore, to leverage the rich semantic information in food images, a Cross-Fusion Module is introduced to efficiently integrate multi-scale information and enhance feature representation. Additionally, we employ a novel loss function to optimize hash learning by ensuring consistency between hash codes and the semantic space, thereby enhancing the learning capability of hash coding. Extensive experiments on three publicly available food datasets demonstrate that FoodHash significantly surpasses existing models in retrieval performance. Specifically, on the ETH Food-101 dataset, FoodHash achieves improvements of 18.1%, 6.7%, 5.2%, and 4.5% over the suboptimal method PTLCH for 16-bits, 32-bits, 48-bits, and 64-bits hash codes, respectively. The source code will be made publicly available upon publication of the article. Pindan Cao, Weiqing Min, Guorui Sheng, Yongqiang Song, Lili Wang 0009, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | DSDGF-Nutri: A Decoupled Self-Distillation Network with Gating Fusion For Food Nutritional AssessmentabstractAccurate assessment of food nutrition is essential for promoting healthy eating habits. While recent deep learning approaches have enhanced vision-based nutritional estimation through RGB-D multi-modal fusion, they often overlook fine-grained surface components (e.g., oil and sugar) that significantly influence nutritional values. Some recent approaches have improved accuracy by incorporating ingredient data, but their reliance on such input during inference limits practical applicability, as ingredient details are often unavailable in real-world settings. To address this limitation, we propose DSDGF-Nutri, a novel Decoupled Self-Distillation network with Gating Fusion for food Nutri tional assessment. Our method leverages ingredient knowledge during training but relies solely on RGB-D inputs at inference. Specifically, DSDGF-Nutri introduces: (1) a self-distillation mechanism with gating fusion that transfers ingredient-aware features to the RGB-D network, enabling robust prediction without test-time ingredient input, and (2) a multi-task decoupling architecture with task-specific decoders to minimize cross-task interference. Extensive evaluations on two benchmark datasets demonstrate DSDGF-Nutri outperforms existing methods, achieving state-of-the-art results. This work establishes a new paradigm of multimodal fusion in nutritional assessment by unifying scientific measurements with scalable computer vision applications. Sujuan Hou, Zhihui Feng, Hao Xiong 0001, Weiqing Min, Peng Li 0081, Shuqiang Jiang |
ACM Multimedia | 4 |
| 2025 | Spatial-Aware Multi-Modal Information Fusion for Food Nutrition EstimationabstractFood nutrition assessment plays a crucial role in maintaining health, preventing diseases, and promoting scientific dietary habits. However, existing nutrition assessment methods often fail to fully consider the relationships between tasks, leading to limited overall performance. Specifically, these methods suffer from three major challenges: (1) task conflicts, where different tasks compete during joint optimization, leading to suboptimal overall performance; (2) varying training difficulties among tasks, leading to imbalanced learning and subpar model generalization; and (3) the small-scale and complex distribution of datasets, which limits the robustness of learned representations. To address these issues, we propose a novel method that reduces interference between tasks, dynamically focuses on more challenging tasks, and incorporates 3D spatial awareness to enhance multi-modal feature representation. First, we decouple the prediction network from the backbone and introduce a CAMTH (Cross-Attention-Based Multi-Task Head Module), effectively mitigating task interference and fully leveraging each task's learning potential. Second, we improve the loss function to adaptively focus on more challenging tasks, improving overall model performance. Third, we design a 3D-FEM (3D Feature Extraction Module) and MMFF (Multi-Modal Feature Fusion Module), enabling the model to fully exploit the spatial information of food and enhance the food's multi-modal feature representation. We validate our method through extensive experiments on the Nutrition5K dataset, comparing it with state-of-the-art (SOTA) models. The results show that our method achieves superior performance in nutrition estimation, demonstrating the effectiveness of our method. Dongjian Yu, Weiqing Min, Xin Jin 0005, Shuqiang Jiang |
ACM Multimedia | 2 |
| 2025 | Cross-Layer and Selective Distillation for Asymmetric Image Retrieval
Weiqing Min, Fangyuan Yao, Guorui Sheng, Shuqiang Jiang |
PRCV (12) | 2 |
| 2025 | Channel grouping vision transformer for lightweight fruit and vegetable recognition
Chengxu Liu 0002, Weiqing Min, Jingru Song, Yancun Yang, Guorui Sheng, Lili Wang 0009, Shuqiang Jiang |
Expert Syst. Appl. | 2 |
| 2025 | SSC-PPI: A Subspace Structure Consistency-Based Method for Protein-Protein Interactions PredictionabstractProtein-protein interactions (PPIs) play an indispensable role in understanding disease-causing mechanisms, and the basic laws of food and drugs on life. Contemporary research on this issue, however, is incapable of guaranteeing structure consistency between extracted features and raw data, and fails to fully investigate the interconnection information of features. Thus, this paper proposes a subspace structure consistency-based method for protein-protein interactions prediction. SSC-PPI is not only capable of investigating the coherent relations between the encoded features generated from amino acid composition and conjoint triad numeric composition of F-vector, composition and transition descriptors, but also fully maintains the latent geometrical structure consistency between feature subspace and data space. Numerous comparative experiments demonstrate its excellent predictable performance with significant accuracies of 100$\%$, 99.95$\%$, 99.98$\%$, 100$\%$ and 100$\%$ respectively on Helicobacter pylori, Human, Saccharomyces cerevisiae (core subset), Human-Bacillus Anthracis and Human-Yersinia pestis datasets, significantly outperforming the comparative models by average increases of 14.39$\%$, 5.45$\%$, 8.10$\%$, 6.05$\%$ and 8.79$\%$ respectively. Additionally, SSC-PPI offers an efficient and reliable framework for large-scale prediction tasks such as drug-drug and drug-food interactions. Ziping Ma 0001, Weiqing Min, Huanpu Zhang, Yulei Huang, Shuqiang Jiang |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | Food3D: Text-Driven Customizable 3D Food Generation With Gaussian SplattingabstractRealistic 3D food creation generation plays a critical role in applications such as nutritional assessment, advertising, and virtual content creation. The existing text-to-3D models typically begin by initializing a 3D representation, which is subsequently refined using supervision from a text-to-image model to obtain the final 3D output. In this work, we present Food3D, a novel framework for 3D food generation designed to address two main limitations of current models. First, the limitation of initialization in 3D generation: poor initialization can result in the generated 3D food lacking crucial details and realism, thereby reducing its quality. To address this issue, we propose a generalized method named Food3D-G, which uses Mamba-based initialization to improve the starting point of the initialization process, thereby enhancing the visual fidelity and quality of the generated 3D food. Second, the limitation of text-to-image models: current text-to-3D models often rely on text-to-image models for supervision. However, a considerable gap persists between the generated images and real-world visuals, particularly when modeling complex food structures. These models fail to accurately capture the fine details and textures, which negatively impacts the quality and realism of the generated 3D food models. To address this limitation, we propose a customizable method for personalized 3D food generation, termed Food3D-C. This method employs a dual-branch diffusion model that effectively captures intricate details, particularly in complex food structures. Within the Food3D framework, both proposed methods incorporate 3D Gaussian splatting (3D GS) and a schedulable interval score matching (S-ISM) algorithm to enhance shape and texture generation. Extensive experiments demonstrate that Food3D achieves state-of-the-art performance, with substantial improvements in detail, shape accuracy, and overall visual realism. Project page and source codes: https://yudongjian.github.io/Food3D/. Dongjian Yu, Weiqing Min, Xin Jin 0005, Shaowen Yao 0001, Shuqiang Jiang |
IEEE Trans. Image Process. | 2 |
| 2025 | Guest Editorial: When Multimedia Meets Food: Multimedia Computing for Food Data Analysis and Applications
Weiqing Min, Shuqiang Jiang, Petia Radeva, Vladimir Pavlovic 0001, Chong-Wah Ngo, Kiyoharu Aizawa, Wanqing Li 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Multimodal Food LearningabstractFood-centered study has received more attention in the multimedia community for its profound impact on our survival, nutrition and health, pleasure, and enjoyment. Our experience of food is typically multi-sensory: We see food objects, smell its odors, taste its flavors, feel its texture, and hear sounds when chewing. Therefore, multimodal food learning is vital in food-centered study, which aims to relate information from multiple food modalities to support various multimedia tasks, ranging from recognition, retrieval, generation, recommendation, and interaction, enabling applications in different fields like healthcare and agriculture. However, there is no surveys on this topic to our knowledge. To fill this gap, this article formalizes multimodal food learning and comprehensively surveys its typical tasks, technical achievements, existing datasets, and applications to provide the blueprint with researchers and practitioners. Based on the current state of the art, we identify both open research issues and promising research directions, such as multimodal food learning benchmark construction, multimodal food foundation model construction, and multimodality diet estimation. We also point out that closer cooperation from researchers between multimedia and food science can handle some existing challenges and meanwhile open up more new opportunities to advance the fast development of multimodal food learning. This is the first comprehensive survey in this topic and we anticipate about 170 reviewed research articles can benefit academia and industry in this community and beyond. Weiqing Min, Xingjian Hong, Yuxin Liu 0009, Leyi Xu, Yilin Wang 0012, Shuqiang Jiang, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Diverse and High-Quality Food Image Generation from Only Food NamesabstractFood image generation holds promising application prospects in food design, advertising, and food education. However, the existing methods rely on information such as recipes, ingredients, or food names, which leads to generated food images with less intra-class diversity. When recipes, ingredients, and food names are identical for the same food, the real-world images may vary significantly in appearance. The question of how to simultaneously ensure the quality and diversity of the generated images is a key issue. To this end, we employ pre-trained diffusion model and Transformer to propose a method for generating diverse and high-quality images of both Chinese and Western food, named CW-Food. Different from previous works that utilize an overall food feature to generate new images, CW-Food first decouples the food images to obtain common intra-class features and private instance features. Additionally, we design a Transformer-based feature fusion module to integrate the common and private features, in order to avoid the shortcomings of conventional methods. Moreover, we also utilize a pre-trained diffusion model as our backbone, which is fine-tuned using LoRA with the fused multi-variate features. Extensive experiments on four datasets demonstrate the advantages of our proposed method, producing diverse and high-quality food images encompassing both Chinese and Western cuisines. To the best of our knowledge, our work is the first attempt to generate Chinese food images using only food names. Dongjian Yu, Weiqing Min, Xin Jin 0005, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Multi-state Ingredient Recognition via Adaptive Multi-centric NetworkabstractIngredient recognition has received significant attention due to its numerous industrial applications, such as intelligent retail terminals and intelligent cooking devices. However, ingredient recognition has the following challenges: 1) dynamic changes in the number of categories; 2) greater diversity and regionality of ingredients; and 3) large visual differences among different states of ingredients. In this article, we propose an adaptive multi-centric network (AdMNet) to solve the problem of ingredient recognition. AdMNet is based on the idea of retrieval, which consists of two main parts, the adaptive multi-centric nearest-neighbor central mean (AdM-NCM) classifier, and the context-aware attentional pooling (CAP) module. The AdM-NCM classifier adaptively establishes category-centric vector groups to recognize ingredients via optimizing the minimum clustering variance, where each state of the ingredient has its corresponding centric vector. The CAP module combines contextual information and multiple attention mechanisms. It captures more focused and discriminative features with higher weights assigned to fine-grained features, which results in better feature representation. In addition, we collect a large-scale ingredient dataset, ISIA Ingredient-201 with 201 classes and 100 442 images. To prove the greater robustness and generalization of our method, we compare the metrics in basic scenarios and realistic scenarios with those of other methods. Specifically, the base scenario is the regular setup, and the real scenario is similar to the class incremental learning setup. The experimental results show that our method reaches the state of the art on both basic scenarios and realistic scenarios with small samples. Jiajun Song, Weiqing Min, Weimin Xiao, Shuqiang Jiang |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Convolution-Enhanced Bi-Branch Adaptive Transformer With Cross-Task Interaction for Food Category and Ingredient RecognitionabstractRecently, visual food analysis has received more and more attention in the computer vision community due to its wide application scenarios, e.g., diet nutrition management, smart restaurant, and personalized diet recommendation. Considering that food images are unstructured images with complex and unfixed visual patterns, mining food-related semantic-aware regions is crucial. Furthermore, the ingredients contained in food images are semantically related to each other due to the cooking habits and have significant semantic relationships with food categories under the hierarchical food classification ontology. Therefore, modeling the long-range semantic relationships between ingredients and the categories-ingredients semantic interactions is beneficial for ingredient recognition and food analysis. Taking these factors into consideration, we propose a multi-task learning framework for food category and ingredient recognition. This framework mainly consists of a food-orient Transformer named Convolution-Enhanced Bi-Branch Adaptive Transformer (CBiAFormer) and a multi-task category-ingredient recognition network called Structural Learning and Cross-Task Interaction (SLCI). In order to capture the complex and unfixed fine-grained patterns of food images, we propose a query-aware data-adaptive attention mechanism called Bi-Branch Adaptive Attention (BiA-Attention) in CBiAFormer, which consists of a local fine-grained branch and a global coarse-grained branch to mine local and global semantic-aware regions for different input images through an adaptive candidate key/value sets assignment for each query. Additionally, a convolutional patch embedding module is proposed to extract the fine-grained features which are neglected by Transformers. To fully utilize the ingredient information, we propose SLCI, which consists of cross-layer attention to model the semantic relationships between ingredients and two cross-task interaction modules to mine the semantic interactions between categories and ingredients. Extensive experiments show that our method achieves competitive performance on three mainstream food datasets (ETH Food-101, Vireo Food-172, and ISIA Food-200). Visualization analyses of CBiAFormer and SLCI on two tasks prove the effectiveness of our method. Codes will be released upon publication. Code and models are available at https://github.com/Liuyuxinict/CBiAFormer. Yuxin Liu 0009, Weiqing Min, Shuqiang Jiang, Yong Rui |
IEEE Trans. Image Process. | 2 |
| 2024 | Synthesizing Knowledge-Enhanced Features for Real-World Zero-Shot Food DetectionabstractFood computing brings various perspectives to computer vision like vision-based food analysis for nutrition and health. As a fundamental task in food computing, food detection needs Zero-Shot Detection (ZSD) on novel unseen food objects to support real-world scenarios, such as intelligent kitchens and smart restaurants. Therefore, we first benchmark the task of Zero-Shot Food Detection (ZSFD) by introducing FOWA dataset with rich attribute annotations. Unlike ZSD, fine-grained problems in ZSFD like inter-class similarity make synthesized features inseparable. The complexity of food semantic attributes further makes it more difficult for current ZSD methods to distinguish various food categories. To address these problems, we propose a novel framework ZSFDet to tackle fine-grained problems by exploiting the interaction between complex attributes. Specifically, we model the correlation between food categories and attributes in ZSFDet by multi-source graphs to provide prior knowledge for distinguishing fine-grained features. Within ZSFDet, Knowledge-Enhanced Feature Synthesizer (KEFS) learns knowledge representation from multiple sources (e.g., ingredients correlation from knowledge graph) via the multi-source graph fusion. Conditioned on the fusion of semantic knowledge representation, the region feature diffusion model in KEFS can generate fine-grained features for training the effective zero-shot detector. Extensive evaluations demonstrate the superior performance of our method ZSFDet on FOWA and the widely-used food dataset UECFOOD-256, with significant improvements by 1.8% and 3.7% ZSD mAP compared with the strong baseline RRFS. Further experiments on PASCAL VOC and MS COCO prove that enhancement of the semantic knowledge can also improve the performance on general ZSD. Code and dataset are available at https://github.com/LanceZPF/KEFS. Weiqing Min, Jiajun Song, Yang Zhang 0117, Shuqiang Jiang |
IEEE Trans. Image Process. | 2 |
| 2024 | Deep Learning for Logo Detection: A SurveyabstractLogo detection has gradually become a research hotspot in the field of computer vision and multimedia for its various applications, such as social media monitoring, intelligent transportation, and video advertising recommendation. Recent advances in this area are dominated by deep learning-based solutions, where many datasets, learning strategies, network architectures, and loss functions have been employed. This article reviews the advance in applying deep learning techniques to logo detection. First, we discuss a comprehensive account of public datasets designed to facilitate performance evaluation of logo detection algorithms, which tend to be more diverse, more challenging, and more reflective of real life. Next, we perform an in-depth analysis of the existing logo detection strategies and their strengths and weaknesses of each learning strategy. Subsequently, we summarize the applications of logo detection in various fields, from intelligent transportation and brand monitoring to copyright and trademark compliance. Finally, we analyze the potential challenges and present the future directions for the development of logo detection. This study aims better to inform readers about the current state of logo detection and encourage more researchers to get involved in logo detection. Sujuan Hou, Weiqing Min, Yanna Zhao, Yuanjie Zheng, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Towards Food Image Retrieval via Generalization-Oriented Sampling and Loss Function DesignabstractFood computing has increasingly received widespread attention in the multimedia field. As a basic task of food computing, food image retrieval has wide applications, that is, food image retrieval can help users to find the desired food from a large number of food images. Besides, the retrieved information can be applied to establish a richer database for the subsequent food content-related recommendation. Food image retrieval aims to achieve better performance on novel categories. Thus, it is worth studying to transfer the embedding ability from the training set to the unseen test set, that is, the generalization of the model. Food is influenced by various factors, such as culture and geography, leading to great differences between domains, such as Asian food and western food. Therefore, it is challenging to study the generalization of the model in food image retrieval. In this article, we improve the classical metric learning framework and propose a generalization-oriented sampling strategy, which boosts the generalization of the model by maximizing the intra-class distance from a proportion of positive pairs to avoid the excessive distance compression in the embedding space. Considering that the existing optimization process is in an opposite direction to our proposed sampling strategy, we further propose an adaptive gradient assignment policy named gradient-adaptive optimization , which can alleviate the intra-class distance compression during optimization by assigning different gradients to different samples. Extensive evaluation on three popular food image datasets demonstrates the effectiveness of the proposed method. We also experiment on three popular general datasets to prove that solving the problem from the generalization can also improve the performance of general image retrieval. Code is available at https://github.com/Jiajun-ISIA/Generalization-oriented-Sampling-and-Loss . Jiajun Song, Weiqing Min, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Lightweight Food Recognition via Aggregation Block and Feature EncodingabstractFood image recognition has recently been given considerable attention in the multimedia field in light of its possible implications on health. The characteristics of the dispersed distribution of ingredients in food images put forward higher requirements on the long-range information extraction ability of neural networks, leading to more complex and deeper models. Nevertheless, the lightweight version of food image recognition is essential for improved implementation on end devices and sustained server-side expansion. To address this issue, we present Aggregation Feature Net (AFNet), a lightweight network that is capable of effectively capturing both global and local features from food images. In AFNet, we develop a novel convolution based on a residual model by encoding global features through row-wise and column-wise information integration. Merging aggregation block with classic local convolution yields a framework that works as the backbone of the network. Based on the efficient use of parameters by the aggregation block, we constructed a lightweight food image recognition network with fewer layers and a smaller scale, assisted by a new type of activation function. Experimental results on four popular food recognition datasets demonstrate that our approach achieves state-of-the-art performance with higher accuracy and fewer FLOPs and parameters. For example, in comparison to the current state-of-the-art model of MobileViTv2, AFNet achieved 88.4% accuracy of the top-1 level on the ETHZ Food-101 dataset, with similar parameters and FLOPs but 1.4% more accuracy. The source code will be provided in supplementary materials. Yancun Yang, Weiqing Min, Jingru Song, Guorui Sheng, Lili Wang 0009, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Toward Egocentric Compositional Action Anticipation with Adaptive Semantic DebiasingabstractPredicting the unknown from the first-person perspective is expected as a necessary step toward machine intelligence, which is essential for practical applications including autonomous driving and robotics. As a human-level task, egocentric action anticipation aims at predicting an unknown action seconds before it is performed from the first-person viewpoint. Egocentric actions are usually provided as verb-noun pairs; however, predicting the unknown action may be trapped in insufficient training data for all possible combinations. Therefore, it is crucial for intelligent systems to use limited known verb-noun pairs to predict new combinations of actions that have never appeared, which is known as compositional generalization. In this article, we are the first to explore the egocentric compositional action anticipation problem, which is more in line with real-world settings but neglected by existing studies. Whereas prediction results are prone to suffer from semantic bias considering the distinct difference between training and test distributions, we further introduce a general and flexible adaptive semantic debiasing framework that is compatible with different deep neural networks. To capture and mitigate semantic bias, we can imagine one counterfactual situation where no visual representations have been observed and only semantic patterns of observation are used to predict the next action. Instead of the traditional counterfactual analysis scheme that reduces semantic bias in a mindless way, we devise a novel counterfactual analysis scheme to adaptively amplify or penalize the effect of semantic experience by considering the discrepancy both among categories and among examples. We also demonstrate that the traditional counterfactual analysis scheme is a special case of the devised adaptive counterfactual analysis scheme. We conduct experiments on three large-scale egocentric video datasets. Experimental results verify the superiority and effectiveness of our proposed solution. Weiqing Min, Shuqiang Jiang, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | A Cross-direction Task Decoupling Network for Small Logo DetectionabstractLogo detection plays an integral role in many applications. However, handling small logos is still difficult since they occupy too few pixels in the image, which burdens the extraction of discriminative features. The aggregation of small logos also brings a great challenge to the classification and localization of logos. To solve these problems, we creatively propose Cross-direction Task Decoupling Network (CTDNet) for small logo detection. We first introduce Cross-direction Feature Pyramid (CFP) to realize cross-direction feature fusion by adopting horizontal transmission and vertical transmission. In addition, Multi-frequency Task Decoupling Head (MTDH) decouples the classification and localization tasks into two branches. A multi-frequency attention convolution branch is designed to achieve more accurate regression by combining discrete cosine transform and convolution creatively. Comprehensive experiments on four logo datasets demonstrate the effectiveness and efficiency of the proposed method. Sujuan Hou, Xingzhuo Li, Weiqing Min, Jing Wang 0138, Yuanjie Zheng, Shuqiang Jiang |
ICME | 3 |
| 2023 | SeeDS: Semantic Separable Diffusion Synthesizer for Zero-shot Food DetectionabstractFood detection is becoming a fundamental task in food computing that supports various multimedia applications, including food recommendation and dietary monitoring. To deal with real-world scenarios, food detection needs to localize and recognize novel food objects that are not seen during training, demanding Zero-Shot Detection (ZSD). However, the complexity of semantic attributes and intra-class feature diversity poses challenges for ZSD methods in distinguishing fine-grained food classes. To tackle this, we propose the Semantic Separable Diffusion Synthesizer (SeeDS) framework for Zero-Shot Food Detection (ZSFD). SeeDS consists of two modules: a Semantic Separable Synthesizing Module (S3M) and a Region Feature Denoising Diffusion Model (RFDDM). The S3M learns the disentangled semantic representation for complex food attributes from ingredients and cuisines, and synthesizes discriminative food features via enhanced semantic information. The RFDDM utilizes a novel diffusion model to generate diversified region features and enhances ZSFD via fine-grained synthesized features. Extensive experiments show the state-of-the-art ZSFD performance of our proposed method on two food datasets, ZSFooD and UECFOOD-256. Moreover, SeeDS also maintains effectiveness on general ZSD datasets, PASCAL VOC and MS COCO. The code and dataset can be found at https://github.com/LanceZPF/SeeDS https://github.com/LanceZPF/SeeDS. Weiqing Min, Yang Zhang 0117, Jiajun Song, Shuqiang Jiang |
ACM Multimedia | 2 |
| 2023 | Dataset Bias in Few-Shot Image RecognitionabstractThe goal of few-shot image recognition (FSIR) is to identify novel categories with a small number of annotated samples by exploiting transferable knowledge from training data (base categories). Most current studies assume that the transferable knowledge can be well used to identify novel categories. However, such transferable capability may be impacted by the dataset bias, and this problem has rarely been investigated before. Besides, most of few-shot learning methods are biased to different datasets, which is also an important issue that needs to be investigated deeply. In this paper, we first investigate the impact of transferable capabilities learned from base categories. Specifically, we use the relevance to measure relationships between base categories and novel categories. Distributions of base categories are depicted via the instance density and category diversity. The FSIR model learns better transferable knowledge from relevant training data. In the relevant data, dense instances or diverse categories can further enrich the learned knowledge. Experimental results on different sub-datasets of Imagenet demonstrate category relevance, instance density and category diversity can depict transferable bias from distributions of base categories. Second, we investigate performance differences on different datasets from the aspects of dataset structures and different few-shot learning methods. Specifically, we introduce image complexity, intra-concept visual consistency, and inter-concept visual similarity to quantify characteristics of dataset structures. We use these quantitative characteristics and eight few-shot learning methods to analyze performance differences on multiple datasets. Based on the experimental analysis, some insightful observations are obtained from the perspective of both dataset structures and few-shot learning methods. We hope these observations are useful to guide future few-shot learning research on new datasets or tasks. Our data is available at http://123.57.42.89/dataset-bias/dataset-bias.html. Shuqiang Jiang, Chenlong Liu, Xinhang Song, Xiangyang Li 0002, Weiqing Min |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Large Scale Visual Food RecognitionabstractFood recognition plays an important role in food choice and intake, which is essential to the health and well-being of humans. It is thus of importance to the computer vision community, and can further support many food-oriented vision and multimodal tasks, e.g., food detection and segmentation, cross-modal recipe retrieval and generation. Unfortunately, we have witnessed remarkable advancements in generic visual recognition for released large-scale datasets, yet largely lags in the food domain. In this paper, we introduce Food2K, which is the largest food recognition dataset with 2,000 categories and over 1 million images. Compared with existing food recognition datasets, Food2K bypasses them in both categories and images by one order of magnitude, and thus establishes a new challenging benchmark to develop advanced models for food visual representation learning. Furthermore, we propose a deep progressive region enhancement network for food recognition, which mainly consists of two components, namely progressive local feature learning and region feature enhancement. The former adopts improved progressive training to learn diverse and complementary local features, while the latter utilizes self-attention to incorporate richer context with multiple scales into local features for further local feature enhancement. Extensive experiments on Food2K demonstrate the effectiveness of our proposed method. More importantly, we have verified better generalization ability of Food2K in various tasks, including food image recognition, food image retrieval, cross-modal recipe retrieval, food detection and segmentation. Food2K can be further explored to benefit more food-relevant tasks including emerging and more complex ones (e.g., nutritional understanding of food), and the trained models on Food2K can be expected as backbones to improve the performance of more food-relevant tasks. We also hope Food2K can serve as a large scale fine-grained visual recognition benchmark, and contributes to the development of large scale fine-grained visual analysis. Weiqing Min, Yuxin Liu 0009, Mengjiang Luo, Liping Kang, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Ingredient Prediction via Context Learning Network With Class-Adaptive Asymmetric LossabstractIngredient prediction has received more and more attention with the help of image processing for its diverse real-world applications, such as nutrition intake management and cafeteria self-checkout system. Existing approaches mainly focus on multi-task food category-ingredient joint learning to improve final recognition by introducing task relevance, while seldom pay attention to making good use of inherent characteristics of ingredients independently. Actually, there are two issues for ingredient prediction. First, compared with fine-grained food recognition, ingredient prediction needs to extract more comprehensive features of the same ingredient and more detailed features of various ingredients from different regions of the food image. Because it can help understand various food compositions and distinguish the differences within ingredient features. Second, the ingredient distributions are extremely unbalanced. Existing loss functions can not simultaneously solve the imbalance between positive-negative samples belonging to each ingredient and significant differences among all classes. To solve these problems, we propose a novel framework named Class-Adaptive Context Learning Network (CACLNet) for ingredient prediction. In order to extract more comprehensive and detailed features, we introduce Ingredient Context Learning (ICL) to reduce the negative impact of complex background in food images and construct internal spatial connections among ingredient regions of food objects in a self-supervised manner, which can strengthen the contacts of the same ingredients through region interactions. In order to solve the imbalance of different classes among ingredients, we propose one novel Class-Adaptive Asymmetric Loss (CAAL) to focus on various ingredient classes adaptively. Besides, considering that the over-suppression of negative samples will over-fit positive samples of those rare ingredients, CAAL alleviates this continuous suppression according to the imbalanced ratios based on gradients while maintaining the contribution of positive samples by lesser suppression. Extensive evaluation on two popular benchmark datasets (Vireo Food-172, UEC Food-100) demonstrates our proposed method achieves the state-of-the-art performance. Further qualitative analysis and visualization show the effectiveness of our method. Code and models are available at https://123.57.42.89/codes/CACLNet/index.html. Mengjiang Luo, Weiqing Min, Jiajun Song, Shuqiang Jiang |
IEEE Trans. Image Process. | 2 |
| 2022 | Rethinking the Optimization of Average Precision: Only Penalizing Negative Instances before Positive Ones Is EnoughabstractOptimising the approximation of Average Precision (AP) has been widely studied for image retrieval. Limited by the definition of AP, such methods consider both negative and positive instances ranking before each positive instance. However, we claim that only penalizing negative instances before positive ones is enough, because the loss only comes from these negative instances. To this end, we propose a novel loss, namely Penalizing Negative instances before Positive ones (PNP), which can directly minimize the number of negative instances before each positive one. In addition, AP-based methods adopt a fixed and sub-optimal gradient assignment strategy. Therefore, we systematically investigate different gradient assignment solutions via constructing derivative functions of the loss, resulting in PNP-I with increasing derivative functions and PNP-D with decreasing ones. PNP-I focuses more on the hard positive instances by assigning larger gradients to them and tries to make all relevant instances closer. In contrast, PNP-D pays less attention to such instances and slowly corrects them. For most real-world data, one class usually contains several local clusters. PNP-I blindly gathers these clusters while PNP-D keeps them as they were. Therefore, PNP-D is more superior. Experiments on three standard retrieval datasets show consistent results with the above analysis. Extensive evaluations demonstrate that PNP-D achieves the state-of-the-art performance. Code is available at https://github.com/interestingzhuo/PNPloss Weiqing Min, Jiajun Song, Liping Kang, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang |
AAAI | 2 |
| 2022 | Ingredient-Guided Region Discovery and Relationship Modeling for Food Category-Ingredient PredictionabstractRecognizing the category and its ingredient composition from food images facilitates automatic nutrition estimation, which is crucial to various health relevant applications, such as nutrition intake management and healthy diet recommendation. Since food is composed of ingredients, discovering ingredient-relevant visual regions can help identify its corresponding category and ingredients. Furthermore, various ingredient relationships like co-occurrence and exclusion are also critical for this task. For that, we propose an ingredient-oriented multi-task food category-ingredient joint learning framework for simultaneous food recognition and ingredient prediction. This framework mainly involves learning an ingredient dictionary for ingredient-relevant visual region discovery and building an ingredient-based semantic-visual graph for ingredient relationship modeling. To obtain ingredient-relevant visual regions, we build an ingredient dictionary to capture multiple ingredient regions and obtain the corresponding assignment map, and then pool the region features belonging to the same ingredient to identify the ingredients more accurately and meanwhile improve the classification performance. For ingredient-relationship modeling, we utilize the visual ingredient representations as nodes and the semantic similarity between ingredient embeddings as edges to construct an ingredient graph, and then learn their relationships via the graph convolutional network to make label embeddings and visual features interact with each other to improve the performance. Finally, fused features from both ingredient-oriented region features and ingredient-relationship features are used in the following multi-task category-ingredient joint learning. Extensive evaluation on three popular benchmark datasets (ETH Food-101, Vireo Food-172 and ISIA Food-200) demonstrates the effectiveness of our method. Further visualization of ingredient assignment maps and attention maps also shows the superiority of our method. Weiqing Min, Liping Kang, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang |
IEEE Trans. Image Process. | 2 |
| 2022 | LogoDet-3K: A Large-scale Image Dataset for Logo DetectionabstractLogo detection has been gaining considerable attention because of its wide range of applications in the multimedia field, such as copyright infringement detection, brand visibility monitoring, and product brand management on social media. In this article, we introduce LogoDet-3K, the largest logo detection dataset with full annotation, which has 3,000 logo categories, about 200,000 manually annotated logo objects, and 158,652 images. LogoDet-3K creates a more challenging benchmark for logo detection, for its higher comprehensive coverage and wider variety in both logo categories and annotated objects compared with existing datasets. We describe the collection and annotation process of our dataset and analyze its scale and diversity in comparison to other datasets for logo detection. We further propose a strong baseline method Logo-Yolo, which incorporates Focal loss and CIoU loss into the basic YOLOv3 framework for large-scale logo detection. It obtains about 4% improvement on the average performance compared with YOLOv3, and greater improvements compared with reported several deep detection models on LogoDet-3K. We perform extensive evaluation on three other existing datasets to further verify on both logo detection and retrieval tasks, and we demonstrate better generalization ability of LogoDet-3K on logo detection and retrieval tasks. The LogoDet-3K dataset is used to promote large-scale logo-related research. The code and LogoDet-3K can be found at https://github.com/Wangjing1551/LogoDet-3K-Dataset. Jing Wang 0138, Weiqing Min, Sujuan Hou, Shengnan Ma, Yuanjie Zheng, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | What If We Could Not See? Counterfactual Analysis for Egocentric Action AnticipationabstractEgocentric action anticipation aims at predicting the near future based on past observation in first-person vision. While future actions may be wrongly predicted due to the dataset bias, we present a counterfactual analysis framework for egocentric action anticipation (CA-EAA) to enhance the capacity. In the factual case, we can predict the upcoming action based on visual features and semantic labels from past observation. Imagining one counterfactual situation where no visual representation had been observed, we would obtain a counterfactual predicted action only using past semantic labels. In this way, we can reduce the side-effect caused by semantic labels via a comparison between factual and counterfactual outcomes, which moves a step towards unbiased prediction for egocentric action anticipation. We conduct experiments on two large-scale egocentric video datasets. Qualitative and quantitative results validate the effectiveness of our proposed CA-EAA. Weiqing Min, Shuqiang Jiang, Yong Rui |
IJCAI | 2 |
| 2021 | FoodLogoDet-1500: A Dataset for Large-Scale Food Logo Detection via Multi-Scale Feature Decoupling NetworkabstractFood logo detection plays an important role in the multimedia for its wide real-world applications, such as food recommendation of the self-service shop and infringement detection on e-commerce platforms. A large-scale food logo dataset is urgently needed for developing advanced food logo detection algorithms. However, there are no available food logo datasets with food brand information. To support efforts towards food logo detection, we introduce the dataset FoodLogoDet-1500, a new large-scale publicly available food logo dataset, which has 1,500 categories, about 100,000 images and about 150,000 manually annotated food logo objects. We describe the collection and annotation process of FoodLogoDet-1500, analyze its scale and diversity, and compare it with other logo datasets. To the best of our knowledge, FoodLogoDet-1500 is the first largest publicly available high-quality dataset for food logo detection. The challenge of food logo detection lies in the large-scale categories and similarities between food logo categories. For that, we propose a novel food logo detection method Multi-scale Feature Decoupling Network (MFDNet), which decouples classification and regression into two branches and focuses on the classification branch to solve the problem of distinguishing multiple food logo categories. Specifically, we introduce the feature offset module, which utilizes the deformation-learning for optimal classification offset and can effectively obtain the most representative features of classification in detection. In addition, we adopt a balanced feature pyramid in MFDNet, which pays attention to global information, balances the multi-scale feature maps, and enhances feature extraction capability. Comprehensive experiments on FoodLogoDet-1500 and other two popular benchmark logo datasets demonstrate the effectiveness of the proposed method. The code and FoodLogoDet-1500 can be found at https://github.com/hq03/FoodLogoDet-1500-Dataset. Weiqing Min, Jing Wang 0138, Sujuan Hou, Yuanjie Zheng, Shuqiang Jiang |
ACM Multimedia | 2 |
| 2021 | Plant Disease Recognition: A Large-Scale Benchmark Dataset and a Visual Region and Loss Reweighting ApproachabstractPlant disease diagnosis is very critical for agriculture due to its importance for increasing crop production. Recent advances in image processing offer us a new way to solve this issue via visual plant disease analysis. However, there are few works in this area, not to mention systematic researches. In this paper, we systematically investigate the problem of visual plant disease recognition for plant disease diagnosis. Compared with other types of images, plant disease images generally exhibit randomly distributed lesions, diverse symptoms and complex backgrounds, and thus are hard to capture discriminative information. To facilitate the plant disease recognition research, we construct a new large-scale plant disease dataset with 271 plant disease categories and 220,592 images. Based on this dataset, we tackle plant disease recognition via reweighting both visual regions and loss to emphasize diseased parts. We first compute the weights of all the divided patches from each image based on the cluster distribution of these patches to indicate the discriminative level of each patch. Then we allocate the weight to each loss for each patch-label pair during weakly-supervised training to enable discriminative disease part learning. We finally extract patch features from the network trained with loss reweighting, and utilize the LSTM network to encode the weighed patch feature sequence into a comprehensive feature representation. Extensive evaluations on this dataset and another public dataset demonstrate the advantage of the proposed method. We expect this research will further the agenda of plant disease recognition in the community of image processing. Xinda Liu, Weiqing Min, Shuhuan Mei, Lili Wang 0006, Shuqiang Jiang |
IEEE Trans. Image Process. | 2 |
| 2021 | Hybrid-Attention Enhanced Two-Stream Fusion Network for Video Venue PredictionabstractVideo venue category prediction has been drawing more attention in the multimedia community for various applications such as personalized location recommendation and video verification. Most of existing works resort to the information from either multiple modalities or other platforms for strengthening video representations. However, noisy acoustic information, sparse textual descriptions and incompatible cross-platform data could limit the performance gain and reduce the universality of the model. Therefore, we focus on discriminative visual feature extraction from videos by introducing a hybrid-attention structure. Particularly, we propose a novel Global-Local Attention Module (GLAM), which can be inserted to neural networks to generate enhanced visual features from video content. In GLAM, the Global Attention (GA) is used to catch contextual scene-oriented information via assigning channels with various weights while the Local Attention (LA) is employed to learn salient object-oriented features via allocating different weights for spatial regions. Moreover, GLAM can be extended to ones with multiple GAs and LAs for further visual enhancement. These two types of features respectively captured by GAs and LAs are integrated via convolution layers, and then delivered into convolutional Long Short-Term Memory (convLSTM) to generate spatial-temporal representations, constituting the content stream. In addition, video motions are explored to learn long-term movement variations, which also contributes to video venue prediction. The content and motion stream constitute our proposed Hybrid-Attention Enhanced Two-Stream Fusion Network (HA-TSFN). HA-TSFN finally merges the features from two streams for comprehensive representations. Extensive experiments demonstrate that our method achieves the state-of-the-art performance in the large-scale dataset Vine. The visualization also shows that the proposed GLAM can capture complementary scene-oriented and object-oriented visual features from videos. Our code is available at:https://github.com/zhangyanchao1014/HA-TSFN. Weiqing Min, Liqiang Nie, Shuqiang Jiang |
IEEE Trans. Multim. | 2 |
| 2021 | Attribute-Guided Feature Learning for Few-Shot Image RecognitionabstractFew-shot image recognition has become an essential problem in the field of machine learning and image recognition, and has attracted more and more research attention. Typically, most few-shot image recognition methods are trained across tasks. However, these methods are apt to learn an embedding network for discriminative representations of training categories, and thus could not distinguish well for novel categories. To establish connections between training and novel categories, we use attribute-related representations for few-shot image recognition and propose an attribute-guided two-layer learning framework, which is capable of learning general feature representations. Specifically, few-shot image recognition trained over tasks and attribute learning trained over images share the same network in a multi-task learning framework. In this way, few-shot image recognition learns feature representations guided by attributes, and is thus less sensitive to novel categories compared with feature representations only using category supervision. Meanwhile, the multi-layer features associated with attributes are aligned with category learning on multiple levels respectively. Therefore we establish a two-layer learning mechanism guided by attributes to capture more discriminative representations, which are complementary compared with a single-layer learning mechanism. Experimental results on CUB-200, AWA and MiniImageNet datasets demonstrate our method effectively improves the performance. Weiqing Min, Shuqiang Jiang |
IEEE Trans. Multim. | 2 |
| 2020 | Logo-2K+: A Large-Scale Logo Dataset for Scalable Logo ClassificationabstractLogo classification has gained increasing attention for its various applications, such as copyright infringement detection, product recommendation and contextual advertising. Compared with other types of object images, the real-world logo images have larger variety in logo appearance and more complexity in their background. Therefore, recognizing the logo from images is challenging. To support efforts towards scalable logo classification task, we have curated a dataset, Logo-2K+, a new large-scale publicly available real-world logo dataset with 2,341 categories and 167,140 images. Compared with existing popular logo datasets, such as FlickrLogos-32 and LOGO-Net, Logo-2K+ has more comprehensive coverage of logo categories and larger quantity of logo images. Moreover, we propose a Discriminative Region Navigation and Augmentation Network (DRNA-Net), which is capable of discovering more informative logo regions and augmenting these image regions for logo classification. DRNA-Net consists of four sub-networks: the navigator sub-network first selected informative logo-relevant regions guided by the teacher sub-network, which can evaluate its confidence belonging to the ground-truth logo class. The data augmentation sub-network then augments the selected regions via both region cropping and region dropping. Finally, the scrutinizer sub-network fuses features from augmented regions and the whole image for logo classification. Comprehensive experiments on Logo-2K+ and other three existing benchmark datasets demonstrate the effectiveness of proposed method. Logo-2K+ and the proposed strong baseline DRNA-Net are expected to further the development of scalable logo image recognition, and the Logo-2K+ dataset can be found at https://github.com/msn199959/Logo-2k-plus-Dataset. Jing Wang 0138, Weiqing Min, Sujuan Hou, Shengnan Ma, Yuanjie Zheng, Haishuai Wang, Shuqiang Jiang |
AAAI | 2 |
| 2020 | Food Computing for MultimediaabstractFood computing applies computational approaches for acquiring and analyzing heterogeneous food data from disparate sources for perception, recognition, retrieval, recommendation, prediction and monitoring of food to address food-related issues in multimedia and beyond. It has received more attention from both academia and industry as one emerging interdiscipline for its various applications, such as improving human health and understanding the culinary culture. Recently, there are more studies on food computing in the multimedia, such as food recognition and multimodal recipe analysis. This tutorial will provide a basic understanding of food computing, and discuss its use in various multimedia tasks, ranging from food recognition, retrieval, recommendation, recipe analysis to cooking behavior understanding. Specifically, we will first introduce food computing, including its method, task and applications. Then we will discuss several typical tasks of food computing in the multimedia including food image recognition, food retrieval and recommendation, multimodal recipe analysis and cooking action anticipation. Finally, we will point out future research directions on food computing in the multimedia. Shuqiang Jiang, Weiqing Min |
ACM Multimedia | 2 |
| 2020 | ISIA Food-500: A Dataset for Large-Scale Food Recognition via Stacked Global-Local Attention NetworkabstractFood recognition has received more and more attention in the multimedia community for its various real-world applications, such as diet management and self-service restaurants. A large-scale ontology of food images is urgently needed for developing advanced large-scale food recognition algorithms, as well as for providing the benchmark dataset for such algorithms. To encourage further progress in food recognition, we introduce the dataset ISIA Food-500 with 500 categories from the list in the Wikipedia and 399,726 images, a more comprehensive food dataset that surpasses existing popular benchmark datasets by category coverage and data volume. Furthermore, we propose a stacked global-local attention network, which consists of two sub-networks for food recognition. One sub-network first utilizes hybrid spatial-channel attention to extract more discriminative features, and then aggregates these multi-scale discriminative features from multiple layers into global-level representation (e.g., texture and shape information about food). The other one generates attentional regions (e.g., ingredient relevant regions) from different regions via cascaded spatial transformers, and further aggregates these multi-scale regional features from different layers into local-level representation. These two types of features are finally fused as comprehensive representation for food recognition. Extensive experiments on ISIA Food-500 and other two popular benchmark datasets demonstrate the effectiveness of our proposed method, and thus can be considered as one strong baseline. The dataset, code and models can be found at http://123.57.42.89/FoodComputing-Dataset/ISIA-Food500.html. Weiqing Min, Linhu Liu, Zhengdong Luo, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang |
ACM Multimedia | 1 |
| 2020 | An Egocentric Action Anticipation Framework via Fusing Intuition and AnalysisabstractIn this paper, we focus on egocentric action anticipation from videos, which enables various applications, such as helping intelligent wearable assistants understand users' needs and enhance their capabilities in the interaction process. It requires intelligent systems to observe from the perspective of the first person and predict an action before it occurs. Owing to the uncertainty of future, it is insufficient to perform action anticipation relying on visual information especially when there exists salient visual difference between past and future. In order to alleviate this problem, which we call visual gap in this paper, we propose one novel Intuition-Analysis Integrated (IAI) framework inspired by psychological research, which mainly consists of three parts: Intuition-based Prediction Network (IPN), Analysis-based Prediction Network (APN) and Adaptive Fusion Network (AFN). To imitate the implicit intuitive thinking process, we model IPN as an encoder-decoder structure and introduce one procedural instruction learning strategy implemented by textual pre-training. On the other hand, we allow APN to process information under designed rules to imitate the explicit analytical thinking, which is divided into three steps: recognition, transitions and combination. Both the procedural instruction learning strategy in IPN and the transition step of APN are crucial to improving the anticipation performance via mitigating the visual gap problem. Considering the complementarity of intuition and analysis, AFN adopts attention fusion to adaptively integrate predictions from IPN and APN to produce the final anticipation results. We conduct experiments on the largest egocentric video dataset. Qualitative and quantitative evaluation results validate the effectiveness of our IAI framework, and demonstrate the advantage of bridging visual gap by utilizing multi-modal information, including both visual features of observed segments and sequential instructions of actions. Weiqing Min, Yong Rui, Shuqiang Jiang |
ACM Multimedia | 2 |
| 2020 | Deep neural networks for emerging multimedia computing and applications
Shuqiang Jiang, Weiqing Min, Yonggang Wen 0001, Qingming Huang, Shuicheng Yan |
Neurocomputing | 2 |
| 2020 | Multi-Scale Multi-View Deep Feature Aggregation for Food RecognitionabstractRecently, food recognition has received more and more attention in image processing and computer vision for its great potential applications in human health. Most of the existing methods directly extracted deep visual features via convolutional neural networks (CNNs) for food recognition. Such methods ignore the characteristics of food images and are, thus, hard to achieve optimal recognition performance. In contrast to general object recognition, food images typically do not exhibit distinctive spatial arrangement and common semantic patterns. In this paper, we propose a multi-scale multi-view feature aggregation (MSMVFA) scheme for food recognition. MSMVFA can aggregate high-level semantic features, mid-level attribute features, and deep visual features into a unified representation. These three types of features describe the food image from different granularity. Therefore, the aggregated features can capture the semantics of food images with the greatest probability. For that solution, we utilize additional ingredient knowledge to obtain mid-level attribute representation via ingredient-supervised CNNs. High-level semantic features and deep visual features are extracted from class-supervised CNNs. Considering food images do not exhibit distinctive spatial layout in many cases, MSMVFA fuses multi-scale CNN activations for each type of features to make aggregated features more discriminative and invariable to geometrical deformation. Finally, the aggregated features are more robust, comprehensive, and discriminative via two-level fusion, namely multi-scale fusion for each type of features and multi-view aggregation for different types of features. In addition, MSMVFA is general and different deep networks can be easily applied into this scheme. Extensive experiments and evaluations demonstrate that our method achieves state-of-the-art recognition performance on three popular large-scale food benchmark datasets in Top-1 recognition accuracy. Furthermore, we expect this paper will further the agenda of food recognition in the community of image processing and computer vision. Shuqiang Jiang, Weiqing Min, Linhu Liu, Zhengdong Luo |
IEEE Trans. Image Process. | 2 |
| 2020 | Multi-Task Deep Relative Attribute Learning for Visual Urban PerceptionabstractVisual urban perception aims to quantify perceptual attributes (e.g., safe and depressing attributes) of physical urban environment from crowd-sourced street-view images and their pairwise comparisons. It has been receiving more and more attention in computer vision for various applications, such as perceptive attribute learning and urban scene understanding. Most existing methods adopt either (i) a regression model trained using image features and ranked scores converted from pairwise comparisons for perceptual attribute prediction or (ii) a pairwise ranking algorithm to independently learn each perceptual attribute. However, the former fails to directly exploit pairwise comparisons while the latter ignores the relationship among different attributes. To address them, we propose a Multi-Task Deep Relative Attribute Learning Network (MTDRALN) to learn all the relative attributes simultaneously via multi-task Siamese networks, where each Siamese network will predict one relative attribute. Combined with deep relative attribute learning, we utilize the structured sparsity to exploit the prior from natural attribute grouping, where all the attributes are divided into different groups based on semantic relatedness in advance. As a result, MTDRALN is capable of learning all the perceptual attributes simultaneously via multi-task learning. Besides the ranking sub-network, MTDRALN further introduces the classification sub-network, and these two types of losses from two sub-networks jointly constrain parameters of the deep network to make the network learn more discriminative visual features for relative attribute learning. In addition, our network can be trained in an end-to-end way to make deep feature learning and multi-task relative attribute learning reinforce each other. Extensive experiments on the large-scale Place Pulse 2.0 dataset validate the advantage of our proposed network. Our qualitative results along with visualization of saliency maps also show that the proposed network is able to learn effective features for perceptual attributes. Weiqing Min, Shuhuan Mei, Linhu Liu, Shuqiang Jiang |
IEEE Trans. Image Process. | 1 |
| 2020 | Food Recommendation: Framework, Existing Solutions, and ChallengesabstractA growing proportion of the global population is becoming overweight or obese, leading to various diseases (e.g., diabetes, ischemic heart disease and even cancer) due to unhealthy eating patterns, such as increased intake of food with high energy and high fat. Food recommendation is of paramount importance to alleviate this problem. Unfortunately, modern multimedia research has enhanced the performance and experience of multimedia recommendation in many fields such as movies and POI, yet largely lags in the food domain. This article proposes a unified framework for food recommendation, and identifies main issues affecting food recommendation including incorporating various context and domain knowledge, building the personal model, and analyzing unique food characteristics. We then review existing solutions for these issues, and finally elaborate research challenges and future directions in this field. To our knowledge, this is the first survey that targets the study of food recommendation in the multimedia field and offers a collection of research studies and technologies to benefit researchers in this field. Weiqing Min, Shuqiang Jiang, Ramesh Jain 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | A Two-Stage Triplet Network Training Framework for Image RetrievalabstractIn this paper, we propose a novel framework for instance-level image retrieval. Recent methods focus on fine-tuning the Convolutional Neural Network (CNN) via a Siamese architecture to improve off-the-shelf CNN features. They generally use the ranking loss to train such networks, and do not take full use of supervised information for better network training, especially with more complex neural architectures. To solve this, we propose a two-stage triplet network training framework, which mainly consists of two stages. First, we propose a Double-Loss Regularized Triplet Network (DLRTN), which extends basic triplet network by attaching the classification sub-network, and is trained via simultaneously optimizing two different types of loss functions. Double-loss functions of DLRTN aim at specific retrieval task and can jointly boost the discriminative capability of DLRTN from different aspects via supervised learning. Second, considering feature maps of the last convolution layer extracted from DLRTN and regions detected from the region proposal network as the input, we then introduce the Regional Generalized-Mean Pooling (RGMP) layer for the triplet network, and re-train this network to learn pooling parameters. Through RGMP, we pool feature maps for each region and aggregate features of different regions from each image to Regional Generalized Activations of Convolutions (R-GAC) as final image representation. R-GAC is capable of generalizing existing Regional Maximum Activations of Convolutions (R-MAC) and is thus more robust to scale and translation. We conduct the experiment on six image retrieval datasets including standard benchmarks and recently introduced INSTRE dataset. Extensive experimental results demonstrate the effectiveness of the proposed framework. Weiqing Min, Shuhuan Mei, Shuqiang Jiang |
IEEE Trans. Multim. | 1 |
| 2020 | Few-shot Food Recognition via Multi-view Representation LearningabstractThis article considers the problem of few-shot learning for food recognition. Automatic food recognition can support various applications, e.g., dietary assessment and food journaling. Most existing works focus on food recognition with large numbers of labelled samples, and fail to recognize food categories with few samples. To address this problem, we propose a Multi-View Few-Shot Learning (MVFSL) framework to explore additional ingredient information for few-shot food recognition. Besides category-oriented deep visual features, we introduce ingredient-supervised deep network to extract ingredient-oriented features. As general and intermediate attributes of food, ingredient-oriented features are informative and complementary to category-oriented features, and thus they play an important role in improving food recognition. Particularly in few-shot food recognition, ingredient information can bridge the gap between disjoint training categories and test categories. To take advantage of ingredient information, we fuse these two kinds of features by first combining their feature maps from their respective deep networks and then convolving combined feature maps. Such convolution is further incorporated into a multi-view relation network, which is capable of comparing pairwise images to enable fine-grained feature learning. MVFSL is trained in an end-to-end fashion for joint optimization on two types of feature learning subnetworks and relation subnetworks. Extensive experiments on different food datasets have consistently demonstrated the advantage of MVFSL in multi-view feature fusion. Furthermore, we extend another two types of networks, namely, Siamese Network and Matching Network, by introducing ingredient information for few-shot food recognition. Experimental results have also shown that introducing ingredient information into these two networks can improve the performance of few-shot food recognition. Shuqiang Jiang, Weiqing Min, Yongqiang Lyu 0002, Linhu Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Ingredient-Guided Cascaded Multi-Attention Network for Food RecognitionabstractRecently, food recognition is gaining more attention in the multimedia community due to its various applications, e.g., multimodal foodlog and personalized healthcare. Most of existing methods directly extract visual features of the whole image using popular deep networks for food recognition without considering its own characteristics. Compared with other types of object images, food images generally do not exhibit distinctive spatial arrangement and common semantic patterns, and thus are very hard to capture discriminative information. In this work, we achieve food recognition by developing an Ingredient-Guided Cascaded Multi-Attention Network (IG-CMAN), which is capable of sequentially localizing multiple informative image regions with multi-scale from category-level to ingredient-level guidance in a coarse-to-fine manner. At the first level, IG-CMAN generates the initial attentional region from the category-supervised network with Spatial Transformer (ST). Taking this localized attentional region as the reference, IG-CMAN combined ST with LSTM to sequentially discover diverse attentional regions with fine-grained scales from ingredient-guided sub-network in the following levels. Furthermore, we introduce a new dataset ISIA Food-200 with 200 food categories from the list in the Wikipedia, about 200,000 food images and 319 ingredients. We conducted extensive experiment on two popular food datasets and newly proposed ISIA Food-200, and verified the effectiveness of our method. Qualitative results along with visualization further show that IG-CMAN can introduce the explainability for localized regions, and is able to learn relevant regions for ingredients. Weiqing Min, Linhu Liu, Zhengdong Luo, Shuqiang Jiang |
ACM Multimedia | 1 |
| 2019 | Session details: Multimedia SearchabstractNo abstract available. Weiqing Min |
MMAsia | 1 |
| 2019 | Hybrid incremental learning of new data and new classes for hand-held object recognition
Chengpeng Chen, Weiqing Min, Shuqiang Jiang |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Instance-level object retrieval via deep region CNN
Shuhuan Mei, Weiqing Min, Hua Duan, Shuqiang Jiang |
Multim. Tools Appl. | 2 |
| 2019 | Hierarchy-Dependent Cross-Platform Multi-View Feature Learning for Venue Category PredictionabstractIn this paper, we focus on visual venue category prediction, which can facilitate various applications for location-based service and personalization. Considering the complementarity of different media platforms, it is reasonable to leverage venue-relevant media data from different platforms to boost the prediction performance. Intuitively, recognizing one venue category involves multiple semantic cues, especially objects and scenes and, thus, they should contribute together to venue category prediction. In addition, these venues can be organized in a natural hierarchical structure, which provides prior knowledge to guide venue category estimation. Taking these aspects into account, we propose a Hierarchy-dependent Cross-platform Multi-view Feature Learning (HCM-FL) framework for venue category prediction from videos by leveraging images from other platforms. HCM-FL includes two major components, namely Cross-Platform Transfer Deep Learning (CPTDL) and Multi-View Feature Learning with the Hierarchical Venue Structure (MVFL-HVS). CPTDL is capable of reinforcing the learned deep network from videos using images from other platforms. Specifically, CPTDL first trained a deep network using videos. These images from other platforms are filtered by the learnt network and these selected images are then fed into this learnt network to enhance it. Two kinds of pre-trained networks on the ImageNet and Places dataset are employed. Therefore, we can harness both object-oriented and scene-oriented deep features through these enhanced deep networks. MVFL-HVS is then developed to enable multi-view feature fusion. It is capable of embedding the hierarchical structure ontology to support more discriminative joint feature learning. We conduct the experiment on videos from Vine and images from Foursquare. These experimental results demonstrate the advantage of our proposed framework in jointly utilizing multi-platform data, multi-view deep features, and hierarchical venue structure knowledge. Shuqiang Jiang, Weiqing Min, Shuhuan Mei |
IEEE Trans. Multim. | 2 |
| 2018 | You Are What You Eat: Exploring Rich Recipe Information for Cross-Region Food AnalysisabstractCuisine is a style of cooking and usually associated with a specific geographic region. Recipes from different cuisines shared on the web are an indicator of culinary cultures in different countries. Therefore, analysis of these recipes can lead to deep understanding of food from the cultural perspective. In this paper, we perform the first cross-region recipe analysis by jointly using the recipe ingredients, food images, and attributes such as the cuisine and course (e.g., main dish and dessert). For that solution, we propose a culinary culture analysis framework to discover the topics of ingredient bases and visualize them to enable various applications. We first propose a probabilistic topic model to discover cuisine-course specific topics. The manifold ranking method is then utilized to incorporate deep visual features to retrieve food images for topic visualization. At last, we applied the topic modeling and visualization method for three applications: 1) multimodal cuisine summarization with both recipe ingredients and images, 2) cuisine-course pattern analysis including topic-specific cuisine distribution and cuisine-specific course distribution of topics, and 3) cuisine recommendation for both cuisine-oriented and ingredient-oriented queries. Through these three applications, we can analyze the culinary cultures at both macro and micro levels. We conduct the experiment on a recipe database Yummly-66K with 66,615 recipes from 10 cuisines in Yummly. Qualitative and quantitative evaluation results have validated the effectiveness of topic modeling and visualization, and demonstrated the advantage of the framework in utilizing rich recipe information to analyze and interpret the culinary cultures from different regions. Weiqing Min, Bing-Kun Bao, Shuhuan Mei, Yong Rui, Shuqiang Jiang |
IEEE Trans. Multim. | 1 |
| 2017 | Dual Track Multimodal Automatic Learning through Human-Robot InteractionabstractHuman beings are constantly improving their cognitive ability via automatic learning from the interaction with the environment. Two important aspects of automatic learning are the visual perception and knowledge acquisition. The fusion of these two aspects is vital for improving the intelligence and interaction performance of robots. Many automatic knowledge extraction and recognition methods have been widely studied. However, little work focuses on integrating automatic knowledge extraction and recognition into a unified framework to enable jointly visual perception and knowledge acquisition. To solve this problem, we propose a Dual Track Multimodal Automatic Learning (DTMAL) system, which consists of two components: Hybrid Incremental Learning (HIL) from the vision track and Multimodal Knowledge Extraction (MKE) from the knowledge track. HIL can incrementally improve recognition ability of the system by learning new object samples and new object concepts. MKE is capable of constructing and updating the multimodal knowledge items based on the recognized new objects from HIL and other knowledge by exploring the multimodal signals. The fusion of the two tracks is a mutual promotion process and jointly devote to the dual track learning. We have conducted the experiments through human-machine interaction and the experimental results validated the effectiveness of our proposed system. Shuqiang Jiang, Weiqing Min, Huayang Wang, Jiaqi Zhou 0001 |
IJCAI | 2 |
| 2017 | A Delicious Recipe Analysis Framework for Exploring Multi-Modal Recipes with Various AttributesabstractHuman beings have developed a diverse food culture. Many factors like ingredients, visual appearance, courses (e.g., breakfast and lunch), flavor and geographical regions affect our food perception and choice. In this work, we focus on multi-dimensional food analysis based on these food factors to benefit various applications like summary and recommendation. For that solution, we propose a delicious recipe analysis framework to incorporate various types of continuous and discrete attribute features and multi-modal information from recipes. First, we develop a Multi-Attribute Theme Modeling (MATM) method, which can incorporate arbitrary types of attribute features to jointly model them and the textual content. We then utilize a multi-modal embedding method to build the correlation between the learned textual theme features from MATM and visual features from the deep learning network. By learning attribute-theme relations and multi-modal correlation, we are able to fulfill different applications, including (1) flavor analysis and comparison for better understanding the flavor patterns from different dimensions, such as the region and course, (2) region-oriented multi-dimensional food summary with both multi-modal and multi-attribute information and (3) multi-attribute oriented recipe recommendation. Furthermore, our proposed framework is flexible and enables easy incorporation of arbitrary types of attributes and modalities. Qualitative and quantitative evaluation results have validated the effectiveness of the proposed method and framework on the collected Yummly dataset. Weiqing Min, Shuqiang Jiang, Shuhui Wang, Shuhuan Mei |
ACM Multimedia | 1 |
| 2017 | A survey on context-aware mobile visual recognition
Weiqing Min, Shuqiang Jiang, Shuhui Wang, Ruihan Xu 0001, Yushan Cao, Luis Herranz, Zhiqiang He 0002 |
Multim. Syst. | 1 |
| 2017 | Being a Supercook: Joint Food Attributes and Multimodal Content Modeling for Recipe Retrieval and ExplorationabstractThis paper considers the problem of recipe-oriented image-ingredient correlation learning with multi-attributes for recipe retrieval and exploration. Existing methods mainly focus on food visual information for recognition while we model visual information, textual content (e.g., ingredients), and attributes (e.g., cuisine and course) together to solve extended recipe-oriented problems, such as multimodal cuisine classification and attribute-enhanced food image retrieval. As a solution, we propose a multimodal multitask deep belief network ($\mathrm{M}^{3}$TDBN) to learn joint image-ingredient representation regularized by different attributes. By grouping ingredients into visible ingredients (which are visible in the food image, e.g., “chicken” and “mushroom”) and nonvisible ingredients (e.g., “salt” and “oil”),$\mathrm{M}^{3}$TDBN is capable of learning both midlevel visual representation between images and visible ingredients and nonvisual representation. Furthermore, in order to utilize different attributes to improve the intermodality correlation,$\mathrm{M}^{3}$TDBN incorporates multitask learning to make different attributes collaborate each other. Based on the proposed$\mathrm{M}^{3}$TDBN, we exploit the derived deep features and the discovered correlations for three extended novel applications: 1) multimodal cuisine classification; 2) attribute-augmented cross-modal recipe image retrieval; and 3) ingredient and attribute inference from food images. The proposed approach is evaluated on the constructed Yummly dataset and the evaluation results have validated the effectiveness of the proposed approach. Weiqing Min, Shuqiang Jiang, Huayang Wang, Xinda Liu, Luis Herranz |
IEEE Trans. Multim. | 1 |
| 2016 | An incremental probabilistic model for temporal theme analysis of landmarks
Weiqing Min, Bing-Kun Bao, Changsheng Xu |
Multim. Syst. | 1 |
| 2015 | Joint Local and Global Consistency on Interdocument and Interword Relationships for Co-ClusteringabstractCo-clustering has recently received a lot of attention due to its effectiveness in simultaneously partitioning words and documents by exploiting the relationships between them. However, most of the existing co-clustering methods neglect or only partially reveal the interword and interdocument relationships. To fully utilize those relationships, the local and global consistencies on both word and document spaces need to be considered, respectively. Local consistency indicates that the label of a word/document can be predicted from its neighbors, while global consistency enforces a smoothness constraint on words/documents labels over the whole data manifold. In this paper, we propose a novel co-clustering method, called co-clustering via local and global consistency, to not only make use of the relationship between word and document, but also jointly explore the local and global consistency on both word and document spaces, respectively. The proposed method has the following characteristics: 1) the word-document relationships is modeled by following information-theoretic co-clustering (ITCC); 2) the local consistency on both interword and interdocument relationships is revealed by a local predictor; and 3) the global consistency on both interword and interdocument relationships is explored by a global smoothness regularization. All the fitting errors from these three-folds are finally integrated together to formulate an objective function, which is iteratively optimized by a convergence provable updating procedure. The extensive experiments on two benchmark document datasets validate the effectiveness of the proposed co-clustering method. Bing-Kun Bao, Weiqing Min, Teng Li 0001, Changsheng Xu |
IEEE Trans. Cybern. | 2 |
| 2015 | Cross-Platform Multi-Modal Topic Modeling for Personalized Inter-Platform RecommendationabstractIn this paper, we investigate a novel cross- platform multimedia problem: given two platforms, Flickr and Foursquare, we conduct the recommendation between these two platforms, namely the photo recommendation from Flickr to Foursquare users and the venue recommendation from Foursquare to Flickr users. Such inter-platform recommendations enable users from one single platform to enjoy different recommendation services effectively . To solve the problem, we propose a cross- platform multi-modal topic model ( CM3TM), which is capable of: 1) differentiating between two kinds of topics, i.e., platform- specific topics only relevant to a certain platform and shared topics characterizing the knowledge shared by different platforms and 2) aligning multiple modalities from different platforms. Specifically, CM3TM can not only split the topic space into the shared topic space and platform-specific topic space and learn them simultaneously, but also enable the alignment among different modalities through the learned topic space. Given the location information, we applied the proposed CM3TM into two inter-platform recommendation applications: 1) personalized venue recommendation from Foursquare to Flickr users and 2) personalized image recommendation from Flickr to Foursquare users. We have conducted experiments on the collected large-scale real-world dataset from Flickr and Foursquare. Qualitative and quantitative evaluation results validate the effectiveness of our method and demonstrate the advantage of connecting different platforms with different modalities for the inter-platform recommendation. Weiqing Min, Bing-Kun Bao, Changsheng Xu, M. Shamim Hossain |
IEEE Trans. Multim. | 1 |
| 2015 | Cross-Platform Emerging Topic Detection and Elaboration from Multimedia StreamsabstractWith the explosive growth of online media platforms in recent years, it becomes more and more attractive to provide users a solution of emerging topic detection and elaboration. And this posts a real challenge to both industrial and academic researchers because of the overwhelming information available in multiple modalities and with large outlier noises. This article provides a method on emerging topic detection and elaboration using multimedia streams cross different online platforms. Specifically, Twitter, New York Times and Flickr are selected for the work to represent the microblog, news portal and imaging sharing platforms. The emerging keywords of Twitter are firstly extracted using aging theory. Then, to overcome the nature of short length message in microblog, Robust Cross-Platform Multimedia Co-Clustering (RCPMM-CC) is proposed to detect emerging topics with three novelties: 1) The data from different media platforms are in multimodalities; 2) The coclustering is processed based on a pairwise correlated structure, in which the involved three media platforms are pairwise dependent; 3) The noninformative samples are automatically pruned away at the same time of coclustering. In the last step of cross-platform elaboration, we enrich each emerging topic with the samples from New York Times and Flickr by computing the implicit links between social topics and samples from selected news and Flickr image clusters, which are obtained by RCPMM-CC. Qualitative and quantitative evaluation results demonstrate the effectiveness of our method. Bing-Kun Bao, Changsheng Xu, Weiqing Min, M. Shamim Hossain |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2014 | Scene and viewpoint based visual summarization for landmarksabstractVisual summarization of landmarks is an important task for applications, such as landmark organization, search and browsing. In this work, we make the first attempt towards landmark summarization by simultaneously considering both the scenes (e.g., sunny view and night view) and viewpoints (e.g., front-side and close-distant viewpoint). In the proposed framework of landmark summarization, we first group images into different clusters by viewpoints, then the distinctive scenes for each viewpoint cluster are discovered by the proposed scene-viewpoint based theme modeling. Compared with the existing topic models, our model is capable of mining scene-viewpoint themes directly from all viewpoint clusters and meanwhile differentiating among these themes by viewpoints. The landmark summary is generated by the discovered scene-viewpoint themes, where each theme is represented by the selected images with one certain scene and viewpoint. The experimental results validate the proposed method and demonstrate its advantage in improving user experience. Weiqing Min, Bing-Kun Bao, Changsheng Xu |
ICIP | 1 |
| 2014 | Mobile Landmark Search with 3D ModelsabstractLandmark search is crucial to improve the quality of travel experience. Smart phones make it possible to search landmarks anytime and anywhere. Most of the existing work computes image features on smart phones locally after taking a landmark image. Compared with sending original image to the remote server, sending computed features saves network bandwidth and consequently makes sending process fast. However, this scheme would be restricted by the limitations of phone battery power and computational ability. In this paper, we propose to send compressed (low resolution) images to remote server instead of computing image features locally for landmark recognition and search. To this end, a robust 3D model based method is proposed to recognize query images with corresponding landmarks. Using the proposed method, images with low resolution can be recognized accurately, even though images only contain a small part of the landmark or are taken under various conditions of lighting, zoom, occlusions and different viewpoints. In order to provide an attractive landmark search result, a 3D texture model is generated to respond to a landmark query. The proposed search approach, which opens up a new direction, starts from a 2D compressed image query input and ends with a 3D model search result. Weiqing Min, Changsheng Xu, Min Xu 0001, Xian Xiao, Bing-Kun Bao |
IEEE Trans. Multim. | 1 |
| 2013 | Social event detection with robust high-order co-clusteringabstractThis paper is devoted to detecting social, real-world events from the sharing images/videos on social media sites like Flickr and YouTube. The fast growing contents make the social media sites become gold mines for social event detection, but we still need to overcome the challenge of processing the associated heterogeneous metadata, such as time-stamp, location, visual content and textual content. Different from the traditional early or late fusion with different types of metadata, we represent them into a star-structured $K$-partite graph, that is, social media itself is regarded as the central vertices set and different types of metadata are treated as the auxiliary vertices sets which are pairwise independent with each other but correlated with the central one. Based on this graph, Social Event Detection with Robust High-Order Co-Clustering (SED-RHOCC) algorithm is proposed and it includes two steps: 1) coarse event detection, 2) clusters and samples refinement. In the first step, by revealing the inter-relationship on the constructed star-structured $K$-partite graph and the intra-relationship within some metadata sets such as time-stamp, we co-cluster social media and the associated metadata separately and iteratively to avoid information loss in early/late fusion. After that, a post process is utilized to refine the clusters and social media samples in the second step. MediaEval Social Event Detection Dataset [1] and its subset are selected to demonstrate the effectiveness of our proposed approach in handling the datasets with and without non-event samples. Bing-Kun Bao, Weiqing Min, Ke Lu 0002, Changsheng Xu |
ICMR | 2 |
| 2013 | Landmark History Visualization
Weiqing Min, Bing-Kun Bao, Changsheng Xu |
MMM (2) | 1 |
| 2013 | Script-to-Movie: A Computational Framework for Story Movie CompositionabstractTraditional movie production has always been a highly professional work that needs team collaboration, advanced devices and techniques, and vast time and money investment. These high threshold requirements not only prevent mass amateur enthusiasts entering this field, but also hinder professionals quickly previewing their conceived story plots. In this paper, we raise a novel application, named script-to-movie (S2M) composition, to automatically produce new movies from existing videos in accordance with user created script. Our motivation is to liberate producers from complex filming and editing operations, thereby people's story idea can be instantly converted into the vivid movie video. To support the novel “What You Dream Is What You See” (WYDIWYS) production mode, we first propose a hierarchical alignment method to automatically construct a video material database with detailed semantic description. Considering diverse story plots in user designed script, the database contains abundant video materials about different characters appearing in various time and places conditions. On this basis, the S2M composition is formulated as a constrained optimization problem, where semantic story plot and syntactic visual content are synthetically considered to identify a group of optimal video segments to narrate the user designed script story. Both quantitative and qualitative experimental results are reported to illustrate the effectiveness of the proposed S2M application. Chao Liang 0001, Changsheng Xu, Jian Cheng 0001, Weiqing Min, Hanqing Lu |
IEEE Trans. Multim. | 4 |
| 2012 | Multimedia news digger on emerging topics from social streamsabstractWith the overwhelming information from social media networks and news portals, it is crucial to provide users a complete package of visual and textual information with popular interests automatically. To this concern, we present a news detection and pushing system, called Me-Digger (Multimedia News Digger), which not only effectively detects emerging topics from social streams but also provides the corresponding information in multiple modalities. Me-digger is the first systematic effort to leverage three sources of data, that is, Twitter, Flickr and Google news, to output with vivid visual and textual contents on emerging topics. Enabled by a novel general-structured high-order co-clustering approach, it has a more accurate detection of emerging topics compared to the existing methods on micro-blog social streams. Bing-Kun Bao, Weiqing Min, Changsheng Xu |
ACM Multimedia | 2 |