VLDB 2026 Research / reviewers in the wild / expert
Guorui Sheng
dblp:119/3189
· DBLP profile ↗
18ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0001-6790-0239ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RFHNet: Relational and Frequency-Aware Hashing Network for Large-Scale Fine-Grained Food Image RetrievalabstractFine-grained food image retrieval is a key task in computational gastronomy, with applications in food traceability, dietary monitoring, and smart catering systems. Although hashing-based retrieval is attractive for large-scale search due to its storage efficiency and fast Hamming-distance computation, existing methods often perform poorly in fine-grained food scenarios, where subtle local semantics and frequency-sensitive visual cues are essential. To address this challenge, we propose RFHNet, a cascaded hierarchical hashing network that captures both global structure and fine-grained local details through multi-level representations. RFHNet includes three components: (1) Fine-grained Relation Modeling (FRM) to capture subtle visual differences among similar food components; (2) Multi-Frequency Modulated Fusion (MFMF) to extract informative multi-frequency features; and (3) Hierarchical Semantic Synergy (HSS) to adaptively integrate multi-level representations and generate discriminative hash codes. Experiments on six food-specific benchmarks show that RFHNet consistently outperforms state-of-the-art hashing methods, with mAP gains of 4.44% to 17.20% at 12 bits. These results validate the effectiveness of RFHNet for large-scale visual food retrieval and smart catering applications. The source code will be released upon publication. Weiqing Min, Guorui Sheng, Lili Wang 0009, Shuqiang Jiang |
ICMR | 3 |
| 2026 | DPFA-net: a lightweight hybrid neural network with dual path feature aggregation for food image recognition
Xiangyi Zhu, Yingnan Sheng, Congrui Lv, Guorui Sheng, Weiqing Min, Shuqiang Jiang |
Multim. Syst. | 5 |
| 2026 | Cross-modal recipe retrieval via multi-granularity alignment
Runqi Zan, Yuxin Yu, Guorui Sheng, Zongchao Huang |
Neural Networks | 4 |
| 2026 | LLM-informed global-local contextualization for zero-shot food detection
Weiqing Min, Guorui Sheng, Jingru Song, Yancun Yang, Shuqiang Jiang |
Pattern Recognit. | 3 |
| 2026 | LIFR-Net: A lightweight hybrid neural network with feature grouping for efficient food image recognition
Qingshuo Sun, Guorui Sheng, Xiangyi Zhu, Jingru Song, Yongqiang Song, Lili Wang 0009 |
Pattern Recognit. Lett. | 2 |
| 2026 | SyMFood: Synergistic Multi-Modal Prompting for Fine-Grained Zero-Shot Food DetectionabstractFine-grained object detection in food computing is severely constrained by the vast diversity of food items and the high cost of data annotation. Existing Zero-Shot Food Detection (ZSFD) methods attempt to solve this by leveraging semantic information, but they suffer from two critical bottlenecks: (1) a "Semantic Dilemma" (SD) where textual descriptions are too ambiguous to distinguish visually similar food categories, and (2) an "Architectural Bottleneck" (AB) due to the granularity mismatch between high-level semantics and low-level visual features. In this article, we propose SyMFood (Synergistic Multi-modal Framework for Food Generalization), a novel ZSFD framework designed to systematically overcome these challenges. To resolve the SD, SyMFood employs a multi-modal prompt system, which combines rich descriptions from Large Language Models (LLMs) with unambiguous visual exemplars to provide precise semantic grounding. To break the AB, SyMFood introduces a “Refine-then-Fuse” architecture. This design first utilizes a Context-Aware Spatial-Channel Refinement (CaSC) block to enhance visual features independently. Subsequently, a Progressive Food Knowledge Fusion (ProFus) module performs bi-directional, iterative co-refinement between the enhanced visual features and multi-modal prompts across all scales. Extensive experiments across four challenging datasets, including food-specific (UEC FOOD 256, FOWA) and general-purpose (PASCAL VOC, MS COCO) benchmarks, validate our approach. The proposed method outperforms baselines, yielding a notable 8.5% improvement in Harmonic Mean on the FOWA dataset in the genearl ZSD (GZSD) setting. The source code will be available at https://github.com/Niko000202/SymFood0202. Weiqing Min, Shoulong Liu, Guorui Sheng, Shuqiang Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Correspondence Calibrating and Dynamic Consistency Learning for Noisy Cross-Modal RetrievalabstractCross-modal retrieval has drawn an increasing amount of attention due to its effective ability for searching semantic relative data points with different modalities. In spite of some progress obtained, such methods often require the data pair maintaining the correct cross-modal correspondence in the training process, which is impractical in real application. To tackle this issue, we propose a Correspondence Calibrating and Dynamic Consistency Learning Network (CCDCL), aiming at optimizing the correspondence of positive samples and deeply investigating the consistency of negative samples. Specifically, to effectively alleviate the false positive issues, co-teaching paradigm is introduced to optimize the correspondence of positive samples by calibrating their confidence scores. To address the false negative sample problem, we propose a Vision-Language Semantic Collaborative Dynamic Margin Adaptation (VLC-DMA) method, which integrates unimodal similarities with calibrated confidence to produce the cross-modal semantic similarities, and consequently employs a dynamic margin function for accurately discriminating false and true negative samples. Experimental results demonstrate that the proposed method effectively improves cross-modal retrieval performance across multiple image-text matching datasets. The code will be available on GitHub upon acceptance of this paper. Yizhen Wu, Linliang Zhang, Guorui Sheng, Ying Li 0016, Qi Tian 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | FoodHash: Context-Aware Proxy Interaction and Fusion for Food Image RetrievalabstractVision-based food image retrieval has garnered significant attention due to its potential for critical applications in dietary and health management. However, food images exhibit more complex feature distributions and lack the geometric regularity and structured patterns typically observed in general image retrieval tasks. This complexity poses a challenge for existing models to extract fine-grained features and semantic information, thereby compromising retrieval performance. To address this challenge, we propose FoodHash, a context-aware proxy interaction and fusion hashing method for food image retrieval. The method incorporates an Aggregation–Interaction–Propagation (AIP) module that facilitates contextual information exchange among patch tokens within the same feature map, guided by proxy tokens, thereby effectively capturing the intricate details of food images. Furthermore, to leverage the rich semantic information in food images, a Cross-Fusion Module is introduced to efficiently integrate multi-scale information and enhance feature representation. Additionally, we employ a novel loss function to optimize hash learning by ensuring consistency between hash codes and the semantic space, thereby enhancing the learning capability of hash coding. Extensive experiments on three publicly available food datasets demonstrate that FoodHash significantly surpasses existing models in retrieval performance. Specifically, on the ETH Food-101 dataset, FoodHash achieves improvements of 18.1%, 6.7%, 5.2%, and 4.5% over the suboptimal method PTLCH for 16-bits, 32-bits, 48-bits, and 64-bits hash codes, respectively. The source code will be made publicly available upon publication of the article. Pindan Cao, Weiqing Min, Guorui Sheng, Yongqiang Song, Lili Wang 0009, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Cross-Layer and Selective Distillation for Asymmetric Image Retrieval
Weiqing Min, Fangyuan Yao, Guorui Sheng, Shuqiang Jiang |
PRCV (12) | 4 |
| 2025 | SAHNet: A Spectral-Aware Hashing Network for Large-Scale Fine-Grained Food Image RetrievalabstractFine-grained food image retrieval is vital for applications such as dietary monitoring and personalized nutrition. While hashing-based methods are popular for their storage and computational efficiency, existing approaches often fail to exploit category-specific spectral-spatial features and tend to preserve redundant representations, thereby limiting their discriminative ability. To address these issues, we propose SAHNet, a hierarchical network with a multi-scale backbone and two key modules: Spectral-Spatial Information Mining (SSIM) for dual-branch spectral/spatial feature extraction, and Selective Feature Filtering (SFF) for enhancing critical cues while suppressing noise. SAHNet generates compact, discriminative hash codes and achieves state-of-the-art performance on four benchmark fine-grained food datasets. Yihan Wang 0023, Bolin Yang, Guorui Sheng, Lili Wang 0009 |
VINCI | 4 |
| 2025 | Channel grouping vision transformer for lightweight fruit and vegetable recognition
Chengxu Liu 0002, Weiqing Min, Jingru Song, Yancun Yang, Guorui Sheng, Lili Wang 0009, Shuqiang Jiang |
Expert Syst. Appl. | 5 |
| 2024 | Cross-modal Semantic Interference Suppression for image-text matching
Shouyong Peng, Yujuan Sun, Guorui Sheng, Haiyan Fu, Xiangwei Kong 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Lightweight Food Recognition via Aggregation Block and Feature EncodingabstractFood image recognition has recently been given considerable attention in the multimedia field in light of its possible implications on health. The characteristics of the dispersed distribution of ingredients in food images put forward higher requirements on the long-range information extraction ability of neural networks, leading to more complex and deeper models. Nevertheless, the lightweight version of food image recognition is essential for improved implementation on end devices and sustained server-side expansion. To address this issue, we present Aggregation Feature Net (AFNet), a lightweight network that is capable of effectively capturing both global and local features from food images. In AFNet, we develop a novel convolution based on a residual model by encoding global features through row-wise and column-wise information integration. Merging aggregation block with classic local convolution yields a framework that works as the backbone of the network. Based on the efficient use of parameters by the aggregation block, we constructed a lightweight food image recognition network with fewer layers and a smaller scale, assisted by a new type of activation function. Experimental results on four popular food recognition datasets demonstrate that our approach achieves state-of-the-art performance with higher accuracy and fewer FLOPs and parameters. For example, in comparison to the current state-of-the-art model of MobileViTv2, AFNet achieved 88.4% accuracy of the top-1 level on the ETHZ Food-101 dataset, with similar parameters and FLOPs but 1.4% more accuracy. The source code will be provided in supplementary materials. Yancun Yang, Weiqing Min, Jingru Song, Guorui Sheng, Lili Wang 0009, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Food recognition via an efficient neural network with transformer groupingabstractRecently, considerable research efforts have been devoted to food recognition for its great potential applications in human health. Much work so far has focused on directly extracted deep visual features via Convolutional Neural Networks, which require significant computational resources and training time. The high requirements on hardware resources severely limit the application of food recognition in mobile devices and the sustainable extension on the server side. Therefore, how to design an efficient and high-performance lightweight neural network for food recognition is the key to solve the problem. In this paper, we propose a Lightweight Transformer-Based Deep Neural Network for food image recognition, which can achieve effective recognition of food images with fewer parameters and lower computational cost. Through Transformer Grouping and Token Shuffling, we construct an efficient food image recognition network that effectively combines the advantages of Transformer to extract global features and MobileNet to extract local features. The proposed network architecture effectively copes with the particularly scattered distribution of salient features in food images, and improves the recognition rate. We conduct extensive experiments on three popular food data sets, demonstrating that our method achieves state-of-the-art performance in applying lightweight neural networks to food image recognition. Guorui Sheng, Shu-Qi Sun, Chengxu Liu 0002, Yancun Yang |
Int. J. Intell. Syst. | 1 |
| 2018 | Direct, Near Real Time Animation of a 3D Tongue Model Using Non-Invasive Ultrasound ImagesabstractA new technique for representing speech articulation with an ultrasound-driven finite element model of the tongue is presented. By using a snake contour extraction algorithm with anatomically motivated constraints and a common coordinate system between the ultrasound and the tongue model, it is possible for the first time to obtain a realistic 3D simulation of the tongue directly from a non-invasive sensor (ultrasound), without mapping through any intermediate sensor modalities, and at near real-time frame rates. Shicheng Chen, Chengrui Wu, Guorui Sheng, Pierre Roussel-Ragot, Bruce Denby |
ICASSP | 4 |
| 2018 | Predicting Tongue Motion in Unlabeled Ultrasound Video Using 3D Convolutional Neural NetworksabstractA 3-dimensional convolutional neural network is trained on unlabeled ultrasound video to predict an upcoming tongue image from previous ones. The network obtains results superior to those of simpler predictors and provides a starting point for exploiting the higher-level representation of the tongue learned by the system in a variety of applications in speech research. This work is believed to be the first application of convolutional neural networks to unlabeled ultrasound video for the purpose of predicting tongue movement. Chengrui Wu, Shicheng Chen, Guorui Sheng, Pierre Roussel-Ragot, Bruce Denby |
ICASSP | 3 |
| 2017 | Detection of content-aware image resizing based on Benford's law
Guorui Sheng, Tao Li 0043, Qingtang Su, Beijing Chen |
Soft Comput. | 1 |
| 2015 | Forensic Detection of Median Filtering in Digital Images Using the Coefficient-Pair Histogram of DCT Value and LBP Pattern
Yun-Ni Lai, Tiegang Gao, Guorui Sheng |
ICIC (1) | 4 |