Dehua Ma

dblp:312/4171 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SCVQ: Sparse-Compensated Vector Quantization for Large Language Models
abstract
Zixuan Zhou, Yujun Diao, Zicheng Kong, Dehua Ma, Zhenbo Xu, Pei Pei Li, Zhaofeng He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yujun Diao, Zicheng Kong, Dehua Ma, Zhenbo Xu, Pei Pei Li, Zhaofeng He 0001
ACL (1)4
2026 Knowledge-guided policy arbitration: A hierarchical cognitive framework for safety-critical decision-making under dynamic conflicting objectives
Fuqing Bie, Xingyang Chang, Leyan Wang, Dehua Ma, Songfu Xu, Shuodi Liu, Yingzhuo Liu, Liuyu Xiang, Zhaofeng He 0001
Expert Syst. Appl.4
2026 BGACNet: boundary-guided cross-semantic attention cascade network for polyp segmentation
Dehua Ma, Mengkun Li
Multim. Syst.3
2025 Improving Food Recognition with Retrieval-Augmented and Domain-Adaptive LVLMs
abstract
Food recognition is pivotal in enhancing intelligent food recommendation systems and nutritional management, contributing to balanced diets and overall health. Although Large Vision-Language Models (LVLMs) have demonstrated impressive performances across various domains, their performance on the food recognition task still lags behind traditional vision models. To bridge this gap, this paper proposes two methods to improve the food recognition capabilities of LVLMs: Retrieval-Augmented Recognition (RAR) and Domain-Adaptive Recognition (DAR). On the one hand, the training-free RAR utilizes a vision model to retrieve relevant image-category pairs from an image-category memory pre-built from the training set, thus incorporating the categorical information into the input of LVLMs to enhance food recognition performance. On the other hand, DAR employs a two-stage training process by first pre-training LVLMs on diverse food analysis tasks and then fine-tuning LVLMs using food recognition data. Extensive evaluations on two large-scale food recognition datasets demonstrate that both RAR and DAR improve the food recognition performance of LVLMs and, compred to RAR, DAR achieves a higher precision that outperforms traditional vision models.
Dehua Ma, Zhenbo Xu, Tianshun Xing, Huijia Wu, Zhaofeng He 0001
ICASSP1
2025 FoodWeight1.4M: A Large-scale Multi-modal Dataset for Weight Estimation
abstract
Large vision language models (VLMs) excel in visual tasks but struggle with weight estimation, hindering 3D perception and embodied intelligence. To address the lack of large-scale weight datasets, we present FoodWeight1.4M, derived from real-world supermarket scenarios. It contains 1.4 million high-quality images across 1,550 food categories, with weights precisely measured and rigorously filtered, making it the first large-scale weight estimation dataset. The weight estimation performance of current VLMs were tested and found to be unsatisfactory, which can be significantly improved by instruction tuning using Food-Weight1.4M. Moreover, we propose two strategies, Category-Guided and Reference Calibration, to enhance weight estimation without fine-tuning. Experiments confirm their effectiveness in improving multi-modal weight perception. Furthermore, experimental results show that pre-training on FoodWeight1.4M can benefit other food analysis tasks. Our dataset will be publicly available soon.
Zhenbo Xu, Dehua Ma, Liuyu Xiang, Huijia Wu, Zhaofeng He 0001
ICME3
2025 RecipeRAG: Advancing Recipe Generation with Reinforced Retrieval Augmented Generation
abstract
Generating accurate recipes from dish images is a challenging task that requires a deep understanding of food categories, ingredient combinations, cooking methods, and context. Current works mainly rely on the two-stage training method or supervised fine-tuning of vision-language models (VLMs). Two-stage models typically first predict ingredients from images and then generate recipes based on both ingredients and images. However, accumulated errors in ingredient prediction often lead to inaccurate recipes. Fine-tuning VLMs only fit the statistical patterns of the training data, lacking deep reasoning capabilities, which leads to severe hallucinations in the generated recipes. In this paper, we introduce a novel reinforced retrieval-augmented generation framework named RecipeRAG for recipe generation, and compare the supervised fine-tuning (SFT) paradigm and the reinforcement fine-tuning (RFT) paradigm. To effectively retrieve recipes relevant to the query image, we improve CLIP to obtain IR-CLIP as both our retriever and re-ranker by integrating metric learning and contrastive learning. The retrieved recipes are then used to enhance the generated results, improving accuracy and reducing hallucinations. However, the SFT VLM often fails to judge the quality of the retrieved recipe information and perform the complex recipe generation. Therefore, we furthermore investigate the two-phase RFT training framework. Firstly, the cold-start phase uses generated Chain-of-Thought (CoT) data for SFT to activate the reasoning capabilities of VLMs. Then, the reinforcement learning phase utilizes Group Relative Policy Optimization (GRPO) to generate multiple reasoning-answer pairs, further enhancing the generalization ability of VLMs in recipe generation tasks. Extensive evaluations on the large-scale Recipe1M dataset demonstrate that RecipeRAG outperforms all previous methods in recipe generation and exhibits strong generalization ability under the RL paradigm.
Zhenbo Xu, Dehua Ma, Fei Liu 0008, Gong Huang, Zhaofeng He 0001
ACM Multimedia3
2025 RANet: A receptive aggregation network for polyp segmentation
Dehua Ma, Yanxiang Li, Wenzhe Meng, Siping Xu
Multim. Syst.1
2025 MCGFF-Net: a multi-scale context-aware and global feature fusion network for enhanced polyp and skin lesion segmentation
Yanxiang Li, Wenzhe Meng, Dehua Ma, Siping Xu
Vis. Comput.3