Shenzhen Huangfu

dblp:371/2604 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0001-5740-6744ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 62% Vision and language · 25% Trustworthy machine learning · 12%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
1.722025
MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation · ACM Multimedia 2025
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs · ICLR 2025
Natural language and speech › Language models and text generation
preference optimization
1.722025
MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation · ACM Multimedia 2025
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs · ICLR 2025
Natural language and speech › Language models and text generation
alignment
0.912025
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs · ICLR 2025
Machine learning › Trustworthy machine learning
hallucination
0.912025
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs · ICLR 2025
Computer vision › Vision and language
image captioning
0.912025
MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation · ACM Multimedia 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs · ICLR 2025

Methods — techniques the papers use, named apart from their topics

hierarchical preference optimization · 0.9direct preference optimization · 0.9cross-modal alignment · 0.9DPO · 0.9
YearPublicationVenuePosition
2025 CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
abstract
Multimodal Large Language Models (MLLMs) still struggle with hallucinations despite their impressive capabilities. Recent studies have attempted to mitigate this by applying Direct Preference Optimization (DPO) to multimodal scenarios using preference pairs from text-based responses. However, our analysis of representation distributions reveals that multimodal DPO struggles to align image and text representations and to distinguish between hallucinated and non-hallucinated descriptions. To address these challenges, In this work, we propose a Cross-modal Hierarchical Direct Preference Optimization (CHiP) to address these limitations. We introduce a visual preference optimization module within the DPO framework, enabling MLLMs to learn from both textual and visual preferences simultaneously. Furthermore, we propose a hierarchical textual preference optimization module that allows the model to capture preferences at multiple granular levels, including response, segment, and token levels. We evaluate CHiP through both quantitative and qualitative analyses, with results across multiple benchmarks demonstrating its effectiveness in reducing hallucinations. On the Object HalBench dataset, CHiP outperforms DPO in hallucination reduction, achieving improvements of 52.7% and 55.5% relative points based on the base model Muffin and LLaVA models, respectively. We make all our datasets and code publicly available.
Jinlan Fu, Shenzhen Huangfu, Hao Fei 0001, Xiaoyu Shen 0001, Bryan Hooi, Xipeng Qiu, See-Kiong Ng
ICLR2
2025 MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
Jinlan Fu, Shenzhen Huangfu, Hao Fei 0001, Yichong Huang, Xiaoyu Shen 0001, Xipeng Qiu, See-Kiong Ng
ACM Multimedia2