Hongtao Deng

dblp:277/1725 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CASA-RAG: Shared-Evidence Arbitration for Training-Free Retrieval-Augmented Generation
Hongtao Deng, Yinxia Lou
KSEM (5)3
2026 Query-Guided Conflict Inference and Incongruity-Aware Alignment for Implicit Hate Speech Detection in Videos
abstract
Implicit hate speech detection in videos is a complex multimodal task that aims to uncover malicious intent masked by coded language, sarcasm, and visual metaphors. However, existing state-of-the-art approaches predominantly rely on symmetric fusion paradigms driven by consistency-seeking objectives, which fundamentally fail to capture the structural conflict inherent in implicit hate; they inadvertently treat the defining cross-modal incongruities, such as a cheerful visual scene contradicting a malicious caption—as alignment noise to be smoothed out, rather than as critical signals to be amplified. To address this “Alignment Trap,” we propose the Temporal-Incongruity Hate Detection (TIHD) framework. Unlike symmetric approaches, TIHD introduces a Query-Guided Conflict Inference Network (QGC-Net) that leverages text as a semantic anchor to explicitly retrieve and amplify contradictory audio-visual features via a learned conflict gate. Furthermore, to capture transient hateful signals without frame-level supervision, we devise an Incongruity-Aware Alignment (IAA) module that performs differentiable soft-alignment scanning with adaptive temporal dynamics. Complemented by a two-stage learning strategy, TIHD effectively learns robust representations and precise decision boundaries. Extensive experiments on the ImpliHateVid and HateMM benchmarks demonstrate that TIHD achieves state-of-the-art performance, significantly outperforming existing baselines in unearthing implicit hate.
Jiakang Yu 0001, Hongtao Deng, Yinxia Lou
ICMR4
2026 CodeMNER: Vision-Language Models are Better Multimodal Named Entity Recognizers via Progressive Vision-Code Alignment
abstract
With the explosive growth of multimedia content on social media, Multimodal Named Entity Recognition (MNER) has garnered significant attention. However, current paradigms predominantly rely on general Vision-Language Models (VLMs) to generate natural language responses. Such unstructured text generation struggles to precisely articulate the complex structured information inherent in MNER tasks, often resulting in outputs that lack logical rigor and explicit structural constraints. To address these limitations, we propose CodeMNER, a novel framework that reformulates MNER tasks as a multimodal code generation problem. By synthesizing executable code instead of natural language, CodeMNER leverages the inherent syntactic rigor and deterministic executability of programming languages, thereby significantly enhancing the model’s capacity for identifying and classifying named entities. Despite the evident advantages of the code generation paradigm, standard VLMs lack the joint alignment between structured code semantics and natural visual representations, making it challenging to directly establish the mapping from visual contexts to executable code. To this end, we design a progressive four-stage training pipeline, encompassing mid-training, supervised fine-tuning, reinforcement learning with verifiable rewards, and downstream adaptation. This pipeline bridges the inherent vision-code alignment gap and augments model performance on MNER. Extensive experiments across standard Twitter-2015 and Twitter-2017 datasets demonstrate that CodeMNER achieves state-of-the-art performance, surpassing existing baselines.
Jiakang Yu 0001, Shizhou Huang, Xiaode Chen, Hongtao Deng, Wang Gao 0002
ICMR4
2025 AMCCL: Adaptive Multi-scale Convolution Fusion Network with Contrastive Learning for Multimodal Sentiment Analysis
Jiakang Yu 0001, Hongtao Deng, Wang Gao 0002
PRICAI (4)3
2025 Fractional Fourier-Enhanced Fusion Network Based on Pareto Optimization for Hyperspectral and LiDAR Data Classification
abstract
In recent years, the utilization of hyperspectral image (HSI) and light detection and ranging (LiDAR) for collaborative classification has emerged as a significant research direction in earth observation tasks, with diverse joint classification algorithms showing promising performance using varying network architectures. However, these methodologies infrequently address the challenge of fusion arising from the substantially larger volume of HSI feature information compared to LiDAR features. Moreover, the effective learning of HSI and LiDAR features while mitigating modality conflicts remains an area that necessitates further investigation. As such, a Fractional Fourier Enhanced Fusion Network based on Pareto Optimization (FrFENet) is proposed for HSI and LiDAR Data classification. To address the disparity in information volume between modalities, a weighted fractional Fourier enhanced fusion module (WFrFEF) is introduced, which applies a weighted fractional Fourier transform to HSI features, enhancing their representations and facilitating balanced fusion with LiDAR features. Furthermore, a Pareto-based soft optimization strategy, HLPareto, is designed to balance learning rates across HSI and LiDAR features in a dual-branch network, effectively avoiding optimization conflicts. Additionally, a spatial-spectral integration module (SSIM) and an elevation information enhancement module (EIEM) are developed to improve feature extraction. The SSIM enables effective spatial-spectral fusion by facilitating token-level interactions, while the EIEM enhances elevation feature representation, preserving spatial geometric information in LiDAR data. Extensive experiments and comparative analyses conducted on three widely utilized HSI and LiDAR datasets have shown that the proposed FrFENet exhibits superior classification performance.
Shou Feng, Hongtao Deng, Yabin Hu, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003
IEEE Trans. Geosci. Remote. Sens.2
2024 Event-centric hierarchical hyperbolic graph for multi-hop question answering over knowledge graphs
Wang Gao 0002, Wenguang Yao, Hongtao Deng
Eng. Appl. Artif. Intell.5
2024 Syntax-based argument correlation-enhanced end-to-end model for scientific relation extraction
Wang Gao 0002, Lang Zhang, Hongtao Deng
Neurocomputing5
2023 Generative non-autoregressive unsupervised keyphrase extraction with neural topic modeling
Yinxia Lou, Wang Gao 0002, Hongtao Deng
Eng. Appl. Artif. Intell.5
2022 Leveraging bilingual-view parallel translation for code-switched emotion detection with adversarial dual-channel encoder
Yinxia Lou, Hongtao Deng, Donghong Ji
Knowl. Based Syst.3