Shizhou Huang

dblp:313/9699 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0004-2057-5271ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 CodeMNER: Vision-Language Models are Better Multimodal Named Entity Recognizers via Progressive Vision-Code Alignment
abstract
With the explosive growth of multimedia content on social media, Multimodal Named Entity Recognition (MNER) has garnered significant attention. However, current paradigms predominantly rely on general Vision-Language Models (VLMs) to generate natural language responses. Such unstructured text generation struggles to precisely articulate the complex structured information inherent in MNER tasks, often resulting in outputs that lack logical rigor and explicit structural constraints. To address these limitations, we propose CodeMNER, a novel framework that reformulates MNER tasks as a multimodal code generation problem. By synthesizing executable code instead of natural language, CodeMNER leverages the inherent syntactic rigor and deterministic executability of programming languages, thereby significantly enhancing the model’s capacity for identifying and classifying named entities. Despite the evident advantages of the code generation paradigm, standard VLMs lack the joint alignment between structured code semantics and natural visual representations, making it challenging to directly establish the mapping from visual contexts to executable code. To this end, we design a progressive four-stage training pipeline, encompassing mid-training, supervised fine-tuning, reinforcement learning with verifiable rewards, and downstream adaptation. This pipeline bridges the inherent vision-code alignment gap and augments model performance on MNER. Extensive experiments across standard Twitter-2015 and Twitter-2017 datasets demonstrate that CodeMNER achieves state-of-the-art performance, surpassing existing baselines.
Jiakang Yu 0001, Shizhou Huang, Xiaode Chen, Hongtao Deng, Wang Gao 0002
ICMR2
2025 GuideNER: Annotation Guidelines Are Better than Examples for In-Context Named Entity Recognition
abstract
Large language models (LLMs) demonstrate impressive performance on downstream tasks through in-context learning(ICL). However, there is a significant gap between their performance in Named Entity Recognition (NER) and in fine-tuning methods. We believe this discrepancy is due to inconsistencies in labeling definitions in NER. In addition, recent research indicates that LLMs do not learn the specific input-label mappings from the demonstrations. Therefore, we argue that using examples to implicitly capture the mapping between inputs and labels in in-context learning is not suitable for NER. Instead, it requires explicitly informing the model of the range of entities contained in the labels, such as annotation guidelines. In this paper, we propose GuideNER, which uses LLMs to summarize concise annotation guidelines as contextual information in ICL. We have conducted experiments on widely used NER datasets, and the experimental results indicate that our method can consistently and significantly outperform state-of-the-art methods, while using shorter prompts. Especially on the GENIA dataset, our model outperforms the previous state-of-the-art model by 12.63 F1 scores.
Shizhou Huang, Bo Xu 0023, Changqun Li, Xin Lin 0001
AAAI1
2025 A Graph Interaction Framework on Relevance for Multimodal Named Entity Recognition with Multiple Images
abstract
Posts containing multiple images have significant research potential in Multimodal Named Entity Recognition nowadays. The previous methods determine whether the images are related to named entities in the text through similarity computation, such as using CLIP. However, it is not effective in some cases and not conducive to task transfer, especially in multi-image scenarios. To address the issue, we propose a graph interaction framework on relevance (GIFR) for Multimodal Named Entity Recognition with multiple images. For humans, they have the abilities to distinguish whether an image is relevant to named entities, but human capabilities are difficult to model. Therefore, we propose using reinforcement learning based on human preference to integrate human abilities into the model to determine whether an image-text pair is relevant, which is referred to as relevance. To better leverage relevance, we construct a heterogeneous graph and introduce graph transformer to enable information interaction. Experiments on benchmark datasets demonstrate that our method achieves the state-of-the-art performance.
Shizhou Huang, Xin Lin 0001
COLING2
2025 Retrieval-Based Multimodal Data Augmentation for Multimodal Information Extraction in Social Media
Shizhou Huang, Bo Xu 0023, Changqun Li, Xin Lin 0001
DASFAA (4)1
2025 Low-Redundancy Knowledge Generation and Modality-Aware Interaction for Multimodal Information Extraction in Social Media
abstract
Multimodal information extraction (MIE) has gained increasing attention, as it helps to accomplish information extraction by adding images as auxiliary information. By acquiring entity-related knowledge, knowledge generation methods can effectively enhance the performance of information extraction models. However, current knowledge generation methods have two weaknesses: (1) they often generate knowledge that includes task-irrelevant information causing redundancy and negatively impacting model performance; (2) they typically concatenate knowledge and text input directly together, ignoring the stylistic and contextual differences arising from their different sources. To address these issues, we propose Low-Redundancy Knowledge Generation and Modality-Aware Interaction (LRKG-MAI). Our approach leverages a large language model to generate task-relevant knowledge with minimal redundancy, while treating knowledge as a distinct modality that interacts with text within its own representation space. Extensive experiments demonstrate the effectiveness of our approach. The source code can be found at https://github.com/JinFish/LRKG-MAI.
Shizhou Huang, Bo Xu 0023, Changqun Li, Xin Lin 0001
ICME1
2024 MNER-MI: A Multi-image Dataset for Multimodal Named Entity Recognition in Social Media
abstract
Recently, multimodal named entity recognition (MNER) has emerged as a vital research area within named entity recognition. However, current MNER datasets and methods are predominantly based on text and a single accompanying image, leaving a significant research gap in MNER scenarios involving multiple images. To address the critical research gap and enhance the scope of MNER for real-world applications, we propose a novel human-annotated MNER dataset with multiple images called MNER-MI. Additionally, we construct a dataset named MNER-MI-Plus, derived from MNER-MI, to ensure its generality and applicability. Based on these datasets, we establish a comprehensive set of strong and representative baselines and we further propose a simple temporal prompt model with multiple images to address the new challenges in multi-image scenarios. We have conducted extensive experiments to demonstrate that considering multiple images provides a significant improvement over a single image and can offer substantial benefits for MNER. Furthermore, our proposed method achieves state-of-the-art results on both MNER-MI and MNER-MI-Plus, demonstrating its effectiveness. The datasets and source code can be found at https://github.com/JinFish/MNER-MI.
Shizhou Huang, Bo Xu 0023, Changqun Li, Jiabo Ye, Xin Lin 0001
LREC/COLING1
2024 A Sentimental Prompt Framework with Visual Text Encoder for Multimodal Sentiment Analysis
abstract
Recently, multimodal sentiment analysis from social media posts has received increasing attention, as it can effectively improve single-modality-based sentiment analysis by leveraging the complementary information between text and images. Despite their success, current methods still suffer from two weaknesses: (1) the current methods for obtaining image representations do not obtain sentiment information, which leads to a significant gap between image representations and results; (2) the current methods ignore the sentiments expressed by the symbols (emoticons, emojis) in the text, but these symbols can effectively reflect the user's sentiments. To address these issues, we propose a sentimental prompt framework with visual text encoder (SPFVTE). Specifically, for the first problem, instead of using the image representation directly, we project the image representation as a prompt and utilize the prompt learning to capture sentimental information in images by learning a sentiment-specific prompt. For the second problem, considering that people get the meanings of emojis and emoticons from their graphics, we propose to render the text as an image and use a visual text encoder to capture the sentiments contained in emojis and emoticons. We have conducted experiments on three public multimodal sentiment datasets, and the experimental results show that our method can significantly and consistently outperform the state-of-the-art methods. The datasets and source code can be found at https://github.com/JinFish/SPFVTE.
Shizhou Huang, Bo Xu 0023, Changqun Li, Jiabo Ye, Xin Lin 0001
ICMR1
2023 A Unified Visual Prompt Tuning Framework with Mixture-of-Experts for Multimodal Information Extraction
Bo Xu 0023, Shizhou Huang, Ming Du 0002, Hongya Wang, Yanghua Xiao, Xin Lin 0001
DASFAA (3)2
2023 Knowledge Graph Enhanced Sentential Relation Extraction via Dual Heterogeneous Graph Context Selection
abstract
Sentential relation extraction is a type of relation extraction task whose goal is to extract semantic relations between entities from a single sentence. Compared with other variants of relation extraction, it often suffers from limitations of semantic contextual information. Due to the presence of knowledge graphs, many approaches propose to augment the semantics of sentences with the knowledge of entities, thus improving the performance of relation extraction. Despite their success, existing methods still suffer from two weaknesses: (1) existing approaches aggregate sentences, entities and their attribute values into a heterogeneous information graph, but do not consider the types of edges; (2) existing methods dynamically select knowledge based only on the structural features of the graph, without considering the features of the nodes themselves. To address these two problems, we propose a dual heterogeneous graph context selection method for knowledge graph enhanced sentential relation extraction. Specifically, to solve the first problem, we employ an edge-aware graph convolutional network to learn the representations of the heterogeneous graph with considering the types of edges. To solve the second problem, we propose dual graph context selection to select the useful context by considering the graph structure and node feature representation together. Experiments conducted on the Wikidata-RE dataset demonstrate the effectiveness of the method.
Bo Xu 0023, Luyi Cheng, Shizhou Huang, Shouang Wei, Ming Du 0002, Hongya Wang
IJCNN4
2023 HSimCSE: Improving Contrastive Learning of Unsupervised Sentence Representation with Adversarial Hard Positives and Dual Hard Negatives
abstract
Recently, contrastive learning (CL) has emerged as the fundamental framework for learning better sentence representations. In the unsupervised sentence representation task, due to the lack of labeled data, current CL-based approaches generally use various methods to generate or select positive and negative samples for the given sentence. Despite their success, existing CL-based unsupervised sentence representation methods underestimate hard positive samples and hard negative samples, which do not fully exploit the power of contrastive learning. In this paper, we argue that we need to focus more on hard positive and hard negative samples. To this end, we propose a novel contrastive learning model, HSimCSE, that extends SimCSE by considering both the hard positive and hard negative samples. Specifically, we first propose a novel adversarial positive sample generation module to generate an adversarial hard positive sample, then we propose a dual negative sample selection module to select hard negative samples from the in-batch samples and the entire training corpus. Finally, we propose a quadruplet loss to minimize the distance between the anchor sample and the adversarial hard positive sample and maximize the distance between the anchor sample and the two hard negative samples. Experiments conducted on seven semantic text similarity tasks demonstrate the effectiveness of our method. The source code can be found at https://github.com/xubodhu/HSimCSE.
Bo Xu 0023, Shouang Wei, Luyi Cheng, Shizhou Huang, Ming Du 0002, Hongya Wang
IJCNN4
2022 Different Data, Different Modalities! Reinforced Data Splitting for Effective Multimodal Information Extraction from Social Media Posts
abstract
Recently, multimodal information extraction from social media posts has gained increasing attention in the natural language processing community. Despite their success, current approaches overestimate the significance of images. In this paper, we argue that different social media posts should consider different modalities for multimodal information extraction. Multimodal models cannot always outperform unimodal models. Some posts are more suitable for the multimodal model, while others are more suitable for the unimodal model. Therefore, we propose a general data splitting strategy to divide the social media posts into two sets so that these two sets can achieve better performance under the information extraction models of the corresponding modalities. Specifically, for an information extraction task, we first propose a data discriminator that divides social media posts into a multimodal and a unimodal set. Then we feed these sets into the corresponding models. Finally, we combine the results of these two models to obtain the final extraction results. Due to the lack of explicit knowledge, we use reinforcement learning to train the data discriminator. Experiments on two different multimodal information extraction tasks demonstrate the effectiveness of our method. The source code of this paper can be found in https://github.com/xubodhu/RDS.
Bo Xu 0023, Shizhou Huang, Ming Du 0002, Hongya Wang, Chaofeng Sha, Yanghua Xiao
COLING2
2022 MAF: A General Matching and Alignment Framework for Multimodal Named Entity Recognition
abstract
In this paper, we study multimodal named entity recognition in social media posts. Existing works mainly focus on using a cross-modal attention mechanism to combine text representation with image representation. However, they still suffer from two weaknesses: (1) the current methods are based on a strong assumption that each text and its accompanying image are matched, and the image can be used to help identify named entities in the text. However, this assumption is not always true in real scenarios, and the strong assumption may reduce the recognition effect of theMNER model; (2) the current methods fail to construct a consistent representation to bridge the semantic gap between two modalities, which prevents the model from establishing a good connection between the text and image. To address these issues, we propose a general matching and alignment framework (MAF) for multimodal named entity recognition in social media posts. Specifically, to solve the first issue, we propose a novel cross-modal matching (CM) module to calculate the similarity score between text and image, and use the score to determine the proportion of visual information that should be retained. To solve the second issue, we propose a novel cross-modal alignment (CA) module to make the representations of the two modalities more consistent. We conduct extensive experiments, ablation studies, and case studies to demonstrate the effectiveness and efficiency of our method.The source code of this paper can be found in https://github.com/xubodhu/MAF.
Bo Xu 0023, Shizhou Huang, Chaofeng Sha, Hongya Wang
WSDM2