EDBT 2026 Demo / reviewers in the wild / expert
Chengcheng Mai
dblp:249/9108
· DBLP profile ↗
10ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0002-6185-8858ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MHier-RAG: Multi-modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-granularity Reasoning
Ziyu Gong 0001, Chengcheng Mai, Yihua Huang 0001 |
ICIC (28) | 2 |
| 2025 | KnowRA: Knowledge Retrieval Augmented Method for Document-level Relation Extraction with Comprehensive Reasoning AbilitiesabstractDocument-level relation extraction (Doc-RE) aims to extract relations between entities across multiple sentences. Therefore, Doc-RE requires more comprehensive reasoning abilities like humans, involving complex cross-sentence interactions between entities, contexts, and external general knowledge, compared to the sentence-level RE. However, most existing Doc-RE methods focus on optimizing single reasoning ability, but lack the ability to utilize external knowledge for comprehensive reasoning on long documents. To solve these problems, a knowledge retrieval augmented method, named KnowRA, was proposed with comprehensive reasoning to autonomously determine whether to accept external knowledge to assist Doc-RE. Firstly, we constructed a document graph for semantic encoding and integrated the co-reference resolution model to augment the co-reference reasoning ability. Then, we expanded the document graph into a document knowledge graph by retrieving the external knowledge base for common-sense reasoning and a novel knowledge filtration method was presented to filter out irrelevant knowledge. Finally, we proposed the axis attention mechanism to build direct and indirect associations with intermediary entities for achieving cross-sentence logical reasoning. Extensive experiments conducted on two datasets verified the effectiveness of our method compared to the state-of-the-art baselines. Our code is available at https://anonymous.4open.science/r/KnowRA. Chengcheng Mai, Yuxiang Wang 0012, Ziyu Gong 0001, Hanxiang Wang, Yihua Huang 0001 |
IJCAI | 1 |
| 2025 | PromptCNER: A Segmentation-based Method for Few-shot Chinese NER with Prompt-tuningabstractRecognizing Chinese entities in low-resource settings is a challenging but promising task, which extracts structured pre-defined entities and corresponding types from unstructured text. Compared with the prosperous Named Entity Recognition (NER) methods for Indo-European languages, such as English, the research on Chinese NER is still in its infancy. The main obstacles to the development of Chinese NER methods include the ambiguity of Chinese entity boundary recognition and limited data resources. To address these issues, in this paper, a word-segmentation-based model is present for few-shot Chinese NER. First, we enumerate all possible candidate entity spans on the character level for accurate entity boundary identification with the proposed word segmentation and combination strategy. Then, one kind of question-answer-based prompt template loaded with the candidate entity spans is proposed to cast entity extraction into the masked token prediction task, for dealing with the low-data problem by taking full advantage of the generality and transferability of the pre-trained language model. The extensive experimental results show that our method outperforms the state-of-the-art baselines in low-data settings and also achieves comparable performance in full-data settings. Chengcheng Mai, Ziyu Gong 0001, Hanxiang Wang, Mengchuan Qiu, Chunfeng Yuan, Yihua Huang 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2024 | AsCL: An Asymmetry-sensitive Contrastive Learning Method for Image-Text Retrieval with Cross-Modal FusionabstractThe image-text retrieval task aims to retrieve relevant information from a given image or text. The main challenge is to unify multimodal representation and distinguish fine-grained differences across modalities, thereby finding similar contents and filtering irrelevant contents. However, existing methods mainly focus on unified semantic representation and concept alignment for multi-modalities, while the fine-grained differences across modalities have rarely been studied before, making it difficult to solve the information asymmetry problem. In this paper, we propose a novel asymmetry-sensitive contrastive learning method. By generating corresponding positive and negative samples for different asymmetry types, our method can simultaneously ensure fine-grained semantic differentiation and unified semantic representation between multi-modalities. Additionally, a hierarchical cross-modal fusion method is proposed, which integrates global and local-level features through a multimodal attention mechanism to achieve concept alignment. Extensive experiments performed on MSCOCO and Flickr30K, demonstrate the effectiveness and superiority of our proposed method. Ziyu Gong 0001, Chengcheng Mai, Yihua Huang 0001 |
ICME | 2 |
| 2024 | ML2MG-VLCR: A Multimodal LLM Guided Zero-shot Method for Visio-linguistic Compositional Reasoning with Autoregressive Generative Language ModelabstractThe visio-linguistic compositional reasoning is an interesting but challenging task aimed at matching two images and two captions, where the two images are different but the two corresponding captions are composed of the same words but in different order. This requires the matching model to have the ability to understand both the composition structure of the image and the order of the description text. However, when faced with compositional reasoning tasks, existing vision-language models are not sensitive to the image structure and text order, acting more like bag-of-words models. To address this challenge, a zero-shot visio-linguistic compositional reasoning method was proposed with the assistance of multimodal LLM and autoregressive generative language model. Given an image and candidate texts with different order compositions, we first leveraged LLaVA to generate descriptive text according to the image, for reflecting the compositional structure of image into text order. Then, an order-sensitive image-text matching method was proposed by calculating the generation probability of the candidate text conditioned on the textualized image information obtained by LLaVA, where autoregressive generative language model explicitly plays an important role in order modeling and evaluating. Experimental results on VG-Relation, VG-Attribution and Flickr30K-Order, demonstrated the superiority of our method in understanding the compositional structure and order of images and texts. Ziyu Gong 0001, Chengcheng Mai, Yihua Huang 0001 |
ICMR | 2 |
| 2024 | iterPrompt: An iterative prompt-tuning method for nested relation extraction with dynamic assignment strategy
Chengcheng Mai, Yuxiang Wang 0012, Ziyu Gong 0001, Hanxiang Wang, Kaiwen Luo, Chunfeng Yuan, Yihua Huang 0001 |
Expert Syst. Appl. | 1 |
| 2023 | Nested relation extraction via self-contrastive learning guided by structure and semantic similarity
Chengcheng Mai, Kaiwen Luo, Yuxiang Wang 0012, Ziyan Peng, Chunfeng Yuan, Yihua Huang 0001 |
Neural Networks | 1 |
| 2022 | Pretraining Multi-modal Representations for Chinese NER Task with Cross-Modality AttentionabstractNamed Entity Recognition (NER) aims to identify the pre-defined entities from the unstructured text. Compared with English NER, Chinese NER faces more challenges: the ambiguity problem in entity boundary recognition due to unavailable explicit delimiters between Chinese characters, and the out-of-vocabulary (OOV) problem caused by rare Chinese characters. However, two important features specific to the Chinese language are ignored by previous studies: glyphs and phonetics, which contain rich semantic information of Chinese. To overcome these issues by exploiting the linguistic potential of Chinese as a logographic language, we present MPM-CNER (short for Multi-modal Pretraining Model for Chinese NER), a model for learning multi-modal representations of Chinese semantics, glyphs, and phonetics, via four pretraining tasks: Radical Consistency Identification (RCI), Glyph Image Classification (GIC), Phonetic Consistency Identification (PCI), and Phonetic Classification Modeling (PCM). Meanwhile, a novel cross-modality attention mechanism is proposed to fuse these multimodal features for further improvement. The experimental results show that our method outperforms the state-of-the-art baseline methods on four benchmark datasets, and the ablation study also verifies the effectiveness of the pre-trained multi-modal representations. Chengcheng Mai, Mengchuan Qiu, Kaiwen Luo, Ziyan Peng, Chunfeng Yuan, Yihua Huang 0001 |
WSDM | 1 |
| 2022 | Pronounce differently, mean differently: A multi-tagging-scheme learning method for Chinese NER integrated with lexicon and phonetic features
Chengcheng Mai, Mengchuan Qiu, Kaiwen Luo, Ziyan Peng, Chunfeng Yuan, Yihua Huang 0001 |
Inf. Process. Manag. | 1 |
| 2021 | TSSE-DMM: Topic Modeling for Short Texts Based on Topic Subdivision and Semantic Enhancement
Chengcheng Mai, Xueming Qiu, Kaiwen Luo, Bo Zhao 0029, Yihua Huang 0001 |
PAKDD (2) | 1 |