Tianwei Yan 0001

dblp:246/5787 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0003-1912-8795ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RICA: Re-ranking with intra-modal and cross-modal alignment for text-based person search
Yu Bai 0022, Wentao Ma 0003, Shan Zhao 0002, Tianwei Yan 0001, Shezheng Song, Chengyu Wang 0008, Qian Wan 0007
Expert Syst. Appl.4
2026 LLaMA-MoT: A cost-effective framework for visual-linguistic instruction tuning based on multi-head adapters and chain-of-thought
Turdi Tohti, Wenpeng Hu, Tianwei Yan 0001, Shaohuang Wang, Askar Hamdulla
Expert Syst. Appl.4
2026 DeepMEL: A multi-agent collaboration framework for multimodal entity linking
abstract
Multimodal Entity Linking (MEL) aims to associate textual and visual mentions with entities in a multimodal knowledge graph. Despite its importance, current methods face challenges such as incomplete contextual information, coarse cross-modal fusion, and the difficulty of jointly large language models (LLMs) and large visual models (LVMs). To address these issues, we propose DeepMEL, a novel framework based on multi-agent collaborative reasoning, which achieves efficient alignment and disambiguation of textual and visual modalities through a role-specialized division strategy. DeepMEL integrates four specialized agents, namely Modal-Fuser, Candidate-Adapter, Entity-Clozer and Role-Orchestrator, to complete end-to-end cross-modal linking through specialized roles and dynamic coordination. DeepMEL adopts a dual-modal alignment path, and combines the fine-grained text semantics generated by the LLM with the structured image representation extracted by the LVM, significantly narrowing the modal gap. We design an adaptive iteration strategy, combines tool-based retrieval and semantic reasoning capabilities to dynamically optimize the candidate set and balance recall and precision. DeepMEL also unifies MEL tasks into a structured cloze prompt to reduce parsing complexity and enhance semantic comprehension. Extensive experiments on five public benchmark datasets demonstrate that DeepMEL achieves state-of-the-art performance, improving ACC by 1 %-57 %. Ablation studies verify the effectiveness of all modules.
Fang Wang 0011, Tianwei Yan 0001, Zonghao Yang, Minghao Hu 0001, Zhunchen Luo, Xiaoying Bai
Inf. Process. Manag.2
2025 M^3EL: A Multi-task Multi-topic Dataset for Multi-modal Entity Linking
abstract
Multi-modal Entity Linking (MEL) is a fundamental component for various downstream tasks. However, existing MEL datasets suffer from small scale, scarcity of topic types and limited coverage of tasks, making them incapable of effectively enhancing the entity linking capabilities of multi-modal models. To address these obstacles, we propose a dataset construction pipeline and publish M^3EL, a large-scale dataset for MEL. M^3EL includes 79,625 instances, covering 9 diverse multi-modal tasks, and 5 different topics. In addition, to further improve the model's adaptability to multi-modal tasks, We propose a modality-augmented training strategy. Utilizing M^3EL as a corpus, train the CLIP_ND model based on CLIP (ViT-B-32), and conduct a comparative analysis with an existing multi-modal baselines. Experimental results show that the existing models perform far below expectations (ACC of 49.4%-75.8%), After analysis, it was obtained that small dataset sizes, insufficient modality task coverage, and limited topic diversity resulted in poor generalization of multi-modal models. Our dataset effectively addresses these issues, and the CLIP_ND model fine-tuned with M^3EL shows a significant improvement in accuracy, with an average improvement of 9.3% to 25% across various tasks. Our dataset publicly available to facilitate future research.
Fang Wang 0011, Shenglin Yin, Xiaoying Bai, Minghao Hu 0001, Tianwei Yan 0001
AAAI5
2025 MSACC: A Unified Multimodal Sentiment Analysis Framework for High Interpretability and Zero-shot Performance
abstract
Compared to large language models, traditional multimodal sentiment analysis frameworks are constrained by their classification heads, resulting in poor performance on zero-shot tasks. Moreover, due to limitations in visual encoders and multimodal fusion modules, most existing frameworks can only process a small number of images, leading to a loss of visual information. In light of these issues, this paper proposes a new framework, MSACC. This framework enhances the model’s zero-shot performance by adopting a contrastive classification method and reduces the loss of visual information through visual relation extraction and three-dimensional sentiment analysis. We conducted extensive experiments on the Yelp dataset. The experimental results show that MSACC outperforms models of the same category in zero-shot MSA tasks, achieving a 48% performance improvement. Furthermore, compared to the large language model ChatGLM2-6B, MSACC still achieved a 7% performance increase while saving 90% of the model size. In addition, in supervised tasks, MSACC also achieved a 3.27% performance improvement compared to the baseline model.
Turdi Tohti, Bo Kong 0002, Dongfang Han, Tianwei Yan 0001, Askar Hamdulla
ICASSP5
2025 DiffMEL: A large-scale difficulty-graded dataset for Multimodal Entity Linking
abstract
Multimodal Large Language Models (MLLMs) have shown tremendous potential in Multimodal Entity Linking (MEL). However, they are still far from achieving the expected effectiveness in practical applications. This could be due to limitations in the MEL dataset used for training. Existing MEL datasets primarily focus on simple tasks and only consider the direct matching of mentions with labeled entities within a multimodal context, ignoring mentions of unmatched entities. Factors such as the presence of the ground-truth entity within the candidate set and its position directly impact the performance of MLLMs on MEL tasks. To tackle these obstacles, we constructed DiffMEL, the first large-scale difficulty-graded dataset for MEL of MLLMs. DiffMEL contains 79,625 instances and 318.5K instance-related high-resolution images, covering 3 various difficulty graded linking tasks and 5 different entity themes. We utilize DiffMEL to train several open-source MLLMs. Experiment results demonstrate DiffMEL empowers MLLMs with stronger capabilities in MEL by a large-margin (5%-56.1%). our dataset is now available at https://github.com/ww-ffff/DiffMEL.
Fang Wang 0011, Xiaoying Bai, Tianwei Yan 0001, Minghao Hu 0001
ICASSP3
2025 Psychologically-Aware Retrieval-Augmented Generation for Coherent Role-Playing in LLMs
Pengyang Shao, Shan Zhao 0002, Shezheng Song, Tianwei Yan 0001, Chengyu Wang 0008
PRCV (4)5
2025 Hierarchical Label-Enhanced Contrastive Learning for Chinese NER
abstract
Recently, character-word lattice structures have achieved promising results for Chinese named entity recognition (NER), reducing word segmentation errors and increasing word boundary information for character sequences. However, constructing the lattice structure is complex and time-consuming, thus these lattice-based models usually suffer from low inference speed. Moreover, the quality of the lexicon affects the accuracy of the NER model. Since noise words can potentially confuse NER, limited coverage of the lexicon can cause lattice-based models to degenerate into partial character-based models. In this article, we propose a hierarchical label-enhanced contrastive learning (HLCL) method for Chinese NER. Instead of relying on the lattice structure, HLCL offers an alternative solution to robustly integrate entity boundary and type information with the help of both labels semantic and contrastive learning. HLCL is empowered by two techniques: 1) sentence-level contrastive learning (SCL) to model global mutual information between two different modalities (e.g., labels and sentences) and 2) token-level contrastive learning (TCL) to close the gap between representations of different characters (e.g., label-enhanced characters and original characters), resulting in local mutual information. With the well-designed contrastive learning scheme and the concise model during inference, HLCL can fully leverage the transferable label semantic and has a superb speed of inference. Experiments on four Chinese NER datasets show that HLCL obtains excellent efficiency as well as performance compared with existing lattice-based approaches.
Chengyu Wang 0008, Shan Zhao 0002, Tianwei Yan 0001, Shezheng Song, Wentao Ma 0003, Kuien Liu, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 HCL: A Hierarchical Contrastive Learning Framework for Zero-Shot Relation Extraction
abstract
Zero-shot relation extraction (ZSRE) is shown to become more significant in the current information extraction system, which aims at predicting relation classes that lack annotations or have just never appeared during training. Previous works focus on projecting sentences with their corresponding relation descriptions to an intermediate semantic space and searching the nearest semantic for predicting unseen classes. Though these methods can achieve sound performance, they only obtain inferior semantic information via a trivial distance metric and neglect the interaction in the instance representations. We are thus motivated to tackle these issues and propose a hierarchical contrastive learning (HCL) framework for ZSRE including projection-level and instance-level modules. Specifically, the projection-level component replaces the distance score function by contrastive loss to connect the input sentence with the relation semantic space. And the instance-level component integrates the external knowledge from sentence entities to establish new contrastive pairs for efficiently learning representations from mutual information. The experimental results on three well-known datasets demonstrate that our model surpasses the existing SOTA by at most 18.97% improvement on the F1 score when unseen classes are 15. Moreover, our model can achieve more competitive performance alone with the increasing number of unseen classes.
Tianwei Yan 0001, Shan Zhao 0002, Minghao Hu 0001, Mengzhu Wang, Xiang Zhang 0008, Zhigang Luo, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 FRCL-MNER: A Finer Grained Rank-Based Contrastive Learning Framework for Multimodal NER
abstract
Multimodal named entity recognition (MNER) is an emerging field that aims to automatically detect named entities and classify their categories, utilizing input text and auxiliary resources such as images. While previous studies have leveraged object detectors to preprocess images and fuse textual semantics with corresponding image features, these methods often overlook the potential finer grained information within each modality and may exacerbate error propagation due to predetection. To address these issues, we propose a finer grained rank-based contrastive learning (FRCL) framework for MNER. This framework employs a global-level contrastive learning to align multimodal semantic features and a Top-K rank-based mask strategy to construct positive-negative pairs, thereby learning a finer grained multimodal interaction representation. Experimental results from three well-known social media datasets reveal that our approach surpasses existing strong baselines, and achieves up to a 1.54% improvement on the Twitter2015 dataset. Extensive discussions further confirm the effectiveness of our approach. We will release the source code on https://github.com/augusyan/FRCL.
Tianwei Yan 0001, Shan Zhao 0002, Wentao Ma 0003, Shezheng Song, Chengyu Wang 0008, Zhibo Rao, Shizhao Chen, Zhigang Luo, Xinwang Liu 0002
IEEE Trans. Neural Networks Learn. Syst.1
2024 A Dual-Way Enhanced Framework from Text Matching Point of View for Multimodal Entity Linking
abstract
Multimodal Entity Linking (MEL) aims at linking ambiguous mentions with multimodal information to entity in Knowledge Graph (KG) such as Wikipedia, which plays a key role in many applications. However, existing methods suffer from shortcomings, including modality impurity such as noise in raw image and ambiguous textual entity representation, which puts obstacles to MEL. We formulate multimodal entity linking as a neural text matching problem where each multimodal information (text and image) is treated as a query, and the model learns the mapping from each query to the relevant entity from candidate entities. This paper introduces a dual-way enhanced (DWE) framework for MEL: (1) our model refines queries with multimodal data and addresses semantic gaps using cross-modal enhancers between text and image information. Besides, DWE innovatively leverages fine-grained image attributes, including facial characteristic and scene feature, to enhance and refine visual features. (2)By using Wikipedia descriptions, DWE enriches entity semantics and obtains more comprehensive textual representation, which reduces between textual representation and the entities in KG. Extensive experiments on three public benchmarks demonstrate that our method achieves state-of-the-art (SOTA) performance, indicating the superiority of our model. The code is released on https://github.com/season1blue/DWE.
Shezheng Song, Shan Zhao 0002, Chengyu Wang 0008, Tianwei Yan 0001, Shasha Li 0001, Xiaoguang Mao, Meng Wang 0001
AAAI4
2023 MCL: Multi-Granularity Contrastive Learning Framework for Chinese NER
abstract
Recently, researchers have applied the word-character lattice framework to integrated word information, which has become very popular for Chinese named entity recognition (NER). However, prior approaches fuse word information by different variants of encoders such as Lattice LSTM or Flat-Lattice Transformer, but are still not data-efficient indeed to fully grasp the depth interaction of cross-granularity and important word information from the lexicon. In this paper, we go beyond the typical lattice structure and propose a novel Multi-Granularity Contrastive Learning framework (MCL), that aims to optimize the inter-granularity distribution distance and emphasize the critical matched words in the lexicon. By carefully combining cross-granularity contrastive learning and bi-granularity contrastive learning, the network can explicitly leverage lexicon information on the initial lattice structure, and further provide more dense interactions of across-granularity, thus significantly improving model performance. Experiments on four Chinese NER datasets show that MCL obtains state-of-the-art results while considering model efficiency. The source code of the proposed method is publicly available at https://github.com/zs50910/MCL
Shan Zhao 0002, Chengyu Wang 0008, Minghao Hu 0001, Tianwei Yan 0001, Meng Wang 0001
AAAI4
2022 AAT: Non-local Networks for Sim-to-Real Adversarial Augmentation Transfer
Mengzhu Wang, Shanshan Wang 0008, Tianwei Yan 0001, Zhigang Luo
ICONIP (4)3
2019 A novel deep residual network-based incomplete information competition strategy for four-players Mahjong games
Tianwei Yan 0001, Mingyuan Luo, Wei Huang 0013
Multim. Tools Appl.2