Tongguan Wang

dblp:359/7646 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
19since 2021 · last 2027
0009-0001-8747-1790ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2027 DFM-NC: Disentangling fine-grained multiplex non-literal cues for multimodal sentiment analysis
Tongguan Wang, Feiyue Xue, Junkai Li, Xiqiao Ba, Yang Xiao 0018, Ying Sha
Expert Syst. Appl.1
2027 SIPTrack: Reliability-aware identity prediction for sparse-interval pig multi-object tracking with a new benchmark
Feiyue Xue, Wangjun Huang, Junkai Li, Tongguan Wang, Huaiping Jin, Ying Sha
Expert Syst. Appl.4
2026 Chinese Two-part Allegorical Sayings Reading Comprehension: Exploration from Reasoning to Metaphor
abstract
The Two-Part Allegorical Saying (TPAS) is a Chinese linguistic phenomenon with a riddle-explanation structure, and an important component of Chinese metaphors. Existing research has primarily used TPAS to assist other semantic tasks, but lacks in-depth exploration of its intrinsic mechanisms: semantic rhetoric, logical reasoning, and metaphorical expression. To address this gap, we construct the first Chinese TPAS Reading Comprehension dataset (CTRC), which contains 18,103 TPASs and 75,296 passages. We frame it as a cloze test where the model selects the most suitable TPAS from candidates to fill passage blanks. To tackle the challenges of this CTRC task, we propose a Multi-view TPAS Contrastive Learning Network (MTCLN). Firstly, the joint vector cross-projection module extracts the rhetorical features of TPAS, such as homophonic puns, through vector space mapping to mitigate the semantic deviations caused by rhetoric. Then, the softened contrastive learning module strengthens the modeling of TPAS logical reasoning through feature association. Finally, the multi-view feature fusion module integrates contextual semantics with diverse TPAS features to facilitate the understanding of metaphorical expressions. Experiments on the CTRC dataset demonstrate that MTCLN achieves an average accuracy of 67.47%, outperforming large language models by 25.48%.
Dongyu Su, Yimin Xiao, Tongguan Wang, Feiyue Xue, Junkai Li, Ying Sha
AAAI3
2026 Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cues
Tongguan Wang, Feiyue Xue, Junkai Li, Ying Sha
WWW2
2026 MePe: Rethinking Multimodal Chinese Idiom Reading Comprehension from a Metaphorical Perspective
abstract
The multimodal Chinese idiom reading comprehension task aims to select the most appropriate idiom from a candidate list via the given text and image. This poses a significant challenge for the model to comprehend each Chinese idiom accurately. Existing multimodal Chinese idiom reading comprehension methods primarily focus on aligning contextual text and images, while overlooking two key attributes of Chinese idioms.(1) There is a discrepancy between the literal and metaphorical meanings of Chinese idioms. (2) The same Chinese idiom has different meanings in different scenarios, which requires targeted understanding by experts who specialize in different fields. To address the above challenges, we rethink the solution to the multimodal idiom reading comprehension task from a metaphorical perspective and propose a framework named MePe. Firstly, we propose a literal metaphorical semantic graph that systematically transforms the implicit discrepancy between the literal and metaphorical meanings of Chinese idioms into structured explicit relationships, thereby making metaphorical meanings more understandable. Then, we propose a mixture of idiom experts consisting of a literal idiom expert and a metaphorical idiom expert. Through division of labor and collaboration among these experts, we achieve an understanding of the dual meanings of Chinese idioms across different scenarios. Finally, we employ the maximum mean discrepancy to adjust the variance between the literal and metaphorical semantic features of Chinese idioms. By mapping these features onto a shared reproducing kernel Hilbert space, the model can better distinguish between the two based on contextual clues. Extensive experiments demonstrate that MePe achieves state-of-the-art performance on the MChIRC dataset.
Tongguan Wang, Junkai Li, Feiyue Xue, Dongyu Su, Wangjun Huang, Ying Sha
WWW1
2026 SCA-Net: Semantic text-enhanced context-aware multimodal framework for fish feeding assessment in aquaculture
Junkai Li, Feiyue Xue, Tongguan Wang, Chunfang Wang, Zongyao Sha, Ying Sha
Expert Syst. Appl.3
2026 A dual-level indicative representation learning method for multimodal sarcasm detection
Yongcheng Zhang, Guixin Su, Tongguan Wang, Mingmin Wu, Xiaomei Wei 0001
Inf. Process. Manag.3
2026 BIG-TM: Bridging Individual Guidance with Trifusion MoPoE for Chinese memes understanding
Tongguan Wang, Junkai Li, Feiyue Xue, Dongyu Su, Guixin Su, Xiaopeng Wen, Ying Sha
Knowl. Based Syst.1
2025 McHirc: A Multimodal Benchmark for Chinese Idiom Reading Comprehension
abstract
The performance of various tasks of natural language processing has greatly improved with the emergence of large language models. However, there is still much room for improvement in understanding certain specific linguistic phenomena, such as Chinese idioms, which are usually composed of four characters. Chinese idioms are difficult to understand due to semantic gaps between their literal and actual meanings. Researchers have proposed the Chinese idiom reading comprehension task to examine the ability of large language models to represent and understand Chinese idioms. The task requires choosing the correct Chinese idiom from a list of candidates to complete the sentence. The current research mainly focuses on text-based idiom comprehension. Nevertheless, there are many idiom application scenarios that combine images and text, and we believe that the corresponding images are beneficial for the model's understanding of the idioms. Therefore, to address the above problems, we first construct a large-scale Multimodal Chinese Idiom Reading Comprehension dataset (MChIRC), which contains a total of 44,433 image-text pairs covering 2,926 idioms. Then, we propose a Dual-Contrastive Idiom Graph Network (DCIGN), which employs a dual-contrastive learning module to align the text and image features corresponding to the same Chinese idiom at both coarse and fine levels, while utilizing a graph structure to capture the semantic relationships between idiom candidates. Finally, we use a cross-attention module to fuse multimodal features with graph features of candidate idioms to predict correct answers. The authoritativeness of MChIRC and the effectiveness of DCIGN are demonstrated through a variety of experiments, which provides a new benchmark for the multimodal Chinese idiom reading comprehension task.
Tongguan Wang, Mingmin Wu, Guixin Su, Dongyu Su, Yuxue Hu, Zhongqiang Huang, Ying Sha
AAAI1
2025 Unified Grid Tagging Scheme for Aspect Sentiment Quad Prediction
abstract
Aspect Sentiment Quad Prediction (ASQP) aims to extract all sentiment elements in quads for a given review to explain the reason for the sentiment. Previous table-filling based methods have achieved promising results by modeling word-pair relations. However, these methods decompose the ASQP task into several subtasks without considering the association between sentiment elements. Most importantly, they fail to tackle the situation where a sentence contains multiple implicit expressions. To address these limitations, we propose a simple yet effective Unified Grid Tagging Scheme (UGTS) to extract sentiment quadruplets in one shot, with two additional special tokens from pre-trained models to represent potential implicit aspect and opinion terms. Based on this, we first introduce the adaptive graph diffusion convolution network to construct the direct connection between explicit and implicit sentiment elements from syntactic and semantic views. Next, we utilize conditional layer normalization to refine the mutual indication effect between words for matching valid aspect-opinion pairs. Finally, we employ the triaffine mechanism to integrate heterogeneous word-pair relations to capture higher-order interactions between sentiment elements. Experimental results on four benchmark datasets show the effectiveness and robustness of our model, which achieves state-of-the-art performance.
Guixin Su, Yongcheng Zhang, Tongguan Wang, Mingmin Wu, Ying Sha
COLING3
2025 GEETI: Graph Embedding-Enhanced Textual Inversion for Chinese Harmful Meme Detection and Identification
Shichao Fu, Tongguan Wang, Ying Sha
ICIC (10)2
2025 Keyword-Oriented Multimodal Modeling for Euphemism Identification
abstract
Euphemism identification deciphers the true meaning of euphemisms, such as linking "weed" (euphemism) to "marijuana" (target keyword) in illicit texts, aiding content moderation and combating underground markets. While existing methods are primarily text-based, the rise of social media highlights the need for multimodal analysis, incorporating text, images, and audio. However, the lack of multimodal datasets for euphemisms limits further research. To address this, we regard euphemisms and their corresponding target keywords as keywords and first introduce a keyword-oriented multimodal corpus of euphemisms (KOM-Euph), involving three datasets (Drug, Weapon, and Sexuality), including text, images, and speech. We further propose a keyword-oriented multimodal euphemism identification method (KOM-EI), which uses cross-modal feature alignment and dynamic fusion modules to explicitly utilize the visual and audio features of the keywords for efficient euphemism identification. Extensive experiments demonstrate that KOM-EI outperforms state-of-the-art models and large language models, and show the importance of our multimodal datasets1.
Yuxue Hu, Junsong Li, Meixuan Chen, Dongyu Su, Tongguan Wang, Ying Sha
ICME5
2025 Incongruity-aware Cross-modal Interaction Network for Multimodal Sarcasm Detection
abstract
Multimodal sarcasm detection aims to identify sarcasm in image-text pairs, essential for reducing malicious social media content. Recent approaches have achieved significant progress by leveraging the powerful pre-trained CLIP model to extract sarcasm cues from both modalities. However, two main issues remain unresolved: insufficient modeling of fine-grained cross-modal inconsistencies and lack of effective mitigation of modality competition. To address these issues, we propose the Incongruity-aware Cross-modal Interaction Network (ICIN). Specifically, we design a feature decomposition and aggregation block to reduce redundant image information from multiple perspectives and integrate semantically enhanced text to capture fine-grained sarcasm cues within each modality. Subsequently, we propose a dynamic multimodal interaction block that captures intermodal inconsistencies via interactive attention and employs an adaptive gating mechanism to balance information from each modality, thus mitigating competition. Finally, a dynamic joint loss optimization strategy is developed to coordinate feature distribution discrepancies. Experimental results on two benchmark datasets show that the proposed model significantly outperforms previous state-of-the-art methods.
Yujun Wu, Meixuan Chen, Tongguan Wang, Ying Sha
ICME4
2025 RCLMuFN: Relational context learning and multiplex fusion network for multimodal sarcasm detection
Tongguan Wang, Junkai Li, Guixin Su, Yongcheng Zhang, Dongyu Su, Yuxue Hu, Ying Sha
Knowl. Based Syst.1
2025 Lightweight image super-resolution reconstruction based on mixed attention and global inductive bias network
Yuxi Cai, Xiaopeng Wen, Tongguan Wang
Multim. Tools Appl.3
2024 Efficient Guided Query Network for Human-Object Interaction Detection
abstract
Recently, Transformer-based one-stage methods have demonstrated excellent efficiency in Human-Object Interaction (HOI) tasks. However, these methods often utilize semantically ambiguous initial queries, thus constraining the model’s ability for set prediction. In addition, currently widely used HOI datasets suffer from long-tail distribution issues, so accurately identifying rare interaction categories remains challenging. To address these challenges, we propose an Efficient Guided Query Network (EGQ-Net). The network introduces a forward-guided relational queries approach, which accurately captures the triplets of interaction relationships by effectively integrating the initial queries predicted by the encoder and the output features of each decoder layer. Furthermore, we used the visual language pre-training models CLIP and BLIP2 to design interaction position query guidance and interaction content query guidance to achieve accurate recognition and localization of interactive areas by queries. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on widely used HOI benchmarks (V-COCO and HICO-DET).
Junkai Li, Huicheng Lai, Tongguan Wang, Hutuo Quan, Dongji Chen
ICME4
2024 UFSRNet: U-shaped face super-resolution reconstruction network based on wavelet transform
Tongguan Wang, Yang Xiao 0018, Yuxi Cai, Guxue Gao, Xiaocong Jin, Huicheng Lai
Multim. Tools Appl.1
2023 Scoreformer: Score Fusion-Based Transformers for Weakly-Supervised Violence Detection
abstract
Violence detection is an application of anomaly detection, which is used to detect violence content in video clips. Using multimodal as input can improve the performance of violence detection. However, the existing MML Transformers-based fusion methods do not take into account the differences between non-homologous modals. The fusion of non-homologous modals makes features become noise between each other. This paper proposes a score fusion-based transformer framework, named Scoreformer. First of all, the optical flow, RGB and audio features pass through the independent self-attention transformer blocks. Second, the optical flow and RGB features pass through the cross-modal transformer blocks, after that they are fused with the audio features through the score fusion block. This method avoids the noise interference caused by the direct fusion of audio features and visual features. Experiments on the XD-Violence dataset show that the proposed method achieves 84.54% of the AP value, which exceeds at least 2.85% compared with the most advanced method (e. g. MSL, CRFD).
Yang Xiao 0018, Tongguan Wang, Huicheng Lai
ICASSP3
2023 Video anomaly detection based on cross-frame prediction mechanism and spatio-temporal memory-enhanced pseudo-3D encoder
Xiaopeng Wen, Huicheng Lai, Guxue Gao, Yang Xiao 0018, Tongguan Wang, Zhenhong Jia
Eng. Appl. Artif. Intell.5