EDBT 2026 Demo / reviewers in the wild / expert
Wenya Guo
dblp:234/4615
· DBLP profile ↗
22ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0001-5609-194XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSR: Structured Subgraph Retrieval for Temporal Knowledge Graph Question Answering with LLMsabstractTemporal Knowledge Graph Question Answering (TKGQA) aims to answer natural language questions based on quadruple facts stored in Temporal Knowledge Graphs (TKGs). Recent studies have integrated Large Language Models (LLMs) to handle the complex semantic reasoning required by temporal questions. They typically linearize TKG quadruples into plain text, reducing TKGQA to a top-n text retrieval and context-based question answering problem. However, this textualization process destroys the original quadruple structure and entangles structural and temporal information with text semantics. Moreover, selecting top-n facts based on semantic similarity inevitably introduces a large amount of irrelevant noise into the LLM input, which degrades reasoning performance. To address these limitations, we propose SSR, a Structured Subgraph Retrieval framework for TKGQA with LLMs. In SSR, a Temporal Question Parser is first introduced to extract structured subgraph patterns and temporal constraints from the input questions, leveraging background context obtained via text retrieval. The Subgraph Retrieval module is then applied to directly filter a relevant subgraph from the TKG that contains the facts required to answer the question. The retrieved subgraph is further temporally compressed and subsequently incorporated into the LLM context for final question answering. Experimental results on two benchmark datasets demonstrate that SSR consistently outperforms strong baselines by a clear margin, achieving state-of-the-art performance. Our code is available at https://github.com/zhangli-coding/SSR. Ying Zhang 0015, Wenya Guo, Shilong Ping, Xinying Qian |
SIGIR | 3 |
| 2026 | Implicit visual knowledge enhanced zero-shot image captioning
Xubo Liu 0002, Wenya Guo, Xumeng Liu, Ruxue Yan, Ying Zhang 0015 |
Expert Syst. Appl. | 2 |
| 2026 | SMIR: Span-based multi-grained information refinement for joint multimodal entity-relation extraction
Xuhui Sui, Ying Zhang 0015, Yu Zhao 0043, Baohang Zhou, Xinying Qian, Wenya Guo, Xiaojie Yuan |
Inf. Process. Manag. | 6 |
| 2026 | ToM: Boosting TextVQA by capturing text-oriented keypoints
Ruxue Yan, Wenya Guo, Xubo Liu 0002, Xumeng Liu, Ying Zhang 0015, Xiaojie Yuan |
Knowl. Based Syst. | 2 |
| 2025 | BiDeV: Bilateral Defusing Verification for Complex Claim Fact-CheckingabstractComplex claim fact-checking performs a crucial role in disinformation detection. However, existing fact-checking methods struggle with claim vagueness, specifically in effectively handling latent information and complex relations within claims. Moreover, evidence redundancy, where non-essential information complicates the verification process, remains a significant issue. To tackle these limitations, we propose Bilateral Defusing Verification (BiDeV), a novel fact-checking working-flow framework integrating multiple role-played LLMs to mimic the human-expert fact-checking process. BiDeV consists of two main modules: Vagueness Defusing identifies latent information and resolves complex relations to simplify the claim, and Redundancy Defusing eliminates redundant content to enhance the evidence quality. Extensive experimental results on two widely used challenging fact-checking benchmarks (Hover and Feverous-s) demonstrate that our BiDeV can achieve the best performance under both gold and open settings. This highlights the effectiveness of BiDeV in handling complex claims and ensuring precise fact-checking. Yuxuan Liu 0009, Hongda Sun 0001, Wenya Guo, Xinyan Xiao, Cunli Mao, Zhengtao Yu 0001, Rui Yan 0001 |
AAAI | 3 |
| 2025 | Zero-shot Document Retrieval with Hybrid Pseudo-document RetrieverabstractThe zero-shot retrieval task aims to retrieve the most relevant documents to a user’s query without relevance labels. Current approaches expand input queries by generating pseudo-documents with large language models (LLMs) and perform document retrieval based on the expanded queries. However, their retrieval methods are limited to either sparse or dense retrieval methods alone. In this paper, we propose a hybrid retriever to further improve the quality of the pseudo-documents and to obtain the relevant information more effectively. Specifically, we use sparse retrievers to obtain keyword information and dense retrievers to obtain contextual information. Then, we introduce reciprocal ranking fusion and weighted scoring fusion into both the pre-retrieval of candidate documents and the final-retrieval of final results to calculate the overall hybrid matching score. Experimental results on TREC DL19/DL20 and several datasets from the BEIR benchmark indicate the superiority of our proposed method. Wenya Guo, Xumeng Liu, Zhaoxiang Hou, Zengxiang Li |
ICASSP | 2 |
| 2025 | Rethinking the Reliability of Evidence in End-to-End Fact-Checking from the Causal PerspectiveabstractEnd-to-end automated fact-checking (AFC) aims to assess the truthfulness of claims using retrieved evidence. Some researchers use crawlers or search APIs to retrieve evidence from the web for veracity classification. However, existing methods indiscriminately rely on the retrieved evidence and overlook that the retrieved results are not always reliable. This unilateral reliance on evidence significantly hampers the performance of fact-checking. In this paper, we account for the diverse reliability levels of retrieved evidence and eliminate the negative impact from the causal perspective. To achieve our goal, we propose a novel Causal intervention and Counterfactual reasoning based Multi-Checker framework (CCMC), which introduces two additional counterfactual fact-checkers to verify claims from the counterfactual perspective. Specifically, we construct two distinct types of counterfactual instances via causal intervention to imitate the situation where the evidence is partially reliable or totally unreliable. Correspondingly, two counterfactual fact-checkers are trained with tailored counterfactual instances by counterfactual reasoning. During inference, the two counterfactual fact-checkers are employed to estimate and eliminate the potential impact of unreliable evidence. Extensive experiments on two real-world datasets demonstrate the superiority of our approach for improving end-to-end AFC. Especially, we surpass existing methods by 3.70% and 5.55% under gold and system evidence on the MOCHEG benchmark, respectively. Our code is available at https://github.com/BeiyuXuboL/CCMC. Xubo Liu 0002, Wenya Guo, Ruxue Yan, Xumeng Liu, Ying Zhang 0015, Ru Zhou |
ACM Multimedia | 2 |
| 2025 | Prompt Compression based on Key-Information Density
Wenya Guo, Ying Zhang 0015, Zengxiang Li |
Expert Syst. Appl. | 2 |
| 2025 | QC2-VQG: Question context complement for visual question generation
Ying Zhang 0015, Xubo Liu 0002, Wenya Guo, Xumeng Liu, Ruxue Yan |
Knowl. Based Syst. | 4 |
| 2024 | Look before You Leap: Dual Logical Verification for Knowledge-based Visual Question GenerationabstractKnowledge-based Visual Question Generation aims to generate visual questions with outside knowledge other than the image. Existing approaches are answer-aware, which incorporate answers into the question-generation process. However, these methods just focus on leveraging the semantics of inputs to propose questions, ignoring the logical coherence among generated questions (Q), images (V), answers (A), and corresponding acquired outside knowledge (K). It results in generating many non-expected questions with low quality, lacking insight and diversity, and some of them are even without any corresponding answer. To address this issue, we inject logical verification into the processes of knowledge acquisition and question generation, which is defined as LVˆ2-Net. Through checking the logical structure among V, A, K, ground-truth and generated Q twice in the whole KB-VQG procedure, LVˆ2-Net can propose diverse and insightful knowledge-based visual questions. And experimental results on two commonly used datasets demonstrate the superiority of LVˆ2-Net. Our code will be released to the public soon. Xumeng Liu, Wenya Guo, Ying Zhang 0015, Xubo Liu 0002, Yu Zhao 0043, Samson Shenglong Yu, Xiaojie Yuan |
LREC/COLING | 2 |
| 2024 | Tracking-forced Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) is a cross-modal task that aims to segment the target object described by language expressions. A video typically consists of multiple frames and existing works conduct segmentation at either the clip-level or the frame-level. Clip-level methods process a clip at once and segment in parallel, lacking explicit inter-frame interactions. In contrast, frame-level methods facilitate direct interactions between frames by processing videos frame by frame, but they are prone to error accumulation. In this paper, we propose a novel tracking-forced framework, introducing high-quality tracking information and forcing the model to achieve accurate segmentation. Concretely, we utilize the ground-truth segmentation of previous frames as accurate inter-frame interactions, providing high-quality tracking references for segmentation in the next frame. This decouples the current input from the previous output, which enables our model to concentrate on accurately segmenting just based on given tracking information, improving training efficiency and preventing error accumulation. For the inference stage without ground-truth masks, we carefully select the beginning frame to construct tracking information, aiming to ensure accurate tracking-based frame-by-frame object segmentation. With these designs, our tracking-forced method significantly outperforms existing methods on 4 widely used benchmarks by at least 3%. Especially, our method achieves 88.3% [email protected] accuracy and 87.6 overall IoU score on the JHMDB-Sentences dataset, surpassing previous best methods by 5.0% and 8.0, respectively. Ruxue Yan, Wenya Guo, Xubo Liu 0002, Xumeng Liu, Ying Zhang 0015, Xiaojie Yuan |
ACM Multimedia | 2 |
| 2024 | SINet: Improving relational features in two-stage referring expression comprehension
Wenya Guo, Ying Zhang 0015, Xiaojie Yuan |
Expert Syst. Appl. | 1 |
| 2023 | CD-BLI: Confidence-Based Dual Refinement for Unsupervised Bilingual Lexicon Induction
Samson Shenglong Yu, Wenya Guo, Ying Zhang 0015, Xiaojie Yuan |
NLPCC (2) | 2 |
| 2023 | Learning Disentangled Representation for Multimodal Cross-Domain Sentiment AnalysisabstractMultimodal cross-domain sentiment analysis aims at transferring domain-invariant sentiment information across datasets to address the insufficiency of labeled data. Existing adaptation methods achieve well performance by remitting the discrepancies in characteristics of multiple modalities. However, the expressive styles of different datasets also contain domain-specific information, which hinders the adaptation performance. In this article, we propose a disentangled sentiment representation adversarial network (DiSRAN) to reduce the domain shift of expressive styles for multimodal cross-domain sentiment analysis. Specifically, we first align the multiple modalities and obtain the joint representation through a cross-modality attention layer. Then, we disentangle sentiment information from the multimodal joint representation that contains domain-specific expressive style by adversarial training. The obtained sentiment representation is domain-invariant, which can better facilitate the sentiment information transfer between different domains. Experimental results on two multimodal cross-domain sentiment analysis tasks demonstrate that the proposed method performs favorably against state-of-the-art approaches. Ying Zhang 0015, Wenya Guo, Xiangrui Cai, Xiaojie Yuan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Dynamic Alternative Attention for Visual Question Answering
Xumeng Liu, Wenya Guo |
WISA | 2 |
| 2022 | A Span-based Multimodal Variational Autoencoder for Semi-supervised Multimodal Named Entity RecognitionabstractMultimodal named entity recognition (MNER) on social media is a challenging task which aims to extract named entities in free text and incorporate images to classify them into userdefined types.The existing semi-supervised named entity recognition methods focus on the text modal and are utilized to reduce labeling costs in traditional NER.However, the previous methods are not efficient for semisupervised MNER.Because the MNER task is defined to combine the text information with image one and needs to consider the mismatch between the posted text and image.To fuse the text and image features for MNER effectively under semi-supervised setting, we propose a novel span-based multimodal variational autoencoder (SMVAE) model for semisupervised MNER.The proposed method exploits modal-specific VAEs to model text and image latent features, and utilizes product-ofexperts to acquire multimodal features.In our approach, the implicit relations between labels and multimodal features are modeled by multimodal VAE.Thus, the useful information of unlabeled data can be exploited in our method under semi-supervised setting.Experimental results on two benchmark datasets demonstrate that our approach not only outperforms baselines under supervised setting, but also improves MNER performance with less labeled data than existing semi-supervised methods. Baohang Zhou, Ying Zhang 0015, Kehui Song, Wenya Guo, Xiaojie Yuan |
EMNLP | 4 |
| 2021 | MTAAL: Multi-Task Adversarial Active Learning for Medical Named Entity Recognition and NormalizationabstractAutomated medical named entity recognition and normalization are fundamental for constructing knowledge graphs and building QA systems. When it comes to medical text, the annotation demands a foundation of expertise and professionalism. Existing methods utilize active learning to reduce costs in corpus annotation, as well as the multi-task learning strategy to model the correlations between different tasks. However, existing models do not take task-specific features for different tasks and diversity of query samples into account. To address these limitations, this paper proposes a multi-task adversarial active learning model for medical named entity recognition and normalization. In our model, the adversarial learning keeps the effectiveness of multi-task learning module and active learning module. The task discriminator eliminates the influence of irregular task-specific features. And the diversity discriminator exploits the heterogeneity between samples to meet the diversity constraint. The empirical results on two medical benchmarks demonstrate the effectiveness of our model against the existing methods. Baohang Zhou, Xiangrui Cai, Ying Zhang 0015, Wenya Guo, Xiaojie Yuan |
AAAI | 4 |
| 2021 | Missing value imputation in multivariate time series with end-to-end generative adversarial networks
Ying Zhang 0015, Baohang Zhou, Xiangrui Cai, Wenya Guo, Xiaoke Ding, Xiaojie Yuan |
Inf. Sci. | 4 |
| 2021 | Re-Attention for Visual Question AnsweringabstractA simultaneous understanding of questions and images is crucial in Visual Question Answering (VQA). While the existing models have achieved satisfactory performance by associating questions with key objects in images, the answers also contain rich information that can be used to describe the visual contents in images. In this paper, we propose a re-attention framework to utilize the information in answers for the VQA task. The framework first learns the initial attention weights for the objects by calculating the similarity of each word-object pair in the feature space. Then, the visual attention map is reconstructed by re-attending the objects in images based on the answer. Through keeping the initial visual attention map and the reconstructed one to be consistent, the learned visual attention map can be corrected by the answer information. Besides, we introduce a gate mechanism to automatically control the contribution of re-attention to model training based on the entropy of the learned initial visual attention maps. We conduct experiments on three benchmark datasets, and the results demonstrate the proposed model performs favorably against state-of-the-art methods. Wenya Guo, Ying Zhang 0015, Jufeng Yang, Xiaojie Yuan |
IEEE Trans. Image Process. | 1 |
| 2021 | LD-MAN: Layout-Driven Multimodal Attention Network for Online News Sentiment RecognitionabstractThe prevailing use of both images and text to express opinions on the web leads to the need for multimodal sentiment recognition. Some commonly used social media data containing short text and few images, such as tweets and product reviews, have been well studied. However, it is still challenging to predict the readers’ sentiment after reading online news articles, since news articles often have more complicated structures, e.g., longer text and more images. To address this problem, we propose a layout-driven multimodal attention network (LD-MAN) to recognize news sentiment in an end-to-end manner. Rather than modeling text and images individually, LD-MAN uses the layout of online news to align images with the corresponding text. Specifically, it exploits a set of distance-based coefficients to model the image locations and measure the contextual relationship between images and text. LD-MAN then learns the affective representations of the articles from the aligned text and images using a multimodal attention mechanism. Considering the lack of relevant datasets in this field, we collect two multimodal online news datasets, containing a total of 14,566 articles with 56,260 images and 251,202 words. Experimental results demonstrate that the proposed method performs favorably compared with state-of-the-art approaches. We will release all the codes, models and datasets to the community. Wenya Guo, Ying Zhang 0015, Xiangrui Cai, Lei Meng 0001, Jufeng Yang, Xiaojie Yuan |
IEEE Trans. Multim. | 1 |
| 2020 | Re-Attention for Visual Question AnsweringabstractVisual Question Answering~(VQA) requires a simultaneous understanding of images and questions. Existing methods achieve well performance by focusing on both key objects in images and key words in questions. However, the answer also contains rich information which can help to better describe the image and generate more accurate attention maps. In this paper, to utilize the information in answer, we propose a re-attention framework for the VQA task. We first associate image and question by calculating the similarity of each object-word pairs in the feature space. Then, based on the answer, the learned model re-attends the corresponding visual objects in images and reconstructs the initial attention map to produce consistent results. Benefiting from the re-attention procedure, the question can be better understood, and the satisfactory answer is generated. Extensive experiments on the benchmark dataset demonstrate the proposed method performs favorably against the state-of-the-art approaches. Wenya Guo, Ying Zhang 0015, Jufeng Yang, Xiangrui Cai, Xiaojie Yuan |
AAAI | 1 |
| 2018 | Jointly Trained Convolutional Neural Networks for Online News Emotion Analysis
Xue Zhao 0001, Ying Zhang 0015, Wenya Guo, Xiaojie Yuan |
WISA | 3 |