VLDB 2026 Research / reviewers in the wild / expert
Jiayuan Xie
dblp:249/8382
· DBLP profile ↗
37ranked-venue papers
13as first author
36since 2021 · last 2026
0000-0002-6833-7879ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 8 first-author · 20 since 2021Artificial intelligence and machine learning · 19 · 8 first-author · 18 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CADMate: Generating CAD Assembly Plan with Geometric Chain-of-Thought and Spatial Physical RewardsabstractJiali Chen, DingBa Fu, Xusen Hei, Yuhang Liu, Yiyang Chen, Jiayuan Xie, Wenqi Fan, Yi Cai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. DingBa Fu, Xusen Hei, Jiayuan Xie, Wenqi Fan, Yi Cai 0001 |
ACL (1) | 6 |
| 2026 | Invert Your Prompt: Editing-Aware Diffusion Inversion
Yangyang Xu 0003, Wenqi Shao, Yong Du 0003, Haiming Zhu, Yang Zhou 0038, Jiayuan Xie, Ping Luo 0002, Shengfeng He |
Int. J. Comput. Vis. | 6 |
| 2026 | Stage-aware industrial defect understanding via multi-agent collaboration
Jiayuan Xie, Xinting Zhang, Yuxi Tu, Yi Cai 0001, Qing Li 0001 |
Knowl. Based Syst. | 1 |
| 2026 | Integrative Prompt Learning for Continual Defect Detection in Industrial Scenarios
Jiayuan Xie, Yuqi Xue, Yi Cai 0001, Qing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Explicitly Guided Difficulty-Controllable Visual Question GenerationabstractVisual question generation (VQG) aims to generate questions from images automatically. While existing studies primarily focus on the quality of generated questions, such as fluency and relevance, the difficulty of the questions is also a crucial factor in assessing their quality. Question difficulty directly impacts the effectiveness of VQG systems in applications like education and human-computer interaction, where appropriately challenging questions can stimulate learning interest and improve interaction experiences. However, accurately defining and controlling question difficulty is a challenging task due to its multidimensional and subjective nature. In this paper, we propose a new definition of the difficulty of questions, i.e., being positively correlated with the number of reasoning steps required to answer a question. For our definition, we construct a corresponding dataset and propose a benchmark as a foundation for future research. Our benchmark is designed to progressively increase the reasoning steps involved in generating questions. Specifically, we first extract the relationships among objects in the image to form a reasoning chain, then gradually increase the difficulty by rewriting the generated question to include more reasoning sub-chains. Experimental results on our constructed dataset show that our benchmark significantly outperforms existing baselines in controlling the reasoning chains of generated questions, producing questions with varying difficulty levels. Jiayuan Xie, Mengqiu Cheng, Xinting Zhang, Yi Cai 0001, Guimin Hu, Mengying Xie, Qing Li 0001 |
AAAI | 1 |
| 2025 | CADReview: Automatically Reviewing CAD Programs with Error Detection and CorrectionabstractJiali Chen, Xusen Hei, HongFei Liu, Yuancheng Wei, Zikun Deng, Jiayuan Xie, Yi Cai, Li Qing. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xusen Hei, Yuancheng Wei, Zikun Deng, Jiayuan Xie, Yi Cai 0001, Qing Li 0001 |
ACL (1) | 6 |
| 2025 | Fine-Grained Features-based Code Search for Precise Query-Code MatchingabstractCode search aims to quickly locate target code snippets from databases using natural language queries, which promotes code reusability. Existing methods can effectively obtain aligned token-level and query word-level features. However, these studies usually represent the semantics of code and query by averaging the features of each token and word respectively, which makes it difficult to accurately capture the code details that are closely related to the query. To address this issue, we propose a fine-grained code search model that consists of a cross-modal encoder, a mapping layer, and a classification layer. Specifically, we utilize a pre-trained model, GraphCodeBERT, in the cross-modal encoder to align features. In the mapping layer, we introduce a co-attention network to capture the fine-grained interactions between code and query, ensuring a model can precisely identify key code segments relevant to the query. Finally, in the classification layer, we incorporate instruction learning techniques that leverage contextual reasoning to improve the accuracy of query-code matching. Experimental results show that our proposed model significantly outperforms existing methods across multiple programming language datasets. Xinting Zhang, Mengqiu Cheng, Mengzhen Wang, Songwen Gong, Jiayuan Xie, Yi Cai 0001, Qing Li 0001 |
COLING | 5 |
| 2025 | Tree-of-Reasoning: Towards Complex Medical Diagnosis via Multi-Agent Reasoning with Evidence TreeabstractLarge language models (LLMs) have shown great potential in the medical domain. However, existing models still fall short when faced with complex medical diagnosis task in the real world. This is mainly because they lack sufficient reasoning depth, which leads to information loss or logical jumps when processing a large amount of specialized medical data, leading to diagnostic errors. To address these challenges, we propose Tree-of-Reasoning (ToR), a novel multi-agent framework designed to handle complex scenarios. Specifically, ToR introduces a tree structure that can clearly record the reasoning path of LLMs and the corresponding clinical evidence. At the same time, we propose a cross-validation mechanism to ensure the consistency of multi-agent decision-making, thereby improving the clinical reasoning ability of multi-agents in complex medical scenarios. Experimental results on real-world medical data show that our framework can achieve better performance than existing baseline methods. Qi Peng 0002, Jialin Cui, Jiayuan Xie, Yi Cai 0001, Qing Li 0001 |
ACM Multimedia | 3 |
| 2025 | ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific ExperimentsabstractExperiment commentary is crucial in describing the experimental procedures, delving into underlying scientific principles, and incorporating content-related safety guidelines. In practice, human teachers rely heavily on subject-specific expertise and invest significant time preparing such commentary. To address this challenge, we introduce the task of automatic commentary generation across multi-discipline scientific experiments. Current LMMs' ability to generate fine-grained and insightful experiment commentary remains largely under-explored. In this paper, we make the following contributions: (i) We construct ExpInstruct, the first dataset tailored for experiment commentary generation, featuring over 7 K step-level commentaries across 21 scientific subjects from 3 core disciplines. (ii) We propose ExpStar, an automatic experiment commentary generation model that leverages a retrieval-augmented mechanism to adaptively access, evaluate, and utilize external knowledge. (iii) Extensive experiments show that our ExpStar substantially outperforms 14 leading LMMs, which highlights the superiority of our dataset and model. We believe that ExpStar holds great potential for advancing AI-assisted scientific experiment instruction. Yujie Jia, Jianpeng Chen, Xusen Hei, Jiayuan Xie, Yi Cai 0001, Qing Li 0001 |
ACM Multimedia | 7 |
| 2025 | From Model Diagram to Code: A Benchmark Dataset and Multi-Agent FrameworkabstractModel Diagram-to-Code Generation aims to translate model diagrams from research papers into implementation code that reconstructs the model's architecture. This task plays a crucial role in accelerating scientific workflows and enhancing the efficiency of industrial model deployment. While recent studies have explored various Image-to-Code Generation tasks using Multimodal Large Language Models (MLLMs), these efforts have primarily focused on reconstructing the visual appearance depicted in input images, leaving this task largely underexplored. The complex structural elements and implicit relationships in model diagrams present greater challenges for MLLMs, particularly in terms of visual reasoning and semantic interpretation. To support this task, we introduce MDCDataset, a dataset designed to evaluate the ability of MLLMs to generate code from model diagrams. It comprises 1,008 instances spanning 16 research domains, each with a model diagram, structured textual content, and the ground-truth code implementation. Furthermore, to address the inherent challenges of this task, we propose MDCAgent, a collaborative multi-agent framework composed of Parsing, Generation, and Check Agents. These agents work in coordination to analyze, extract, and verify complex elements and implicit relationships within model diagrams, thereby enhancing the visual architecture-aware reasoning capabilities of MLLMs. Our extensive experiments confirm the effectiveness of the framework. Mengzhen Wang, Xunbin Huang, Jiayuan Xie, Shukai Ma, Jiale Men, Dayong Liang, Yi Cai 0001 |
ACM Multimedia | 3 |
| 2025 | Visual defect detection for historical building preservation
Mengqiu Cheng, Xinting Zhang, Leihua Xia, Jiayuan Xie, Zongfang Ma, Qing Li 0001 |
Expert Syst. Appl. | 5 |
| 2025 | CKE-Former: A clinical knowledge-enhanced transformer for disease classification in telemedicine
Qi Peng 0002, Yi Cai 0001, Jiankun Liu, Jiayuan Xie, Qing Li 0001 |
Knowl. Based Syst. | 6 |
| 2025 | Explicitly diverse visual question generation
Jiayuan Xie, Jiasheng Zheng, Wenhao Fang, Yi Cai 0001, Qing Li 0001 |
Neural Networks | 1 |
| 2025 | Integration of Multi-Source Medical Data for Medical Diagnosis Question AnsweringabstractMedical question answering aims to enhance diagnostic support, improve patient education, and assist in clinical decision-making by automatically answering medical-related queries, which is an important foundation for realizing intelligent healthcare. Existing methods predominantly focus on extracting key information from a single data source, e.g., CT image, for answering. However, these methods are not enough to promote the development of intelligent healthcare, because they lack comprehensive medical diagnosis capabilities, which usually require the integration of multi-source data (e.g., laboratory tests, radiology images, pathology images, etc.) for processing. To address these limitations, our paper introduces the extended task of medical question answering, named medical diagnosis question answering MedDQA. MedDQA task aims to answer questions related to medical diagnosis based on multi-source data. Specifically, we introduce a corresponding dataset that incorporates multi-source diagnostic information from 250,917 patients in clinical data from hospital records, and utilize a large-scale model for constructing Q&A pairs. We propose a novel system based on large language models, named medical multi-agent (MMA) system, which includes a mechanism of multiple agents to handle different medical tasks. Each agent is specifically tailored to process various modalities of data and provide outputs in a uniform textual modality. Experimental results demonstrate that the MMA system's architecture significantly enhances the handling of multi-source data, thereby improving medical diagnosis, establishing a robust baseline for future research. Qi Peng 0002, Yi Cai 0001, Jiankun Liu, Jiayuan Xie, Qing Li 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2024 | Automated Defect Report Generation for Enhanced Industrial Quality ControlabstractDefect detection is a pivotal aspect ensuring product quality and production efficiency in industrial manufacturing. Existing studies on defect detection predominantly focus on locating defects through bounding boxes and classifying defect types. However, their methods can only provide limited information and fail to meet the requirements for further processing after detecting defects. To this end, we propose a novel task called defect detection report generation, which aims to provide more comprehensive and informative insights into detected defects in the form of text reports. For this task, we propose some new datasets, which contain 16 different materials and each defect contains a detailed report of human constructs. In addition, we propose a knowledge-aware report generation model as a baseline for future research, which aims to incorporate additional knowledge to generate detailed analysis and subsequent processing related to defect in images. By constructing defect report datasets and proposing corresponding baselines, we chart new directions for future research and practical applications of this task. Jiayuan Xie, Zhiping Zhou, Xinting Zhang, Jiexin Wang 0002, Yi Cai 0001, Qing Li 0001 |
AAAI | 1 |
| 2024 | Knowledge-Guided Cross-Topic Visual Question GenerationabstractVisual question generation (VQG) task aims to generate high-quality questions based on the input image. Current methods primarily focus on generating questions containing specified content utilizing answers or question types as constraints. However, these constraints make it challenging to control the topic of generated questions (e.g., conversation or test subject topics) for various applications. Thus, it is necessary to utilize topics as constraints to guide question generation. Considering that there are many topics and it is almost impossible for human annotations to cover them, we propose the cross-topic learning VQG (CTL-VQG) task, which aims to generate questions related to unseen topics in cross-topic scenarios. In this paper, we propose a knowledge-guided cross-topic visual question generation (KC-VQG) model to extract unseen topic-related information for question generation. Specifically, an image-topic feature extractor is introduced in our model to extract topic-related intuitive visual features; an image-topic knowledge extractor is used to extract and select the most appropriate topic-related implicit knowledge from large language models for generating questions. Extensive experiments show that our model outperforms baselines and can effectively generate unseen topic-related questions in cross-topic scenarios. Guohua Wang 0003, Jiayuan Xie, Wenhao Fang, Yi Cai 0001 |
LREC/COLING | 3 |
| 2024 | Deconfounded Emotion Guidance Sticker Selection with Causal InferenceabstractWith the increasing popularity of online social applications, stickers have become common in online chats. Teaching a model to select the appropriate sticker from a set of candidate stickers based on dialogue context is important for optimizing the user experience. Existing methods have proposed leveraging emotional information to facilitate the selection of appropriate stickers. However, considering the frequent co-occurrence among sticker images, words with emotional preference in the dialogue and emotion labels, these methods tend to over-rely on such dataset bias, inducing spurious correlations during training. As a result, these methods may select inappropriate stickers that do not match users' intended expression. In this paper, we introduce a causal graph to explicitly identify the spurious correlations in the sticker selection task. Building upon the analysis, we propose a Causal Knowledge-Enhanced Sticker Selection model to mitigate spurious correlations. Specifically, we design a knowledge-enhanced emotional utterance extractor to identify emotional information within dialogues. Then an interventional visual feature extractor is employed to obtain unbiased visual features, aligning them with the emotional utterances representation. Finally, a standard transformer encoder fuses the multimodal information for emotion recognition and sticker selection. Extensive experiments on the MOD dataset show that our CKS model significantly outperforms the baseline models. Yi Cai 0001, Ruohang Xu, Jiexin Wang 0002, Jiayuan Xie, Qing Li 0001 |
ACM Multimedia | 5 |
| 2024 | Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning DistractorabstractLarge multimodal models (LMMs) have shown remarkable performance in the visual commonsense reasoning (VCR) task, which aims to answer a multiple-choice question based on visual commonsense within an image. However, the ability of LMMs to correct potential visual commonsense errors in the distractor upon their occurrence is yet under-explored. Drawing inspiration from how a human teacher crafts challenging distractors to test students' comprehension of the concepts or skills and assists them in identifying and correcting errors toward the answer, we are the pioneering research for LMMs to simulate this error correction process. To this end, we employ GPT-4 as a ''teacher'' to collect the explainable feedback dataset VCR-DF for error correction, which serves as a benchmark to evaluate the ability of LMMs to identify misconceptions and clarify reasons behind the error in VCR distractors toward final answers. In addition, we propose an LMM-based Pedagogical Expert Instructed Feedback Generation (PEIFG) model to incorporate the learnable expert prompts and multimodal instruction as guidance for feedback generation. Experimental results show that our PEIFG significantly outperforms existing LMMs. We believe that our benchmark provides a new direction for evaluating the capabilities of LMMs. Xusen Hei, Yuqi Xue, Yuancheng Wei, Jiayuan Xie, Yi Cai 0001, Qing Li 0001 |
ACM Multimedia | 5 |
| 2024 | Generalized News Event Discovery via Dynamic Augmentation and Entropy OptimizationabstractNews event discovery refers to the identification and detection of news events using multimodal data on social media. Currently, most works assume that the test set consists of known events. However, in real life, the emergence of new events is more frequent, which invalidates this assumption. In this paper, we propose a Dynamic Augmentation and Entropy Optimization (DAEO) model to address the scenario of generalized news event discovery, which requires the model to not only identify known events but also distinguish various new events. Specifically, we first introduce a multimodal augmentation module, which utilizes adversarial learning to enhance the multimodal representation capability. Secondly, we design an adaptive entropy optimization strategy combined with a self-distillation method, which uses multi-view pseudo-label consistency to improve the model's performance on both known and new events. In addition, we collect a multimodal news event discovery (MNED) dataset of 161,350 samples annotated with 66 real-world events. Extensive experimental results on the MNED dataset demonstrate the effectiveness of our proposed method. Our dataset is available on https://github.com/RetrainIt/MNED. Zehang Lin, Jiayuan Xie, Zhenguo Yang, Yi Yu 0001, Qing Li 0001 |
ACM Multimedia | 2 |
| 2024 | Multi-modal news event detection with external knowledge
Zehang Lin, Jiayuan Xie, Qing Li 0001 |
Inf. Process. Manag. | 2 |
| 2024 | Context-Aware Dynamic Word Embeddings for Aspect Term ExtractionabstractThe aspect term extraction (ATE) task aims to extract aspect terms describing a part or an attribute of a product from review sentences. Most existing works rely on either general or domain embedding to address this problem. Despite the promising results, the importance of general and domain embeddings is still ignored by most methods, resulting in degraded performances. Besides, word embedding is also related to downstream tasks, and how to regularize word embeddings to capture context-aware information is an unresolved problem. To solve these issues, we first propose context-aware dynamic word embedding (CDWE), which could simultaneously consider general meanings, domain-specific meanings, and the context information of words. Based on CDWE, we propose an attention-based convolution neural network, called ADWE-CNN for ATE, which could adaptively capture the previous meanings of words by utilizing an attention mechanism to assign different importance to the respective embeddings. The experimental results show that ADWE-CNN achieves a comparable performance with the state-of-the-art approaches. Various ablation studies have been conducted to explore the benefit of each component. Our code is publicly available athttp://github.com/xiejiajia2018/ADWE-CNN. Jiayuan Xie, Yi Cai 0001, Zehang Lin, Ho-fung Leung, Qing Li 0001, Tat-Seng Chua |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Video Question Generation for Dynamic ChangesabstractVideo question generation task aims to generate meaningful questions about a video targeting an answer. Existing methods merely focus on the static appearance features in the image frames or simply identify a motion in the video to ask general questions. However, a video contains dynamically changing visual content that deserves to be questioned, e.g., changes in object motions, object states and relationships among objects, which is more practical and closer to the dynamic world we live in. In this paper, we propose a difference-aware video question generation model that aims to generate questions about temporal differences in the video, i.e., capturing the dynamic changes between image frames of a video to ask questions. To capture the dynamic changes between image frames, we utilize a temporal difference extractor to localize the differences for each frame pair of a video through an attention mechanism. Then, we introduce an answer-aware module to capture the answer-related image frame pair containing their differences for question generation, which aims to guide our model to focus on answer-related content for questioning. Finally, the output of the answer-aware module is sent to a decoder module to generate questions. Extensive experiments on SVQA and MSVD-QA datasets show that the proposed model outperforms state-of-the-art models, e.g., our model achieves at least 17.1% improvement over existing models in the SVQA dataset. This is because our model can generate questions similar to ground truths that involve changes between image frames in videos. Our code is available at https://github.com/Gary-code/D-VQG. Jiayuan Xie, Yi Cai 0001, Qingbao Huang, Qing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Knowledge-Augmented Visual Question Answering With Natural Language ExplanationabstractVisual question answering with natural language explanation (VQA-NLE) is a challenging task that requires models to not only generate accurate answers but also to provide explanations that justify the relevant decision-making processes. This task is accomplished by generating natural language sentences based on the given question-image pair. However, existing methods often struggle to ensure consistency between the answers and explanations due to their disregard of the crucial interactions between these factors. Moreover, existing methods overlook the potential benefits of incorporating additional knowledge, which hinders their ability to effectively bridge the semantic gap between questions and images, leading to less accurate explanations. In this paper, we present a novel approach denoted the knowledge-based iterative consensus VQA-NLE (KICNLE) model to address these limitations. To maintain consistency, our model incorporates an iterative consensus generator that adopts a multi-iteration generative method, enabling multiple iterations of the answer and explanation in each generation. In each iteration, the current answer is utilized to generate an explanation, which in turn guides the generation of a new answer. Additionally, a knowledge retrieval module is introduced to provide potentially valid candidate knowledge, guide the generation process, effectively bridge the gap between questions and images, and enable the production of high-quality answer-explanation pairs. Extensive experiments conducted on three different datasets demonstrate the superiority of our proposed KICNLE model over competing state-of-the-art approaches. Our code is available at https://github.com/Gary-code/KICNLE. Jiayuan Xie, Yi Cai 0001, Ruohang Xu, Jiexin Wang 0002, Qing Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Diverse Visual Question Generation Based on Multiple Objects SelectionabstractVisual question generation task aims at generating high-quality questions about a given image. To make this tak applicable to various scenarios, e.g., the growing demand for exams, it is important to generate diverse questions. The existing methods for this task control diverse question generation based on different question types, e.g., “what” and “when.” Although different question types lead to description diversity, they cannot guarantee semantic diversity when asking the same objects. Research in the field of psychology shows that humans pay attention to different objects in an image based on their preferences, which is beneficial to constructing semantically diverse questions. According to the research, we propose a multi-selector visual question generation (MS-VQG) model that aims to focus on different objects to generate diverse questions. Specifically, our MS-VQG model employs multiple selectors to imitate different humans to select different objects in a given image. Based on these different selected objects, our MS-VQG model can generate diverse questions corresponding to each selector. Extensive experiments on two datasets show that our proposed model outperforms the baselines in generating diverse questions. Wenhao Fang, Jiayuan Xie, Yi Cai 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Category-Guided Visual Question Generation (Student Abstract)abstractVisual question generation aims to generate high-quality questions related to images. Generating questions based only on images can better reduce labor costs and thus be easily applied. However, their methods tend to generate similar general questions that fail to ask questions about the specific content of each image scene. In this paper, we propose a category-guided visual question generation model that can generate questions with multiple categories that focus on different objects in an image. Specifically, our model first selects the appropriate question category based on the objects in the image and the relationships among objects. Then, we generate corresponding questions based on the selected question categories. Experiments conducted on the TDIUC dataset show that our proposed model outperforms existing models in terms of diversity and quality. Wenhao Fang, Jiayuan Xie, Yi Cai 0001 |
AAAI | 4 |
| 2023 | Deconfounded Visual Question Generation with Causal InferenceabstractVisual Question Generation (VQG) task aims to generate meaningful and logically reasonable questions about the given image targeting an answer. Existing methods mainly focus on the visual concepts present in the image for question generation and have shown remarkable performance in VQG. However, these models frequently learn highly co-occurring object relationships and attributes, which is an inherent bias in question generation. This previously overlooked bias causes models to over-exploit the spurious correlations among visual features, the target answer, and the question. Therefore, they may generate inappropriate questions that contradict the visual content or facts. In this paper, we first introduce a causal perspective on VQG and adopt the causal graph to analyze spurious correlations among variables. Building on the analysis, we propose a Knowledge Enhanced Causal Visual Question Generation (KECVQG) model to mitigate the impact of spurious correlations in question generation. Specifically, an interventional visual feature extractor (IVE) is introduced in KECVQG, which aims to obtain unbiased visual features by disentangling. Then a knowledge-guided representation extractor (KRE) is employed to align unbiased features with external knowledge. Finally, the output features from KRE are sent into a standard transformer decoder to generate questions. Extensive experiments on the VQA v2.0 and OKVQA datasets show that KECVQG significantly outperforms existing models. Zhenjun Guo, Jiayuan Xie, Yi Cai 0001, Qing Li 0001 |
ACM Multimedia | 3 |
| 2023 | Visual question generation for explicit questioning purposes based on target objects
Jiayuan Xie, Wenhao Fang, Yi Cai 0001, Qing Li 0001 |
Neural Networks | 1 |
| 2023 | Enhancing Paraphrase Question Generation With Prior KnowledgeabstractParaphrase question generation (PQG) aims to rewrite a given original question to a new paraphrase question, where the paraphrase question needs to have the same expressed meaning as the original question, but have a difference in expression form. Existing methods on PQG mainly focus on synonym substitution or word order adjustment based on the original question. However, rewriting based on the word-level may not guarantee the difference between paraphrase questions and original questions. In this paper, we propose a knowledge-aware paraphrase question generation model. Our model first employs a knowledge extractor to extract the prior knowledge related to the original question from the knowledge base. Then an attention mechanism and a gate mechanism are introduced in our model to selectively utilize the extracted prior knowledge for rewriting, which helps to expand the content of the generated question to maximize the difference. Additionally, we use a discriminator module to promote the generated paraphrase to be semantically close to the original question and the ground truth. Specifically, the loss function of the discriminator penalizes the excessive distance between the representation of the paraphrase question and the ground truth. Extensive experiments on the Quora dataset show that the proposed model outperforms the baselines. Further, our model is applied to the SQuAD dataset, which proves the generalization ability of our model in the existing QA dataset. Jiayuan Xie, Wenhao Fang, Qingbao Huang, Yi Cai 0001, Tao Wang 0036 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Scene-Text Oriented Referring Expression ComprehensionabstractReferring expression comprehension (REC) aims to identify and locate a specific object in visual scenes referred to by a natural language expression. Existing studies of REC only focus on basic visual attributes and neglect scene text. Since scene text has the functions of object identification and disambiguation, it is naturally and frequently used to refer to objects. However, existing methods do not explicitly recognize text in images and fail to align scene text mentioned in expressions with the text shown in images, resulting in object localization errors. This article takes the first step toward addressing these limitations. First, we introduce a new task called scene-text oriented referring expression comprehension, which aims to align visual cues and textual semantics of scene text with referring expressions and visual contents. Second, we propose a scene text awareness network that can bridge the gap between texts from two modalities by grounding visual representations of expression-correlated scene texts. Specifically, we propose a correlated text extraction module to solve the problem of lacking semantic understanding, and a correlated region activation module to address the fixed alignment problem and absent alignment problem. These modules ensure that the proposed method focuses on local regions that are most relevant to scene text, thus mitigating the misalignment of scene text with irrelevant regions. Third, to conduct quantitative evaluations, we establish a new benchmark dataset called RefText. Experimental results demonstrate that the proposed method can effectively comprehend scene-text oriented referring expressions and achieves excellent performance. Yuqi Bu, Liuwu Li, Jiayuan Xie, Qiong Liu 0006, Yi Cai 0001, Qingbao Huang, Qing Li 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Visual Paraphrase Generation with Key Information RetainedabstractVisual paraphrase generation task aims to rewrite a given image-related original sentence into a new paraphrase, where the paraphrase needs to have the same expressed meaning as the original sentence but have a difference in expression form. Existing studies mainly extract two semantic vectors to represent the entire image and the entire original sentence, respectively, for paraphrase generation. However, these semantic vectors for an image or a sentence may lead to the model failing to focus on some key objects in the original sentence, which may generate semantically inconsistent sentences by changing key object information. In this article, we propose an object-level paraphrase generation model, which generates paraphrases by adjusting the permutation of key objects and modifying their associated descriptions. To adjust the permutation of key objects, an object-sorting module aims to obtain new object sequences based on the key object information and original sentences. Then, a sequence generation module sequentially generates paraphrases based on the permutation of the newly object sequences. Each generation step focuses on different image features associated with different key objects to generate descriptions with differences. Furthermore, we use a semantic discriminator module to promote the generated paraphrase to be semantically close to the original sentence. Specifically, the loss function of the discriminator penalizes the excessive distance between the paraphrase and the original sentence. Extensive experiments on the MS COCO dataset show that the proposed model outperforms the baselines. Jiayuan Xie, Yi Cai 0001, Qingbao Huang, Qing Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Graph convolutional network for difficulty-controllable visual question generation
Jiayuan Xie, Yi Cai 0001, Zehang Lin, Qing Li 0001, Tao Wang 0036 |
World Wide Web (WWW) | 2 |
| 2022 | Bridging the Gap between Expression and Scene Text for Referring Expression Comprehension (Student Abstract)abstractReferring expression comprehension aims at grounding the object in an image referred to by the expression. Scene text that serves as an identifier has a natural advantage in referring to objects. However, existing methods only consider the text in the expression, but ignore the text in the image, leading to a mismatch. In this paper, we propose a novel model that can recognize the scene text. We assign the extracted scene text to its corresponding visual region and ground the target object guided by expression. Experimental results on two benchmarks demonstrate the effectiveness of our model. Yuqi Bu, Jiayuan Xie, Liuwu Li, Qiong Liu 0006, Yi Cai 0001 |
AAAI | 2 |
| 2022 | Diverse Distractor Generation for Constructing High-Quality Multiple Choice QuestionsabstractDistractor generation task aims to generate incorrect options (i.e., distractors) for multiple choice questions from an article.Existing methods for this task often utilize a standard encoder-decoder framework. However, these methods often tend to generate semantically similar distractors, since the same article representations are used to generate different distractors. Multiple generated distractors with similar semantics are considered equivalent. Because the correct answer is unique, students can eliminate these distractors even without reading the article. In this paper, we propose a multi-selector generation network (MSG-Net) that generates distractors with rich semantics based on different sentences in an article. MSG-Net adopts a multi-selector mechanism to select multiple different sentences in an article that are useful to generate diverse distractors. Specifically, a question-aware and answer-aware mechanism are introduced to assist in selecting useful key sentences, where each key sentence is coherent with the question and not equivalent to the answer. MSG-Net can generate diverse distractors based on each selected key sentence with different semantics. Extensive experiments on the RACE dataset and Cosmos QA dataset show that the proposed model outperforms the state-of-the-art models in generating diverse distractors. Jiayuan Xie, Ningxin Peng, Yi Cai 0001, Tao Wang 0036, Qingbao Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Knowledge-Based Visual Question GenerationabstractVisual question generation task aims to generate meaningful questions about an image targeting an answer. Existing methods focus on the visual concepts in the image for question generation. However, humans inevitably use their knowledge related to visual objects in images to construct questions. In this paper, we propose a knowledge-based visual question generation model that can integrate visual concepts and non-visual knowledge to generate questions. To obtain visual concepts, we utilize a pre-trained object detection model to obtain object-level features of each object in the image. To obtain useful non-visual knowledge, we first retrieve the knowledge from the knowledge-base related to the visual objects in the image. Considering that not all retrieved knowledge is helpful for this task, we introduce an answer-aware module to capture the candidate knowledge related to the answer from the retrieved knowledge, which ensures that the generated content can be targeted at the answer. Finally, object-level representations containing visual concepts and non-visual knowledge are sent to a decoder module to generate questions. Extensive experiments on the FVQA and KBVQA datasets show that the proposed model outperforms the state-of-the-art models. Jiayuan Xie, Wenhao Fang, Yi Cai 0001, Qingbao Huang, Qing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | A Double Phases Generation Network for Yes or No Question Generation (Student Abstract)abstractThis paper aims to solve the task of generating yes or no questions, which generates yes/no questions based on given passages. These questions can be used for evaluation automatically. We propose a double phases generation network that can identify specific phrases related to facts from the input passage and use them as auxiliary information for generation. Specifically, the 1st-phase prediction uses the extracted phrases as assistance to generate an initial question. Then, the 2nd-phase prediction utilizes an attention network to focus on the relevant phrases related to the initial question in the passage to generate questions that are more relevant to the specific facts contained in the initial question. Extensive experiments we performed on BoolQ dataset demonstrate the effectiveness of our framework. Jiayuan Xie, Yi Cai 0001, Zehang Lin |
AAAI | 1 |
| 2021 | Multiple Objects-Aware Visual Question GenerationabstractVisual question generation task aims to generate meaningful questions about an image according to a target answer. Existing studies mainly focus on merely one object related to the target answer in an image to generate a question. However, a target answer is often related to multiple key objects in an image, which focuses on only one object may mislead its model to generate questions that are only related to partial fragments of the answer. To address this problem, we propose a multi-objects aware generation model to capture all key objects related to an answer and generate the corresponding question. We first introduce a co-attention network to capture the relationship between each object in an image and the answer, and then extract the key objects that are related to the answer. Then, a graph network is introduced to capture the relationships between the key objects and other objects in the image that are not related to the answer, which helps generate questions that involve more visual content. Finally, the learned information from the graph network is fed into a standard decoder module to produce questions. Extensive experiments on the VQA v2.0 dataset show that the proposed model outperforms the state-of-the-art models. Jiayuan Xie, Yi Cai 0001, Qingbao Huang, Tao Wang 0036 |
ACM Multimedia | 1 |
| 2019 | A Distant Supervised Relation Extraction Model with Two Denoising StrategiesabstractDistant supervised relation extraction has been an effective way to find relational facts from text. However, distant supervised method inevitably accompanies with wrongly labeled sentences. Noisy sentences lead to poor performance of relation extraction models. Though existing piecewise convolutional neural network model with sentence-level attention (PCNN+ATT) is an effective way to reduce the effect of noisy sentences, it still has two limitations. On one hand, it adopts a PCNN module as sentence encoder, which only captures local contextual features of words and might lose important information. On the other hand, it neglects the fact that not all words contribute equally to the semantics of sentences. To address these two issues, we propose a hierarchical attention-based bidirectional GRU (HA-BiGRU) model. For the first limitation, our model utilizes a BiGRU module in place of PCNN, so as to extract global contextual information. For the second limitation, our model combines word-level and sentence-level attention mechanisms, which help get accurate sentence representations. To further alleviate the wrongly labeling problem, we first calculate the co-occurrence probabilities (CP) between the shortest dependency path (SDP) and the relation labels. Based on these co-occurrence probabilities, two denoising strategies are proposed to reduce noise interference respectively from aspect of filtering labeled data and integrating CP information into model. Experimental results on the corpus of Freebase and New York Times (Freebase+NYT) show that the HA-BiGRU model outperforms baseline models, and the two co-occurrence probabilities based denoising strategies can improve robustness of HA-BiGRU model. Zikai Zhou, Yi Cai 0001, Jiayuan Xie, Qing Li 0001, Haoran Xie 0001 |
IJCNN | 4 |