Jiuxiang You

dblp:355/8626 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
0009-0004-0093-1128ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Visual Entity-Centric Prompting for Knowledge Retrieval in Knowledge-based VQA
abstract
External knowledge provides critical clues for knowledge-based visual question answering (KB-VQA), while the implicit knowledge in images is difficult to capture in order to construct effective queries for knowledge bases. To this end, we propose a visual entity-centric prompting for knowledge retrieval (VEPR) to bridge the gap between the implicit and explicit knowledge driven by visual entities via large language models. More specifically, a visual entity question answering (EQ) module is devised to localize the critical entities from the given images and questions to generate entity-centric questions via a large language model. In particular, EQ obtains several entity-centric question-answer pairs via a visual language model. Furthermore, a question-answer-enhanced retrieval (ER) module is devised to construct a query by summarizing the text including question-answer pairs, captions and questions, in order to require explicit knowledge items. Finally, a multi-branch reader (MR) module is designed to encode the given questions, visual content and retrieved knowledge items, which are decoded to make answer predictions. Extensive experiments conducted on two public datasets demonstrate the effectiveness of the VEPR.
Jiuxiang You, Ziyue Qiu, Guobo Xie, Yi Yu 0001, Zhenguo Yang
ICASSP2
2025 A Multi-Hop Graph Reasoning Network for Knowledge-Based VQA
abstract
Knowledge-based visual question answering (KB-VQA) requires reasoning about the visual grounding relations between the images and questions by incorporating external knowledge. Existing works typically retrieve knowledge from knowledge graphs by leveraging global multimodal representations of image–text pairs for graph convolution, which neglect contextual clues at hop granularity, resulting in suboptimal spreading and leveraging of contextual information. To this end, we propose a multi-hop graph reasoning network (MGRN) for KB-VQA, which consists of a knowledge graph constructor (KGC) module, a semantic-instructed graph reasoning (SGR) module, and an answering module. MGRN exploits multimodal semantics from given images and questions as instructions for graph reasoning to obtain the knowledge representation from either the scene graph or knowledge base. Specifically, KGC fuses the scene graph with triplets from ConceptNet and Comet to construct a contextual knowledge graph for retrieving knowledge representation. Furthermore, SGR conducts multi-hop graph reasoning to select top- K knowledge items for answering by passing and filtering interplay messages on contextual knowledge graphs under the guidance of multimodal semantic representation. Extensive experiments conducted on two public datasets show the effectiveness and outperformance of our method.
Jiuxiang You, Zhenguo Yang, Xiaoping Li 0001, Haoran Xie 0001, Qing Li 0001, Wenyin Liu
ACM Trans. Intell. Syst. Technol.2
2024 Modality-specific and -shared Contrastive Learning for Sentiment Analysis
abstract
In this paper, we propose a two-stage network with modality-specific and -shared contrastive learning (MMCL) for multimodal sentiment analysis. MMCL comprises a category-aware modality-specific contrastive (CMC) module and a self-decoupled modality-shared contrastive (SMC) module. In the first stage, the CMC module guides the encoders to extract modality-specific representations by constructing positive-negative pairs according to sample categories. In the second stage, the SMC module guides the encoders to extract modality-shared representations by constructing positive-negative pairs based on modalities and decoupling the self-contrast of all modalities. In the aforementioned modules, we leverage self-modulation factors to focus more on hard positive pairs through assigning different loss weights to positive pairs depending on their distance. In particular, we introduce a dynamic routing algorithm to cluster the inputs of the contrastive modules during training, where a gradient stopping strategy is utilized to isolate the backpropagation process of the CMC and SMC modules. Extensive experiments on the CMU-MOSI and CMU-MOSEI datasets show that MMCL achieves the state-of-the-art performance.
Dahuang Liu, Jiuxiang You, Guobo Xie, Lap-Kei Lee, Fu Lee Wang, Zhenguo Yang
ICMR2
2023 Multimodal Conditional VAE for Zero-Shot Real-World Event Discovery
Zhuopan Yang, Jiuxiang You, Zhenguo Yang
ADMA (2)3
2023 A Retriever-Reader Framework with Visual Entity Linking for Knowledge-Based Visual Question Answering
abstract
In this paper, we propose a Retriever-Reader framework with Visual Entity Linking (RR-VEL) for knowledge-based visual question answering. Given images and original questions, the visual entity linking (VEL) module extracts key entities in images to replace the question referents for semantic disambiguation, achieving entity-oriented queries with explicit entities. Furthermore, the Retriever encodes the queries and knowledge items by Bert with a feed-forward layer, and obtains a set of knowledge candidates. The Reader encodes the questions with image captions and knowledge candidates in two branches, which avoids their interference during self-attentive encoding. Finally, the decoder of Reader fuses the encoded features to generate answers. Extensive experiments conducted on the two public datasets show that our method significantly outperforms the existing baselines.
Jiuxiang You, Zhenguo Yang, Qing Li 0001, Wenyin Liu
ICME1
2023 Meta Perturbation Generation Network for Text-Based CAPTCHA
Zhuoting Wu, Zhiwei Guo 0001, Jiuxiang You, Zhenguo Yang, Qing Li 0001, Wenyin Liu
SecureComm (1)3
2023 Event-Oriented Visual Question Answering: The E-VQA Dataset and Benchmark
abstract
Visual question answering (VQA) is a challenging task that reasons over questions on images with knowledge. A prerequisite for VQA is the availability of annotated datasets, while the available datasets have several limitations. 1) The diversity of questions and answers are limited to a few question categories and certain concepts (e.g., objects, relations, actions.) with somewhat mechanical answers. 2) The availability of background knowledge or context information has been disregarded with just images, questions and answers being provided. 3) The timeliness of knowledge has not been examined, though some works may introduce factual or commonsense knowledge bases, e.g., ConceptNet, DBPedia. In this paper, we provide an Event-oriented Visual Question Answering (E-VQA) dataset including free-form questions and answers for real-world event concepts, which provides context information of events as domain knowledge in addition to images. E-VQA consists of 2,690 social media images, 9,088 questions, 5,479 answers, and 1,157 news media articles for references being annotated to 182 real-world events, covering a wide range of topics, such as armed conflicts and attacks, disasters and accidents, law and crime. For comparisons, we investigate 10 state-of-the-art VQA methods as benchmarks.
Zhenguo Yang, Jiale Xiang, Jiuxiang You, Qing Li 0001, Wenyin Liu
IEEE Trans. Knowl. Data Eng.3