EDBT 2026 Demo / reviewers in the wild / expert
Ning Xu 0003
dblp:04/5856-3
· DBLP profile ↗
56ranked-venue papers
16as first author
41since 2021 · last 2027
0000-0002-7526-4356ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 13 first-author · 31 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 9 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Beyond content: A dual-channel approach for social bot detection via unmasking behavioral sequence camouflage
Hongshuo Tian, Jinlin Guo, Xianzhu Liu, Ning Xu 0003, Lanjun Wang |
Expert Syst. Appl. | 6 |
| 2026 | Towards social-aware image captioning via chain-of-thought prompting
Shenyuan Zhang, Ning Xu 0003, Quanhan Wu, Jinlin Guo, Hongshuo Tian, Anan Liu |
Expert Syst. Appl. | 2 |
| 2026 | Medical VLP Model Is Vulnerable: Toward Multimodal Adversarial Attack on Large Medical Vision-Language ModelsabstractMedical Visual Question Answering (Medical VQA) is an essential task that facilitates the automated interpretation of complex clinical imagery with corresponding textual questions, thereby supporting both clinicians and patients in making informed medical decisions. With the rapid progress of Vision-Language Pretraining (VLP) in general domains, the development of medical VLP models has emerged as a rapidly growing interdisciplinary area at the intersection of artificial intelligence (AI) and healthcare. However, few works have been proposed to evaluate the adversarial robustness of medical VLP models, which faces two primary challenges: (1) the complexity of medical texts, stemming from the presence of terminologies, poses significant challenges for models in comprehending the text for adversarial attack; (2) the diversity of medical images arises from the variety of anatomical regions depicted, which requires models to determine critical anatomical regions for attack. In this paper, we propose a novel multimodal adversarial attack generator for evaluating the robustness of medical VLP models. Specifically, for the complexity of medical texts, we integrate medical knowledge when crafting text adversarial samples, which can facilitate the terminologies understanding and adversarial strength; for the diversity of medical images, we divide the anatomical regions into either global or local regions in medical images, which are determined by learned balance weights for perturbations. Our experimental study not only provides a quantitative understanding in medical VLP models, but also underscores the critical need for thorough safety evaluations before implementing them in real-world medical applications. Zimu Lu, Ning Xu 0003, Hongshuo Tian, Lanjun Wang, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Constituency-Tree-Induced Vision-Language Alignment for Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs) integrate sophisticated large vision models (LVMs) to empower large language models (LLMs) with vision ability to perceive, reason, and interact in vision-language (V-L) tasks, while the modality bridge between two specialists becomes the bottleneck that translates visual signals into linguistic representations. However, most of the existing methods train the modality bridge with coarse-grained image-text pairs, neglecting the structural mapping between V-L semantics that facilitates modality translation from LVMs to LLMs. To mitigate this, we propose a Constituency-Tree-Induced Multimodal Bridging mechanism (CTIMB) that learns the fine-grained connection from LVMs to LLMs by the structural guidance from multi-modal constituency tree. Our approach consists of: 1) the multi-modal constituency-tree parser that jointly exploits the semantic structure of vision and language; 2) the lightweight connector that translates visual signals into linguistic representation and re-arranges them according to the constituency-tree structure; 3) the dynamic construction loss that aids in aligning the semantic structures derived from the tree parser and the connector. The CTIMB can learn the fine-grained mapping between visual and linguistic semantics, seamlessly bridge the LVMs and LLMs to enhance V-L tasks, and is more cost-efficient compared with current methods. Extensive experiments have demonstrated that our method more accurately interprets the visual features, enabling LLMs to conduct downstream tasks more effectively, and achieve superior performance with less training cost. Yingchen Zhai, Ning Xu 0003, Hongshuo Tian, Bolun Zheng, Chenggang Yan 0001, Jinbo Cao, Rongbao Kang, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | MMToT: Multi-Modal Token-of-Thought Reasoning for Large ModelsabstractWith the development of Large models (LMs), recent methods tend to leverage them to complete various downstream tasks, like VQA and image caption. A typical method is Chain-of-Thought (CoT) prompting, which improves the reasoning abilities of LMs by providing intermediate steps. However, existing CoT prompts have two main drawbacks: 1) They typically represent CoT through multiple sentences, which often introduce irrelevant textual or visual context that may confuse LMs. 2) The current CoT methods fail to consider how the contributions of different tokens vary for answer inference. For this, we propose the Multi-Modal Token-of-Thought (MMToT), a novel token-level prompt method to improve LMs' multi-modal reasoning capabilities. Furthermore, MMToT stress on two strengths against CoT: 1) To prevent LMs from being affected by irrelevant contexts, we propose to extract explicit multi-modal tokens rather than sentences to construct MMToT, enhancing the reliability of generated answers. 2) To ensure that LMs prioritize tokens with high contribution scores during answer generation, we propose a confident decision-making module to evaluate and integrate each token's contribution in MMToT. Compared to existing methods, the proposed MMToT demonstrates superior performance on Science-QA, MATH, OKVQA, and VQA-introspect datasets. Furthermore, ablation studies and visualization results validate the effectiveness and interpretability of MMToT. Ning Xu 0003, Zimu Lu, Hongshuo Tian, Bolun Zheng, Jinbo Cao, Anan Liu |
IEEE Trans. Multim. | 1 |
| 2026 | How to Understand Named Entities: Using Commonsense for News CaptioningabstractNews captioning aims to describe an image with its news article body as input. It greatly relies on a set of detected named entities, including real-world people, organizations, and places. This article exploits commonsense knowledge to understand named entities for news captioning. By “understand,” we mean correlating the news content with commonsense in the wild, which helps an agent to (1) distinguish semantically similar named entities and (2) describe named entities using words outside of training corpora. Our approach consists of three modules: (a) Filter Module aims to clarify the commonsense concerning a named entity from two aspects: what does it mean ? and what is it related to ?, which divide the commonsense into explanatory knowledge and relevant knowledge , respectively. (b) Distinguish Module aggregates explanatory knowledge from node-degree , dependency , and distinguish three aspects to distinguish semantically similar named entities. (c) Enrich Module attaches relevant knowledge to named entities to enrich the entity description by commonsense information (e.g., identity and social position). Finally, all of information is integrated into the large multimodal model to generate the news caption. Extensive experiments on two challenging datasets (i.e., GoodNews and NYTimes) demonstrate the superiority of our method. Ablation studies and visualization further validate its effectiveness in understanding named entities. Shenyuan Zhang, Ning Xu 0003, Yanhui Wang 0001, Tongle Ma, Wu Liu 0005, Jinlin Guo, Anan Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2026 | Knowledge and multi-detail enhanced GAN for human-driven text-to-image synthesisabstractHuman-driven text-to-image synthesis aims to create controllable images, which not only adhere to the semantic of given text but also incorporate the visual characteristics of given human. For example, given “a man on the beach” (text) along with a photo of human, the model aims to generate an image depicting the human on the beach. Although current diffusion-based methods have shown promise in this task, they face two major limitations: (1) The generated images appear to be a bit stiff and unnatural, almost like collages of human and backgrounds; (2) The details of human in the generated image are inconsistent with those in the input, losing the original identity. To address these issues, we present the Knowledge and Multi-Detail Enhanced GAN for the task of human-driven text-to-image synthesis. It employs external knowledge as references to improve the harmony between human and backgrounds, and uses CLIP’s multi-layer features to intensify human details. First, we search the database to retrieve external images that are similar to the given text, serving as our knowledge. Second, to preserve the human details, we present the Multi-Detail Enhancer, which uses the image encoder of CLIP to extract human representation at multiple levels. Third, to enhance the human-background naturalness, we present the Knowledge Attention Enhancer, which can seamlessly blend human, text, and knowledge by attentively retain useful information and filter out noise from knowledge. Finally, we introduce the dual discriminators to guide the entire network, which can facilitate the accurate capture of human details and generation of images. Extensive experiments demonstrate the superiority of our method with its efficiency and lower computational demands. It is about 300 times faster than diffusion-based models, uses only 5% of the parameters, and completes training in just two days on three V100 GPUs. Ning Xu 0003, Zhewen Shen, Hongshuo Tian, Bolun Zheng, Chenggang Yan 0001, Jinbo Cao, Rongbao Kang, Anan Liu |
Vis. Informatics | 1 |
| 2025 | SMTPD: A New Benchmark for Temporal Prediction of Social Media PopularityabstractSocial media popularity prediction task aims to predict the popularity of posts on social media platforms, which has a positive driving effect on application scenarios such as content optimization, digital marketing and online advertising. Though many studies have made significant progress, few of them pay much attention to the integration between popularity prediction with temporal alignment. In this paper, with exploring YouTube’s multilingual and multi-modal content, we construct a new social media temporal popularity prediction benchmark, namely SMTPD, and suggest a baseline framework for temporal popularity prediction. Through data analysis and experiments, we verify that temporal alignment and early popularity play crucial roles in social media popularity prediction for not only deepening the understanding of temporal dynamics of popularity in social media but also offering a suggestion about developing more effective prediction models in this field. Code is available at https://github.com/zhuwei321/SMTPD Yijie Xu, Bolun Zheng, Hangjia Pan, Yuchen Yao, Ning Xu 0003, Anan Liu, Chenggang Yan 0001 |
CVPR | 6 |
| 2025 | Mixture of causal experts: A causal perspective to build dual-level mixture-of-experts models
Ning Xu 0003, Hongshuo Tian, Bolun Zheng, Jinbo Cao, Anan Liu |
Expert Syst. Appl. | 2 |
| 2025 | Counterfactual GAN for debiased text-to-image synthesis
Xianghua Kong, Ning Xu 0003, Zefang Sun, Zhewen Shen, Bolun Zheng, Chenggang Yan 0001, Jinbo Cao, Rongbao Kang, Anan Liu |
Multim. Syst. | 2 |
| 2025 | Enriched Image Captioning Based on Knowledge Divergence and FocusabstractImage captioning is a fundamental task in computer vision that aims to generate precise and comprehensive descriptions of images automatically. Intuitively, humans initially rely on the image content, e.g., “cake on a plate”, to gradually gather relevant knowledge facts e.g., “birthday party”, “candles”, which is a process referred to as divergence. Then, we perform step-by-step reasoning based on the images to refine, and rearrange these knowledge facts for explicit sentence generation, a process referred to as focus. However, existing image captioning methods mainly rely on the encode-decode framework that does not well fit the “divergence-focus” nature of the task. To this end, we propose the knowledge “divergence-focus” method for Image Captioning (K-DFIC) to gather and polish knowledge facts for image understanding, which consists of two components: (a) Knowledge Divergence Module aims to leverage the divergence capability of large-scale pre-trained model to acquire knowledge facts relevant to the image content. To achieve this, we design a scene-graph-aware prompt that serves as a “trigger” for GPT-3.5, encouraging it to “diverge” and generate more sophisticated, human-like knowledge. (b) Knowledge Focus Module aims to refine acquired knowledge facts and further rearrange them in a coherent manner. We design the interactive refining network to encode knowledge, which is refined with the visual features to remove irrelevant words. Then, to generate fluent image descriptions, we design the large-scale pre-trained model-based rearrangement method to estimate the importance of each knowledge word for an image. Finally, we fuse the refined knowledge and visual features to assist the decoder in generating captions. We demonstrate the superiority of our approach through extensive experiments on the MSCOCO dataset. Our approach surpasses state-of-the-art performance across all metrics in the Karpathy split. For example, our model obtains the best CIDEr-D score of 148.4%. Additional ablation studies and visualization further validate our effectiveness. Anan Liu, Quanhan Wu, Ning Xu 0003, Hongshuo Tian, Lanjun Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Model Can Be Subtle: Two Important Mechanisms for Social Media Popularity PredictionabstractSocial media popularity prediction is an important channel to explore content sharing and communication on social networks. It aims to capture informative cues by analyzing multi-type data (such as user profile, image, and text) to decide the popularity of a specified post. In this article, we divide social network users into two categories (i.e., active and inactive users) and find a dilemma in existing models: If an active user publishes the low-popularity post, the model will habitually predict the high score. On the contrary, if an inactive user provides the high-popularity post, the model still gives the low score incorrectly. Therefore, how to make the model more subtle to users is important. Comparing to existing methods that directly leverage multi-modal features for regression training, this article stresses more on two novel mechanisms. The first method aims to prevent the over-fitting on user IDs. We propose the attribute-sensitive interactive mechanism (M1) by incorporating explicit user-attribute and post-attribute interaction. It can analyze which type of features a user cares the most and weaken the model’s dependence on user IDs. The second method aims to strengthen the influence of post content. We propose the knowledge embedding mechanism (M2) to revise the popularity scores in existing models by fusing the statistical frequency over multi-type data. Note that both mechanisms are model-agnostic, which can be applicable in any popularity prediction model. Extensive experiments conducted on the Social Media Prediction Dataset further validate the effectiveness. Ning Xu 0003, Jing Liu 0002, Lanjun Wang, Xuanya Li, Mengxiao Zhu 0001, Yongdong Zhang 0001, Anan Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Rumor Detection Framework Based on Multi-source Knowledge AdaptationabstractRumors proliferating on social media pose significant risks to politics, economics, and society. A considerable amount of research focuses on improving automated rumor detection models on dedicated datasets, but their performance often plummets when applied to novel, unforeseen events due to varying data distributions. Inspired by domain adaptation and few-shot learning, we propose the Multi-Source Knowledge Adaptation framework MSKA. It is a plug-and-play framework where any rumor detection model can be integrated to enhance its performance in detecting contemporary events. Specifically, MSKA relies on event-level knowledge adaptation and cross-domain category alignment to compel the rumor detection model to learn more discriminative knowledge representations. Experiments on the public dataset demonstrate that MSKA improves the performance of the base rumor detection model in detecting emerging events and outperforms other baselines. In terms of accuracy, MSKA achieves at least a 7% improvement over the base rumor detection model. Ning Xu 0003, Jingqiu Li, Lanjun Wang, Anan Liu |
ICME | 1 |
| 2024 | Cross-Modal Coherence-Enhanced Feedback Prompting for News CaptioningabstractNews Captioning involves generating the descriptions for news images based on the detailed content of related news articles. Given that these articles often contain extensive information not directly related to the image, captions may end up misaligned with the visual content. To mitigate this issue, we propose the novel cross-modal coherence-enhanced feedback prompting method to clarify the crucial elements that align closely with the visual content for news captioning. Specifically, we first adapt CLIP to develop a news-specific image-text matching module, enriched with insights from language model MPNet using a matching-score comparative loss, which facilitates effective cross-modal knowledge distillation. This module enhance the coherence between images and each news sentences via rating confidence. Then, we design confidence-aware prompts to fine-tune LLaVA model with by LoRa strategy, focusing on essential details in extensive articles. Lastly, we evaluate the generated news caption with refined CLIP, constructing confidence-feedback prompts to further enhance LLaVA through feedback learning, which iteratively refine captions to improve its accuracy. Extensive experiments conduct on two public datasets, GoodNews and NYTimes800k, have validated the effectiveness of our method. Ning Xu 0003, Hongshuo Tian, Anan Liu |
ACM Multimedia | 1 |
| 2024 | Prior knowledge guided text to image generation
Anan Liu, Zefang Sun, Ning Xu 0003, Rongbao Kang, Jinbo Cao, Weijun Qin, Shenyuan Zhang, Xuanya Li |
Pattern Recognit. Lett. | 3 |
| 2024 | Gaussian Distribution-Aware Commonsense Knowledge Learning for Scene Graph GenerationabstractKnowledge-based Scene Graph Generation (SGG) requires external commonsense knowledge beyond the visual scene to infer the relation between objects. Such knowledge can be obtained in a variety of forms, such as vision, text, and graph. However, there are two drawbacks as follows: 1) commonsense knowledge essentially has uncertainty, but current works usually represent knowledge in a deterministic manner, which is not well matched to its nature, 2) using commonsense knowledge without denoising will introduce irrelevant information. This can increase the burden on the relation classifier and only obtain marginal gains over a large amount of data. In this paper, we propose a novel Gaussian distribution-aware commonsense knowledge learning method for SGG. First, we associate each object pair with a Gaussian distribution, which parametrizes visual context and commonsense as mean and variance, respectively. We prove that Gaussian modeling can provide a probabilistic soft space to measure the uncertainty of external knowledge, which allows diverse predictions. Second, to reduce semantic noise in commonsense, we sample multiple variables from the Gaussian distribution and train multi-expert classifiers, which can be dynamically examined for the ensemble softmax classification. Extensive comparative experiments on two benchmarks confirm that our method can achieve competitive performance against the state-of-the-art. Ablation studies verify the essential roles of individual components. Moreover, the visualization of multi-expert classifiers confirms our ability to integrate commonsense for relation inference. Hongshuo Tian, Ning Xu 0003, Mohan Kankanhalli, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Rule-Driven News CaptioningabstractNews captioning task aims to generate sentences by describing named entities or concrete events for an image with its news article. Existing methods have achieved remarkable results by relying on the large-scale pre-trained models, which primarily focus on the correlations between the input news content and the output predictions. However, the news captioning requires adhering to some fundamental rules of news reporting, such as accurately describing the individuals and actions associated with the event. In this paper, we propose the rule-driven news captioning method, which can generate image descriptions following designated rule signal. Specifically, we first design the news-aware semantic rule for the descriptions. This rule incorporates the primary action depicted in the image (e.g., “performing”) and the roles played by named entities involved in the action (e.g., “Agent” and “Place”). Second, we inject this semantic rule into the large-scale pre-trained model, BART, with the prefix-tuning strategy, where multiple encoder layers are embedded with news-aware semantic rule. Finally, we can effectively guide BART to generate news sentences that comply with the designated rule. Extensive experiments on two widely used datasets (i.e., GoodNews and NYTimes800k) demonstrate the effectiveness of our method. Ning Xu 0003, Hongshuo Tian, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Multi-Modal Validation and Domain Interaction Learning for Knowledge-Based Visual Question AnsweringabstractKnowledge-based Visual Question Answering (KB-VQA) aims to answer the image-aware question via the external knowledge, which requires an agent to not only understand images but also explicitly retrieve and integrate knowledge facts. Intuitively, to accurately answer the question, we humans can validate the retrieved knowledge based on our memory, and then align the knowledge facts with the image regions to infer answers. However, most existing methods ignore the process of knowledge validation and alignment. In this paper, we propose the Multi-Modal Validation and Domain Interaction Learning method, which consists of two components: 1) Multi-modal validation for knowledge retrieval. We propose the multi-modal validation module (MMV) to evaluate the confidence of each retrieved knowledge fact via images and questions, which preserves knowledge candidates effective for inferring answers. 2) Domain interaction for knowledge integration. We propose the Domain Interaction TRansformer module (DI-TR) to align visual regions with knowledge facts by the interaction learning in the improved transformer. Specifically, the inter-domain and intra-domain masks are injected into each self-attention layer to control the integration scope. The proposed method outperforms several strong baselines on three widely-used knowledge-based datasets: KRVQA, OK-VQA and VQA2.0. Extensive experiments and ablation studies demonstrate the effectiveness of multi-modal knowledge validation and domain interaction learning. Ning Xu 0003, Anan Liu, Hongshuo Tian, Yongdong Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Counterfactual Visual Dialog: Robust Commonsense Knowledge Learning From Unbiased TrainingabstractVisual Dialog (VD) requires an agent to answer the current question by engaging in a conversation with humans referring to an image. Despite the recent progress, it is beneficial to introduce external commonsense knowledge to fully understand the given image and dialog history. However, the existing knowledge-based VD models are inclined to rely on severe learning bias brought by commonsense, e.g., the retrieved$< {\mathtt{{bus}}}, {\mathtt{capable\;of}}, {\mathtt{transport\;people}}>$,$< {\mathtt{{bus}}}, {\mathtt{is\;a}}, {\mathtt{public\;transport}}>$, and$< {\mathtt{{bus}}}, {\mathtt{is\;a}}, {\mathtt{car}}>$can induce a spurious correlation between the question “What is the bus used for?” and the false answer “City bus”. There are two challenges to make commonsense learning more robust against spurious correlations: 1) how to disentangle the true effect of “good” commonsense knowledge from the whole, and 2) how to estimate and remove the effect of “bad” commonsense bias on answers. In this article, we propose a novel CounterFactual Commonsense learning scheme for the Visual Dialog task (CFC-VD). First, comparing with the causal graph of existing VD models, we add one new commonsense node and one new link to multi-modal information from history, question, and image. Since the retrieved knowledge prior is subtle and uncontrollable, we consider it as an unobserved confounder in the commonsense node, which leads to spurious correlations for the answer inference. Then, to remove the effect of the confounder, we formulate it as the direct causal effect of commonsense on answers and remove the direct language effect by subtracting it from the total causal effect via counterfactual reasoning. Experimental results certify the effectiveness of our method on the prevailing Visdial v0.9 and Visdial v1.0 datasets. Anan Liu, Ning Xu 0003, Hongshuo Tian, Jing Liu 0002, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Event-Aware Retrospective Learning for Knowledge-Based Image CaptioningabstractExternal knowledge has been widely applied in image captioning tasks to enrich the generated sentences. However, existing methods retrieve knowledge by considering only semantic relevance while ignoring whether they are useful for captioning. For example, when querying “person” in external knowledge, the most relevant concepts may be “wearing shirt” or “riding horse” statistically, which are not consistent with image contents and introduce noise to generated sentences. Intuitively, we humans can iteratively correlate visual clues with corresponding knowledge to distinguish useful clues from noise. Therefore, we propose an event-aware retrospective learning network for knowledge-based image captioning, which employs a retrospective validation mechanism on captioning models to align the retrieved knowledge with visual contents. This approach is an event-aware perspective and helps select useful knowledge that corresponds to visual facts. To better align images and knowledge, 1) we design an event-aware retrieval algorithm that clusters word-centered knowledge into triplet-centered knowledge (i.e., from “” to “- edge -”, which provides an event context to facilitate knowledge retrieval and validation. 2) We revisit image contents to retrospectively validate retrieved knowledge by aligning the visual representation between knowledge and image. We summarize the visual characteristics of each knowledge event from the visual genome dataset to help learn which knowledge does not exist in the visual scene and should be discarded. 3) We adopt a dynamic knowledge fusion module that calibrates image and knowledge representations for sentence generation, which includes a knowledge-controlled gate unit that jointly calculates visual and semantic features in event-aware patterns. Compared to current knowledge-based captioning methods, the proposed network retrospectively learns the visual facts by event-aware retrieval and knowledge-image visual alignment, which regularizes the knowledge-incorporated captioning with visual evidence. Extensive experiments on the MS-COCO dataset demonstrate the effectiveness of our method. Ablation studies and visualization demonstrate the advantages of each component of the proposed model. Anan Liu, Yingchen Zhai, Ning Xu 0003, Hongshuo Tian, Weizhi Nie, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Learning to Supervise Knowledge Retrieval Over a Tree Structure for Visual Question AnsweringabstractKnowledge-based visual question answering (KBVQA) aims to retrieve the external knowledge out of images to answer questions. However, current methods always introduce various irrelevant knowledge due to two drawbacks: (1) Synonymy issue. Existing methods heavily rely on words from questions or object labels in images to match knowledge from databases, which disregards the same word may hold multiple meanings within different contexts. (2) Knowledge uncertainty issue. Due to the absence of supervisory signals, recent methods can not determine which knowledge is applicable for answer inference, which can mislead to admit useless knowledge. To address these two problems, we propose to supervise the process of knowledge retrieval over a tree structure for KB-VQA task. For the synonymy issue, we construct a hierarchical knowledge tree to capture the subordination information between knowledge facts, mitigating the impact of synonyms on knowledge retrieval. For the knowledge uncertainty issue, we use the retrieval history as the ground truth to supervise the knowledge retrieval, which facilitates the QA model to form an explicit path of knowledge facts for answer understanding. Finally, we integrate the image, question, and retrieved knowledge into a variant of transformer to predict answers. Experimental results validate the effectiveness of the proposed method on KR-VQA, OK-VQA and VQA v2 datasets. Ning Xu 0003, Zimu Lu, Hongshuo Tian, Rongbao Kang, Jinbo Cao, Yongdong Zhang 0001, Anan Liu |
IEEE Trans. Multim. | 1 |
| 2024 | Multi-stage reasoning on introspecting and revising bias for visual question answeringabstractVisual Question Answering (VQA) is a task that involves predicting an answer to a question depending on the content of an image. However, recent VQA methods have relied more on language priors between the question and answer rather than the image content. To address this issue, many debiasing methods have been proposed to reduce language bias in model reasoning. However, the bias can be divided into two categories: good bias and bad bias. Good bias can benefit to the answer prediction, while the bad bias may associate the models with the unrelated information. Therefore, instead of excluding good and bad bias indiscriminately in existing debiasing methods, we proposed a bias discrimination module to distinguish them. Additionally, bad bias may reduce the model’s reliance on image content during answer reasoning and thus attend little on image features updating. To tackle this, we leverage Markov theory to construct a Markov field with image regions and question words as nodes. This helps with feature updating for both image regions and question words, thereby facilitating more accurate and comprehensive reasoning about both the image content and question. To verify the effectiveness of our network, we evaluate our network on VQA v2 and VQA cp v2 datasets and conduct extensive quantity and quality studies to verify the effectiveness of our proposed network. Experimental resu- lts show that our network achieves significant performance against the previous state-of-the-art methods. Anan Liu, Zimu Lu, Ning Xu 0003, Min Liu 0008, Chenggang Yan 0001, Bolun Zheng, Yulong Duan, Xuanya Li |
ACM Trans. Web | 3 |
| 2024 | A fine-grained deconfounding study for knowledge-based visual dialogabstractKnowledge-based Visual Dialog is a challenging vision-language task, where an agent engages in dialog to answer questions with humans based on the input image and corresponding commonsense knowledge. The debiasing methods based on causal graphs have gradually sparked much attention in the field of Visual Dialog (VD), yielding impressive achievements. However, existing studies focus on the coarse-grained deconfounding, which lacks a principled analysis of the bias. In this paper, we propose a fined-grained study of deconfounding on: (1) We define the confounder from two perspectives. The first is user preference (denoted as U h ), derived from human-annotated dialog history, which may introduce spurious correlations between questions and answers. The second is commonsense language bias (denoted as U c ), where certain words appear so frequently in the retrieved commonsense knowledge that the model tends to memorize these patterns, thereby establishing spurious correlations between the commonsense knowledge and the answers. (2) Given that the current question directly influences answer generation, we further decompose the confounders into U h 1 , U h 2 and U c 1 , U c 2 , based on their relevance to the current question. Specifically, U h 1 and U c 1 represent dialog history and high-frequency words that are highly correlated with the current question, while U h 2 and U c 2 are sampled from dialog history and words with low relevance to the current question. Through a comprehensive evaluation and comparison of all components, we demonstrate the necessity of jointly considering both U h and U c . Fine-grained deconfounding, particularly with respect to the current question, proves to be more effective. Ablation studies, quantitative results, and visualizations further confirm the effectiveness of the proposed method. Anan Liu, Quanhan Wu, Xianzhu Liu, Ning Xu 0003 |
Vis. Informatics | 6 |
| 2023 | Towards Confidence-Aware Commonsense Knowledge Integration for Scene Graph GenerationabstractCommonsense knowledge has been widely explored to improve Scene Graph Generation (SGG). Existing methods simply incorporate the described relations of knowledge bases into each part of the scene for a concrete understanding. However, they ignore the discussion about whether a visual scene needs to associate commonsense knowledge for making inferences. Specifically, the difficulty of relation recognition varies from its type. Some frequent spatial relations (e.g. on) usually produce less perception error even without any prior information, while others involved many rules and patterns (e.g. throwing) possess few samples and require to combine with some commonsense knowledge as supplementary. In this paper, we propose a novel confidence-aware commonsense knowledge integration for SGG. Firstly, we depend on mutual information maximization to design a hybrid-attention module, which decreases the uncertainty in representation learning given external knowledge. Second, we introduce an extra branch for SGG network to perform confidence estimation independent of any ground truth labels, in which the output scalar explicitly reflects the difficulty of visual recognition. This value is equipped with the ability to balance the demand for commonsense knowledge in a given scene. Experiments are conducted with the backbone of MOTIFS on Visual Genome (VG) and our method effectively promotes the metric of mRecall with little performance hit for metric Recall, especially for predicting unseen relations. Hongshuo Tian, Ning Xu 0003, Yanhui Wang 0001, Chenggang Yan 0001, Bolun Zheng, Xuanya Li, Anan Liu |
ICME | 2 |
| 2023 | Knowledge Prompt Makes Composed Pre-Trained Models Zero-Shot News CaptionerabstractNews image captioning aims to generate descriptions containing concrete named entities for news images by leveraging relevant news articles. However, existing approaches suffer from two shortcomings: 1) lack of commonsense knowledge required to understand named entities, and 2) limited multimodal context modeling capabilities. In this paper, we propose to migrate the ability of large-scale pre-trained models for news image captioning. To acquire factual knowledge for describing named entities, we induce a pre-trained language model for commonsense knowledge reasoning using context-aware knowledge prompts. To compose a new multimodal context modeling capability, we coordinate pre-trained models by a unified language representation and constrain joint multimodal context reasoning with cross-modal consistency objective. Experimental results on GoodNews and NYTimes datasets show that our proposed method exhibits considerable captioning capabilities even without training on news data. Yanhui Wang 0001, Ning Xu 0003, Hongshuo Tian, Yulong Duan, Xuanya Li, Anan Liu |
ICME | 2 |
| 2023 | Exploring visual relationship for social media popularity prediction
Anan Liu, Ning Xu 0003, Shenyuan Zhang, Yejun Tang, Xuanya Li |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | SMPC: boosting social media popularity prediction with caption
Anan Liu, Ning Xu 0003, Jing Liu 0002, Yuting Su 0001, Shenyuan Zhang, Yejun Tang, Junbo Guo, Guoqing Jin, Xuanya Li |
Multim. Syst. | 3 |
| 2023 | A comprehensive survey on deep-learning-based visual captioning
Bowen Xin, Ning Xu 0003, Yingchen Zhai, Zimu Lu, Jing Liu 0002, Weizhi Nie, Xuanya Li, Anan Liu |
Multim. Syst. | 2 |
| 2022 | Closed-loop reasoning with graph-aware dense interaction for visual dialog
Anan Liu, Ning Xu 0003, Junbo Guo, Guoqing Jin, Xuanya Li |
Multim. Syst. | 3 |
| 2022 | Region-Aware Image Captioning via Interaction LearningabstractImage captioning is one of the primary goals in computer vision which aims to automatically generate natural descriptions for images. Intuitively, human visual system can notice some stimulating regions at first glance, and then volitionally focus on interesting objects within the region. For example, to generate a free-form sentence about “boy-catch-baseball”, the visual region involving “boy” and “baseball” could be first attended and then guide the salient object discovery for the word-by-word generation. Till now, previous captioning works mainly rely on the object-wise modeling and ignore the rich regional patterns. To mitigate the drawback, this paper proposes the region-aware interaction learning method, which aims to explicitly capture the semantic correlations in the region and object dimensions for the word inference. First, given an image, we extract a set of regions which contain diverse objects and their relations. Second, we present the spatial-GCN interaction refining structure which can establish the connection between regions and objects to effectively capture contextual information. Third, we design the dual-attention interaction inference procedure, which enables attention to be calculated in region and object dimensions jointly for the word generation. Specifically, the guidance mechanism is proposed to selectively emphasize semantic inter-dependencies from region to object attentions. Extensive experiments on the MSCOCO dataset demonstrate the superiority of the proposed method. Additional ablation studies and visualization further validate its effectiveness. Anan Liu, Yingchen Zhai, Ning Xu 0003, Weizhi Nie, Wenhui Li 0001, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | High-Order Interaction Learning for Image CaptioningabstractImage captioning aims at understanding various semantic concepts (e.g., objects and relationships) from an image and integrating them in a sentence-level description. Hence, it is necessary to learn the interaction among these concepts. If we define the context of the interaction to be involved in thesubject-predicate-objecttriplet, most current methods only focus on the single triplet for the first-order interaction to generate sentences. Intuitively, we humans are able to perceive the high-order interaction among concepts from two or more triplets to describe an image. For example, when we see the tripletsman-cutting-sandwichandman-with-knife, it is natural to integrate and predict the sentenceman cutting sandwich with knife. This depends on the high-order interaction betweencuttingandknifein different triplets. Therefore, exploiting high-order interaction is expected to benefit image captioning and focus on reasoning. In this paper, we introduce the novel high-order interaction learning method over detected objects and relationships for image captioning under the umbrella of the encoder-decoder framework. We first extract a set of object and relationship features in an image. During the encoding stage, the interactive refining network is proposed to learn high-order representations by modeling intra- and inter-object feature interaction in the self-attention fashion. During the decoding stage, the interactive fusion network is proposed to integrate object and relationship information by strengthening their high-order interaction based on language context for sentence generation. In this way, we learn the object-relationship dependencies in different stages, which can provide abundant cues for both visual understanding and caption generation. Extensive experiments show that the proposed method can achieve competitive performances against the state-of-the-art methods on MSCOCO dataset. Additional ablation studies further validate its effectiveness. Yanhui Wang 0001, Ning Xu 0003, Anan Liu, Wenhui Li 0001, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Toward Region-Aware Attention Learning for Scene Graph GenerationabstractScene graph generation (SGGen) is a challenging task due to a complex visual context of an image. Intuitively, the human visual system can volitionally focus on attended regions by salient stimuli associated with visual cues. For example, to infer the relationship between man and horse, the interaction between human leg and horseback can provide strong visual evidence to predict the predicate ride. Besides, the attended region face can also help to determine the object man. Till now, most of the existing works studied the SGGen by extracting coarse-grained bounding box features while understanding fine-grained visual regions received limited attention. To mitigate the drawback, this article proposes a region-aware attention learning method. The key idea is to explicitly construct the attention space to explore salient regions with the object and predicate inferences. First, we extract a set of regions in an image with the standard detection pipeline. Each region regresses to an object. Second, we propose the object-wise attention graph neural network (GNN), which incorporates attention modules into the graph structure to discover attended regions for object inference. Third, we build the predicate-wise co-attention GNN to jointly highlight subject's and object's attended regions for predicate inference. Particularly, each subject-object pair is connected with one of the latent predicates to construct one triplet. The proposed intra-triplet and inter-triplet learning mechanism can help discover the pair-wise attended regions to infer predicates. Extensive experiments on two popular benchmarks demonstrate the superiority of the proposed method. Additional ablation studies and visualization further validate its effectiveness. Anan Liu, Hongshuo Tian, Ning Xu 0003, Weizhi Nie, Yongdong Zhang 0001, Mohan Kankanhalli |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | A review of feature fusion-based media popularity prediction methodsabstractWith the popularization of social media, the way of information transmission has changed, and the prediction of information popularity based on social media platforms has attracted extensive attention. Feature fusion-based media popularity prediction methods focus on the multi-modal features of social media, which aim at exploring the key factors affecting media popularity. Meanwhile, the methods make up for the deficiency in feature utilization of traditional methods based on information propagation processes. In this paper, we review feature fusion-based media popularity prediction methods from the perspective of feature extraction and predictive model construction. Before that, we analyze the influencing factors of media popularity to provide intuitive understanding. We further argue about the advantages and disadvantages of existing methods and datasets to highlight the future directions. Finally, we discuss the applications of popularity prediction. To the best of our knowledge, this is the first survey reporting feature fusion-based media popularity prediction methods. Anan Liu, Ning Xu 0003, Junbo Guo, Guoqing Jin, Yejun Tang, Shenyuan Zhang |
Vis. Informatics | 3 |
| 2021 | Triangle-Reward Reinforcement Learning: A Visual-Linguistic Semantic Alignment for Image CaptioningabstractImage captioning aims to generate a sentence consisting of sequential linguistic words, to describe visual units (i.e., objects, relationships, and attributes) in a given image. Most of existing methods rely on the prevalent supervised learning with cross-entropy (XE) function to transfer visual units into a sequence of linguistic words. However, we argue that the XE objective is not sensitive to visual-linguistic alignment, which cannot discriminately penalize the semantic inconsistency and shrink the context gap. To solve these problems, we propose the Triangle-Reward Reinforcement Learning (TRRL) method. TRRL uses the scene graph (G)---objects as nodes and relationships as edges---to represent images, generated sentences, and ground truth sentences individually, and mutually align them during the training process. Specifically, TRRL formulates the image captioning into cooperative agents, where the first agent aims to extract visual scene graph (Gimg) from image (I) and the second agent translates this graph into sentence (S). To discriminately penalize the visual-linguistic inconsistency, TRRL proposes the novel triangle-reward function: 1) the generated sentence and its corresponding ground truth are decomposed into the linguistic scene graph (Gsen) and ground-truth scene graph (Ggt), respectively; 2) Gimg, Gsen, and Ggt are paired to calculate the semantic similarity scores which are proportionally assigned to reward each agent. Meanwhile, to make the training objective sensitive to context changes, we propose the node-level and triplet-level scoring methods to jointly measure the visual-linguistic graph correlations. Extensive experiments on the MSCOCO dataset demonstrate the superiority of TRRL. Additional ablation studies further validate its effectiveness. Weizhi Nie, Jiesi Li, Ning Xu 0003, Anan Liu, Xuanya Li, Yongdong Zhang 0001 |
ACM Multimedia | 3 |
| 2021 | Mask and Predict: Multi-step Reasoning for Scene Graph GenerationabstractScene Graph Generation (SGG) aims to parse the image as a set of semantics, containing objects and their relations. Currently, the SGG methods only stay at presenting the intuitive detection in the image, such as the triplet "logo on board". Intuitively, we humans can further refine these intuitive detections as rational descriptions like "flower painted on surfboard". However, most of existing methods always formulate SGG as a straightforward task, only limited by the manner of one-time prediction, which focuses on a single-pass pipeline and predicts all the semantic. Therefore, to handle this problem, we propose a novel multi-step reasoning manner for SGG. Concretely, we break SGG into two explicit learning stages, including intuitive training stage (ITS) and rational training stage (RTS). In the first stage, we follow the traditional SGG processing to detect objects and relationships, yielding an intuitive scene graph. In the second stage, we perform multi-step reasoning to refine the intuitive scene graph. For each step of reasoning, it consists of two kinds of operations: mask and predict. According to primary predictions and their confidences, we constantly select and mask the low-confidence predictions, which features are optimized and predicted again. After several iterations, all of intuitive semantics will gradually tend to be revised with high confidences, yielding a rational scene graph. Extensive experiments on Visual Genome prove the superiority of the proposed method. Additional ablation studies and visualization cases further validate its effectiveness. Hongshuo Tian, Ning Xu 0003, Anan Liu, Chenggang Yan 0001, Zhendong Mao 0001, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2021 | Multi-type decision fusion network for visual Q&A
Anan Liu, Zimu Lu, Ning Xu 0003, Weizhi Nie, Wenhui Li 0001 |
Image Vis. Comput. | 3 |
| 2021 | Coupled-dynamic learning for vision and language: Exploring Interaction between different tasks
Ning Xu 0003, Hongshuo Tian, Yanhui Wang 0001, Weizhi Nie, Dan Song 0006, Anan Liu, Wu Liu 0005 |
Pattern Recognit. | 1 |
| 2021 | Scene-Graph-Guided message passing network for dense captioning
Anan Liu, Yanhui Wang 0001, Ning Xu 0003, Xuanya Li |
Pattern Recognit. Lett. | 3 |
| 2021 | Scene Graph Inference via Multi-Scale Context ModelingabstractThe scene graph generated for an image structurally represents its object interactions and it substantially aids image scene understanding. To the best of our knowledge, most current works on scene graph generation chiefly focus on pairwise object regions for object and relation inference while ignoring the global visual context outside of these regions. Guided by the intuition that object/relation inference can benefit from the visual context within an image, this paper proposes a multi-scale context modeling method, which can jointly discover and integrate the complementary object-centric and region-centric context for scene graph inference. While both the object-centric and region-centric contexts are separately modeled by their individual modules, a bi-directional message propagation strategy is designed to mutually reinforce the context modeling. A context-fused inference is then proposed to integrate the multi-scale context to guide scene graph inference. Extensive experiments establish that this method can achieve competitive performance compared to the state-of-the-art methods on three benchmarks. Additional ablation studies further validate its effectiveness. Code has been made available at: https://github.com/ningxu1990/MSCM. Ning Xu 0003, Anan Liu, Yongkang Wong, Weizhi Nie, Yuting Su 0001, Mohan Kankanhalli |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Adaptively Clustering-Driven Learning for Visual Relationship DetectionabstractVisual relationship detection aims to describe the interactions between pairs of objects, such asperson-ride-bikeandbike-next to-cartriplets. In reality, it is often the case that there exist some groups of strongly correlated relationships, while others are weakly related. Intuitively, the common relationships can be roughly categorized into several types such as geometric (e.g., next to), action (e.g., ride), and so on. However, previous studies ignore the relatedness discovery among multiple relationships, which only lie on a unified space to leverage visual features or statistical dependencies into categories. To tackle this problem, we propose an adaptively clustering-driven network for visual relationship detection, which can implicitly divide the unified relationship space into several subspaces with specific characteristics. Particularly, we propose two novel modules to discover the common distribution space and latent relationship association, respectively, which map pairs of object features into translation subspaces to induce the discriminative relationship clustering. Then, a fused inference is designed to integrate the group-induced representations with the language prior to facilitate the predicate inference. Especially, we design the Frobenius-norm regularization to boost the clustering. To the best of our knowledge, the proposed method is the first supervised framework to realizesubject-predicate-objectrelationship-aware clustering for visual relationship detection. Extensive experiments show that the proposed method can achieve competing performances against the state-of-the-art methods on the Visual Genome dataset. Additional ablation studies further validate its effectiveness. Anan Liu, Yanhui Wang 0001, Ning Xu 0003, Weizhi Nie, Jie Nie, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Image Captioning with multi-level similarity-guided semantic matchingabstractImage Captioning is a cross-modal task that needs to automatically generate coherent natural sentences to describe the image contents. Due to the large gap between vision and language modalities, most of the existing methods have the problem of inaccurate semantic matching between images and generated captions. To solve the problem, this paper proposes a novel multi-level similarity-guided semantic matching method for image captioning, which can fuse local and global semantic similarities to learn the latent semantic correlation between images and generated captions. Specifically, we extract the semantic units containing fine-grained semantic information of images and generated captions, respectively. Based on the comparison of the semantic units, we design a local semantic similarity evaluation mechanism. Meanwhile, we employ the CIDEr score to characterize the global semantic similarity. The local and global two-level similarities are finally fused using the reinforcement learning theory, to guide the model optimization to obtain better semantic matching. The quantitative and qualitative experiments on large-scale MSCOCO dataset illustrate the superiority of the proposed method, which can achieve fine-grained semantic matching of images and generated captions. Jiesi Li, Ning Xu 0003, Weizhi Nie, Shenyuan Zhang |
Vis. Informatics | 2 |
| 2020 | Part-Aware Interactive Learning for Scene Graph GenerationabstractGenerating scene graph to describe the whereabouts and interactions of objects in an image has attracted increasing attention of researchers. Most existing methods explore object-level visual context or bodypart-object cooperation with the message passing structure, which can not meet the part-aware interaction nature of scene graph. Normally, a subject interacts with an object through crucial parts in each other. Besides, the correlation among parts within an identical object can also help predicting objects and their relationships. Hence, both of subject and object parts and their intra- and inter-object correlations should be fully considered for scene graph generation. In this paper, we propose a part-aware interactive learning method, which are divided into the intra-object and inter-object scenarios. First, we detect objects from an image and further decompose each one into a set of parts. Second, the part-aware graph attention module is proposed to refine part features via the intra-object message passing, and the refined features are incorporated for object inference. Third, the visual mutual attention module is designed to discover part-aware correlated visual cues precisely for predicate inference. It can highlight the subject-related object parts and the object-related subject parts during inter-object interactive learning. We demonstrate the superiority of our method against the state of the arts on Visual Genome. Ablation studies and visualization further validate its effectiveness. Hongshuo Tian, Ning Xu 0003, Anan Liu, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2020 | 3D Model classification based on few-shot learning
Jie Nie, Ning Xu 0003, Ge Yan 0010, Zhiqiang Wei 0002 |
Neurocomputing | 2 |
| 2020 | Corrigendum to "3D model classification based on few-shot learning" [Neurocomputing 398 (2020) 539-546/21103]
Jie Nie, Ning Xu 0003, Ge Yan 0010, Zhiqiang Wei 0002 |
Neurocomputing | 2 |
| 2020 | Hierarchical Deep Neural Network for Image Captioning
Yuting Su 0001, Yuqian Li 0003, Ning Xu 0003, Anan Liu |
Neural Process. Lett. | 3 |
| 2020 | Multi-Level Policy and Reward-Based Deep Reinforcement Learning Framework for Image CaptioningabstractImage captioning is one of the most challenging tasks in AI because it requires an understanding of both complex visuals and natural language. Because image captioning is essentially a sequential prediction task, recent advances in image captioning have used reinforcement learning (RL) to better explore the dynamics of word-by-word generation. However, the existing RL-based image captioning methods rely primarily on a single policy network and reward function-an approach that is not well matched to the multi-level (word and sentence) and multi-modal (vision and language) nature of the task. To solve this problem, we propose a novel multi-level policy and reward RL framework for image captioning that can be easily integrated with RNN-based captioning models, language metrics, or visual-semantic functions for optimization. Specifically, the proposed framework includes two modules: 1) a multi-level policy network that jointly updates the word- and sentence-level policies for word generation; and 2) a multi-level reward function that collaboratively leverages both a vision-language reward and a language-language reward to guide the policy. Furthermore, we propose a guidance term to bridge the policy and the reward for RL optimization. The extensive experiments on the MSCOCO and Flickr30k datasets and the analyses show that the proposed framework achieves competitive performances on a variety of evaluation metrics. In addition, we conduct ablation studies on multiple variants of the proposed framework and explore several representative image captioning models and metrics for the word-level policy network and the language-language reward function to evaluate the generalization ability of the proposed framework. Ning Xu 0003, Hanwang Zhang, Anan Liu, Weizhi Nie, Yuting Su 0001, Jie Nie, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Scene graph captioner: Image captioning based on structural visual representation
Ning Xu 0003, Anan Liu, Jing Liu 0002, Weizhi Nie, Yuting Su 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Multi-guiding long short-term memory for video captioning
Ning Xu 0003, Anan Liu, Weizhi Nie, Yuting Su 0001 |
Multim. Syst. | 1 |
| 2019 | Dual-Stream Recurrent Neural Network for Video CaptioningabstractRecent progress in using recurrent neural networks (RNNs) for video description has attracted an increasing interest, due to its capability to encode a sequence of frames for caption generation. While existing methods have studied various features (e.g., CNN, 3D CNN, and semantic attributes) for visual encoding, the representation and fusion of heterogeneous information from multi-modal spaces have not fully explored. Consider that different modalities are often asynchronous, frame-level multi-modal fusion (e.g., concatenation and linear fusion) will negatively influence each modality. In this paper, we propose a dual-stream RNN (DS-RNN) framework to jointly discover and integrate the hidden states of both visual and semantic streams for video caption generation. First, an encoding RNN is used for each stream to flexibly exploit the hidden states of respective modality. Specifically, we proposed an attentive multi-grained encoder module to enhance the local feature learning with global semantics feature. Then, a dual-stream decoder is deployed to integrate the asynchronous yet complementary sequential hidden states from both streams for caption generation. Extensive experiments on three benchmark datasets, namely, MSVD, MSR-VTT, and MPII-MD, show that DS-RNN achieves competitive performance against the state-of-the-art. Additional ablation studies were conducted on various variants of the proposed DS-RNN. Ning Xu 0003, Anan Liu, Yongkang Wong, Yongdong Zhang 0001, Weizhi Nie, Yuting Su 0001, Mohan Kankanhalli |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Multi-Domain and Multi-Task Learning for Human Action RecognitionabstractDomain-invariant (view-invariant & modalityinvariant) feature representation is essential for human action recognition. Moreover, given a discriminative visual representation, it is critical to discover the latent correlations among multiple actions in order to facilitate action modeling. To address these problems, we propose a multi-domain & multi-task learning (MDMTL) method to (1) extract domain-invariant information for multi-view and multi-modal action representation and (2) explore the relatedness among multiple action categories. Specifically, we present a sparse transfer learning-based method to co-embed multi-domain (multi-view & multi-modality) data into a single common space for discriminative feature learning. Additionally, visual feature learning is incorporated into the multitask learning framework, with the Frobenius-norm regularization term and the sparse constraint term, for joint task modeling and task relatedness-induced feature learning. To the best of our knowledge, MDMTL is the first supervised framework to jointly realize domain-invariant feature learning and task modeling for multi-domain action recognition. Experiments conducted on the INRIA Xmas Motion Acquisition Sequences (IXMAS) dataset, the MSR Daily Activity 3D (DailyActivity3D) dataset, and the Multi-modal & Multi-view & Interactive (M2I) dataset, which is the most recent and largest multi-view and multi-model action recognition dataset, demonstrate the superiority of MDMTL over the state-of-the-art approaches. Anan Liu, Ning Xu 0003, Weizhi Nie, Yuting Su 0001, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Multi-Level Policy and Reward Reinforcement Learning for Image CaptioningabstractImage captioning is one of the most challenging hallmark of AI, due to its complexity in visual and natural language understanding. As it is essentially a sequential prediction task, recent advances in image captioning use Reinforcement Learning (RL) to better explore the dynamics of word-by-word generation. However, existing RL-based image captioning methods mainly rely on a single policy network and reward function that does not well fit the multi-level (word and sentence) and multi-modal (vision and language) nature of the task. To this end, we propose a novel multi-level policy and reward RL framework for image captioning. It contains two modules: 1) Multi-Level Policy Network that can adaptively fuse the word-level policy and the sentence-level policy for the word generation; and 2) Multi-Level Reward Function that collaboratively leverages both vision-language reward and language-language reward to guide the policy. Further, we propose a guidance term to bridge the policy and the reward for RL optimization. Extensive experiments and analysis on MSCOCO and Flickr30k show that the proposed framework can achieve competing performances with respect to different evaluation metrics. Anan Liu, Ning Xu 0003, Hanwang Zhang, Weizhi Nie, Yuting Su 0001, Yongdong Zhang 0001 |
IJCAI | 2 |
| 2018 | Attention-in-Attention Networks for Surveillance Video Understanding in Internet of ThingsabstractIn this paper, we propose an approach to generate the comprehensive video interpretation for the surveillance video understanding in Internet of Things. The key problem of many visual learning tasks is to adaptively select and fuse diverse and complimentary features for video representation. We design the attention-in-attention (AIA) network to hierarchically explore the attention fusion in an end-to-end manner, and demonstrate the value of this model on the multievent recognition and video captioning challenges. Particularly, it consists of multiple encoder attention modules (EAMs) and a fusion attention module (FAM). Each EAM aims to highlight the space-specific features by selecting the most salient visual features or semantic attributes and averages them into one attentive feature. The FAM can suppress or enhance the activation of multispace attentive features and adaptively co-embed them for comprehensive video representation. Then, one long short-term memory unit decodes the video representations to generate multiple event labels or video captions. This architecture is capable of: 1) adaptively learning the salient space-specific feature representation and 2) co-embedding multispace attentive features into one space for feature fusion. Experiments conducted on the surveillance video dataset (concurrent event dataset) and the popular video captioning datasets (Microsoft Research Video Description Corpus and MSR-Video to Text). It shows that the proposed AIA can achieve competitive performances against the state of the arts. Ning Xu 0003, Anan Liu, Weizhi Nie, Yuting Su 0001 |
IEEE Internet Things J. | 1 |
| 2017 | Hierarchical & multimodal video captioning: Discovering and transferring multimodal knowledge for vision to language
Anan Liu, Ning Xu 0003, Yongkang Wong, Junnan Li 0001, Yuting Su 0001, Mohan Kankanhalli |
Comput. Vis. Image Underst. | 2 |
| 2017 | Benchmarking a Multimodal and Multiview and Interactive Dataset for Human Action RecognitionabstractHuman action recognition is an active research area in both computer vision and machine learning communities. In the past decades, the machine learning problem has evolved from conventional single-view learning problem, to cross-view learning, cross-domain learning and multitask learning, where a large number of algorithms have been proposed in the literature. Despite having large number of action recognition datasets, most of them are designed for a subset of the four learning problems, where the comparisons between algorithms can further limited by variances within datasets, experimental configurations, and other factors. To the best of our knowledge, there exists no dataset that allows concurrent analysis on the four learning problems. In this paper, we introduce a novel multimodal and multiview and interactive (M2I) dataset, which is designed for the evaluation of human action recognition methods under all four scenarios. This dataset consists of 1760 action samples from 22 action categories, including nine person-person interactive actions and 13 person-object interactive actions. We systematically benchmark state-of-the-art approaches on M2I dataset on all four learning problems. Overall, we evaluated 13 approaches with nine popular feature and descriptor combinations. Our comprehensive analysis demonstrates that M2I dataset is challenging due to significant intraclass and view variations, and multiple similar action categories, as well as provides solid foundation for the evaluation of existing state-of-the-art algorithms. Anan Liu, Ning Xu 0003, Weizhi Nie, Yuting Su 0001, Yongkang Wong, Mohan Kankanhalli |
IEEE Trans. Cybern. | 2 |
| 2015 | Multi-modal & Multi-view & Interactive Benchmark Dataset for Human Action RecognitionabstractHuman action recognition is one of the most active research areas in both computer vision and machine learning communities. Several methods for human action recognition have been proposed in the literature and promising results have been achieved on the popular datasets. However, the comparison of existing methods is often limited given the different datasets, experimental settings, feature representations, and so on. In particularly, there are no human action dataset that allow concurrent analysis on three popular scenarios, namely single view, cross view, and cross domain. In this paper, we introduce a Multi-modal & Multi-view & Interactive (M2I) dataset, which is designed for the evaluation of the performances of human action recognition under multi-view scenario. This dataset consists of 1760 action samples, including 9 person-person interaction actions and 13 person-object interaction actions. Moreover, we respectively evaluate three representative methods for the single-view, cross-view, and cross domain human action recognition on this dataset with the proposed evaluation protocol. It is experimentally demonstrated that this dataset is extremely challenging due to large intraclass variation, multiple similar actions, significant view difference. This benchmark can provide solid basis for the evaluation of this task and will benefit advancing related computer vision and machine learning research topics. Ning Xu 0003, Anan Liu, Weizhi Nie, Yongkang Wong, Fuwu Li, Yuting Su 0001 |
ACM Multimedia | 1 |
| 2015 | Single/multi-view human action recognition via regularized multi-task learning
Anan Liu, Ning Xu 0003, Yuting Su 0001, Zhaoxuan Yang |
Neurocomputing | 2 |