EDBT 2026 Demo / reviewers in the wild / expert
Pijian Li
dblp:294/0958
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2026
0009-0005-1924-5248ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Knowledge-enhanced Chinese multimodal hate speech detection
Qingbao Huang, Pijian Li, Xingmao Zhang, Shizhen Chen, Haonan Cheng, Zhiyue Liu |
Expert Syst. Appl. | 2 |
| 2026 | Anchor-Based Multimodal Verification: A Dynamic Query Framework for Fake News Forensics in Short VideosabstractThe proliferation of maliciously altered short videos on social media platforms poses a significant threat to information security ecosystems, eroding public trust in digital media. Despite recent advancements in detecting fake video news, significant challenges remain in the forensic analysis of short videos, leading to issues of bias. First, as technology rapidly advances, fake videos are becoming increasingly semantically convincing, undermining the effectiveness of current classification methods. Second, the heterogeneous nature of video modalities (visual, textual, audio) creates critical challenges for models to learn discriminative feature representations. To address these challenges, we propose a dynamic query framework for fake news forensics in short videos, termed the Semantic Guided Adaptive Network (SGAN). Our approach is motivated by the need to utilize superficial alignment to identify suspicious manipulations through anchor-based verification and to leverage the adaptive capability of learnable queries to learn the heterogeneous boundary in each modality. Specifically, SGAN comprises a verification module and a flexible query learning module. The verification module employs text as the anchor to verify detailed context, mining fine-grained information while emphasizing key features, thereby providing candidate manipulations for downstream modules. The query learning module leverages learnable queries to map heterogeneous forensic features and integrates them through multi-level fusion for decision-making. Extensive experiments conducted on two widely used datasets demonstrate the effectiveness and generalization of the proposed method. Pijian Li, Qingbao Huang, Feng Shuang 0002, Yi Cai 0001, Haonan Cheng, Qing Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2026 | DESSM: Dual Encoder-Based State Space Model for Image InpaintingabstractImage inpainting represents a fundamental and challenging problem in computer vision, requiring the synthesis of visually plausible content for missing regions while preserving both textural details and structural coherence. While current approaches employ auxiliary networks and attention mechanisms to capture structural priors and expand receptive fields, they remain constrained by two fundamental limitations: (1) insufficient interaction between textural and structural priors, and (2) the quadratic computational complexity inherent in attention operations. To overcome these challenges, we present DESSM, an innovative Dual Encoder-based State space Model for image inpainting that achieves efficient global context modeling with linear computational complexity. Our DESSM integrates three synergistically designed modules: (1) a dual-branch encoder for complementary learning of textural patterns and structural priors, (2) a Feature Cross Fusion Block (FCFB) enabling dynamic feature interaction while adaptively suppressing redundant information, and (3) a Spatial-Channel joint Selective scan Block (SCSB) for efficient long-range dependency modeling. Comprehensive evaluations across four standard benchmarks (i.e., CelebA, CelebA-HQ, Places2, and Paris StreetView) demonstrate that our DESSM achieves state-of-the-art performance in both visual fidelity and computational efficiency. Rongrong Zhou, Peizhou Cai, Pijian Li, Qingbao Huang |
IEEE Trans. Multim. | 4 |
| 2026 | Metaphorical Visual Question Answering: Benchmark and Knowledge-Enhanced Metaphor Understanding MethodabstractFact and common-sense reasoning grounded in metaphorical imagery constitute a more challenging form of visual question answering (VQA). Under this form, models typically cannot obtain answers directly from images. Models first need to identify and comprehend the metaphorical components within the image, subsequently integrating prior knowledge to establish the mappings between the target and source domains depicted in the image. To evaluate this capability, we propose a VQA benchmark based on metaphorical images (METAVQA), which measures the understanding of the model of metaphorical images in the form of VQA. Experimental results indicate that interpreting metaphorical images requires robust prior knowledge and a strong ability to understand abstract components, which remains a challenge for LLMs. To improve the metaphorical understanding capability of LLMs, we propose a knowledge-enhanced method (KEMU). This method utilizes LLMs to extract image triplets, retrieve implicit metaphorical knowledge, and reason through questions step-by-step using chain of thought prompting, and pretrained classifier to match the refined reasoning output to the provided answer choices. Validated on the METAVQA dataset, KEMU outperforms LLMs in both one-hop and multi-hop questions. Our code and benchmark can be seen inhttps://github.com/VILAN-Lab/METAVQA. Qingbao Huang, Peihang He, Pijian Li, Yang Tian 0008, Yi Cai 0001, Qing Li 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Multi-modal Metaphor Explanation: A Dataset and Benchmark
Wenye Zhao, Pijian Li, Qingbao Huang |
PRCV (12) | 4 |
| 2024 | Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by PointabstractAs a fundamental and challenging task in the vision and language domain, Referring Expression Comprehension (REC) has shown impressive improvements recently. However, for a complex task that couples the comprehension of abstract concepts and the localization of concrete instances, one-stage approaches are bottlenecked by computing and data resources. To obtain a low-cost solution, the prevailing two-stage approaches decouple REC into localization (region proposal) and comprehension (region-expression matching) at region-level, but the solution based on isolated regions cannot sufficiently utilize the context and is usually limited by the quality of proposals. Therefore, it is necessary to rebuild an efficient two-stage solution system. In this paper, we propose a point-based two-stage framework for REC, in which the two stages are redefined as point-based cross-modal comprehension and point-based instance localization. Specifically, we reconstruct the raw bounding box and segmentation mask into center and mass scores as soft ground-truth for measuring point-level cross-modal correlations. With the soft ground-truth, REC can be approximated as a binary classification problem, which fundamentally avoids the impact of isolated regions on the optimization process. Remarkably, the consistent metrics between center and mass scores allow our system to directly optimize grounding and segmentation by utilizing the same architecture. Experiments on multiple benchmarks show the feasibility and potential of our point-based paradigm. Our code available at https://github.com/VILAN-Lab/PBREC-MT. Peizhi Zhao, Shiyi Zheng, Wenye Zhao, Dongsheng Xu 0001, Pijian Li, Yi Cai 0001, Qingbao Huang |
AAAI | 5 |
| 2024 | Multi-Granularity Feature Fusion for Image-Guided Story Ending GenerationabstractImage-guided Story Ending Generation aims at generating a reasonable and logical ending given a story context and an ending-related image. The existing models have achieved some success by fusing global image features with story context through an attention mechanism. However, they ignore the logical relationship between the story context and the image regions, and have not considered the high-level semantic features of the image such as visual sentiment. This may cause the generated ending inconsistent with the logic or sentiment of the given information. In this paper, we propose aMulti-Granularity featureFusion (MGF) model to solve this problem. Concretely, we first employ an image sentiment extractor to grasp the sentiment features of the image as part of the global image features. We then design a scene subgraph selector to capture the image features of the key region by picking the scene subgraph most relevant to the context. Finally, we fuse the textual and visual features from object level, region level, and global level, respectively. Our model is thereby capable of effectively capturing the key region features and visual sentiment of the image, so as to generate a more logical and sentimental ending. Experimental results show that our MGF model outperforms the state-of-the-art models on most metrics. Pijian Li, Qingbao Huang, Yi Cai 0001, Feng Shuang 0002, Qing Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | Region-Focused Network for Dense CaptioningabstractDense captioning is a very critical but under-explored task, which aims to densely detect localized regions-of-interest (RoIs) and describe them with natural language in a given image. Although recent studies tried to fuse multi-scale features from different visual instances to generate more accurate descriptions, their methods still suffer from the lack of exploration of relation semantic information in images, leading to less informative descriptions. Furthermore, indiscriminately fusing all visual instance features will introduce redundant information, resulting in poor matching between descriptions and corresponding regions. In this work, we propose a Region-Focused Network (RFN) to address these issues. Specifically, to fully comprehend the images, we first extract the object-level features, and encode the interaction and position relations between objects to enhance the object representations. Then, to decrease the interference from redundant information about the target region, we extract the most relevant information to the region. Finally, a region-based Transformer is employed to compose and align the previous mined information and generate the corresponding descriptions. Extensive experiments on Visual Genome V1.0 and V1.2 datasets show that our RFN model outperforms the state-of-the-art methods, thus verifying its effectiveness. Our code is available at https://github.com/VILAN-Lab/DesCap . Qingbao Huang, Pijian Li, Youji Huang, Feng Shuang 0002, Yi Cai 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Linking People across Text and Images Based on Social Relation Reasoning
Peizhi Zhao, Pijian Li, Yi Cai 0001, Qingbao Huang |
AAAI | 3 |
| 2023 | Scene-text Oriented Visual Entailment: Task, Dataset and SolutionabstractVisual Entailment (VE) is a fine-grained reasoning task aiming to predict whether the image semantically entails a hypothesis in textual form.Existing studies of VE only focus on basic visual attributes but largely overlook the importance of scene text, which usually entails rich semantic information and crucial clues (e.g., time, place, affiliation, and topic), leading to superficial design of hypothesis or incorrect entailment prediction. To fill this gap, we propose a new task called scene-text oriented Visual Entailment (STOVE), which requires models to predict whether an image semantically entails the corresponding hypothesis designed based on the scene text-centered visual information.STOVE task challenges a model to deeply understand the interplay between language and images containing scene text, requiring aligning hypotheses tokens, scene text, and visual contents.To support the researches on STOVE, we further collect a dataset termed TextVE, consisting of 23,864 images and 47,728 hypotheses related to scene text, which is constructed with the strategy of minimizing biases.Additionally, we present a baseline named MMTVE applying a multimodal transformer to model the spatial, semantic, and visual reasoning relations between multiple scene text tokens, hypotheses, and visual features.Experimental results illustrate that our model is effective in comprehending STOVE and achieves outstanding performance.Our codes are available at https://github.com/VISLANG-Lab/TextVE. Nan Li 0055, Pijian Li, Dongsheng Xu 0001, Wenye Zhao, Yi Cai 0001, Qingbao Huang |
ACM Multimedia | 2 |
| 2022 | Suppressing Biased Samples for Robust VQAabstractMost existing visual question answering (VQA) models strongly rely on language bias to answer questions, i.e., they always tend to fit question-answer pairs on the train split and perform poorly on the test spilt when the answer distributions are different. This behavior makes them hard to be applied in real scenarios. To reduce the language biases, previous studies mainly integrate modules to overcome language priors (ensemble-based methods) or generate additional training data to balance dataset biases (data-balanced methods). However, all the existing ensemble-based methods drop their accuracies on the VQA v2 dataset, while data-balanced methods may introduce new biases and cannot guarantee the quality of the generated data. In this paper, we propose a model-agnostic training scheme called Suppressing Biased Samples (SBS) to overcome language priors. SBS consists of two collaborative parts, i.e., a Data Classifier Module to divide the dataset into biased samples and unbiased samples by utilizing the similarity in the semantic space, and a Bias Penalty Module to suppress the biased samples to weaken their influence. As a new way of balancing data to address language bias, SBS overcomes the shortcomings of previous data-balanced methods. Experimental results show that our method can be merged into other bias-reduction methods and achieves a new state-of-the-art performance on the commonly used VQA-CP v2 dataset. Ninglin Ouyang, Qingbao Huang, Pijian Li, Yi Cai 0001, Bin Liu 0053, Ho-fung Leung, Qing Li 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Entity Guided Question Generation with Contextual Structure and Sequence Information CapturingabstractQuestion generation is a challenging task and has attracted widespread attention in recent years. Although previous studies have made great progress, there are still two main shortcomings: First, previous work did not simultaneously capture the sequence information and structure information hidden in the context, which results in poor results of the generated questions. Second, the generated questions cannot be answered by the given context. To tackle these issues, we propose an entity guided question generation model with contextual structure information and sequence information capturing. We use a Graph Convolutional Network and a Bidirectional Long Short Term Memory Network to capture the structure information and sequence information of the context, simultaneously. In addition, to improve the answerability of the generated questions, we use an entity-guided approach to obtain question type from the answer, and jointly encode the answer and question type. Both automatic and manual metrics show that our model can generate comparable questions with state-of-the-art models. Our code is available at https://github.com/VISLANG-Lab/EGSS. Qingbao Huang, Mingyi Fu, Linzhang Mo, Yi Cai 0001, Pijian Li, Qing Li 0001, Ho-fung Leung |
AAAI | 6 |
| 2021 | Story Ending Generation with Multi-Level Graph Convolutional Networks over Dependency TreesabstractAs an interesting and challenging task, story ending generation aims at generating a reasonable and coherent ending for a given story context. The key challenge of the task is to comprehend the context sufficiently and capture the hidden logic information effectively, which has not been well explored by most existing generative models. To tackle this issue, we propose a context-aware Multi-level Graph Convolutional Networks over Dependency Parse (MGCN-DP) trees to capture dependency relations and context clues more effectively. We utilize dependency parse trees to facilitate capturing relations and events in the context implicitly, and Multi-level Graph Convolutional Networks to update and deliver the representation crossing levels to obtain richer contextual information. Both automatic and manual evaluations show that our MGCN-DP can achieve comparable performance with state-of-the-art models. Our source code is available at https://github.com/VISLANG-Lab/MLGCN-DP. Qingbao Huang, Linzhang Mo, Pijian Li, Yi Cai 0001, Qingguang Liu, Jielong Wei, Qing Li 0001, Ho-fung Leung |
AAAI | 3 |