VLDB 2026 Research / reviewers in the wild / expert
Binbin Li 0003
dblp:06/8137-3
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0002-4886-1267ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seedcap: Semantic Expansion and Entity-Driven Zero-Shot Image CaptioningabstractZero-shot image captioning (ZIC) generates natural language descriptions for images using only textual data during training. Recent progress leverages text-to-image models to construct pseudo image-text pairs, alleviating the modality mismatch between text-only training and image-based inference. However, the distribution gap between synthetic and real images can compromise performance when models trained on generated images are deployed in real-world scenarios. To address this challenge, we propose SEEDCap, a zero-shot captioning framework featuring SEED (Semantic Expansion and Entity-Driven), a parameter-free module that performs dual-level semantic expansion. SEED enhances the correspondence between image features and retrieved textual cues at both entity and sentence levels, producing enriched semantic embeddings that effectively reduce the distributional gap between synthetic and real inputs. These enhanced features are then integrated with the global visual representation through a modality fusion module, which is subsequently decoded by a large language model to generate accurate and contextually rich captions. Extensive experiments demonstrate that SEEDCap achieves state-of-the-art zero-shot results on both in-domain and cross-domain benchmarks, underscoring its robustness and practical utility. Binbin Li 0003, Shupei Xiao, Dayan Wu, Gengqi Yang, Siyu Jia, Zisen Qi |
ICMR | 1 |
| 2025 | Gather and Trace: Rethinking Video TextVQA from an Instance-oriented PerspectiveabstractVideo text-based visual question answering (Video TextVQA) aims to answer questions by explicitly reading and reasoning about the text involved in a video. Most works in this field follow a frame-level framework which suffers from redundant text entities and implicit relation modeling, resulting in limitations in both accuracy and efficiency. In this paper, we rethink the Video TextVQA task from an instance-oriented perspective and propose a novel model termed GAT (Gather and Trace). First, to obtain accurate reading result for each video text instance, a context-aggregated instance gathering module is designed to integrate the visual appearance, layout characteristics, and textual contents of the related entities into a unified textual representation. Then, to capture dynamic evolution of text in the video flow, an instance-focused trajectory tracing module is utilized to establish spatio-temporal relationships between instances and infer the final answer. Extensive experiments on several public Video TextVQA datasets validate the effectiveness and generalization of our framework. GAT outperforms existing Video TextVQA methods, video-language pretraining methods, and video large language models in both accuracy and inference speed. Notably, GAT surpasses the previous state-of-the-art Video TextVQA methods by 3.86% in accuracy and achieves ten times of faster inference speed than video large language models. The source code is available at https://github.com/zhangyan-ucas/GAT. Gangyan Zeng, Daiqing Wu, Huawen Shen, Binbin Li 0003, Yu Zhou 0015, Can Ma, Xiaojun Bi 0002 |
ACM Multimedia | 5 |
| 2025 | Uni-DocDiff: A Unified Document Restoration Model Based on DiffusionabstractRemoving various degradations from damaged documents greatly benefits digitization, downstream document analysis, and readability. Previous methods often treat each restoration task independently with dedicated models, leading to a cumbersome and highly complex document processing system. Although recent studies attempt to unify multiple tasks, they often suffer from limited scalability due to handcrafted prompts and heavy preprocessing, and fail to fully exploit inter-task synergy within a shared architecture. To address the aforementioned challenges, we propose Uni-DocDiff, a Unified and highly scalable Doc ument restoration model based on Dif fusion. Uni-DocDiff develops a learnable task prompt design, ensuring exceptional scalability across diverse tasks. To further enhance its multi-task capabilities and address potential task interference, we devise a novel Prior Pool, a simple yet comprehensive mechanism that combines both local high-frequency features and global low-frequency features. Additionally, we design the Prior Fusion Module (PFM), which enables the model to adaptively select the most relevant prior information for each specific task. Extensive experiments show that the versatile Uni-DocDiff achieves performance comparable or even superior performance compared with task-specific expert models, and simultaneously holds the task scalability for seamless adaptation to new tasks. Fangmin Zhao, Weichao Zeng, Zhenhang Li, Dongbao Yang, Binbin Li 0003, Xiaojun Bi 0002, Yu Zhou 0015 |
ACM Multimedia | 5 |
| 2025 | DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual PromptsabstractRecent lightweight retrieval-augmented image caption models often utilize retrieved data solely as text prompts, thereby creating a semantic gap by leaving the original visual features unenhanced, particularly for object details or complex scenes. To address this limitation, we propose DualCap, a novel approach that enriches the visual representation by generating a visual prompt from retrieved similar images. Our model employs a dual retrieval mechanism, using standard image-to-text retrieval for text prompts and a novel image-to-image retrieval to source visually analogous scenes. Specifically, salient keywords and phrases are derived from the captions of visually similar scenes to capture key objects and similar details. These textual features are then encoded and integrated with the original image features through a lightweight, trainable feature fusion network. Extensive experiments demonstrate that our method achieves competitive performance while requiring fewer trainable parameters compared to previous visual-prompting captioning approaches. The source code is available at https://github.com/mungeryang/DualCap. Binbin Li 0003, Guimiao Yang, Zisen Qi, Haiping Wang 0003 |
MMAsia | 1 |
| 2024 | Triple GNNs: Introducing Syntactic and Semantic Information for Conversational Aspect-Based Quadruple Sentiment AnalysisabstractConversational Aspect-Based Sentiment Analysis (DiaASQ) aims to detect quadruples {target, aspect, opinion, sentiment polarity} from given dialogues. In DiaASQ, elements constituting these quadruples are not necessarily confined to individual sentences but may span across multiple utterances within a dialogue. This necessitates a dual focus on both the syntactic information of individual utterances and the semantic interaction among them. However, previous studies have primarily focused on coarse-grained relationships between utterances, thus overlooking the potential benefits of detailed intra-utterance syntactic information and the granularity of inter-utterance relationships. This paper introduces the Triple GNNs network to enhance DiaAsQ. It employs a Graph Convolutional Network (GCN) for modeling syntactic dependencies within utterances and a Dual Graph Attention Network (DualGATs) to construct interactions between utterances. Experiments on two standard datasets reveal that our model significantly outperforms stateof-the-art baselines. The code is available at https://github.com/ nlperi2b/Triple-GNNs-. Binbin Li 0003, Siyu Jia, Bingnan Ma, Zisen Qi, Xingbang Tan, Menghan Guo, Shenghui Liu |
CSCWD | 1 |
| 2024 | Dynamic Multi-Scale Context Aggregation for Conversational Aspect-Based Sentiment Quadruple AnalysisabstractConversational aspect-based sentiment quadruple analysis, namely DiaASQ, aims to extract the quadruple of target-aspect-opinion-sentiment within a dialogue . In DiaASQ, a quadruple’s elements often cross multiple utterances. This situation complicates the extraction process, emphasizing the need for an adequate understanding of conversational context and interactions. However, existing work independently encodes each utterance, thereby struggling to capture long-range conversational context and overlooking the deep inter-utterance dependencies. In this work, we propose a novel Dynamic Multi-scale Context Aggregation network (DMCA) to address the challenges. Specifically, we first utilize dialogue structure to generate multi-scale utterance windows for capturing rich contextual information. After that, we design a Dynamic Hierarchical Aggregation module(DHA) to integrate progressive cues between them. In addition, we form a multi-stage loss strategy to improve model performance and generalization ability. Extensive experimental results show that the DMCA model outperforms baselines significantly and achieves state-of-the-art performance1. Wenyuan Zhang 0002, Binbin Li 0003, Siyu Jia, Zisen Qi, Xingbang Tan |
ICASSP | 3 |
| 2023 | UCWSC: A unified cross-modal weakly supervised classification ensemble frameworkabstractIn recent years, Internet data has grown exponentially, but due to the lack of labels, the data that can be used is still relatively small. To solve this problem, research on weak supervision has emerged. However, common weakly supervised research often focuses on either single-modal data or multi-modal data research, which cannot be compatible with both types of data at the same time. Motivated by this observation, we propose a unified cross-modal weakly supervised classification ensemble framework (UCWSC) to tackle this issue. Especially, our proposed framework is based on high-order feature information of different modes. First, We introduce a feature fusion method based on high-order features to increase the amount of acquired information. Then we propose a modified Feature MixMatch algorithm with learning from feature representations. We propose feature fusion and decision fusion methods for weakly supervised classification of multi-modal data with voting and weighting mechanisms as discriminators to obtain the final classification results, respectively. We demonstrate the compatibility of these techniques, our classification accuracy can reach around 99% on the Wikipedia dataset and 78% on the MVSA-Multiple dataset. Huiyang Chang, Binbin Li 0003, Guangjun Wu, Haiping Wang 0003, Siyu Jia, Zisen Qi, Xiaohua Jiang |
CSCWD | 2 |
| 2022 | Edge Federated Learning for Social Profit Optimality: A Cooperative Game Approach
Wenyuan Zhang 0002, Guangjun Wu, Yongfei Liu, Binbin Li 0003 |
CollaborateCom (1) | 4 |