VLDB 2026 Research / reviewers in the wild / expert
Bo Wang 0134
dblp:72/6811-134
· DBLP profile ↗
13ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0001-8102-5346ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EXCEEDS: Extracting Complex Events via Nugget-based Grid Modeling in Scientific DomainabstractIt is crucial to understand a specific domain by events.Extensive event extraction research has been conducted in many domains such as news, finance, and biology.However, event extraction in scientific domain is still insufficiently supported by comprehensive datasets and tailored methods.Compared with other domains, scientific domain has two characteristics: (1) denser nuggets and events, and (2) more complex information forms.To solve the above problem, considering these two characteristics, we first construct SciEvents, a large-scale multi-event document-level dataset with a schema tailored for scientific domain.It consists of 2,508 documents and 24,381 events under multi-stage manual annotation and quality control.Then, we propose EXCEEDS, an end-to-end scientific event extraction framework by encoding dense nuggets into a grid matrix and simplifying complex event extraction as a nugget-based grid modeling task.Experiments on SciEvents demonstrate state-of-the-art performances of EXCEEDS.Both the SciEvents dataset and the EXCEEDS framework are released publicly to facilitate future research. Yi-Fan Lu, Xianling Mao, Bo Wang 0134, Xiao Liu 0029, Heyan Huang |
ACL (1) | 3 |
| 2025 | Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document RetrievalabstractDocument retrieval in real-world scenarios faces significant challenges due to diverse document formats and modalities. Traditional text-based approaches rely on tailored parsing techniques that disregard layout information and are prone to errors, while recent parsing-free visual methods often struggle to capture fine-grained textual semantics in text-rich scenarios. To address these limitations, we propose \textbf{Unveil}, a novel visual-textual embedding framework that effectively integrates textual and visual features for robust document representation. Through knowledge distillation, we transfer the semantic understanding capabilities from the visual-textual embedding model to a purely visual model, enabling efficient parsing-free retrieval while preserving semantic fidelity. Experimental results demonstrate that our visual-textual embedding method surpasses existing approaches, while knowledge distillation successfully bridges the performance gap between visual-textual and visual-only methods, improving both retrieval accuracy and efficiency. Hao Sun 0015, Yingyan Hou, Jiayan Guo, Bo Wang 0134, Chunyu Yang 0005, Jinsong Ni, Yan Zhang 0117 |
ACL (1) | 4 |
| 2025 | Mitigating the Discrepancy Between Video and Text Temporal Sequences: A Time-Perception Enhanced Video Grounding method for LLMabstractExisting video LLMs typically excel at capturing the overall description of a video but lack the ability to demonstrate an understanding of temporal dynamics and a fine-grained grasp of localized content within the video. In this paper, we propose a Time-Perception Enhanced Video Grounding via Boundary Perception and Temporal Reasoning aimed at mitigating LLMs’ difficulties in understanding the discrepancies between video and text temporality. Specifically, to address the inherent biases in current datasets, we design a series of boundary-perception tasks to enable LLMs to capture accurate video temporality. To tackle LLMs’ insufficient understanding of temporal information, we develop specialized tasks for boundary perception and temporal relationship reasoning to deepen LLMs’ perception of video temporality. Our experimental results show significant improvements across three datasets: ActivityNet, Charades, and DiDeMo (achieving up to 11.2% improvement on [email protected]), demonstrating the effectiveness of our proposed temporal awareness-enhanced data construction method. Xuefen Li, Bo Wang 0134, Ge Shi 0002, Chong Feng 0001, Jiahao Teng |
COLING | 2 |
| 2025 | DecoupleSearch: Decouple Planning and Search via Hierarchical Reward ModelingabstractRetrieval-Augmented Generation (RAG) systems have emerged as a pivotal methodology for enhancing Large Language Models (LLMs) through the dynamic integration of external knowledge.To further improve RAG's flexibility, Agentic RAG introduces autonomous agents into the workflow.However, Agentic RAG faces several challenges: (1) the success of each step depends on both high-quality planning and accurate search, (2) the lack of supervision for intermediate reasoning steps, and (3) the exponentially large candidate space for planning and searching.To address these challenges, we propose DecoupleSearch, a novel framework that decouples planning and search processes using dual value models, enabling independent optimization of plan reasoning and search grounding.Our approach constructs a reasoning tree, where each node represents planning and search steps.We leverage Monte Carlo Tree Search to assess the quality of each step.During inference, Hierarchical Beam Search iteratively refines planning and search candidates with dual value models.Extensive experiments across policy models of varying parameter sizes, demonstrate the effectiveness of our method. Hao Sun 0015, Zile Qiao, Bo Wang 0134, Guoxin Chen, Yingyan Hou, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Yan Zhang 0117 |
EMNLP | 3 |
| 2025 | ClearView: A Quality-aware Cross-modal Alignment Framework for CT Report GenerationabstractWhile automated CT report generation (CTRG) systems promise to enhance clinical workflow efficiency, current solutions, even those based on advanced multi-modal large language models (MLLMs), face fundamental challenges in ensuring report quality and reliability. Through systematic analysis of representation dynamics in MLLM-based CTRG models, we discover that existing systems lack robust quality discrimination capabilities, manifesting in two critical limitations: representation entanglement between reports of varying quality levels in the feature space, and quality-insensitive generation due to conventional training paradigms focusing solely on ground-truth reports. To address these limitations, we propose CT-ClearView, a quality-aware cross-modal alignment framework consisting of two key innovations: (1) a systematic methodology for constructing clinically-relevant hard negative examples using GPT-4, which introduces subtle but significant clinical errors while maintaining report structure and plausibility, and (2) a contrastive learning framework that leverages these examples to effectively disentangle representations of varying quality reports and enhance the model's sensitivity to clinical details. Extensive experiments on CTRG-Chest-548K and CTRG-Brain-263K datasets demonstrate significant performance improvements in natural language generation (NLG) metrics compared to existing approaches (e.g., increases of 11% in BLEU-1 and 16% in both BLEU-4 and ROUGE-L, on the CTRG-Chest-548K datasets). Qingyong Su, Chong Feng 0001, Bo Wang 0134, Ge Shi 0002, Yan Zhuang 0012 |
ICMR | 3 |
| 2025 | ALOHA: Adapting Local Spatio-Temporal Context to Enhance the Audio-Visual Semantic SegmentationabstractAudio-Visual Semantic Segmentation (AVSS) plays a crucial role in pixel-level multi-modal perception for real-world applications such as robotic navigation and autonomous driving. Existing methods typically rely on global spatio-temporal modules to fuse audio and visual representations, which aids in generating pixel-level semantic masks. However, these approaches often overlook the importance of local spatio-temporal context in understanding semantics, leading to suboptimal performance. This limitation makes it difficult for models to accurately distinguish sound-emitting objects from irrelevant background noise, resulting in erroneous segmentation across the spatio-temporal dimension. To address this issue, we propose the ALOHA framework, which A dapts LO cal spatio-temporal context to en HA nce AVSS. The framework introduces two key components designed to leverage and enhance local spatio-temporal context information: the LOHA adapter and the Selective Context Enhancement (SCE) module. Specifically, the LOHA adapter adaptively captures essential modality information across spatio-temporal dimensions, while implicitly learning fine-grained local context through the local attention mechanism. Furthermore, the SCE module selectively enhances the local context related to the semantics, thereby facilitating the distinction between the sounding object and irrelevant background and improving segmentation accuracy. Moreover, to better adapt to embodied AI systems, our framework utilizes a parameter-shared encoder and applies the adapters in a staged manner. This design significantly reduces the number of trainable parameters, making it more parameter-efficient. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance on the AVSBench-Semantic benchmark dataset and shows competitive results on the AVSBench-Object benchmark, while exhibiting broad adaptability across different visual backbone networks. Yang-Hao Zhou, Heyan Huang, Cunhan Guo, Rongcheng Tu, Zeyu Xiao 0002, Bo Wang 0134, Xianling Mao |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative LearningabstractRecent advancements in large language models (LLMs) have been remarkable.Users face a choice between using cloud-based LLMs for generation quality and deploying local-based LLMs for lower computational cost.The former option is typically costly and inefficient, while the latter usually fails to deliver satisfactory performance for reasoning steps requiring deliberate thought processes.In this work, we propose a novel LLM utilization paradigm that facilitates the collaborative operation of large cloud-based LLMs and smaller local-deployed LLMs.Our framework comprises two primary modules: the local agent instantiated with a relatively smaller LLM, handling less complex reasoning steps, and the cloud agent equipped with a larger LLM, managing more intricate reasoning steps.This collaborative processing is enabled through an adaptive mechanism where the local agent introspectively identifies errors and proactively seeks assistance from the cloud agent, thereby effectively integrating the strengths of both locally-deployed and cloudbased LLMs, resulting in significant enhancements in task completion performance and efficiency.We evaluate ADASWITCH across 7 benchmarks, ranging from mathematical reasoning and complex question answering, using various types of LLMs to instantiate the local and cloud agents.The empirical results show that ADASWITCH effectively improves the performance of the local agent, and sometimes achieves competitive results compared to the cloud agent while utilizing much less computational overhead.Question: Joy is 2 years older than twice the age of Tom.If Tom is 10 years old, how old is Joy?Thought: Calculate twice the age of Tom.Action: R1 = Calculator(2 * 2) Observation: 4 Reflection: Is previous step wrong: Yes Thought: Calculate twice the age of Tom who is 10.Action: R1 = Calculator(10 * 2) Observation: 20 Reflection: Is previous step wrong: No Thought: Calculate the age of Joy. Hao Sun 0015, Jiayi Wu 0001, Hengyi Cai, Xiaochi Wei, Yue Feng 0002, Bo Wang 0134, Shuaiqiang Wang, Yan Zhang 0117, Dawei Yin 0001 |
EMNLP | 6 |
| 2024 | Towards Verifiable Text Generation with Evolving Memory and Self-ReflectionabstractDespite the remarkable ability of large language models (LLMs) in language comprehension and generation, they often suffer from producing factually incorrect information, also known as hallucination.A promising solution to this issue is verifiable text generation, which prompts LLMs to generate content with citations for accuracy verification.However, verifiable text generation is non-trivial due to the focus-shifting phenomenon, the intricate reasoning needed to align the claim with correct citations, and the dilemma between the precision and breadth of retrieved documents.In this paper, we present VTG, an innovative framework for Verifiable Text Generation with evolving memory and self-reflection.VTG introduces evolving long short-term memory to retain both valuable documents and recent documents.A two-tier verifier equipped with an evidence finder is proposed to rethink and reflect on the relationship between the claim and citations.Furthermore, active retrieval and diverse query generation are utilized to enhance both the precision and breadth of the retrieved documents.We conduct extensive experiments on five datasets across three knowledge-intensive tasks and the results reveal that VTG significantly outperforms baselines. Hao Sun 0015, Hengyi Cai, Bo Wang 0134, Yingyan Hou, Xiaochi Wei, Shuaiqiang Wang, Yan Zhang 0117, Dawei Yin 0001 |
EMNLP | 3 |
| 2024 | Retrieved In-Context Principles from Previous MistakesabstractIn-context learning (ICL) has been instrumental in adapting large language models (LLMs) to downstream tasks using correct input-output examples.Recent advances have attempted to improve model performance through principles derived from mistakes, yet these approaches suffer from lack of customization and inadequate error coverage.To address these limitations, we propose Retrieved In-Context Principles (RICP), a novel teacherstudent framework.In RICP, the teacher model analyzes mistakes from the student model to generate reasons and insights for preventing similar mistakes.These mistakes are clustered based on their underlying reasons for developing task-level principles, enhancing the error coverage of principles.During inference, the most relevant mistakes for each question are retrieved to create question-level principles, improving the customization of the provided guidance.RICP is orthogonal to existing prompting methods and does not require intervention from the teacher model during inference.Experimental results across seven reasoning benchmarks reveal that RICP effectively enhances performance when applied to various prompting strategies. Hao Sun 0015, Yong Jiang 0005, Bo Wang 0134, Yingyan Hou, Yan Zhang 0117, Pengjun Xie, Fei Huang 0002 |
EMNLP | 3 |
| 2024 | One for All: A Unified Generative Framework for Image Emotion ClassificationabstractImage Emotion Classification (IEC) is an essential research area, offering valuable insights into user emotional states for a wide range of applications, including opinion mining, recommendation systems, and mental health treatment. The challenges associated with IEC are mainly attributed to the complexity and ambiguity of human emotions, the lack of a universally accepted emotion model, and excessive dependence on prior knowledge. To address these challenges, we propose a novel Unified Generative framework for Image Emotion Classification (UGRIE), which is capable of simultaneously modeling various emotion models and capturing intricate semantic relationships between emotion labels. Our approach employs a flexible natural language template, converting the IEC task into a template-filling process that can be easily adapted to accommodate a diverse range of IEC tasks. To further enhance the performance, we devise a mapping mechanism to seamlessly integrate the multimodal pre-training model CLIP with the text generation pre-training model BART, thus leveraging the strengths of both models. A comprehensive set of experiments conducted on multiple public datasets demonstrates that our proposed method consistently outperforms existing approaches to a large margin in supervised settings, exhibits remarkable performance in low-resource scenarios, and unifies distinct emotion models within a single, versatile framework. Ge Shi 0002, Sinuo Deng, Bo Wang 0134, Chong Feng 0001, Yan Zhuang 0012 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Dynamic Prefix-Tuning for Generative Template-based Event ExtractionabstractWe consider event extraction in a generative manner with template-based conditional generation.Although there is a rising trend of casting the task of event extraction as a sequence generation problem with prompts, these generation-based methods have two significant challenges, including using suboptimal prompts and static event type information.In this paper, we propose a generative templatebased event extraction method with dynamic prefix (GTEE-DYNPREF) by integrating context information with type-specific prefixes to learn a context-specific prefix for each context.Experimental results show that our model achieves competitive results with the state-ofthe-art classification-based model ONEIE on ACE 2005 and achieves the best performances on ERE.Additionally, our model is proven to be portable to new types of events effectively. Xiao Liu 0029, Heyan Huang, Ge Shi 0002, Bo Wang 0134 |
ACL (1) | 4 |
| 2022 | BIT-WOW at NLPCC-2022 Task5 Track1: Hierarchical Multi-label Classification via Label-Aware Graph Convolutional Network
Bo Wang 0134, Yi-Fan Lu, Xiaochi Wei, Xiao Liu 0029, Ge Shi 0002, Changsen Yuan, Heyan Huang, Chong Feng 0001, Xianling Mao |
NLPCC (2) | 1 |
| 2021 | BIT-Event at NLPCC-2021 Task 3: Subevent Identification via Adversarial Training
Xiao Liu 0029, Ge Shi 0002, Bo Wang 0134, Changsen Yuan, Heyan Huang, Chong Feng 0001, Lifang Wu |
NLPCC (2) | 3 |