VLDB 2026 Research / reviewers in the wild / expert
Bowen Wang 0002
dblp:64/4732-2
· DBLP profile ↗
19ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-2911-5595ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can We Trust LLMs for Medical Diagnosis? Evaluating Robustness of Clinical Reasoning Under Perturbation
Zhaofeng Niu, Chongxin Du, Bowen Wang 0002, Yang Song 0010, Liangzhi Li 0001 |
ICIC (23) | 3 |
| 2026 | Can Large Language Models Make In-Character Decisions? Evaluating Role Consistency in Open-Ended Decision Making
Zhaofeng Niu, Luhui Li, Bowen Wang 0002, Yang Song 0010, Liangzhi Li 0001 |
ICIC (23) | 3 |
| 2026 | LegalKQA: A Dataset for Legal Question Rewriter Optimization with Automatic Evaluation Signals
Zhaofeng Niu, Bowen Wang 0002, Wuyunzhaola Borjigin, Liangzhi Li 0001 |
ICIC (22) | 3 |
| 2025 | Putting People in LLMs' Shoes: Generating Better Answers via Question RewriterabstractLarge Language Models (LLMs) have demonstrated significant capabilities, particularly in the domain of question answering (QA). However, their effectiveness in QA is often undermined by the vagueness of user questions. To address this issue, we introduce single-round instance-level prompt optimization, referred to as question rewriter. By enhancing the intelligibility of human questions for black-box LLMs, our question rewriter improves the quality of generated answers. The rewriter is optimized using direct preference optimization based on feedback collected from automatic criteria for evaluating generated answers; therefore, its training does not require costly human annotations. The experiments across multiple black-box LLMs and long-form question answering (LFQA) datasets demonstrate the efficacy of our method. This paper provides a practical framework for training question rewriters and sets a precedent for future explorations in prompt optimization within LFQA tasks. Bowen Wang 0002, Zhouqiang Jiang, Yuta Nakashima |
AAAI | 2 |
| 2025 | ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-trainingabstractRecent approaches for visually-rich document understanding (VrDU) uses manually annotated semantic groups, where a semantic group encompasses all semantically relevant but not obviously grouped words. As OCR tools are unable to automatically identify such grouping, we argue that current VrDU approaches are unrealistic. We thus introduce a new variant of the VrDU task, real-world visually-rich document understanding (ReVrDU), that does not allow for using manually annotated semantic groups. We also propose a new method, ReLayout, compliant with the ReVrDU scenario, which learns to capture semantic grouping through arranging words and bringing the representations of words that belong to the potential same semantic group closer together. Our experimental results demonstrate the performance of existing methods is deteriorated with the ReVrDU task, while ReLayout shows superiour performance. Zhouqiang Jiang, Bowen Wang 0002, Yuta Nakashima |
COLING | 2 |
| 2025 | Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the UnknownabstractThe real value of knowledge lies not just in its accumulation, but in its potential to be harnessed effectively to conquer the unknown. Although recent multimodal large language models (MLLMs) exhibit impressing multimodal capabilities, they often fail in rarely encountered domain-specific tasks due to limited relevant knowledge. To explore this, we adopt visual game cognition as a testbed and select Monster Hunter: World as the target to construct a multimodal knowledge graph (MH-MMKG), which incorporates multi-modalities and intricate entity relations. We also design a series of challenging queries based on MH-MMKG to evaluate the models' ability for complex knowledge retrieval and reasoning. Furthermore, we propose a multi-agent retriever that enables a model to autonomously search relevant knowledge without additional training. Experimental results show that our approach significantly enhances the performance of MLLMs, providing a new perspective on multimodal knowledge-augmented reasoning and laying a solid foundation for future research. Bowen Wang 0002, Zhouqiang Jiang, Yasuaki Susumu, Shotaro Miwa, Tianwei Chen 0001, Yuta Nakashima |
ICCV | 1 |
| 2025 | Towards Open-World Video Segmentation via Iterative Automatic PromptingabstractLarge-scale pre-trained visual foundation models, such as the Segment Anything Model 2 (SAM2), demonstrate strong performance in video segmentation. However, they require multiple iterations of sophisticated manual prompts to achieve satisfactory results. This paper introduces a method that integrates existing visual foundation models without the need for additional training, named IAP-SAM2, which enables iterative automatic prompting for open-world video segmentation. An innovative automatic prompting mechanism is designed to allow SAM2 to segment target object on videos. Additionally, we propose a multi-round iterative prompt generation strategy based on feature similarity, along with a voting mechanism to refine object segmentation and address occlusion issues in video segmentation. Experimental results show that IAP-SAM2 outperforms existing open-world segmentation approaches on the DAVIS and LVOS datasets, particularly in handling complex videos with multiple targets and object occlusions, while maintaining robust segmentation performance. In the era of emerging foundation models, this work unlocks the potential of these models for automated video segmentation and expands the pathway for leveraging combined foundation models to address real-world challenges. Liangzhi Li 0001, Zhouqiang Jiang, Xingfu Cheng, Zhaofeng Niu, Bowen Wang 0002, Guangshun Li |
IJCNN | 6 |
| 2025 | Role-Playing in Vision-Language Models: A Comprehensive Evaluation of Image Description PerformanceabstractDespite the significant advances made in large language models (LLMs) and vision-language models (VLMs), research on role-playing (RP) within VLMs remains in its nascent stages, with a conspicuous lack of systematic evaluations of their role-playing capabilities. This study aims to address this gap by exploring how specific prompts related to different roles influence VLM performance in image description tasks. We propose a comprehensive evaluation framework specifically designed to assess the role-playing abilities of VLMs, encompassing classification accuracy, semantic similarity, lexical diversity, and potential hazards of generated content. Our findings indicate that as the age of the roles increases, the performance of VLMs improves significantly; models portraying older roles produce descriptions that are semantically more accurate and contextually richer. Furthermore, the introduction of domain-specific roles markedly enhances model performance, particularly when expert knowledge aligns with task requirements. This study not only underscores the necessity for a systematic assessment of role-playing capabilities in VLMs but also provides valuable insights for the development of multimodal systems that exhibit contextual awareness and moral responsibility across various applications. Zhaofeng Niu, Xiaoya Chang, Bowen Wang 0002, Xingfu Cheng, Guangshun Li, Liangzhi Li 0001 |
IJCNN | 3 |
| 2024 | A Semantic Segmentation Method for Skin Lesion Images Based on ViT
Zhaofeng Niu, Zhouqiang Jiang, Bowen Wang 0002, Guangshun Li, Liangzhi Li 0001 |
ICONIP (8) | 4 |
| 2024 | DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language ModelsabstractLarge language models (LLMs) have recently showcased remarkable capabilities, spanning a wide range of tasks and applications, including those in the medical domain. Models like GPT-4 excel in medical question answering but may face challenges in the lack of interpretability when handling complex tasks in real clinical settings. We thus introduce the diagnostic reasoning dataset for clinical notes (DiReCT), aiming at evaluating the reasoning ability and interpretability of LLMs compared to human doctors. It contains 511 clinical notes, each meticulously annotated by physicians, detailing the diagnostic reasoning process from observations in a clinical note to the final diagnosis. Additionally, a diagnostic knowledge graph is provided to offer essential knowledge for reasoning, which may not be covered in the training data of existing LLMs. Evaluations of leading LLMs on DiReCT bring out a significant gap between their reasoning ability and that of human doctors, highlighting the critical need for models that can reason effectively in real-world clinical scenarios. Bowen Wang 0002, Jiuyang Chang, Yiming Qian, Guoxin Chen, Zhouqiang Jiang, Yuta Nakashima, Hajime Nagahara |
NeurIPS | 1 |
| 2024 | Instruct Me More! Random Prompting for Visual In-Context LearningabstractLarge-scale models trained on extensive datasets, have emerged as the preferred approach due to their high generalizability across various tasks. In-context learning (ICL), a popular strategy in natural language processing, uses such models for different tasks by providing instructive prompts but without updating model parameters. This idea is now being explored in computer vision, where an input-output image pair (called an in-context pair) is supplied to the model with a query image as a prompt to exemplify the desired output. The efficacy of visual ICL often depends on the quality of the prompts. We thus introduce a method coined Instruct Me More (InMeMo), which augments in-context pairs with a learnable perturbation (prompt), to explore its potential. Our experiments on mainstream tasks reveal that InMeMo surpasses the current state-of-the-art performance. Specifically, compared to the baseline without learnable prompt, InMeMo boosts mIoU scores by 7.35 and 15.13 for foreground segmentation and single object detection tasks, respectively. Our findings suggest that InMeMo offers a versatile and efficient way to enhance the performance of visual ICL with lightweight training. Code is available at https://github.com/Jackieam/InMeMo. Bowen Wang 0002, Liangzhi Li 0001, Yuta Nakashima, Hajime Nagahara |
WACV | 2 |
| 2024 | Improving facade parsing with vision transformers and line integration
Bowen Wang 0002, Jiaxin Zhang 0018, Yunqin Li, Liangzhi Li 0001, Yuta Nakashima |
Adv. Eng. Informatics | 1 |
| 2023 | Learning Bottleneck Concepts in Image ClassificationabstractInterpreting and explaining the behavior of deep neural networks is critical for many tasks. Explainable AI provides a way to address this challenge, mostly by providing per-pixel relevance to the decision. Yet, interpreting such explanations may require expert knowledge. Some recent attempts toward interpretability adopt a concept-based framework, giving a higher-level relationship between some concepts and model decisions. This paper proposes Bottleneck Concept Learner (BotCL), which represents an image solely by the presence/absence of concepts learned through training over the target task without explicit supervision over the concepts. It uses self-supervision and tailored regularizers so that learned concepts can be human-understandable. Using some image classification tasks as our testbed, we demonstrate BotCL's potential to rebuild neural networks for better interpretability11Code is avaliable at https://github.com/wbw520/BotCL and a simple demo is available at https://botcl.liangzhili.com/. Bowen Wang 0002, Liangzhi Li 0001, Yuta Nakashima, Hajime Nagahara |
CVPR | 1 |
| 2023 | Explaining Federated Learning Through Concepts in Image Classification
Jiaxin Shen, Xiaoyi Tao, Liangzhi Li 0001, Zhiyang Li 0001, Bowen Wang 0002 |
ICA3PP (5) | 5 |
| 2023 | CARE-MI: Chinese Benchmark for Misinformation Evaluation in Maternity and Infant CareabstractThe recent advances in natural language processing (NLP), have led to a new trend of applying large language models (LLMs) to real-world scenarios. While the latest LLMs are astonishingly fluent when interacting with humans, they suffer from the misinformation problem by unintentionally generating factually false statements. This can lead to harmful consequences, especially when produced within sensitive contexts, such as healthcare. Yet few previous works have focused on evaluating misinformation in the long-form (LF) generation of LLMs, especially for knowledge-intensive topics. Moreover, although LLMs have been shown to perform well in different languages, misinformation evaluation has been mostly conducted in English. To this end, we present a benchmark, CARE-MI, for evaluating LLM misinformation in: 1) a sensitive topic, specifically the maternity and infant care domain; and 2) a language other than English, namely Chinese. Most importantly, we provide an innovative paradigm for building LF generation evaluation benchmarks that can be transferred to other knowledge-intensive domains and low-resourced languages. Our proposed benchmark fills the gap between the extensive usage of LLMs and the lack of datasets for assessing the misinformation generated by these models. It contains 1,612 expert-checked questions, accompanied with human-selected references. Using our benchmark, we conduct extensive experiments and found that current Chinese LLMs are far from perfect in the topic of maternity and infant care. In an effort to minimize the reliance on human resources for performance evaluation, we offer off-the-shelf judgment models for automatically assessing the LF output of LLMs given benchmark questions. Moreover, we compare potential solutions for LF generation evaluation and provide insights for building better automated metrics. Tong Xiang, Liangzhi Li 0004, Wangyue Li, Mingbai Bai, Bowen Wang 0002, Noa Garcia |
NeurIPS | 6 |
| 2023 | Panoptic-aware Image-to-Image TranslationabstractDespite remarkable progress in image translation, the complex scene with multiple discrepant objects remains a challenging problem. The translated images have low fidelity and tiny objects in fewer details causing unsatisfactory performance in object recognition. Without thorough object perception (i.e., bounding boxes, categories, and masks) of images as prior knowledge, the style transformation of each object will be difficult to track in translation. We propose panoptic-aware generative adversarial networks (PanopticGAN) for image-to-image translation together with a compact panoptic segmentation dataset. The panoptic perception (i.e., foreground instances and background semantics of the image scene) is extracted to achieve alignment between object content codes of the input domain and panoptic-level style codes sampled from the target style space, then refined by a proposed feature masking module for sharping object boundaries. The image-level combination between content and sampled style codes is also merged for higher fidelity image generation. Our proposed method was systematically compared with different competing methods and obtained significant improvement in both image quality and object recognition performance. Photchara Ratsamee, Bowen Wang 0002, Zhaojie Luo, Yuuki Uranishi, Manabu Higashida, Haruo Takemura |
WACV | 3 |
| 2023 | Match them up: visually explainable few-shot image classificationabstractAbstract Few-shot learning (FSL) approaches, mostly neural network-based, assume that pre-trained knowledge can be obtained from base (seen) classes and transferred to novel (unseen) classes. However, the black-box nature of neural networks makes it difficult to understand what is actually transferred, which may hamper FSL application in some risk-sensitive areas. In this paper, we reveal a new way to perform FSL for image classification, using a visual representation from the backbone model and patterns generated by a self-attention based explainable module. The representation weighted by patterns only includes a minimum number of distinguishable features and the visualized patterns can serve as an informative hint on the transferred knowledge. On three mainstream datasets, experimental results prove that the proposed method can enable satisfying explainability and achieve high classification results. Code is available at https://github.com/wbw520/MTUNet . Bowen Wang 0002, Liangzhi Li 0001, Manisha Verma, Yuta Nakashima, Ryo Kawasaki, Hajime Nagahara |
Appl. Intell. | 1 |
| 2021 | SCOUTER: Slot Attention-based Classifier for Explainable Image RecognitionabstractExplainable artificial intelligence has been gaining attention in the past few years. However, most existing methods are based on gradients or intermediate features, which are not directly involved in the decision-making process of the classifier. In this paper, we propose a slot attention-based classifier called SCOUTER for transparent yet accurate classification. Two major differences from other attention-based methods include: (a) SCOUTER’s explanation is involved in the final confidence for each category, offering more intuitive interpretation, and (b) all the categories have their corresponding positive or negative explanation, which tells "why the image is of a certain category" or "why the image is not of a certain category." We design a new loss tailored for SCOUTER that controls the model’s behavior to switch between positive and negative explanations, as well as the size of explanatory regions. Experimental results show that SCOUTER can give better visual explanations in terms of various metrics while keeping good accuracy on small and medium-sized datasets. Code is available1. Liangzhi Li 0001, Bowen Wang 0002, Manisha Verma, Yuta Nakashima, Ryo Kawasaki, Hajime Nagahara |
ICCV | 2 |
| 2021 | Image Retrieval by Hierarchy-aware Deep Hashing Based on Multi-task LearningabstractDeep hashing has been widely used to approximate nearest-neighbor search for image retrieval tasks. Most of them are trained with image-label pairs without any inter-label relationship, which may not make full use of the real-world data. This paper presents deep hashing, named HA2SH, that leverages multiple types of labels with hierarchical structures that an ethnological museum assigns to their artifacts. We experimentally prove that HA2SH can learn to generate hashes that give a better retrieval performance. Our code is available at https://github.com/wbw520/minpaku. Bowen Wang 0002, Liangzhi Li 0001, Yuta Nakashima, Takehiro Yamamoto, Hiroaki Ohshima, Yoshiyuki Shoji, Kenro Aihara, Noriko Kando |
ICMR | 1 |