VLDB 2026 Research / reviewers in the wild / expert
Liangzhi Li 0001
dblp:169/4123
· DBLP profile ↗
24ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0002-8879-5957ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can We Trust LLMs for Medical Diagnosis? Evaluating Robustness of Clinical Reasoning Under Perturbation
Zhaofeng Niu, Chongxin Du, Bowen Wang 0002, Yang Song 0010, Liangzhi Li 0001 |
ICIC (23) | 6 |
| 2026 | Can Large Language Models Make In-Character Decisions? Evaluating Role Consistency in Open-Ended Decision Making
Zhaofeng Niu, Luhui Li, Bowen Wang 0002, Yang Song 0010, Liangzhi Li 0001 |
ICIC (23) | 6 |
| 2026 | LegalKQA: A Dataset for Legal Question Rewriter Optimization with Automatic Evaluation Signals
Zhaofeng Niu, Bowen Wang 0002, Wuyunzhaola Borjigin, Liangzhi Li 0001 |
ICIC (22) | 6 |
| 2026 | LIBS: Instructional Action Quality Assessment via Supervoxel-Based Fine-Grained AttributionabstractThe lack of actionable guidance is a fundamental limitation in Action Quality Assessment (AQA), as traditional methods provide overall scores without offering specific insights for improvement. Moreover, existing interpretable approaches often rely on expensive supervised spatial annotations or yield noisy, unsigned saliency maps. To address these challenges, we propose Learning Interpretability Based Supervoxels (LIBS), a novel framework for generating instructional feedback. Distinguishing itself from fully supervised methods, LIBS employs an unsupervised soft-clustering mechanism to segment videos into coherent supervoxels without requiring pixel-level mask annotations. This allows for scalable, fine-grained spatio-temporal analysis while preserving action continuity. Furthermore, we introduce a sensitivity propensity analysis to quantify the contribution of each supervoxel. Unlike traditional attribution methods, this mechanism explicitly decomposes the quality score into positive (strengths) and negative (flaws) components, enabling the system to decode abstract scores into concrete, actionable instructions. Experimental validation across multiple datasets demonstrates that LIBS achieves superior interpretability and efficiency compared to state-of-the-art baselines, marking an improvement from diagnostic to instructional AQA applications. Xiaoyi Tao, Dongxu Ma, Liangzhi Li 0001, Manisha Verma, Lei Chen 0091, Xin Xie 0001, Sheng Chen 0015, Wenxin Li 0001, Jien Kato, Bing Zhang 0015, Xiulong Liu 0001 |
IEEE Trans. Computers | 3 |
| 2025 | Towards Open-World Video Segmentation via Iterative Automatic PromptingabstractLarge-scale pre-trained visual foundation models, such as the Segment Anything Model 2 (SAM2), demonstrate strong performance in video segmentation. However, they require multiple iterations of sophisticated manual prompts to achieve satisfactory results. This paper introduces a method that integrates existing visual foundation models without the need for additional training, named IAP-SAM2, which enables iterative automatic prompting for open-world video segmentation. An innovative automatic prompting mechanism is designed to allow SAM2 to segment target object on videos. Additionally, we propose a multi-round iterative prompt generation strategy based on feature similarity, along with a voting mechanism to refine object segmentation and address occlusion issues in video segmentation. Experimental results show that IAP-SAM2 outperforms existing open-world segmentation approaches on the DAVIS and LVOS datasets, particularly in handling complex videos with multiple targets and object occlusions, while maintaining robust segmentation performance. In the era of emerging foundation models, this work unlocks the potential of these models for automated video segmentation and expands the pathway for leveraging combined foundation models to address real-world challenges. Liangzhi Li 0001, Zhouqiang Jiang, Xingfu Cheng, Zhaofeng Niu, Bowen Wang 0002, Guangshun Li |
IJCNN | 1 |
| 2025 | Role-Playing in Vision-Language Models: A Comprehensive Evaluation of Image Description PerformanceabstractDespite the significant advances made in large language models (LLMs) and vision-language models (VLMs), research on role-playing (RP) within VLMs remains in its nascent stages, with a conspicuous lack of systematic evaluations of their role-playing capabilities. This study aims to address this gap by exploring how specific prompts related to different roles influence VLM performance in image description tasks. We propose a comprehensive evaluation framework specifically designed to assess the role-playing abilities of VLMs, encompassing classification accuracy, semantic similarity, lexical diversity, and potential hazards of generated content. Our findings indicate that as the age of the roles increases, the performance of VLMs improves significantly; models portraying older roles produce descriptions that are semantically more accurate and contextually richer. Furthermore, the introduction of domain-specific roles markedly enhances model performance, particularly when expert knowledge aligns with task requirements. This study not only underscores the necessity for a systematic assessment of role-playing capabilities in VLMs but also provides valuable insights for the development of multimodal systems that exhibit contextual awareness and moral responsibility across various applications. Zhaofeng Niu, Xiaoya Chang, Bowen Wang 0002, Xingfu Cheng, Guangshun Li, Liangzhi Li 0001 |
IJCNN | 6 |
| 2024 | Can Multiple-choice Questions Really Be Useful in Detecting the Abilities of LLMs?abstractMultiple-choice questions (MCQs) are widely used in the evaluation of large language models (LLMs) due to their simplicity and efficiency. However, there are concerns about whether MCQs can truly measure LLM’s capabilities, particularly in knowledge-intensive scenarios where long-form generation (LFG) answers are required. The misalignment between the task and the evaluation method demands a thoughtful analysis of MCQ’s efficacy, which we undertake in this paper by evaluating nine LLMs on four question-answering (QA) datasets in two languages: Chinese and English. We identify a significant issue: LLMs exhibit an order sensitivity in bilingual MCQs, favoring answers located at specific positions, i.e., the first position. We further quantify the gap between MCQs and long-form generation questions (LFGQs) by comparing their direct outputs, token logits, and embeddings. Our results reveal a relatively low correlation between answers from MCQs and LFGQs for identical questions. Additionally, we propose two methods to quantify the consistency and confidence of LLMs’ output, which can be generalized to other QA evaluation benchmarks. Notably, our analysis challenges the idea that the higher the consistency, the greater the accuracy. We also find MCQs to be less reliable than LFGQs in terms of expected calibration error. Finally, the misalignment between MCQs and LFGQs is not only reflected in the evaluation performance but also in the embedding space. Our code and models can be accessed at https://github.com/Meetyou-AI-Lab/Can-MC-Evaluate-LLMs. Wangyue Li, Liangzhi Li 0001, Tong Xiang, Noa Garcia |
LREC/COLING | 2 |
| 2024 | Estimating Socioeconomic Proxy Variables Using Multimodal Deep Learning Models
Yanbing Bai, Zelan Zhu, Huixue Su, Liangzhi Li 0001 |
ICIC (13) | 5 |
| 2024 | A Semantic Segmentation Method for Skin Lesion Images Based on ViT
Zhaofeng Niu, Zhouqiang Jiang, Bowen Wang 0002, Guangshun Li, Liangzhi Li 0001 |
ICONIP (8) | 6 |
| 2024 | Retrieving Emotional Stimuli in ArtworksabstractWe introduce an emotional stimuli retrieval task that targets extracting emotional regions that evoke people's emotions (i.e., emotional stimuli) in artworks. This task offers new challenges to the community because of the diversity of artwork styles and the subjectivity of emotions, which can be a suitable testbed for benchmarking the capability of the current neural networks to deal with human emotion. For this task, we construct a dataset called APOLO for quantifying emotional stimuli retrieval performance in artworks by crowd-sourcing pixel-level annotation of emotional stimuli. APOLO contains 6,781 emotional stimuli in 4,718 artworks for validation and testing. We also evaluate eight baseline methods, including a dedicated one, to show the difficulties of the task and the limitations of the current techniques through qualitative and quantitative experiments. Our data and methods are available in https://github.com/Tianwei3989/apolo. Tianwei Chen 0001, Noa Garcia, Liangzhi Li 0001, Yuta Nakashima |
ICMR | 3 |
| 2024 | Instruct Me More! Random Prompting for Visual In-Context LearningabstractLarge-scale models trained on extensive datasets, have emerged as the preferred approach due to their high generalizability across various tasks. In-context learning (ICL), a popular strategy in natural language processing, uses such models for different tasks by providing instructive prompts but without updating model parameters. This idea is now being explored in computer vision, where an input-output image pair (called an in-context pair) is supplied to the model with a query image as a prompt to exemplify the desired output. The efficacy of visual ICL often depends on the quality of the prompts. We thus introduce a method coined Instruct Me More (InMeMo), which augments in-context pairs with a learnable perturbation (prompt), to explore its potential. Our experiments on mainstream tasks reveal that InMeMo surpasses the current state-of-the-art performance. Specifically, compared to the baseline without learnable prompt, InMeMo boosts mIoU scores by 7.35 and 15.13 for foreground segmentation and single object detection tasks, respectively. Our findings suggest that InMeMo offers a versatile and efficient way to enhance the performance of visual ICL with lightweight training. Code is available at https://github.com/Jackieam/InMeMo. Bowen Wang 0002, Liangzhi Li 0001, Yuta Nakashima, Hajime Nagahara |
WACV | 3 |
| 2024 | Improving facade parsing with vision transformers and line integration
Bowen Wang 0002, Jiaxin Zhang 0018, Yunqin Li, Liangzhi Li 0001, Yuta Nakashima |
Adv. Eng. Informatics | 5 |
| 2023 | Learning Bottleneck Concepts in Image ClassificationabstractInterpreting and explaining the behavior of deep neural networks is critical for many tasks. Explainable AI provides a way to address this challenge, mostly by providing per-pixel relevance to the decision. Yet, interpreting such explanations may require expert knowledge. Some recent attempts toward interpretability adopt a concept-based framework, giving a higher-level relationship between some concepts and model decisions. This paper proposes Bottleneck Concept Learner (BotCL), which represents an image solely by the presence/absence of concepts learned through training over the target task without explicit supervision over the concepts. It uses self-supervision and tailored regularizers so that learned concepts can be human-understandable. Using some image classification tasks as our testbed, we demonstrate BotCL's potential to rebuild neural networks for better interpretability11Code is avaliable at https://github.com/wbw520/BotCL and a simple demo is available at https://botcl.liangzhili.com/. Bowen Wang 0002, Liangzhi Li 0001, Yuta Nakashima, Hajime Nagahara |
CVPR | 2 |
| 2023 | Explaining Federated Learning Through Concepts in Image Classification
Jiaxin Shen, Xiaoyi Tao, Liangzhi Li 0001, Zhiyang Li 0001, Bowen Wang 0002 |
ICA3PP (5) | 3 |
| 2023 | Match them up: visually explainable few-shot image classificationabstractAbstract Few-shot learning (FSL) approaches, mostly neural network-based, assume that pre-trained knowledge can be obtained from base (seen) classes and transferred to novel (unseen) classes. However, the black-box nature of neural networks makes it difficult to understand what is actually transferred, which may hamper FSL application in some risk-sensitive areas. In this paper, we reveal a new way to perform FSL for image classification, using a visual representation from the backbone model and patterns generated by a self-attention based explainable module. The representation weighted by patterns only includes a minimum number of distinguishable features and the visualized patterns can serve as an informative hint on the transferred knowledge. On three mainstream datasets, experimental results prove that the proposed method can enable satisfying explainability and achieve high classification results. Code is available at https://github.com/wbw520/MTUNet . Bowen Wang 0002, Liangzhi Li 0001, Manisha Verma, Yuta Nakashima, Ryo Kawasaki, Hajime Nagahara |
Appl. Intell. | 2 |
| 2022 | One-shot pruning of gated recurrent unit neural network by sensitivity for time-series predictionabstractAlthough deep learning models have been successfully adopted in many applications, they are facing challenges to be deployed on energy-limited devices (e.g., some mobile devices, etc.) due to their high computation complexity. In this paper, we focus on reducing the costs of Gated Recurrent Units (GRUs) for time-series prediction tasks and we propose a new pruning method that can recognize and remove the neural connections that have little influence on the network loss, using a controllable threshold on the absolute value of the pre-trained GRU weights. This is different from existing approaches which usually try to find and preserve the connections with large weight values. We further propose a sparse-connection GRU model (SCGRU) that only needs a one-time pruning (with fine-tuning), rather than using multiple prune-retrain cycles. A large number of experimental results demonstrate that the proposed method is able to largely reduce the storage and computation costs while achieving the state-of-arts performance in two datasets. Code is available ( https://github.com/imLingo/SCGRU). Xiangzheng Ling, Liangzhi Li 0001, Liyan Xiong, Xiaohui Huang 0003 |
Neurocomputing | 3 |
| 2021 | SCOUTER: Slot Attention-based Classifier for Explainable Image RecognitionabstractExplainable artificial intelligence has been gaining attention in the past few years. However, most existing methods are based on gradients or intermediate features, which are not directly involved in the decision-making process of the classifier. In this paper, we propose a slot attention-based classifier called SCOUTER for transparent yet accurate classification. Two major differences from other attention-based methods include: (a) SCOUTER’s explanation is involved in the final confidence for each category, offering more intuitive interpretation, and (b) all the categories have their corresponding positive or negative explanation, which tells "why the image is of a certain category" or "why the image is not of a certain category." We design a new loss tailored for SCOUTER that controls the model’s behavior to switch between positive and negative explanations, as well as the size of explanatory regions. Experimental results show that SCOUTER can give better visual explanations in terms of various metrics while keeping good accuracy on small and medium-sized datasets. Code is available1. Liangzhi Li 0001, Bowen Wang 0002, Manisha Verma, Yuta Nakashima, Ryo Kawasaki, Hajime Nagahara |
ICCV | 1 |
| 2021 | Image Retrieval by Hierarchy-aware Deep Hashing Based on Multi-task LearningabstractDeep hashing has been widely used to approximate nearest-neighbor search for image retrieval tasks. Most of them are trained with image-label pairs without any inter-label relationship, which may not make full use of the real-world data. This paper presents deep hashing, named HA2SH, that leverages multiple types of labels with hierarchical structures that an ethnological museum assigns to their artifacts. We experimentally prove that HA2SH can learn to generate hashes that give a better retrieval performance. Our code is available at https://github.com/wbw520/minpaku. Bowen Wang 0002, Liangzhi Li 0001, Yuta Nakashima, Takehiro Yamamoto, Hiroaki Ohshima, Yoshiyuki Shoji, Kenro Aihara, Noriko Kando |
ICMR | 2 |
| 2020 | IterNet: Retinal Image Segmentation Utilizing Structural Redundancy in Vessel NetworksabstractRetinal vessel segmentation is of great interest for diagnosis of retinal vascular diseases. To further improve the performance of vessel segmentation, we propose IterNet, a new model based on UNet [1], with the ability to find obscured details of the vessel from the segmented vessel image itself, rather than the raw input image. IterNet consists of multiple iterations of a mini-UNet, which can be 4× deeper than the common UNet. IterNet also adopts the weight-sharing and skip-connection features to facilitate training; therefore, even with such a large architecture, IterNet can still learn from merely 10~20 labeled images, without pre-training or any prior knowledge. IterNet achieves AUCs of 0.9816, 0.9851, and 0.9881 on three mainstream datasets, namely DRIVE, CHASE-DB1, and STARE, respectively, which currently are the best scores in the literature. The source code is available1. Liangzhi Li 0001, Manisha Verma, Yuta Nakashima, Hajime Nagahara, Ryo Kawasaki |
WACV | 1 |
| 2019 | Sustainable CNN for Robotic: An Offloading Game in the 3D Vision ComputationabstractThree-dimensional (3D) scene understanding is of great significance to many robotic applications. With the huge development of the deep learning methods, especially the convolutional neural network (CNN), 3D robotic vision has achieved a satisfactory performance. However, in most scenarios, sustainability becomes a severe problem, and few existing approaches pay enough attention to energy consumption. In this paper, we propose an energy-aware system for sustainable robotic 3D vision. Our contributions mainly include: 1) an effective CNN model for the 3D scene understanding; and 2) an offloading strategy to make the deep model more sustainable. First, we design a deep CNN model to analyze the 3D point cloud data. The proposed model contains 92 layers for a state-of-the-art recognition accuracy, which, however, bring a big burden to the computing hardware. Then, we formulate this deep learning computation problem as a non-cooperative game, and adopt a heuristic algorithm to balance the local computing and cloud offloading, in order to obtain an optimal solution, in which both the efficiency and energy-saving are taken into account. Simulations demonstrate that our approach is robust and efficient, and outperforms the state-of-the-art in several related tasks. Liangzhi Li 0001, Kaoru Ota, Mianxiong Dong |
IEEE Trans. Sustain. Comput. | 1 |
| 2018 | Enabling 60 GHz Seamless Coverage for Mobile Devices: A Motion Learning ApproachabstractDespite all the benefits 60 GHz networks bring about, such as high network bandwidth, effective data rates, etc., one of its main application scenarios, Line-of- Sight (LOS) communications, still has troubles in actual indoor environments due to its high directionality. Traditional beam training methods are inaccurate and time-wasting, leading to unstable and inefficient wireless networks. Therefore, in this paper, we attempt to address this problem from a new aspect, i.e., assisting the signal adaptation with human mobility prediction. A state-of-the-art long short-term memory (LSTM) model is adopted to analyze the past trajectories and predict the future position, which can serve as an important reference for the transmitters to proactively adjust their beams and provide seamless coverage. In addition, we also design an algorithm to optimize the beam selection problem and improve the network quality. To the best of our knowledge, this is the first work in the field to use deep learning models for the beam selection problem. Simulations demonstrate that our approach is robust and efficient, and outperforms the state-of-the-art in several related tasks. Liangzhi Li 0001, Kaoru Ota, Mianxiong Dong, Christos V. Verikoukis |
GLOBECOM | 1 |
| 2018 | Human in the Loop: Distributed Deep Model for Mobile CrowdsensingabstractWith the proliferation of mobile devices, crowdsensing has become an appealing technique to collect and process big data. Meanwhile, the rise of fifth generation wireless systems, especially the new cellular base stations with computing ability, brings about the revolutionary edge computing. Although many approaches regarding the mobile crowdsensing have emerged in the last few years, very few of them are focused on the combination of edge computing and crowdsensing. In this paper, we adopt the state-of-the-art edge computing method to solve the crowdsensing problem with the real-time sensing data, and more importantly, make human be in the loop again, in order to respect the users’ willing and privacy. A distributed deep learning model is adopted to extract features from the captured data, which is not only a compression process to reduce the communication cost, but an encryption procedure for safety protection. The proposed model enables the crowdsensing system to fully harness the computing capacity of edge nodes and devices, and obtain a strong data analysis ability to process the captured data. Simulations demonstrate that our approach is robust and efficient, and outperforms other strategies in several related tasks. Liangzhi Li 0001, Kaoru Ota, Mianxiong Dong |
IEEE Internet Things J. | 1 |
| 2018 | Deep Learning for Smart Industry: Efficient Manufacture Inspection System With Fog ComputingabstractWith the rapid development of Internet of things devices and network infrastructure, there have been a lot of sensors adopted in the industrial productions, resulting in a large size of data. One of the most popular examples is the manufacture inspection, which is to detect the defects of the products. In order to implement a robust inspection system with higher accuracy, we propose a deep learning based classification model in this paper, which can find the possible defective products. As there may be many assembly lines in one factory, one huge problem in this scenario is how to process such big data in real time. Therefore, we design our system with the concept of fog computing. By offloading the computation burden from the central server to the fog nodes, the system obtains the ability to deal with extremely large data. There are two obvious advantages in our system. The first one is that we adapt the convolutional neural network model to the fog computing environment, which significantly improves its computing efficiency. The other one is that we work out an inspection model, which can simultaneously indicate the defect type and its degree. The experiments well prove that the proposed method is robust and efficient. Liangzhi Li 0001, Kaoru Ota, Mianxiong Dong |
IEEE Trans. Ind. Informatics | 1 |
| 2017 | Eyes in the Dark: Distributed Scene Understanding for Disaster ManagementabstractRobotic is a great substitute for human to explore the dangerous areas, and will also be a great help for disaster management. Although the rise of depth sensor technologies gives a huge boost to robotic vision research, traditional approaches cannot be applied to disaster-handling robots directly due to some limitations. In this paper, we focus on the 3D robotic perception, and propose a view-invariant Convolutional Neural Network (CNN) Model for scene understanding in disaster scenarios. The proposed system is highly distributed and parallel, which is of great help to improve the efficiency of network training. In our system, two individual CNNs are used to, respectively, propose objects from input data and classify their categories. We attempt to overcome the difficulties and restrictions caused by disasters using several specially-designed multi-task loss functions. The most significant advantage in our work is that the proposed method can learn a view-invariant feature with no requirement on RGB data, which is essential for harsh, disordered and changeable environments. Additionally, an effective optimization algorithm to accelerate the learning process is also included in our work. Simulations demonstrate that our approach is robust and efficient, and outperforms the state-of-the-art in several related tasks. Liangzhi Li 0001, Kaoru Ota, Mianxiong Dong, Wuyunzhaola Borjigin |
IEEE Trans. Parallel Distributed Syst. | 1 |