VLDB 2026 Research / reviewers in the wild / expert
Bangzheng Li
dblp:261/9826
· DBLP profile ↗
8ranked-venue papers
2as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Trustworthy machine learning · 54% Information extraction and text analysis · 20% Vision and language · 13% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.8 | 1 | 2024 | BLINK: Multimodal Large Language Models Can See but Not Perceive · ECCV (23) 2024 |
Machine learning › Trustworthy machine learning
counterfactual data augmentation |
0.6 | 1 | 2022 | Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity Typing · EMNLP 2022 |
Natural language and speech › Information extraction and text analysis
entity typing |
0.6 | 1 | 2022 | Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity Typing · EMNLP 2022 |
Machine learning › Trustworthy machine learning › fairness
model bias |
0.6 | 1 | 2022 | Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity Typing · EMNLP 2022 |
Machine learning › Trustworthy machine learning › robustness
spurious correlation |
0.6 | 1 | 2022 | Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity Typing · EMNLP 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation › concept representation
concept mapping |
0.4 | 1 | 2020 | Neural Concept Map Generation for Effective Document Classification with Interpretable Structured Summarization · SIGIR 2020 |
Machine learning › Trustworthy machine learning › interpretability › explainable AI
interpretable classification |
0.4 | 1 | 2020 | Neural Concept Map Generation for Effective Document Classification with Interpretable Structured Summarization · SIGIR 2020 |
Natural language and speech › Information extraction and text analysis
text classification |
0.4 | 1 | 2020 | Neural Concept Map Generation for Effective Document Classification with Interpretable Structured Summarization · SIGIR 2020 |
Machine learning › Trustworthy machine learning
interpretability |
0.2 | 1 | 2024 | BLINK: Multimodal Large Language Models Can See but Not Perceive · ECCV (23) 2024 |
Computer vision › 3D vision
multimodal perception |
0.2 | 1 | 2024 | BLINK: Multimodal Large Language Models Can See but Not Perceive · ECCV (23) 2024 |
Natural language and speech › Language models and text generation › text summarization
document summarization |
0.1 | 1 | 2020 | Neural Concept Map Generation for Effective Document Classification with Interpretable Structured Summarization · SIGIR 2020 |
Natural language and speech › Information extraction and text analysis › document understanding
structured summarization |
0.1 | 1 | 2020 | Neural Concept Map Generation for Effective Document Classification with Interpretable Structured Summarization · SIGIR 2020 |
Methods — techniques the papers use, named apart from their topics
red teaming · 0.8multimodal prompting · 0.8benchmark evaluation · 0.8counterfactual data augmentation · 0.6weakly supervised learning · 0.4graph neural network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | BLINK: Multimodal Large Language Models Can See but Not Perceive
Yushi Hu, Bangzheng Li, Yu Feng 0013, Haoyu Wang 0005, Xudong Lin 0003, Dan Roth 0001, Noah A. Smith, Wei-Chiu Ma, Ranjay Krishna |
ECCV (23) | 3 |
| 2024 | Red Teaming Language Models for Processing Contradictory DialoguesabstractMost language models currently available are prone to self-contradiction during dialogues.To mitigate this issue, this study explores a novel contradictory dialogue processing task that aims to detect and modify contradictory statements in a conversation.This task is inspired by research on context faithfulness and dialogue comprehension, which have demonstrated that the detection and understanding of contradictions often necessitate detailed explanations.We develop a dataset comprising contradictory dialogues, in which one side of the conversation contradicts itself.Each dialogue is accompanied by an explanatory label that highlights the location and details of the contradiction.With this dataset, we present a Red Teaming framework for contradictory dialogue processing.The framework detects and attempts to explain the dialogue, then modifies the existing contradictory content using the explanation.Our experiments demonstrate that the framework improves the ability to detect contradictory dialogues and provides valid explanations.Additionally, it showcases distinct capabilities for modifying such dialogues.Our study highlights the importance of the logical inconsistency problem in conversational AI 1 Prompts Instructions Explanation Xiaofei Wen, Bangzheng Li, Tenghao Huang, Muhao Chen 0001 |
EMNLP | 2 |
| 2024 | Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?abstractBangzheng Li, Ben Zhou, Fei Wang, Xingyu Fu, Dan Roth, Muhao Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Bangzheng Li, Ben Zhou, Fei Wang 0060, Dan Roth 0001, Muhao Chen 0001 |
NAACL-HLT | 1 |
| 2022 | Does Your Model Classify Entities Reasonably? Diagnosing and Mitigating Spurious Correlations in Entity TypingabstractEntity typing aims at predicting one or more words that describe the type(s) of a specific mention in a sentence.Due to shortcuts from surface patterns to annotated entity labels and biased training, existing entity typing models are subject to the problem of spurious correlations.To comprehensively investigate the faithfulness and reliability of entity typing methods, we first systematically define distinct kinds of model biases that are reflected mainly from spurious correlations.Particularly, we identify six types of existing model biases, including mention-context bias, lexical overlapping bias, named entity bias, pronoun bias, dependency bias, and overgeneralization bias.To mitigate model biases, we then introduce a counterfactual data augmentation method.By augmenting the original training set with their debiased counterparts, models are forced to fully comprehend sentences and discover the fundamental cues for entity typing, rather than relying on spurious correlations for shortcuts.Experimental results on the UFET dataset show our counterfactual data augmentation approach helps improve generalization of different entity typing models with consistently better performance on both the original and debiased test sets 1 .Input: Last week I stayed in Treasure Island for two nights when visiting Las Vegas.Gold labels: hotel, resort, location, place Pred labels: island, land, location, place Input: Next day (-> Next twenty-four hour period), after the Slovaks captured Rymanow Zdroj and the the Germans seized Krosno, the Brigade was ordered to withdraw to Sanok and leave Dukla.Gold labels: day, time, event, date, year Pred labels: day, time, date -> time, hour period Input: Kevin Donovan (-> Brennan), after seeing Michael Caine movie about the Zulu uprising, decided to form the Universal Zulu Nation, an organization based on merits derived from art and achievements Nan Xu 0014, Fei Wang 0060, Bangzheng Li, Mingtao Dong, Muhao Chen 0001 |
EMNLP | 3 |
| 2022 | Unified Semantic Typing with Meaningful Label InferenceabstractJames Y. Huang, Bangzheng Li, Jiashu Xu, Muhao Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. James Y. Huang, Bangzheng Li, Muhao Chen 0001 |
NAACL-HLT | 2 |
| 2022 | Ultra-fine Entity Typing with Indirect Supervision from Natural Language InferenceabstractAbstract The task of ultra-fine entity typing (UFET) seeks to predict diverse and free-form words or phrases that describe the appropriate types of entities mentioned in sentences. A key challenge for this task lies in the large number of types and the scarcity of annotated data per type. Existing systems formulate the task as a multi-way classification problem and train directly or distantly supervised classifiers. This causes two issues: (i) the classifiers do not capture the type semantics because types are often converted into indices; (ii) systems developed in this way are limited to predicting within a pre-defined type set, and often fall short of generalizing to types that are rarely seen or unseen in training. This work presents LITE🍻, a new approach that formulates entity typing as a natural language inference (NLI) problem, making use of (i) the indirect supervision from NLI to infer type information meaningfully represented as textual hypotheses and alleviate the data scarcity issue, as well as (ii) a learning-to-rank objective to avoid the pre-defining of a type set. Experiments show that, with limited training data, LITE obtains state-of-the-art performance on the UFET task. In addition, LITE demonstrates its strong generalizability by not only yielding best results on other fine-grained entity typing benchmarks, more importantly, a pre-trained LITE system works well on new data containing unseen types.1 Bangzheng Li, Wenpeng Yin 0001, Muhao Chen 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Fine-Grained Named Entity Recognition with Distant Supervision in COVID-19 LiteratureabstractBiomedical named entity recognition (BioNER) is a fundamental step for mining COVID-19 literature. Existing BioNER datasets cover a few common coarse-grained entity types (e.g., genes, chemicals, and diseases), which cannot be used to recognize highly domain-specific entity types (e.g., animal models of diseases) or emerging ones (e.g., coronaviruses) for COVID-19 studies. We present CORD-NER, a fine-grained named entity recognized dataset of COVID-19 literature (up until May 19, 2020). CORD-NER contains over 12 million sentences annotated via distant supervision. Also included in CORD-NER are 2,000 manually-curated sentences as a test set for performance evaluation. CORD-NER covers 75 fine-grained entity types. In addition to the common biomedical entity types, it covers new entity types specifically related to COVID-19 studies, such as coronaviruses, viral proteins, evolution, and immune responses. The dictionaries of these fine-grained entity types are collected from existing knowledge bases and human-input seed sets. We further present DISTNER, a distantly supervised NER model that relies on a massive unlabeled corpus and a collection of dictionaries to annotate the COVID-19 corpus. DISTNER provides a benchmark performance on the CORD-NER test set for future research. Xuan Wang 0008, Xiangchen Song, Bangzheng Li, Kang Zhou 0002, Qi Li 0012, Jiawei Han 0001 |
BIBM | 3 |
| 2020 | Neural Concept Map Generation for Effective Document Classification with Interpretable Structured SummarizationabstractConcept maps provide concise structured representations for documents regarding their important concepts and interaction links, which have been widely used for document summarization and downstream tasks. However, the construction of concept maps often relies heavily on heuristic design and auxiliary tools. Recent popular neural network models, on the other hand, are shown effective in tasks across various domains, but are short in interpretability and prone to overfitting. In this work, we bridge the gap between concept map construction and neural network models, by designing doc2graph, a novel weakly-supervised text-to-graph neural network, which generates concept maps in the middle and is trained towards document-level tasks like document classification. In our experiments, doc2graph outperforms both its traditional baselines and neural counterparts by significant margins in document classification, while producing high-quality interpretable concept maps as document structured summarization. Carl Yang 0001, Jieyu Zhang 0001, Bangzheng Li, Jiawei Han 0001 |
SIGIR | 4 |