Bei Yan

dblp:64/9772 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A Survey of Multimodal Hallucination Evaluation and Detection
Yuecong Min, Jie Zhang 0071, Bei Yan, Shiguang Shan
Int. J. Comput. Vis.4
2026 MM-MoralBench: A multimodal moral evaluation benchmark for large vision-language models
Bei Yan, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001
Pattern Recognit.1
2025 Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs
abstract
Currently many benchmarks have been proposed to evaluate the perception ability of the Large Vision-Language Models (LVLMs). However, most benchmarks conduct questions by selecting images from existing datasets, resulting in the potential data leakage. Besides, these benchmarks merely focus on evaluating LVLMs on the realistic style images and clean scenarios, leaving the multi-stylized images and noisy scenarios unexplored. In response to these challenges, we propose a dynamic and scalable benchmark named Dysca for evaluating LVLMs by leveraging synthesis images. Specifically, we leverage Stable Diffusion and design a rule-based method to dynamically generate novel images, questions and the corresponding answers. We consider 51 kinds of image styles and evaluate the perception capability in 20 subtasks. Moreover, we conduct evaluations under 4 scenarios (i.e., Clean, Corruption, Print Attacking and Adversarial Attacking) and 3 question types (i.e., Multi-choices, True-or-false and Free-form). Thanks to the generative paradigm, Dysca serves as a scalable benchmark for easily adding new subtasks and scenarios. A total of 24 advanced open-source LVLMs and 2 close-source LVLMs are evaluated on Dysca, revealing the drawbacks of current LVLMs. The benchmark is released in anonymous github page \url{https://github.com/Benchmark-Dysca/Dysca}.
Jie Zhang 0071, Mengqi Lei, Zheng Yuan 0005, Bei Yan, Shiguang Shan, Xilin Chen 0001
ICLR5
2025 SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
abstract
Despite rapid advances, Large Vision-Language Models (LVLMs) still suffer from hallucinations, i.e., generating content inconsistent with input or established world knowledge, which correspond to faithfulness and factuality hallucinations, respectively. Prior studies primarily evaluate faithfulness hallucination at a rather coarse level (e.g., object-level) and lack fine-grained analysis. Additionally, existing benchmarks often rely on costly manual curation or reused public datasets, raising concerns about scalability and data leakage. To address these limitations, we propose an automated data construction pipeline that produces scalable, controllable, and diverse evaluation data. We also design a hierarchical hallucination induction framework with input perturbations to simulate realistic noisy scenarios. Integrating these designs, we construct SHALE, a Scalable HALlucination Evaluation benchmark designed to assess both faithfulness and factuality hallucinations via a fine-grained hallucination categorization scheme. SHALE comprises over 30K image-instruction pairs spanning 12 representative visual perception aspects for faithfulness and 6 knowledge domains for factuality, considering both clean and noisy scenarios. Extensive experiments on over 20 mainstream LVLMs reveal significant factuality hallucinations and high sensitivity to semantic perturbations.
Bei Yan, Yuecong Min, Jie Zhang 0071, Shiguang Shan
ACM Multimedia1
2025 Content Moderation and Hate Speech on Alternative Platforms: A Case Study of BitChute
abstract
Frustration with mainstream social media platforms and their content moderation decisions have prompted many users to search for "anti-censorship" alternatives, such as those in alt-tech, which may be used to share extreme or hateful content. The tension in alt-tech between limiting perceived censorship and reducing hate speech means that content moderation policies are essential but also controversial. Nevertheless, alt-tech content moderation policies are understudied. To address this research gap, we leverage quasi-experimental design to measure the impact of an "incitement to hatred" policy change using 5.2 million comments and the metadata of 800 thousand videos from the alt-tech platform BitChute. We uncover evidence for a "backlash effect," finding that after the implementation of the policy, hate speech increased significantly for comments and video metadata. This study contributes to the literature on content moderation policies in a challenging context where users may not be receptive to perceived impositions.
Jacob Erickson, Bei Yan
Proc. ACM Hum. Comput. Interact.2
2024 Affective Design: The Influence of Facebook Reactions on the Emotional Expression of the 114th US Congress
abstract
Political communication is critical for democracy, but polarized emotions in communication may make careful deliberation difficult. Much of modern political communication occurs on social media, which may exacerbate these challenges. This study examines how the design of social media features impact political communication. We examined how the introduction of Facebook Reactions influenced the posts of the 114th US Congress on the platform. We start by analyzing the emotional content of posts, finding that politicians generally increased their usage of negative emotions in their posts after the feature’s launch. Further analysis showed that increased user engagement preceded the rise in negative emotions, suggesting that politicians were making adjustments based on user feedback. Our results show that the design features of social media can shape online political communication.
Jacob Erickson, Bei Yan
CHI2
2023 Adaptive Adversarial Patch Attack on Face Recognition Models
abstract
Face recognition models have become widely used for identity authentication in scenarios such as cell phone unlocking and financial payment, but they are vulnerable to adversarial examples. Due to the realizability in the physical world, adversarial patch attack has emerged as a significant security threat. However, most existing adversarial patch attack methods focus on only one aspect of patch generation, such as patch location or shape. To overcome this limitation, we propose a novel unified Adaptive Adversarial Patch (AAP) attack framework for targeted attack on face recognition models. Our method comprehensively considers various factors during patch generation, including location, shape, and number. Our approach adaptively selects patch location and number based on saliency map and clustering, while simultaneously deforming patch shape and optimizing perturbations. Extensive experiments under both white-box and black-box settings demonstrate that our proposed method achieves higher attack success rates compared to SOTA methods.
Bei Yan, Jie Zhang 0071, Zheng Yuan 0005, Shiguang Shan
IJCB1
2017 Crowd Diversity and Performance in Wikipedia: The Mediating Effects of Task Conflict and Communication
abstract
Crowd diversity is a key attribute that impacts crowd performance in online collaboration systems. As a structural composition of a crowd, diversity is likely to influence crowd performance through communication processes during collaboration. This study examined how diversity influenced crowd performance under different conditions of task conflict and communication in Wikipedia article production. With a sample of 5,899 articles, we found that contribution diversity positively predicted crowd performance, whereas experience diversity was negatively related to performance. In addition, task communication and conflict partially mediated the relationship between crowd diversity and performance. Task communication positively predicted performance for both forms of diversity. Task conflict, on the other hand, was positively predicted by expertise diversity, but had negative associations with contribution diversity and performance. The findings help unpack the reasons for differential effects of diversity on crowd performance, and demonstrate the importance of including communication variables when studying online crowd collaboration.
Ruqin Ren, Bei Yan
CHI2