EDBT 2026 Demo / reviewers in the wild / expert
Yeshuang Zhu
dblp:121/4904
· DBLP profile ↗
15ranked-venue papers
2as first author
12since 2021 · last 2025
0009-0009-7918-2962ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Investigating Context-Aware Collaborative Text Entry on Smartphones using Large Language ModelsabstractText entry is a fundamental and ubiquitous task, but users often face challenges such as situational impairments or difficulties in sentence formulation.Motivated by this, we explore the potential of large language models (LLMs) to assist with text entry in realworld contexts.We propose a collaborative smartphone-based text entry system, CATIA, that leverages LLMs to provide text suggestions based on contextual factors, including screen content, time, location, activity, and more.In a 7-day in-the-wild study with 36 participants, the system offered appropriate text suggestions in over 80% of cases.Users exhibited different collaborative behaviors depending on whether they were composing text for interpersonal communication or information services.Additionally, the relevance Yuanchun Shi, Weinan Shi, Meizhu Chen, Yeshuang Zhu, Jinchao Zhang 0001, Chun Yu |
CHI | 8 |
| 2025 | Secret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution AnalysisabstractThe advancement of Generative Adversarial Networks (GANs) and diffusion models significantly enhances the realism of synthetic images, driving progress in image processing and creative design. However, this progress also necessitates the development of effective detection methods, as synthetic images become increasingly difficult to distinguish from real ones. This difficulty leads to societal issues, such as the spread of misinformation, identity theft, and online fraud. While previous detection methods perform well on public benchmarks, they struggle with our benchmark, FakeART, particularly when dealing with the latest models and cross-domain tasks (e.g., photo-to-painting). To address this challenge, we develop a new synthetic image detection technique based on color distribution. Unlike real images, synthetic images often exhibit uneven color distribution. By employing color quantization and restoration techniques, we analyze the color differences before and after image restoration. We discover and prove that these differences closely relate to the uniformity of color distribution. Based on this finding, we extract effective color features and combine them with image features to create a detection model with only 1.4 million parameters. This model achieves state-of-the-art results across various evaluation benchmarks, including the challenging FakeART dataset. Zexi Jia, Chuanwei Huang, Yeshuang Zhu, Hongyan Fei, Xiaoyue Duan, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016 |
CVPR | 3 |
| 2025 | Semantic to Structure: Learning Structural Representations for Infringement DetectionabstractStructural information in images is crucial for aesthetic assessment, and it is widely recognized in the artistic field that imitating the structure of other works significantly infringes on creators’ rights. The advancement of diffusion models has led to AI-generated content imitating artists’ structural creations, yet effective detection methods are still lacking. In this paper, we define this phenomenon as "structural infringement" and propose a corresponding detection method. Additionally, we develop quantitative metrics and create manually annotated datasets for evaluation: the SIA dataset of synthesized data, and the SIR dataset of real data. Due to the current lack of datasets for structural infringement detection, we propose a new data synthesis strategy based on diffusion models and LLM, successfully training a structural infringement detection model. Experimental results show that our method can successfully detect structural infringements and achieve notable improvements on annotated test sets. Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jinchao Zhang 0001, Jie Zhou 0016 |
ICASSP | 4 |
| 2025 | ILDiff: Generate Transparent Animated Stickers by Implicit Layout DistillationabstractHigh-quality animated stickers usually contain transparent channels, which are often ignored by current video generation models. To generate fine-grained animated transparency channels, existing methods can be roughly divided into video matting algorithms and diffusion-based algorithms. The methods based on video matting have poor performance in dealing with semi-open areas in stickers, while diffusion-based methods are often used to model a single image, which will lead to local flicker when modeling animated stickers. In this paper, we firstly propose an ILDiff method to generate animated transparent channels through implicit layout distillation, which solves the problems of semi-open area collapse and no consideration of temporal information in existing methods. Secondly, we create the Transparent Animated Sticker Dataset (TASD), which contains 0.32M high-quality samples with transparent channel, to provide data support for related fields. Extensive experiments demonstrate that ILDiff can produce finer and smoother transparent channels compared to other methods such as Matting Anything and Layer Diffusion. Our code and dataset will be released at link https://xiaoyuan1996.github.io. Yeshuang Zhu, Jie Zhou 0016, Jinchao Zhang 0001 |
ICASSP | 3 |
| 2025 | MCID: Multi-aspect Copyright Infringement Detection for Generated Images
Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jiapei Zhang, Xiaoyue Duan, Jinchao Zhang 0001, Jie Zhou 0016 |
ICCV | 4 |
| 2025 | A Visual Leap in Clip Compositionality Reasoning Through Generation of Counterfactual SetsabstractVision-language models (VLMs) often struggle with compositional reasoning due to insufficient high-quality image-text data. To tackle this challenge, we propose a novel block-based diffusion approach that automatically generates counterfactual datasets without manual annotation. Our method utilizes large language models to identify entities and their spatial relationships. It then independently generates image blocks as "puzzle pieces" coherently arranged according to specified compositional rules. This process creates diverse, high-fidelity counterfactual image-text pairs with precisely controlled variations. In addition, we introduce a specialized loss function that differentiates inter-set from intra-set samples, enhancing training efficiency and reducing the need for negative samples. Experiments demonstrate that fine-tuning VLMs with our counterfactual datasets significantly improves visual reasoning performance. Our approach achieves state-of-the-art results across multiple benchmarks while using substantially less training data than existing methods. Zexi Jia, Chuanwei Huang, Hongyan Fei, Yeshuang Zhu, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016 |
ICCV | 4 |
| 2025 | From Imitation to Innovation: The Emergence of Ai's Unique Artistic Styles and the Challenge of Copyright Protection
Zexi Jia, Chuanwei Huang, Yeshuang Zhu, Hongyan Fei, Jiapei Zhang, Jinchao Zhang 0001, Jie Zhou 0016 |
ICCV | 3 |
| 2025 | Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step Alignment
Renye Yan, Jikang Cheng, Yaozhong Gan, Shikun Sun, Yunfan Yang, Ling Liang 0003, Jinlong Lin, Yeshuang Zhu, Jie Zhou 0001, Junliang Xing, Yimao Cai, Ru Huang 0001 |
ICCV | 9 |
| 2025 | WalkVLM: Aid Visually Impaired People Walking by Vision Language ModelabstractApproximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people. With the recent progress of vision-language models (VLMs), employing VLMs to improve this field has emerged as a popular research topic. However, most existing methods are studied on self-built question-answering datasets, lacking a unified training and testing benchmark for walk guidance. Moreover, in blind walking task, it is necessary to perform real-time streaming video parsing and generate concise yet informative reminders, which poses a great challenge for VLMs that suffer from redundant responses and low inference efficiency. In this paper, we firstly release a diverse, extensive, and unbiased walking awareness dataset, containing 12k video-manual annotation pairs from Europe and Asia to provide a fair training and testing benchmark for blind walking task. Furthermore, a WalkVLM model is proposed, which employs chain of thought for hierarchical planning to generate concise but informative reminders and utilizes temporal-aware adaptive prediction to reduce the temporal redundancy of reminders. Finally, we have established a solid benchmark for blind walking task and verified the advantages of WalkVLM in stream video processing for this task compared to other VLMs. Our dataset and code will be released at anonymous link https://walkvlm2024.github.io. Yeshuang Zhu, Jiapei Zhang, Zexi Jia, Peixiang Luo, Xiaoyue Duan, Jie Zhou 0016, Jinchao Zhang 0001 |
ICCV | 3 |
| 2025 | VSD2M: Large-scale Vision-language Sticker Dataset for Multi-frame Animated Sticker Generationabstractmedia, Nowadays, advanced text-to-video algorithms have spawned numerous general video generation systems that allow users to customize high-quality, photo-realistic videos by only providing simple text prompts. However, creating customized animated stickers, which have lower frame rates and more abstract semantics than videos, is greatly hindered by difficulties in data acquisition and incomplete benchmarks. To facilitate the exploration of researchers in animated sticker generation (ASG) field, we construct the currently largest vision-language sticker dataset named "VSD2M" at a two-million scale that contains static and animated stickers. Furthermore, to improve the performance of traditional video generation methods on ASG tasks with discrete characteristics, we propose a Spatial Temporal Interaction layer that utilizes semantic interaction and detail preservation to address the issue of insufficient information utilization. To our knowledge, this is the most comprehensive benchmark for multi-frame ASG task, and we hope it can provide valuable inspiration for other scholars in intelligent creation. Jiapei Zhang, Yeshuang Zhu, Jie Zhou 0016, Jinchao Zhang 0001 |
ICME | 4 |
| 2025 | ArtFRD: A Fisher-Rao Mixture Metric for Generative Model Aesthetic EvaluationabstractRecent advances in generative modeling have enabled the synthesis of high-quality artistic images. Nevertheless, systematic evaluation of generative models from an aesthetic standpoint is still lacking, which hinders progress in artistic image synthesis. Existing evaluation metrics, such as Fréchet Inception Distance (FID) and CMMD, struggle with aesthetic assessment: they rely on pretrained visual features that overlook nuanced artistic attributes and employ distance functions ill-suited for modeling the diverse, multi-modal distribution of artistic styles. To address these limitations, we propose ArtFRD, a metric specifically designed for generative aesthetic evaluation. Grounded in aesthetic theory, ArtFRD extracts visual features along four key aesthetic dimensions-brushstroke, composition, lighting, and color-to capture fine-grained artistic properties. To model the multi-modal nature of artistic styles, we adopt a Gaussian Mixture Model assumption and derive an efficient approximation of the Fisher-Rao distance, which serves as the final evaluation score. Extensive experiments demonstrate that ArtFRD aligns significantly better with human aesthetic judgments than existing metrics, even across a wide range of artistic styles. These results highlight its potential as a robust and interpretable foundation for future research in generative aesthetic evaluation. Chuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu, Jinchao Zhang 0001, Jie Zhou 0016 |
ACM Multimedia | 4 |
| 2025 | Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive BenchmarkabstractMultimodal language analysis is a rapidly evolving field that leverages multiple modalities to enhance the understanding of high-level semantics underlying human conversational utterances. Despite its significance, little research has investigated the capability of multimodal large language models (MLLMs) to comprehend cognitive-level semantics. In this paper, we introduce MMLA, a comprehensive benchmark specifically designed to address this gap. MMLA comprises over 61K multimodal utterances drawn from both staged and real-world scenarios, covering six core dimensions of multimodal semantics: intent, emotion, dialogue act, sentiment, speaking style, and communication behavior. We evaluate eight mainstream branches of LLMs and MLLMs using three methods: zero-shot inference, supervised fine-tuning, and instruction tuning. Extensive experiments reveal that even fine-tuned models achieve only about 60~70% accuracy, underscoring the limitations of current MLLMs in understanding complex human language. We believe that MMLA will serve as a solid foundation for exploring the potential of large language models in multimodal language analysis and provide valuable resources to advance this field. The datasets and code are open-sourced at https://github.com/thuiar/MMLA. Hanlei Zhang, Hua Xu 0003, Yeshuang Zhu, Peiwu Wang, Haige Zhu, Jie Zhou 0016, Jinchao Zhang 0001 |
NeurIPS | 4 |
| 2019 | QuizBot: A Dialogue-based Adaptive Learning System for Factual KnowledgeabstractAdvances in conversational AI have the potential to enable more engaging and effective ways to teach factual knowledge. To investigate this hypothesis, we created QuizBot, a dialogue-based agent that helps students learn factual knowledge in science, safety, and English vocabulary. We evaluated QuizBot with 76 students through two within-subject studies against a flashcard app, the traditional medium for learning factual knowledge. Though both systems used the same algorithm for sequencing materials, QuizBot led to students recognizing (and recalling) over 20% more correct answers than when students used the flashcard app. Using a conversational agent is more time consuming to practice with, but in a second study, of their own volition, students spent 2.6x more time learning with QuizBot than with flashcards and reported preferring it strongly for casual learning. Our results in this second study showed QuizBot yielded improved learning gains over flashcards on recall. These results suggest that educational chatbot systems may have beneficial use, particularly for learning outside of traditional settings. Sherry Ruan, Justin Xu, Bryce Joe-Kun Tham, Zhengneng Qiu, Yeshuang Zhu, Elizabeth L. Murnane, Emma Brunskill, James A. Landay |
CHI | 6 |
| 2017 | ViVo: Video-Augmented Dictionary for Vocabulary LearningabstractResearch on Computer-Assisted Language Learning (CALL) has shown that the use of multimedia materials such as images and videos can facilitate interpretation and memorization of new words and phrases by providing richer cues than text alone. We present ViVo, a novel video-augmented dictionary that provides an inexpensive, convenient, and scalable way to exploit huge online video resources for vocabulary learning. ViVo automatically generates short video clips from existing movies with the target word highlighted in the subtitles. In particular, we apply a word sense disambiguation algorithm to identify the appropriate movie scenes with adequate contextual information for learning. We analyze the challenges and feasibility of this approach and describe our interaction design. A user study showed that learners were able to retain nearly 30% more new words with ViVo than with a standard bilingual dictionary days after learning. They preferred our video-augmented dictionary for its benefits in memorization and enjoyable learning experience. Yeshuang Zhu, Yuntao Wang 0001, Chun Yu, Shaoyun Shi, Yankai Zhang, Shuang He, Peijun Zhao, Xiaojuan Ma, Yuanchun Shi |
CHI | 1 |
| 2017 | CEPT: Collaborative Editing Tool for Non-Native AuthorsabstractDue to language deficiencies, individual non-native speakers (NNS) face many difficulties while writing. In this paper, we propose to build a collaborative editing system that aims to facilitate the sharing of language knowledge among non-native co-authors, with the ultimate goal of improving writing quality. We describe CEPT, which allows individual co-authors to generate their own revisions as well as incorporating edits from others to achieve mutual inspiration. The main technical challenge is how to aggregate edits of multiple co-authors and present them in an easy-to-understand way. After iterative design, CEPT highlights three novel features: 1) cross-version sentence mapping for edit tracking, 2) summarization of edits from multiple co-authors, and 3) a collaborative editing interface that enables co-authors to examine, comment on, and borrow edits of others. A preliminary lab study showed that CEPT could significantly improve both the language quality and collaboration experience of NNS writers, due to its efficacy for sharing language knowledge. Yeshuang Zhu, Shichao Yue, Chun Yu, Yuanchun Shi |
CSCW | 1 |