Yudong Li 0001

dblp:212/6739-1 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-6779-8836ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 OncoCoT: A Temporal-causal Chain-of-Thought Dataset for Oncologic Decision-Making
abstract
Long Chain-of-Thought (CoT) reasoning has shown great promise in complex reasoning tasks, but its application to medical decision-making presents unique challenges. Unlike structured tasks relying on static verification frameworks, medical decision-making requires dynamic validation through longitudinal clinical outcomes, exhibiting temporal-causal dependencies that complicate the verification of reasoning processes. Therefore, we introduce a novel data construction framework specifically designed for medical decision-making. First, the framework analyzes real-world clinical cases to construct a timeline of medical events and identify critical decision points, including examination, diagnosis, and treatment. Subsequently, it employs a clinical causality-aware strategy to generate decision-making questions at the identified points, along with reasoning traces and corresponding answers. Finally, information drawn from future nodes serves as clinical logic-constrained criteria to re-evaluate and refine the soundness of the generated reasoning and responses. Building on this, we present OncoCoT, an oncologic decision-making dataset derived from clinical records over the past four years across eight common cancer types. Furthermore, we distill a subset of OncoCoT into a dedicated benchmark, OncoEval, to facilitate systematic evaluation of clinical reasoning capabilities in LLMs. Evaluation results show that existing state-of-the-art reasoning models, such as Deepseek-r1 and GPT-o3, exhibit limited capability in addressing clinical problems in OncoEval, highlighting the need for further improvement.
Peiru Yang, Yudong Li 0001, Shiting Wang, Haotian Gan, Xintian Li, Qingyu Gao, Yongfeng Huang 0001
AAAI2
2026 KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates
abstract
Standard Large Language Model (LLM) pretraining typically treats corpora as flattened token sequences, often overlooking the realworld context that humans naturally rely on to contextualize information.To bridge this gap, we introduce Knowledge Coordinate Conditioning (KoCo), a simple method that maps every document into a three-dimensional semantic coordinate.By prepending these coordinates as textual prefixes for pre-training, we aim to equip the model with explicit contextual awareness to learn the documents within the real-world knowledge structure.Experiment results demonstrate that KoCo significantly enhances performance across 10 downstream tasks and accelerates pre-training convergence by approximately 30%.Furthermore, our analysis indicates that explicitly modeling knowledge coordinates helps the model distinguish stable facts from noise, effectively mitigating hallucination in generated outputs.# Role: Knowledge Taxonomist & Data Curator # Core Task: Your task is to analyze the provided [TEXT] and its source [URL], and generate a single, concise "Knowledge Context Meta-Tag."This tag serves solely to inform the training model of the "position", "sourcle", and "attribute" of the text it will read within the broader human knowledge system.Note that your goal is to **classify text**, not **summarize the content of the text**.Do not expand or modify existing information.-# Step Guide: 1. **Analysis [URL]:** * Check the domain name (e.g
Yudong Li 0001, Jiawei Cai, LinLin Shen
ACL (1)1
2025 FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs
abstract
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely unexplored. To address this gap, we introduce FaceBench, a dataset featuring hierarchical multi-view and multi-level attributes specifically designed to assess the comprehensive face perception abilities of MLLMs. Initially, we construct a hierarchical facial attribute structure, which encompasses five views with up to three levels of attributes, totaling over 210 attributes and 700 attribute values. Based on the structure, the proposed FaceBench consists of 49,919 visual questionanswering (VQA) pairs for evaluation and 23,841 pairs for fine-tuning. Moreover, we further develop a robust face perception MLLM baseline, Face-LLaVA, by training with our proposed face VQA data. Extensive experiments on various mainstream MLLMs and Face-LLaVA are conducted to test their face perception ability, with results also compared against human performance. The results reveal that, the existing MLLMs are far from satisfactory in understanding the fine-grained facial attributes, while our Face-LLaVA significantly outperforms existing open-source models with a small amount of training data and is comparable to commercial ones like GPT-4o and Gemini. The dataset will be released at https://github.com/CVI-SZU/FaceBench
Xusen Ma, Xianxu Hou, Meidan Ding, Yudong Li 0001, Junliang Chen 0002, Wenting Chen, Xiaoyang Peng, LinLin Shen
CVPR5
2025 Aligned or Apart? Multi-Agent Insights into Consumer and Brand Messaging Discrepancies
abstract
In the digital age, brand meaning is increasingly shaped through user participation and content sharing on social media platforms. However, significant perceptual gaps often exist between official brand narratives and consumer interpretations. These multimodal and cognitively nuanced gaps are challenging to detect and model using traditional analytical methods. To address this, we propose a multi-agent framework that metaphorically models perception as an optical process-propagation, interference, and measurement---termed OPIM. We construct a novel dual-perspective dataset from representative social media platforms, integrating text and image content from both user-generated and official brand communications. We evaluate brand perception along six psychological dimensions. Experiments across 15 brands demonstrate that our framework effectively captures key perception gaps, particularly in sincerity, professionalism, and attractiveness. In contrast, materialism and sophistication exhibit higher alignment between brand messaging and consumer perception. Our framework enhances the cognitive alignment and multimodal interpretability of large language models, offering actionable insights for brand strategy and bridging computational modeling with human-centric understanding. The dataset will be available at https://github.com/htgan-ai/OPIM.
Haotian Gan, Yudong Li 0001, Weidong Tang
ACM Multimedia2
2024 Dynamic Data Sampler for Cross-Language Transfer Learning in Large Language Models
abstract
Large Language Models (LLMs) have gained significant attention in the field of natural language processing (NLP) due to their wide range of applications. However, training LLMs for languages other than English poses significant challenges, due to the difficulty in acquiring large-scale corpus and the requisite computing resources. In this paper, we propose ChatFlow, a cross-language transfer-based LLM, to address these challenges and train large Chinese language models in a cost-effective manner. We employ a mix of Chinese, English, and parallel corpus to continuously train the LLaMA2 model, aiming to align cross-language representations and facilitate the knowledge transfer specifically to the Chinese language model. In addition, we use a dynamic data sampler to progressively transition the model from unsupervised pre-training to supervised fine-tuning. Experimental results demonstrate that our approach accelerates model convergence and achieves superior performance. We evaluate ChatFlow on popular Chinese and English benchmarks, the results indicate that it outperforms other Chinese models post-trained on LLaMA-2-7B.
Yudong Li 0001, Zhe Zhao 0006, LinLin Shen, Cheng Hou, Xianxu Hou
ICASSP1
2024 FLIP-80M: 80 Million Visual-Linguistic Pairs for Facial Language-Image Pre-Training
abstract
While significant progress has been made in multi-modal learning driven by large-scale image-text datasets, there is still a noticeable gap in the availability of such datasets within the facial domain. To facilitate and advance the field of facial representation learning, we present FLIP-80M, a large-scale visual-linguistic dataset comprising over 80 million face images paired with text descriptions. FLIP-80M is constructed by leveraging the large openly available image-text-pair dataset LAION-5B and a mixed-method approach to filter face-related pairs from both visual and linguistic perspectives. Our curation process involves face detection, face caption classification, text de-noising, and synthesis-based image augmentation. As a result, FLIP-80M stands as the largest face-text dataset to date. To evaluate the potential of our dataset, we fine-tune the CLIP model using the proposed FLIP-80M, to create FLIP (Facial Language-Image Pretraining) and assess its representation capabilities across various downstream tasks. Our experiments demonstrate that our FLIP model achieves state-of-the-art results in a range of face analysis tasks, including face parsing, face alignment, and face attribute classification. The dataset and models are available at https://github.com/ydli-ai/FLIP.
Yudong Li 0001, Xianxu Hou, Dezhi Zheng, LinLin Shen, Zhe Zhao 0006
ACM Multimedia1
2023 Learning Visual Prior via Generative Pre-Training
abstract
Various stuff and things in visual data possess specific traits, which can be learned by deep neural networks and are implicitly represented as the visual prior, e.g., object location and shape, in the model. Such prior potentially impacts many vision tasks. For example, in conditional image synthesis, spatial conditions failing to adhere to the prior can result in visually inaccurate synthetic results. This work aims to explicitly learn the visual prior and enable the customization of sampling. Inspired by advances in language modeling, we propose to learn Visual prior via Generative Pre-Training, dubbed VisorGPT. By discretizing visual locations, e.g., bounding boxes, human pose, and instance masks, into sequences, VisorGPT can model visual prior through likelihood maximization. Besides, prompt engineering is investigated to unify various visual locations and enable customized sampling of sequential outputs from the learned prior. Experimental results demonstrate the effectiveness of VisorGPT in modeling visual prior and extrapolating to novel scenes, potentially motivating that discrete visual locations can be integrated into the learning paradigm of current language models to further perceive visual world. Code is available at https://sierkinhane.github.io/visor-gpt.
Jinheng Xie, Kai Ye 0004, Yudong Li 0001, Yuexiang Li, Qinghong Lin, Yefeng Zheng 0001, LinLin Shen, Zheng Shou 0001
NeurIPS3
2023 TextFace: Text-to-Style Mapping Based Face Generation and Manipulation
abstract
As a subtopic of text-to-image synthesis, text-to-face generation has great potential in face-related applications. In this paper, we propose a generic text-to-face framework, namely, TextFace, to achieve diverse and high-quality face image generation from text descriptions. We introduce text-to-style mapping, a novel method where the text description can be directly encoded into the latent space of a pretrained StyleGAN. Guided by our text-image similarity matching and face captioning-based text alignment, the textual latent code can be fed into the generator of a well-trained StyleGAN to produce diverse face images with high resolution (1024×1024). Furthermore, our model inherently supports semantic face editing using text descriptions. Finally, experimental results quantitatively and qualitatively demonstrate the superior performance of our model.
Xianxu Hou, Yudong Li 0001, LinLin Shen
IEEE Trans. Multim.3
2022 CSL: A Large-scale Chinese Scientific Literature Dataset
abstract
Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development of Chinese scientific NLP. In this work, we present CSL, a large-scale Chinese Scientific Literature dataset, which contains the titles, abstracts, keywords and academic fields of 396k papers. To our knowledge, CSL is the first scientific document dataset in Chinese. The CSL can serve as a Chinese corpus. Also, this semi-structured data is a natural annotation that can constitute many supervised NLP tasks. Based on CSL, we present a benchmark to evaluate the performance of models across scientific domain tasks, i.e., summarization, keyword generation and text classification. We analyze the behavior of existing text-to-text models on the evaluation tasks and reveal the challenges for Chinese scientific NLP tasks, which provides a valuable reference for future research. Data and code will be publicly available.
Yudong Li 0001, Zhe Zhao 0006, LinLin Shen, Weijie Liu 0002, Weiquan Mao
COLING1
2022 Talk2Face: A Unified Sequence-based Framework for Diverse Face Generation and Analysis Tasks
abstract
Facial analysis is an important domain in computer vision and has received extensive research attention. For numerous downstream tasks with different input/output formats and modalities, existing methods usually design task-specific architectures and train them using face datasets collected in the particular task domain. In this work, we proposed a single model, Talk2Face, to simultaneously tackle a large number of face generation and analysis tasks, e.g. text guided face synthesis, face captioning and age estimation. Specifically, we cast different tasks into a sequence-to-sequence format with the same architecture, parameters and objectives. While text and facial images are tokenized to sequences, the annotation labels of faces for different tasks are also converted to natural languages for unified representation. We collect a set of 2.3M face-text pairs from available datasets across different tasks, to train the proposed model. Uniform templates are then designed to enable the model to perform different downstream tasks, according to the task context and target. Experiments on different tasks show that our model achieves better face generation and caption performances than SOTA approaches. On age estimation and multi-attribute classification, our model reaches competitive performance with those models specially designed and trained for these particular tasks. In practice, our model is much easier to be deployed to different facial analysis related tasks. Code and dataset will be available at https://github.com/ydli-ai/Talk2Face.
Yudong Li 0001, Xianxu Hou, Zhe Zhao 0006, LinLin Shen, Xuefeng Yang, Kimmo Yan
ACM Multimedia1
2020 CLUE: A Chinese Language Understanding Evaluation Benchmark
abstract
Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, Zhenzhong Lan. Proceedings of the 28th International Conference on Computational Linguistics. 2020.
Liang Xu 0011, Hai Hu 0001, Xuanwei Zhang, Chenjie Cao, Yudong Li 0001, Yechen Xu, Kai Sun 0006, Dian Yu 0001, Cong Yu 0010, Yin Tian, Qianqian Dong, Weitang Liu, Yiming Cui 0001, Rongzhao Wang, Weijian Xie, Yina Patterson, Zuoyu Tian, Shaoweihua Liu, Zhe Zhao 0006, Qipeng Zhao, Cong Yue, Zhengliang Yang, Kyle Richardson 0001, Zhen-Zhong Lan
COLING6