Yujian Yuan

dblp:327/2729 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Persistent Autoregressive Mapping with Traffic Rules for Autonomous Driving
abstract
Safe autonomous driving requires both accurate HD map construction and persistent awareness of traffic rules, even when their associated signs are no longer visible. However, existing methods either focus solely on geometric elements or treat rules as temporary classifications, failing to capture their persistent effectiveness across extended driving sequences. In this paper, we present PAMR (Persistent Autoregressive Mapping with Traffic Rules), a novel framework that performs autoregressive co-construction of lane vectors and traffic rules from visual observations. Our approach introduces two key mechanisms: Map-Rule Co-Construction for processing driving scenes in temporal segments, and Map-Rule Cache for maintaining rule consistency across these segments. To properly evaluate continuous and consistent map generation, we develop MapDRv2, featuring improved lane geometry annotations. Extensive experiments demonstrate that PAMR achieves superior performance in joint vector-rule mapping tasks, while maintaining persistent rule effectiveness throughout extended driving sequences.
Shiyi Liang, Xinyuan Chang, Changjie Wu, Huiyuan Yan, Yifan Bai 0001, Yujian Yuan, Shuang Zeng, Mu Xu, Xing Wei 0001
AAAI8
2026 UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data
abstract
Large-scale map construction is foundational for critical applications such as autonomous driving and navigation systems. Traditional large-scale map construction approaches mainly rely on costly and inefficient special data collection vehicles and labor-intensive annotation processes. While existing satellite-based methods have demonstrated promising potential in enhancing the efficiency and coverage of map construction, they exhibit two major limitations: (1) inherent drawbacks of satellite data (e.g., occlusions, outdatedness) and (2) inefficient vectorization from perception-based methods, resulting in discontinuous and rough roads that require extensive post-processing. This paper presents a novel generative framework, UniMapGen, for large-scale map construction, offering three key innovations: (1) representing lane lines as discrete sequence and establishing an iterative strategy to generate more complete and smooth map vectors than traditional perception-based methods. (2) proposing a flexible architecture that supports multi-modal inputs, enabling dynamic selection among BEV, PV, and text prompt, to overcome the drawbacks of satellite data. (3) developing a state update strategy for global continuity and consistency of the constructed large-scale map. UniMapGen achieves state-of-the-art performance on the OpenSatMap dataset. Furthermore, UniMapGen can infer occluded roads and predict roads missing from dataset annotations.
Yujian Yuan, Changjie Wu, Xinyuan Chang, Sijin Wang, Shiyi Liang, Shuang Zeng, Mu Xu
AAAI1
2026 PriorDrive: Enhancing Online HD Mapping with Unified Vector Priors
abstract
High-Definition Maps (HD maps) are essential for the precise navigation and decision-making of autonomous vehicles, yet their creation and upkeep present significant cost and timeliness challenges. The online construction of HD maps using on-board sensors has emerged as a promising solution; however, these methods can be impeded by incomplete data due to occlusions and inclement weather, while their performance in distant regions remains unsatisfying. This paper proposes PriorDrive to address these limitations by directly harnessing the power of various vectorized prior maps, significantly enhancing the robustness and accuracy of online HD map construction. Our approach integrates a variety of prior maps uniformly, such as OpenStreetMap's Standard Definition Maps (SD maps), outdated HD maps from vendors, and locally constructed maps from historical vehicle data. To effectively integrate such prior information into online mapping models, we introduce a Hybrid Prior Representation (HPQuery) that standardizes the representation of diverse map elements. We further propose a Unified Vector Encoder (UVE), which employs fused prior embedding and a dual encoding mechanism to encode vector data. To improve the UVE's generalizability and performance, we propose a segment-level and point-level pre-training strategy that enables the UVE to learn the prior distribution of vector data. Through extensive testing on the nuScenes, Argoverse 2 and OpenLane-V2, we demonstrate that PriorDrive is highly compatible with various online mapping models and substantially improves map prediction capabilities. The integration of prior maps through PriorDrive offers a robust solution to the challenges of single-perception data, paving the way for more reliable autonomous driving.
Shuang Zeng, Xinyuan Chang, Yujian Yuan, Shiyi Liang, Mu Xu, Xing Wei 0001
AAAI4
2026 Domain adaptive object detection via CLIP-space guidance and LoRA fine-tuning
Enze Qi, Kan Chang, Qingzhi Zhang, Xueyu Zhang, Yehua Ling, Yujian Yuan, Zan Gao 0001
Expert Syst. Appl.7
2026 UECNet: A unified framework for exposure correction utilizing region-level prompts
Shucheng Xia, Kan Chang, Xuxin Tai, Yehua Ling, Yujian Yuan, Zan Gao 0001
Knowl. Based Syst.7
2025 Prompt-Based Two-Stage Enhancement for Low-Light Object Detection
abstract
Detecting objects in low-light conditions is challenging since the primary features of objects may be buried in dark regions. To tackle this issue, we propose a prompt-based two-stage enhancement (PTE) method for low-light object detection. Our approach utilizes a prompt generator to produce prompts from a support set established based on the input low-light image. These prompts encode the specific degradation information of the input, and thus they are further utilized to guide both image-level and feature-level enhancements for object detection. This allows our PTE to adapt to various lighting conditions. With these prompts, the image-level enhancement adjusts the low-light image using an estimated reciprocal curve, while the feature-level enhancement adaptively modulates the multi-scale features extracted by the detection backbone. Experimental results demonstrate that our method outperforms existing methods in terms of detection accuracy, while maintaining a low number of parameters and high running speed. Our source code and pre-trained model will be available at https://github.com/xbh301/PTE-YOLO.
Bohan Xiong, Kan Chang, Shilin Huang, Shucheng Xia, Yujian Yuan
ICME6
2025 Exp-VQA: Fine-grained facial expression analysis via visual question answering
Yujian Yuan, Jiabei Zeng, Shiguang Shan
Pattern Recognit.1
2025 Multi-View Facial Expressions Analysis of Autistic Children in Social Play
abstract
Atypical facial expressions during interaction are among the early symptoms of autism spectrum disorder (ASD) and are included in standard diagnostic assessments. However, current methods rely on subjective human judgments, introducing bias and limiting objectivity. This paper proposes an automated framework for objective and quantitative assessment of autistic children's facial expressions during social play. Initially, we utilize four synchronized cameras to record interactions between ASD children and teachers during structured activities dominated by the teacher. To address challenges posed by head movements and occluded faces, we introduce a multi-view facial expression recognition strategy. Its effectiveness is demonstrated by experiments in real-world applications. To quantify the patterns of affect status and the dynamic complexity of facial expressions, we use the temporally accumulated distribution of the basic facial expressions and the multi-dimensional multiscale entropy of the facial expression sequence. Analysis of these features revealed significant differences between ASD and TD groups. Experimental results, derived from our quantified features, confirm conclusions drawn from previous research and experiential observations. With these facial expression features, ASD and typically developing (TD) children are accurately classified (accuracy 92.1%, precision 94.4%, 89.5% sensitivity, 94.7% specificity) in empirical experiments, suggesting the potential of our framework for improved ASD assessment.
Jiabei Zeng, Yujian Yuan, Lu Qu, Fei Chang, Xuran Sun, Jinqiuyu Gong, Xuling Han, Qiaoyun Liu, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Affect. Comput.2
2025 Benchmarking Radiology Report Generation From Noisy Free-Texts
abstract
Automatic radiology report generation can enhance diagnostic efficiency and accuracy. However, clean open-source imaging scan-report pairs are limited in scale and variety. Moreover, the vast amount of radiological texts available online is often too noisy to be directly employed. To address this challenge, we introduce a novel task called Noisy Report Refinement (NRR), which generates radiology reports from noisy free-texts. To achieve this, we propose a report refinement pipeline that leverages large language models (LLMs) enhanced with guided self-critique and report selection strategies. To address the inability of existing radiology report generation metrics in measuring cleanliness, radiological usefulness, and factual correctness across various modalities of reports in NRR task, we introduce a new benchmark, NRRBench, for NRR evaluation. This benchmark includes two online-sourced datasets and four clinically explainable LLM-based metrics: two metrics evaluate the matching rate of radiology entities and modality-specific template attributes respectively, one metric assesses report cleanliness, and a combined metric evaluates overall NRR performance. Experiments demonstrate that guided self-critique and report selection strategies significantly improve the quality of refined reports. Additionally, our proposed metrics show a much higher correlation with noisy rate and error count of reports than radiology report generation metrics in evaluating NRR.
Yujian Yuan, Yanting Zheng, Liangqiong Qu
IEEE J. Biomed. Health Informatics1
2024 Fine-grained image classification based on TinyVit object location and graph convolution network
Shijie Zheng, Gaocai Wang, Yujian Yuan, Shuqiang Huang
J. Vis. Commun. Image Represent.3
2023 Describe Your Facial Expressions by Linking Image Encoders and Large Language Models
Yujian Yuan, Jiabei Zeng, Shiguang Shan
BMVC1
2023 Constraint-Based Adversarial Networks for Unsupervised Abstract Text Summarization
abstract
Abstract text summarization is a classic sequence-to-sequence natural language generation task. In order to improve the quality of unsupervised abstract text summarization in unsupervised mode, we propose two constraints for training text summarization model, embedding space constraint and information ratio constraint. We construct a generative adversarial network with two discriminators based on these two constraints (TC-SUM-GAN). We use unsupervised and supervised methods to train the model in the experiment. Experimental results show that the ROUGE-1 value of the unsupervised TC-SUM-GAN increases by [Formula: see text] points compared with the basic model and at least 1.96 points compared with other comparative models. The ROUGE scores of the supervised TC-SUM-GAN are also improved. TC-SUM-GAN achieves very competitive results for the metrics of ROUGE-1 and ROUGE-2. In addition, the abstracts generated by our model are closer to those generated manually.
Liwei Jing, Yujian Yuan, Zuqiang Meng, Yifeng Tan, Patrick Shen-Pei Wang, Xichun Li
Int. J. Pattern Recognit. Artif. Intell.3
2022 FFP: joint Fast Fourier transform and fractal dimension in amino acid property-aware phylogenetic analysis
abstract
BACKGROUND: Amino acid property-aware phylogenetic analysis (APPA) refers to the phylogenetic analysis method based on amino acid property encoding, which is used for understanding and inferring evolutionary relationships between species from the molecular perspective. Fast Fourier transform (FFT) and Higuchi's fractal dimension (HFD) have excellent performance in describing sequences' structural and complexity information for APPA. However, with the exponential growth of protein sequence data, it is very important to develop a reliable APPA method for protein sequence analysis. RESULTS: Consequently, we propose a new method named FFP, it joints FFT and HFD. Firstly, FFP is used to encode protein sequences on the basis of the important physicochemical properties of amino acids, the dissociation constant, which determines acidity and basicity of protein molecules. Secondly, FFT and HFD are used to generate the feature vectors of encoded sequences, whereafter, the distance matrix is calculated from the cosine function, which describes the degree of similarity between species. The smaller the distance between them, the more similar they are. Finally, the phylogenetic tree is constructed. When FFP is tested for phylogenetic analysis on four groups of protein sequences, the results are obviously better than other comparisons, with the highest accuracy up to more than 97%. CONCLUSION: FFP has higher accuracy in APPA and multi-sequence alignment. It also can measure the protein sequence similarity effectively. And it is hoped to play a role in APPA's related research.
Yujian Yuan, Xichun Li, Zuqiang Meng
BMC Bioinform.4