Yujia Wu

dblp:187/9394 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
abstract
Multimodal Large Language Models (MLLMs) with unified architectures excel across a wide range of vision-language tasks, yet aligning them with personalized image generation remains a significant challenge. Existing methods for MLLMs are frequently subject-specific, demanding a data-intensive fine-tuning process for every new subject, which limits their scalability. In this paper, we introduce MM-R1, a framework that integrates a cross-modal Chain-of-Thought (X-CoT) reasoning strategy to unlock the inherent potential of unified MLLMs for personalized image generation. Specifically, we structure personalization as an integrated visual reasoning and generation process: (1) grounding subject concepts by interpreting and understanding user-provided images and contextual cues, and (2) generating personalized images conditioned on both the extracted subject representations and user prompts. To further enhance the reasoning capability, we adopt Grouped Reward Proximal Policy Optimization(GRPO) to explicitly align the generation. Experiments demonstrate that MM-R1 unleashes the personalization capability of unified MLLMs to generate images with high subject fidelity and strong text alignment in a zero-shot manner.
Yujia Wu, Kuncheng Li, Jiwei Wei, Shiyuan He, Jinyu Guo, Ning Xie 0003
AAAI2
2026 Vision-G1: Towards General Reasoning Vision-Language Models via Reinforcement Learning
abstract
Recent vision-language models (VLMs) show strong reasoning capabilities through training with reinforcement learning from verifiable rewards (RLVR). Despite their impressive capabilities, current VLMs focus on a limited range of reasoning tasks, such as mathematical and logical reasoning, due to the lack of readily available verifiable reward data in broader domains. As a result, these models struggle to generalize their reasoning abilities to the wide variety of challenges encountered in real-world environments. To address this limitation, we collect and assemble a comprehensive RL-ready visual reasoning training dataset encompassing 46 datasets across 13 dimensions of 5 domains, covering a wide range of realistic scenarios such as infographic reasoning, mathematical reasoning, spatial reasoning, and general science reasoning. Based on this dataset, we propose an influence function-based data filtering strategy and a multi-round data curriculum method to iteratively strengthen general visual reasoning abilities. Using this approach, we train a general reasoning VLM, namely Vision-G1. Our 7B model achieves state-of-the-art performance across nine visual reasoning benchmarks, surpassing previous similar-sized VLMs and even GPT-4o and Gemini-1.5 Flash.
Yuheng Zha, Kun Zhou 0002, Yujia Wu, Yushu Wang, Shibo Hao, Zhengzhong Liu 0001, Eric P. Xing, Zhiting Hu
AAAI3
2026 LoLDU: Low-Rank Adaptation via Lower-Diag-Upper Decomposition for Parameter-Efficient Fine-Tuning
abstract
The rapid growth of model scale has necessitated substantial computational resources for fine-tuning. Existing approach such as low-rank adaptation (LoRA) has sought to address the problem of handling the large updated parameters in full fine-tuning (FT). However, LoRA utilize random initialization and optimization of low-rank matrices to approximate updated weights, which can result in suboptimal convergence and an accuracy gap compared to full fine-tuning (FT). To address these issues, we propose low-rank LDU (LoLDU), a parameter-efficient fine-tuning (PEFT) approach that significantly reduces trainable parameters by 2600 times compared to regular PEFT methods while maintaining comparable performance. LoLDU leverages lower-diag-upper (LDU) decomposition to initialize low-rank matrices for faster convergence and nonsingularity. We focus on optimizing the diagonal matrix for scaling transformations. To the best of our knowledge, LoLDU has the fewest parameters among all PEFT approaches. We conducted extensive experiments across 4 instruction-following datasets, six natural language understanding (NLU) datasets, eight image classification datasets, and image generation datasets with multiple model types [LLaMA2, RoBERTa, ViT, and stable diffusion (SD)], providing a comprehensive and detailed analysis. Our open-source code can be accessed at https://anonymous.4open.science/r/LoLDU-B5A6.
Yiming Shi, Yujia Wu, Jiwei Wei, Ran Ran 0001, Cheng-Wei Sun, Shiyuan He, Yang Yang 0002
IEEE Trans. Neural Networks Learn. Syst.2
2025 Conditional Prediction ROC Bands for Graph Classification
abstract
Graph classification in medical imaging and drug discovery requires accuracy and robust uncertainty quantification. To address this need, we introduce Conditional Prediction ROC (CP-ROC) bands, offering uncertainty quantification for ROC curves and robustness to distributional shifts in test data. Although developed for Tensorized Graph Neural Networks (TGNNs), CP-ROC is adaptable to general Graph Neural Networks (GNNs) and other machine learning models. We establish statistically guaranteed coverage for CP-ROC under a local exchangeability condition. This addresses uncertainty challenges for ROC curves under non-iid setting, ensuring reliability when test graph distributions differ from training data. Empirically, to establish local exchangeability for TGNNs, we introduce a data-driven approach to construct local calibration sets for graphs. Comprehensive evaluations show that CP-ROC significantly improves prediction reliability across diverse tasks. This method enhances uncertainty quantification efficiency and reliability for ROC curves, proving valuable for real-world applications with non-iid objects.
Yujia Wu, Elynn Y. Chen, Zheshi Zheng
AISTATS1
2025 Multi-Agent Transformer-based Automated Imbalanced Time Series Classification with Hyperparameter Optimization
abstract
Time series classification is a prevalent research topic with wide applications, and Imbalanced Time Series Classification (ITSC) has emerged with high importance due to the nature of imbalanced class distribution in real-world applications. However, most existing ITSC methods rely heavily on manual adjustments that limit their scalability and performance. To address this, we introduce Multi-Agent Transformer-based Automated Imbalanced Time Series Classification (MAT-AITSC), a novel framework that includes an ITSC pipeline and a Multi-Agent Transformer-based Hyperparameter Optimization method named MAT-HPO method named MAT-HPO. Our pipeline uses cost-sensitive learning to address class imbalance while preserving the distribution of time series. MAT-HPO sequentially optimizes model architecture, class weights, and training hyperparameters to enhance classification performance on imbalanced time series datasets. Extensive evaluations on 52 diverse time series datasets demonstrate that MAT-AITSC significantly outperforms existing methods. To the best of our knowledge, this is the first work that offers an automated solution for imbalanced time series classification, reducing the need for manual intervention and paving the way for broader real-world applications.
Nai-Hsin Cheng, Yujia Wu, Vincent S. Tseng
IJCNN2
2025 Tensor-Fused Multi-view Graph Contrastive Learning
Yujia Wu, Junyi Mo, Elynn Y. Chen
PAKDD (7)1
2025 A survey of text classification based on pre-trained language model
Yujia Wu, Jun Wan 0005
Neurocomputing1
2024 Word and Character Semantic Fusion by Pretrained Language Models for Text Classification
abstract
The utilization of fine-tuned pre-trained language models (PLMs) in text classification has become widespread and has achieved remarkable performance. However, a significant limitation of these models is that they tend to mark special symbols that are not registered in PLMs as [UNK], which ultimately reduces the text classification accuracy. Although some PLMs have attempted to overcome this limitation by learning subwords or using character input during model training, effectively fusing subword and character features remains a challenge. To address this challenge, we propose a hybrid network architecture called SemFusion that collaborates two PLMs to learn rich semantic information and improve classification accuracy. Specifically, we use two different Transformer Encoders to extract features for each subword and character in the text sequence. These features include the [CLS] feature, which represents the overall meaning of the text sequence, and the token feature, which represents each subword or character. We then utilize a multi-head self-attention mechanism to model the correlation between text sequences at different granularities for enhancing semantic expression, thereby achieving more accurate text classification. Our experimental results on four benchmark text datasets demonstrate that our proposed method outperforms the current state-of-the-art methods in the literature.
Yujia Wu, Jun Wan 0005
IJCNN1
2024 Improving Zero-Shot Image Captioning Efficiency with Metropolis-Hastings
Dehu Du, Yujia Wu
PRCV (7)2
2024 Improving Text Classification Performance Through Multimodal Representation
Yujia Wu, Hong Ren
PRCV (7)1
2024 Precise facial landmark detection by Dynamic Semantic Aggregation Transformer
Jun Wan 0005, Yujia Wu, Zhihui Lai 0001, Wenwen Min, Jun Liu 0036
Pattern Recognit.3
2023 ParaNet:Parallel Networks with Pre-trained Models for Text Classification
Yujia Wu, Xingli Chen
ADMA (3)1
2023 CharCaps: Character-Level Text Classification Using Capsule Networks
Yujia Wu, Kangning Zhan
ICIC (2)1
2021 A hybrid neural network approach to combine textual information and rating information for item recommendation
Donghua Liu, Jing Li 0055, Bo Du 0001, Rong Gao 0001, Yujia Wu
Knowl. Inf. Syst.6
2020 Text Classification using Triplet Capsule Networks
abstract
Most existing methods only consider the local features of the samples, and their experimental results show better performance than traditional Non-deep learning methods. However, in these methods, the global features of the sample space are usually ignored, and these ignored global features will affect the classification accuracy. To solve this problem, a novel triple capsule network framework is proposed to text classification. The training in the first stage, to obtain a basic capsule network for obtaining local features. Then, three capsule networks sharing parameters are combined spatially, and the triplet loss function is used in the second stage of training. By comparative learning, the capsule network can learn global features that can represent the spatial distance between different categories. Through comparison experiments on six datasets and ten general benchmark algorithms, the results show that our results is the first in the four datasets.
Yujia Wu, Jing Li 0055, Zhiquan Ding
IJCNN1
2020 Siamese capsule networks with global and local features for text classification
Yujia Wu, Jing Li 0055, Jia Wu 0001
Neurocomputing1
2019 Face alignment by Component Adaptive Mechanism
Jun Wan 0005, Jing Li 0055, Yujia Wu, Yafu Xiao, Xuefei Li 0001
Neurocomputing4
2019 Research on Gas Pipeline Multi-Point Leak Signal Processing and Source Locating Using VMD, BSS and Relative Entropy
abstract
The early multi-point leakage source signals of urban gas pipeline are weak and can be easily affected by environmental noise and signal interference between adjacent sources, which causes large leakage positioning error. In this paper, an integrated signal processing method combining VMD, BSS and Relative entropy for multi-point pipeline leakage signal and source positioning is presented. Firstly, VMD and Relative entropy were employed to obtain effective IMF mode components and their features during the decomposition for pipeline leakage signal. Relative entropy was used to improve signal-to-noise ratio and extract the features of leakage signal. Then, BSS was used to decompose multi-point mixed leakage signals so as to obtain independent signal components. Finally, the time difference and wave velocity were, respectively, obtained by calculating the time domain distribution of the independent signal components and the main modal guided wave, so the precise positioning of the pipeline leakage was realized. The results show that the combined method proposed can not only select and extract leakage signal adaptively but also separate single independent signal from multi-point mixed signal, which helps to locate pipeline multi-point leakage sources more accurately.
Yongmei Hao, Ni Qin, Zhixiang Xing, Yujia Wu, Yunfei Yue
Int. J. Pattern Recognit. Artif. Intell.4
2018 Three-Dimensional Reconstruction of Target Self-Calibrating System with Nonlinear Optimization Technique
abstract
In this paper, the three-dimensional (3D) reconstruction of target self-calibrating system for the guidance system of air-to-air missile is researched. The basic ideology of self-calibrating theory is studied in depth and also the advantages and disadvantages of traditional calibration method, which is based on active vision and target self-calibrating method, are listed for comparison. The mathematical model of the perspective camera is established, and on this basis, the camera parameters are figured out combining with LM optimization algorithm. The reconstruction is conducted by the method of stratified calibrating. It is proved that the theory of 3D reconstruction of target self-calibrating system in air to air missile is available according to the experimental results. It puts forward a new research approach for the guidance system of air to air missile to identify the target characteristic information in different azimuths.
Qiangfeng Wang, Yan Cao 0003, Yu Bai 0006, Yujia Wu, Qingyun Wu
Int. J. Pattern Recognit. Artif. Intell.4
2018 On the string matching with k mismatches
Yangjun Chen, Yujia Wu
Theor. Comput. Sci.2
2017 BWT Arrays and Mismatching Trees: A New Way for String Matching with k Mismatches
abstract
In this paper, we discuss an efficient and effective index mechanism to do the string matching with k mismatches, by which we will find all the substrings in a target string s having at most k positions different from a pattern string r. The main idea is to transform s to a BWT-array as index, denoted as BWT(s), and search r against it. During the process, the precomputed mismatch information of r will be utilized to speed up the BWT(s)'s navigation. In this way, the time complexity can be reduced to O(kn' + n + mlogm), where m = |r|, n = |s|, and n' is the number of leaf nodes of a tree structure, called a mismatching tree, produced during a search of BWT(s). Extensive experiments have been conducted, which show that our method for this problem is promising.
Yangjun Chen, Yujia Wu
ICDE2