Dawei Dai

dblp:79/4849 · DBLP profile ↗
← Back
24ranked-venue papers
12as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 FDSRM: A Feature-Driven Style-Agnostic Foundation Model for Sketch-Less Facial Image Retrieval
abstract
Sketch-less facial image retrieval (SLFIR) framework efficiently retrieves target images with minimal strokes through human-computer interaction, thus overcoming the traditional model's reliance on high-quality sketch images. However, the variability in sketching styles and the randomness of stroke placement during the drawing process pose challenges in matching target images. To address this issue, we propose a feature-driven foundation model for sketch-less facial image retrieval (FDSRM), which is designed to be independent of the sketch style and comprises two core components: the feature observer and the adaptive fusion adapter (AFA). First, to address the diversity of sketch styles, we design the feature observer module (FOM). It employs multiple experts focused on extracting key features and semantic information common to various sketch styles and the target image. This helps the model to precisely identify crucial features for effective matching in stylistically diverse sketches. Second, to address the randomness of stroke placement, we introduce prior knowledge of sketching and, in conjunction with the AFA component, dynamically learn and adjust the fusion strategy of sketches and text based on the current state of sketch strokes. This enables more accurate and targeted feature fusion throughout the sketching process. Furthermore, we train a facial image-text alignment pretraining (FAIP) model on a large-scale facial dataset and use it as the backbone of FDSRM, which significantly improved the model's robustness to unknown facial features. Extensive experiments demonstrate that our method exhibits significant advantages in terms of accuracy in early retrieval and system generalization capabilities. Even without additional auxiliary information, it outperforms state-of-the-art methods in both qualitative and quantitative measures in multistyle application scenarios.
Yingge Liu, Dawei Dai, Shuyin Xia, Guoyin Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation
abstract
Facial images have extensive practical applications. Although the current large-scale text-image diffusion models exhibit strong generation capabilities, it is challenging to generate the desired facial images using only text prompt. Image prompts are a logical choice. However, current methods of this type generally focus on general domain. In this paper, we aim to optimize image makeup techniques to generate the desired facial images. Specifically, (1) we built a dataset of 4 million high-quality face image-text pairs based on the FaceCaption-15M and LAION-Face to train our Face-MakeUp model; (2) to maintain consistency with the reference facial image, we extract/learn multi-scale content features and pose features for the facial image, integrating these into the diffusion model to enhance the preservation of facial identity features for diffusion models. Validation on two face-related test datasets demonstrates that our Face-MakeUp can achieve the best comprehensive performance. All codes, data, and model checkpoints are available at: https://github.com/ddw2AIGROUP2CQUPT/Face-MakeUp.
Dawei Dai, Yinxiu Zhou, Hang Xing, Chenghang Li
ECAI1
2025 From Sparse to Complete: Semantic Understanding Based on Stroke Evolution in On-the-fly Sketch-based Image Retrieval
abstract
In contrast with human sketching, which pre-conceptualizes outlines and features, conventional sketch retrieval models rely primarily rely on pixel-level processing and feature extraction, limiting their ability to capture early sketch intent. Consequently, these models are susceptible to subjective stroke noise, reducing retrieval accuracy. To address this issue, we propose a novel on-the-fly noise stroke retrieval framework designed to align with human sketch-drawing cognition. The proposed framework introduces two core innovations. (i) A stroke consistency detection module that effectively discriminates and suppresses noise strokes by quantifying the structural similarity between the current stroke and the target image, as well as its alignment with key skeletal components. (ii) An adaptive gated mixture of experts module that dynamically selects and integrates features from multiple expert networks during the early, sparse stages of sketching, thereby capturing relevant information with greater precision. Experimental results across diverse sketch datasets demonstrate that the proposed method effectively identifies and suppresses early noise strokes, significantly enhances sketch retrieval performance, and exhibits strong robustness across varying sketch styles.
Yingge Liu, Dawei Dai, Xiangling Hou, Shilin Zhao, Guoyin Wang 0001
IJCAI2
2025 Multi-granularity representation learning for sketch-based dynamic face image retrieval
Dawei Dai, Shiyu Fu
Appl. Intell.2
2025 Multivariate Feedback-Based Image-Text Joint Learning for Sketch-Less Facial Image Retrieval
abstract
Sketch-Less Facial Image Retrieval (SLFIR) framework facilitates the retrieval of target images with minimal strokes through a human-computer interactive approach, thereby circumventing the need for high-quality sketches required by traditional frameworks. The primary approach utilizes a contrastive learning framework that minimizes the distance between sketch images and their target images in the embedding space, while maximizing the distance from non-target images, thus efficiently learning representations of sketches and images. However, during the initial stages of sketching, the sparse strokes that capture only partial facial features can inadvertently match non-target facial images, blurring the distinctions between positive and negative samples and impairing early retrieval performance. To overcome this challenge, we introduce a multimodal retrieval model based on diversified feedback reinforcement learning, which not only enhances the semantic integrity of sketches but also optimally ranks the sketches corresponding to positive samples using diversified feedback. Specifically, (1) we developed a Facial Language-Image Pre-training (FLIP) model and, leveraging this model, constructed an on-the-fly multimodal retrieval model that excels in recognizing sparse and exaggerated sketches by extracting and integrating multiscale features from both sketches and textual descriptions. (2) Furthermore, we implemented a novel reward mechanism that adjusts the rewards for target images, accommodating reasonable fluctuations in sketch rankings on actual images. This mechanism effectively differentiates similar images during retrieval, ensuring a more consistent and progressively improving ranking list. Extensive experiments validate that our proposed method significantly enhances early retrieval accuracy and generalization capability.
Yingge Liu, Dawei Dai, Guoyin Wang 0001, Shuyin Xia
IEEE Trans. Circuits Syst. Video Technol.2
2025 An Adaptive Multi-Granularity Graph Representation of Image via Granular-ball Computing
abstract
Graph neural networks (GNNs) encounter challenges in establishing deep structures and managing a large number of parameters effectively to learn node features comprehensively. Consequently, in vision tasks, GNNs often struggle to achieve high classification accuracy compared to convolutional neural networks. Nonetheless, GNNs retain crucial advantages and potential, particularly in lightweight network scale and efficient, reliable decision-making. Thus, improving GNN performance in vision tasks remains a significant research endeavor, with numerous important works exploring the application of GNN models in such contexts, where the graph representation of images poses a key challenge. Existing methods often fall short in adaptively generating blocks of different sizes and their corresponding edges to form graph representations according to graph semantics. To address this issue, we propose a novel method to convert images into graphical forms using granular-ball computing. Our approach does not rely on manual annotation or other learning methods, yet it can dynamically generate block nodes of varying sizes and corresponding edges. Compared to other state-of-the-art methods, our approach better captures semantic information within the graph. Despite having fewer parameters, our method significantly enhances accuracy. Overall, our work holds substantial implications for improving the performance of graph neural networks in vision tasks.
Dawei Dai, Fan Chen 0010, Shuyin Xia, Guoyin Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2024 PA-LLaVA: A Large Language-Vision Assistant for Human Pathology Image Understanding
abstract
The previous advancements in pathology image understanding primarily involved developing models tailored to specific tasks. Recent studies has demonstrated that the large vision-language model can enhance the performance of various downstream tasks in medical image understanding. In this study, we developed a domain-specific large language-vision assistant (PA-LLaVA) for pathology image understanding. Specifically, (1) we first construct a human pathology image-text dataset by cleaning the public medical image-text data for domain-specific alignment; (2) Using the proposed image-text data, we first train a pathology language-image pretraining (PLIP) model as the specialized visual encoder for pathology image, and then we developed scale-invariant connector to avoid the information loss caused by image scaling; (3) We adopt two-stage learning to train PA-LLaVA, first stage for domain alignment, and second stage for end to end visual question & answering (VQA) task. In experiments, we evaluate our PA-LLaVA on both supervised and zero-shot VQA datasets, our model achieved the best overall performance among multimodal models of similar scale. The ablation experiments also confirmed the effectiveness of our design. We posit that our PA-LLaVA model and the datasets presented in this work can promote research in field of computational pathology. All codes are available at: https://github.com/ddw2AIGROUP2CQUPT/PA-LLaVA
Dawei Dai, Qianlan Yang, Xiaojing Shen, Shuyin Xia, Guoyin Wang 0001
BIBM1
2024 GraphConvNet: A Dual Network Utilizing Local Features Coupled with Structural Information for Predicting Knee Osteoarthritis
abstract
Knee Osteoarthritis (KOA) is a common joint disease that severely affects the normal lives of patients. In clinical practice, the severity of KOA is commonly evaluated by observing radiographs of the knee joint However, this approach heavily relies on a doctor’s clinical experience and exhibits a certain degree of subjectivity. In previous studies, various advanced deep convolutional neural network (CNN) models have been used to diagnose KOA. As known, CNN models often focus on learning the local detailed features for decision-making and lack attention to global structural information. In this study, we propose a dual network called GraphConvNet that integrates a visual graph neural network with a deep CNN to enhance representation learning by leveraging both local detailed features and global structural information. Our proposed method was evaluated using an Osteoarthritis Initiative (OAI) dataset and achieved an overall accuracy, recall, precision, and mean absolute error (MAE) of 75.24%, 77.75%, 74.33% and 0.283, respectively. Experiments demonstrate that our proposed method significantly improves the performance and achieves state-of-the-art performance. All codes are available at https://github.com/ddw2AIGROUP2CQUPT/GraphConvNet.
Pengju Tang, Dawei Dai, Xionghui Yang, Guoyin Wang 0001
BIBM2
2024 Multimodal Image-Text Representation Learning for Sketch-Less Facial Image Retrieval
abstract
Sketch-less facial image retrieval (SLFIR) framework aims to break the barriers that drawing a high-quality facial sketch requires excellent skills and substantial time, it performs the retrieval using a partial sketch with as few strokes as possible. However, such early-stage sketches often contain only local details, resulting in poor retrieval performance. In this study, we propose learning of the representation by fusing the sketches with prior human semantic knowledge to improve the early retrieval performance. Specifically, (1) based on the LAION-Face dataset, a facial language-image pretraining (FLIP) model is constructed to learn the aligned representations of facial image and text; (2) subsequently, using FLIP as the backbone, multiscale features of sketch and text are extracted and fused to learn the efficient representation for the final retrieval. The proposed method achieves state-of-the-art early retrieval performance on all two public datasets and exhibits a good generalization ability in practical testing.
Dawei Dai, Yingge Liu, Shiyu Fu, Guoyin Wang 0001
ICME1
2024 Granular-Ball Representation Learning for Deep CNN on Learning with Label Noise
Dawei Dai, Shuyin Xia, Guoyin Wang 0001
ICONIP (4)1
2024 Prior semantic-embedding representation learning for on-the-fly FG-SBIR
Yingge Liu, Dawei Dai, Kenan Zou, Xiufang Tan, Yiqiao Wu, Guoyin Wang 0001
Expert Syst. Appl.2
2024 LGRL: Local-Global Representation Learning for On-the-Fly FG-SBIR
abstract
On-the-fly Fine-grained sketch-based image retrieval (On-the-fly FG-SBIR) framework aim to break the barriers that sketch drawing requires excellent skills and is time-consuming. Considering such problems, a partial sketch with fewer strokes contains only the little local information, and the drawing process may show great difference among users, resulting in poor performance at the early retrieval. In this study, we developed a local-global representation learning (LGRL) method, in which we learn the representations for both the local and global regions of the partial sketch and its target photos. Specifically, we first designed a triplet network to learn the joint embedding space shared between the local and global regions of the entire sketch and its corresponding region of the photo. Then, we divided each partial sketch in the sketch-drawing episode into several local regions; Another learnable module following the triplet network was designed to learn the representations for the local regions of the partial sketch. Finally, by combining both the local and global regions of the sketches and photos, the final distance was determined. In the experiments, our method outperformed state-of-the-art baseline methods in terms of early retrieval efficiency on two publicly sketch-retrieval datasets and the practice test. Codes are available on:https://github.com/ddw2AIGROUP2CQUPT/LGRL.
Dawei Dai, Yingge Liu, Yutang Li, Shiyu Fu, Shuyin Xia, Guoyin Wang 0001
IEEE Trans. Big Data1
2023 Sketch Less Face Image Retrieval: A New Challenge
abstract
In some specific scenarios, face sketch was used to identify a person. However, drawing a complete face sketch often needs skills and takes time, which hinder its widespread applicability in the practice. In this study, we proposed a new task named sketch less face image retrieval (SLFIR), in which the retrieval was carried out at each stroke and aim to retrieve the target face photo using a partial sketch with as few strokes as possible (see Fig. 1). Firstly, we developed a method to generate the data of sketch with drawing process, and opened such dataset; Secondly, we proposed a two-stage method as the baseline for SLFIR that (1) a triplet network, was first adopt to learn the joint embedding space shared between the complete sketch and its target face photo; (2) regarding the sketch drawing episode as a sequence, we designed a LSTM module to optimize the representation of the incomplete face sketch. Experiments indicate that the new framework can finish the retrieval using a partial or poor drawing sketch. (https://github.com/ddw2AIGROUP2CQUPT/SLFIR)
Dawei Dai, Yutang Li, Shiyu Fu, Shuyin Xia, Guoyin Wang 0001
ICASSP1
2023 Ensemble learning framework for image retrieval via deep hash ranking
Donggen Li, Dawei Dai, Jiancu Chen, Shuyin Xia, Guoyin Wang 0001
Knowl. Based Syst.2
2022 Ensemble Ranking for Image Retrieval via Deep Hash
Donggen Li, Dawei Dai, Hongyuan Shan, Shunyin Xia, Yulong Xia
ICANN (4)2
2022 Double embedding and bidirectional sentiment dependence detector for aspect sentiment triplet extraction
Dawei Dai, Shuyin Xia, Guoyin Wang 0001, Zizhong Chen
Knowl. Based Syst.1
2022 Multi-granularity Association Learning for On-the-fly Fine-grained Sketch-based Image Retrieval
Dawei Dai, Yingge Liu, Shuyin Xia, Guoyin Wang 0001
Knowl. Based Syst.1
2021 Designing a Partially Understandable Neural Network through Semantic Embedding
abstract
In recent years, the convolution neural network (CNN) has been successfully applied in numerous fields as a machine learning model. However, such neural models are still considered to be “black box” for most tasks. The fundamental issue underlying this problem is that the information and knowledge learned by a neural network during the training process are unknown and unpredictable. In this study, we attempted to design a partially understandable neural network through semantic embedding. Firstly, we selected several understandable feature extraction operators as expected information. Secondly, we embedded these operators into the hierarchical layers of a neural network at the beginning of its training process. Finally, these embedded operators were only involved in forward calculation and remained unchanged in the error back-propagation of the training process. We applied our method to the ResNet and DenseNet models in image classification tasks. In the experiments, our new models achieved almost the same performance as the original ones, but decreased performance was exhibited when shutting the embedded parts, and the features in the embedded parts were understandable. The experiments verified that the embedded parts not only contributed to the classification tasks, but also caused the neural network model to be partially understandable.
Chengfu Tang, Dawei Dai, Haoyue Bai 0001, Guoyin Wang 0001, Shuyin Xia, Feng Hu 0001
IJCNN2
2021 Building partially understandable convolutional neural networks by differentiating class-related neural nodes
Dawei Dai, Chengfu Tang, Guoyin Wang 0001, Shuyin Xia
Neurocomputing1
2020 Parameters Sharing in Residual Neural Networks
Dawei Dai, Hui Wei 0001
Neural Process. Lett.1
2018 A Bio-Feasible Computational Circuit for Neural Activities Persisting and Decaying
Dawei Dai, Hui Wei 0001, Su Zihao
ICANN (2)1
2018 Balanced Cortical Microcircuitry-Based Network for Working Memory
Hui Wei 0001, Su Zihao, Dawei Dai
ICANN (1)3
2017 A Plausible Micro Neural Circuit for Decision-Making
Hui Wei 0001, Dawei Dai, Yijie Bu
CogSci2
1986 The Weak Generative Capacity of Parenthesis-Free Categorial Grammars
Joyce Friedman, Dawei Dai, Weiguo Wang
COLING2