Dakui Wang

dblp:142/0190 · DBLP profile ↗
← Back
28ranked-venue papers
0as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Security and privacy · 5 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 ReTD: Reconstruction-Based Traceability Detection for Generated Images
abstract
The objective of generated image traceability is to accurately identify and locate the source models. In this paper, we propose ReTD (Reconstruction-Based Traceability Detection), a generalized model for generated image traceability detection. Firstly, we use VAE to reconstruct images which are compared with the original ones to extract numerical distinguishing features. Secondly, we use Vision Transformer to learn the fine-grained distribution features to realize the generated image traceability classification. Finally, we conduct traceability experiments using images generated by ten GAN and Diffusion models. The experimental results demonstrate that ReTD only training a unified classifier improves accuracy by 9.4% compared to the state-of-the-art method. The ReTD-related code is availble at https://github.com/chenweizhuo/ReTD.
Weizhuo Chen, Fangfang Yuan, Cong Cao 0001, Dakui Wang, Yanbing Liu 0007
ICASSP5
2025 See Better, Say Better: Vision-Augmented Decoding for Mitigating Hallucinations in Large Vision-Language Models
Xinyi Sun, Diandian Guo, Cong Cao 0001, Fangfang Yuan, Dakui Wang, Yanbing Liu 0007
NLPCC (1)5
2024 Multi-Level Graph Convolutional Network for Document Information Extraction
abstract
Document information extraction aims to identify and extract entities and other essential information from documents. Its performance is significantly dependent on the learned relationship between the text and its corresponding layout, which is not fully exploited by existing methods. Therefore, we propose a multi-level graph convolutional network to improve the capability for modeling document information. Firstly, we integrate text embeddings with corresponding image and layout features as text features, which are used to construct a text-level graph to extract semantic relationships as word features through graph convolutions. We then construct a layout-level graph using region layout features as nodes, extracting structure relationships through further graph convolutions. This multi-level graph structure allows our model to fuse fine-grained word features with coarse-grained region features for effective sequence labeling. Comprehensive experiments on various datasets consistently demonstrate the effectiveness of our method, and the results show that our model outperforms the state-of-the-art methods with the aid of multi-level features.
Fangfang Yuan, Dakui Wang, Cong Cao 0001, Yanbing Liu 0007
ICTAI5
2024 Efficient One-Shot Pruning of Large Language Models with Low-Rank Approximation
abstract
Model pruning, as an effective method for compressing large language models (LLMs), has recently attracted considerable attention in the field of natural language processing. However, existing LLM pruning methods have two main drawbacks: (1) Iterative pruning for LLMs with over a billion parameters requires retraining, which leads to significant pruning costs. (2) LLMs Pruning is formalized as a weight reconstruction problem that necessitates second-order information, incurring expensive computations. To address these issues, we propose a novel pruning method named Eplra: efficient one-shot pruning of large language models with low-rank approximation, which efficiently identifies sparse networks in LLMs. Specifically, we design a novel pruning metric based on input activations for the rapid one-shot compression of LLMs. We first incorporate input activations into the calculation of weight importance to promote precise pruning of low-priority weights. Then, we perform local weight comparisons across each output of linear layers to induce uniform sparsity. Next, we expand Eplra into semi-structured pruning patterns to accommodate various acceleration scenarios. Finally, we employ low-rank parametrized update matrices to fine-tune the pruned model, facilitating a swift recovery of model performance. Experimental results on various language benchmark datasets demonstrate that Eplra outperforms the state-of-the-art methods.
Yangyan Xu, Cong Cao 0001, Fangfang Yuan, Rongxin Mi, Nannan Sun, Dakui Wang, Yanbing Liu 0007
SMC6
2024 Enhancing GPT-3.5 for Knowledge-Based VQA with In-Context Prompt Learning and Image Captioning
abstract
Traditional visual question answering (VQA) often falls short as merely relying on image information is insufficient to answer given questions. Therefore, Knowledge-Based Visual Question Answering (KB-VQA) has emerged. Typically, KB-VQA involves first retrieving knowledge from external knowledge bases, then using the retrieved knowledge in conjunction with the understanding of visual content for joint reasoning to predict answers. However, current models often suffer from weak visual perception capabilities when processing image information. Additionally, due to the incompleteness of external knowledge bases, retrieved knowledge may contain noise or even irrelevant information. Moreover, the re-embedding of knowledge text features during the model's reasoning process may deviate from the original meanings in the knowledge base. To address these challenges, we propose a method for Knowledge-Based Visual Question Answering (KB-VQA) using GPT-3.5, leveraging image captions and in-context prompts. We utilize an advanced captioning model to convert images into accurate textual representations, enhancing the large language model's understanding of visual information. Moreover, we eliminate the need for additional knowledge bases by directly employing GPT-3.5 as a knowledge base for knowledge retrieval and generate logically consistent text during inference to predict answers. Furthermore, we enhance GPT-3.5's question-answering capability for VQA through in-context prompt learning. Experiments on the public OK-VQA dataset demonstrate the superior performance of our model.
Yuling Yang, Cong Cao 0001, Fangfang Yuan, Dakui Wang, Yanbing Liu 0007
SMC5
2023 Few-shot Malicious Domain Detection on Heterogeneous Graph with Meta-learning
abstract
The Domain Name System (DNS), one of the essential basic services on the Internet, is often abused by attackers to launch various cyber attacks, such as phishing and spamming. Researchers have proposed many machine learning-based and deep learning-based methods to detect malicious domains. However, these methods rely on a large-scale dataset with labeled samples for model training. The fact is that the labeled domain samples are limited in the real-world DNS dataset. In this paper, we propose a few-shot malicious domain detection model named MetaDom, which employs a meta-learning algorithm for model optimization. Specifically, We first model the DNS scenario as a heterogeneous graph to capture richer information by analysing the complex relations among domains, IP addresses and clients. Then, we learn the domain representations with a heterogeneous graph neural network on the DNS HG. Finally, considering that only few labeled data are available in the real-world DNS scenario, a meta-learning algorithm with knowledge distillation is introduced to optimize the model. Extensive experiments on the real DNS dataset show that MetaDom outperforms other state-of-the-art methods.
Fangfang Yuan, Cong Cao 0001, Majing Su, Dakui Wang, Yanbing Liu 0007
CSCWD5
2023 MetaBERT: Collaborative Meta-Learning for Accelerating BERT Inference
abstract
Early exit methods are used to accelerate inference in pre-trained language models and maintain competitive performance on resource-constrained devices. However, existing methods for training early exit classifiers suffer from the problem of poor classifier representations in different layers, leading to difficulties in adapting to diverse natural language processing tasks. To address this issue, we propose MetaBERT: collaborative Meta-learning for accelerating BERT inference. The main goal of MetaBERT is to train early exit classifiers through collaborative meta-learning, in which case, few gradient updates can be quickly adapted to new tasks. Moreover, this novel meta-training approach produces good generalization performance, thus achieving an effective balance between the inference result and efficiency. Extensive experimental results show that our approach outperforms previous training methods by a large margin, and achieves state-of-the-art results compared to other competitive models.
Yangyan Xu, Fangfang Yuan, Cong Cao 0001, Majing Su, Dakui Wang, Yanbing Liu 0007
CSCWD6
2023 Robust Malicious Domain Detection Against Adversarial Attacks on Heterogeneous Graph
abstract
Domain Name System (DNS) is a crucial infrastructure of the Internet, yet it is also a primary medium for disseminating illicit information. Researchers have proposed numerous methods to detect malicious domains, among which heterogeneous graph (HG) based models have demonstrated good performance. However, their success may also motivate attackers to defeat HG based models in order to evade detection. In this paper, we propose a novel malicious domain detection model named RoDom, which is robust against adversarial attacks on HG. Firstly, we introduce different perturbations to construct multiple attacked graphs, which are designed to simulate different types of adversarial attacks on the HG. Secondly, we design a discriminator to perform robust representation learning on the HG by discriminating the original graph from attacked graphs. Finally, we introduce a classification selector to further improve the model's robustness by automatically combining domain representations of multiple HGs for domain classification. The experimental results show that RoDom out-performs other state-of-the-art methods and exhibits stronger robustness against adversarial attacks on the HG.
Zhiping Li, Fangfang Yuan, Dakui Wang, Cong Cao 0001, Yanbing Liu 0007
SMC5
2022 Noise Suppression with Label Graph in Distantly Supervised Relation Extraction
abstract
Distantly supervised relation extraction suffers from the influence of noise data. Some works solved this issue with relation-aware attention based on multi-instance learning to reduce the weights of noise data in sentence bag, which achieved remarkable results. However, the relationships of label-label and label-sentence in semantic space that can enhance label representation effectively was ignored. The enhanced label representation can keep the representation of noise data that does not contain the corresponding relation away from the label in semantic space. In view of this, we propose a novel method to capture the two relationships above to suppress noise in distantly supervised relation extraction with a label graph. To be specific, the single label of each sentence is first expanded to multi-label and the label graph is built based on the label hierarchy to capture implicit relation among them. Then the relation-aware attention in semantic space is deployed to assign low weights to noise sentences for the connection of label-sentence. Finally, the distantly supervised relation extraction task is optimized as a multi-label multi-classification problem. Experiment results on New York Times indicate that our method shows significantly improvement compared with the strong baseline methods.
Dakui Wang, Yangyang Ding, Xiaojun Chen 0004
CSCWD2
2022 DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints
abstract
Backdoor attack is a type of serious security threat to deep learning models. An adversary can provide users with a model trained on poisoned data to manipulate prediction behavior in test stage using a backdoor. The backdoored models behave normally on clean images, yet can be activated and output incorrect prediction if the input is stamped with a specific trigger pattern. Most existing backdoor attacks focus on manually defining imperceptible triggers in input space without considering the abnormality of triggers' latent representations in the poisoned model. These attacks are susceptible to backdoor detection algorithms and even visual inspection. In this paper, We propose a novel and stealthy backdoor attack - DEFEAT. It poisons the clean data using adaptive imperceptible perturbation and restricts latent representation during training process to strengthen our attack's stealthiness and resistance to defense algorithms. We conduct extensive experiments on multiple image classifiers using real-world datasets to demonstrate that our attack can 1) hold against the state-of-the-art defenses, 2) deceive the victim model with high attack success without jeopardizing model utility, and 3) provide practical stealthiness on image data.
Zhendong Zhao, Xiaojun Chen 0004, Yuexin Xuan, Ye Dong, Dakui Wang, Kaitai Liang
CVPR5
2022 KAFNN: A Knowledge Augmentation Framework to Graph Neural Networks
abstract
The semi-supervised node classification task is a basic problem in graph neural networks(GNNs). GNNs have shown their superiority in graph datasets over traditional neural networks such as Multilayer Perceptron. However, due to the limitation of Weisfeiler-Lehman, the existing GNNs will discard some prior knowledge, which is hard to be coped with, such as Dropout skill, etc. In this paper, we proposed a framework called KAFNN to introduce knowledge discarded obliviously to enhance data representation. KAFNN, based on the Siamese network, introduces the framework of combining GNNs and deep neural networks(DNNs) to capture the data presentation as whole as possible, which will inject more knowledge into GNNs. Extensive experiments based on seven public datasets and seven GNN models have shown that KAFNN has promoted presentation of several state-of-the-art GNN models in a competitive performance.
Bisheng Tang, Xiaojun Chen 0004, Dakui Wang, Zhendong Zhao
IJCNN3
2022 Rethinking the Feature Iteration Process of Graph Convolution Networks
abstract
Node classification is a fundamental research problem in graph neural networks(GNNs), which uses node's feature and label to capture node embedding in a low dimension. The existing graph node classification approaches mainly focus on GNNs from global and local perspectives. The relevant research is relatively insufficient for the micro perspective, which refers to the feature itself. In this paper, we prove that deeper GCNs' features will be updated with the same coefficient in the same dimension, limiting deeper GCNs' expression. To overcome the limits of the deeper GCN model, we propose a zero feature (k-ZF) method to train GCNs. Specifically, k-ZF randomly sets the initial k feature value to zero, acting as a data rectifier and augmenter, and is also a skill equipped with GCNs models and other GCNs skills. Extensive experiments based on three public datasets show that k-ZF significantly improves GCNs in the feature aspect and achieves competitive accuracy.
Bisheng Tang, Xiaojun Chen 0004, Dakui Wang, Zhendong Zhao
IJCNN3
2022 Application of distributed motion estimation for swarm MAVs in a GPS-restricted environment based on a wireless sensor network
Wenzhong Lou, Jinkui Wang, Zilong Su, Dakui Wang
J. Supercomput.4
2021 FLOD: Oblivious Defender for Private Byzantine-Robust Federated Learning with Dishonest-Majority
Ye Dong, Xiaojun Chen 0004, Kaiyun Li, Dakui Wang
ESORICS (1)4
2021 Enhancing Label Representations with Relational Inductive Bias Constraint for Fine-Grained Entity Typing
abstract
Fine-Grained Entity Typing (FGET) is a task that aims at classifying an entity mention into a wide range of entity label types. Recent researches improve the task performance by imposing the label-relational inductive bias based on the hierarchy of labels or label co-occurrence graph. However, they usually overlook explicit interactions between instances and labels which may limit the capability of label representations. Therefore, we propose a novel method based on a two-phase graph network for the FGET task to enhance the label representations, via imposing the relational inductive biases of instance-to-label and label-to-label. In the phase 1, instance features will be introduced into label representations to make the label representations more representative. In the phase 2, interactions of labels will capture dependency relationships among them thus make label representations more smooth. During prediction, we introduce a pseudo-label generator for the construction of the two-phase graph. The input instances differ from batch to batch so that the label representations are dynamic. Experiments on three public datasets verify the effectiveness and stability of our proposed method and achieve state-of-the-art results on their testing sets.
Xiaojun Chen 0004, Dakui Wang
IJCAI3
2021 A Differential Privacy Collaborative Deep Learning Algorithm in Pervasive Edge Computing Environment
abstract
With the development of 5G technology and intelligent terminals, the future direction of the Industrial Internet of Things (IIoT) evolution is Pervasive Edge Computing (PEC). In the pervasive edge computing environment, intelligent terminals can perform calculations and data processing. By migrating part of the original cloud computing model's calculations to intelligent terminals, the intelligent terminal can complete model training without uploading local data to a remote server. Pervasive edge computing solves the problem of data islands and is also successfully applied in scenarios such as vehicle interconnection and video surveillance. However, pervasive edge computing is facing great security problems. Suppose the remote server is honest but curious. In that case, it can still design algorithms for the intelligent terminal to execute and infer sensitive content such as their identity data and private pictures through the information returned by the intelligent terminal. In this paper, we research the problem of honest but curious remote servers infringing intelligent terminal privacy and propose a differential privacy collaborative deep learning algorithm in the pervasive edge computing environment. We use a Gaussian mechanism that meets the differential privacy guarantee to add noise on the first layer of the neural network to protect the data of the intelligent terminal and use analytical moments accountant technology to track the cumulative privacy loss. Experiments show that with the Gaussian mechanism, the training data of intelligent terminals can be protected reduction inaccuracy.
Dayin Zhang, Xiaojun Chen 0004, Jinqiao Shi, Dakui Wang
TrustCom4
2021 Robust node embedding against graph structural perturbations
Zhendong Zhao, Xiaojun Chen 0004, Dakui Wang, Yuexin Xuan
Inf. Sci.3
2020 A Shared-Word Sensitive Sequence-to-Sequence Features Extractor for Sentences Matching
abstract
Sentences matching is a basic task in Natural Language Processing (NLP). Interaction-based methods, which employ interactions between words of two sentences and construct word-level matching features to classify, are generally used due to their fine-grained features. However, they have many invalid interactions that may affect matching precision. In this paper, we limit the objects of interacting to shared words4 of two sentences. On the one hand, they can reduce invalid interactions. On the other hand, because of the different context semantics, the representation of the same word may be quite different, conversely, the representation difference can also be used to reflect the semantic difference of different contexts. To better extract global features of shared words, we introduce a sequence-to-sequence features extractor to force decoder to learn more contextual information from encoder. We implement the method based on Transformer[28], with syntactic parsing as additional knowledge. Our proposed method achieved better performance than strong baselines and the experiment results also demonstrate the efficiency of sequence-to-sequence features extractor and significance of the shared words.
Dakui Wang, Xiaojun Chen 0004, Pencheng Liao, Shujuan Chen
ECAI2
2020 Generate Images with Obfuscated Attributes for Private Image Classification
Dakui Wang, Xiaojun Chen 0004
MMM (2)2
2020 EaSTFLy: Efficient and secure ternary federated learning
Ye Dong, Xiaojun Chen 0004, Liyan Shen, Dakui Wang
Comput. Secur.4
2019 Privacy-Preserving Distributed Machine Learning Based on Secret Sharing
Ye Dong, Xiaojun Chen 0004, Liyan Shen, Dakui Wang
ICICS4
2018 Efficient and Private Set Intersection of Human Genomes
Liyan Shen, Xiaojun Chen 0004, Dakui Wang, Binxing Fang, Ye Dong
BIBM3
2018 Improve Word Mover's Distance with Part-of-Speech Tagging
abstract
Word Mover's Distance (WMD) is a document distance metric with free parameter, intelligible interpretation and unprecedented accuracy on document classification. WMD is on the basis of word embedding and largely focuses on semantic relationships rather than syntactic relationships, which would bring some limitations on measuring document distance. To enhance the impact of syntactic information, we proposed a new method called WMD with Part-of-Speech (PWMD) that integrates part-of-speech (POS) into the original WMD model. POS is a kind of syntactic information, providing more valuable features combined with WMD in document distance metric. Two combination strategies of the POS tagging are provided in “WMD, “word level” and “document level”. The results of contrastive experiments have shown that the PWMD is able to get better document distance than WMD.
Xiaojun Chen 0004, Li Bai 0004, Dakui Wang, Jinqiao Shi
ICPR3
2017 Efficient and Scalable Privacy-Preserving Similar Document Detection
abstract
Similar document detection has been well studied for many applications, such as file management systems, plagiarism and double submission detection. Traditional detection algorithms are challenged by the privacy-preserving problems. Recently, privacy-preserving similar document detection between two parties gains more attention. However, most of the existing works mainly focus on computing similarity between two documents, and they are inefficient with O(n2) computation complexity when processing secure comparison between two n-document sets. Focusing on this problem, this paper presents a new efficient and scalable privacy-preserving similar document detection protocol based on oblivious multi-garbled Bloom filter intersection and MinHash algorithm. Experimental evaluation shows that when processing large document sets, our protocol still remains linear computation complexity with the scale of document sets increasing and achieves overwhelming computational performance improvement against other major approaches.
Xiaojie Yu, Xiaojun Chen 0004, Jinqiao Shi, Liyan Shen, Dakui Wang
GLOBECOM5
2017 Improving Password Guessing Using Byte Pair Encoding
Dakui Wang, Xiaojun Chen 0004, Jinqiao Shi, Li Guo 0001
ISC2
2015 Modeling nuclear shape with boundary representation of object
abstract
The nucleus shape is of greatly importance for disease diagnosis and cell function. In this paper, a new model for nucleus shape is proposed to more accurately represent it, which mainly contains two components: boundary axis representation and weighted least-squares B-spline fitting. Boundary axis representation is achieved by segmenting the nuclear object extracted from original images into two parts with the major axis and recognizing the object boundary of each part. Then weighted least-squares B-spline fitting is employed to obtain two curves to represent the nucleus shape. Experiments and comparisons demonstrate the effectiveness of the proposed model. Moreover, outliers can be discarded by using weighted idea to remarkably improve the fitting accuracy.
Shuliang Wang 0001, Ying Li 0014, Dakui Wang, Hanning Yuan
BIBM3
2014 Topology Potential-Based Parameter Selecting for Support Vector Machine
Shuliang Wang 0001, Long Zhao 0002, Dakui Wang
ADMA4
2014 ELMDF: A new classification algorithm based on Data Field
abstract
In this paper, a novel classification algorithm, ELMDF (Extreme Learning Machine based on Data Field), is proposed to solve the problem of estimating the number of hidden layer neurons in typical ELM. For constructing ELMDF, a new theory based on data field, FMDF (Fundamental Matrix of Data Field) is proposed in this paper. The breast cancer cell image dataset, and the genome dataset are used to test and illustrate the proposed method. The experimental case demonstrates that ELMDF performs better than other six typical supervised learning algorithms on different datasets.
Shuliang Wang 0001, Dakui Wang
BIBM2