Xin Wang 0086

dblp:10/5630-86 · DBLP profile ↗
← Back
33ranked-venue papers
1as first author
22since 2021 · last 2026
0000-0002-1069-4830ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 8 since 2021Human-computer interaction and ubiquitous computing · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 PSPO: Trainable Potential-Based Reward Shaping with Internal Model Signals for Post-Training Policy Optimization of Large Language Models
Miaobo Hu, Bokun Wang, Shuhao Hu, Xin Wang 0086, Daren Zha, Jun Xiao 0005
ICIC (5)5
2025 EventPuzzle: A Benchmark for Multi-Perspective Event Prediction Based on Event Arguments
abstract
Event prediction is a critical task in natural language processing, aimed at reasoning and forecasting future events based on known event texts. This paper introduces EventPuzzle, a benchmark designed to evaluate the event prediction capabilities of large language models based on event arguments. By introducing argument points, we design tasks and evaluation methods to assess models' ability to predict events from different argument perspectives. EventPuzzle consists of both closed-ended and open-ended tasks. In the closed-ended task, models select the correct argument point from causal chains, while in the open-ended task, models generate event descriptions using two strategies: Argument-based Generation and Direct Generation. We construct an argument point dataset and evaluate multiple LLMs, demonstrating the models' performance across various tasks. Our experimental analysis reveals the strengths and limitations of current models and suggests future directions for improving event prediction.
Guoxuan Ding, Junhao Zhou, Xin Wang 0086, Daren Zha
CIKM5
2025 CoMuS-KG: A Collaborative Framework of Multimodal Unstructured Data and Knowledge Graph
abstract
Large language models (LLMs) have demonstrated remarkable capabilities in many fields, especially in complex neural lagnuage processing tasks. Despite their impressive performance, the content generated by LLMs still suffers from the problem of hallucination, particularly in tasks that require real-time data or specialized domain knowledge. Knowledge graphs and multimodal unstructured data serve as important sources of knowledge that can help address the hallucination issues in LLMs. However, existing methods mostly utilize knowledge graphs or multimodal unstructured data in isolation, neglecting the interaction between the two and it is the interaction that contributes to the extraction of deep knowledge in the knowledge base. In this paper, we propose a novel framework called the Collaborative Framework of Multimodal Unstructured Data and Knowledge Graph (CoMuS-KG). This framework enhances the reasoning capabilities of LLMs by enabling interaction between multimodal unstructured data and knowledge graphs, extracting deep knowledge from unstructured data, and completing missing information in knowledge graphs. Specifically, CoMuS-KG first decompose the question posed to the LLMs into multiple sub-questions and convert these sub-questions into knowledge graph triplets with missing head entity, tail entity, or relation. And then the knowledge graph and multimodal unstructured data are used to complete these triplets. Finally, we use the completed triplets to answer the original question and the completed triplets can be updated back into the knowledge graph to assist in other reasoning tasks. Extensive experiments on three KGQA benchmark datasets demonstrate the question-answering performance and reasoning capabilities of CoMuS-KG. Our code is publicly available at: https://github.comlGuChongAnlCoMuS-KG
Shuhao Hu, Xin Wang 0086, Ji Xiang, Lei Wang 0135, Jiahui Shen
CSCWD2
2025 Dual-Layer Meta-Learning for Few-Shot Named Entity Recognition
abstract
We propose a Dual-Layer Meta-Learning Network for Few-Shot Named Entity Recognition, where the network can selectively retain positive training signals from the memory chain to enhance the meta-model's learning capability and filter out interference from non-positive signals. Additionally, to mitigate the parameter explosion caused by the dual-layer network, we further use Chebyshev polynomials to fit the token classification function for entity span detection and employ the Kolmogorov-Arnold Network to fit the prototype-oriented classification function for entity span classification. This effectively reduces the runtime and GPU usage of the dual-layer structure.
Lei Wang 0135, Yange Wang, Xin Wang 0086
CSCWD5
2025 FlexFFN: Hierarchical Dynamic Selection of Feedforward Networks for Large Language Models
abstract
Optimizing the efficiency and adaptability of large language models (LLMs) for diverse downstream tasks remains a critical challenge. We propose FlexFFN (Flexible Feedforward Network), a novel framework that introduces hierarchical dynamic selection to enhance computational efficiency, flexibility, and performance in LLMs. At the macro level, FlexFFN leverages a Mixture of Experts (MoE) architecture to dynamically activate distinct FFN modules based on input characteristics. At the micro level, within each FFN module, a dynamic switching mechanism selects between KAN and traditional MLP, combining the rapid convergence capabilities of MLPs with the compositional learning and interpretability strengths of KAN. Additionally, FlexFFN integrates QLoRA (Quantized Low-Rank Adaptation) to significantly reduce memory requirements and computational costs during fine-tuning. By introducing these innovations, FlexFFN achieves a fine-grained balance between computational cost and model expressiveness, making it well-suited for large-scale training and deployment. Experimental results demonstrate that FlexFFN outperforms traditional architectures by reducing computational overhead while improving task-specific adaptability and model efficiency.
Miaobo Hu, Bokun Wang, Haoyuan Teng, Daren Zha, Xin Wang 0086, Jun Xiao 0005, Lei Wang 0135
IJCNN6
2025 IKG-Agent: Intent-driven Knowledge Graph Agent for Adaptive Workflow Reasoning
abstract
In this paper, we propose IKG-Agent, an intent-driven knowledge graph agent framework designed to perform adaptive workflow reasoning for complex question answering. Unlike traditional approaches, IKG-Agent integrates semantic parsing (SP), subgraph retrieval (SR), and large language models (LLMs) into a unified system. When a user submits a query, IKG-Agent first identifies the underlying intent and constructs a tailored reasoning workflow. Based on task complexity and query requirements, the framework dynamically selects optimal reasoning paths and tools, adjusting workflows during execution. By leveraging a shared knowledge memory system to continuously evaluate information sufficiency at each step, IKG-Agent mitigates error accumulation in traditional SP/SR-based reasoning—particularly for long relational paths and complex multi-hop inference. Experimental results demonstrate that IKG-Agent outperforms state-of-the-art methods on multiple public datasets, achieving significant improvements in accuracy and reliability for tasks requiring multi-level reasoning. Our code and data will be publicly released.
Yunzhi Liang, SiYang Tao, Haoyuan Teng, Xin Wang 0086, Lei Wang 0135, Ji Xiang
IJCNN5
2024 Common Forgery Artifact Driven Deepfake Face Detection
abstract
Given the substantial security risks associated with Deepfake technology, the identification of manipulated facial images has become a focal point of research. Regrettably, the majority of current Deepfake detection methods struggle to effectively discern forgery artifacts across various resolutions. Variations in image or video resolutions present substantial challenges to maintaining identity security in cooperative work environments. In this study, we introduce a Deepfake face detection model that relies on the identification of common forgery artifacts. Our model utilizes CFNet (Common Forgery Artifact Extraction Network) to automatically filter regions containing forged artifacts. These common forged artifacts are found in images of various resolutions, substantially enhancing the model’s accuracy in low-resolution images. Furthermore, our custom-designed multi-modal features ensure the model excels in high-resolution scenarios. Comprehensive experiments validate the efficacy of our model, achieving accuracy rates of 90.464% for Deepfakes, 75.520% for Face2Face, and 83.536% for FaceSwap within the Low Quality (LQ) category of the FF+ dataset.
Haotian Wu 0007, Xin Wang 0086, Ji Xiang, Liyue Ren
CSCWD2
2024 DST-FRD: A Distillation Method of Swin Transformer for Facial Reenactment Detection
abstract
In recent times, transformer-based deepfake detection networks have exhibited remarkable performance. However, the computational complexity and the number of parameters have constrained the practical application of these networks. To address these issues, we propose a knowledge distillation method for the Swin Transformer network. Specifically, this method utilizes the region prediction results of face images to distill the knowledge of the Swin Transformer in subregions, compensating for the deficiency of the small window size of the Swin Transformer in the early stage. Extensive experiments have demonstrated that our proposed distillation method not only reduces the parameters and computational effort of the model but also surpasses the teacher network in accuracy on low-resolution images. Our student network exhibits significantly lower computational complexity and fewer parameters than the teacher network, with reductions of only 17.96% and 20.44%, respectively. Despite this reduction in complexity and parameters, our student network has achieved state-of-the-art results on the FaceForensics++ dataset, surpassing the teacher network by 0.071%, 0.86%, and 8.339% on Raw/Raw, Raw/C23, and Raw/C40, respectively.
Haotian Wu 0007, Xin Wang 0086, Ji Xiang, Liyue Ren
CSCWD3
2024 CA-GCN: A Confidence-aware uncertain knowledge graph embedding model based on graph convolutional networks
abstract
Compared with deterministic knowledge graphs, uncertain knowledge graphs can better describe the inherent uncertainty of relations facts in the real world. However, existing embedding models for uncertain knowledge graphs can not capture the direct effect of the confidence score on information propagation between different entities across different relations. This paper proposes a confidence-aware knowledge graph embedding model based on graph convolutional networks (CA-GCN). It enables the propagation of confidence information within the graph structure, allowing confidence scores to directly influence message passing, and ultimately improves the performance of unseen relation facts confidence prediction. To mitigate the adverse effects of false-negative samples during training, we design a collaborative mechanism between a confidence generator and a relation discriminator. The relation discriminator is used to determine if a given triplet is valid, and the confidence generator is used to generate confidence scores. In this way, confidence information of positive and negative samples can play unique roles in different encoders at different stages to avoid the influence of false negatives. Experiments on three widely used datasets show promising results with an obvious performance improvement compared to existing uncertain knowledge embedding models.
Lin Zhao 0006, Ji Xiang, Yunzhi Liang, Zeyi Liu 0002, Xin Wang 0086
CSCWD6
2024 Adaptive Spatial-Temporal Hypergraph Fusion Learning for Next POI Recommendation
abstract
Next point-of-interest (POI) recommendation has been a trending task to provide next POI suggestions. Most existing sequential-based and graph-based methods have endeavored to model user visiting behaviors and achieved considerable performances. However, they have either modeled user interests at a coarse-grained interaction level or ignored complex high-order feature interactions through general heuristic message passing scheme, making it challenging to capture complementary effects. To tackle these challenges, we propose a novel framework Adaptive Spatial-Temporal Hypergraph Fusion Learning (ASTHL) for next POI recommendation. Specifically, we design disentangled POI-centric learning to decouple spatial-temporal factors and utilize cross-view contrastive learning to enhance the quality of POI representations. Furthermore, we propose multi-semantic enhanced hypergraph learning to adaptively fuse spatial-temporal factors through well-designed aggregation and propagation scheme. Extensive experiments on three real-world datasets validate the superiority of our proposal over various state-of-the-arts. To facilitate future research, our code is available at https://github.com/icmpnorequest/ICASSP2024_ASTHL.
Yantong Lai, Yijun Su, Lingwei Wei, Daren Zha, Xin Wang 0086
ICASSP6
2024 A Redundant Relation Reduced Bidirectional Extraction Framework Based on SpanBERT for Relational Triple Extraction
Ji Xiang, Lei Wang 0135, Xin Wang 0086
ICIC (13)5
2024 A Meta-pattern-enhanced Generative Few-shot Attribute Extraction Framework for Open-world Sparse Corpora
abstract
Open-world attribute extraction is one of the most important tasks of information extraction aiming to mine all the valuable attributes of entities and their corresponding values from unstructured texts, usually in the form of (entity, attribute, value) triplets. However, existing methods have difficulty extracting attribute triplets from open-world sparse corpora where the attribute names are not previously given, especially in the few- shot scenario with only few manual annotations available. To solve the above problems, we propose a two-stage Meta-pattern-Enhanced Generative Few-shot Attribute Extraction (MEGFAE) framework which can be used to discover utmost valuable attribute triplets from open-world sparse corpora in a generative manner. For evaluation on open-world sparse corpora, we introduce a benchmark dataset called OSN-51511The dataset is available in https://github.com/sunshower-liu/OSN-515.. Experimental results verifies the effectiveness of our framework and inspires future explorations on the text mining on sparse corpora.
Xiyu Liu 0003, Xin Wang 0086, Zeyi Liu 0002, Nan Mu, Tianshu Fu, Ji Xiang
MSN3
2024 Modeling the Uncertainty of Information Propagation for Rumor Detection: A Neuro-Fuzzy Approach
abstract
Automatic rumor detection is critical for maintaining a healthy social media environment. The mainstream methods generally learn rich features from information cascades by modeling the cascade as a tree or graph structure where edges are built based on interactions between a tweet and retweets. Some psychology studies have empirically shown that users' various subjective factors always cause the uncertainty of interactions such as differences among interactive behavior activation thresholds or semantic relevancy. However, previous works model interactions by employing a simple fully connected layer on fixed edge weights in the graph and cannot reasonably describe this inherent uncertainty of complex interactions. In this article, inspired by the fuzzy theory, we propose a novel neuro-fuzzy method, fuzzy graph convolutional networks (FGCNs), to sufficiently understand uncertain interactions in the information cascade in a fuzzy perspective. Specifically, a new strategy of graph construction is first designed to convert each information cascade into a heterogeneous graph structure with the consideration of explicit interactive behaviors between a tweet and its retweet, as well as implicit interactive behaviors among retweets, enriching more structural clues in the graph. Then, we improve graph convolutional networks by incorporating edge fuzzification (EF) modules. The EFs adapt edge weights according to predefined membership to enhance message passing in the graph. The proposed model can provide a stronger relational inductive bias for expressing uncertain interactions and capture more discriminative and robust structural features for rumor detection. Extensive experiments demonstrate the effectiveness and superiority of FGCN on both rumor detection and early rumor detection.
Lingwei Wei, Dou Hu 0001, Wei Zhou 0019, Xin Wang 0086, Songlin Hu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 RZSR Randomly Initialized Zero-Shot Method for Blind Super-Resolution
abstract
When the unknown degradation is mixed with unknown blurry kernels, how to perform super-resolution operation is an open issue. The mean idea of the existing zero-shot and non-zero-shot methods is to estimate blurry kernel. The effects of these methods depend on the accuracy of the deduced blurry kernel. In this paper, we propose Randomly initialized Zero-Shot Super-Resolution (RZSR) training strategy. RZSR is a zero-shot training method and it allows the network to extract low-resolution image features and generate its counterpart high-resolution images under the interference of degradation algorithms. We further propose two model-agnostic modules which are Adaptive Information Extraction Module (AIEM) and knowledge dictionary. They respectively assist the network to extract features and well fit the data distribution of clear images. RZSR can be applied to any single image super-resolution and video super-resolution models. We prove the generalization ability and superiority of RZSR through a series of experiments.
Tianshu Fu, Guanqun Liu 0002, Xin Wang 0086, Daren Zha, Jiahui Shen
CSCWD3
2023 RW-MMDCG: Muti-modal via Rolling-Window Directed Graph Network for Conversational Emotion Recognition
abstract
Multimodal Conversational Emotion recognition (MMCER) aims to detect the muti-emotion label for each utterance from heterogeneous visual, text and audio modalities. In this paper, we focus on applying multi-modal graph data structures to conversational emotion recognition and use a novel and efficient graph—MMDCGs to better integrate multi-modal contextual information into conversations. MMDCG provides a new way of encoding intrinsic structural connectivity. Besides, inspired by time series analysis, we set a rolling time window as the receptive field, which can reduce the interference of remote information on the current utterances detection and achieve the purpose of data enhancement. We innovatively ensemble such graph structures with transformers, named rolling-windows MMDCGs (RW-MMDCG). Comprehensive experiments are performed on two representative multi-modal datasets, IEMOCAP and MELD, and we compare them with existing baselines, demonstrating the great advantages and effectiveness of RW-MMDCG.
Daren Zha, Qingfei Zhao, Yuanye He, Xin Wang 0086
CSCWD5
2022 A Noise-Aware Framework for Blind Image Super-Resolution
abstract
The real-world image degradation in the super-resolution task is recently considered as a combination of Gaussian blur, down-sampling, and additional white Gaussian noise. To han-dle this degradation, previous methods estimate the Gaussian blur kernel or model the degradation based on a randomly selected image patch. However, these methods cannot han-dle degradations with high-level noise well as they ignore the spatial variability or even the existence of noise. Moreover, using image denoising networks to preprocess low-resolution images also fails due to the loss of important high-frequency information. In this paper, we propose a framework called EASE to flexibly handle real-world degradations. Specifi-cally, we develop a lightweight module to erase noise and blur simultaneously by learning from an image denoising and an image restoration network, which adapts to existing net-works that focus on handling bicubic down-sampling. Exten-sive experiments prove the superiority of our method, espe-cially when handling degradations with high-level noise.
Guanqun Liu 0002, Xin Wang 0086, Lei Wang 0135, Daren Zha, Lin Zhao 0006, Zhe Kong, Peng Qi 0005
ICME2
2022 Searching Models with Nested Attention for Blind Super-Resolution
abstract
Blind super-resolution task aims to restore low-resolution im“ages with unknown degradations to high-resolution counter-parts. Existing methods rely on degradations estimation to re-construct high-resolution images. However, they need human involvement to obtain the best results as they treat unknown types of degradations as known conditions and manually select corresponding trained models. Moreover, they cannot fully use estimated degradations and generate blurry artifacts as they ignore that the impact of degradations on images is re-lated to images contents. In this paper, we propose HIS-NEST which contains an automatic search strategy HIS and a net-work structure NEST. Specifically, to bypass manual partici-pation, HIS automatically selects the clearest image by esti-mating the qualities of generated images. Furthermore, NEST protects the connection between degradations and images by using no loss functions to limit the degradations estimation and analyzing degradations from the perspective of channel and space. Extensive experiments show that our method out-performs state-of-the-art methods.
Guanqun Liu 0002, Xin Wang 0086, Lei Wang 0135, Daren Zha, Lin Zhao 0006, Zhe Kong, Peng Qi 0005
ICME2
2022 Multi-level Fusion of Multi-modal Semantic Embeddings for Zero Shot Learning
abstract
Zero shot learning aims to recognize objects whose instances may not be covered by the training data. To generalize knowledge from seen classes to the novel ones, semantic space is built to embed knowledge from various views into multi-modal semantic embeddings. Existing semantic embeddings neglect the relationships between classes which are essential to transfer knowledge between classes. Moreover, existing zero shot learning models ignore the complementarity between semantic embeddings from different modalities. To tackle these problems, in this work, we resort to graph theory to explicitly model the interdependence between classes and then obtain new modal semantic embeddings. Furthermore, we pioneer to propose a multi-level fusion model to effectively combine knowledge encoded in multi-modal semantic embeddings together. By the virtue of subsequent fusion block, the results of multi-level fusion can be furtherly enriched and fused. Experiments show that our model could achieve promising results on various datasets. Ablation study suggests that our method is well suited for zero shot learning.
Zhe Kong, Xin Wang 0086, Neng Gao, Yuhan Liu 0012, Chenyang Tu
ICMI2
2022 GGViT: Multistream Vision Transformer Network in Face2Face Facial Reenactment Detection
abstract
Detecting manipulated facial images and videos on social networks has been an urgent problem to be solved. The compression of videos on social media has destroyed some pixel details that could be used to detect forgeries. Hence, it is crucial to detect manipulated faces in videos of different quality. We propose a new multi-stream network architecture named GGViT, which utilizes global information to improve the generalization of the model. The embedding of the whole face extracted by ViT will guide each stream network. Through a large number of experiments, we have proved that our proposed model achieves state-of-the-art classification accuracy on FF++ dataset, and has been greatly improved on scenarios of different compression rates. The accuracy of Raw/C23, Raw/C40 and C23/C40 was increased by 24.34%, 15.08% and 10.14% respectively.
Haotian Wu 0007, Xin Wang 0086, Ji Xiang
ICPR3
2021 Entity and Relation Matching Consensus for Entity Alignment
abstract
Entity alignment aims to match synonymous entities across different knowledge graphs, which is a fundamental task for knowledge integration. Recently, researchers have devoted to leveraging rich information within relations to enhance entity alignment. They explicitly incorporate relations in entity representation and alignment, demonstrating remarkable results. However, affected by the semantic assumptions from early works, these works represent a relation by combining all the entities it connects, ignoring the semantic independence between entity and relation. Moreover, since these works perform alignment by comparing embedding similarity, they fail to consider a graph level alignment and tend to find local false correspondences.
Jinzhu Yang, Wei Zhou 0019, Wanhui Qian, Xin Wang 0086, Jizhong Han, Songlin Hu 0001
CIKM5
2021 Efficient, Low-Cost, Real-Time Video Super-Resolution Network
Guanqun Liu 0002, Xin Wang 0086, Daren Zha, Lei Wang 0135, Lin Zhao 0006
ICONIP (4)2
2021 FRAGAN-VSR: Frame-Recurrent Attention Generative Adversarial Network for Video Super-Resolution
abstract
Video super resolution (SR) is an important task, which recovers high-resolution (HR) frames from consecutive low-resolution (LR) couterparts. The most advanced works achieved good performance to this day. However, most of them has largely focussed on making a breakthrough in accuracy and speed, which has neglect that how to recover the finer texture details. Therefore, in this paper, we first present an Video SR model combined generative adversarial network and recurrent neural network (GAN-RNN) structure. It is forced by the self-attention mechanism to pay great attention to the high-frequency information of the LR frames. The perceptual loss is introduced to retain the high-frequency detail which is different from other video SR network. A great deal of evaluations and comparisons with previous methods have confirmed the merits of the proposed framework which can significantly outperform the current state of the art.
Guanqun Liu 0002, Daren Zha, Xin Wang 0086, Lin Zhao 0006, Lei Wang 0135
ICTAI5
2020 SIDGAN: Single Image Dehazing without Paired Supervision
abstract
Single image dehazing is challenging without scene airlight and transmission map. Most of existing dehazing algorithms tend to estimate key parameters based on manual designed priors or statistics, which may be invalid in some scenarios. Although deep learning-based dehazing methods provide an effective solution, most of them rely on paired training datasets, which are prohibitively difficult to be collected in real world. In this paper, we propose an effective end-to-end generative adversarial network for single image dehazing, named SIDGAN. The proposed SIDGAN adopts a U-net architecture with a novel color-consistency loss derived from dark channel prior and perceptual loss, which can be trained in an unsupervised fashion without paired synthetic datasets. We create a RealHaze dataset for network training, including 4,000 outdoor hazy images and 4,000 haze-free images. Extensive experiments demonstrate that our proposed SIDGAN achieves better performance than existing state-of-the-art methods on both synthetic datasets and real-world datasets in terms of PSNR, SSIM, and subjective visual experience.
Pan Wei, Xin Wang 0086, Lei Wang 0135, Ji Xiang
ICPR2
2020 Hierarchical Interaction Networks with Rethinking Mechanism for Document-Level Sentiment Analysis
Lingwei Wei, Dou Hu 0001, Wei Zhou 0019, Xuehai Tang, Xiaodan Zhang 0004, Xin Wang 0086, Jizhong Han, Songlin Hu 0001
ECML/PKDD (3)6
2020 TransMVG: Knowledge Graph Embedding Based on Multiple-Valued Gates
Neng Gao, Jun Yuan 0008, Xin Wang 0086, Lei Wang 0135
WISE (1)4
2020 Network representation learning with ensemble methods
Ji Xiang, Xin Wang 0086
Neurocomputing3
2019 STNet: A Style Transformation Network for Deep Image Steganography
Neng Gao, Xin Wang 0086, Ji Xiang, Guanqun Liu 0002
ICONIP (2)3
2019 TransI: Translating Infinite Dimensional Embeddings Based on Trend Smooth Distance
Neng Gao, Lei Wang 0135, Xin Wang 0086
KSEM (1)4
2019 HidingGAN: High Capacity Information Hiding with Generative Adversarial Network
abstract
Abstract Image steganography is the technique of hiding secret information within images. It is an important research direction in the security field. Benefitting from the rapid development of deep neural networks, many steganographic algorithms based on deep learning have been proposed. However, two problems remain to be solved in which the most existing methods are limited by small image size and information capacity. In this paper, to address these problems, we propose a high capacity image steganographic model named HidingGAN. The proposed model utilizes a new secret information preprocessing method and Inception‐ResNet block to promote better integration of secret information and image features. Meanwhile, we introduce generative adversarial networks and perceptual loss to maintain the same statistical characteristics of cover images and stego images in the high‐dimensional feature space, thereby improving the undetectability. Through these manners, our model reaches higher imperceptibility, security, and capacity. Experiment results show that our HidingGAN achieves the capacity of 4 bits‐per‐pixel (bpp) at 256 × 256 pixels, improving over the previous best result of 0.4 bpp at 32 × 32 pixels.
Neng Gao, Xin Wang 0086, Ji Xiang, Daren Zha
Comput. Graph. Forum3
2018 SSteGAN: Self-learning Steganography Based on Generative Adversarial Networks
Neng Gao, Xin Wang 0086, Xuexin Qu
ICONIP (2)3
2018 Structure, Attribute and Homophily Preserved Social Network Embedding
Xiang Li 0045, Jiahui Shen, Xin Wang 0086
ICONIP (6)4
2018 Perceptual-DualGAN: Perceptual Losses for Image to Image Translation with Generative Adversarial Nets
abstract
Thinking about cross-domain image-to-image translation problems, where an input image belonging to domain U is transformed into an output image belonging to another domain V. A series of typical tasks, such as style transformation, colorization, super-resolution, can be seen as cross-domain image-to-image translation tasks. Recent methods such as Conditional Generative Adversarial Networks (cGANs) make big progress in this field, but they require paired image data, which is hard to obtain. The DualGAN (Unsupervised Dual Learning for Image-to-Image Translation) architecture was proposed to solve the issue of lack of paired data. But the pixel-level reconstruction losses of DualGAN are simple. In this paper, we replace the pixel-level reconstruction losses with the perceptual reconstruction losses, and propose a more advanced framework for cross-domain image-to-image translation named perceptual-DualGAN. The perceptual reconstruction losses consist of feature reconstruction losses and style reconstruction losses, both of them are computed from pretrained loss networks. Experiments on multiple image translation tasks show that our framework almost performs superior to other methods. And the results of experiments illustrate that our framework can generate more realistic and more natural photos.
Xuexin Qu, Xin Wang 0086, Lei Wang 0135, Lingchen Zhang
IJCNN2
2016 Novel MITM Attacks on Security Protocols in SDN: A Feasibility Study
Xin Wang 0086, Neng Gao, Lingchen Zhang, Zongbin Liu, Lei Wang 0135
ICICS1