VLDB 2026 Research / reviewers in the wild / expert
Xin Wang 0045
dblp:10/5630-45
· DBLP profile ↗
48ranked-venue papers
4as first author
37since 2021 · last 2026
0000-0002-7528-2407ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 3 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RL-I2IT: Image-to-image translation with deep reinforcement learning
Jing Hu 0009, Ziwei Luo 0002, Chengming Feng, Shu Hu 0001, Bin B. Zhu, Xi Wu 0004, Xin Li 0005, Hongtu Zhu, Siwei Lyu, Xin Wang 0045 |
Neural Networks | 10 |
| 2025 | A Sequential Multi-Stage Approach for Code Vulnerability Detection via Confidence- and Collaboration-based Decision MakingabstractWhile large language models (LLMs) have shown strong capabilities across diverse domains, their application to code vulnerability detection holds great potential for identifying security flaws and improving software safety.In this paper, we propose a sequential multi-stage approach via confidence-and collaboration-based decision making (ConColl).The system adopts a three-stage sequential classification framework, proceeding through a single agent, retrieval-augmented generation (RAG) with external examples, and multi-agent reasoning enhanced with RAG.The decision process selects among these strategies to balance performance and cost, with the process terminating at any stage where a high-certainty prediction is achieved.Experiments on a benchmark dataset and a low-resource language demonstrate the effectiveness of our framework in enhancing code vulnerability detection performance. Chung-Nan Tsai, Xin Wang 0045, Cheng-Hsiung Lee, Chingsheng Lin |
EMNLP | 2 |
| 2025 | Preserving AUC Fairness in Learning with Noisy Protected GroupsabstractThe Area Under the ROC Curve (AUC) is a key metric for classification, especially under class imbalance, with growing research focus on optimizing AUC over accuracy in applications like medical image analysis and deepfake detection. This leads to fairness in AUC optimization becoming crucial as biases can impact protected groups. While various fairness mitigation techniques exist, fairness considerations in AUC optimization remain in their early stages, with most research focusing on improving AUC fairness under the
assumption of clean protected groups. However, these studies often overlook the impact of noisy protected groups, leading to fairness violations in practice. To address this, we propose the first robust AUC fairness approach under noisy protected groups with fairness theoretical guarantees using distributionally robust optimization. Extensive experiments on tabular and image datasets show that our method outperforms state-of-the-art approaches in preserving AUC fairness. The code is in https://github.com/Purdue-M2/AUC_Fairness_with_Noisy_Groups. Wenbin Zhang 0002, Xin Wang 0045, Zhenhuan Yang, Shu Hu 0001 |
ICML | 4 |
| 2025 | RLMiniStyler: Light-weight RL Style Agent for Arbitrary Sequential Neural Style GenerationabstractArbitrary style transfer aims to apply the style of any given artistic image to another content image. Still, existing deep learning-based methods often require significant computational costs to generate diverse stylized results. Motivated by this, we propose a novel reinforcement learning-based framework for arbitrary style transfer RLMiniStyler. This framework leverages a unified reinforcement learning policy to iteratively guide the style transfer process by exploring and exploiting stylization feedback, generating smooth sequences of stylized results while achieving model lightweight. Furthermore, we introduce an uncertainty-aware multi-task learning strategy that automatically adjusts loss weights to adapt to the content and style balance requirements at different training stages, thereby accelerating model convergence. Through a series of experiments across image various resolutions, we have validated the advantages of RLMiniStyler over other state-of-the-art methods in generating high-quality, diverse artistic image sequences at a lower cost. Codes are available at https://github.com/fengxiaoming520/RLMiniStyler. Jing Hu 0009, Chengming Feng, Shu Hu 0001, Ming-Ching Chang, Xin Li 0005, Xi Wu 0004, Xin Wang 0045 |
IJCAI | 7 |
| 2025 | LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language ModelsabstractAccurate and efficient question-answering systems are essential for high-quality patient care in the medical field. While Large Language Models (LLMs) have made remarkable strides across various domains, they still face challenges in medical question answering, particularly in understanding domain-specific terminology and performing complex reasoning, limiting their effectiveness in critical applications. To address this, we propose a multi-agent medical question-answering (MedQA) system incorporating similar case generation. We leverage the Llama3.1:70B model in a multi-agent architecture to enhance enhance zero-shot classification on the MedQA dataset, utilizing the model’s inherent medical knowledge and reasoning capabilities without additional training data. Experimental results show substantial gains over existing benchmark models, with improvements of 7% in both accuracy and F1-score across various medical QA tasks. Furthermore, we examine the model’s interpretability and reliability in addressing complex medical queries. This research not only offers a robust solution for medical question answering but also establishes a foundation for broader applications of LLMs in the medical domain. Yineng Chen, Chingsheng Lin, Shu Hu 0001, Jinrong Hu, Xi Wu 0004, Xin Wang 0045 |
IJCNN | 9 |
| 2025 | VB-KGN: Variational Bayesian Kernel Generation Networks for Motion Image DeblurringabstractMotion blur estimation is a critical and fundamental task in scene analysis and image restoration. While most state-of-the-art deep learning-based methods for single-image motion image deblurring focus on constructing deep networks or developing training strategies, the characterization of motion blur has received less attention. In this paper, we innovatively propose a non-parametric Variational Bayesian Kernel Generation Network (VB-KGN) for characterizing motion blur in a single image. To solve this model, we employ the variational inference framework to approximate the expected statistical distribution of motion blur images in a data-driven manner. The qualitative and quantitative evaluations of our experimental results demonstrate that our proposed model can generate highly accurate motion blur kernels, significantly improving motion image deblurring performance and substantially reducing the need for extensive training sample preprocessing for deblurring tasks. Ying Fu 0003, Xiaojie Li 0001, Xin Wang 0045, Xi Wu 0004, Shu Hu 0001, Siwei Lyu, Wei Liu 0044 |
IEEE Trans. Multim. | 4 |
| 2025 | Spotting the Fakes: A Deep Dive into GAN-Generated Face DetectionabstractGenerative Adversarial Networks (GANs) have enabled the creation of highly authentic facial images, which are increasingly used in deceptive social media profiles and other forms of disinformation, resulting in serious consequences. Significant progress has been made in developing GAN-generated face detection systems to identify these fake images. This study offers a comprehensive review of recent advancements in GAN-generated face detection, focusing on techniques that detect facial images generated by GAN models. We categorize detection methods into three groups: (1) deep learning-based approaches, (2) physics-based methods, and (3) physiology-based methods. We summarize key concepts in each category, connecting them to relevant implementations, datasets, and evaluation metrics. Additionally, we provide a comparative analysis between automated detection and human visual performance to highlight the strengths and weaknesses of both approaches. Furthermore, we review related surveys, including detecting morphed faces, manipulated faces, DeepFake, and faces generated by diffusion models. Finally, we discuss unresolved challenges and suggest potential directions for future research. Xin Wang 0045, Ting Yu Tsai, Shu Hu 0001, Ming-Ching Chang, Pradeep K. Atrey, Siwei Lyu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Robust CLIP-Based Detector for Exposing Diffusion Model-Generated ImagesabstractDiffusion models (DMs) have revolutionized image generation, producing high-quality images with applications spanning various fields. However, their ability to create hyper-realistic images poses significant challenges in distinguishing between real and synthetic content, raising concerns about digital authenticity and potential misuse in creating deepfakes. This work introduces a robust detection framework that integrates image and text features extracted by CLIP model with a Multilayer Perceptron (MLP) classifier. We propose a novel loss that can improve the detector’s robustness and handle imbalanced datasets. Additionally, we flatten the loss landscape during the model training to improve the detector’s generalization capabilities. The effectiveness of our method, which outperforms traditional detection techniques, is demonstrated through extensive experiments, underscoring its potential to set a new state-of-the-art approach in DM-generated image detection. The code is available at https://github.com/Purdue-M2/RobustDM_Generated_Image_Detection. Santosh, Irene Amerini, Xin Wang 0045, Shu Hu 0001 |
AVSS | 4 |
| 2024 | RepMedGraf: Re-parameterization Medical Generated Radiation Field for Improved 3D Image ReconstructionabstractIn the field of medical imaging, 3D image reconstruction has emerged as a crucial technique for accurate disease diagnosis. Neural Radiance Field (NeRF) has shown promise in generating high-quality 3D models through deep learning, but its application in medical imaging remains limited. This paper presents RepMedGraf, an improved model addressing the limitations of NeRF. RepMedGraf utilizes a generator based on lightweight RepVGG Blocks instead of MedNeRF model. By doing so, the risk of overfitting is reduced, and training efficiency is improved. The proposed method is trained on publicly available chest and knee datasets. Comparative evaluations are conducted based on the generated radiation fields, demonstrating the effectiveness and quality of the 3D models produced by RepMedGraf. The results highlight the potential of RepMedGraf as a valuable tool in medical imaging for enhanced diagnostic accuracy and improved patient care. Ruotong Sun, Fayang Liao, Qinrui Fan, Shu Hu 0001, Xin Wang 0045, Jing Hu 0009 |
AVSS | 8 |
| 2024 | CGD-Net: A Hybrid End-to-end Network with gating decoding for Liver Tumor Segmentation from CT ImagesabstractLiver tumor segmentation plays a crucial role in the diagnosis and treatment of hepatic lesions. However, accurate tumor segmentation remains a challenging task due to the fuzzy boundaries of liver tumors and the uncertainty in shape, size, and location. In this paper, we propose a new end-to-end segmentation network called CGD-Net, which incorporates Transformer and frequency-domain features into a convolutional network, and proposes a new decoder structure to automatically learn from Segmentation of liver tumors in CT images. The proposed CGD-Net consists of a Transformer encoder, a frequency domain information fusion module, a gated decoder and three skip connections. Using the powerful feature extraction capability of the Transformer encoder to extract multi-feature information.CGD(Control Gate Decoder)blocks gradually restore the feature information lost in the encoding process by emphasizing the original information.In order to fully utilize the information of the original image, three skip connections are used to connect each encoder layer and its corresponding decoder layer, and a FEM (Frequency-domain Enhance module)is built in the third skip connection to fuse frequency domain features. Experiments on the LiTS dataset validate that the proposed CGD-Net can effectively segment liver tumors from CT images in an end-to-end manner, with segmentation accuracy exceeding many existing methods. Xiaogang Zhu 0003, Ziqiu Liu, Ouyang Shaobo, Xin Wang 0045, Shu Hu 0001, Feng Ding 0007 |
AVSS | 5 |
| 2024 | QRNN-Transformer: Recognizing Textual EntailmentabstractIn recent years, the Transformer model based on the self-attention mechanism has made significant progress in natural language processing and has also been applied in the text implication recognition task, achieving excellent results. However, the Transformer model still has deficiencies in modeling local information in the text. To improve the Transformer model, the QRNN-Transformer was proposed, which uses the QRNN network to divide the input text sequence into local short sequences to capture the local information of the input text. The self-attention is improved by combining the gating mechanism to make the model select tasks-related words or features. Extensive experiments have demonstrated the QRNN-Transformer model can effectively improve the accuracy of entailment relationship recognition on both English and Chinese datasets.l Xiaogang Zhu 0003, Zhihan Yan, Wenzhan Wang, Shu Hu 0001, Xin Wang 0045, Chunnian Liu |
AVSS | 5 |
| 2024 | Preserving Fairness Generalization in Deepfake DetectionabstractAlthough effective deepfake detection models have been developed in recent years, recent studies have revealed that these models can result in unfair performance disparities among demographic groups, such as race and gender. This can lead to particular groups facing unfair targeting or exclusion from detection, potentially allowing misclassified deepfakes to manipulate public opinion and undermine trust in the model. The existing method for addressing this problem is providing a fair loss function. It shows good fairness performance for intra-domain evaluation but does not maintain fairness for cross-domain testing. This highlights the significance of fairness generalization in the fight against deepfakes. In this work, we propose the first method to address the fairness generalization problem in deepfake detection by simultaneously considering features, loss, and optimization aspects. Our method employs disentanglement learning to extract demographic and domain-agnostic forgery features, fusing them to encourage fair learning across a flattened loss landscape. Extensive experiments on prominent deepfake datasets demonstrate our method's effectiveness, surpassing state-of-the-art approaches in preserving fairness during cross-domain deepfake detection. The code is available at https://github.com/Purdue-M2/Fairness-Generalization. Xinan He, Yan Ju, Xin Wang 0045, Feng Ding 0007, Shu Hu 0001 |
CVPR | 4 |
| 2024 | Masked Conditional Diffusion Model for Enhancing Deepfake DetectionabstractRecent studies on deepfake detection have achieved promising results when training and testing faces are from the same dataset. However, their results severely degrade when confronted with forged samples that the model has not yet seen during training. In this paper, deepfake data to help detect deepfakes. this paper present we put a new insight into diffusion model-based data augmentation, and propose a Masked Conditional Diffusion Model (MCDM) for enhancing deepfake detection. It generates a variety of forged faces from a masked pristine one, encouraging the deepfake detection model to learn generic and robust representations without overfitting to special artifacts. Extensive experiments demonstrate that forgery images generated with our method are of high quality and helpful to improve the performance of deepfake detection models. Tiewen Chen, Shanmin Yang, Shu Hu 0001, Zhenghan Fang, Ying Fu 0003, Xi Wu 0004, Xin Wang 0045 |
IJCNN | 7 |
| 2024 | Uncertainty-Aware Explainable Recommendation with Large Language ModelsabstractProviding explanations within the recommendation system would boost user satisfaction and foster trust, especially by elaborating on the reasons for selecting recommended items tailored to the user. The predominant approach in this domain revolves around generating text-based explanations, with a notable emphasis on applying large language models (LLMs). However, refining LLMs for explainable recommendations proves impractical due to time constraints and computing resource limitations. As an alternative, the current approach involves training the prompt rather than the LLM. In this study, we developed a model that utilizes the ID vectors of user and item inputs as prompts for GPT-2. We employed a joint training mechanism within a multi-task learning framework to optimize both the recommendation task and explanation task. This strategy enables a more effective exploration of users’ interests, improving recommendation effectiveness and user satisfaction. Through the experiments, our method achieving 1.59 DIV, 0.57 USR and 0.41 FCR on the Yelp, TripAdvisor and Amazon dataset respectively, demonstrates superior performance over four SOTA methods in terms of explainability evaluation metric. In addition, we identified that the proposed model is able to ensure stable textual quality on the three public datasets. Yicui Peng, Chingsheng Lin, Guo Huang, Jinrong Hu, Bin Kong 0001, Shu Hu 0001, Xi Wu 0004, Xin Wang 0045 |
IJCNN | 10 |
| 2024 | X-Transfer: A Transfer Learning-Based Framework for GAN-Generated Fake Image DetectionabstractGenerative adversarial networks (GANs) have remarkably advanced in diverse domains, especially image generation and editing. However, the misuse of GANs for generating deceptive images, such as face replacement, raises significant security concerns, which have gained widespread attention. Therefore, it is urgent to develop effective detection methods to distinguish between real and fake images. Current research centers around the application of transfer learning. Nevertheless, it encounters challenges such as knowledge forgetting from the original dataset and inadequate performance when dealing with imbalanced data during training. To alleviate this issue, this paper introduces a novel GAN-generated image detection algorithm called X-Transfer, which enhances transfer learning by utilizing two neural networks that employ interleaved parallel gradient transmission. In addition, we combine AUC loss and cross-entropy loss to improve the model’s performance. We carry out comprehensive experiments on multiple facial image datasets. The results show that our model outperforms the general transferring approach, and the best metric achieves 99.04%, which is increased by approximately 10%. Furthermore, we demonstrate excellent performance on non-face datasets, validating its generality and broader application prospects. Shu Hu 0001, Bin B. Zhu, Chingsheng Lin, Xi Wu 0004, Jinrong Hu, Xin Wang 0045 |
IJCNN | 8 |
| 2024 | Robustly Optimized Deep Feature Decoupling Network for Fatty Liver Diseases Detection
Shu Hu 0001, Bo Peng 0006, Jiashu Zhang, Xi Wu 0004, Xin Wang 0045 |
MICCAI (1) | 6 |
| 2024 | Semi-Supervised Thyroid Nodule Detection in Ultrasound VideosabstractDeep learning techniques have been investigated for the computer-aided diagnosis of thyroid nodules in ultrasound images. However, most existing thyroid nodule detection methods were simply based on static ultrasound images, which cannot well explore spatial and temporal information following the clinical examination process. In this paper, we propose a novel video-based semi-supervised framework for ultrasound thyroid nodule detection. Especially, considering clinical examinations that need to detect thyroid nodules at the ultrasonic probe positions, we first construct an adjacent frame guided detection backbone network by using adjacent supporting reference frames. To further reduce the labour-intensive thyroid nodule annotation in ultrasound videos, we extend the video-based detection in a semi-supervised manner by using both labeled and unlabeled videos. Based on the detection consistency in sequential neighbouring frames, a pseudo label adaptation strategy is proposed for the refinement of unpredicted frames. The proposed framework is validated on 996 transverse viewed and 1088 longitudinal viewed ultrasound videos. Experimental results demonstrated the superior performance of our proposed method in the ultrasound video-based detection of thyroid nodules. Zhongyu Li 0002, Canhua Xu, Bite Zhang, Jihua Zhu, Xin Wang 0045, Meng Yang 0026, Shi Chang |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Outlier Robust Adversarial Training
Shu Hu 0001, Zhenhuan Yang, Xin Wang 0045, Yiming Ying, Siwei Lyu |
ACML | 3 |
| 2023 | GAN-Generated Faces Detection: A Survey and New PerspectivesabstractGenerative Adversarial Networks (GAN) have led to the generation of very realistic face images, which have been used in fake social media accounts and other disinformation matters that can generate profound impacts. Therefore, the corresponding GAN-face detection techniques are under active development that can examine and expose such fake faces. In this work, we aim to provide a comprehensive review of recent progress in GAN-face detection. We focus on methods that can detect face images that are generated or synthesized from GAN models. We classify the existing detection works into four categories: (1) deep learning-based, (2) physical-based, (3) physiological-based methods, and (4) evaluation and comparison against human visual performance. For each category, we summarize the key ideas and connect them with method implementations. We also discuss open problems and suggest future research directions. Xin Wang 0045, Shu Hu 0001, Ming-Ching Chang, Siwei Lyu |
ECAI | 1 |
| 2023 | Detection of Real-Time Deepfakes in Video Conferencing with Active Probing and Corneal ReflectionabstractThe COVID pandemic has led to the wide adoption of online video calls in recent years. However, the increasing reliance on video calls provides opportunities for new impersonation attacks by fraudsters using the advanced real-time DeepFakes. Real-time DeepFakes pose new challenges to detection methods, which have to run in real-time as a video call is ongoing. In this paper, we describe a new active forensic method to detect real-time DeepFakes. Specifically, we authenticate video calls by displaying a distinct pattern on the screen and using the corneal reflection extracted from the images of the call participant’s face. This pattern can be induced by a call participant displaying on a shared screen or directly integrated into the video-call client. In either case, no specialized imaging or lighting hardware is required. Through large-scale simulations, we evaluate the reliability of this approach under a range in a variety of real-world imaging scenarios. Xin Wang 0045, Siwei Lyu |
ICASSP | 2 |
| 2023 | Controlling Neural Style Transfer with Deep Reinforcement LearningabstractControlling the degree of stylization in the Neural Style Transfer (NST) is a little tricky since it usually needs hand-engineering on hyper-parameters. In this paper, we propose the first deep Reinforcement Learning (RL) based architecture that splits one-step style transfer into a step-wise process for the NST task. Our RL-based method tends to preserve more details and structures of the content image in early steps, and synthesize more style patterns in later steps. It is a user-easily-controlled style-transfer method. Additionally, as our RL-based model performs the stylization progressively, it is lightweight and has lower computational complexity than existing one-step Deep Learning (DL) based models. Experimental results demonstrate the effectiveness and robustness of our method. Chengming Feng, Jing Hu 0009, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Xi Wu 0004, Hongtu Zhu, Siwei Lyu |
IJCAI | 3 |
| 2023 | RMBench: Benchmarking Deep Reinforcement Learning for Robotic Manipulator ControlabstractReinforcement learning is used to tackle complex tasks with high-dimensional sensory inputs. Over the past decade, a wide range of reinforcement learning algorithms have been developed, with recent progress benefiting from deep learning for raw sensory signal representation. This raises a natural question: how well do these algorithms perform across different robotic manipulation tasks? To objectively compare algorithms, benchmarks use performance metrics. Benchmarks use objective performance metrics to offer a scientific way to compare algorithms. In this paper, we introduce RMBench, the first benchmark for robotic manipulations with high-dimensional continuous action and state spaces. We implement and evaluate reinforcement learning algorithms that take observed pixels as inputs and report their average performance and learning curves to demonstrate their performance and training stability. Our study concludes that none of the evaluated algorithms can handle all tasks well, with soft Actor-Critic outperforming most algorithms in terms of average reward and stability, and an algorithm combined with data augmentation potentially facilitating learning policies. Our code is publicly available at https://github.com/xiangyanfei212/RMBench-2022.git, including all benchmark tasks and studied algorithms. Yanfei Xiang, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Xiaomeng Huang, Xi Wu 0004, Siwei Lyu |
IROS | 2 |
| 2023 | TransDose: Transformer-based radiotherapy dose prediction from CT images guided by super-pixel-level GCN classification
Zhengyang Jiao, Xingchen Peng, Yan Wang 0015, Jianghong Xiao, Dong Nie, Xi Wu 0004, Xin Wang 0045, Jiliu Zhou, Dinggang Shen |
Medical Image Anal. | 7 |
| 2023 | Rank-Based Decomposable Losses in Machine Learning: A SurveyabstractRecent works have revealed an essential paradigm in designing loss functions that differentiate individual losses versus aggregate losses. The individual loss measures the quality of the model on a sample, while the aggregate loss combines individual losses/scores over each training sample. Both have a common procedure that aggregates a set of individual values to a single numerical value. The ranking order reflects the most fundamental relation among individual values in designing losses. In addition, decomposability, in which a loss can be decomposed into an ensemble of individual terms, becomes a significant property of organizing losses/scores. This survey provides a systematic and comprehensive review of rank-based decomposable losses in machine learning. Specifically, we provide a new taxonomy of loss functions that follows the perspectives of aggregate loss and individual loss. We identify the aggregator to form such losses, which are examples of set functions. We organize the rank-based decomposable losses into eight categories. Following these categories, we review the literature on rank-based aggregate losses and rank-based individual losses. We describe general formulas for these losses and connect them with existing research topics. We also suggest future research directions spanning unexplored, remaining, and emerging issues in rank-based decomposable losses. Shu Hu 0001, Xin Wang 0045, Siwei Lyu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Stochastic Planner-Actor-Critic for Unsupervised Deformable Image RegistrationabstractLarge deformations of organs, caused by diverse shapes and nonlinear shape changes, pose a significant challenge for medical image registration. Traditional registration methods need to iteratively optimize an objective function via a specific deformation model along with meticulous parameter tuning, but which have limited capabilities in registering images with large deformations. While deep learning-based methods can learn the complex mapping from input images to their respective deformation field, it is regression-based and is prone to be stuck at local minima, particularly when large deformations are involved. To this end, we present Stochastic Planner-Actor-Critic (spac), a novel reinforcement learning-based framework that performs step-wise registration. The key notion is warping a moving image successively by each time step to finally align to a fixed image. Considering that it is challenging to handle high dimensional continuous action and state spaces in the conventional reinforcement learning (RL) framework, we introduce a new concept `Plan' to the standard Actor-Critic model, which is of low dimension and can facilitate the actor to generate a tractable high dimensional action. The entire framework is based on unsupervised training and operates in an end-to-end manner. We evaluate our method on several 2D and 3D medical image datasets, some of which contain large deformations. Our empirical results highlight that our work achieves consistent, significant gains and outperforms state-of-the-art methods. Ziwei Luo 0002, Jing Hu 0009, Xin Wang 0045, Shu Hu 0001, Bin Kong 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
AAAI | 3 |
| 2022 | Synergistic Network Learning and Label Correction for Noise-Robust Image ClassificationabstractLarge training datasets almost always contain examples with inaccurate or incorrect labels. Deep Neural Networks (DNNs) tend to overfit training label noise, resulting in poorer model performance in practice. To address this problem, we propose a robust label correction framework combining the ideas of small loss selection and noise correction, which learns network parameters and reassigns ground truth labels iteratively. Taking the expertise of DNNs to learn meaningful patterns before fitting noise, our framework first trains two networks over the current dataset with small loss selection. Based on the classification loss and agreement loss of two networks, we can measure the confidence of training data. More and more confident samples are selected for label correction during the learning process. We demonstrate our method on both synthetic and real-world datasets with different noise types and rates, including CIFAR-10, CIFAR-100 and Clothing1M, where our method outperforms the baseline approaches. Bin Kong 0001, Eric J. Seibel, Xin Wang 0045, Youbing Yin, Qi Song 0001 |
ICASSP | 4 |
| 2022 | Eyes Tell All: Irregular Pupil Shapes Reveal GAN-Generated FacesabstractGenerative adversarial network (GAN) generated high-realistic human faces are visually challenging to discern from real ones. They have been used as profile images for fake social media accounts, which leads to high negative social impacts. In this work, we show that GAN-generated faces can be exposed via irregular pupil shapes. This phenomenon is caused by the lack of physiological constraints in the GAN models. We demonstrate that such artifacts exist widely in high-quality GAN-generated faces. We design an automatic method to segment the pupils from the eyes and analyze their shapes to distinguish GAN-generated faces from real ones. Qualitative and quantitative evaluations of our method on the Flickr-Faces-HQ dataset and a StyleGAN2 generated face dataset demonstrate the effectiveness and simplicity of our method. Shu Hu 0001, Xin Wang 0045, Ming-Ching Chang, Siwei Lyu |
ICASSP | 3 |
| 2022 | Contrastive Class-Specific Encoding for Few-Shot Object DetectionabstractIn this paper, we propose a new few-shot object detection (FSOD) framework that introduces a new contrastive branch to extract the class representation of images, which improves the generalization performance of the detection model for novel classes. Additionally, we investigate the effectiveness of both self-supervised and supervised contrastive losses for class-specific encoding in our framework. Experimental results on the benchmark datasets indicate that our proposed method archives the state-of-the-art performance compared with existing FSOD methods. Dizhong Lin, Ying Fu 0003, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
ICME | 3 |
| 2022 | Sum of Ranked Range Loss for Supervised LearningabstractIn forming learning objectives, one oftentimes needs to aggregate a set of individual values to a single output. Such cases occur in the aggregate loss, which combines individual losses of a learning model over each training sample, and in the individual loss for multi-label learning, which combines prediction scores over all class labels. In this work, we introduce the sum of ranked range (SoRR) as a general approach to form learning objectives. A ranked range is a consecutive sequence of sorted values of a set of real numbers. The minimization of SoRR is solved with the difference of convex algorithm (DCA). We explore two applications in machine learning of the minimization of the SoRR framework, namely the AoRR aggregate loss for binary/multi-class classification at the sample level and the TKML individual loss for multi-label/multi-class classification at the label level. A combination loss of AoRR and TKML is proposed as a new learning objective for improving the robustness of multi-label learning in the face of outliers in sample and labels alike. Our empirical results highlight the effectiveness of the proposed optimization frameworks and demonstrate the applicability of proposed losses using synthetic and real data sets. Shu Hu 0001, Yiming Ying, Xin Wang 0045, Siwei Lyu |
J. Mach. Learn. Res. | 3 |
| 2022 | Learning a deep dual-level network for robust DeepFake detection
Wenbo Pu, Jing Hu 0009, Xin Wang 0045, Yuezun Li, Shu Hu 0001, Bin B. Zhu, Rui Song 0006, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
Pattern Recognit. | 3 |
| 2022 | DE-GAN: Domain Embedded GAN for High Quality Face Image Inpainting
Xian Zhang 0008, Xin Wang 0045, Canghong Shi, Xiaojie Li 0001, Bin Kong 0001, Siwei Lyu, Bin B. Zhu, Jiancheng Lv 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004, Imran Mumtaz |
Pattern Recognit. | 2 |
| 2021 | NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and LocalizationabstractFor iris recognition in non-cooperative environments, iris segmentation has been regarded as the first most important challenge still open to the biometric community, affecting all downstream tasks from normalization to recognition. In recent years, deep learning technologies have gained significant popularity among various computer vision tasks and also been introduced in iris biometrics, especially iris segmentation. To investigate recent developments and attract more interest of researchers in the iris segmentation method, we organized the 2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization (NIR-ISL 2021) at the 2021 International Joint Conference on Biometrics (IJCB 2021). The challenge was used as a public platform to assess the performance of iris segmentation and localization methods on Asian and African NIR iris images captured in non-cooperative environments. The three best-performing entries achieved solid and satisfactory iris segmentation and localization results in most cases, and their code and models have been made publicly available for reproducibility research. Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Jawad Muhammad, Qi Zhang 0015, Qichuan Tian, Zhaofeng He 0001, Zhenan Sun, Tianbao Liu, Wei Yang 0006, Dongliang Wu, Yingfeng Liu, Ruiye Zhou, Huihai Wu, Junbao Wang, Wantong Xiong, Xueyu Shi, Shao Zeng, Peihua Li, Huijie Wu, Xinhui Zhang, Menghan Zhang, Fadi Boutros, Naser Damer, Arjan Kuijper, Juan E. Tapia, Andres Valenzuela, Christoph Busch 0001, Gourav Gupta, Kiran B. Raja, Xi Wu 0004, Xiaojie Li 0001, Jingfu Yang, Hongyan Jing, Xin Wang 0045, Bin Kong 0001, Youbing Yin, Qi Song 0001, Siwei Lyu, Shu Hu 0001, Leon Premk, Matej Vitek, Vitomir Struc, Peter Peer, Jalil Nourmohammadi-Khiarak, Farhang Jaryani, Samaneh Salehi Nasab, Seyed Naeim Moafinejad, Yasin Amini, Morteza Noshad |
IJCB | 48 |
| 2021 | TkML-AP: Adversarial Attacks to Top-k Multi-Label LearningabstractTop-k multi-label learning, which returns the top-k predicted labels from an input, has many practical applications such as image annotation, document analysis, and web search engine. However, the vulnerabilities of such algorithms with regards to dedicated adversarial perturbation attacks have not been extensively studied previously. In this work, we develop methods to create adversarial perturbations that can be used to attack top-k multi-label learning-based image annotation systems (TkML-AP). Our methods explicitly consider the top-k ranking relation and are based on novel loss functions. Experimental evaluations on large-scale benchmark datasets including PASCAL VOC and MS COCO demonstrate the effectiveness of our methods in reducing the performance of state-of-the-art top-k multi-label learning methods, under both untargeted and targeted attacks. Shu Hu 0001, Lipeng Ke, Xin Wang 0045, Siwei Lyu |
ICCV | 3 |
| 2021 | Imperceptible Adversarial Examples For Fake Image DetectionabstractFooling people with highly realistic fake images generated with Deepfake or GANs brings a great social disturbance to our society. Many methods have been proposed to detect fake images, but they are vulnerable to adversarial perturbations – intentionally designed noises that can lead to the wrong prediction. Existing methods of attacking fake image detectors usually generate adversarial perturbations to perturb almost the entire image. This is redundant and increases the perceptibility of perturbations. In this paper, we propose a novel method to disrupt the fake image detection by determining key pixels to a fake image detector and attacking only the key pixels, which results in the L0and the L2norms of adversarial perturbations much less than those of existing works. Experiments on two public datasets with three fake image detectors indicate that our proposed method achieves state-of the-art performance in both white-box and black-box attacks. Quanyu Liao, Yuezun Li, Xin Wang 0045, Bin Kong 0001, Bin B. Zhu, Siwei Lyu, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
ICIP | 3 |
| 2021 | Transferable Adversarial Examples for Anchor Free Object DetectionabstractDeep neural networks have been demonstrated to be vulnerable to adversarial attacks: subtle perturbation can completely change prediction result. The vulnerability has led to a surge of research in this direction, including adversarial attacks on object detection networks. However, previous studies are dedicated to attacking anchor-based object detectors. In this paper, we present the first adversarial attack on anchor-free object detectors. It conducts category-wise, instead of previously instance-wise, attacks on object detectors, and leverages high-level semantic information to efficiently generate transferable adversarial examples, which can also be transferred to attack other object detectors, even anchor-based detectors such as Faster R-CNN. Experimental results on two benchmark datasets demonstrate that our proposed method achieves state-of-the-art performance and transferability. Quanyu Liao, Xin Wang 0045, Bin Kong 0001, Siwei Lyu, Bin B. Zhu, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
ICME | 2 |
| 2021 | Stochastic Actor-Executor-Critic for Image-to-Image TranslationabstractTraining a model-free deep reinforcement learning model to solve image-to-image translation is difficult since it involves high-dimensional continuous state and action spaces. In this paper, we draw inspiration from the recent success of the maximum entropy reinforcement learning framework designed for challenging continuous control problems to develop stochastic policies over high dimensional continuous spaces including image representation, generation, and control simultaneously. Central to this method is the Stochastic Actor-Executor-Critic (SAEC) which is an off-policy actor-critic model with an additional executor to generate realistic images. Specifically, the actor focuses on the high-level representation and control policy by a stochastic latent action, as well as explicitly directs the executor to generate low-level actions to manipulate the state. Experiments on several image-to-image translation tasks have demonstrated the effectiveness and robustness of the proposed SAEC when facing high-dimensional continuous space problems. Ziwei Luo 0002, Jing Hu 0009, Xin Wang 0045, Siwei Lyu, Bin Kong 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
IJCAI | 3 |
| 2021 | End-to-end multimodal image registration via reinforcement learning
Jing Hu 0009, Ziwei Luo 0002, Xin Wang 0045, Shanhui Sun, Youbing Yin, Kunlin Cao, Qi Song 0001, Siwei Lyu, Xi Wu 0004 |
Medical Image Anal. | 3 |
| 2020 | Fast Local Attack: Generating Local Adversarial Examples for Object DetectorsabstractThe deep neural network is vulnerable to adversarial examples. Adding imperceptible adversarial perturbations to images is enough to make them fail. Most existing research focuses on attacking image classifiers or anchor-based object detectors, but they generate globally perturbation on the whole image, which is unnecessary. In our work, we leverage higher-level semantic information to generate high aggressive local perturbations for anchor-free object detectors. As a result, it is less computationally intensive and achieves a higher black-box attack as well as transferring attack performance. The adversarial examples generated by our method are not only capable of attacking anchor-free object detectors, but also able to be transferred to attack anchor-based object detector. Quanyu Liao, Xin Wang 0045, Bin Kong 0001, Siwei Lyu, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
IJCNN | 2 |
| 2020 | Learning by Minimizing the Sum of Ranked RangeabstractIn forming learning objectives, one oftentimes needs to aggregate a set of individual values to a single output. Such cases occur in the aggregate loss, which combines individual losses of a learning model over each training sample, and in the individual loss for multi-label learning, which combines prediction scores over all class labels. In this work, we introduce the sum of ranked range (SoRR) as a general approach to form learning objectives. A ranked range is a consecutive sequence of sorted values of a set of real numbers. The minimization of SoRR is solved with the difference of convex algorithm (DCA). We explore two applications in machine learning of the minimization of the SoRR framework, namely the AoRR aggregate loss for binary classification and the TKML individual loss for multi-label/multi-class classification. Our empirical results highlight the effectiveness of the proposed optimization framework and demonstrate the applicability of proposed losses using synthetic and real datasets. Shu Hu 0001, Yiming Ying, Xin Wang 0045, Siwei Lyu |
NeurIPS | 3 |
| 2020 | Learning physical properties in complex visual scenes: An intelligent machine for perceiving blood flow dynamics from static CT angiography imaging
Zhifan Gao, Xin Wang 0045, Shanhui Sun, Dan Wu 0002, Youbing Yin, Xin Liu 0023, Heye Zhang, Victor Hugo C. de Albuquerque |
Neural Networks | 2 |
| 2019 | A Multi-modality Network for Cardiomyopathy Death Risk Prediction with CMR Images and Clinical Information
Chaoyang Xia, Xiaojie Li 0001, Xin Wang 0045, Bin Kong 0001, Yucheng Chen 0003, Youbing Yin, Kunlin Cao, Qi Song 0001, Siwei Lyu, Xi Wu 0004 |
MICCAI (2) | 3 |
| 2019 | ACNET: Attention-based Convolution Network with Additional Discriminative Features for DCM Classification (S)abstractFor dilated cardiomyopathy (DCM) patients, immediate emergency diagnosis and treatment are critical for life saving and later recovery.T1 mapping is a non-invasive and effective diagnostic imaging approach to detect DCM.However, it is a demanding and time-consuming approach.In this paper, we propose an attention-based network structure, which can automatically identify DCM patients in a speedy manner to prioritize their treatment.In the proposed method, we adopt attention modules to generate attention-aware features.Inside each attention module, a bottom-up top-down feed-forward structure is used to unfold the feed-forward and feed-back attention processes into a single feed-forward process.It allows the network to focus more on determining useful information about the current output that is significant in the input data.Moreover, inspired by the residual network idea, we make full use of the characteristics of the original data.Combined residual block, we design down-residual modules for classification tasks.It consists of seven convolution layers and three layers of residual blocks.Our network achieves the most advanced recognition performance on cardiac datasets.We evaluated our approach on CMR(cardiac magnetic resonance) T1 mapping images with lower PSNR(peak signal to noise ratio), and the results demonstrate that our architecture outperforms previous approaches. Xin Wang 0045, Xiaojie Li 0001, Yucheng Chen 0003, Jiliu Zhou, Kunlin Cao, Qi Song 0001, Xi Wu 0004, Youbing Yin |
SEKE | 2 |
| 2019 | Efficient algorithms for graph regularized PLSA for probabilistic topic modeling
Xin Wang 0045, Ming-Ching Chang, Siwei Lyu |
Pattern Recognit. | 1 |
| 2018 | Holistic and Deep Feature Pyramids for Saliency Detection
Shizhong Dong, Zhifan Gao, Shanhui Sun, Xin Wang 0045, Ming Li 0005, Heye Zhang, Guang Yang 0006, Huafeng Liu 0003, Shuo Li 0001 |
BMVC | 4 |
| 2018 | Invasive Cancer Detection Utilizing Compressed Convolutional Neural Network and Transfer Learning
Bin Kong 0001, Shanhui Sun, Xin Wang 0045, Qi Song 0001, Shaoting Zhang 0001 |
MICCAI (2) | 3 |
| 2016 | Co-Regularized PLSA for Multi-Modal LearningabstractMany learning problems in real world applications involve rich datasets comprising multiple information modalities. In this work, we study co-regularized PLSA(coPLSA) as an efficient solution to probabilistic topic analysis of multi-modal data. In coPLSA, similarities between topic compositions of a data entity across different data modalities are measured with divergences between discrete probabilities, which are incorporated as a co-regularizer to augment individual PLSA models over each data modality. We derive efficient iterative learning algorithms for coPLSA with symmetric KL, L2 and L1 divergences as co-regularizers, in each case the essential optimization problem affords simple numerical solutions that entail only matrix arithmetic operations and numerical solution of 1D nonlinear equations. We evaluate the performance of the coPLSA algorithms on text/image cross-modal retrieval tasks, on which they show competitive performance with state-of-the-art methods. Xin Wang 0045, Ming-Ching Chang, Yiming Ying, Siwei Lyu |
AAAI | 1 |
| 2015 | Fast Online Upper Body Pose Estimation from VideoabstractEstimation of human body poses from video is an important problem in computer vision with many applications. Most existing methods for video pose estimation are offline in nature, where all frames in the video are used in the process to estimate the body pose in each frame. In this work, we describe a fast online video upper body pose estimation method (CDBN-MODEC) that is based on a conditional dynamic Bayesian network model, which predicts upper body pose in a frame without using information from future frames. Our method combines fast single image based pose estimation methods with the temporal correlation of poses between frames. We collect a new high frame rate upper body pose dataset that better reflects practical scenarios calling for fast online video pose estimation. When evaluated on this dataset and the VideoPose2 benchmark dataset, CDBN-MODEC achieves improvements in both performance and running efficiency over several state-of-art online video pose estimation methods. Ming-Ching Chang, Honggang Qi, Xin Wang 0045, Hong Cheng 0002, Siwei Lyu |
BMVC | 3 |
| 2013 | On Algorithms for Sparse Multi-factor NMFabstractNonnegative matrix factorization (NMF) is a popular data analysis method, the objective of which is to decompose a matrix with all nonnegative components into the product of two other nonnegative matrices. In this work, we describe a new simple and efficient algorithm for multi-factor nonnegative matrix factorization problem ({mfNMF}), which generalizes the original NMF problem to more than two factors. Furthermore, we extend the mfNMF algorithm to incorporate a regularizer based on Dirichlet distribution over normalized columns to encourage sparsity in the obtained factors. Our sparse NMF algorithm affords a closed form and an intuitive interpretation, and is more efficient in comparison with previous works that use fix point iterations. We demonstrate the effectiveness and efficiency of our algorithms on both synthetic and real data sets. Siwei Lyu, Xin Wang 0045 |
NIPS | 2 |