VLDB 2026 Research / reviewers in the wild / expert
Shu Hu 0001
dblp:169/9795-1
· DBLP profile ↗
43ranked-venue papers
7as first author
41since 2021 · last 2026
0000-0003-1446-4140ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 6 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 24 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RL-I2IT: Image-to-image translation with deep reinforcement learning
Jing Hu 0009, Ziwei Luo 0002, Chengming Feng, Shu Hu 0001, Bin B. Zhu, Xi Wu 0004, Xin Li 0005, Hongtu Zhu, Siwei Lyu, Xin Wang 0045 |
Neural Networks | 4 |
| 2025 | Improving Generalization for AI-Synthesized Voice DetectionabstractAI-synthesized voice technology has the potential to create realistic human voices for beneficial applications, but it can also be misused for malicious purposes. While existing AI-synthesized voice detection models excel in intra-domain evaluation, they face challenges in generalizing across different domains, potentially becoming obsolete as new voice generators emerge. Current solutions use diverse data and advanced machine learning techniques (e.g., domain-invariant representation, self-supervised learning), but are limited by predefined vocoders and sensitivity to factors like background noise and speaker identity. In this work, we introduce an innovative disentanglement framework aimed at extracting domain-agnostic artifact features related to vocoders. Utilizing these features, we enhance model learning in a flat loss landscape, enabling escape from suboptimal solutions and improving generalization. Extensive experiments on benchmarks show our approach outperforms state-of-the-art methods, achieving up to 5.12% improvement in the equal error rate metric in intra-domain and 7.59% in cross-domain evaluations. Hainan Ren, Chun-Hao Liu, Shu Hu 0001 |
AAAI | 5 |
| 2025 | Preserving AUC Fairness in Learning with Noisy Protected GroupsabstractThe Area Under the ROC Curve (AUC) is a key metric for classification, especially under class imbalance, with growing research focus on optimizing AUC over accuracy in applications like medical image analysis and deepfake detection. This leads to fairness in AUC optimization becoming crucial as biases can impact protected groups. While various fairness mitigation techniques exist, fairness considerations in AUC optimization remain in their early stages, with most research focusing on improving AUC fairness under the
assumption of clean protected groups. However, these studies often overlook the impact of noisy protected groups, leading to fairness violations in practice. To address this, we propose the first robust AUC fairness approach under noisy protected groups with fairness theoretical guarantees using distributionally robust optimization. Extensive experiments on tabular and image datasets show that our method outperforms state-of-the-art approaches in preserving AUC fairness. The code is in https://github.com/Purdue-M2/AUC_Fairness_with_Noisy_Groups. Wenbin Zhang 0002, Xin Wang 0045, Zhenhuan Yang, Shu Hu 0001 |
ICML | 6 |
| 2025 | RLMiniStyler: Light-weight RL Style Agent for Arbitrary Sequential Neural Style GenerationabstractArbitrary style transfer aims to apply the style of any given artistic image to another content image. Still, existing deep learning-based methods often require significant computational costs to generate diverse stylized results. Motivated by this, we propose a novel reinforcement learning-based framework for arbitrary style transfer RLMiniStyler. This framework leverages a unified reinforcement learning policy to iteratively guide the style transfer process by exploring and exploiting stylization feedback, generating smooth sequences of stylized results while achieving model lightweight. Furthermore, we introduce an uncertainty-aware multi-task learning strategy that automatically adjusts loss weights to adapt to the content and style balance requirements at different training stages, thereby accelerating model convergence. Through a series of experiments across image various resolutions, we have validated the advantages of RLMiniStyler over other state-of-the-art methods in generating high-quality, diverse artistic image sequences at a lower cost. Codes are available at https://github.com/fengxiaoming520/RLMiniStyler. Jing Hu 0009, Chengming Feng, Shu Hu 0001, Ming-Ching Chang, Xin Li 0005, Xi Wu 0004, Xin Wang 0045 |
IJCAI | 3 |
| 2025 | Towards Fairness with Limited Demographics via Disentangled LearningabstractFairness in artificial intelligence has garnered increasing attention due to concerns about discriminatory AI-based decision-making, prompting the development of numerous mitigation approaches. However, most existing methods assume that demographic information is readily available, which may not align with real-world scenarios where such information is often incomplete. To this end, this paper tackles the pervasive yet overlooked challenge of developing fair machine learning algorithms with limited demographics. Specifically, we explore leveraging limited demographic information to accurately infer missing demographics while simultaneously evaluating and optimizing model fairness. We argue that this approach better aligns with common real-world socially sensitive scenarios involving limited demographics. Extensive experiments on three benchmark datasets highlight the effectiveness of the proposed method, surpassing state-of-the-art with significant gains in fairness while maintaining comparable utility. Zichong Wang, Anqi Wu, Nuno Moniz, Shu Hu 0001, Bart P. Knijnenburg, Xingquan Zhu 0001, Wenbin Zhang 0002 |
IJCAI | 4 |
| 2025 | Improving Generalization of Medical Image Registration Foundation ModelabstractDeformable registration is a fundamental task in medical image processing, aiming to achieve precise alignment by establishing nonlinear correspondences between images. Traditional methods offer good adaptability and interpretability but are limited by computational efficiency. Although deep learning approaches have significantly improved registration speed and accuracy, they often lack flexibility and generalizability across different datasets and tasks. In recent years, foundation models have emerged as a promising direction, leveraging large and diverse datasets to learn universal features and transformation patterns for image registration, thus demonstrating strong cross-task transferability. However, these models still face challenges in generalization and robustness when encountering novel anatomical structures, varying imaging conditions, or unseen modalities. To address these limitations, this paper incorporates Sharpness-Aware Minimization (SAM) into foundation models to enhance their generalization and robustness in medical image registration. By optimizing the flatness of the loss landscape, SAM improves model stability across diverse data distributions and strengthens its ability to handle complex clinical scenarios. Experimental results show that foundation models integrated with SAM achieve significant improvements in cross-dataset registration performance, offering new insights for the advancement of medical image registration technology. Our code is available at https://github.com/Promise13/fm_sam. Jing Hu 0009, Kaiwei Yu, Hongjiang Xian, Shu Hu 0001 |
IJCNN | 4 |
| 2025 | LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language ModelsabstractAccurate and efficient question-answering systems are essential for high-quality patient care in the medical field. While Large Language Models (LLMs) have made remarkable strides across various domains, they still face challenges in medical question answering, particularly in understanding domain-specific terminology and performing complex reasoning, limiting their effectiveness in critical applications. To address this, we propose a multi-agent medical question-answering (MedQA) system incorporating similar case generation. We leverage the Llama3.1:70B model in a multi-agent architecture to enhance enhance zero-shot classification on the MedQA dataset, utilizing the model’s inherent medical knowledge and reasoning capabilities without additional training data. Experimental results show substantial gains over existing benchmark models, with improvements of 7% in both accuracy and F1-score across various medical QA tasks. Furthermore, we examine the model’s interpretability and reliability in addressing complex medical queries. This research not only offers a robust solution for medical question answering but also establishes a foundation for broader applications of LLMs in the medical domain. Yineng Chen, Chingsheng Lin, Shu Hu 0001, Jinrong Hu, Xi Wu 0004, Xin Wang 0045 |
IJCNN | 6 |
| 2025 | Rethinking Individual Fairness in Deepfake DetectionabstractGenerative AI models have substantially improved the realism of synthetic media, yet their misuse through sophisticated DeepFakes poses significant risks. Despite recent advances in deepfake detection, fairness remains inadequately addressed, enabling deepfake markers to exploit biases against specific populations. While previous studies have emphasized group-level fairness, individual fairness (i.e., ensuring similar predictions for similar individuals) remains largely unexplored. In this work, we identify for the first time that the original principle of individual fairness fundamentally fails in the context of deepfake detection, revealing a critical gap previously unexplored in the literature. To mitigate it, we propose the first generalizable framework that can be integrated into existing deepfake detectors to enhance individual fairness and generalization. Extensive experiments conducted on leading deepfake datasets demonstrate that our approach significantly improves individual fairness while maintaining robust detection performance, outperforming state-of-the-art methods. The code is available at: https://github.com/Purdue-M2/Individual-Fairness-Deepfake-Detection. Aryana Hou, Justin Li, Shu Hu 0001 |
ACM Multimedia | 4 |
| 2025 | Redefining Fairness: A Multi-dimensional Perspective and Integrated Evaluation Framework
Zichong Wang, Zhipeng Yin, Zhen Liu 0017, Roland H. C. Yap, Xiaocai Zhang, Shu Hu 0001, Wenbin Zhang 0002 |
ECML/PKDD (1) | 6 |
| 2025 | VB-KGN: Variational Bayesian Kernel Generation Networks for Motion Image DeblurringabstractMotion blur estimation is a critical and fundamental task in scene analysis and image restoration. While most state-of-the-art deep learning-based methods for single-image motion image deblurring focus on constructing deep networks or developing training strategies, the characterization of motion blur has received less attention. In this paper, we innovatively propose a non-parametric Variational Bayesian Kernel Generation Network (VB-KGN) for characterizing motion blur in a single image. To solve this model, we employ the variational inference framework to approximate the expected statistical distribution of motion blur images in a data-driven manner. The qualitative and quantitative evaluations of our experimental results demonstrate that our proposed model can generate highly accurate motion blur kernels, significantly improving motion image deblurring performance and substantially reducing the need for extensive training sample preprocessing for deblurring tasks. Ying Fu 0003, Xiaojie Li 0001, Xin Wang 0045, Xi Wu 0004, Shu Hu 0001, Siwei Lyu, Wei Liu 0044 |
IEEE Trans. Multim. | 6 |
| 2025 | Spotting the Fakes: A Deep Dive into GAN-Generated Face DetectionabstractGenerative Adversarial Networks (GANs) have enabled the creation of highly authentic facial images, which are increasingly used in deceptive social media profiles and other forms of disinformation, resulting in serious consequences. Significant progress has been made in developing GAN-generated face detection systems to identify these fake images. This study offers a comprehensive review of recent advancements in GAN-generated face detection, focusing on techniques that detect facial images generated by GAN models. We categorize detection methods into three groups: (1) deep learning-based approaches, (2) physics-based methods, and (3) physiology-based methods. We summarize key concepts in each category, connecting them to relevant implementations, datasets, and evaluation metrics. Additionally, we provide a comparative analysis between automated detection and human visual performance to highlight the strengths and weaknesses of both approaches. Furthermore, we review related surveys, including detecting morphed faces, manipulated faces, DeepFake, and faces generated by diffusion models. Finally, we discuss unresolved challenges and suggest potential directions for future research. Xin Wang 0045, Ting Yu Tsai, Shu Hu 0001, Ming-Ching Chang, Pradeep K. Atrey, Siwei Lyu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Enhancing Vehicle Re-identification and Matching for Weaving AnalysisabstractVehicle weaving on highways contributes to traffic congestion, raises safety issues, and underscores the need for sophisticated traffic management systems. Current tools are inadequate in offering precise and comprehensive data on lane-specific weaving patterns. This paper introduces an innovative method for collecting non-overlapping video data in weaving zones, enabling the generation of quantitative insights into lane-specific weaving behaviors. Our experimental results confirm the efficacy of this approach, delivering critical data that can assist transportation authorities in enhancing traffic control and roadway infrastructure. More information about our dataset can be found here: VeID-Weaving. Mei Qiu, Stanley Y. P. Chien, Lauren A. Christopher, Yaobin Chen, Shu Hu 0001 |
AVSS | 6 |
| 2024 | Robust CLIP-Based Detector for Exposing Diffusion Model-Generated ImagesabstractDiffusion models (DMs) have revolutionized image generation, producing high-quality images with applications spanning various fields. However, their ability to create hyper-realistic images poses significant challenges in distinguishing between real and synthetic content, raising concerns about digital authenticity and potential misuse in creating deepfakes. This work introduces a robust detection framework that integrates image and text features extracted by CLIP model with a Multilayer Perceptron (MLP) classifier. We propose a novel loss that can improve the detector’s robustness and handle imbalanced datasets. Additionally, we flatten the loss landscape during the model training to improve the detector’s generalization capabilities. The effectiveness of our method, which outperforms traditional detection techniques, is demonstrated through extensive experiments, underscoring its potential to set a new state-of-the-art approach in DM-generated image detection. The code is available at https://github.com/Purdue-M2/RobustDM_Generated_Image_Detection. Santosh, Irene Amerini, Xin Wang 0045, Shu Hu 0001 |
AVSS | 5 |
| 2024 | RepMedGraf: Re-parameterization Medical Generated Radiation Field for Improved 3D Image ReconstructionabstractIn the field of medical imaging, 3D image reconstruction has emerged as a crucial technique for accurate disease diagnosis. Neural Radiance Field (NeRF) has shown promise in generating high-quality 3D models through deep learning, but its application in medical imaging remains limited. This paper presents RepMedGraf, an improved model addressing the limitations of NeRF. RepMedGraf utilizes a generator based on lightweight RepVGG Blocks instead of MedNeRF model. By doing so, the risk of overfitting is reduced, and training efficiency is improved. The proposed method is trained on publicly available chest and knee datasets. Comparative evaluations are conducted based on the generated radiation fields, demonstrating the effectiveness and quality of the 3D models produced by RepMedGraf. The results highlight the potential of RepMedGraf as a valuable tool in medical imaging for enhanced diagnostic accuracy and improved patient care. Ruotong Sun, Fayang Liao, Qinrui Fan, Shu Hu 0001, Xin Wang 0045, Jing Hu 0009 |
AVSS | 7 |
| 2024 | CGD-Net: A Hybrid End-to-end Network with gating decoding for Liver Tumor Segmentation from CT ImagesabstractLiver tumor segmentation plays a crucial role in the diagnosis and treatment of hepatic lesions. However, accurate tumor segmentation remains a challenging task due to the fuzzy boundaries of liver tumors and the uncertainty in shape, size, and location. In this paper, we propose a new end-to-end segmentation network called CGD-Net, which incorporates Transformer and frequency-domain features into a convolutional network, and proposes a new decoder structure to automatically learn from Segmentation of liver tumors in CT images. The proposed CGD-Net consists of a Transformer encoder, a frequency domain information fusion module, a gated decoder and three skip connections. Using the powerful feature extraction capability of the Transformer encoder to extract multi-feature information.CGD(Control Gate Decoder)blocks gradually restore the feature information lost in the encoding process by emphasizing the original information.In order to fully utilize the information of the original image, three skip connections are used to connect each encoder layer and its corresponding decoder layer, and a FEM (Frequency-domain Enhance module)is built in the third skip connection to fuse frequency domain features. Experiments on the LiTS dataset validate that the proposed CGD-Net can effectively segment liver tumors from CT images in an end-to-end manner, with segmentation accuracy exceeding many existing methods. Xiaogang Zhu 0003, Ziqiu Liu, Ouyang Shaobo, Xin Wang 0045, Shu Hu 0001, Feng Ding 0007 |
AVSS | 6 |
| 2024 | QRNN-Transformer: Recognizing Textual EntailmentabstractIn recent years, the Transformer model based on the self-attention mechanism has made significant progress in natural language processing and has also been applied in the text implication recognition task, achieving excellent results. However, the Transformer model still has deficiencies in modeling local information in the text. To improve the Transformer model, the QRNN-Transformer was proposed, which uses the QRNN network to divide the input text sequence into local short sequences to capture the local information of the input text. The self-attention is improved by combining the gating mechanism to make the model select tasks-related words or features. Extensive experiments have demonstrated the QRNN-Transformer model can effectively improve the accuracy of entailment relationship recognition on both English and Chinese datasets.l Xiaogang Zhu 0003, Zhihan Yan, Wenzhan Wang, Shu Hu 0001, Xin Wang 0045, Chunnian Liu |
AVSS | 4 |
| 2024 | Preserving Fairness Generalization in Deepfake DetectionabstractAlthough effective deepfake detection models have been developed in recent years, recent studies have revealed that these models can result in unfair performance disparities among demographic groups, such as race and gender. This can lead to particular groups facing unfair targeting or exclusion from detection, potentially allowing misclassified deepfakes to manipulate public opinion and undermine trust in the model. The existing method for addressing this problem is providing a fair loss function. It shows good fairness performance for intra-domain evaluation but does not maintain fairness for cross-domain testing. This highlights the significance of fairness generalization in the fight against deepfakes. In this work, we propose the first method to address the fairness generalization problem in deepfake detection by simultaneously considering features, loss, and optimization aspects. Our method employs disentanglement learning to extract demographic and domain-agnostic forgery features, fusing them to encourage fair learning across a flattened loss landscape. Extensive experiments on prominent deepfake datasets demonstrate our method's effectiveness, surpassing state-of-the-art approaches in preserving fairness during cross-domain deepfake detection. The code is available at https://github.com/Purdue-M2/Fairness-Generalization. Xinan He, Yan Ju, Xin Wang 0045, Feng Ding 0007, Shu Hu 0001 |
CVPR | 6 |
| 2024 | Synthesizing Black-Box Anti-Forensics Deepfakes With High Visual QualityabstractDeepFake, an AI technology for creating facial forgeries, has garnered global attention. Amid such circumstances, forensics researchers focus on developing defensive algorithms to counter these threats. In contrast, there are techniques developed for enhancing the aggressiveness of DeepFake, e.g., through anti-forensics attacks, to disrupt forensic detectors. However, such attacks often sacrifice image visual quality for improved undetectability. To address this issue, we propose a method to generate novel adversarial sharpening masks for launching black-box anti-forensics attacks. Unlike many existing arts, with such perturbations injected, DeepFakes could achieve high anti-forensics performance while exhibiting pleasant sharpening visual effects. After experimental evaluations, we prove that the proposed method could successfully disrupt the state-of-the-art DeepFake detectors. Besides, compared with the images processed by existing DeepFake anti-forensics methods, the visual qualities of antiforensics DeepFakes rendered by the proposed method are significantly refined. Bing Fan, Shu Hu 0001, Feng Ding 0007 |
ICASSP | 2 |
| 2024 | Contextual Reinforcement Learning for Unsupervised Deformable Multimodal Medical Images RegistrationabstractMultimodal deformable image registration refers to the process of finding the spatial correspondence between pairs of images with multimodal and mapping them onto the same coordinate system. Most of the deep learning-based registration methods are one-shot registration, which is difficult to handle images with significant deformations or displacements. Reinforcement learning can handle these challenges by viewing registration as a strategic decision-making process which is a step-by-step registration. However, it faces challenges with high-dimensional and continuous deformation fields. To overcome this, we introduce a planner network that maps high-dimensional input state to low-dimensional plan, guiding the actor to generate continuous actions. In order to handle complex multi-modal registration, we propose a multi-frame plan module which encourages artificial agent to explicitly utilize the redundant states in the registration process and learn more accurate registration actions from the generated state frames. To facilitate the training and convergence of the model, we define an unsupervised reward function and incorporate spectral normalization layers. The entire framework is a fully unsupervised registration framework and training in an end-to-end manner. We evaluated our method on publicly available T1w and T2w brain datasets, and the results indicate that our method has excellent deformable registration capability for multimodal images. Hongjiang Xian, Zhikun Shuai, Jing Hu 0009, Shu Hu 0001 |
IJCB | 6 |
| 2024 | Masked Conditional Diffusion Model for Enhancing Deepfake DetectionabstractRecent studies on deepfake detection have achieved promising results when training and testing faces are from the same dataset. However, their results severely degrade when confronted with forged samples that the model has not yet seen during training. In this paper, deepfake data to help detect deepfakes. this paper present we put a new insight into diffusion model-based data augmentation, and propose a Masked Conditional Diffusion Model (MCDM) for enhancing deepfake detection. It generates a variety of forged faces from a masked pristine one, encouraging the deepfake detection model to learn generic and robust representations without overfitting to special artifacts. Extensive experiments demonstrate that forgery images generated with our method are of high quality and helpful to improve the performance of deepfake detection models. Tiewen Chen, Shanmin Yang, Shu Hu 0001, Zhenghan Fang, Ying Fu 0003, Xi Wu 0004, Xin Wang 0045 |
IJCNN | 3 |
| 2024 | Efficient Image Super-Resolution via Symmetric Visual Attention NetworkabstractIn recent years, efficient super-resolution research has focused on reducing model complexity and improving efficiency by leveraging deep small-kernel convolution, but it has the problem of a small receptive field, which leads to a limited ability of the network to reconstruct details. Large kernel convolution can provide a large receptive field and lead to a substantial enhancement in the quality of image reconstruction, but its computational cost is too high. To minimize the model’s parameter count and achieve efficient super-resolution reconstruction, this study introduces a symmetric visual attention network. The network decomposes the large kernel convolution into three different lightweight and efficient convolutions. It then forms a bottleneck structure by leveraging the varied receptive field sizes of these convolutions in combination. The attention mechanism is integrated to create a bottleneck attention module, enhancing the network’s feature awareness. Furthermore, the bottleneck attention modules are symmetrically arranged to construct a symmetric large kernel attention block, thereby further enhancing the network’s capability to extract deep features. The experimental results demonstrate that the proposed model achieves competitive quantitative metrics when compared to other lightweight super-resolution methods, and the details of the reconstructed images are enhanced. With only 183K parameters, the model achieves a lightweight yet high-quality super-resolution model, offering a novel solution approach for efficient super-resolution. Qinrui Fan, Chengxu Wu, Shu Hu 0001, Xi Wu 0004, Xin Wang 0001, Jing Hu 0009 |
IJCNN | 3 |
| 2024 | Uncertainty-Aware Explainable Recommendation with Large Language ModelsabstractProviding explanations within the recommendation system would boost user satisfaction and foster trust, especially by elaborating on the reasons for selecting recommended items tailored to the user. The predominant approach in this domain revolves around generating text-based explanations, with a notable emphasis on applying large language models (LLMs). However, refining LLMs for explainable recommendations proves impractical due to time constraints and computing resource limitations. As an alternative, the current approach involves training the prompt rather than the LLM. In this study, we developed a model that utilizes the ID vectors of user and item inputs as prompts for GPT-2. We employed a joint training mechanism within a multi-task learning framework to optimize both the recommendation task and explanation task. This strategy enables a more effective exploration of users’ interests, improving recommendation effectiveness and user satisfaction. Through the experiments, our method achieving 1.59 DIV, 0.57 USR and 0.41 FCR on the Yelp, TripAdvisor and Amazon dataset respectively, demonstrates superior performance over four SOTA methods in terms of explainability evaluation metric. In addition, we identified that the proposed model is able to ensure stable textual quality on the three public datasets. Yicui Peng, Chingsheng Lin, Guo Huang, Jinrong Hu, Bin Kong 0001, Shu Hu 0001, Xi Wu 0004, Xin Wang 0045 |
IJCNN | 8 |
| 2024 | X-Transfer: A Transfer Learning-Based Framework for GAN-Generated Fake Image DetectionabstractGenerative adversarial networks (GANs) have remarkably advanced in diverse domains, especially image generation and editing. However, the misuse of GANs for generating deceptive images, such as face replacement, raises significant security concerns, which have gained widespread attention. Therefore, it is urgent to develop effective detection methods to distinguish between real and fake images. Current research centers around the application of transfer learning. Nevertheless, it encounters challenges such as knowledge forgetting from the original dataset and inadequate performance when dealing with imbalanced data during training. To alleviate this issue, this paper introduces a novel GAN-generated image detection algorithm called X-Transfer, which enhances transfer learning by utilizing two neural networks that employ interleaved parallel gradient transmission. In addition, we combine AUC loss and cross-entropy loss to improve the model’s performance. We carry out comprehensive experiments on multiple facial image datasets. The results show that our model outperforms the general transferring approach, and the best metric achieves 99.04%, which is increased by approximately 10%. Furthermore, we demonstrate excellent performance on non-face datasets, validating its generality and broader application prospects. Shu Hu 0001, Bin B. Zhu, Chingsheng Lin, Xi Wu 0004, Jinrong Hu, Xin Wang 0045 |
IJCNN | 3 |
| 2024 | Robustly Optimized Deep Feature Decoupling Network for Fatty Liver Diseases Detection
Shu Hu 0001, Bo Peng 0006, Jiashu Zhang, Xi Wu 0004, Xin Wang 0045 |
MICCAI (1) | 2 |
| 2024 | Improving Fairness in Deepfake DetectionabstractDespite the development of effective deepfake detectors in recent years, recent studies have demonstrated that biases in the data used to train these detectors can lead to disparities in detection accuracy across different races and genders. This can result in different groups being unfairly targeted or excluded from detection, allowing undetected deepfakes to manipulate public opinion and erode trust in a deepfake detection model. While existing studies have focused on evaluating fairness of deepfake detectors, to the best of our knowledge, no method has been developed to encourage fairness in deepfake detection at the algorithm level. In this work, we make the first attempt to improve deepfake detection fairness by proposing novel loss functions that handle both the setting where demographic information (e.g., annotations of race and gender) is available as well as the case where this information is absent. Fundamentally, both approaches can be used to convert many existing deepfake detectors into ones that encourages fairness. Extensive experiments on four deepfake datasets and five deepfake detectors demonstrate the effectiveness and flexibility of our approach in improving deep-fake detection fairness. Our code is available at https://github.com/littlejuyan/DF_Fairness. Yan Ju, Shu Hu 0001, Shan Jia, George H. Chen, Siwei Lyu |
WACV | 2 |
| 2024 | Fairness in Survival Analysis with Distributionally Robust OptimizationabstractWe propose a general approach for encouraging fairness in survival analysis models that is based on minimizing a worst-case error across all subpopulations that are “large enough” (occurring with at least a user-specified probability threshold). This approach can be used to convert a wide variety of existing survival analysis models into ones that simultaneously encourage fairness, without requiring the user to specify which attributes or features to treat as sensitive in the training loss function. From a technical standpoint, our approach applies recent methodological developments of distributionally robust optimization (DRO) to survival analysis. The complication is that existing DRO theory uses a training loss function that decomposes across contributions of individual data points, i.e., any term that shows up in the loss function depends only on a single training point. This decomposition does not hold for commonly used survival loss functions, including for the standard Cox proportional hazards model, its deep neural network variants, and many other recently developed survival analysis models that use loss functions involving ranking or similarity score calculations. We address this technical hurdle using a sample splitting strategy. We demonstrate our sample splitting DRO approach by using it to create fair versions of a diverse set of existing survival analysis models including the classical Cox model (and its deep neural network variant DeepSurv), the discrete-time model DeepHit, and the neural ODE model SODEN. We also establish a finite-sample theoretical guarantee to show what our sample splitting DRO loss converges to. Specifically for the Cox model, we further derive an exact DRO approach that does not use sample splitting. For all the survival models that we convert into DRO variants, we show that the DRO variants often score better on recently established fairness metrics (without incurring a significant drop in accuracy) compared to existing survival analysis fairness regularization techniques, including ones which directly use sensitive demographic information in their training loss functions. Shu Hu 0001, George H. Chen |
J. Mach. Learn. Res. | 1 |
| 2023 | Outlier Robust Adversarial Training
Shu Hu 0001, Zhenhuan Yang, Xin Wang 0045, Yiming Ying, Siwei Lyu |
ACML | 1 |
| 2023 | GAN-Generated Faces Detection: A Survey and New PerspectivesabstractGenerative Adversarial Networks (GAN) have led to the generation of very realistic face images, which have been used in fake social media accounts and other disinformation matters that can generate profound impacts. Therefore, the corresponding GAN-face detection techniques are under active development that can examine and expose such fake faces. In this work, we aim to provide a comprehensive review of recent progress in GAN-face detection. We focus on methods that can detect face images that are generated or synthesized from GAN models. We classify the existing detection works into four categories: (1) deep learning-based, (2) physical-based, (3) physiological-based methods, and (4) evaluation and comparison against human visual performance. For each category, we summarize the key ideas and connect them with method implementations. We also discuss open problems and suggest future research directions. Xin Wang 0045, Shu Hu 0001, Ming-Ching Chang, Siwei Lyu |
ECAI | 3 |
| 2023 | Controlling Neural Style Transfer with Deep Reinforcement LearningabstractControlling the degree of stylization in the Neural Style Transfer (NST) is a little tricky since it usually needs hand-engineering on hyper-parameters. In this paper, we propose the first deep Reinforcement Learning (RL) based architecture that splits one-step style transfer into a step-wise process for the NST task. Our RL-based method tends to preserve more details and structures of the content image in early steps, and synthesize more style patterns in later steps. It is a user-easily-controlled style-transfer method. Additionally, as our RL-based model performs the stylization progressively, it is lightweight and has lower computational complexity than existing one-step Deep Learning (DL) based models. Experimental results demonstrate the effectiveness and robustness of our method. Chengming Feng, Jing Hu 0009, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Xi Wu 0004, Hongtu Zhu, Siwei Lyu |
IJCAI | 4 |
| 2023 | RMBench: Benchmarking Deep Reinforcement Learning for Robotic Manipulator ControlabstractReinforcement learning is used to tackle complex tasks with high-dimensional sensory inputs. Over the past decade, a wide range of reinforcement learning algorithms have been developed, with recent progress benefiting from deep learning for raw sensory signal representation. This raises a natural question: how well do these algorithms perform across different robotic manipulation tasks? To objectively compare algorithms, benchmarks use performance metrics. Benchmarks use objective performance metrics to offer a scientific way to compare algorithms. In this paper, we introduce RMBench, the first benchmark for robotic manipulations with high-dimensional continuous action and state spaces. We implement and evaluate reinforcement learning algorithms that take observed pixels as inputs and report their average performance and learning curves to demonstrate their performance and training stability. Our study concludes that none of the evaluated algorithms can handle all tasks well, with soft Actor-Critic outperforming most algorithms in terms of average reward and stability, and an algorithm combined with data augmentation potentially facilitating learning policies. Our code is publicly available at https://github.com/xiangyanfei212/RMBench-2022.git, including all benchmark tasks and studied algorithms. Yanfei Xiang, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Xiaomeng Huang, Xi Wu 0004, Siwei Lyu |
IROS | 3 |
| 2023 | Rank-Based Decomposable Losses in Machine Learning: A SurveyabstractRecent works have revealed an essential paradigm in designing loss functions that differentiate individual losses versus aggregate losses. The individual loss measures the quality of the model on a sample, while the aggregate loss combines individual losses/scores over each training sample. Both have a common procedure that aggregates a set of individual values to a single numerical value. The ranking order reflects the most fundamental relation among individual values in designing losses. In addition, decomposability, in which a loss can be decomposed into an ensemble of individual terms, becomes a significant property of organizing losses/scores. This survey provides a systematic and comprehensive review of rank-based decomposable losses in machine learning. Specifically, we provide a new taxonomy of loss functions that follows the perspectives of aggregate loss and individual loss. We identify the aggregator to form such losses, which are examples of set functions. We organize the rank-based decomposable losses into eight categories. Following these categories, we review the literature on rank-based aggregate losses and rank-based individual losses. We describe general formulas for these losses and connect them with existing research topics. We also suggest future research directions spanning unexplored, remaining, and emerging issues in rank-based decomposable losses. Shu Hu 0001, Xin Wang 0045, Siwei Lyu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Stochastic Planner-Actor-Critic for Unsupervised Deformable Image RegistrationabstractLarge deformations of organs, caused by diverse shapes and nonlinear shape changes, pose a significant challenge for medical image registration. Traditional registration methods need to iteratively optimize an objective function via a specific deformation model along with meticulous parameter tuning, but which have limited capabilities in registering images with large deformations. While deep learning-based methods can learn the complex mapping from input images to their respective deformation field, it is regression-based and is prone to be stuck at local minima, particularly when large deformations are involved. To this end, we present Stochastic Planner-Actor-Critic (spac), a novel reinforcement learning-based framework that performs step-wise registration. The key notion is warping a moving image successively by each time step to finally align to a fixed image. Considering that it is challenging to handle high dimensional continuous action and state spaces in the conventional reinforcement learning (RL) framework, we introduce a new concept `Plan' to the standard Actor-Critic model, which is of low dimension and can facilitate the actor to generate a tractable high dimensional action. The entire framework is based on unsupervised training and operates in an end-to-end manner. We evaluate our method on several 2D and 3D medical image datasets, some of which contain large deformations. Our empirical results highlight that our work achieves consistent, significant gains and outperforms state-of-the-art methods. Ziwei Luo 0002, Jing Hu 0009, Xin Wang 0045, Shu Hu 0001, Bin Kong 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
AAAI | 4 |
| 2022 | Eyes Tell All: Irregular Pupil Shapes Reveal GAN-Generated FacesabstractGenerative adversarial network (GAN) generated high-realistic human faces are visually challenging to discern from real ones. They have been used as profile images for fake social media accounts, which leads to high negative social impacts. In this work, we show that GAN-generated faces can be exposed via irregular pupil shapes. This phenomenon is caused by the lack of physiological constraints in the GAN models. We demonstrate that such artifacts exist widely in high-quality GAN-generated faces. We design an automatic method to segment the pupils from the eyes and analyze their shapes to distinguish GAN-generated faces from real ones. Qualitative and quantitative evaluations of our method on the Flickr-Faces-HQ dataset and a StyleGAN2 generated face dataset demonstrate the effectiveness and simplicity of our method. Shu Hu 0001, Xin Wang 0045, Ming-Ching Chang, Siwei Lyu |
ICASSP | 2 |
| 2022 | Contrastive Class-Specific Encoding for Few-Shot Object DetectionabstractIn this paper, we propose a new few-shot object detection (FSOD) framework that introduces a new contrastive branch to extract the class representation of images, which improves the generalization performance of the detection model for novel classes. Additionally, we investigate the effectiveness of both self-supervised and supervised contrastive losses for class-specific encoding in our framework. Experimental results on the benchmark datasets indicate that our proposed method archives the state-of-the-art performance compared with existing FSOD methods. Dizhong Lin, Ying Fu 0003, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
ICME | 4 |
| 2022 | Differentially private SGDA for minimax problemsabstractStochastic gradient descent ascent (SGDA) and its variants have been the workhorse for solving minimax problems. However, in contrast to the well-studied stochastic gradient descent (SGD) with differential privacy (DP) constraints, there is little work on understanding the generalization (utility) of SGDA with DP constraints. In this paper, we use the algorithmic stability approach to establish the generalization (utility) of DP-SGDA in different settings. In particular, for the convex-concave setting, we prove that the DP-SGDA can achieve an optimal utility rate in terms of the weak primal-dual population risk in both smooth and non-smooth cases. To our best knowledge, this is the first-ever-known result for DP-SGDA in the non-smooth case. We further provide its utility analysis in the nonconvex-strongly-concave setting which is the first-ever-known result in terms of the primal population risk. The convergence and generalization results for this nonconvex setting are new even in the non-private setting. Finally, numerical experiments are conducted to demonstrate the effectiveness of DP-SGDA for both convex and nonconvex cases. Zhenhuan Yang, Shu Hu 0001, Yunwen Lei, Kush R. Varshney, Siwei Lyu, Yiming Ying |
UAI | 2 |
| 2022 | Network traffic analysis over clustering-based collective anomaly detection
Chonghua Wang, Zhiqiang Hao, Shu Hu 0001, Bo Jiang 0013, Xuehong Chen |
Comput. Networks | 4 |
| 2022 | Sum of Ranked Range Loss for Supervised LearningabstractIn forming learning objectives, one oftentimes needs to aggregate a set of individual values to a single output. Such cases occur in the aggregate loss, which combines individual losses of a learning model over each training sample, and in the individual loss for multi-label learning, which combines prediction scores over all class labels. In this work, we introduce the sum of ranked range (SoRR) as a general approach to form learning objectives. A ranked range is a consecutive sequence of sorted values of a set of real numbers. The minimization of SoRR is solved with the difference of convex algorithm (DCA). We explore two applications in machine learning of the minimization of the SoRR framework, namely the AoRR aggregate loss for binary/multi-class classification at the sample level and the TKML individual loss for multi-label/multi-class classification at the label level. A combination loss of AoRR and TKML is proposed as a new learning objective for improving the robustness of multi-label learning in the face of outliers in sample and labels alike. Our empirical results highlight the effectiveness of the proposed optimization frameworks and demonstrate the applicability of proposed losses using synthetic and real data sets. Shu Hu 0001, Yiming Ying, Xin Wang 0045, Siwei Lyu |
J. Mach. Learn. Res. | 1 |
| 2022 | Learning a deep dual-level network for robust DeepFake detection
Wenbo Pu, Jing Hu 0009, Xin Wang 0045, Yuezun Li, Shu Hu 0001, Bin B. Zhu, Rui Song 0006, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
Pattern Recognit. | 5 |
| 2021 | Exposing GAN-Generated Faces Using Inconsistent Corneal Specular HighlightsabstractSophisticated generative adversary network (GAN) models are now able to synthesize highly realistic human faces that are difficult to discern from real ones visually. In this work, we show that GAN synthesized faces can be exposed with the inconsistent corneal specular highlights between two eyes. The inconsistency is caused by the lack of physical/physiological constraints in the GAN models. We show that such artifacts exist widely in high-quality GAN synthesized faces and further describe an automatic method to extract and compare corneal specular highlights from two eyes. Qualitative and quantitative evaluations of our method suggest its simplicity and effectiveness in distinguishing GAN synthesized faces. Shu Hu 0001, Yuezun Li, Siwei Lyu |
ICASSP | 1 |
| 2021 | NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and LocalizationabstractFor iris recognition in non-cooperative environments, iris segmentation has been regarded as the first most important challenge still open to the biometric community, affecting all downstream tasks from normalization to recognition. In recent years, deep learning technologies have gained significant popularity among various computer vision tasks and also been introduced in iris biometrics, especially iris segmentation. To investigate recent developments and attract more interest of researchers in the iris segmentation method, we organized the 2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization (NIR-ISL 2021) at the 2021 International Joint Conference on Biometrics (IJCB 2021). The challenge was used as a public platform to assess the performance of iris segmentation and localization methods on Asian and African NIR iris images captured in non-cooperative environments. The three best-performing entries achieved solid and satisfactory iris segmentation and localization results in most cases, and their code and models have been made publicly available for reproducibility research. Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Jawad Muhammad, Qi Zhang 0015, Qichuan Tian, Zhaofeng He 0001, Zhenan Sun, Tianbao Liu, Wei Yang 0006, Dongliang Wu, Yingfeng Liu, Ruiye Zhou, Huihai Wu, Junbao Wang, Wantong Xiong, Xueyu Shi, Shao Zeng, Peihua Li, Huijie Wu, Xinhui Zhang, Menghan Zhang, Fadi Boutros, Naser Damer, Arjan Kuijper, Juan E. Tapia, Andres Valenzuela, Christoph Busch 0001, Gourav Gupta, Kiran B. Raja, Xi Wu 0004, Xiaojie Li 0001, Jingfu Yang, Hongyan Jing, Xin Wang 0045, Bin Kong 0001, Youbing Yin, Qi Song 0001, Siwei Lyu, Shu Hu 0001, Leon Premk, Matej Vitek, Vitomir Struc, Peter Peer, Jalil Nourmohammadi-Khiarak, Farhang Jaryani, Samaneh Salehi Nasab, Seyed Naeim Moafinejad, Yasin Amini, Morteza Noshad |
IJCB | 53 |
| 2021 | TkML-AP: Adversarial Attacks to Top-k Multi-Label LearningabstractTop-k multi-label learning, which returns the top-k predicted labels from an input, has many practical applications such as image annotation, document analysis, and web search engine. However, the vulnerabilities of such algorithms with regards to dedicated adversarial perturbation attacks have not been extensively studied previously. In this work, we develop methods to create adversarial perturbations that can be used to attack top-k multi-label learning-based image annotation systems (TkML-AP). Our methods explicitly consider the top-k ranking relation and are based on novel loss functions. Experimental evaluations on large-scale benchmark datasets including PASCAL VOC and MS COCO demonstrate the effectiveness of our methods in reducing the performance of state-of-the-art top-k multi-label learning methods, under both untargeted and targeted attacks. Shu Hu 0001, Lipeng Ke, Xin Wang 0045, Siwei Lyu |
ICCV | 1 |
| 2020 | Learning by Minimizing the Sum of Ranked RangeabstractIn forming learning objectives, one oftentimes needs to aggregate a set of individual values to a single output. Such cases occur in the aggregate loss, which combines individual losses of a learning model over each training sample, and in the individual loss for multi-label learning, which combines prediction scores over all class labels. In this work, we introduce the sum of ranked range (SoRR) as a general approach to form learning objectives. A ranked range is a consecutive sequence of sorted values of a set of real numbers. The minimization of SoRR is solved with the difference of convex algorithm (DCA). We explore two applications in machine learning of the minimization of the SoRR framework, namely the AoRR aggregate loss for binary classification and the TKML individual loss for multi-label/multi-class classification. Our empirical results highlight the effectiveness of the proposed optimization framework and demonstrate the applicability of proposed losses using synthetic and real datasets. Shu Hu 0001, Yiming Ying, Xin Wang 0045, Siwei Lyu |
NeurIPS | 1 |
| 2019 | A Fast PC Algorithm for High Dimensional Causal Discovery with Multi-Core PCsabstractDiscovering causal relationships from observational data is a crucial problem and it has applications in many research areas. The PC algorithm is the state-of-the-art constraint based method for causal discovery. However, runtime of the PC algorithm, in the worst-case, is exponential to the number of nodes (variables), and thus it is inefficient when being applied to high dimensional data, e.g., gene expression datasets. On another note, the advancement of computer hardware in the last decade has resulted in the widespread availability of multi-core personal computers. There is a significant motivation for designing a parallelized PC algorithm that is suitable for personal computers and does not require end users' parallel computing knowledge beyond their competency in using the PC algorithm. In this paper, we develop parallel-PC, a fast and memory efficient PC algorithm using the parallel computing technique. We apply our method to a range of synthetic and real-world high dimensional datasets. Experimental results on a dataset from the DREAM 5 challenge show that the original PC algorithm could not produce any results after running more than 24 hours; meanwhile, our parallel-PC algorithm managed to finish within around 12 hours with a 4-core CPU computer, and less than six hours with a 8-core CPU computer. Furthermore, we integrate parallel-PC into a causal inference method for inferring miRNA-mRNA regulatory relationships. The experimental results show that parallel-PC helps improve both the efficiency and accuracy of the causal inference algorithm. Thuc Duy Le, Tao Hoang, Jiuyong Li, Lin Liu 0003, Huawen Liu, Shu Hu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |