EDBT 2026 Demo / reviewers in the wild / expert
Siwei Lyu
dblp:51/4482
· DBLP profile ↗
186ranked-venue papers
22as first author
86since 2021 · last 2026
0000-0002-0992-685XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 115 · 16 first-author · 48 since 2021Graphics, computer vision, multimedia, augmented reality and games · 113 · 8 first-author · 53 since 2021Security and privacy · 11 · 1 first-author · 7 since 2021Systems, architecture and hardware · 6 · 1 since 2021Computer networks · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DICE: Distilling Classifier-Free Guidance into Text EmbeddingsabstractText-to-image diffusion models are capable of generating high-quality images, but suboptimal pre-trained text representations often result in these images failing to align closely with the given text prompts. Classifier-free guidance (CFG) is a popular and effective technique for improving text-image alignment in the generative process. However, CFG introduces significant computational overhead. In this paper, we present DIstilling CFG by sharpening text Embeddings (DICE) that replaces CFG in the sampling process with half the computational complexity while maintaining similar generation quality. DICE distills a CFG-based text-to-image diffusion model into a CFG-free version by refining text embeddings to replicate CFG-based directions. In this way, we avoid the computational drawbacks of CFG, enabling high-quality, well-aligned image generation at a fast sampling speed. Furthermore, examining the enhancement pattern, we identify the underlying mechanism of DICE that sharpens specific components of text embeddings to preserve semantic information while enhancing fine-grained details. Extensive experiments on multiple Stable Diffusion v1.5 variants, SDXL, and PixArt-\alpha demonstrate the effectiveness of our method. Defang Chen 0001, Can Wang 0001, Chun Chen 0001, Siwei Lyu |
AAAI | 5 |
| 2026 | ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speechabstractASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce the ASVspoof 5 database which is generated in a crowdsourced fashion from data collected in diverse acoustic conditions (cf. studio-quality data for earlier ASVspoof databases) and from ∼ 2,000 speakers (cf. ∼ 100 earlier). The database contains attacks generated with 32 different algorithms, also crowdsourced, and optimised to varying degrees using new surrogate detection models. Among them are attacks generated with a mix of legacy and contemporary text-to-speech synthesis and voice conversion models, in addition to adversarial attacks which are incorporated for the first time. ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They include two distinct partitions for the training of different sets of attack models, two more for the development and evaluation of surrogate detection models, and then three additional partitions which comprise the ASVspoof 5 training, development and evaluation sets. An auxiliary set of data collected from an additional 30k speakers can also be used to train speaker encoders for the implementation of attack algorithms. Also described herein is an experimental validation of the new ASVspoof 5 database using a set of automatic speaker verification and spoof/deepfake baseline detectors. With the exception of protocols and tools for the generation of spoofed/deepfake speech, the resources described in this paper, already used by participants of the ASVspoof 5 challenge in 2024, are now all freely available to the community. Xin Wang 0037, Héctor Delgado, Hemlata Tak, Jee-Weon Jung, Hye-Jin Shim, Massimiliano Todisco, Ivan Kukanov, Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen, Nicholas W. D. Evans, Kong-Aik Lee, Junichi Yamagishi, Myeonghun Jeong, Yongyi Zang, Soumi Maiti, Florian Lux, Nicolas Müller, Wangyou Zhang, Chengzhe Sun 0001, Shuwei Hou, Siwei Lyu, Sébastien Le Maguer, Hanjie Guo, Vishwanath Pratap Singh |
Comput. Speech Lang. | 24 |
| 2026 | Assessing the ability of neural TTS systems to model consonant-induced f0 perturbation
Chengzhe Sun 0001, Phil Rose, Cassandra L. Jacobs, Siwei Lyu |
Comput. Speech Lang. | 5 |
| 2026 | Attacks in Adversarial Machine Learning: A Systematic Survey from the Lifecycle Perspective
Baoyuan Wu, Zihao Zhu 0001, Li Liu 0036, Qingshan Liu 0001, Zhaofeng He 0001, Siwei Lyu |
Int. J. Comput. Vis. | 6 |
| 2026 | RL-I2IT: Image-to-image translation with deep reinforcement learning
Jing Hu 0009, Ziwei Luo 0002, Chengming Feng, Shu Hu 0001, Bin B. Zhu, Xi Wu 0004, Xin Li 0005, Hongtu Zhu, Siwei Lyu, Xin Wang 0045 |
Neural Networks | 9 |
| 2025 | A Landscape Survey of Tools and Real-World Cases of AI-Generated MediaabstractGenerative AI technologies have undergone significant and rapid advances in recent years. The impact of AI-generated media (AIGM) as a tool for disinformation, deception, and defamation has been widely reported. At the same time, numerous methods for detecting AIGM have emerged in research, leveraging data generated and collected in controlled environments. A comprehensive understanding of its real-world use cases is crucial to evaluating the negative impact of AIGM and developing counter technologies. More importantly, it is essential to assess the capabilities of publicly and commercially available tools currently in generating AIGM. This paper aims to address these concerns. First, we introduce a definitional framework categorizing different types of AIGM tasks. Based on this framework, we extensively survey existing tools capable of generating AIGM. Furthermore, drawing on our experience of interacting with practitioners and analyzing real-world misuse incidents using the deepfake-o-meter.org platform [9], we examine AIGM misuse and discuss the best practices identified through our investigations. Jialing Cai, Chengzhe Sun 0001, Soumyya Kanti Datta, Siwei Lyu |
AVSS | 4 |
| 2025 | HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language ModelsabstractWe introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a diverse range of conditional signals (e.g. text, objects, partial sequences). At its core, HOIGPT utilizes a large language model to predict the bidrectional transformation between HOI sequences and natural language descriptions. Given text inputs, HOIGPT generates a sequence of hand and object meshes; given (partial) HOI sequences, HOIGPT generates text descriptions and completes the sequences. To facilitate HOI understanding with a large language model, this paper introduces two key innovations: (1) a novel physically grounded HOI tokenizer, the hand-object decomposed VQ-VAE, for discretizing HOI sequences, and (2) a motion-aware language model trained to process and generate both text and HOI tokens. Extensive experiments demonstrate that HOIGPT sets new state-of-the-art performance on both text generation (+2.01% R Precision) and HOI generation (-2.56 FID) across multiple tasks and benchmarks. Mingzhen Huang, Fu-Jen Chu, Bugra Tekin, Kevin J. Liang, Weiyao Wang 0001, Pierre Gleize, Hongfei Xue, Siwei Lyu, Kris Makoto Kitani, Matt Feiszli |
CVPR | 10 |
| 2025 | Your Text Encoder Can Be an Object-Level Watermarking Controller
Naresh Kumar Devulapally, Mingzhen Huang, Vishal Asnani, Shruti Agarwal, Siwei Lyu, Vishnu Suresh Lokhande |
ICCV | 5 |
| 2025 | Knowledge Distillation with Refined LogitsabstractRecent research on knowledge distillation has increasingly focused on logit distillation because of its simplicity, effectiveness, and versatility in model compression. In this paper, we introduce Refined Logit Distillation (RLD) to address the limitations of current logit distillation methods. Our approach is motivated by the observation that even high-performing teacher models can make incorrect predictions, creating an exacerbated divergence between the standard distillation loss and the cross-entropy loss, which can undermine the consistency of the student model's learning objectives. Previous attempts to use labels to empirically correct teacher predictions may undermine the class correlations. In contrast, our RLD employs labeling information to dynamically refine teacher logits. In this way, our method can effectively eliminate misleading information from the teacher while preserving crucial class correlations, thus enhancing the value and efficiency of distilled knowledge. Experimental results on CIFAR-100 and ImageNet demonstrate its superiority over existing methods. Our code is available at https://github.com/zju-SWJ/RLD. Wujie Sun, Defang Chen 0001, Siwei Lyu, Genlang Chen, Chun Chen 0001, Can Wang 0001 |
ICCV | 3 |
| 2025 | Modality-Agnostic Deepfakes DetectionabstractAs AI-generated content (AIGC) thrives, deepfakes have expanded from single-modality falsification to cross-modal fake content creation, where either audio or visual components can be manipulated.While using two unimodal detectors can detect audio-visual deepfakes, cross-modal forgery clues could be overlooked.Existing multimodal deepfake detectors typically establish correspondence between the audio and visual modalities for binary real/fake classification and require the co-occurrence of both modalities.However, in real-world multi-modal applications, missing modality scenarios may occur where either modality is unavailable.In such cases, audio-visual detection methods are less practical than two independent unimodal methods.Consequently, the detector can not always obtain the number or type of manipulated modalities beforehand, necessitating a fake-modality-agnostic audio-visual detector.In this work, we introduce a comprehensive framework that is agnostic to fake modalities, which facilitates the identification of multimodal deepfakes and handles situations with missing modalities, regardless of the manipulations embedded in audio, video, or even cross-modal forms.To enhance the modeling of cross-modal forgery clues, we employ audio-visual speech recognition (AVSR) Jin Liu 0020, Jiao Dai, Xi Wang 0014, Shan Jia, Siwei Lyu, Jizhong Han |
IH&MMSec | 8 |
| 2025 | X2-DFD: A framework for explainable and extendable Deepfake DetectionabstractThis paper proposes **$\mathcal{X}^2$-DFD**, an **e$\mathcal{X}$plainable** and **e$\mathcal{X}$tendable** framework based on multimodal large-language models (MLLMs) for deepfake detection, consisting of three key stages.
The first stage, *Model Feature Assessment*, systematically evaluates the detectability of forgery-related features for the MLLM, generating a prioritized ranking of features based on their intrinsic importance to the model.
The second stage, *Explainable Dataset Construction*, consists of two key modules: *Strong Feature Strengthening*, which is designed to enhance the model’s existing detection and explanation capabilities by reinforcing its well-learned features, and *Weak Feature Supplementing*, which addresses gaps by integrating specific feature detectors (e.g., low-level artifact analyzers) to compensate for the MLLM’s limitations.
The third stage, Fine-tuning and Inference, involves fine-tuning the MLLM on the constructed dataset and deploying it for final detection and explanation.
By integrating these three stages, our approach enhances the MLLM's strengths while supplementing its weaknesses, ultimately improving both the detectability and explainability.
Extensive experiments and ablations, followed by a comprehensive human study, validate the improved performance of our approach compared to the original MLLMs.
More encouragingly, our framework is designed to be plug-and-play, allowing it to seamlessly integrate with future more advanced MLLMs and specific feature detectors, leading to continual improvement and extension to face the challenges of rapidly evolving deepfakes. Code can be found on https://github.com/chenyize111/X2DFD. Yize Chen, Zhiyuan Yan 0002, Kangran Zhao, Siwei Lyu, Baoyuan Wu |
NeurIPS | 5 |
| 2025 | Semantics-Oriented Multitask Learning for DeepFake Detection: A Joint Embedding ApproachabstractIn recent years, the multimedia forensics and security community has seen remarkable progress in multitask learning for DeepFake (i.e., face forgery) detection. The prevailing approach has been to frame DeepFake detection as a binary classification problem augmented by manipulation-oriented auxiliary tasks. This scheme focuses on learning features specific to face manipulations with limited generalizability. In this paper, we delve deeper into semantics-oriented multitask learning for DeepFake detection, capturing the relationships among face semantics via joint embedding. We first propose an automated dataset expansion technique that broadens current face forgery datasets to support semantics-oriented DeepFake detection tasks at both the global face attribute and local face region levels. Furthermore, we resort to the joint embedding of face images and labels (depicted by text descriptions) for prediction. This approach eliminates the need for manually setting task-agnostic and task-specific parameters, which is typically required when predicting multiple labels directly from images. In addition, we employ bi-level optimization to dynamically balance the fidelity loss weightings of various tasks, making the training process fully automated. Extensive experiments on six DeepFake datasets show that our method improves the generalizability of DeepFake detection and renders some degree of model interpretation by providing human-understandable explanations. Mian Zou, Baosheng Yu, Yibing Zhan, Siwei Lyu, Kede Ma |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face DetectionabstractFace-swapping DeepFakes have become an escalating societal concern, attracting increasing attention in recent years. To counter this, we investigate a new proactive defense framework to prevent individuals from being victimized in DeepFake videos. The core idea of this framework is to contaminate the inputs of DeepFake models by disrupting face detectors, based on the observation that face detectors are commonly used to automatically extract victim faces in most DeepFake techniques. Once the face detectors malfunction, the faces will not be correctly extracted, thereby impairing the training or synthesis stages of DeepFake models. To achieve this, we describe a strategy named FacePoison, which fools face detectors by adding dedicated adversarial perturbations to video frames. Building upon this, we introduce VideoFacePoison, an extended strategy that can efficiently propagate FacePoison across video frames instead of applying it individually to each frame, thus significantly reducing the computational overhead while retaining favorable attack performance. This framework is validated on five face detectors, and extensive experiments against eleven different DeepFake models demonstrate the effectiveness of disrupting face detectors to hinder DeepFake generation. The source code is publicly available at: https://github.com/OUC-VAS/FacePoison. Delong Zhu 0002, Yuezun Li, Baoyuan Wu, Jiaran Zhou, Zhibo Wang 0001, Siwei Lyu |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Semantic Contextualization of Face Forgery: A New Definition, Dataset, and Detection MethodabstractIn recent years, deep learning has greatly streamlined the process of manipulating photographic face images. Aware of the potential dangers, researchers have developed various tools to spot these counterfeits. Yet, none asks the fundamental question:What digital manipulations make a real photographic face image fake, while others do not? In this paper, we put face forgery in a semantic context and define thatcomputational methods that alter semantic face attributes to exceed human discrimination thresholds are sources of face forgery. Following our definition, we construct a large face forgery image dataset, where each image is associated with a set of labels organized in a hierarchical graph. Our dataset enables two new testing protocols to probe the generalizability of face forgery detectors. Moreover, we propose a semantics-oriented face forgery detection method that captures label relations and prioritizes the primary task (i.e., real or fake face detection). We show that the proposed dataset successfully exposes the weaknesses of current detectors as the test set and consistently improves their generalizability as the training set. Additionally, we demonstrate the superiority of our semantics-oriented method over traditional binary and multi-class classification-based detectors. Mian Zou, Baosheng Yu, Yibing Zhan, Siwei Lyu, Kede Ma |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | VB-KGN: Variational Bayesian Kernel Generation Networks for Motion Image DeblurringabstractMotion blur estimation is a critical and fundamental task in scene analysis and image restoration. While most state-of-the-art deep learning-based methods for single-image motion image deblurring focus on constructing deep networks or developing training strategies, the characterization of motion blur has received less attention. In this paper, we innovatively propose a non-parametric Variational Bayesian Kernel Generation Network (VB-KGN) for characterizing motion blur in a single image. To solve this model, we employ the variational inference framework to approximate the expected statistical distribution of motion blur images in a data-driven manner. The qualitative and quantitative evaluations of our experimental results demonstrate that our proposed model can generate highly accurate motion blur kernels, significantly improving motion image deblurring performance and substantially reducing the need for extensive training sample preprocessing for deblurring tasks. Ying Fu 0003, Xiaojie Li 0001, Xin Wang 0045, Xi Wu 0004, Shu Hu 0001, Siwei Lyu, Wei Liu 0044 |
IEEE Trans. Multim. | 8 |
| 2025 | Generating Higher-Quality Anti-Forensics DeepFakes with Adversarial Sharpening MaskabstractDeepFake, an AI technology that can automatically synthesize facial forgeries, has recently attracted worldwide attention. While DeepFakes can be entertaining, they can also be used to spread falsified information or be weaponized as cognition warfare. Forensic researchers have been dedicated to designing defensive algorithms to combat such disinformation. However, attacking technologies have been developed to make DeepFake products more aggressive. For example, by launching anti-forensics and adversarial attacks, DeepFakes can be disguised as authentic media to evade forensic detectors. However, such manipulations often sacrifice image quality for satisfactory undetectability. To address this issue, we propose a method to generate a novel adversarial sharpening mask for launching black-box anti-forensics attacks. Unlike many existing methods, our approach injects perturbations that allow DeepFakes to achieve high anti-forensics performance while maintaining pleasant sharpening visual effects. Experimental evaluations demonstrate that our method successfully disrupts state-of-the-art DeepFake detectors. Moreover, compared to images processed by existing DeepFake anti-forensics methods, our method’s quality of anti-forensics DeepFakes rendered is significantly improved. Our code is available at https://github.com/fb-reps/HQ-AF_GAN . Bing Fan, Feng Ding 0007, Guopu Zhu, Jiwu Huang, Sam Kwong, Pradeep K. Atrey, Siwei Lyu |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Spotting the Fakes: A Deep Dive into GAN-Generated Face DetectionabstractGenerative Adversarial Networks (GANs) have enabled the creation of highly authentic facial images, which are increasingly used in deceptive social media profiles and other forms of disinformation, resulting in serious consequences. Significant progress has been made in developing GAN-generated face detection systems to identify these fake images. This study offers a comprehensive review of recent advancements in GAN-generated face detection, focusing on techniques that detect facial images generated by GAN models. We categorize detection methods into three groups: (1) deep learning-based approaches, (2) physics-based methods, and (3) physiology-based methods. We summarize key concepts in each category, connecting them to relevant implementations, datasets, and evaluation metrics. Additionally, we provide a comparative analysis between automated detection and human visual performance to highlight the strengths and weaknesses of both approaches. Furthermore, we review related surveys, including detecting morphed faces, manipulated faces, DeepFake, and faces generated by diffusion models. Finally, we discuss unresolved challenges and suggest potential directions for future research. Xin Wang 0045, Ting Yu Tsai, Shu Hu 0001, Ming-Ching Chang, Pradeep K. Atrey, Siwei Lyu |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2024 | Transcending Forgery Specificity with Latent Space Augmentation for Generalizable Deepfake DetectionabstractDeepfake detection faces a critical generalization hurdle, with performance deteriorating when there is a mismatch between the distributions of training and testing data. A broadly received explanation is the tendency of these detectors to be overfitted to forgery-specific artifacts, rather than learning features that are widely applicable across various forgeries. To address this issue, we propose a simple yet effective detector called LSDA (Latent Space Data Augmentation), which is based on a heuristic idea: representations with a wider variety of forgeries should be able to learn a more generalizable decision boundary, thereby mitigating the overfitting of method-specific features (see Fig. 1). Following this idea, we propose to enlarge the forgery space by constructing and simulating variations within and across forgery features in the latent space. This approach encompasses the acquisition of enriched, domain-specific features and the facilitation of smoother transitions between different forgery types, effectively bridging domain gaps. Our approach culminates in refining a binary classifier that leverages the distilled knowledge from the enhanced features, striving for a generalizable deepfake detector. Comprehensive experiments show that our proposed method is surprisingly effective and transcends state-of-the-art detectors across several widely used benchmarks. Zhiyuan Yan 0002, Yuhao Luo 0002, Siwei Lyu, Qingshan Liu 0001, Baoyuan Wu |
CVPR | 3 |
| 2024 | Enhancing Adversarial Robustness of DNNS Via Weight Decorrelation in TrainingabstractDeep Neural Networks (DNNs) are vulnerable to adversarial perturbations, raising significant concerns about their security. Numerous methods have been proposed to enhance DNN robustness. However, many methods, including adversarial training and noise injection, improve robustness by incorporating external data into the network. Exploring the network’s inherent potential is crucial to improve adversarial robustness. Inspired by principles in physical chemistry, where increased disorder leads to greater energetic stability, we introduce the Weight Decorrelation Loss. This method is simple but effective, enhancing robustness by disrupting the feature space’s ordered structure. The proposed loss achieves substantial performance improvements and state-of-the-art performance after being combined with Gaussian noise. We conduct comprehensive experiments on five datasets, comparing our approach to state-of-the-art defense methods. The results demonstrate our method’s effectiveness against several powerful white-box and black-box attacks. Yuezun Li, Honggang Qi, Siwei Lyu |
ICASSP | 4 |
| 2024 | Exposing Text-Image Inconsistency Using Diffusion ModelsabstractIn the battle against widespread online misinformation, a growing problem is text-image inconsistency, where images are misleadingly paired with texts with different intent or meaning. Existing classification-based methods for text-image inconsistency can identify contextual inconsistencies but fail to provide explainable justifications for their decisions that humans can understand. Although more nuanced, human evaluation is impractical at scale and susceptible to errors. To address these limitations, this study introduces D-TIIL (Diffusion-based Text-Image Inconsistency Localization), which employs text-to-image diffusion models to localize semantic inconsistencies in text and image pairs. These models, trained on large-scale datasets act as ``omniscient" agents that filter out irrelevant information and incorporate background knowledge to identify inconsistencies. In addition, D-TIIL uses text embeddings and modified image regions to visualize these inconsistencies. To evaluate D-TIIL's efficacy, we introduce a new TIIL dataset containing 14K consistent and inconsistent text-image pairs. Unlike existing datasets, TIIL enables assessment at the level of individual words and image regions and is carefully designed to represent various inconsistencies. D-TIIL offers a scalable and evidence-based approach to identifying and localizing text-image inconsistency, providing a robust framework for future research combating misinformation. Mingzhen Huang, Shan Jia, Zhou Zhou 0009, Yan Ju, Jialing Cai, Siwei Lyu |
ICLR | 6 |
| 2024 | Exposing Lip-syncing Deepfakes from Mouth InconsistenciesabstractA lip-syncing deepfake is a digitally manipulated video in which a person’s lip movements are created convincingly using AI models to match altered or entirely new audio. Lipsyncing deepfakes are a dangerous type of deepfakes as the artifacts are limited to the lip region and more difficult to discern. In this paper, we describe a novel approach, LIP-syncing detection based on mouth INConsistency (LIPINC), for lip-syncing deepfake detection by identifying temporal inconsistencies in the mouth region. These inconsistencies are seen in the adjacent frames and throughout the video. Our model can successfully capture these irregularities and outperforms the state-of-the-art methods on several benchmark deepfake datasets. Code is available at https://github.com/skrantidatta/LIPINC. Soumyya Kanti Datta, Shan Jia, Siwei Lyu |
ICME | 3 |
| 2024 | Explicit Correlation Learning for Generalizable Cross-Modal Deepfake DetectionabstractWith the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific modalities, traditional detection methods fall short in addressing the generalizability of detection across diverse cross-modal deepfakes. This paper aims to explicitly learn potential cross-modal correlation to enhance deepfake detection towards various generation scenarios. Our approach introduces a correlation distillation task, which models the inherent cross-modal correlation based on content information. This strategy helps to prevent the model from overfitting merely to audio-visual synchronization. Additionally, we present the Cross-Modal Deepfake Dataset (CMDFD), a comprehensive dataset with four generation methods to evaluate the detection of diverse cross-modal deepfakes. The experimental results on CMDFD and FakeAVCeleb datasets demonstrate the superior generalizability of our method over existing state-of-the-art methods. Our code and data can be found at https://github.com/ljj898/CMDFD-Dataset-and-Deepfake-Detection. Cai Yu, Shan Jia, Xiaomeng Fu, Jin Liu 0020, Jiao Dai, Xi Wang 0014, Siwei Lyu, Jizhong Han |
ICME | 8 |
| 2024 | On the Trajectory Regularity of ODE-based Diffusion SamplingabstractDiffusion-based generative models use stochastic differential equations (SDEs) and their equivalent ordinary differential equations (ODEs) to establish a smooth connection between a complex data distribution and a tractable prior distribution. In this paper, we identify several intriguing trajectory properties in the ODE-based sampling process of diffusion models. We characterize an implicit denoising trajectory and discuss its vital role in forming the coupled sampling trajectory with a strong shape regularity, regardless of the generated content. We also describe a dynamic programming-based scheme to make the time schedule in sampling better fit the underlying trajectory structure. This simple strategy requires minimal modification to any given ODE-based numerical solvers and incurs negligible computational cost, while delivering superior performance in image generation, especially in $5\sim 10$ function evaluations. Defang Chen 0001, Can Wang 0001, Chunhua Shen, Siwei Lyu |
ICML | 5 |
| 2024 | ParallelEdits: Efficient Multi-Aspect Text-Driven Image Editing with Attention GroupingabstractText-driven image synthesis has made significant advancements with the development of diffusion models, transforming how visual content is generated from text prompts. Despite these advances, text-driven image editing, a key area in computer graphics, faces unique challenges. A major challenge is making simultaneous edits across multiple objects or attributes. Applying these methods sequentially for multi-attribute edits increases computational demands and efficiency losses.
In this paper, we address these challenges with significant contributions. Our main contribution is the development of ParallelEdits, a method that seamlessly manages simultaneous edits across multiple attributes. In contrast to previous approaches, ParallelEdits not only preserves the quality of single attribute edits but also significantly improves the performance of multitasking edits. This is achieved through innovative attention distribution mechanism and multi-branch design that operates across several processing heads.
Additionally, we introduce the PIE-Bench++ dataset, an expansion of the original PIE-Bench dataset, to better support evaluating image-editing tasks involving multiple objects and attributes simultaneously. This dataset is a benchmark for evaluating text-driven image editing methods in multifaceted scenarios. Mingzhen Huang, Jialing Cai, Shan Jia, Vishnu Suresh Lokhande, Siwei Lyu |
NeurIPS | 5 |
| 2024 | First-Order Minimax Bilevel OptimizationabstractMulti-block minimax bilevel optimization has been studied recently due to its great potential in multi-task learning, robust machine learning, and few-shot learning. However, due to the complex three-level optimization structure, existing algorithms often suffer from issues such as high computing costs due to the second-order model derivatives or high memory consumption in storing all blocks' parameters. In this paper, we tackle these challenges by proposing two novel fully first-order algorithms named FOSL and MemCS. FOSL features a fully single-loop structure by updating all three variables simultaneously, and MemCS is a memory-efficient double-loop algorithm with cold-start initialization. We provide a comprehensive convergence analysis for both algorithms under full and partial block participation, and show that their sample complexities match or outperform those of the same type of methods in standard bilevel optimization. We evaluate our methods in two applications: the recently proposed multi-task deep AUC maximization and a novel rank-based robust meta-learning. Our methods consistently improve over existing methods with better performance over various datasets. Zhaofeng Si, Siwei Lyu, Kaiyi Ji |
NeurIPS | 3 |
| 2024 | Simple and Fast Distillation of Diffusion ModelsabstractDiffusion-based generative models have demonstrated their powerful performance across various tasks, but this comes at a cost of the slow sampling speed. To achieve both efficient and high-quality synthesis, various distillation-based accelerated sampling methods have been developed recently. However, they generally require time-consuming fine tuning with elaborate designs to achieve satisfactory performance in a specific number of function evaluation (NFE), making them difficult to employ in practice. To address this issue, we propose **S**imple and **F**ast **D**istillation (SFD) of diffusion models, which simplifies the paradigm used in existing methods and largely shortens their fine-tuning time up to $1000\times$. We begin with a vanilla distillation-based sampling method and boost its performance to state of the art by identifying and addressing several small yet vital factors affecting the synthesis efficiency and quality. Our method can also achieve sampling with variable NFEs using a single distilled model. Extensive experiments demonstrate that SFD strikes a good balance between the sample quality and fine-tuning costs in few-step image generation task. For example, SFD achieves 4.53 FID (NFE=2) on CIFAR-10 with only **0.64 hours** of fine-tuning on a single NVIDIA A100 GPU. Defang Chen 0001, Can Wang 0001, Chun Chen 0001, Siwei Lyu |
NeurIPS | 5 |
| 2024 | Improving Fairness in Deepfake DetectionabstractDespite the development of effective deepfake detectors in recent years, recent studies have demonstrated that biases in the data used to train these detectors can lead to disparities in detection accuracy across different races and genders. This can result in different groups being unfairly targeted or excluded from detection, allowing undetected deepfakes to manipulate public opinion and erode trust in a deepfake detection model. While existing studies have focused on evaluating fairness of deepfake detectors, to the best of our knowledge, no method has been developed to encourage fairness in deepfake detection at the algorithm level. In this work, we make the first attempt to improve deepfake detection fairness by proposing novel loss functions that handle both the setting where demographic information (e.g., annotations of race and gender) is available as well as the case where this information is absent. Fundamentally, both approaches can be used to convert many existing deepfake detectors into ones that encourages fairness. Extensive experiments on four deepfake datasets and five deepfake detectors demonstrate the effectiveness and flexibility of our approach in improving deep-fake detection fairness. Our code is available at https://github.com/littlejuyan/DF_Fairness. Yan Ju, Shu Hu 0001, Shan Jia, George H. Chen, Siwei Lyu |
WACV | 5 |
| 2024 | LandmarkBreaker: A proactive method to obstruct DeepFakes via disrupting facial landmark extraction
Yuezun Li, Pu Sun 0001, Honggang Qi, Siwei Lyu |
Comput. Vis. Image Underst. | 4 |
| 2024 | AdaNI: Adaptive Noise Injection to improve adversarial robustness
Yuezun Li, Honggang Qi, Siwei Lyu |
Comput. Vis. Image Underst. | 4 |
| 2024 | COMICS: End-to-End Bi-Grained Contrastive Learning for Multi-Face Forgery DetectionabstractDeepFakes have raised serious societal concerns, leading to a great surge in detection-based forensics methods in recent years. Face forgery recognition is a standard detection method that usually follows a two-phase pipeline,i.e., it extracts the face first and then determines its authenticity by classification. While those methods perform well in ideal experimental environment, they face challenges when dealing with DeepFakes in the wild involving complex background and multiple faces of varying sizes. Moreover, most face forgery recognition methods can only process one face at a time. One straightforward way to address this issue is to simultaneous process multi-face by integrating face extraction and forgery detection in an end-to-end fashion by adapting advanced object detection architectures. However, as these object detection architectures are designed to capture the discriminative features of different object categories rather than the subtle forgery traces among the faces, the direct adaptation suffers from limited representation ability. In this paper, we propose Contrastive Multi-FaceForensics (COMICS), an end-to-end framework for multi-face forgery detection. COMICS integrates face extraction and forgery detection in a seamless manner and adapts to the advanced object detection architectures. The core of the proposed framework is a bi-grained contrastive learning approach that explores face forgery traces at both the coarse- and fine-grained levels. Specifically, coarse-grained level contrastive learning captures the discriminative features among positive and negative proposal pairs at multiple layers produced by the proposal generator, and the fine-grained level contrastive learning captures the pixel-wise discrepancy between the forged and original areas of the same face and the pixel-wise content inconsistency among different faces. Extensive experiments on the OpenForensics and FFIW datasets demonstrate that our method outperforms other counterparts and shows great potential for being integrated into various architectures. Codes are available at https://github.com/zhangconghhh/COMICS. Honggang Qi, Shuhui Wang, Yuezun Li, Siwei Lyu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | ForensicsForest Family: A Series of Multi-Scale Hierarchical Cascade Forests for Detecting GAN-Generated FacesabstractThe prominent progress in generative models has significantly improved the authenticity of generated faces, raising serious concerns in society. To combat GAN-generated faces, many countermeasures based on Convolutional Neural Networks (CNNs) have been spawned due to their strong learning capabilities. In this paper, we rethink this problem and explore a new approach based on forest models instead of CNNs. Concretely, we describe a simple and effective forest-based method set, termed ForensicsForest Family, to detect GAN-generate faces. The ForensicsForest family is composed of three variants: ForensicsForest, Hybrid ForensicsForest and Divide-and-Conquer ForensicsForest. ForenscisForest is a novel Multi-scale Hierarchical Cascade Forest that takes appearance, frequency, and biological features as input, hierarchically cascades different levels of features for authenticity prediction, and employs a multi-scale ensemble scheme to consider different levels of information comprehensively for further performance improvement. Building upon ForensicsForest, we create Hybrid ForensicsForest, an extended version that integrates the CNN layers into models, to further enhance the efficacy of augmented features. Furthermore, to reduce memory usage during training, we introduce Divide-and-Conquer ForensicsForest, which can construct a forest model using only a portion of training samplings. In the training stage, we train several candidate forest models using the subsets of training samples. Then, a ForensicsForest is assembled by selecting suitable components from these candidate forest models. Our method is validated on state-of-the-art GAN-generated face datasets and compared with several CNN models, demonstrating the surprising effectiveness of our method in detecting GAN-generated faces. Jiucui Lu, Jiaran Zhou, Junyu Dong, Bin Li 0011, Siwei Lyu, Yuezun Li |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | GLFF: Global and Local Feature Fusion for AI-Synthesized Image DetectionabstractWith the rapid development of deep generative models (such as Generative Adversarial Networks and Diffusion models), AI-synthesized images are now of such high quality that humans can hardly distinguish them from pristine ones. Although existing detection methods have shown high performance in specific evaluation settings,e.g., on images from seen models or on images without real-world post-processing, they tend to suffer serious performance degradation in real-world scenarios where testing images can be generated by more powerful generation models or combined with various post-processing operations. To address this issue, we propose a Global and Local Feature Fusion (GLFF) framework to learn rich and discriminative representations by combining multi-scale global features from the whole image with refined local features from informative patches for AI-synthesized image detection. GLFF fuses information from two branches: the global branch to extract multi-scale semantic features and the local branch to select informative patches for detailed local artifacts extraction. Due to the lack of a synthesized image dataset simulating real-world applications for evaluation, we further create a challenging fake image dataset, named DeepFakeFaceForensics ($DF^{3}$), which contains 6 state-of-the-art generation models and a variety of post-processing techniques to approach the real-world scenarios. Experimental results demonstrate the superiority of our method to the state-of-the-art methods on the proposed$DF^{3}$dataset and three other open-source datasets. Yan Ju, Shan Jia, Jialing Cai, Haiying Guan, Siwei Lyu |
IEEE Trans. Multim. | 5 |
| 2023 | Outlier Robust Adversarial Training
Shu Hu 0001, Zhenhuan Yang, Xin Wang 0045, Yiming Ying, Siwei Lyu |
ACML | 5 |
| 2023 | Tracking Multiple Deformable Objects in Egocentric VideosabstractMost existing multiple object tracking (MOT) methods that solely rely on appearance features struggle in tracking highly deformable objects. Other MOT methods that use motion clues to associate identities across frames have difficulty handling egocentric videos effectively or efficiently. In this work, we present DogThruGlasses, a large-scale deformable multi-object tracking dataset, with 150 videos and 73K annotated frames, which is collected exclusively by smart glasses. We also propose DETracker, a new MOT method that jointly detects and tracks deformable objects in egocentric videos. DETracker uses three novel modules, namely the motion disentanglement network (MDN), the patch association network (PAN) and the patch memory network (PMN), to explicitly tackle severe ego motion and track fast morphing target objects. DETracker is end-to-end trainable and achieves near real-time speed, which outperforms existing state-of-the-art method on DogThruGlasses and YouTube-Hand. Mingzhen Huang, Xiaoxing Li, Honghong Peng, Siwei Lyu |
CVPR | 5 |
| 2023 | GAN-Generated Faces Detection: A Survey and New PerspectivesabstractGenerative Adversarial Networks (GAN) have led to the generation of very realistic face images, which have been used in fake social media accounts and other disinformation matters that can generate profound impacts. Therefore, the corresponding GAN-face detection techniques are under active development that can examine and expose such fake faces. In this work, we aim to provide a comprehensive review of recent progress in GAN-face detection. We focus on methods that can detect face images that are generated or synthesized from GAN models. We classify the existing detection works into four categories: (1) deep learning-based, (2) physical-based, (3) physiological-based methods, and (4) evaluation and comparison against human visual performance. For each category, we summarize the key ideas and connect them with method implementations. We also discuss open problems and suggest future research directions. Xin Wang 0045, Shu Hu 0001, Ming-Ching Chang, Siwei Lyu |
ECAI | 5 |
| 2023 | Detection of Real-Time Deepfakes in Video Conferencing with Active Probing and Corneal ReflectionabstractThe COVID pandemic has led to the wide adoption of online video calls in recent years. However, the increasing reliance on video calls provides opportunities for new impersonation attacks by fraudsters using the advanced real-time DeepFakes. Real-time DeepFakes pose new challenges to detection methods, which have to run in real-time as a video call is ongoing. In this paper, we describe a new active forensic method to detect real-time DeepFakes. Specifically, we authenticate video calls by displaying a distinct pattern on the screen and using the corneal reflection extracted from the images of the call participant’s face. This pattern can be induced by a call participant displaying on a shared screen or directly integrated into the video-call client. In either case, no specialized imaging or lighting hardware is required. Through large-scale simulations, we evaluate the reliability of this approach under a range in a variety of real-world imaging scenarios. Xin Wang 0045, Siwei Lyu |
ICASSP | 3 |
| 2023 | Face Poison: Obstructing DeepFakes by Disrupting Face DetectionabstractRecent years have seen fast development in synthesizing realistic human faces using AI-based forgery technique called DeepFake, which can be weaponized to cause negative personal and social impacts. In this work, we develop a defense method, namely FacePosion, to prevent individuals from becoming victims of DeepFake videos by sabotaging would-be training data. This is achieved by disrupting face detection, a prerequisite step to prepare victim faces for training DeepFake model. Once the training faces are wrongly extracted, the DeepFake model can not be well trained. Specifically, we propose a multi-scale feature-level adversarial attack to disrupt the intermediate features of face detectors using different scales. Extensive experiments are conducted on seven various DeepFake models using six face detection methods, empirically showing that disrupting face detectors using our method can effectively obstruct DeepFakes. Yuezun Li, Jiaran Zhou, Siwei Lyu |
ICME | 3 |
| 2023 | Forensics Forest: Multi-scale Hierarchical Cascade Forest for Detecting GAN-generated FacesabstractWe describe a simple and effective method called ForensicsForest to detect GAN-generate faces. Instead of using the commonly used CNN models, we describe a novel multi-scale hierarchical cascade forest, which takes semantic and frequency features as input, and hierarchically cascades different levels of features for authenticity prediction. We then propose a multi-scale ensemble, which comprehensively considers different levels of information to improve the performance further. Our method is validated on state-of-the-art GAN-generated face datasets in comparison with several CNN models, which demonstrates the surprising effectiveness of our method in detecting GAN-generated faces. Jiucui Lu, Yuezun Li, Jiaran Zhou, Bin Li 0011, Siwei Lyu |
ICME | 5 |
| 2023 | Controlling Neural Style Transfer with Deep Reinforcement LearningabstractControlling the degree of stylization in the Neural Style Transfer (NST) is a little tricky since it usually needs hand-engineering on hyper-parameters. In this paper, we propose the first deep Reinforcement Learning (RL) based architecture that splits one-step style transfer into a step-wise process for the NST task. Our RL-based method tends to preserve more details and structures of the content image in early steps, and synthesize more style patterns in later steps. It is a user-easily-controlled style-transfer method. Additionally, as our RL-based model performs the stylization progressively, it is lightweight and has lower computational complexity than existing one-step Deep Learning (DL) based models. Experimental results demonstrate the effectiveness and robustness of our method. Chengming Feng, Jing Hu 0009, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Xi Wu 0004, Hongtu Zhu, Siwei Lyu |
IJCAI | 8 |
| 2023 | RMBench: Benchmarking Deep Reinforcement Learning for Robotic Manipulator ControlabstractReinforcement learning is used to tackle complex tasks with high-dimensional sensory inputs. Over the past decade, a wide range of reinforcement learning algorithms have been developed, with recent progress benefiting from deep learning for raw sensory signal representation. This raises a natural question: how well do these algorithms perform across different robotic manipulation tasks? To objectively compare algorithms, benchmarks use performance metrics. Benchmarks use objective performance metrics to offer a scientific way to compare algorithms. In this paper, we introduce RMBench, the first benchmark for robotic manipulations with high-dimensional continuous action and state spaces. We implement and evaluate reinforcement learning algorithms that take observed pixels as inputs and report their average performance and learning curves to demonstrate their performance and training stability. Our study concludes that none of the evaluated algorithms can handle all tasks well, with soft Actor-Critic outperforming most algorithms in terms of average reward and stability, and an algorithm combined with data augmentation potentially facilitating learning policies. Our code is publicly available at https://github.com/xiangyanfei212/RMBench-2022.git, including all benchmark tasks and studied algorithms. Yanfei Xiang, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Xiaomeng Huang, Xi Wu 0004, Siwei Lyu |
IROS | 7 |
| 2023 | Language-guided Human Motion Synthesis with Atomic ActionsabstractLanguage-guided human motion synthesis has been a challenging task due to the inherent complexity and diversity of human behaviors. Previous methods face limitations in generalization to novel actions, often resulting in unrealistic or incoherent motion sequences. In this paper, we propose ATOM (ATomic mOtion Modeling) to mitigate this problem, by decomposing actions into atomic actions, and employing a curriculum learning strategy to learn atomic action composition. First, we disentangle complex human motions into a set of atomic actions during learning, and then assemble novel actions using the learned atomic actions, which offers better adaptability to new actions. Moreover, we introduce a curriculum learning training strategy that leverages masked motion modeling with a gradual increase in the mask ratio, and thus facilitates atomic action assembly. This approach mitigates the overfitting problem commonly encountered in previous methods while enforcing the model to learn better motion representations. We demonstrate the effectiveness of ATOM through extensive experiments, including text-to-motion and action-to-motion synthesis tasks. We further illustrate its superiority in synthesizing plausible and coherent text-guided human motion sequences. Yuanhao Zhai 0001, Mingzhen Huang, Tianyu Luan, Lu Dong 0004, Ifeoma Nwogu, Siwei Lyu, David S. Doermann, Junsong Yuan 0001 |
ACM Multimedia | 6 |
| 2023 | DeepfakeBench: A Comprehensive Benchmark of Deepfake DetectionabstractA critical yet frequently overlooked challenge in the field of deepfake detection is the lack of a standardized, unified, comprehensive benchmark. This issue leads to unfair performance comparisons and potentially misleading results. Specifically, there is a lack of uniformity in data processing pipelines, resulting in inconsistent data inputs for detection models. Additionally, there are noticeable differences in experimental settings, and evaluation strategies and metrics lack standardization. To fill this gap, we present the first comprehensive benchmark for deepfake detection, called \textit{DeepfakeBench}, which offers three key contributions: 1) a unified data management system to ensure consistent input across all detectors, 2) an integrated framework for state-of-the-art methods implementation, and 3) standardized evaluation metrics and protocols to promote transparency and reproducibility. Featuring an extensible, modular-based codebase, \textit{DeepfakeBench} contains 15 state-of-the-art detection methods, 9 deepfake datasets, a series of deepfake detection evaluation protocols and analysis tools, as well as comprehensive evaluations. Moreover, we provide new insights based on extensive analysis of these evaluations from various perspectives (\eg, data augmentations, backbones). We hope that our efforts could facilitate future research and foster innovation in this increasingly critical domain. All codes, evaluations, and analyses of our benchmark are publicly available at \url{https://github.com/SCLBD/DeepfakeBench}. Zhiyuan Yan 0002, Yong Zhang 0034, Xinhang Yuan, Siwei Lyu, Baoyuan Wu |
NeurIPS | 4 |
| 2023 | Rank-Based Decomposable Losses in Machine Learning: A SurveyabstractRecent works have revealed an essential paradigm in designing loss functions that differentiate individual losses versus aggregate losses. The individual loss measures the quality of the model on a sample, while the aggregate loss combines individual losses/scores over each training sample. Both have a common procedure that aggregates a set of individual values to a single numerical value. The ranking order reflects the most fundamental relation among individual values in designing losses. In addition, decomposability, in which a loss can be decomposed into an ensemble of individual terms, becomes a significant property of organizing losses/scores. This survey provides a systematic and comprehensive review of rank-based decomposable losses in machine learning. Specifically, we provide a new taxonomy of loss functions that follows the perspectives of aggregate loss and individual loss. We identify the aggregator to form such losses, which are examples of set functions. We organize the rank-based decomposable losses into eight categories. Following these categories, we review the literature on rank-based aggregate losses and rank-based individual losses. We describe general formulas for these losses and connect them with existing research topics. We also suggest future research directions spanning unexplored, remaining, and emerging issues in rank-based decomposable losses. Shu Hu 0001, Xin Wang 0045, Siwei Lyu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | ExS-GAN: Synthesizing Anti-Forensics Images via Extra Supervised GANabstractSo far, researchers have proposed many forensics tools to protect the authenticity and integrity of digital information. However, with the explosive development of machine learning, existing forensics tools may compromise against new attacks anytime. Hence, it is always necessary to investigate anti-forensics to expose the vulnerabilities of forensics tools. It is beneficial for forensics researchers to develop new tools as countermeasures. To date, one of the potential threats is the generative adversarial networks (GANs), which could be employed for fabricating or forging falsified data to attack forensics detectors. In this article, we investigate the anti-forensics performance of GANs by proposing a novel model, the ExS-GAN, which features an extra supervision system. After training, the proposed model could launch anti-forensics attacks on various manipulated images. Evaluated by experiments, the proposed method could achieve high anti-forensics performance while preserving satisfying image quality. We also justify the proposed extra supervision via an ablation study. Feng Ding 0007, Zhangyi Shen, Guopu Zhu, Sam Kwong, Yicong Zhou, Siwei Lyu |
IEEE Trans. Cybern. | 6 |
| 2023 | Robust Scene Parsing by Mining Supportive Knowledge From DatasetabstractScene parsing, or semantic segmentation, aims at labeling all pixels in an image with the predefined categories of things and stuff. Learning a robust representation for each pixel is crucial for this task. Existing state-of-the-art (SOTA) algorithms employ deep neural networks to learn (discover) the representations needed for parsing from raw data. Nevertheless, these networks discover desired features or representations only from the given image (content), ignoring more generic knowledge contained in the dataset. To overcome this deficiency, we make the first attempt to explore the meaningful supportive knowledge, including general visual concepts (i.e., the generic representations for objects and stuff) and their relations from the whole dataset to enhance the underlying representations of a specific scene for better scene parsing. Specifically, we propose a novel supportive knowledge mining module (SKMM) and a knowledge augmentation operator (KAO), which can be easily plugged into modern scene parsing networks. By taking image-specific content and dataset-level supportive knowledge into full consideration, the resulting model, called knowledge augmented neural network (KANN), can better understand the given scene and provide greater representational power. Experiments are conducted on three challenging scene parsing and semantic segmentation datasets: Cityscapes, Pascal-Context, and ADE20K. The results show that our KANN is effective and achieves better results than all existing SOTA methods. Ao Luo, Fan Yang 0054, Xin Li 0079, Yuezun Li, Zhicheng Jiao, Hong Cheng 0002, Siwei Lyu |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | Towards To-a-T Spatio-Temporal Focus for Skeleton-Based Action RecognitionabstractGraph Convolutional Networks (GCNs) have been widely used to model the high-order dynamic dependencies for skeleton-based action recognition. Most existing approaches do not explicitly embed the high-order spatio-temporal importance to joints’ spatial connection topology and intensity, and they do not have direct objectives on their attention module to jointly learn when and where to focus on in the action sequence. To address these problems, we propose the To-a-T Spatio-Temporal Focus (STF), a skeleton-based action recognition framework that utilizes the spatio-temporal gradient to focus on relevant spatio-temporal features. We first propose the STF modules with learnable gradient-enforced and instance-dependent adjacency matrices to model the high-order spatio-temporal dynamics. Second, we propose three loss terms defined on the gradient-based spatio-temporal focus to explicitly guide the classifier when and where to look at, distinguish confusing classes, and optimize the stacked STF modules. STF outperforms the state-of-the-art methods on the NTU RGB+D 60, NTU RGB+D 120, and Kinetics Skeleton 400 datasets in all 15 settings over different views, subjects, setups, and input modalities, and STF also shows better accuracy on scarce data and dataset shifting settings. Lipeng Ke, Kuan-Chuan Peng, Siwei Lyu |
AAAI | 3 |
| 2022 | Stochastic Planner-Actor-Critic for Unsupervised Deformable Image RegistrationabstractLarge deformations of organs, caused by diverse shapes and nonlinear shape changes, pose a significant challenge for medical image registration. Traditional registration methods need to iteratively optimize an objective function via a specific deformation model along with meticulous parameter tuning, but which have limited capabilities in registering images with large deformations. While deep learning-based methods can learn the complex mapping from input images to their respective deformation field, it is regression-based and is prone to be stuck at local minima, particularly when large deformations are involved. To this end, we present Stochastic Planner-Actor-Critic (spac), a novel reinforcement learning-based framework that performs step-wise registration. The key notion is warping a moving image successively by each time step to finally align to a fixed image. Considering that it is challenging to handle high dimensional continuous action and state spaces in the conventional reinforcement learning (RL) framework, we introduce a new concept `Plan' to the standard Actor-Critic model, which is of low dimension and can facilitate the actor to generate a tractable high dimensional action. The entire framework is based on unsupervised training and operates in an end-to-end manner. We evaluate our method on several 2D and 3D medical image datasets, some of which contain large deformations. Our empirical results highlight that our work achieves consistent, significant gains and outperforms state-of-the-art methods. Ziwei Luo 0002, Jing Hu 0009, Xin Wang 0045, Shu Hu 0001, Bin Kong 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
AAAI | 9 |
| 2022 | Adaptive Face Forgery Detection in Cross Domain
Luchuan Song, Xiaoyi Dong, Zhenchao Jin, Yuefeng Chen, Siwei Lyu |
ECCV (34) | 7 |
| 2022 | Vocbench: A Neural Vocoder Benchmark for Speech SynthesisabstractNeural vocoders, used for converting the spectral representations of an audio signal to the waveforms, are a commonly used component in speech synthesis pipelines. It focuses on synthesizing waveforms from low-dimensional representation, such as Mel-Spectrograms. In recent years, different approaches have been introduced to develop such vocoders. However, it becomes more challenging to assess these new vocoders and compare their performance to previous ones. To address this problem, we present VocBench, a framework that benchmark the performance of state-of-the-art neural vocoders. VocBench uses a systematic study to evaluate different neural vocoders in a shared environment that enables a fair comparison between them. In our experiments, we use the same setup for datasets, training pipeline, and evaluation metrics for all neural vocoders. We perform a subjective and objective evaluation to compare the performance of each vocoder along a different axis. Our results demonstrate that the framework can show competitive efficacy and quality of the synthesized samples for each vocoder. VocBench framework is available at https://github.com/facebookresearch/vocoder-benchmark. Ehab A. AlBadawy, Andrew Gibiansky, Jilong Wu, Ming-Ching Chang, Siwei Lyu |
ICASSP | 6 |
| 2022 | Eyes Tell All: Irregular Pupil Shapes Reveal GAN-Generated FacesabstractGenerative adversarial network (GAN) generated high-realistic human faces are visually challenging to discern from real ones. They have been used as profile images for fake social media accounts, which leads to high negative social impacts. In this work, we show that GAN-generated faces can be exposed via irregular pupil shapes. This phenomenon is caused by the lack of physiological constraints in the GAN models. We demonstrate that such artifacts exist widely in high-quality GAN-generated faces. We design an automatic method to segment the pupils from the eyes and analyze their shapes to distinguish GAN-generated faces from real ones. Qualitative and quantitative evaluations of our method on the Flickr-Faces-HQ dataset and a StyleGAN2 generated face dataset demonstrate the effectiveness and simplicity of our method. Shu Hu 0001, Xin Wang 0045, Ming-Ching Chang, Siwei Lyu |
ICASSP | 5 |
| 2022 | Text-Image De-Contextualization Detection Using Vision-Language ModelsabstractText-image de-contextualization, which uses inconsistent image-text pairs, is an emerging form of misinformation and drawing increasing attention due to the great threat to information authenticity. With real content but semantic mismatch in multiple modalities, the detection of de-contextualization is a challenging problem in media forensics. Inspired by the recent advances in vision-language models with powerful relationship learning between images and texts, we leverage the vision-language models to the media de-contextualization detection task. Two popular models, namely CLIP and VinVL, are evaluated and compared on several news and social media datasets to show their performance in detecting image-text inconsistency in de-contextualization. We also summarize interesting observations and shed lights to the use of vision-language models in de-contextualization detection. Mingzhen Huang, Shan Jia, Ming-Ching Chang, Siwei Lyu |
ICASSP | 4 |
| 2022 | DFGC 2022: The Second DeepFake Game CompetitionabstractThis paper presents the summary report on our DFGC 2022 competition. The DeepFake is rapidly evolving, and realistic face-swaps are becoming more deceptive and difficult to detect. On the other hand, methods for detecting DeepFakes are also improving. There is a two-party game between DeepFake creators and defenders. This competition provides a common platform for benchmarking the game between the current state-of-the-arts in Deep-Fake creation and detection methods. The main research question to be answered by this competition is the current state of the two adversaries when competed with each other. This is the second edition after the last year's DFGC 2021, with a new, more diverse video dataset, a more realistic game setting, and more reasonable evaluation metrics. With this competition, we aim to stimulate research ideas for building better defenses against the DeepFake threats. We also release our DFGC 2022 dataset contributed by both our participants and ourselves to enrich the DeepFake data resources for the research community (https://github.com/NiCE-X/DFGC-2022). Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Zhenan Sun, Zhen Lei 0001, Siwei Lyu |
IJCB | 8 |
| 2022 | Uncertainty Aware Multitask Pyramid Vision Transformer for UAV-Based Object Re-IdentificationabstractObject Re-IDentification (ReID), one of the most significant problems in biometrics and surveillance systems, has been extensively studied by image processing and computer vision communities in the past decades. Learning a robust and discriminative feature representation is a crucial challenge for object ReID. The problem is even more challenging in ReID based on Unmanned Aerial Vehicle (UAV) as the images are characterized by continuously varying camera parameters (e.g., view angle, altitude, etc.) of a flying drone. To address this challenge, multiscale feature representation has been considered to characterize images captured from UAV flying at different altitudes. In this work, we propose a multitask learning approach, which employs a new multiscale architecture without convolution, Pyramid Vision Transformer (PVT), as the backbone for UAV-based object ReID. By uncertainty modeling of intraclass variations, our proposed model can be jointly optimized using both uncertainty-aware object ID and camera ID information. Experimental results are reported on PRAI and VRAI, two ReID data sets from aerial surveillance, to verify the effectiveness of our proposed approach. Syeda Nyma Ferdous, Xin Li 0005, Siwei Lyu |
ICIP | 3 |
| 2022 | Model Attribution of Face-Swap Deepfake VideosabstractAI-created face-swap videos, commonly known as Deepfakes, have attracted wide attention as powerful impersonation attacks. Existing research on Deepfakes mostly focuses on binary detection to distinguish between real and fake videos. However, it is also important to determine the specific generation model for a fake video, which can help attribute it to the source for forensic investigation. In this paper, we fill this gap by studying the model attribution problem of Deepfake videos. We first introduce a new dataset with DeepFakes from Different Models (DFDM) based on several Autoencoder models. Specifically, five generation models with variations in encoder, decoder, intermediate layer, input resolution, and compression ratio have been used to generate a total of 6, 450 Deepfake videos based on the same input. Then we take Deepfakes model attribution as a multiclass classification task and propose a spatial and temporal attention based method to explore the differences among Deep-fakes in the new dataset. Experimental evaluation shows that most existing Deepfakes detection methods failed in Deep-fakes model attribution, while the proposed method achieved over 70% accuracy on the high-quality DFDM dataset1. Shan Jia, Xin Li 0005, Siwei Lyu |
ICIP | 3 |
| 2022 | Fusing Global and Local Features for Generalized AI-Synthesized Image DetectionabstractWith the development of the Generative Adversarial Networks (GANs) and DeepFakes, AI-synthesized images are now of such high quality that humans can hardly distinguish them from real images. It is imperative for media forensics to develop detectors to expose them accurately. Existing detection methods have shown high performance in generated images detection, but they tend to generalize poorly in the real-world scenarios, where the synthetic images are usually generated with unseen models using unknown source data. In this work, we emphasize the importance of combining information from the whole image and informative patches in improving the generalization ability of AI-synthesized image detection. Specifically, we design a two-branch model to combine global spatial information from the whole image and local informative features from multiple patches selected by a novel patch selection module. Multi-head attention mechanism is further utilized to fuse the global and local features. We collect a highly diverse dataset synthesized by 19 models with various objects and resolutions to evaluate our model. Experimental results demonstrate the high accuracy and good generalization ability of our method in detecting generated images. Our code is available at https://github.com/littlejuyan/FusingGlobalandLocal. Yan Ju, Shan Jia, Lipeng Ke, Hongfei Xue, Koki Nagano, Siwei Lyu |
ICIP | 6 |
| 2022 | Faketracer: Exposing Deepfakes with Training Data ContaminationabstractWe describe a proactive defense method to expose Deep-Fakes with training data contamination. Note that the existing methods usually focus on defending from general DeepFakes, which are synthesized by GAN using random noise. In contrast, our method is dedicated to defending from native Deep-Fakes, which is synthesized by auto-encoder that involves face swapping and encoding-decoding process that general DeepFakes do not have. Specifically, we design two types of traces namely sustainable traces and erasable traces, which are added on the faces to manipulate the training of DeepFake models. Consequently, the trained DeepFake model can synthesize faces with sustainable traces but no erasable traces. With the help of these two traces, we can expose DeepFakes proactively. Our method is compared with recent passive and proactive defense methods, which corroborates the efficacy of our method. Pu Sun 0001, Yuezun Li, Honggang Qi, Siwei Lyu |
ICIP | 4 |
| 2022 | Contrastive Class-Specific Encoding for Few-Shot Object DetectionabstractIn this paper, we propose a new few-shot object detection (FSOD) framework that introduces a new contrastive branch to extract the class representation of images, which improves the generalization performance of the detection model for novel classes. Additionally, we investigate the effectiveness of both self-supervised and supervised contrastive losses for class-specific encoding in our framework. Experimental results on the benchmark datasets indicate that our proposed method archives the state-of-the-art performance compared with existing FSOD methods. Dizhong Lin, Ying Fu 0003, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
ICME | 8 |
| 2022 | Counterfactual Image Enhancement for Explanation of Face Swap Deepfakes
Bo Peng 0002, Siwei Lyu, Wei Wang 0025, Jing Dong 0003 |
PRCV (2) | 2 |
| 2022 | Differentially private SGDA for minimax problemsabstractStochastic gradient descent ascent (SGDA) and its variants have been the workhorse for solving minimax problems. However, in contrast to the well-studied stochastic gradient descent (SGD) with differential privacy (DP) constraints, there is little work on understanding the generalization (utility) of SGDA with DP constraints. In this paper, we use the algorithmic stability approach to establish the generalization (utility) of DP-SGDA in different settings. In particular, for the convex-concave setting, we prove that the DP-SGDA can achieve an optimal utility rate in terms of the weak primal-dual population risk in both smooth and non-smooth cases. To our best knowledge, this is the first-ever-known result for DP-SGDA in the non-smooth case. We further provide its utility analysis in the nonconvex-strongly-concave setting which is the first-ever-known result in terms of the primal population risk. The convergence and generalization results for this nonconvex setting are new even in the non-private setting. Finally, numerical experiments are conducted to demonstrate the effectiveness of DP-SGDA for both convex and nonconvex cases. Zhenhuan Yang, Shu Hu 0001, Yunwen Lei, Kush R. Varshney, Siwei Lyu, Yiming Ying |
UAI | 5 |
| 2022 | Simultaneous multi-person tracking and activity recognition based on cohesive cluster search
Wenbo Li 0001, Yi Wei 0006, Siwei Lyu, Ming-Ching Chang |
Comput. Vis. Image Underst. | 3 |
| 2022 | Sum of Ranked Range Loss for Supervised LearningabstractIn forming learning objectives, one oftentimes needs to aggregate a set of individual values to a single output. Such cases occur in the aggregate loss, which combines individual losses of a learning model over each training sample, and in the individual loss for multi-label learning, which combines prediction scores over all class labels. In this work, we introduce the sum of ranked range (SoRR) as a general approach to form learning objectives. A ranked range is a consecutive sequence of sorted values of a set of real numbers. The minimization of SoRR is solved with the difference of convex algorithm (DCA). We explore two applications in machine learning of the minimization of the SoRR framework, namely the AoRR aggregate loss for binary/multi-class classification at the sample level and the TKML individual loss for multi-label/multi-class classification at the label level. A combination loss of AoRR and TKML is proposed as a new learning objective for improving the robustness of multi-label learning in the face of outliers in sample and labels alike. Our empirical results highlight the effectiveness of the proposed optimization frameworks and demonstrate the applicability of proposed losses using synthetic and real data sets. Shu Hu 0001, Yiming Ying, Xin Wang 0045, Siwei Lyu |
J. Mach. Learn. Res. | 4 |
| 2022 | Average Top-k Aggregate Loss for Supervised LearningabstractIn this work, we introduce theaverage top-$k$k($\mathrm {AT}_k$) loss, which is the average over the$k$largest individual losses over a training data, as a new aggregate loss for supervised learning. We show that the$\mathrm {AT}_k$loss is a natural generalization of the two widely used aggregate losses, namely the average loss and the maximum loss. Yet, the$\mathrm {AT}_k$loss can better adapt to different data distributions because of the extra flexibility provided by the different choices of$k$. Furthermore, it remains a convex function over all individual losses and can be combined with different types of individual loss without significant increase in computation. We then provide interpretations of the$\mathrm {AT}_k$loss from the perspective of the modification of individual loss and robustness to training data distributions. We further study the classification calibration of the$\mathrm {AT}_k$loss and the error bounds of$\mathrm {AT}_k$-SVM model. We demonstrate the applicability of minimum average top-$k$learning for supervised learning problems including binary/multi-class classification and regression, using experiments on both synthetic and real datasets. Siwei Lyu, Yanbo Fan, Yiming Ying, Bao-Gang Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Learning a deep dual-level network for robust DeepFake detection
Wenbo Pu, Jing Hu 0009, Xin Wang 0045, Yuezun Li, Shu Hu 0001, Bin B. Zhu, Rui Song 0006, Qi Song 0001, Xi Wu 0004, Siwei Lyu |
Pattern Recognit. | 10 |
| 2022 | DE-GAN: Domain Embedded GAN for High Quality Face Image Inpainting
Xian Zhang 0008, Xin Wang 0045, Canghong Shi, Xiaojie Li 0001, Bin Kong 0001, Siwei Lyu, Bin B. Zhu, Jiancheng Lv 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004, Imran Mumtaz |
Pattern Recognit. | 7 |
| 2022 | LandmarkGAN: Synthesizing faces from landmarks
Pu Sun 0001, Yuezun Li, Honggang Qi, Siwei Lyu |
Pattern Recognit. Lett. | 4 |
| 2022 | A Unified Framework for High Fidelity Face Swap and Expression ReenactmentabstractFace manipulation techniques improve fast with the development of powerful image generation models. Two particular face manipulation methods, namely face swap and expression reenactment attract much attention for their flexibility and ease to generate high quality synthesis results. Recently, these two subjects are actively studied. However, most existing methods treat the two tasks separately, ignoring their underlying similarity. In this paper, we propose to tackle the two problems within a unified framework that achieves high quality synthesis results. The enabling component for our unified framework is the clean disentanglement of 3D pose, shape, and expression factors and then recombining them for different tasks accordingly. We then use the same set of 2D representations for face swap and expression reenactment tasks that are input to a common image translation model to directly generate the final synthetic images. Once trained, the proposed model can accomplish both face swap and expression reenactment tasks for previously unseen subjects. Comprehensive experiments and comparisons show that the proposed method achieves high fidelity results in multiple aspects, and it is especially good at faithfully preserving source facial shape in the face swap task, and accurately transferring facial movements in the expression reenactment task. Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Siwei Lyu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | DetPoseNet: Improving Multi-Person Pose Estimation via Coarse-Pose FilteringabstractHuman detection and pose estimation are essential for understanding human activities in images and videos. Mainstream multi-human pose estimation methods take a top-down approach, where human detection is first performed, then each detected person bounding box is fed into a pose estimation network. This top-down approach suffers from the early commitment of initial detections in crowded scenes and other cases with ambiguities or occlusions, leading to pose estimation failures. We propose the DetPoseNet, an end-to-end multi-human detection and pose estimation framework in a unified three-stage network. Our method consists of a coarse-pose proposal extraction sub-net, a coarse-pose based proposal filtering module, and a multi-scale pose refinement sub-net. The coarse-pose proposal sub-net extracts whole-body bounding boxes and body keypoint proposals in a single shot. The coarse-pose filtering step based on the person and keypoint proposals can effectively rule out unlikely detections, thus improving subsequent processing. The pose refinement sub-net performs cascaded pose estimation on each refined proposal region. Multi-scale supervision and multi-scale regression are used in the pose refinement sub-net to simultaneously strengthen context feature learning. Structure-aware loss and keypoint masking are applied to further improve the pose refinement robustness. Our framework is flexible to accept most existing top-down pose estimators as the role of the pose refinement sub-net in our approach. Experiments on COCO and OCHuman datasets demonstrate the effectiveness of the proposed framework. The proposed method is computationally efficient (5-6x speedup) in estimating multi-person poses with refined bounding boxes in sub-seconds. Lipeng Ke, Ming-Ching Chang, Honggang Qi, Siwei Lyu |
IEEE Trans. Image Process. | 4 |
| 2022 | Anti-Forensics for Face Swapping Videos via Adversarial TrainingabstractGenerating falsified faces by artificial intelligence, widely known as DeepFake, has attracted attention worldwide since 2017. Given the potential threat brought by this novel technique, forensics researchers dedicated themselves to detect the video forgery. Except for exposing falsified faces, there could be extended research directions for DeepFake such as anti-forensics. It can disclose the vulnerability of current DeepFake forensics methods. Besides, it could also enable DeepFake videos as tactical weapons if the falsified faces are more subtle to be detected. In this paper, we propose a GAN model to behave as an anti-forensics tool. It features a novel architecture with additional supervising modules for enhancing image visual quality. Besides, a loss function is designed to improve the efficiency of the proposed model. After experimental evaluations, we show that the DeepFake forensics detectors are susceptible to attacks launched by the proposed method. Besides, the proposed method can efficiently produce anti-forensics videos in satisfying visual quality without noticeable artifacts. Compared with the other anti-forensics approaches, this is tremendous progress achieved for DeepFake anti-forensics. The attack launched by our proposed method can be truly regarded as DeepFake anti-forensics as it can fool detecting algorithms and human eyes simultaneously. Feng Ding 0007, Guopu Zhu, Yingcan Li, Xinpeng Zhang 0001, Pradeep K. Atrey, Siwei Lyu |
IEEE Trans. Multim. | 6 |
| 2021 | Stability and Differential Privacy of Stochastic Gradient Descent for Pairwise Learning with Non-Smooth LossabstractPairwise learning has recently received increasing attention since it subsumes many important machine learning tasks (e.g. AUC maximization and metric learning) into a unifying framework. In this paper, we give the first-ever-known stability and generalization analysis of stochastic gradient descent (SGD) for pairwise learning with non-smooth loss functions, which are widely used (e.g. Ranking SVM with the hinge loss). We introduce a novel decomposition in its stability analysis to decouple the pairwisely dependent random variables, and derive generalization bounds consistent with pointwise learning. Furthermore, we apply our stability analysis to develop differentially private SGD for pairwise learning, for which our utility bounds match with the state-of-the-art output perturbation method (Huai et al., 2020) with smooth losses. Finally, we illustrate the results using specific examples of AUC maximization and similarity metric learning. As a byproduct, we provide an affirmative solution to an open question on the advantage of the nuclear-norm constraint over Frobenius norm constraint in similarity metric learning. Zhenhuan Yang, Yunwen Lei, Siwei Lyu, Yiming Ying |
AISTATS | 3 |
| 2021 | FlagDetSeg: Multi-Nation Flag Detection and Segmentation in the WildabstractWe present a simple and effective flag detection approach for multi-nation flag instance segmentation in-the-wild based on data augmentation and Mask-RCNN PointRend. To the best of our knowledge, this is the first multi-nation flag detection work incorporating recent deep object detection with code and dataset that will be released for public use. Flag images with binary segmentation are collected from public domain including the Open Image V6 and annotated for up to 225 countries. Additional flag images are generated from template flag images with cropping, warping, masking, and color adaption to hallucinate realistic-looking flag images for training and testing. Data augmentation is performed by fusing and transforming the segmented flags on top of natural image backgrounds to synthesize new images. To cope with the large variability of flags with the lack of authentic annotated flags, we combine the trained binary Mask-RCNN segmentation weights with the new multi-nation classifier for fine-tuning. For evaluation, the proposed model is compared with other popular detectors and instance segmentation methods including YOLACT++. Results show the efficacy of the proposed approach. Shou-Fang Wu, Ming-Ching Chang, Siwei Lyu, Cheng-Shih Wong, Abhineet Kumar Pandey, Po-Chi Su |
AVSS | 3 |
| 2021 | A Video Analytic System for Rail Crossing Point ProtectionabstractWith the rise of AI deep learning, video surveillance based on deep neural networks can provide real-time detection and tracking of vehicles and pedestrians. We present a video analytic system for monitoring railway crossing and providing security protection for rail intersections. Our system can automatically determine the rail-crossing gate status via visual detection and analyze traffic by detecting and tracking passing vehicles, thus to oversee a set of rail-transportation related safety events. Assuming a fixed camera view, each gate RoI can be manually annotated once for each site during system setup, and then gate status can be automatically detected afterwards. Vehicles are detected using YOLOv4 and multi-target tracking is performed using DeepSORT. Safety-related events including trespassing are continuously monitored using rule-based triggering. Experimental evaluation is performed on a Youtube rail crossing dataset as well as a private dataset. On the private dataset of 76 total minutes from 38 videos, our system can successfully detect all 56 events out of 58 annotated events. On the public dataset of 14.21 hrs of videos, it detects 58 out of 62 events. Guangliang Zhao, Abhineet Kumar Pandey, Ming-Ching Chang, Siwei Lyu |
AVSS | 4 |
| 2021 | Multi-Teacher Single-Student Visual Transformer with Multi-Level Attention for Face Spoofing Detection
Yao-Hui Huang, Jun-Wei Hsieh, Ming-Ching Chang, Lipeng Ke, Siwei Lyu, Arpita Samanta Santra |
BMVC | 5 |
| 2021 | Learnable Discrete Wavelet Pooling (LDW-Pooling) for Convolutional Networks
Bor-Shiun Wang, Jun-Wei Hsieh, Ping-Yang Chen, Ming-Ching Chang, Lipeng Ke, Siwei Lyu |
BMVC | 6 |
| 2021 | Detection, Tracking, and Counting Meets Drones in Crowds: A BenchmarkabstractTo promote the developments of object detection, tracking and counting algorithms in drone-captured videos, we construct a benchmark with a new drone-captured large-scale dataset, named as DroneCrowd, formed by 112 video clips with 33, 600 HD frames in various scenarios. Notably, we annotate 20, 800 people trajectories with 4.8 million heads and several video-level attributes. Meanwhile, we design the Space-Time Neighbor-Aware Network (STNNet) as a strong baseline to solve object detection, tracking and counting jointly in dense crowds. STNNet is formed by the feature extraction module, followed by the density map estimation heads, and localization and association subnets. To exploit the context information of neighboring objects, we design the neighboring context loss to guide the association subnet training, which enforces consistent relative position of nearby objects in temporal domain. Extensive experiments on our DroneCrowd dataset demonstrate that STNNet performs favorably against the state-of-the-arts. Longyin Wen, Dawei Du, Pengfei Zhu 0001, Qinghua Hu, Qilong Wang 0001, Liefeng Bo, Siwei Lyu |
CVPR | 7 |
| 2021 | Exposing GAN-Generated Faces Using Inconsistent Corneal Specular HighlightsabstractSophisticated generative adversary network (GAN) models are now able to synthesize highly realistic human faces that are difficult to discern from real ones visually. In this work, we show that GAN synthesized faces can be exposed with the inconsistent corneal specular highlights between two eyes. The inconsistency is caused by the lack of physical/physiological constraints in the GAN models. We show that such artifacts exist widely in high-quality GAN synthesized faces and further describe an automatic method to extract and compare corneal specular highlights from two eyes. Qualitative and quantitative evaluations of our method suggest its simplicity and effectiveness in distinguishing GAN synthesized faces. Shu Hu 0001, Yuezun Li, Siwei Lyu |
ICASSP | 3 |
| 2021 | DFGC 2021: A DeepFake Game CompetitionabstractThis paper presents a summary of the DeepFake Game Competition (DFGC) 20211. DeepFake technology is developing fast, and realistic face-swaps are increasingly deceiving and hard to detect. At the same time, DeepFake detection methods are also improving. There is a two-party game between DeepFake creators and detectors. This competition provides a common platform for benchmarking the adversarial game between current state-of-the-art DeepFake creation and detection methods. In this paper, we present the organization, results and top solutions of this competition and also share our insights obtained during this event. We also release the DFGC-21 testing dataset collected from our participants to further benefit the research community2. Bo Peng 0002, Hongxing Fan, Wei Wang 0025, Jing Dong 0003, Yuezun Li, Siwei Lyu, Qi Li 0005, Zhenan Sun, Baoying Chen, Yanjie Hu, Shenghai Luo, Junrui Huang, Yutong Yao, Boyuan Liu, Changtao Miao, Changlei Lu, Wanyi Zhuang |
IJCB | 6 |
| 2021 | NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and LocalizationabstractFor iris recognition in non-cooperative environments, iris segmentation has been regarded as the first most important challenge still open to the biometric community, affecting all downstream tasks from normalization to recognition. In recent years, deep learning technologies have gained significant popularity among various computer vision tasks and also been introduced in iris biometrics, especially iris segmentation. To investigate recent developments and attract more interest of researchers in the iris segmentation method, we organized the 2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization (NIR-ISL 2021) at the 2021 International Joint Conference on Biometrics (IJCB 2021). The challenge was used as a public platform to assess the performance of iris segmentation and localization methods on Asian and African NIR iris images captured in non-cooperative environments. The three best-performing entries achieved solid and satisfactory iris segmentation and localization results in most cases, and their code and models have been made publicly available for reproducibility research. Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Jawad Muhammad, Qi Zhang 0015, Qichuan Tian, Zhaofeng He 0001, Zhenan Sun, Tianbao Liu, Wei Yang 0006, Dongliang Wu, Yingfeng Liu, Ruiye Zhou, Huihai Wu, Junbao Wang, Wantong Xiong, Xueyu Shi, Shao Zeng, Peihua Li, Huijie Wu, Xinhui Zhang, Menghan Zhang, Fadi Boutros, Naser Damer, Arjan Kuijper, Juan E. Tapia, Andres Valenzuela, Christoph Busch 0001, Gourav Gupta, Kiran B. Raja, Xi Wu 0004, Xiaojie Li 0001, Jingfu Yang, Hongyan Jing, Xin Wang 0045, Bin Kong 0001, Youbing Yin, Qi Song 0001, Siwei Lyu, Shu Hu 0001, Leon Premk, Matej Vitek, Vitomir Struc, Peter Peer, Jalil Nourmohammadi-Khiarak, Farhang Jaryani, Samaneh Salehi Nasab, Seyed Naeim Moafinejad, Yasin Amini, Morteza Noshad |
IJCB | 52 |
| 2021 | TkML-AP: Adversarial Attacks to Top-k Multi-Label LearningabstractTop-k multi-label learning, which returns the top-k predicted labels from an input, has many practical applications such as image annotation, document analysis, and web search engine. However, the vulnerabilities of such algorithms with regards to dedicated adversarial perturbation attacks have not been extensively studied previously. In this work, we develop methods to create adversarial perturbations that can be used to attack top-k multi-label learning-based image annotation systems (TkML-AP). Our methods explicitly consider the top-k ranking relation and are based on novel loss functions. Experimental evaluations on large-scale benchmark datasets including PASCAL VOC and MS COCO demonstrate the effectiveness of our methods in reducing the performance of state-of-the-art top-k multi-label learning methods, under both untargeted and targeted attacks. Shu Hu 0001, Lipeng Ke, Xin Wang 0045, Siwei Lyu |
ICCV | 4 |
| 2021 | Invisible Backdoor Attack with Sample-Specific TriggersabstractRecently, backdoor attacks pose a new security threat to the training process of deep neural networks (DNNs). Attackers intend to inject hidden backdoors into DNNs, such that the attacked model performs well on benign samples, whereas its prediction will be maliciously changed if hidden backdoors are activated by the attacker-defined trigger. Existing backdoor attacks usually adopt the setting that triggers are sample-agnostic, i.e., different poisoned samples contain the same trigger, resulting in that the attacks could be easily mitigated by current backdoor defenses. In this work, we explore a novel attack paradigm, where backdoor triggers are sample-specific. In our attack, we only need to modify certain training samples with invisible perturbation, while not need to manipulate other training components (e.g., training loss, and model structure) as required in many existing attacks. Specifically, inspired by the recent advance in DNN-based image steganography, we generate sample-specific invisible additive noises as backdoor triggers by encoding an attacker-specified string into benign images through an encoder-decoder network. The mapping from the string to the target label will be generated when DNNs are trained on the poisoned dataset. Extensive experiments on benchmark datasets verify the effectiveness of our method in attacking models with or without defenses. The code will be available at https://github.com/yuezunli/ISSBA. Yuezun Li, Yiming Li 0004, Baoyuan Wu, Longkang Li, Ran He 0001, Siwei Lyu |
ICCV | 6 |
| 2021 | Imperceptible Adversarial Examples For Fake Image DetectionabstractFooling people with highly realistic fake images generated with Deepfake or GANs brings a great social disturbance to our society. Many methods have been proposed to detect fake images, but they are vulnerable to adversarial perturbations – intentionally designed noises that can lead to the wrong prediction. Existing methods of attacking fake image detectors usually generate adversarial perturbations to perturb almost the entire image. This is redundant and increases the perceptibility of perturbations. In this paper, we propose a novel method to disrupt the fake image detection by determining key pixels to a fake image detector and attacking only the key pixels, which results in the L0and the L2norms of adversarial perturbations much less than those of existing works. Experiments on two public datasets with three fake image detectors indicate that our proposed method achieves state-of the-art performance in both white-box and black-box attacks. Quanyu Liao, Yuezun Li, Xin Wang 0045, Bin Kong 0001, Bin B. Zhu, Siwei Lyu, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
ICIP | 6 |
| 2021 | Face Forgery Detection Based On Segmentation NetworkabstractRecent progress in facial manipulation technologies have made it hard to distinguish the sophisticated face swapped images/videos. Due to the diversity of generation software and data sources, it is extremely challenging to devise an efficient generality framework. Instead of regarding the detection process as a vanilla binary classification task, we proposed a detection framework based on pixel-level classification. Considering that the acquisition of real pixel-level ground-truth is somehow expensive or even impractical, we proposed a pseudo ground-truth generation pipeline with prior knowledge of facial manipulation. Besides, we added a new module into the neural network to capture frequency clues, while the ablation experiment verified the effectiveness of this module. The experimental results on several public datasets demonstrated that our proposed framework is effective and superior to other existing similar detection networks. Yingbin Zhou, Anwei Luo, Xiangui Kang, Siwei Lyu |
ICIP | 4 |
| 2021 | Transferable Adversarial Examples for Anchor Free Object DetectionabstractDeep neural networks have been demonstrated to be vulnerable to adversarial attacks: subtle perturbation can completely change prediction result. The vulnerability has led to a surge of research in this direction, including adversarial attacks on object detection networks. However, previous studies are dedicated to attacking anchor-based object detectors. In this paper, we present the first adversarial attack on anchor-free object detectors. It conducts category-wise, instead of previously instance-wise, attacks on object detectors, and leverages high-level semantic information to efficiently generate transferable adversarial examples, which can also be transferred to attack other object detectors, even anchor-based detectors such as Faster R-CNN. Experimental results on two benchmark datasets demonstrate that our proposed method achieves state-of-the-art performance and transferability. Quanyu Liao, Xin Wang 0045, Bin Kong 0001, Siwei Lyu, Bin B. Zhu, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
ICME | 4 |
| 2021 | Stochastic Actor-Executor-Critic for Image-to-Image TranslationabstractTraining a model-free deep reinforcement learning model to solve image-to-image translation is difficult since it involves high-dimensional continuous state and action spaces. In this paper, we draw inspiration from the recent success of the maximum entropy reinforcement learning framework designed for challenging continuous control problems to develop stochastic policies over high dimensional continuous spaces including image representation, generation, and control simultaneously. Central to this method is the Stochastic Actor-Executor-Critic (SAEC) which is an off-policy actor-critic model with an additional executor to generate realistic images. Specifically, the actor focuses on the high-level representation and control policy by a stochastic latent action, as well as explicitly directs the executor to generate low-level actions to manipulate the state. Experiments on several image-to-image translation tasks have demonstrated the effectiveness and robustness of the proposed SAEC when facing high-dimensional continuous space problems. Ziwei Luo 0002, Jing Hu 0009, Xin Wang 0045, Siwei Lyu, Bin Kong 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
IJCAI | 4 |
| 2021 | TransRPN: Towards the Transferable Adversarial Perturbations using Region Proposal Networks and Beyond
Yuezun Li, Ming-Ching Chang, Pu Sun 0001, Honggang Qi, Junyu Dong, Siwei Lyu |
Comput. Vis. Image Underst. | 6 |
| 2021 | End-to-end multimodal image registration via reinforcement learning
Jing Hu 0009, Ziwei Luo 0002, Xin Wang 0045, Shanhui Sun, Youbing Yin, Kunlin Cao, Qi Song 0001, Siwei Lyu, Xi Wu 0004 |
Medical Image Anal. | 8 |
| 2021 | Fast Online Video Pose Estimation by Dynamic Bayesian Modeling of Mode TransitionsabstractWe propose a fast online video pose estimation method to detect and track human upper-body poses based on a conditional dynamic Bayesian modeling of pose modes without referring to future frames. The estimation of human body poses from videos is an important task with many applications. Our method extends fast image-based pose estimation to live video streams by leveraging the temporal correlation of articulated poses between frames. Video pose estimation is inferred over a time window using a conditional dynamic Bayesian network (CDBN), which we term time-windowed CDBN. Specifically, latent pose modes and their transitions are modeled and co-determined from the combination of three modules: 1) inference based on current observations; 2) the modeling of mode-to-mode transitions as a probabilistic prior; and 3) the modeling of state-to-mode transitions using a multimode softmax regression. Given the predicted pose modes, the body poses in terms of arm joint locations can then be determined more accurately and robustly. Our method is suitable to investigate high frame rate (HFR) scenarios, where pose mode transitions can effectively capture action-related temporal information to boost performance. We evaluate our method on a newly collected HFR-Pose dataset and four major video pose datasets (VideoPose2, TUM Kitchen, FLIC, and Penn_Action). Our method achieves improvements in both accuracy and efficiency over existing online video pose estimation methods. Ming-Ching Chang, Lipeng Ke, Honggang Qi, Longyin Wen, Siwei Lyu |
IEEE Trans. Cybern. | 5 |
| 2020 | 3D Single-Person Concurrent Activity Detection Using Stacked Relation NetworkabstractWe aim to detect real-world concurrent activities performed by a single person from a streaming 3D skeleton sequence. Different from most existing works that deal with concurrent activities performed by multiple persons that are seldom correlated, we focus on concurrent activities that are spatio-temporally or causally correlated and performed by a single person. For the sake of generalization, we propose an approach based on a decompositional design to learn a dedicated feature representation for each activity class. To address the scalability issue, we further extend the class-level decompositional design to the postural-primitive level, such that each class-wise representation does not need to be extracted by independent backbones, but through a dedicated weighted aggregation of a shared pool of postural primitives. There are multiple interdependent instances deriving from each decomposition. Thus, we propose Stacked Relation Networks (SRN), with a specialized relation network for each decomposition, so as to enhance the expressiveness of instance-wise representations via the inter-instance relationship modeling. SRN achieves state-of-the-art performance on a public dataset and a newly collected dataset. The relation weights within SRN are interpretable among the activity contexts. The new dataset and code are available at https://github.com/weiyi1991/UA_Concurrent/ Yi Wei 0006, Wenbo Li 0001, Yanbo Fan, Linghan Xu, Ming-Ching Chang, Siwei Lyu |
AAAI | 6 |
| 2020 | MagGAN: High-Resolution Face Attribute Editing with Mask-Guided Generative Adversarial Network
Yi Wei 0006, Zhe Gan, Wenbo Li 0001, Siwei Lyu, Ming-Ching Chang, Lei Zhang 0001, Jianfeng Gao 0001, Pengchuan Zhang |
ACCV (4) | 4 |
| 2020 | Celeb-DF: A Large-Scale Challenging Dataset for DeepFake ForensicsabstractAI-synthesized face-swapping videos, commonly known as DeepFakes, is an emerging problem threatening the trustworthiness of online information. The need to develop and evaluate DeepFake detection algorithms calls for datasets of DeepFake videos. However, current DeepFake datasets suffer from low visual quality and do not resemble DeepFake videos circulated on the Internet. We present a new large-scale challenging DeepFake video dataset, Celeb-DF, which contains 5,639 high-quality DeepFake videos of celebrities generated using improved synthesis process. We conduct a comprehensive evaluation of DeepFake detection methods and datasets to demonstrate the escalated level of challenges posed by Celeb-DF. Yuezun Li, Xin Yang 0008, Pu Sun 0001, Honggang Qi, Siwei Lyu |
CVPR | 5 |
| 2020 | Learning Semantic Neural Tree for Human Parsing
Ruyi Ji, Dawei Du, Libo Zhang 0001, Longyin Wen, Chen Zhao 0024, Feiyue Huang, Siwei Lyu |
ECCV (13) | 8 |
| 2020 | Cascade Graph Neural Networks for RGB-D Salient Object Detection
Ao Luo, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao, Hong Cheng 0002, Siwei Lyu |
ECCV (12) | 6 |
| 2020 | Fast Portrait Segmentation With Highly Light-Weight NetworkabstractIn this paper, we describe a fast and light-weight portrait segmentation method based on a new highly light-weight backbone (HLB) architecture. The core element of HLB is a bottleneck-based factorized block (BFB) that has much fewer parameters than existing alternatives while keeping good learning capacity. Consequently, the HLB-based portrait segmentation method can run faster than the existing methods yet retaining the competitive accuracy performance with state-of-the-arts. Experiments conducted on two benchmark datasets demonstrate the effectiveness and efficiency of our method. Yuezun Li, Ao Luo, Siwei Lyu |
ICIP | 3 |
| 2020 | Fast Local Attack: Generating Local Adversarial Examples for Object DetectorsabstractThe deep neural network is vulnerable to adversarial examples. Adding imperceptible adversarial perturbations to images is enough to make them fail. Most existing research focuses on attacking image classifiers or anchor-based object detectors, but they generate globally perturbation on the whole image, which is unnecessary. In our work, we leverage higher-level semantic information to generate high aggressive local perturbations for anchor-free object detectors. As a result, it is less computationally intensive and achieves a higher black-box attack as well as transferring attack performance. The adversarial examples generated by our method are not only capable of attacking anchor-free object detectors, but also able to be transferred to attack anchor-based object detector. Quanyu Liao, Xin Wang 0045, Bin Kong 0001, Siwei Lyu, Youbing Yin, Qi Song 0001, Xi Wu 0004 |
IJCNN | 4 |
| 2020 | Voice Conversion Using Speech-to-Speech Neuro-Style Transfer
Ehab A. AlBadawy, Siwei Lyu |
INTERSPEECH | 2 |
| 2020 | Explainable and Efficient Sequential Correlation Network for 3D Single Person Concurrent Activity DetectionabstractWe present the sequential correlation network (SCN) to improve concurrent activity detection. SCN combines a recurrent neural network and a correlation model hierarchically to model the complex correlations and temporal dynamics of concurrent activities. SCN has several advantages that enable effective learning even from a small dataset for real-world deployment. Unlike the majority of approaches assuming that each subject performs one activity at a time, SCN is end-to- end trainable, i.e., it can automatically learn the inclusive or exclusive relations of concurrent activities. SCN is lightweight in design using only a small set of learnable parameters to model the spatio-temporal correlations of activities. This also enhances the explainability of the learned parameters. Furthermore, the learning of SCN can benefit from the initialization using semantically meaningful priors. We evaluate the proposed method against the state-of-the-art method on two benchmark datasets with human skeletal data, SCN achieves comparable performance to the SOTA but with much faster inference speed and less memory usage. Yi Wei 0006, Wenbo Li 0001, Ming-Ching Chang, Hongxia Jin, Siwei Lyu |
IROS | 5 |
| 2020 | Guided Attention Network for Object Detection and Counting on DronesabstractObject detection and counting are related but challenging problems, especially for drone based scenes with small objects and cluttered background. In this paper, we propose a new Guided Attention network (GAnet) to deal with both object detection and counting tasks based on the feature pyramid. Different from the previous methods relying on unsupervised attention modules, we fuse different scales of feature maps by using the proposed weakly-supervised Background Attention (BA) between the background and objects for more semantic feature representation. Then, the Foreground Attention (FA) module is developed to consider both global and local appearance of the object to facilitate accurate localization. Moreover, the new data argumentation strategy is designed to train a robust model in the drone based scenes with various illumination conditions. Extensive experiments on three challenging benchmarks (i.e., UAVDT, CARPK and PUCPR+) show the state-of-the-art detection and counting performance of the proposed method compared with existing methods. Code can be found at https://isrc.iscas.ac.cn/gitlab/research/ganet. Yuanqiang Cai, Dawei Du, Libo Zhang 0001, Longyin Wen, Weiqiang Wang 0001, Siwei Lyu |
ACM Multimedia | 7 |
| 2020 | Learning by Minimizing the Sum of Ranked RangeabstractIn forming learning objectives, one oftentimes needs to aggregate a set of individual values to a single output. Such cases occur in the aggregate loss, which combines individual losses of a learning model over each training sample, and in the individual loss for multi-label learning, which combines prediction scores over all class labels. In this work, we introduce the sum of ranked range (SoRR) as a general approach to form learning objectives. A ranked range is a consecutive sequence of sorted values of a set of real numbers. The minimization of SoRR is solved with the difference of convex algorithm (DCA). We explore two applications in machine learning of the minimization of the SoRR framework, namely the AoRR aggregate loss for binary classification and the TKML individual loss for multi-label/multi-class classification. Our empirical results highlight the effectiveness of the proposed optimization framework and demonstrate the applicability of proposed losses using synthetic and real datasets. Shu Hu 0001, Yiming Ying, Xin Wang 0045, Siwei Lyu |
NeurIPS | 4 |
| 2020 | UA-DETRAC: A new benchmark and protocol for multi-object detection and tracking
Longyin Wen, Dawei Du, Zhaowei Cai, Zhen Lei 0001, Ming-Ching Chang, Honggang Qi, Jongwoo Lim, Ming-Hsuan Yang 0001, Siwei Lyu |
Comput. Vis. Image Underst. | 9 |
| 2019 | Scale Invariant Fully Convolutional Network: Detecting Hands EfficientlyabstractExisting hand detection methods usually follow the pipeline of multiple stages with high computation cost, i.e., feature extraction, region proposal, bounding box regression, and additional layers for rotated region detection. In this paper, we propose a new Scale Invariant Fully Convolutional Network (SIFCN) trained in an end-to-end fashion to detect hands efficiently. Specifically, we merge the feature maps from high to low layers in an iterative way, which handles different scales of hands better with less time overhead comparing to concatenating them simply. Moreover, we develop the Complementary Weighted Fusion (CWF) block to make full use of the distinctive features among multiple layers to achieve scale invariance. To deal with rotated hand detection, we present the rotation map to get rid of complex rotation and derotation layers. Besides, we design the multi-scale loss scheme to accelerate the training process significantly by adding supervision to the intermediate layers of the network. Compared with the state-of-the-art methods, our algorithm shows comparable accuracy and runs a 4.23 times faster speed on the VIVA dataset and achieves better average precision on Oxford hand detection dataset at a speed of 62.5 fps. Dawei Du, Libo Zhang 0001, Tiejian Luo, Feiyue Huang, Siwei Lyu |
AAAI | 7 |
| 2019 | Learning Non-Uniform Hypergraph for Multi-Object TrackingabstractThe majority of Multi-Object Tracking (MOT) algorithms based on the tracking-by-detection scheme do not use higher order dependencies among objects or tracklets, which makes them less effective in handling complex scenarios. In this work, we present a new near-online MOT algorithm based on non-uniform hypergraph, which can model different degrees of dependencies among tracklets in a unified objective. The nodes in the hypergraph correspond to the tracklets and the hyperedges with different degrees encode various kinds of dependencies among them. Specifically, instead of setting the weights of hyperedges with different degrees empirically, they are learned automatically using the structural support vector machine algorithm (SSVM). Several experiments are carried out on various challenging datasets (i.e., PETS09, ParkingLot sequence, SubwayFace, and MOT16 benchmark), to demonstrate that our method achieves favorable performance against the state-of-the-art MOT methods. Longyin Wen, Dawei Du, Shengkun Li, Xiao Bian, Siwei Lyu |
AAAI | 5 |
| 2019 | Graph-to-Graph Energy Minimization for Video Object SegmentationabstractWe describe a new unsupervised video object segmentation (VOS) method based on the graph-to-graph energy minimization, which focuses on exploiting the mutual bootstrapping information between bottom-up (i.e., using pixel/superpixel attributes) and top-down (i.e., using learned appearance and motion cues) processes in a unified framework. Specifically, we construct a graph-to-graph energy function to encode the spatial similarities among superpixels (superpixel-graph) and temporal consistency among regions (region-graph). An efficient heuristic iterative algorithm is used to minimize the energy function to get the optimal assignment of superpixel and region labels to complete the VOS task. Experiments on two challenging benchmarks (i.e., SegTrack v2 and DAVIS) show that the proposed method achieves favorable performance against the state-of-the-art unsupervised VOS methods and comparable performance with the state-of-the-art semi-supervised methods. Yuezun Li, Longyin Wen, Ming-Ching Chang, Siwei Lyu |
AVSS | 4 |
| 2019 | Exploring the Vulnerability of Single Shot Module in Object Detectors via Imperceptible Background Patches
Yuezun Li, Xiao Bian, Ming-Ching Chang, Siwei Lyu |
BMVC | 4 |
| 2019 | Object-Driven Text-To-Image Synthesis via Adversarial TrainingabstractIn this paper, we propose Object-driven Attentive Generative Adversarial Newtorks (Obj-GANs) that allow attention-driven, multi-stage refinement for synthesizing complex images from text descriptions. With a novel object-driven attentive generative network, the Obj-GAN can synthesize salient objects by paying attention to their most relevant words in the text descriptions and their pre-generated class label. In addition, a novel object-wise discriminator based on the Fast R-CNN model is proposed to provide rich object-wise discrimination signals on whether the synthesized object matches the text description and the pre-generated class label. The proposed Obj-GAN significantly outperforms the previous state of the art in various metrics on the large-scale MS-COCO benchmark, increasing the inception score by 27% and decreasing the FID score by 11%. A thorough comparison between the classic grid attention and the new object-driven attention is provided through analyzing their mechanisms and visualizing their attention layers, showing insights of how the proposed model generates complex scenes in high quality. Wenbo Li 0001, Pengchuan Zhang, Lei Zhang 0001, Qiuyuan Huang, Xiaodong He 0001, Siwei Lyu, Jianfeng Gao 0001 |
CVPR | 6 |
| 2019 | Dual-stream CNN for Structured Time Series ClassificationabstractThe structured time series (STS) classification problem requires the modeling of interweaved spatiotemporal dependency. Most previous methods model these two dependencies independently. Due to the complexity of the STS data, we argue that a desirable method should be a holistic framework that is adaptive and flexible. This motivates us to design a deep neural network with such merits. Inspired by the dual-stream hypothesis in neural science, we propose a novel dual-stream framework for modeling the interweaved spatiotemporal dependency, and develop a convolutional neural network within this framework that aims to achieve high adaptability and flexibility in STS configurations of sequential order and dependency range. Our model is highly modularized and scalable, making it easy to be adapted to specific tasks. The effectiveness of our model is demonstrated through experiments on benchmark datasets for skeleton based activity recognition. Shuchen Weng, Wenbo Li 0001, Yi Zhang 0070, Siwei Lyu |
ICASSP | 4 |
| 2019 | Exposing Deep Fakes Using Inconsistent Head PosesabstractIn this paper, we propose a new method to expose AI-generated fake face images or videos (commonly known as the Deep Fakes). Our method is based on the observations that Deep Fakes are created by splicing synthesized face region into the original image, and in doing so, introducing errors that can be revealed when 3D head poses are estimated from the face images. We perform experiments to demonstrate this phenomenon and further develop a classification method based on this cue. Using features based on this cue, an SVM classifier is evaluated using a set of real face images and Deep Fakes. Xin Yang 0008, Yuezun Li, Siwei Lyu |
ICASSP | 3 |
| 2019 | De-identification Without Losing FacesabstractTraining of deep learning models for computer vision requires large image or video datasets from real world. Often, in collecting such datasets, we also need to protect the privacy of the people captured in the images or videos, while still preserve useful attributes such as facial expressions. In this work, we describe a new face de-identification method to achieve this, which is based on a face attribute transfer model (FATM). FATM is a deep neural network model trained to map non-identity related facial attributes to the face of donors, who are a small number of consented subjects. Using the donors' faces ensures the natural appearance of the synthesized faces, and FATM blends the donors' facial attributes to those of the original faces to diversify the appearance of the synthesized faces. Experimental results on several sets of images and videos demonstrate the effectiveness of our face de-ID algorithm. Yuezun Li, Siwei Lyu |
IH&MMSec | 2 |
| 2019 | Exposing GAN-synthesized Faces Using Landmark LocationsabstractGenerative adversary networks (GANs) have recently led to highly realistic image synthesis results. In this work, we describe a new method to expose GAN-synthesized images using the locations of the facial landmark points. Our method is based on the observations that the facial parts configuration generated by GAN models are different from those of the real faces, due to the lack of global constraints. We perform experiments demonstrating this phenomenon, and show that an SVM classifier trained using the locations of facial landmark points is sufficient to achieve good classification performance for GAN-synthesized faces. Xin Yang 0008, Yuezun Li, Honggang Qi, Siwei Lyu |
IH&MMSec | 4 |
| 2019 | Deep Correlated Predictive Subspace Learning for Incomplete Multi-View Semi-Supervised ClassificationabstractIncomplete view information often results in failure cases of the conventional multi-view methods. To address this problem, we propose a Deep Correlated Predictive Subspace Learning (DCPSL) method for incomplete multi-view semi-supervised classification. Specifically, we integrate semi-supervised deep matrix factorization, correlated subspace learning, and multi-view label prediction into a unified framework to jointly learn the deep correlated predictive subspace and multi-view shared and private label predictors. DCPSL is able to learn proper subspace representation that is suitable for class label prediction, which can further improve the performance of classification. Extensive experimental results on various practical datasets demonstrate that the proposed method performs favorably against the state-of-the-art methods. Zhe Xue, Junping Du 0001, Dawei Du, Wenqi Ren, Siwei Lyu |
IJCAI | 5 |
| 2019 | A Multi-modality Network for Cardiomyopathy Death Risk Prediction with CMR Images and Clinical Information
Chaoyang Xia, Xiaojie Li 0001, Xin Wang 0045, Bin Kong 0001, Yucheng Chen 0003, Youbing Yin, Kunlin Cao, Qi Song 0001, Siwei Lyu, Xi Wu 0004 |
MICCAI (2) | 9 |
| 2019 | Data Priming Network for Automatic Check-OutabstractAutomatic Check-Out (ACO) receives increased interests in recent years. An important component of the ACO system is the visual item counting, which recognizes the categories and counts of the items chosen by the customers. However, the training of such a system is challenged by the domain adaptation problem, in which the training data are images from isolated items while the testing images are for collections of items. Existing methods solve this problem with data augmentation using synthesized images, but the image synthesis leads to unreal images that affect the training process. In this paper, we propose a new data priming method to solve the domain adaptation problem. Specifically, we first use pre-augmentation data priming, in which we remove distracting background from the training images using the coarse-to-fine strategy and select images with realistic view angles by the pose pruning method. In the post-augmentation step, we train a data priming network using detection and counting collaborative learning, and select more reliable images from testing data to fine-tune the final visual item tallying network. Experiments on the large scale Retail Product Checkout (RPC) dataset demonstrate the superiority of the proposed method, i.e., we achieve 80.51% checkout accuracy compared with 56.68% of the baseline methods. The source codes can be found in https://isrc.iscas.ac.cn/gitlab/research/acm-mm-2019-ACO. Dawei Du, Libo Zhang 0001, Tiejian Luo, Qi Tian 0001, Longyin Wen, Siwei Lyu |
ACM Multimedia | 8 |
| 2019 | Single-Shot Scale-Aware Network for Real-Time Face Detection
Longyin Wen, Hailin Shi, Zhen Lei 0001, Siwei Lyu, Stan Z. Li |
Int. J. Comput. Vis. | 5 |
| 2019 | Deep low-rank subspace ensemble for multi-view clustering
Zhe Xue, Junping Du 0001, Dawei Du, Siwei Lyu |
Inf. Sci. | 4 |
| 2019 | Efficient algorithms for graph regularized PLSA for probabilistic topic modeling
Xin Wang 0045, Ming-Ching Chang, Siwei Lyu |
Pattern Recognit. | 4 |
| 2019 | Deep Constrained Low-Rank Subspace Learning for Multi-View Semi-Supervised ClassificationabstractSemi-supervised classification receives increasing interests because it can predict class labels based on both limited labeled and sufficient unlabeled data. In this letter, we propose a deep constrained low-rank subspace learning (DCLSL) method for multi-view semi-supervised classification. Specifically, we integrate deep constrained matrix factorization, low-rank subspace learning, and class label learning into a unified objective function to jointly learn data similarity matrices and class label matrix. DCLSL is able to obtain the discriminative subspace representation of each view and effectively aggregate similarity matrices of multiple views, resulting in better classification performance. Experimental results on various datasets demonstrate the effectiveness of our method. Zhe Xue, Junping Du 0001, Dawei Du, Guorong Li, Qingming Huang, Siwei Lyu |
IEEE Signal Process. Lett. | 6 |
| 2018 | Evolvement Constrained Adversarial Learning for Video Style Transfer
Wenbo Li 0001, Longyin Wen, Xiao Bian, Siwei Lyu |
ACCV (1) | 4 |
| 2018 | Pixel Offset Regression (POR) for Single-shot Instance SegmentationabstractState-of-the-art instance segmentation methods including Mask-RCNN and MNC are multi-shot, as multiple region of interest (ROI) forward passes are required to distinguish candidate regions. Multi-shot architectures usually achieve good performance on public benchmarks. However, hundreds of ROI forward passes in sequel limits their running efficiency, which is a critical point in several utilities such as vehicle surveillance. As such, we arrange our focus on seeking a well trade-off between performance and efficiency. In this paper, we introduce a novel Pixel Offset Regression (POR) scheme which can simply extend single-shot object detector to single-shot instance segmentation system, i.e., segmenting all instances in a single pass. Our framework is based on VGG161with following four parts: (1) a single-shot detection branch to generate object detections, (2) a segmentation branch to estimate foreground masks, (3) a pixel offset regression branch to effectively estimate the distance and orientation from each pixel to the respective object center and (4) a merging process combining output of each branch to obtain instances. Our framework is evaluated on Berkeley-BDD, KITTI and PASCAL VOC2012 validation set, with comparison against several VGG16 based multi-shot methods. Without whistles and bells, our framework exhibits decent performance, which shows good potential for fast speed required applications. Yuezun Li, Xiao Bian, Ming-Ching Chang, Longyin Wen, Siwei Lyu |
AVSS | 5 |
| 2018 | UA-DETRAC 2018: Report of AVSS2018 & IWT4S Challenge on Advanced Traffic MonitoringabstractA desirable smart traffic-monitoring and street-safety system can elicit and support the intervention of law enforcement agencies or medical staff. Recently, there has been a dramatically higher demand for such smart systems. To this end, the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S) was organized in conjunction with the 15th IEEE International Conference on Advanced Video and Signal-based Surveillance (AVSS 2018). Our goal is to advance the state-of-the-art detection and tracking algorithms and provide a comprehensive performance evaluation for them. We evaluate 5 submitted detection and 7 submitted tracking methods on the large-scale UA-DETRAC benchmark, and the results are shared publicly on the website http://detrac-db. rit.albany.edu. We expect this challenge to advance the research and development of new detection and tracking methods for transportation applications. Siwei Lyu, Ming-Ching Chang, Dawei Du, Wenbo Li 0001, Yi Wei 0006, Marco Del Coco, Pierluigi Carcagnì, Arne Schumann, Bharti Munjal, Dinh-Quoc-Trung Dang, Doo-Hyun Choi, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Guna Seetharaman, Jang-Woon Baek, Jong Taek Lee, Kannappan Palaniappan, Kil-Taek Lim, Kiyoung Moon, Kwang-Ju Kim, Lars Wilko Sommer, Meltem Brandlmaier, Minsung Kang, Moongu Jeon, Noor Al-Shakarji, Oliver Acatay, Pyong-Kun Kim, Sikandar Amin, Thomas Sikora, Tien Ba Dinh, Tobias Senst, Vu-Gia-Hy Che, Young-Chul Lim, Yun-Su Chung |
AVSS | 1 |
| 2018 | Robust Adversarial Perturbation on Deep Proposal-based Models
Yuezun Li, Daniel Tian, Ming-Ching Chang, Xiao Bian, Siwei Lyu |
BMVC | 5 |
| 2018 | Tagging Like Humans: Diverse and Distinct Image AnnotationabstractIn this work we propose a new automatic image annotation model, dubbed diverse and distinct image annotation (D2IA). The generative model D2IA is inspired by the ensemble of human annotations, which create semantically relevant, yet distinct and diverse tags. In D2IA, we generate a relevant and distinct tag subset, in which the tags are relevant to the image contents and semantically distinct to each other, using sequential sampling from a determinantal point process (DPP) model. Multiple such tag subsets that cover diverse semantic aspects or diverse semantic levels of the image contents are generated by randomly perturbing the DPP sampling process. We leverage a generative adversarial network (GAN) model to train D2IA. Extensive experiments including quantitative and qualitative comparisons, as well as human subject studies, on two benchmark datasets demonstrate that the proposed model can produce more diverse and distinct tags than the state-of-the-arts. Baoyuan Wu, Peng Sun 0011, Wei Liu 0005, Bernard Ghanem, Siwei Lyu |
CVPR | 6 |
| 2018 | Multi-Scale Structure-Aware Network for Human Pose Estimation
Lipeng Ke, Ming-Ching Chang, Honggang Qi, Siwei Lyu |
ECCV (2) | 4 |
| 2018 | Multi-Scale Supervised Network for Human Pose EstimationabstractHuman pose estimation is an important topic in computer vision with many applications including gesture and activity recognition. However, pose estimation from image is challenging due to appearance variations, occlusions, clutter background, and complex activities. To alleviate these problems, we develop a robust pose estimation method based on the recent deep conv-deconv modules with two improvements: (1) multi -scale supervision of body keypoints, and (2) a global regression to improve structural consistency of keypoints. We refine keypoint detection heatmaps using layer-wise multi-scale supervision to better capture local contexts. Pose inference via keypoint association is optimized globally using a regression network at the end. Our method can effectively disambiguate keypoint matches in close proximity including the mismatch of left-right body parts, and better infer occluded parts. Experimental results show that our method achieves competitive performance among state-of-the-art methods on the MPII and FLIC datasets. Lipeng Ke, Honggang Qi, Ming-Ching Chang, Siwei Lyu |
ICIP | 4 |
| 2018 | Stochastic Proximal Algorithms for AUC MaximizationabstractStochastic optimization algorithms such as SGDs update the model sequentially with cheap per-iteration costs, making them amenable for large-scale data analysis. However, most of the existing studies focus on the classification accuracy which can not be directly applied to the important problems of maximizing the Area under the ROC curve (AUC) in imbalanced classification and bipartite ranking. In this paper, we develop a novel stochastic proximal algorithm for AUC maximization which is referred to as SPAM. Compared with the previous literature, our algorithm SPAM applies to a non-smooth penalty function, and achieves a convergence rate of O(log t/t) for strongly convex functions while both space and per-iteration costs are of one datum. Michael Natole, Yiming Ying, Siwei Lyu |
ICML | 3 |
| 2018 | Explain Black-box Image Classifications Using Superpixel-based InterpretationabstractHow to best understand and interpret the decisions of deep neural networks is a crucial topic, as the impact of intelligent deep network systems is prevalent in many applications. We propose a superpixel based method to interpret and explain the results of black-box deep networks in the widely-applied image classification tasks. We perform probabilistic prediction difference analysis upon one or more superpixels clustered from image pixels. Our method generates a superpixel score map visualization that can provide rich interpretation regarding image components. Such interpretation provides supportive/unsupportive likelihood of image regions upon the decisions performed by the black-box classifier. We compare our method against state-of-art pixelwise interpretation methods over the latest deep neural network classifiers on the ImageNet dataset. Results show that our method produces more consistent interpretations in less computation time. Our method also supports interactive interpretation, where users can acquire explanations on specified regions through a convenient interface for a prompt reaction. Yi Wei 0006, Ming-Ching Chang, Yiming Ying, Ser-Nam Lim, Siwei Lyu |
ICPR | 5 |
| 2018 | Global Contrast Enhancement Detection via Deep Multi-Path NetworkabstractIdentifying global contrast enhancement in an image is an important task in forensics estimation. Several previous methods analyze the “peak-gap” fingerprints in graylevel histograms. However, images in real scenarios are often stored in the JPEG format with middle/low compression quality, resulting in less obvious “peak-gap” effect and then unsatisfactory performance. In this paper, we propose a novel deep Multi-Path Network (MPNet) based approach to learn discriminative features from graylevel histograms. Specifically, given the histograms, their high-level peaks and gaps information can be exploited effectively after several shared convolutional layers in the network, even in middle/low quality compressed images. Moreover, the proposed multi-path module is able to focus on dealing with specific forensics operations for more robustness on image compression. The experiments on three challenging datasets (i.e., Dresden, RAISE and UCID) demonstrate the effectiveness of the proposed method compared to existing methods. Dawei Du, Lipeng Ke, Honggang Qi, Siwei Lyu |
ICPR | 5 |
| 2018 | Efficient State Estimation with Constrained Rao-Blackwellized Particle FilterabstractDue to the limitations of the robotic sensors, during a robotic manipulation task, the acquisition of the object's state can be unreliable and noisy. Combining an accurate model of multi-body dynamic system with Bayesian filtering methods has been shown to be able to filter out noise from the object's observed states. However, efficiency of these filtering methods suffers from samples that violate the physical constraints, e.g., no penetration constraint. In this paper, we propose a Rao-Blackwellized Particle Filter (RBPF) that samples the contact states and updates the object's poses using Kalman filters. This RBPF also enforces the physical constraints on the samples by solving a quadratic programming problem. By comparing our method with methods that does not consider physical constraints, we show that our proposed RBPF is not only able to estimate the object's states, e.g., poses, more accurately but also able to infer unobserved states, e.g., velocities, with higher precision. Shuai Li 0015, Siwei Lyu, Jeffrey C. Trinkle |
IROS | 2 |
| 2018 | A Univariate Bound of Area Under ROC
Siwei Lyu, Yiming Ying |
UAI | 1 |
| 2018 | Countering JPEG anti-forensics based on noise level estimation
Hui Zeng 0002, Xiangui Kang, Siwei Lyu |
Sci. China Inf. Sci. | 4 |
| 2018 | Multi-label Learning with Missing Labels Using Mixed Dependency Graphs
Baoyuan Wu, Wei Liu 0005, Bernard Ghanem, Siwei Lyu |
Int. J. Comput. Vis. | 5 |
| 2018 | Iterative Graph Seeking for Object TrackingabstractTo effectively solve the challenges in object tracking, such as large deformation and severe occlusion, many existing methods use graph-based models to capture target part relations, and adopt a sequential scheme of target part selection, part matching, and state estimation. However, such methods have two major drawbacks: 1) inaccurate part selection leads to performance deterioration of part matching and state estimation and 2) there are insufficient effective global constraints for local part selection and matching. In this paper, we propose a new object tracking method based on iterative graph seeking, which integrate target part selection, part matching, and state estimation using a unified energy minimization framework. Our method also incorporates structural information in local parts variations using the global constraint. We devise an alternative iteration scheme to minimize the energy function for searching the most plausible target geometric graph. Experimental results on several challenging benchmarks (i.e., VOT2015, OTB2013, and OTB2015) demonstrate improved performance and robustness in comparison with existing algorithms. Dawei Du, Longyin Wen, Honggang Qi, Qingming Huang, Qi Tian 0001, Siwei Lyu |
IEEE Trans. Image Process. | 6 |
| 2018 | Contrast Enhancement Estimation for Digital Image ForensicsabstractInconsistency in contrast enhancement can be used to expose image forgeries. In this work, we describe a new method to estimate contrast enhancement operations from a single image. Our method takes advantage of the nature of contrast enhancement as a mapping between pixel values and the distinct characteristics it introduces to the image pixel histogram. Our method recovers the original pixel histogram and the contrast enhancement simultaneously from a single image with an iterative algorithm. Unlike previous works, our method is robust in the presence of additive noise perturbations that are used to hide the traces of contrast enhancement. Furthermore, we also develop an effective method to detect image regions undergone contrast enhancement transformations that are different from the rest of the image, and we use this method to detect composite images. We perform extensive experimental evaluations to demonstrate the efficacy and efficiency of our method. Longyin Wen, Honggang Qi, Siwei Lyu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2017 | Unsupervised Learning of Multi-Level Descriptors for Person Re-IdentificationabstractIn this paper, we propose a novel coding method named weighted linear coding (WLC) to learn multi-level (e.g., pixel-level, patch-level and image-level) descriptors from raw pixel data in an unsupervised manner. It guarantees the property of saliency with a similarity constraint. The resulting multi-level descriptors have a good balance between the robustness and distinctiveness. Based on WLC, all data from the same region can be jointly encoded. Consequently, when we extract the holistic image features, it is able to preserve the spatial consistency. Furthermore, we apply PCA to these features and compact person representations are then achieved. During the stage of matching persons, we exploit the complementary information resided in multi-level descriptors via a score-level fusion strategy. Experiments on the challenging person re-identification datasets - VIPeR and CUHK 01, demonstrate the effectiveness of our method. Yang Yang 0062, Longyin Wen, Siwei Lyu, Stan Z. Li |
AAAI | 3 |
| 2017 | UA-DETRAC 2017: Report of AVSS2017 & IWT4S Challenge on Advanced Traffic MonitoringabstractThe rapid advances of transportation infrastructure have led to a dramatic increase in the demand for smart systems capable of monitoring traffic and street safety. Fundamental to these applications are a community-based evaluation platform and benchmark for object detection and multi-object tracking. To this end, we organize the AVSS2017 Challenge on Advanced Traffic Monitoring, in conjunction with the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S), to evaluate the state-of-the-art object detection and multi-object tracking algorithms in the relevance of traffic surveillance. Submitted algorithms are evaluated using the large-scale UA-DETRAC benchmark and evaluation protocol. The benchmark, the evaluation toolkit and the algorithm performance are publicly available from the website http://detrac-db.rit.albany.edu. Siwei Lyu, Ming-Ching Chang, Dawei Du, Longyin Wen, Honggang Qi, Yuezun Li, Yi Wei 0006, Lipeng Ke, Tao Hu 0011, Marco Del Coco, Pierluigi Carcagnì, Dmitriy Anisimov, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Hao Ye 0005, Hong Wang 0014, Kannappan Palaniappan, Koray Ozcan, Li Wang 0033, Liang Wang 0001, Martin Lauer, Nattachai Watcharapinchai, Nenghui Song, Noor Al-Shakarji, Sikandar Amin, Sitapa Watcharapinchai, Tatiana Khanova, Thomas Sikora, Tino Kutschbach, Volker Eiselein, Wei Tian 0001, Xiangyang Xue 0001, Xiaoyi Yu, Yao Lu 0028, Yingbin Zheng, Yongzhen Huang, Yuqi Zhang 0001 |
AVSS | 1 |
| 2017 | Adaptive RNN Tree for Large-Scale Human Action RecognitionabstractIn this work, we present the RNN Tree (RNN-T), an adaptive learning framework for skeleton based human action recognition. Our method categorizes action classes and uses multiple Recurrent Neural Networks (RNNs) in a treelike hierarchy. The RNNs in RNN-T are co-trained with the action category hierarchy, which determines the structure of RNN-T. Actions in skeletal representations are recognized via a hierarchical inference process, during which individual RNNs differentiate finer-grained action classes with increasing confidence. Inference in RNN-T ends when any RNN in the tree recognizes the action with high confidence, or a leaf node is reached. RNN-T effectively addresses two main challenges of large-scale action recognition: (i) able to distinguish fine-grained action classes that are intractable using a single network, and (ii) adaptive to new action classes by augmenting an existing model. We demonstrate the effectiveness of RNN-T/ACH method and compare it with the state-of-the-art methods on a large-scale dataset and several existing benchmarks. Wenbo Li 0001, Longyin Wen, Ming-Ching Chang, Ser-Nam Lim, Siwei Lyu |
ICCV | 5 |
| 2017 | Hybrid structure hypergraph for online deformable object trackingabstractRecent advances in visual tracking field design part-based model to handle the deformation and occlusion challenges. Previous methods only consider the sole degree of dependencies (e.g., pairwise or high-order dependencies) between object parts in consecutive frames. However, the degree of dependencies of different object parts in consecutive frames are not consistent, especially when large deformation and occlusion happen. To that end, we design a hybrid structure hypergraph based tracker, which use a non-uniform hypergraph to model the dependencies among object parts. The tracking task is further formulated as the dense structures extracting problem on the non-uniform hypergraph, which is solved by an approximate algorithm efficiently. Several experiments are carried out on publicly available online deformable object tracking dataset, i.e., Deform-SOT dataset, to demonstrate the favorable performance of the proposed method against the state-of-the-art online tracking methods. Shengkun Li, Dawei Du, Longyin Wen, Ming-Ching Chang, Siwei Lyu |
ICIP | 5 |
| 2017 | LSTM with working memoryabstractPrevious RNN architectures have largely been superseded by LSTM, or “Long Short-Term Memory”. Since its introduction, there have been many variations on this simple design. However, it is still widely used and we are not aware of a gated-RNN architecture that outperforms LSTM in a broad sense while still being as simple and efficient. In this paper we propose a modified LSTM-like architecture. Our architecture is still simple and achieves better performance on the tasks that we tested on. We also introduce a new RNN performance benchmark that uses the handwritten digits and stresses several important network capabilities. Andrew Pulver, Siwei Lyu |
IJCNN | 2 |
| 2017 | Learning with Average Top-k LossabstractIn this work, we introduce the average top-$k$ (\atk) loss as a new ensemble loss for supervised learning. The \atk loss provides a natural generalization of the two widely used ensemble losses, namely the average loss and the maximum loss. Furthermore, the \atk loss combines the advantages of them and can alleviate their corresponding drawbacks to better adapt to different data distributions. We show that the \atk loss affords an intuitive interpretation that reduces the penalty of continuous and convex individual losses on correctly classified data. The \atk loss can lead to convex optimization problems that can be solved effectively with conventional sub-gradient based method. We further study the Statistical Learning Theory of \matk by establishing its classification calibration and statistical consistency of \matk which provide useful insights on the practical choice of the parameter $k$. We demonstrate the applicability of \matk learning combined with different individual loss functions for binary and multi-class classification and regression using synthetic and real datasets. Yanbo Fan, Siwei Lyu, Yiming Ying, Bao-Gang Hu |
NIPS | 2 |
| 2017 | Multi-Camera Multi-Target Tracking with Space-Time-View Hyper-graph
Longyin Wen, Zhen Lei 0001, Ming-Ching Chang, Honggang Qi, Siwei Lyu |
Int. J. Comput. Vis. | 5 |
| 2017 | Geometric Hypergraph Learning for Visual TrackingabstractGraph-based representation is widely used in visual tracking field by finding correct correspondences between target parts in different frames. However, most graph-based trackers consider pairwise geometric relations between local parts. They do not make full use of the target's intrinsic structure, thereby making the representation easily disturbed by errors in pairwise affinities when large deformation or occlusion occurs. In this paper, we propose a geometric hypergraph learning-based tracking method, which fully exploits high-order geometric relations among multiple correspondences of parts in different frames. Then visual tracking is formulated as the mode-seeking problem on the hypergraph in which vertices represent correspondence hypotheses and hyperedges describe high-order geometric relations among correspondences. Besides, a confidence-aware sampling method is developed to select representative vertices and hyperedges to construct the geometric hypergraph for more robustness and scalability. The experiments are carried out on three challenging datasets (VOT2014, OTB100, and Deform-SOT) to demonstrate that our method performs favorably against other existing trackers. Dawei Du, Honggang Qi, Longyin Wen, Qi Tian 0001, Qingming Huang, Siwei Lyu |
IEEE Trans. Cybern. | 6 |
| 2016 | Co-Regularized PLSA for Multi-Modal LearningabstractMany learning problems in real world applications involve rich datasets comprising multiple information modalities. In this work, we study co-regularized PLSA(coPLSA) as an efficient solution to probabilistic topic analysis of multi-modal data. In coPLSA, similarities between topic compositions of a data entity across different data modalities are measured with divergences between discrete probabilities, which are incorporated as a co-regularizer to augment individual PLSA models over each data modality. We derive efficient iterative learning algorithms for coPLSA with symmetric KL, L2 and L1 divergences as co-regularizers, in each case the essential optimization problem affords simple numerical solutions that entail only matrix arithmetic operations and numerical solution of 1D nonlinear equations. We evaluate the performance of the coPLSA algorithms on text/image cross-modal retrieval tasks, on which they show competitive performance with state-of-the-art methods. Xin Wang 0045, Ming-Ching Chang, Yiming Ying, Siwei Lyu |
AAAI | 4 |
| 2016 | Constrained Submodular Minimization for Missing Labels and Class Imbalance in Multi-label LearningabstractIn multi-label learning, there are two main challenges: missing labels and class imbalance (CIB). The former assumes that only a partial set of labels are provided for each training instance while other labels are missing. CIB is observed from two perspectives: first, the number of negative labels of each instance is much larger than its positive labels; second, the rate of positive instances (i.e. the number of positive instances divided by the total number of instances) of different classes are significantly different. Both missing labels and CIB lead to significant performance degradation. In this work, we propose a new method to handle these two challenges simultaneously. We formulate the problem as a constrained submodular minimization that is composed of a submodular objective function that encourages label consistency and smoothness, as well as, class cardinality bound constraints to handle class imbalance. We further present a convex approximation based on the Lovasz extension of submodular functions, leading to a linear program, which can be efficiently solved by the alternative direction method of multipliers (ADMM). Experimental results on several benchmark datasets demonstrate the improved performance of our method over several state-of-the-art methods. Baoyuan Wu, Siwei Lyu, Bernard Ghanem |
AAAI | 2 |
| 2016 | Fast Convergence of Online Pairwise Learning AlgorithmsabstractPairwise learning usually refers to a learning task which involves a loss function depending on pairs of examples, among which most notable ones are bipartite ranking, metric learning and AUC maximization. In this paper, we focus on online learning algorithms for pairwise learning problems without strong convexity, for which all previously known algorithms achieve a convergence rate of \mathcalO(1/\sqrtT) after T iterations. In particular, we study an online learning algorithm for pairwise learning with a least-square loss function in an unconstrained setting. We prove that the convergence of its last iterate can converge to the desired minimizer at a rate arbitrarily close to \mathcalO(1/T) up to logarithmic factor. The rates for this algorithm are established in high probability under the assumptions of polynomially decaying step sizes. Martin Boissier 0002, Siwei Lyu, Yiming Ying, Ding-Xuan Zhou |
AISTATS | 2 |
| 2016 | Efficient large-scale photometric reconstruction using Divide-Recon-Fuse 3D Structure from MotionabstractWe propose an efficient framework for large-scale 3D reconstruction from a large set of photos following the Structure-from-Motion (SfM) paradigm with divide-conquer and fusion. Our main novelty is to ensure commonality from overlaps between image sets corresponding to their reconstructions, which facilitates effective stitching and fusion. Specifically, such commonality is ensured by selecting a set of duplicated images (which are termed anchor images) in adjacent image sets prior to the 3D reconstruction. The anchor images can assist accurate fusion of the 3D point clouds. We describe an efficient RANSAC scheme for pairwise stitching. Our method is intuitively scalable to large site reconstruction via subdivision and fusion following a graph construct. We further describe another RANSAC algorithm to improve loop closure in our anchor image approach. Experimental results on reconstructing a large portion of a university campus demonstrate the efficacy of our method. Yueming Yang, Ming-Ching Chang, Longyin Wen, Peter H. Tu, Honggang Qi, Siwei Lyu |
AVSS | 6 |
| 2016 | Stochastic Online AUC MaximizationabstractArea under ROC (AUC) is a metric which is widely used for measuring the classification performance for imbalanced data. It is of theoretical and practical interest to develop online learning algorithms that maximizes AUC for large-scale data. A specific challenge in developing online AUC maximization algorithm is that the learning objective function is usually defined over a pair of training examples of opposite classes, and existing methods achieves on-line processing with higher space and time complexity. In this work, we propose a new stochastic online algorithm for AUC maximization. In particular, we show that AUC optimization can be equivalently formulated as a convex-concave saddle point problem. From this saddle representation, a stochastic online algorithm (SOLAM) is proposed which has time and space complexity of one datum. We establish theoretical convergence of SOLAM with high probability and demonstrate its effectiveness and efficiency on standard benchmark datasets. Yiming Ying, Longyin Wen, Siwei Lyu |
NIPS | 3 |
| 2016 | Exploiting Hierarchical Dense Structures on Hypergraphs for Multi-Object TrackingabstractMost multi-object tracking algorithms are developed within the tracking-by-detection framework that consider the pairwise appearance similarities between detection responses or tracklets within a limited temporal window, and thus less effective in handling long-term occlusions or distinguishing spatially close targets with similar appearance in crowded scenes. In this work, we propose an algorithm that formulates the multi-object tracking task as one to exploit hierarchical dense structures on an undirected hypergraph constructed based on tracklet affinity. The dense structures indicate a group of vertices that are inter-connected with a set of hyperedges with high affinity values. The appearance and motion similarities among multiple tracklets across the spatio-temporal domain are considered globally by exploiting high-order similarities rather than pairwise ones, thereby facilitating distinguish spatially close targets with similar appearance. In addition, the hierarchical design of the optimization process helps the proposed tracking algorithm handle long-term occlusions robustly. Extensive experiments on various challenging datasets of both multi-pedestrian and multi-face tracking tasks, demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods. Longyin Wen, Zhen Lei 0001, Siwei Lyu, Stan Z. Li, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Online Deformable Object Tracking Based on Structure-Aware Hyper-GraphabstractRecent advances in online visual tracking focus on designing part-based model to handle the deformation and occlusion challenges. However, previous methods usually consider only the pairwise structural dependences of target parts in two consecutive frames rather than the higher order constraints in multiple frames, making them less effective in handling large deformation and occlusion challenges. This paper describes a new and efficient method for online deformable object tracking. Different from most existing methods, this paper exploits higher order structural dependences of different parts of the tracking target in multiple consecutive frames. We construct a structure-aware hyper-graph to capture such higher order dependences, and solve the tracking problem by searching dense subgraphs on it. Furthermore, we also describe a new evaluating data set for online deformable object tracking (the Deform-SOT data set), which includes 50 challenging sequences with full annotations that represent realistic tracking challenges, such as large deformations and severe occlusions. The experimental result of the proposed method shows considerable improvement in performance over the state-of-the-art tracking methods. Dawei Du, Honggang Qi, Wenbo Li 0001, Longyin Wen, Qingming Huang, Siwei Lyu |
IEEE Trans. Image Process. | 6 |
| 2016 | Measuring and Predicting Visual Importance of Similar ObjectsabstractSimilar objects are ubiquitous and abundant in both natural and artificial scenes. Determining the visual importance of several similar objects in a complex photograph is a challenge for image understanding algorithms. This study aims to define the importance of similar objects in an image and to develop a method that can select the most important instances for an input image from multiple similar objects. This task is challenging because multiple objects must be compared without adequate semantic information. This challenge is addressed by building an image database and designing an interactive system to measure object importance from human observers. This ground truth is used to define a range of features related to the visual importance of similar objects. Then, these features are used in learning-to-rank and random forest to rank similar objects in an image. Importance predictions were validated on 5,922 objects. The most important objects can be identified automatically. The factors related to composition (e.g., size, location, and overlap) are particularly informative, although clarity and color contrast are also important. We demonstrate the usefulness of similar object importance on various applications, including image retargeting, image compression, image re-attentionizing, image admixture, and manipulation of blindness images. Yan Kong, Weiming Dong, Xing Mei, Chongyang Ma, Tong-Yee Lee, Siwei Lyu, Feiyue Huang, Xiaopeng Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2015 | Fast Online Upper Body Pose Estimation from VideoabstractEstimation of human body poses from video is an important problem in computer vision with many applications. Most existing methods for video pose estimation are offline in nature, where all frames in the video are used in the process to estimate the body pose in each frame. In this work, we describe a fast online video upper body pose estimation method (CDBN-MODEC) that is based on a conditional dynamic Bayesian network model, which predicts upper body pose in a frame without using information from future frames. Our method combines fast single image based pose estimation methods with the temporal correlation of poses between frames. We collect a new high frame rate upper body pose dataset that better reflects practical scenarios calling for fast online video pose estimation. When evaluated on this dataset and the VideoPose2 benchmark dataset, CDBN-MODEC achieves improvements in both performance and running efficiency over several state-of-art online video pose estimation methods. Ming-Ching Chang, Honggang Qi, Xin Wang 0045, Hong Cheng 0002, Siwei Lyu |
BMVC | 5 |
| 2015 | UniHIST: A unified framework for image restoration with marginal histogram constraintsabstractMarginal histograms provide valuable information for various computer vision problems. However, current image restoration methods do not fully exploit the potential of marginal histograms, in particular, their role as ensemble constraints on the marginal statistics of the restored image. In this paper, we introduce a new framework, UniHIST, to incorporate marginal histogram constraints into image restoration. The key idea of UniHIST is to minimize the discrepancy between the marginal histograms of the restored image and the reference histograms in pixel or gradient domains using the quadratic Wasserstein (W2) distance. The W2distance can be computed directly from data without resorting to density estimation. It provides a differentiable metric between marginal histograms and allows easy integration with existing image restoration methods. We demonstrate the effectiveness of UniHIST through denoising of pattern images and non-blind deconvolution of natural images. We show that UniHIST enhances restoration performance and leads to visual and quantitative improvements over existing state-of-the-art methods. Xing Mei, Weiming Dong, Bao-Gang Hu, Siwei Lyu |
CVPR | 4 |
| 2015 | Category-Blind Human Action Recognition: A Practical Recognition SystemabstractExisting human action recognition systems for 3D sequences obtained from the depth camera are designed to cope with only one action category, either single-person action or two-person interaction, and are difficult to be extended to scenarios where both action categories co-exist. In this paper, we propose the category-blind human recognition method (CHARM) which can recognize a human action without making assumptions of the action category. In our CHARM approach, we represent a human action (either a single-person action or a two-person interaction) class using a co-occurrence of motion primitives. Subsequently, we classify an action instance based on matching its motion primitive co-occurrence patterns to each class representation. The matching task is formulated as maximum clique problems. We conduct extensive evaluations of CHARM using three datasets for single-person actions, two-person interactions, and their mixtures. Experimental results show that CHARM performs favorably when compared with several state-of-the-art single-person action and two-person interaction based methods without making explicit assumptions of action category. Wenbo Li 0001, Longyin Wen, Mooi Choo Chuah, Siwei Lyu |
ICCV | 4 |
| 2015 | Improving Image Restoration with Soft-RoundingabstractSeveral important classes of images such as text, barcode and pattern images have the property that pixels can only take a distinct subset of values. This knowledge can benefit the restoration of such images, but it has not been widely considered in current restoration methods. In this work, we describe an effective and efficient approach to incorporate the knowledge of distinct pixel values of the pristine images into the general regularized least squares restoration framework. We introduce a new regularizer that attains zero at the designated pixel values and becomes a quadratic penalty function in the intervals between them. When incorporated into the regularized least squares restoration framework, this regularizer leads to a simple and efficient step that resembles and extends the rounding operation, which we term as soft-rounding. We apply the soft-rounding enhanced solution to the restoration of binary text/barcode images and pattern images with multiple distinct pixel values. Experimental results show that soft-rounding enhanced restoration methods achieve significant improvement in both visual quality and quantitative measures (PSNR and SSIM). Furthermore, we show that this regularizer can also benefit the restoration of general natural images. Xing Mei, Honggang Qi, Bao-Gang Hu, Siwei Lyu |
ICCV | 4 |
| 2015 | ML-MG: Multi-label Learning with Missing Labels Using a Mixed GraphabstractThis work focuses on the problem of multi-label learning with missing labels (MLML), which aims to label each test instance with multiple class labels given training instances that have an incomplete/partial set of these labels (i.e. some of their labels are missing). To handle missing labels, we propose a unified model of label dependencies by constructing a mixed graph, which jointly incorporates (i) instance-level similarity and class co-occurrence as undirected edges and (ii) semantic label hierarchy as directed edges. Unlike most MLML methods, We formulate this learning problem transductively as a convex quadratic matrix optimization problem that encourages training label consistency and encodes both types of label dependencies (i.e. undirected and directed edges) using quadratic terms and hard linear constraints. The alternating direction method of multipliers (ADMM) can be used to exactly and efficiently solve this problem. To evaluate our proposed method, we consider two popular applications (image and video annotation), where the label hierarchy can be derived from Wordnet. Experimental results show that our method achieves a significant improvement over state-of-the-art methods in performance and robustness to missing labels. Baoyuan Wu, Siwei Lyu, Bernard Ghanem |
ICCV | 2 |
| 2015 | Seeing as it happens: Real time 3D video event visualizationabstractWe present a video event visualization system that can render steerable 3D views of tracked targets onto a reconstructed 3D site representation. The framework takes object tracking meta-data generated from a multi-camera event tracking system as input and produces an immersive 3D playback as a representation of the observation. This 3D representation can provide users seeing as it happens of the events for surveillance applications. Our system can further “animate” the virtual viewing camera and generate a first-person immersive playback of the event, either from the trajectory of a specified real-world target or from a virtual avatar. Such synthetic view can provide additional insights for event recognition for on-line monitoring, investigation, and forensic applications. Yueming Yang, Ming-Ching Chang, Peter H. Tu, Siwei Lyu |
ICIP | 4 |
| 2015 | State estimation for dynamic systems with intermittent contactabstractDynamic system states estimation, such as object pose and contact states estimation, is essential for robots to perform manipulation tasks. In order to make accurate estimation, the state transition model needs to be physically correct. Complementarity formulations of the dynamics are widely used for describing rigid body physical behaviors in the simulation field, which makes it a good state transition model for dynamic system states estimation problem. However, the non-smoothness of complementarity models and the high dimensionality of the dynamic system make the estimation problem challenging. In this paper, we propose a particle filtering framework that solves the estimation problem by sampling the discrete contact states using contact graphs and collision detection algorithms, and by estimating the continuous states through a Kalman filter. This method exploits the piecewise continuous property of complementarity problems and reduces the dimension of the sampling space compared with sampling the high dimensional continuous states space. We demonstrate that this method makes stable and reliable estimation in physical experiments. Shuai Li 0015, Siwei Lyu, Jeffrey C. Trinkle |
ICRA | 2 |
| 2015 | A comparative study of contact models for contact-aware state estimationabstractWe study the contact-aware state estimation (CASE) problem, i.e., the problem of estimating the state of an object while it is being actively manipulated by a robot. Several researchers have developed particle filters for this problem. They estimate the state (pose and velocity) of manipulated objects, some physical properties (such as mass and shape), and contact information (such as, gain or loss of contact and transitions between sliding and sticking). However, the effects of various contact and noise models, which can have a huge impact on the estimation results, are obfuscated by implementation details. In this paper, we study the CASE problem arising from a simple pushing task with the goal of shedding light on the fundamental contact modeling choices. Specifically, we evaluate four particle filters based upon four probabilistic state transition models generated from a deterministic multibody dynamics models with rigid or compliant contacts, each of which is augmented by one of two different noise models. Comparisons of these state transition models are carried out through the analysis of real and simulated experiments, the results of which, provide guidance to filter designers. Shuai Li 0015, Siwei Lyu, Jeffrey C. Trinkle, Wolfram Burgard |
IROS | 2 |
| 2015 | Multi-label learning with missing labels for image annotation and facial action unit recognition
Baoyuan Wu, Siwei Lyu, Bao-Gang Hu |
Pattern Recognit. | 2 |
| 2014 | Using Projection Kurtosis Concentration of Natural Images for Blind Noise Covariance Matrix EstimationabstractKurtosis of 1D projections provides important statistical characteristics of natural images. In this work, we first provide a theoretical underpinning to a recently observed phenomenon known as projection kurtosis concentration that the kurtosis of natural images over different band-pass channels tend to concentrate around a typical value. Based on this analysis, we further describe a new method to estimate the covariance matrix of correlated Gaussian noise from a noise corrupted image using random band-pass filters. We demonstrate the effectiveness of our blind noise covariance matrix estimation method on natural images. Siwei Lyu |
CVPR | 2 |
| 2014 | Variational EM Learning of DSBNs with Conditional Deep Boltzmann Machines
Siwei Lyu |
ICANN | 2 |
| 2014 | Non-blind image restoration with symmetric generalized Pareto priorsabstractThis paper presents a new non-blind image restoration method based on the symmetric generalized Pareto (SGP) prior, which models the heavy-tailed distributions of gradients for natural images. Through experiments we show that the SGP model achieves log likelihood scores comparable to the hyper-Laplacian model when fitted to gradients and other band-pass filter responses. More importantly, when incorporated into a Bayesian MAP framework for non-blind image restoration, the SGP model leads to a closed-form solution for a per-pixel subproblem, which affords computational advantages in comparison with the numerical solutions induced from the hyper-Laplacian model. Experimental results show that our method is comparable to existing methods in restoration quality and processing speed. Xing Mei, Bao-Gang Hu, Siwei Lyu |
ICIP | 3 |
| 2014 | Blind estimation of pixel brightness transformabstractPixel brightness transforms (PBT), examples of which include gamma correction, sigmoid stretching and histogram equalization, are common operations on digital images, and it is practically useful to estimate such transforms directly from an image. In this work, we describe an effective and efficient method to estimate PBT from images, which takes advantage of the nature of PBT as a mapping between integral pixel values and the distinct characteristics it introduces to the pixel value (PV) histograms of the transformed image. Our method recovers the original PV histogram and the PBT simultaneously with an efficient iterative algorithm, and can effectively handle perturbations due to noise and compression. We perform experimental evaluation to demonstrate the efficacy and efficiency of the proposed method. Siwei Lyu |
ICIP | 2 |
| 2014 | Exposing Region Splicing Forgeries with Blind Local Noise Estimation
Siwei Lyu, Xunyu Pan |
Int. J. Comput. Vis. | 1 |
| 2013 | Simultaneous Clustering and Tracklet Linking for Multi-face Tracking in VideosabstractWe describe a novel method that simultaneously clusters and associates short sequences of detected faces (termed as face track lets) in videos. The rationale of our method is that face track let clustering and linking are related problems that can benefit from the solutions of each other. Our method is based on a hidden Markov random field model that represents the joint dependencies of cluster labels and track let linking associations. We provide an efficient algorithm based on constrained clustering and optimal matching for the simultaneous inference of cluster labels and track let associations. We demonstrate significant improvements on the state-of-the-art results in face tracking and clustering performances on several video datasets. Baoyuan Wu, Siwei Lyu, Bao-Gang Hu |
ICCV | 2 |
| 2013 | A dynamic Bayesian approach to real-time estimation and filtering in grasp acquisitionabstractIn this work, we develop a general solution to a broad class of grasping and manipulation problems that we term as C-SLAM for contact simultaneous localization and modeling, where the robots need to accurately track the motions of the contacted bodies and the locations of contacts, while simultaneously estimating important system parameters, such as body dimensions, masses and friction coefficients between contacting surfaces. Our solution framework is based on a dynamic Bayesian inference framework, and hence, we refer to it as Dynamic Bayesian C-SLAM (DBC-SLAM). DBC-SLAM combines an NCP-based dynamic model with the dynamic Bayesian network, and incorporates model parameter estimation as an intrinsic part of the overall inference procedure. We show two preliminary “proof-of-concept” examples that demonstrate the use of DBC-SLAM in robotic contact tasks. Li Zhang 0130, Siwei Lyu, Jeffrey C. Trinkle |
ICRA | 2 |
| 2013 | Deep Feature Learning Using Target Priors with Applications in ECoG Signal Decoding for BCI
Zuoguan Wang, Siwei Lyu, Gerwin Schalk |
IJCAI | 2 |
| 2013 | On Algorithms for Sparse Multi-factor NMFabstractNonnegative matrix factorization (NMF) is a popular data analysis method, the objective of which is to decompose a matrix with all nonnegative components into the product of two other nonnegative matrices. In this work, we describe a new simple and efficient algorithm for multi-factor nonnegative matrix factorization problem ({mfNMF}), which generalizes the original NMF problem to more than two factors. Furthermore, we extend the mfNMF algorithm to incorporate a regularizer based on Dirichlet distribution over normalized columns to encourage sparsity in the obtained factors. Our sparse NMF algorithm affords a closed form and an intuitive interpretation, and is more efficient in comparison with previous works that use fix point iterations. We demonstrate the effectiveness and efficiency of our algorithms on both synthetic and real data sets. Siwei Lyu, Xin Wang 0045 |
NIPS | 1 |
| 2012 | Boosting with Side Information
Jixu Chen, Xiaoming Liu 0002, Siwei Lyu |
ACCV (1) | 3 |
| 2012 | Detecting splicing in digital audios using local noise level estimationabstractOne common form of tampering in digital audio signals is known as splicing, where sections from one audio is inserted to another audio. In this paper, we propose an effective splicing detection method for audios. Our method achieves this by detecting abnormal differences in the local noise levels in an audio signal. This estimation of local noise levels is based on an observed property of audio signals that they tend to have kurtosis close to a constant in the band-pass filtered domain. We demonstrate the efficacy and robustness of the proposed method using both synthetic and realistic audio splicing forgeries. Xunyu Pan, Siwei Lyu |
ICASSP | 3 |
| 2012 | Exposing image splicing with inconsistent local noise variancesabstractImage splicing is a simple and common image tampering operation, where a selected region from an image is pasted into another image with the aim to change its content. In this paper, based on the fact that images from different origins tend to have different amount of noise introduced by the sensors or post-processing steps, we describe an effective method to expose image splicing by detecting inconsistencies in local noise variances. Our method estimates local noise variances based on an observation that kurtosis values of natural images in band-pass filtered domains tend to concentrate around a constant value, and is accelerated by the use of integral image. We demonstrate the efficacy and robustness of our method based on several sets of forged images generated with image splicing. Xunyu Pan, Siwei Lyu |
ICCP | 3 |
| 2012 | Learning with Target PriorabstractIn the conventional approaches for supervised parametric learning, relations between data and target variables are provided through training sets consisting of pairs of corresponded data and target variables. In this work, we describe a new learning scheme for parametric learning, in which the target variables $\y$ can be modeled with a prior model $p(\y)$ and the relations between data and target variables are estimated through $p(\y)$ and a set of uncorresponded data $\x$ in training. We term this method as learning with target priors (LTP). Specifically, LTP learning seeks parameter $\t$ that maximizes the log likelihood of $f_\t(\x)$ on a uncorresponded training set with regards to $p(\y)$. Compared to the conventional (semi)supervised learning approach, LTP can make efficient use of prior knowledge of the target variables in the form of probabilistic distributions, and thus removes/reduces the reliance on training data in learning. Compared to the Bayesian approach, the learned parametric regressor in LTP can be more efficiently implemented and deployed in tasks where running efficiency is critical, such as on-line BCI signal decoding. We demonstrate the effectiveness of the proposed approach on parametric regression tasks for BCI signal decoding and pose estimation from video. Zuoguan Wang, Siwei Lyu, Gerwin Schalk |
NIPS | 2 |
| 2011 | Unifying Non-Maximum Likelihood Learning Objectives with Minimum KL ContractionabstractWhen used to learn high dimensional parametric probabilistic models, the clas- sical maximum likelihood (ML) learning often suffers from computational in- tractability, which motivates the active developments of non-ML learning meth- ods. Yet, because of their divergent motivations and forms, the objective func- tions of many non-ML learning methods are seemingly unrelated, and there lacks a unified framework to understand them. In this work, based on an information geometric view of parametric learning, we introduce a general non-ML learning principle termed as minimum KL contraction, where we seek optimal parameters that minimizes the contraction of the KL divergence between the two distributions after they are transformed with a KL contraction operator. We then show that the objective functions of several important or recently developed non-ML learn- ing methods, including contrastive divergence [12], noise-contrastive estimation [11], partial likelihood [7], non-local contrastive objectives [31], score match- ing [14], pseudo-likelihood [3], maximum conditional likelihood [17], maximum mutual information [2], maximum marginal likelihood [9], and conditional and marginal composite likelihood [24], can be unified under the minimum KL con- traction framework with different choices of the KL contraction operators. Siwei Lyu |
NIPS | 1 |
| 2011 | Dependency Reduction with Divisive Normalization: Justification and EffectivenessabstractEfficient coding transforms that reduce or remove statistical dependencies in natural sensory signals are important for both biology and engineering. In recent years, divisive normalization (DN) has been advocated as a simple and effective nonlinear efficient coding transform. In this work, we first elaborate on the theoretical justification for DN as an efficient coding transform. Specifically, we use the multivariate t model to represent several important statistical properties of natural sensory signals and show that DN approximates the optimal transforms that eliminate statistical dependencies in the multivariate t model. Second, we show that several forms of DN used in the literature are equivalent in their effects as efficient coding transforms. Third, we provide a quantitative evaluation of the overall dependency reduction performance of DN for both the multivariate t models and natural sensory signals. Finally, we find that statistical dependencies in the multivariate t model and natural sensory signals are increased by the DN transform with low-input dimensions. This implies that for DN to be an effective efficient coding transform, it has to pool over a sufficiently large number of inputs. Siwei Lyu |
Neural Comput. | 1 |
| 2010 | Nonnegative matrix factorization with matrix exponentiationabstractNonnegative matrix factorization (NMF) has been successfully applied to different domains as a technique able to find part-based linear representations for nonnegative data. However, when extra constraints are incorporated into NMF, simple gradient descent optimization can be inefficient for high-dimensional problems, due to the overhead to enforce the nonnegativity constraints. We describe an alternative formulation based on matrix exponentiation, where the nonnegativity constraints are enforced implicitly, and a direct gradient descent algorithm can have better efficiency. In numerical experiments, such a reformulation leads to significant improvement in running time. Siwei Lyu |
ICASSP | 1 |
| 2010 | Detecting image region duplication using SIFT featuresabstractRegion duplication is a common form of image manipulation where part of an image is pasted to another location to conceal undesirable contents. Most existing methods to detect region duplication are based on finding exact copies of pixel blocks, which cannot handle cases when a region is scaled or rotated before pasted to a new location. In this work, we describe a new detection method based on matching image SIFT features. The robustness of the SIFT features with regards to local transforms renders this method able to detect general region duplications with efficient computation. The effectiveness of this method is demonstrated with experimental results, both qualitatively and quantitatively in terms of the detection accuracy and the false positive rate. Xunyu Pan, Siwei Lyu |
ICASSP | 2 |
| 2010 | Divisive Normalization: Justification and Effectiveness as Efficient Coding TransformabstractDivisive normalization (DN) has been advocated as an effective nonlinear {\em efficient coding} transform for natural sensory signals with applications in biology and engineering. In this work, we aim to establish a connection between the DN transform and the statistical properties of natural sensory signals. Our analysis is based on the use of multivariate {\em t} model to capture some important statistical properties of natural sensory signals. The multivariate {\em t} model justifies DN as an approximation to the transform that completely eliminates its statistical dependency. Furthermore, using the multivariate {\em t} model and measuring statistical dependency with multi-information, we can precisely quantify the statistical dependency that is reduced by the DN transform. We compare this with the actual performance of the DN transform in reducing statistical dependencies of natural sensory signals. Our theoretical analysis and quantitative evaluations confirm DN as an effective efficient coding transform for natural sensory signals. On the other hand, we also observe a previously unreported phenomenon that DN may increase statistical dependencies when the size of pooling is small. Siwei Lyu |
NIPS | 1 |
| 2010 | Region Duplication Detection Using Image Feature MatchingabstractRegion duplication is a simple and effective operation to create digital image forgeries, where a continuous portion of pixels in an image, after possible geometrical and illumination adjustments, are copied and pasted to a different location in the same image. Most existing region duplication detection methods are based on directly matching blocks of image pixels or transform coefficients, and are not effective when the duplicated regions have geometrical or illumination distortions. In this work, we describe a new region duplication detection method that is robust to distortions of the duplicated regions. Our method starts by estimating the transform between matched scale invariant feature transform (SIFT) keypoints, which are insensitive to geometrical and illumination distortions, and then finds all pixels within the duplicated regions after discounting the estimated transforms. The proposed method shows effective detection on an automatically synthesized forgery image database with duplicated and distorted regions. We further demonstrate its practical performance with several challenging forgery images created with state-of-the-art tools. Xunyu Pan, Siwei Lyu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2009 | An implicit Markov random field model for the multi-scale oriented representations of natural imagesabstractIn this paper, we describe a new Markov random field (MRF) model for natural images in multiscale oriented representations. The MRF in this model is specified with the singleton conditional densities (the density of one subband coefficient given its Markovian neighbors), while the clique potentials and joint density of this model are implicitly defined. The singleton conditional densities are chosen to have maximum entropy and consistent with observed statistical properties of natural images. We then describe parameter learning for this model, and a sparse prior to choose optimal model structure. Using this model as image prior, we develop an iterative image denoising method, and a solution to restoring images with missing blocks of subband coefficients. Siwei Lyu |
CVPR | 1 |
| 2009 | Interpretation and Generalization of Score Matching
Siwei Lyu |
UAI | 1 |
| 2009 | Nonlinear Extraction of Independent Components of Natural Images Using Radial GaussianizationabstractWe consider the problem of efficiently encoding a signal by transforming it to a new representation whose components are statistically independent. A widely studied linear solution, known as independent component analysis (ICA), exists for the case when the signal is generated as a linear transformation of independent nongaussian sources. Here, we examine a complementary case, in which the source is nongaussian and elliptically symmetric. In this case, no invertible linear transform suffices to decompose the signal into independent components, but we show that a simple nonlinear transformation, which we call radial gaussianization (RG), is able to remove all dependencies. We then examine this methodology in the context of natural image statistics. We first show that distributions of spatially proximal bandpass filter responses are better described as elliptical than as linearly transformed independent sources. Consistent with this, we demonstrate that the reduction in dependency achieved by applying RG to either nearby pairs or blocks of bandpass filter responses is significantly greater than that achieved by ICA. Finally, we show that the RG transformation may be closely approximated by divisive normalization, which has been used to model the nonlinear response properties of visual neurons. Siwei Lyu, Eero P. Simoncelli |
Neural Comput. | 1 |
| 2009 | Modeling Multiscale Subbands of Photographic Images with Fields of Gaussian Scale MixturesabstractThe local statistical properties of photographic images, when represented in a multi-scale basis, have been described using Gaussian scale mixtures. Here, we use this local description as a substrate for constructing a global field of Gaussian scale mixtures (FoGSMs). Specifically, we model multi-scale subbands as a product of an exponentiated homogeneous Gaussian Markov random field (hGMRF) and a second independent hGMRF. We show that parameter estimation for this model is feasible, and that samples drawn from a FoGSM model have marginal and joint statistics similar to subband coefficients of photographic images. We develop an algorithm for removing additive Gaussian white noise based on the FoGSM model, and demonstrate denoising performance comparable with state-of-the-art methods. Siwei Lyu, Eero P. Simoncelli |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Nonlinear image representation using divisive normalizationabstractIn this paper, we describe a nonlinear image representation based on divisive normalization that is designed to match the statistical properties of photographic images, as well as the perceptual sensitivity of biological visual systems. We decompose an image using a multi-scale oriented representation, and use Student's t as a model of the dependencies within local clusters of coefficients. We then show that normalization of each coefficient by the square root of a linear combination of the amplitudes of the coefficients in the cluster reduces statistical dependencies. We further show that the resulting divisive normalization transform is invertible and provide an efficient iterative inversion algorithm. Finally, we probe the statistical and perceptual advantages of this image representation by examining its robustness to added noise, and using it to enhance image contrast. Siwei Lyu, Eero P. Simoncelli |
CVPR | 1 |
| 2008 | Reducing statistical dependencies in natural signals using radial GaussianizationabstractWe consider the problem of efficiently encoding a signal by transforming it to a new representation whose components are statistically independent. A widely studied linear solution, independent components analysis (ICA), exists for the case when the signal is generated as a linear transformation of independent non- Gaussian sources. Here, we examine a complementary case, in which the source is non-Gaussian but elliptically symmetric. In this case, no linear transform suffices to properly decompose the signal into independent components, but we show that a simple nonlinear transformation, which we call radial Gaussianization (RG), is able to remove all dependencies. We then demonstrate this methodology in the context of natural signal statistics. We first show that the joint distributions of bandpass filter responses, for both sound and images, are better described as elliptical than linearly transformed independent sources. Consistent with this, we demonstrate that the reduction in dependency achieved by applying RG to either pairs or blocks of bandpass filter responses is significantly greater than that achieved by PCA or ICA. Siwei Lyu, Eero P. Simoncelli |
NIPS | 1 |
| 2006 | Statistical Modeling of Images with Fields of Gaussian Scale MixturesabstractThe local statistical properties of photographic images, when represented in a multi-scale basis, have been described using Gaussian scale mixtures (GSMs). Here, we use this local description to construct a global field of Gaussian scale mixtures (FoGSM). Specifically, we model subbands of wavelet coefficients as a product of an exponentiated homogeneous Gaussian Markov random field (hGMRF) and a second independent hGMRF. We show that parameter estimation for FoGSM is feasible, and that samples drawn from an estimated FoGSM model have marginal and joint statistics similar to wavelet coefficients of photographic images. We develop an algorithm for image denoising based on the FoGSM model, and demonstrate substantial improvements over current state-ofthe-art denoising method based on the local GSM model. Many successful methods in image processing and computer vision rely on statistical models for images, and it is thus of continuing interest to develop improved models, both in terms of their ability to precisely capture image structures, and in terms of their tractability when used in applications. Constructing such a model is difficult, primarily because of the intrinsic high dimensionality of the space of images. Two simplifying assumptions are usually made to reduce model complexity. The first is Markovianity: the density of a pixel conditioned on a small neighborhood, is assumed to be independent from the rest of the image. The second assumption is homogeneity: the local density is assumed to be independent of its absolute position within the image. The set of models satisfying both of these assumptions constitute the class of homogeneous Markov random fields (hMRFs). Over the past two decades, studies of photographic images represented with multi-scale multiorientation image decompositions (loosely referred to as "wavelets") have revealed striking nonGaussian regularities and inter and intra-subband dependencies. For instance, wavelet coefficients generally have highly kurtotic marginal distributions [1, 2], and their amplitudes exhibit strong correlations with the amplitudes of nearby coefficients [3, 4]. One model that can capture the nonGaussian marginal behaviors is a product of non-Gaussian scalar variables [5]. A number of authors have developed non-Gaussian MRF models based on this sort of local description [6, 7, 8], among which the recently developed fields of experts model [7] has demonstrated impressive performance in denoising (albeit at an extremely high computational cost in learning model parameters). An alternative model that can capture non-Gaussian local structure is a scale mixture model [9, 10, 11]. An important special case is Gaussian scale mixtures (GSM), which consists of a Gaussian random vector whose amplitude is modulated by a hidden scaling variable. The GSM model provides a particularly good description of local image statistics, and the Gaussian substructure of the model leads to efficient algorithms for parameter estimation and inference. Local GSM-based methods represent the current state-of-the-art in image denoising [12]. The power of GSM models should be substantially improved when extended to describe more than a small neighborhood of wavelet coefficients. To this end, several authors have embedded local Gaussian mixtures into tree-structured MRF models [e.g., 13, 14]. In order to maintain tractability, these models are arranged such that coefficients are grouped in non-overlapping clusters, allowing a graphical probability model with no loops. Despite their global consistency, the artificially imposed cluster boundaries lead to substantial artifacts in applications such as denoising. In this paper, we use a local GSM as a basis for a globally consistent and spatially homogeneous field of Gaussian scale mixtures (FoGSM). Specifically, the FoGSM is formulated as the product of two mutually independent MRFs: a positive multiplier field obtained by exponentiating a homogeneous Gaussian MRF (hGMRF), and a second hGMRF. We develop a parameter estimation procedure, and show that the model is able to capture important statistical regularities in the marginal and joint wavelet statistics of a photographic image. We apply the FoGSM to image denoising, demonstrating substantial improvement over the previous state-of-the-art results obtained with a local GSM model. 1 Gaussian scale mixtures A GSM random vector x is formed as the product of a zero-mean Gaussian random vector u and an d d independent random variable z, as x = zu, where = denotes equality in distribution. The density of x is determined by the covariance of the Gaussian vector, , and the density of the multiplier, p z (z), through the integral - T -1 p z z xx 1 exp (1) p(x) = Nx (0, z) pz (z)dz z (z)d z. 2z z|| A key property of GSMs is that when z determines the scale of the conditional variance of x given z, wich is a Gaussian variable with zero mean and covariance z. In addition, the normalized variable h x z is a zero mean Gaussian with covariance matrix . The GSM model has been used to describe the marginal and joint densities of local clusters of wavelet coefficients, both within and across subbands [9], where the embedded Gaussian structure affords simple and efficient computation. This local GSM model has been be used for denoising, by independently estimating each coefficient conditioned on its surrounding cluster [12]. This method achieves state-of-the-art performances, despite the fact that treating overlapping clusters as independent does not give rise to a globally consistent statistical model that satisfies all the local constraints. 2 Fields of Gaussian scale mixtures In this section, we develop fields of Gaussian scale mixtures (FoGSM) as a framework for modeling wavelet coefficients of photographic images. Analogous to the local GSM model, we use a latent multiplier field to modulate a homogeneous Gaussian MRF (hGMRF). Formally, we define a FoGSM x as the product of two mutually independent MRFs, d x = u z, (2) where u is a zero-mean hGMRF, and z is a field of positive multipliers that control the local coefficient variances. The operator denotes element-wise multiplication, and the square root operation is applied to each component. Note that x has a one-dimensional GSM marginal distributions, while its components have dependencies captured by the MRF structures of u and z. Analogous to the local GSM, when conditioned on z, x is an inhomogeneous GMRF | | - - -1 x = 1 T D -1 1 Qu | Qu | p(x|z) x (x z)T Qu (x i exp z Qu D z i exp zi 2 zi 2 , z) (3) where Qu is the inverse covariance matrix of u (also known as the precision matrix), and D() denotes the operator that form a diagonal matrix from an input vector. Note also that the elementwise division of the two fields, x z, yields a hGMRF with precision matrix Q u . To complete the FoGSM model, we need to specify the structure of the multiplier field z. For tractability, we use another hGMRF as a substrate, and map it into positive values by exponentiation, Siwei Lyu, Eero P. Simoncelli |
NIPS | 1 |
| 2006 | Steganalysis using higher-order image statisticsabstractTechniques for information hiding (steganography) are becoming increasingly more sophisticated and widespread. With high-resolution digital images as carriers, detecting hidden messages is also becoming considerably more difficult. We describe a universal approach to steganalysis for detecting the presence of hidden messages embedded within digital images. We show that, within multiscale, multiorientation image decompositions (e.g., wavelets), first- and higher-order magnitude and phase statistics are relatively consistent across a broad range of images, but are disturbed by the presence of embedded hidden messages. We show the efficacy of our approach on a large collection of images, and on eight different steganographic embedding algorithms. Siwei Lyu, Hany Farid |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2005 | Infomax BoostingabstractIn this paper, we described an efficient feature pursuit scheme for boosting. The proposed method is based on the infomax principle, which seeks optimal feature that achieves maximal mutual information with class labels. Direct feature pursuit with infomax is computationally prohibitive, so an efficient gradient ascent algorithm is further proposed, based on the quadratic mutual information, non-parametric density estimation and fast Gauss transform. The feature pursuit process is integrated into a boosting framework as infomax boosting. The performance of a face detector based on infomax boosting is reported. Siwei Lyu |
CVPR (1) | 1 |
| 2005 | Mercer Kernels for Object Recognition with Local FeaturesabstractA new class of kernels for object recognition based on local image feature representations are introduced in this paper. These kernels satisfy the Mercer condition and incorporate multiple types of local features and semilocal constraints between them. Experimental results of SVM classifiers coupled with the proposed kernels are reported on recognition tasks with the COIL-100 database and compared with existing methods. The proposed kernels achieved competitive performance and were robust to changes in object configurations and image degradations. Siwei Lyu |
CVPR (2) | 1 |
| 2005 | A Kernel Between Unordered Sets of Data: The Gaussian Mixture Approach
Siwei Lyu |
ECML | 1 |
| 2005 | Automatic image orientation determination with natural image statisticsabstractIn this paper, we propose a new method for automatically determining image orientations. This method is based on a set of natural image statistics collected from a multi-scale multi-orientation image decomposition (e.g., wavelets). From these statistics, a two-stage hierarchal classification with multiple binary SVM classifiers is employed to determine image orientation. The proposed method is evaluated and compared to existing methods with experiments performed on 18040 natural images, where it showed promising performance. Siwei Lyu |
ACM Multimedia | 1 |