Yue Wu 0001

dblp:41/5979-1 · also Rex Yue Wu · DBLP profile ↗
← Back
54ranked-venue papers
25as first author
17since 2021 · last 2024
0000-0003-0126-3614ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 12 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 13 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author
YearPublicationVenuePosition
2024 SmartPlay : A Benchmark for LLMs as Intelligent Agents
abstract
Recent large language models (LLMs) have demonstrated great potential toward intelligent agents and next-gen automation, but there currently lacks a systematic benchmark for evaluating LLMs' abilities as agents. We introduce SmartPlay: both a challenging benchmark and a methodology for evaluating LLMs as agents. SmartPlay consists of 6 different games, including Rock-Paper-Scissors, Tower of Hanoi, Minecraft. Each game features a unique setting, providing up to 20 evaluation settings and infinite environment variations. Each game in SmartPlay uniquely challenges a subset of 9 important capabilities of an intelligent LLM agent, including reasoning with object dependencies, planning ahead, spatial reasoning, learning from history, and understanding randomness. The distinction between the set of capabilities each game test allows us to analyze each capability separately. SmartPlay serves not only as a rigorous testing ground for evaluating the overall performance of LLM agents but also as a road-map for identifying gaps in current methodologies. We release our benchmark at https://github.com/microsoft/SmartPlay
Yue Wu 0001, Tom M. Mitchell, Yuanzhi Li
ICLR1
2024 Consistent Arbitrary Style Transfer Using Consistency Training and Self-Attention Module
abstract
Arbitrary style transfer (AST) has garnered considerable attention for its ability to transfer styles infinitely. Although existing methods have achieved impressive results, they may overlook style consistencies and fail to capture crucial style patterns, leading to inconsistent style transfer (ST) caused by minor disturbances. To tackle this issue, we conduct a mathematical analysis of inconsistent ST and develop a style inconsistency measure (SIM) to quantify the inconsistencies between generated images. Moreover, we propose a consistent AST (CAST) framework that effectively captures and transfers essential style features into content images. The proposed CAST framework incorporates an intersection-of-union-preserving crop (IoUPC) module to obtain style pairs with minor disturbance, a self-attention (SA) module to learn the crucial style features, and a style inconsistency loss regularization (SILR) to facilitate consistent feature learning for consistent stylization. Our proposed framework not only provides an optimal solution for consistent ST but also outperforms existing methods when embedded into the CAST framework. Extensive experiments demonstrate that the proposed CAST framework can effectively transfer style patterns while preserving consistency and achieve the state-of-the-art performance.
Yue Wu 0001, Yicong Zhou
IEEE Trans. Neural Networks Learn. Syst.2
2023 User-Controllable Arbitrary Style Transfer via Entropy Regularization
abstract
Ensuring the overall end-user experience is a challenging task in arbitrary style transfer (AST) due to the subjective nature of style transfer quality. A good practice is to provide users many instead of one AST result. However, existing approaches require to run multiple AST models or inference a diversified AST (DAST) solution multiple times, and thus they are either slow in speed or limited in diversity. In this paper, we propose a novel solution ensuring both efficiency and diversity for generating multiple user-controllable AST results by systematically modulating AST behavior at run-time. We begin with reformulating three prominent AST methods into a unified assign-and-mix problem and discover that the entropies of their assignment matrices exhibit a large variance. We then solve the unified problem in an optimal transport framework using the Sinkhorn-Knopp algorithm with a user input ε to control the said entropy and thus modulate stylization. Empirical results demonstrate the superiority of the proposed solution, with speed and stylization quality comparable to or better than existing AST and significantly more diverse than previous DAST works. Code is available at https://github.com/cplusx/eps-Assign-and-Mix.
Jiaxin Cheng, Yue Wu 0001, Ayush Jaiswal, Xu Zhang 0022, Pradeep Natarajan, Premkumar Natarajan
AAAI2
2023 Graph Generative Model for Benchmarking Graph Neural Networks
abstract
As the field of Graph Neural Networks (GNN) continues to grow, it experiences a corresponding increase in the need for large, real-world datasets to train and test new GNN models on challenging, realistic problems. Unfortunately, such graph datasets are often generated from online, highly privacy-restricted ecosystems, which makes research and development on these datasets hard, if not impossible. This greatly reduces the amount of benchmark graphs available to researchers, causing the field to rely only on a handful of publicly-available datasets. To address this problem, we introduce a novel graph generative model, Computation Graph Transformer (CGT) that learns and reproduces the distribution of real-world graphs in a privacy-controlled way. More specifically, CGT (1) generates effective benchmark graphs on which GNNs show similar task performance as on the source graphs, (2) scales to process large-scale graphs, (3) incorporates off-the-shelf privacy modules to guarantee end-user privacy of the generated graph. Extensive experiments across a vast body of graph generative models show that only our model can successfully generate privacy-controlled, synthetic substitutes of large-scale real-world graphs that can be effectively used to benchmark GNN models.
Minji Yoon, Yue Wu 0001, John Palowitch, Bryan Perozzi, Ruslan Salakhutdinov
ICML2
2023 Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals
abstract
High sample complexity has long been a challenge for RL. On the other hand, humans learn to perform tasks not only from interaction or demonstrations, but also by reading unstructured text documents, e.g., instruction manuals. Instruction manuals and wiki pages are among the most abundant data that could inform agents of valuable features and policies or task-specific environmental dynamics and reward structures. Therefore, we hypothesize that the ability to utilize human-written instruction manuals to assist learning policies for specific tasks should lead to a more efficient and better-performing agent. We propose the Read and Reward framework. Read and Reward speeds up RL algorithms on Atari games by reading manuals released by the Atari game developers. Our framework consists of a QA Extraction module that extracts and summarizes relevant information from the manual and a Reasoning module that evaluates object-agent interactions based on information from the manual. An auxiliary reward is then provided to a standard A2C RL agent, when interaction is detected. Experimentally, various RL algorithms obtain significant improvement in performance and training speed when assisted by our design. Code at github.com/Holmeswww/RnR
Yue Wu 0001, Yewen Fan, Paul Pu Liang, Amos Azaria, Yuanzhi Li, Tom M. Mitchell
NeurIPS1
2023 SPRING: Studying Papers and Reasoning to play Games
abstract
Open-world survival games pose significant challenges for AI algorithms due to their multi-tasking, deep exploration, and goal prioritization requirements. Despite reinforcement learning (RL) being popular for solving games, its high sample complexity limits its effectiveness in complex open-world games like Crafter or Minecraft. We propose a novel approach, SPRING, to read Crafter's original academic paper and use the knowledge learned to reason and play the game through a large language model (LLM). Prompted with the LaTeX source as game context and a description of the agent's current observation, our SPRING framework employs a directed acyclic graph (DAG) with game-related questions as nodes and dependencies as edges. We identify the optimal action to take in the environment by traversing the DAG and calculating LLM responses for each node in topological order, with the LLM's answer to final node directly translating to environment actions. In our experiments, we study the quality of in-context "reasoning" induced by different forms of prompts under the setting of the Crafter environment. Our experiments suggest that LLMs, when prompted with consistent chain-of-thought, have great potential in completing sophisticated high-level trajectories. Quantitatively, SPRING with GPT-4 outperforms all state-of-the-art RL baselines, trained for 1M steps, without any training. Finally, we show the potential of Crafter as a test bed for LLMs. Code at github.com/holmeswww/SPRING
Yue Wu 0001, So Yeon Min, Shrimai Prabhumoye, Yonatan Bisk, Ruslan Salakhutdinov, Amos Azaria, Tom M. Mitchell, Yuanzhi Li
NeurIPS1
2022 Enhancing Fairness in Face Detection in Computer Vision Systems by Demographic Bias Mitigation
abstract
Fairness has become an important agenda in computer vision and artificial intelligence. Recent studies have shown that many computer vision models and datasets exhibit demographic biases and proposed mitigation strategies. These works attempt to address accuracy disparity, spurious correlations, or unbalanced representations in datasets in tasks such as face recognition, verification and expression and attribute classification. These tasks, however, all require face detection as the first preprocessing step, and surprisingly, there has been little effort in identifying or mitigating biases in face detection. Biased face detectors themselves pose a threat against fair and ethical AI systems, and their biases may be further passed on to subsequent downstream tasks such as face recognition in a computer vision pipeline. This paper therefore investigates the problem of biases in face detection, focusing on accuracy disparity of detectors between demographic groups including gender, age group, and skin tone. We collect perceived demographic attributes on a popular face detection benchmark dataset, WIDER FACE, report skewed demographic distributions, and compare detection performance between groups. In order to mitigate the biases, we apply three mitigation methods that have been introduced in the recent literature and also propose two novel methods. Experimental results show that these methods are effective in reducing demographic biases. We also discuss how the effectiveness varies by demographic attributes, detection easiness, and multiple detectors, which will shed light on this new topic of addressing face detection bias.
Jianwei Feng, Prateek Singhal, Vivek Yadav, Yue Wu 0001, Pradeep Natarajan, Varsha Hedau, Jungseock Joo
AIES6
2022 FashionVLP: Vision Language Transformer for Fashion Retrieval with Feedback
abstract
Fashion image retrieval based on a query pair of reference image and natural language feedback is a challenging task that requires models to assess fashion related information from visual and textual modalities simultaneously. We propose a new vision-language transformer based model, FashionVLP, that brings the prior knowledge contained in large image-text corpora to the domain of fashion image retrieval, and combines visual information from multiple levels of context to effectively capture fashion-related information. While queries are encoded through the transformer layers, our asymmetric design adopts a novel attention-based approach for fusing target image features without involving text or transformer layers in the process. Extensive results show that FashionVLP achieves the state-of-the-art performance on benchmark datasets, with a large 23% relative improvement on the challenging FashionIQ dataset, which contains complex natural language feedback.
Sonam Goenka, Zhaoheng Zheng, Ayush Jaiswal, Rakesh Chada, Yue Wu 0001, Varsha Hedau, Pradeep Natarajan
CVPR5
2022 Few-Shot Gaze Estimation with Model Offset Predictors
abstract
Due to the variance of optical properties across different people, the performance of a person-agnostic gaze estimation model may not generalize well on a specific person. Though one may achieve better performance by training a person-specific model, it typically requires a large number of samples which is not available in real-life scenarios. Hence, few-shot gaze estimation method is preferred for the small number of samples from a target person. However, the key question is how to close the performance gap between a "few-shot" model and the "many-shot" model. In this paper, we propose to learn a person-specific offset predictor which outputs the difference between the person-agnostic model and the many-shot person-specific model with as few as one training sample. We adapt the knowledge to a new person by using the average of meta-learned offset predictors parameters as the initialization of the new offset predictor. Experiments show that the proposed few-shot person-specific model is not only closer to the corresponding many-shot person-specific model but also has better accuracy than the SOTA few-shot gaze estimation methods in multiple gaze datasets.
Jiawei Ma, Xu Zhang 0022, Yue Wu 0001, Varsha Hedau, Shih-Fu Chang
ICASSP3
2022 Neural Style Transfer With Adaptive Auto-Correlation Alignment Loss
abstract
The neural style transfer has achieved a significant improvement with deep learning methods. However, the existing methods are susceptible to lack the ability for handling the texture style transfer because of their less consideration of the textural structure from style images. To overcome this drawback, this letter presents a simple method to capture the textural structure by using an adaptive auto-correlation alignment loss function. Furthermore, we also introduce three metrics to quantitatively evaluate the performance. We qualitatively and quantitatively evaluate the proposed methods. The experimental results demonstrate the superiority of the proposed method and our method can synthesize the stylized images with rich texture style patterns.
Yue Wu 0001, Xiaofei Yang 0002, Yicong Zhou
IEEE Signal Process. Lett.2
2021 Style-Aware Normalized Loss for Improving Arbitrary Style Transfer
abstract
Neural Style Transfer (NST) has quickly evolved from single-style to infinite-style models, also known as Arbitrary Style Transfer (AST). Although appealing results have been widely reported in literature, our empirical studies on four well-known AST approaches (GoogleMagenta [14], AdaIN [19], LinearTransfer [29], and SANet [37]) show that more than 50% of the time, AST stylized images are not acceptable to human users, typically due to under- or over-stylization. We systematically study the cause of this imbalanced style transferability (IST ) and propose a simple yet effective solution to mitigate this issue. Our studies show that the IST issue is related to the conventional AST style loss, and reveal that the root cause is the equal weightage of training samples irrespective of the properties of their corresponding style images, which biases the model towards certain styles. Through investigation of the theoretical bounds of the AST style loss, we propose a new loss that largely overcomes IST . Theoretical analysis and experimental results validate the effectiveness of our loss, with over 80% relative improvement in style deception rate and 98% relatively higher preference in human evaluation.
Jiaxin Cheng, Ayush Jaiswal, Yue Wu 0001, Pradeep Natarajan, Premkumar Natarajan
CVPR3
2021 Adversarial Mask Generation for Preserving Visual Privacy
abstract
We present a privacy preserving machine learning method for images that separates task-relevant information from task-irrelevant information. Our primary hypothesis is that by revealing the minimal number of pixels required for a task we can provide the most privacy preserving guarantees. Specifically, we propose an adversarial method that masks out task-irrelevant information from an image for preserving privacy. The proposed method only uses task-specific label information and no privacy annotations such as identity of the subject, gender, race, etc., are required. We validate the proposed method on face attribute prediction on the CelebA dataset and emotion recognition on the FER+ dataset, showing that we can preserve visual privacy with little degradation in the task performance.
Ayush Jaiswal, Yue Wu 0001, Vivek Yadav, Pradeep Natarajan
FG3
2021 SauvolaNet: Learning Adaptive Sauvola Network for Degraded Document Binarization
Deng Li 0002, Yue Wu 0001, Yicong Zhou
ICDAR (4)2
2021 Linecounter: Learning Handwritten Text Line Segmentation By Counting
abstract
Handwritten Text Line Segmentation (HTLS) is a low-level but important task for many higher-level document processing tasks like handwritten text recognition. It is often formulated in terms of semantic segmentation or object detection in deep learning. However, both formulations have serious shortcomings. The former requires heavy post-processing of splitting/merging adjacent segments, while the latter may fail on dense or curved texts. In this paper, we propose a novel Line Counting formulation for HTLS – that involves counting the number of text lines from the top at every pixel location. This formulation helps learn an end-to-end HTLS solution that directly predicts per-pixel line number for a given document image. Furthermore, we propose a deep neural network (DNN) model LineCounter to perform HTLS through the Line Counting formulation. Our extensive experiments on the three public datasets (ICDAR2013-HSC [1], HIT-MW [2], and VML-AHTE [3]) demonstrate that LineCounter outperforms state-of-the-art HTLS approaches. Source code is available at https://github.com/Leedeng/LineCounter.
Deng Li 0002, Yue Wu 0001, Yicong Zhou
ICIP2
2021 Self-supervised Learning from a Multi-view Perspective
Yao-Hung Tsai, Yue Wu 0001, Ruslan Salakhutdinov, Louis-Philippe Morency
ICLR2
2021 Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning
abstract
Offline Reinforcement Learning promises to learn effective policies from previously-collected, static datasets without the need for exploration. However, existing Q-learning and actor-critic based off-policy RL algorithms fail when bootstrapping from out-of-distribution (OOD) actions or states. We hypothesize that a key missing ingredient from the existing methods is a proper treatment of uncertainty in the offline setting. We propose Uncertainty Weighted Actor-Critic (UWAC), an algorithm that detects OOD state-action pairs and down-weights their contribution in the training objectives accordingly. Implementation-wise, we adopt a practical and effective dropout-based uncertainty estimation method that introduces very little overhead over existing RL algorithms. Empirically, we observe that UWAC substantially improves model stability during training. In addition, UWAC out-performs existing offline RL methods on a variety of competitive tasks, and achieves significant performance gains over the state-of-the-art baseline on datasets with sparse demonstrations collected from human experts.
Yue Wu 0001, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind, Jian Zhang 0050, Ruslan Salakhutdinov, Hanlin Goh
ICML1
2021 Class-agnostic Object Detection
abstract
Object detection models perform well at localizing and classifying objects that they are shown during training. However, due to the difficulty and cost associated with creating and annotating detection datasets, trained models detect a limited number of object types with unknown objects treated as background content. This hinders the adoption of conventional detectors in real-world applications like large-scale object matching, visual grounding, visual relation prediction, obstacle detection (where it is more important to determine the presence and location of objects than to find specific types), etc. We propose class-agnostic object detection as a new problem that focuses on detecting objects irrespective of their object-classes. Specifically, the goal is to predict bounding boxes for all objects in an image but not their object-classes. The predicted boxes can then be consumed by another system to perform application-specific classification, retrieval, etc. We propose training and eval uation protocols for benchmarking class-agnostic detectors to advance future research in this domain. Finally, we propose (1) baseline methods and (2) a new adversarial learning framework for class-agnostic detection that forces the model to exclude class-specific information from features used for predictions. Experimental results show that adversarial learning improves class-agnostic detection efficacy.
Ayush Jaiswal, Yue Wu 0001, Pradeep Natarajan, Premkumar Natarajan
WACV2
2020 Online optimal consensus control of unknown linear multi-agent systems via time-based adaptive dynamic programming
Tieshan Li 0001, Qi-He Shan, Renhai Yu, Yue Wu 0001, C. L. Philip Chen
Neurocomputing5
2019 ManTra-Net: Manipulation Tracing Network for Detection and Localization of Image Forgeries With Anomalous Features
abstract
To fight against real-life image forgery, which commonly involves different types and combined manipulations, we propose a unified deep neural architecture called ManTraNet. Unlike many existing solutions, ManTra-Net is an end-to-end network that performs both detection and localization without extra preprocessing and postprocessing. ManTra-Net is a fully convolutional network and handles images of arbitrary sizes and many known forgery types such splicing, copy-move, removal, enhancement, and even unknown types. This paper has three salient contributions. We design a simple yet effective self-supervised learning task to learn robust image manipulation traces from classifying 385 image manipulation types. Further, we formulate the forgery localization problem as a local anomaly detection problem, design a Z-score feature to capture local anomaly, and propose a novel long short-term memory solution to assess local anomalies. Finally, we carefully conduct ablation experiments to systematically optimize the proposed network design. Our extensive experimental results demonstrate the generalizability, robustness and superiority of ManTra-Net, not only in single types of manipulations/forgeries, but also in their complicated combinations.
Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan
CVPR1
2019 QATM: Quality-Aware Template Matching for Deep Learning
abstract
Finding a template in a search image is one of the core problems in many computer vision applications, such as template matching, image semantic alignment, image-to-GPS verification \etc. In this paper, we propose a novel quality-aware template matching method, which is not only used as a standalone template matching algorithm, but also a trainable layer that can be easily plugged in any deep neural network. Specifically, we assess the quality of a matching pair as its soft-ranking among all matching pairs, and thus different matching scenarios like 1-to-1, 1-to-many, and many-to-many will be all reflected to different values. Our extensive studies in the classic template matching problem and deep learning tasks demonstrate the effectiveness of QATM: it not only outperforms SOTA template matching methods when used alone, but also largely improves existing DNN solutions when used in DNN.
Jiaxin Cheng, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan
CVPR2
2019 AIRD: Adversarial Learning Framework for Image Repurposing Detection
abstract
Image repurposing is a commonly used method for spreading misinformation on social media and online forums, which involves publishing untampered images with modified metadata to create rumors and further propaganda. While manual verification is possible, given vast amounts of verified knowledge available on the internet, the increasing prevalence and ease of this form of semantic manipulation call for the development of robust automatic ways of assessing the semantic integrity of multimedia data. In this paper, we present a novel method for image repurposing detection that is based on the real-world adversarial interplay between a bad actor who repurposes images with counterfeit metadata and a watchdog who verifies the semantic consistency between images and their accompanying metadata, where both players have access to a reference dataset of verified content, which they can use to achieve their goals. The proposed method exhibits state-of-the-art performance on location-identity, subject-identity and painting-artist verification, showing its efficacy across a diverse set of scenarios.
Ayush Jaiswal, Yue Wu 0001, Wael Abd-Almageed, Iacopo Masi, Premkumar Natarajan
CVPR2
2019 Layout-aware Subfigure Decomposition for Complex Figures in the Biomedical Literature
abstract
Published scientific figure is a valuable information resource, but often occur as composite images. The ImageCLEF meeting presented a shared evaluation in 2016 to use machine learning to split these composite figures into components automatically. We adapted an existing high-performance object detection method to analyze the substructure of published biomedical figures by developing a novel multi-branch output convolution neural network to predict irregular panel layouts and provide augmented training data to drive learning. Our system has an accuracy of 86.8% on the 2016 ImageCLEF Medical dataset and 83.1% on a new dataset derived from open access papers from the INTACT database of molecular interactions.
Xiangyang Shi, Yue Wu 0001, Huaigu Cao, Gully A. P. C. Burns, Premkumar Natarajan
ICASSP2
2019 A Study of Script Language Effects in Deep Neural-Network-Based Scene Text Detection
abstract
This study is different from most of the recent text detection work which focuses on creating a robust text detector system. In this work we studied how script languages affect a text detector's performance by using a multi-language synthetic dataset-namely, the Synthetic Octa-Language (SOL) dataset. The effect of script languages continues to be largely unexplored. Previously, this kind of experiment was infeasible because too many factors influence the performance of a text detector. We really cannot tell what role the factor X plays, neither positive nor negative. To overcome these difficulties, we used controlled synthesized data, which allows us to explicitly control factors such as base image, script language, text content, text color, font face, and font size. With the SOL dataset, we were able to investigate the effect that script languages have on on deep neural-network (DNN)-based methods under different scenarios. Moreover, this dataset can be used in other script-language-related text detection research as well.
Jiaxin Cheng, Achin Gupta, Yue Wu 0001, Premkumar Natarajan
ICDAR3
2019 Learning Pose-Aware Models for Pose-Invariant Face Recognition in the Wild
abstract
We propose a method designed to push the frontiers of unconstrained face recognition in the wild with an emphasis on extreme out-of-plane pose variations. Existing methods either expect a single model to learn pose invariance by training on massive amounts of data or else normalize images by aligning faces to a single frontal pose. Contrary to these, our method is designed to explicitly tackle pose variations. Our proposed Pose-Aware Models (PAM) process a face image using several pose-specific, deep convolutional neural networks (CNN). 3D rendering is used to synthesize multiple face poses from input images to both train these models and to provide additional robustness to pose variations at test time. Our paper presents an extensive analysis of the IARPA Janus Benchmark A (IJB-A), evaluating the effects that landmark detection accuracy, CNN layer selection, and pose model selection all have on the performance of the recognition pipeline. It further provides comparative evaluations on IJB-A and the PIPA dataset. These tests show that our approach outperforms existing methods, even surprisingly matching the accuracy of methods that were specifically fine-tuned to the target dataset. Parts of this work previously appeared in [1] and [2].
Iacopo Masi, Feng-Ju Chang, Jongmoo Choi, Shai Harel, Jungyeon Kim, KangGeon Kim, Jatuporn Toy Leksut, Stephen Rawls, Yue Wu 0001, Tal Hassner, Wael Abd-Almageed, Gérard G. Medioni, Louis-Philippe Morency, Premkumar Natarajan, Ramakant Nevatia
IEEE Trans. Pattern Anal. Mach. Intell.9
2018 Image-to-GPS Verification Through a Bottom-Up Pattern Matching Network
Jiaxin Cheng, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan
ACCV (5)2
2018 Bidirectional Conditional Generative Adversarial Networks
Ayush Jaiswal, Wael Abd-Almageed, Yue Wu 0001, Premkumar Natarajan
ACCV (3)3
2018 Weighted Feature Pooling Network in Template-Based Recognition
Zekun Li 0007, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan
ACCV (5)2
2018 BusterNet: Detecting Copy-Move Image Forgery with Source/Target Localization
Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan
ECCV (6)1
2018 Deep Multimodal Image-Repurposing Detection
abstract
Nefarious actors on social media and other platforms often spread rumors and falsehoods through images whose metadata (e.g., captions) have been modified to provide visual substantiation of the rumor/falsehood. This type of modification is referred to as image repurposing, in which often an unmanipulated image is published along with incorrect or manipulated metadata to serve the actor's ulterior motives. We present the Multimodal Entity Image Repurposing (MEIR) dataset, a substantially challenging dataset over that which has been previously available to support research into image repurposing detection. The new dataset includes location, person, and organization manipulations on real-world data sourced from Flickr. We also present a novel, end-to-end, deep multimodal learning model for assessing the integrity of an image by combining information extracted from the image with related information from a knowledge base. The proposed method is compared against state-of-the-art techniques on existing datasets as well as MEIR, where it outperforms existing methods across the board, with AUC improvement up to 0.23.
Ekraam Sabir, Wael Abd-Almageed, Yue Wu 0001, Premkumar Natarajan
ACM Multimedia3
2018 Unsupervised Adversarial Invariance
abstract
Data representations that contain all the information about target variables but are invariant to nuisance factors benefit supervised learning algorithms by preventing them from learning associations between these factors and the targets, thus reducing overfitting. We present a novel unsupervised invariance induction framework for neural networks that learns a split representation of data through competitive training between the prediction task and a reconstruction task coupled with disentanglement, without needing any labeled information about nuisance factors or domain knowledge. We describe an adversarial instantiation of this framework and provide analysis of its working. Our unsupervised model outperforms state-of-the-art methods, which are supervised, at inducing invariance to inherent nuisance factors, effectively using synthetic data augmentation to learn invariance, and domain adaptation. Our method can be applied to any prediction task, eg., binary/multi-class classification or regression, without loss of generality.
Ayush Jaiswal, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan
NeurIPS2
2018 Image Copy-Move Forgery Detection via an End-to-End Deep Neural Network
abstract
In this paper, for the first time, we introduce a new end-to-end deep neural network predicting forgery masks to the image copy-move forgery detection problem. Specifically, we use a convolutional neural network to extract block-like features from an image, compute self-correlations between different blocks, use a pointwise feature extractor to locate matching points, and reconstruct a forgery mask through a deconvolutional network. Unlike classic solutions requiring multiple stages of training and parameter tuning, ranging from feature extraction to postprocessing, the proposed solution is fully trainable and can be jointly optimized for the forgery mask reconstruction loss. Our experimental results demonstrate that the proposed method achieves better forgery detection performance than classic approaches relying on different features and matching schemes, and it is more robust against various known attacks like affine transformation, JPEG compression, blurring, etc.
Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan
WACV1
2018 Facial Landmark Detection with Tweaked Convolutional Neural Networks
abstract
This paper concerns the problem of facial landmark detection. We provide a unique new analysis of the features produced at intermediate layers of a convolutional neural network (CNN) trained to regress facial landmark coordinates. This analysis shows that while being processed by the CNN, face images can be partitioned in an unsupervised manner into subsets containing faces in similar poses (i.e., 3D views) and facial properties (e.g., presence or absence of eye-wear). Based on this finding, we describe a novel CNN architecture, specialized to regress the facial landmark coordinates of faces in specific poses and appearances. To address the shortage of training data, particularly in extreme profile poses, we additionally present data augmentation techniques designed to provide sufficient training examples for each of these specialized sub-networks. The proposed Tweaked CNN (TCNN) architecture is shown to outperform existing landmark detection methods in an extensive battery of tests on the AFW, ALFW, and 300W benchmarks. Finally, to promote reproducibility of our results, we make code and trained models publicly available through our project webpage.
Yue Wu 0001, Tal Hassner, KangGeon Kim, Gérard G. Medioni, Premkumar Natarajan
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Designing Hyperchaotic Cat Maps With Any Desired Number of Positive Lyapunov Exponents
abstract
Generating chaotic maps with expected dynamics of users is a challenging topic. Utilizing the inherent relation between the Lyapunov exponents (LEs) of the Cat map and its associated Cat matrix, this paper proposes a simple but efficient method to construct an -dimensional ( -D) hyperchaotic Cat map (HCM) with any desired number of positive LEs. The method first generates two basic -D Cat matrices iteratively and then constructs the final -D Cat matrix by performing similarity transformation on one basic -D Cat matrix by the other. Given any number of positive LEs, it can generate an -D HCM with desired hyperchaotic complexity. Two illustrative examples of -D HCMs were constructed to show the effectiveness of the proposed method, and to verify the inherent relation between the LEs and Cat matrix. Theoretical analysis proves that the parameter space of the generated HCM is very large. Performance evaluations show that, compared with existing methods, the proposed method can construct -D HCMs with lower computation complexity and their outputs demonstrate strong randomness and complex ergodicity.
Zhongyun Hua, Yicong Zhou, Chengqing Li, Yue Wu 0001
IEEE Trans. Cybern.5
2017 EPAT: Euclidean Perturbation Analysis and Transform - An Agnostic Data Adaptation Framework for Improving Facial Landmark Detectors
abstract
We propose EPAT, (Euclidean Perturbation Analysis and Transform) a novel unsupervised adaptation approach for improving the accuracy of any facial landmark detector by characterizing the stability of landmark prediction on test images. In EPAT, a test image is transformed several times using a set of Euclidean transforms, producing several perturbed images. The black box landmark detector is used to find facial landmarks on each perturbed version of the test image. Subsequently, inverse transforms are applied to the corresponding landmarks in order to map them back to the original image. Mean and variance are calculated for all inversely transformed detection. Mean and variance represent the new ensemble prediction and the sensitivity of the underlying landmark detector, respectively. We also introduce affine variance (AV) of facial landmarks. AV is used as a measure of the stability of the predicted landmarks and a criterion for selecting a good data adaptation model which effectively addresses potential mismatches between test and training data of the underlying landmark detector. EPAT is evaluated using four state-of-the-art landmark detectors on the standard 300W dataset and also incorporated into a face recognition pipeline to show improved recognition accuracy on the challenging IJB-A dataset.
Yue Wu 0001, Wael Abd-Almageed, Stephen Rawls, Premkumar Natarajan
FG1
2017 Self-Organized Text Detection with Minimal Post-processing via Border Learning
abstract
In this paper we propose a new solution to the text detection problem via border learning. Specifically, we make four major contributions: 1) We analyze the insufficiencies of the classic non-text and text settings for text detection. 2) We introduce the border class to the text detection problem for the first time, and validate that the decoding process is largely simplified with the help of text border. 3) We collect and release a new text detection PPT dataset containing 10,692 images with non-text, border, and text annotations. 4) We develop a lightweight (only 0.28M parameters), fully convolutional network (FCN) to effectively learn borders in text images. The results of our extensive experiments show that the proposed solution achieves comparable performance, and often outperforms state-of-theart approaches on standard benchmarks-even though our solution only requires minimal post-processing to parse a bounding box from a detected text map, while others often require heavy post-processing.
Yue Wu 0001, Premkumar Natarajan
ICCV1
2017 Deep Matching and Validation Network: An End-to-End Solution to Constrained Image Splicing Localization and Detection
abstract
Image splicing is a very common image manipulation technique that is sometimes used for malicious purposes. A splicing detection and localization algorithm usually takes an input image and produces a binary decision indicating whether the input image has been manipulated, and also a segmentation mask that corresponds to the spliced region. Most existing splicing detection and localization pipelines suffer from two main shortcomings: 1) they use handcrafted features that are not robust against subsequent processing (e.g., compression), and 2) each stage of the pipeline is usually optimized independently. In this paper we extend the formulation of the underlying splicing problem to consider two input images, a query image and a potential donor image. Here the task is to estimate the probability that the donor image has been used to splice the query image, and obtain the splicing masks for both the query and donor images. We introduce a novel deep convolutional neural network architecture, called Deep Matching and Validation Network (DMVN), which simultaneously localizes and detects image splicing. The proposed approach does not depend on handcrafted features and uses raw input images to create deep learned representations. Furthermore, the DMVN is end-to-end optimized to produce the probability estimates and the segmentation masks. Our extensive experiments demonstrate that this approach outperforms state-of-the-art splicing detection methods by a large margin in terms of both AUC score and speed.
Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan
ACM Multimedia1
2016 Learning document image binarization from data
abstract
We present a fully trainable solution for binarization of degraded document images using extremely randomized trees. Unlike previous attempts that often use simple features, our method encodes all heuristics about whether or not a pixel is foreground text into a high-dimensional feature vector and learns a more complicated decision function. We introduce two novel features, the Logarithm Intensity Percentile (LIP) and the Relative Darkness Index (RDI), and combine them with low level features, and reformulated features from existing binarization methods. Experimental results show that using small sample size (about 1.5% of all available training data), we can achieve a binarization performance comparable to manually-tuned, state-of-the-art methods. Additionally, the trained document binarization classifier shows good generalization capabilities on out-of-domain data.
Yue Wu 0001, Premkumar Natarajan, Stephen Rawls, Wael Abd-Almageed
ICIP1
2016 Computationally efficient template-based face recognition
abstract
Classically, face recognition depends on computing the similarity (or distance) between a pair of face images and/or their respective representations, where each subject is represented by one image. Template-based face recognition was introduced by the release of IARPA's Janus Benchmark-A (IJB-A) dataset, in which each enrolled subject is represented by a group of one or more images, called a template. The group of images comprising a template might have been acquired using different head poses, illuminations, ages and facial expressions. Template images could come from still images or video frames. Therefore, measuring the similarity between templates representing two subjects significantly increases the number of pairwise image comparisons (i.e., O(NM), where N and M are the number of image templates being compared). As the number of enrolled subjects, K, increases, both computational and space requirements become computationally prohibitive. To address this challenge, we present a novel approximate nearest-neighbor (ANN) search-based solution. Given a query template, ANN methods are used to find similar face images. Retrieved images are used to construct a template pool that is used to find the correct identity of the query subject. The proposed approach largely reduces the number of imposter template-pair comparisons. Experimental results on the IJB-A dataset show that the proposed approach achieves significant speed-up and storage savings, without sacrificing accuracy.
Yue Wu 0001, Wael Abd-Almageed, Stephen Rawls, Premkumar Natarajan
ICPR1
2016 Face recognition using deep multi-pose representations
abstract
We introduce our method and system for face recognition using multiple pose-aware deep learning models. In our representation, a face image is processed by several pose-specific deep convolutional neural network (CNN) models to generate multiple pose-specific features. 3D rendering is used to generate multiple face poses from the input image. Sensitivity of the recognition system to pose variations is reduced since we use an ensemble of pose-specific CNN features. The paper presents extensive experimental results on the effect of landmark detection, CNN layer selection and pose model selection on the performance of the recognition pipeline. Our novel representation achieves better results than the state-of-the-art on IARPA's CS2 and NIST's IJB-A in both verification and identification (i.e. search) tasks.
Wael Abd-Almageed, Yue Wu 0001, Stephen Rawls, Shai Harel, Tal Hassner, Iacopo Masi, Jongmoo Choi, Jatuporn Toy Leksut, Jungyeon Kim, Premkumar Natarajan, Ramakant Nevatia, Gérard G. Medioni
WACV2
2016 2D Sudoku associated bijections for image scrambling
Yue Wu 0001, Yicong Zhou, Sos S. Agaian, Joseph P. Noonan
Inf. Sci.1
2016 n-Dimensional Discrete Cat Map Generation Using Laplace Expansions
abstract
Different from existing methods that use matrix multiplications and have high computation complexity, this paper proposes an efficient generation method of${n}$-dimensional (${n}\text{D}$) Cat maps using Laplace expansions. New parameters are also introduced to control the spatial configurations of the${n}\text{D}$Cat matrix. Thus, the proposed method provides an efficient way to mix dynamics of all dimensions at one time. To investigate its implementations and applications, we further introduce a fast implementation algorithm of the proposed method with time complexity${O(n^{4})}$and a pseudorandom number generator using the Cat map generated by the proposed method. The experimental results show that, compared with existing generation methods, the proposed method has a larger parameter space and simpler algorithm complexity, generates${n}\text{D}$Cat matrices with a lower inner correlation, and thus yields more random and unpredictable outputs of${n}\text{D}$Cat maps.
Yue Wu 0001, Zhongyun Hua, Yicong Zhou
IEEE Trans. Cybern.1
2014 Confusion Network Based Recurrent Neural Network Language Modeling for Chinese OCR Error Detection
abstract
This paper presents a new framework for OCR error detection, which uses a conditional random field model to combine rich features from multiple sources, including confusion networks (c-nets), lexical local context and recurrent neural network language model (RNNLM)1. We propose a novel, efficient method for computing character-level c-net based RNNLM scores by using dynamic programming and c-net partial unfolding. Our experiments show that our error detection model has consistent observable improvements over a high baseline employed by our current OCR demo system, as measured by average precision and detection error trade-off curve on two test sets of Chinese image documents. Both linguistic and recognition features contribute to the high performance, with the former especially informative. In addition, we show that the new feature we proposed, the c-net RNNLM feature, plays a remarkable beneficial role in improving error detection rate. These results suggest that applications on top of image text recognition can benefit substantially from a hybrid strategy that combines techniques from optical character recognition and natural language processing.
Jinying Chen, Yue Wu 0001, Huaigu Cao, Premkumar Natarajan
ICPR2
2014 Design of image cipher using latin squares
Yue Wu 0001, Yicong Zhou, Joseph P. Noonan, Sos S. Agaian
Inf. Sci.1
2014 A constrained optimization approach to combining multiple non-local means denoising estimates
Brian Tracey, Eric L. Miller 0001, Yue Wu 0001, Pradeep Natarajan, Joseph P. Noonan
Signal Process.3
2014 Fast blockwise SURE shrinkage for image denoising
Yue Wu 0001, Brian Tracey, Premkumar Natarajan, Joseph P. Noonan
Signal Process.1
2014 A symmetric image cipher using wave perturbations
Yue Wu 0001, Yicong Zhou, Sos S. Agaian, Joseph P. Noonan
Signal Process.1
2013 Local Shannon entropy measure with statistical tests for image randomness
Yue Wu 0001, Yicong Zhou, George Saveriades, Sos S. Agaian, Joseph P. Noonan, Premkumar Natarajan
Inf. Sci.1
2013 James-Stein Type Center Pixel Weights for Non-Local Means Image Denoising
abstract
Non-Local Means (NLM) and its variants have proven to be effective and robust in many image denoising tasks. In this letter, we study approaches to selecting center pixel weights (CPW) in NLM. Our key contributions are 1) we give a novel formulation of the CPW problem from a statistical shrinkage perspective; 2) we construct the James–Stein shrinkage estimator in the CPW context; and 3) we propose a new local James–Stein type CPW (LJSCPW) that is locally tuned for each image pixel. Our experimental results showed that compared to existing CPW solutions, the LJSCPW is more robust and effective under various noise levels. In particular, the NLM with the LJSCPW attains higher means with smaller variances in terms of the peak signal and noise ratio (PSNR) and structural similarity (SSIM), implying it improves the NLM denoising performance and makes the denoising less sensitive to parameter changes.
Yue Wu 0001, Brian Tracey, Premkumar Natarajan, Joseph P. Noonan
IEEE Signal Process. Lett.1
2013 Probabilistic Non-Local Means
abstract
In this letter, we propose a so-called probabilistic non-local means (PNLM) method for image denoising. Our main contributions are: 1) we point out defects of the weight function used in the classic NLM; 2) we successfully derive all theoretical statistics of patch-wise differences for Gaussian noise; and 3) we employ this prior information and formulate the probabilistic weights truly reflecting the similarity between two noisy patches. Our simulation results indicate the PNLM outperforms the classic NLM and many NLM recent variants in terms of the peak signal noise ratio (PSNR) and the structural similarity (SSIM) index. Encouraging improvements are also found when we replace the NLM weights with the PNLM weights in tested NLM variants.
Yue Wu 0001, Brian Tracey, Premkumar Natarajan, Joseph P. Noonan
IEEE Signal Process. Lett.1
2011 Wavelet Band-pass Filters for Matching Multiple Templates in Real-time
abstract
Many applications in image processing and computer vision require finding a particular template in an image or a video, that is, template matching. Given a template and an input, the matching algorithm finds the region of interest (ROI) that most closely matches the template in terms of some similarity measurement. According to the way similarity measurements are performed, the template matching methods can roughly be classified into two groups: 1) patch matching schemes, such as the sum of absolute difference (SAD) , the sum of squared difference (SSD) [1], or cross correlation (XCORR), where the similarity measurement directly relies on pixel information from the patch of interest; and 2) feature matching schemes, such as invariant features [4] and bags of features [5], where similarity measurement relies on features describing the template and the frame. Patch matching methods are not robust, especially when noise, skew, or errors occur. Further, they consume a large amount of time, because of expensive sliding window search for calculating the similarity score over all possible locations. Several techniques have been explored for accelerating such matching methods, including early rejections and correlation techniques [1]. However, the computation cost could still be unaffordable when the frame size is large. Typically other techniques, like frame difference, are used to reduce the search space in applications. Feature matching methods process the template and describe it with features, which are ideally invariant to rotation, skew, noise etc. However in many cases, the use of a more complicated model for similarity measurement results in higher computational cost. Further, sliding window search is also a costly stage for such methods. While there exist known algorithms for fast search of object instances in an image using branch-and-bound techniques, in our particular problem, methods of this type have two crucial limitations. First, they require a large number of training samples for each class to learn robust classifiers. Second, interest point detectors like SIFT [5] typically do not generate sufficient number of feature points, because of the small size of the provided logo, large homogenous regions and degradations. Wavelets based approaches have been extensively used in object detection and recognition. In [3], wavelet coefficients based image histogram are collected in bins and are used for classifying logos. In [6], wavelet coefficients are directly used and trained for pedestrian detection. In [8], wavelet coefficients are selected to form rotation-invariant features by using the angular-radial transform. However, matching logos within frames using [3, 6, 8] still requires expensive window searching and thus are not appropriate for real-time processing. In this paper, we propose a new matching method using the wavelet based band-pass filters (WBPFs). Instead of using direct distance measurement requiring expensive window search, the similarity is measured in the indirect way involving two stages. In the stage of offline template processing (see Figure 1), a template is automatically described by a set of three directional WBPFs, where only salient wavelet frequency components of the template are allowed to pass. In the stage of online frame processing (see Figure 2), a frame is transformed to the wavelet domain and its sub-bands are filtered with respect to the corresponding template WBFPs. Finally, the detection is made with respect to the region of the densest responses under spatial constraints [2, 4]. We show that the proposed template matching system has a very low computational cost, which is 50 times faster than the correlation based SSD [1] and 10 times faster than the orthogonal Haar transform (OHT) based SSD [7]. Further, the proposed method does not trade-off accuracy, since the use of subtemplate information makes it robust to skew and camera view change. Experimental results demonstrate our method for real-time logo detection in broadcast videos.
Yue Wu 0001, Pradeep Natarajan, Joseph P. Noonan, Rohit Prasad, Premkumar Natarajan
BMVC1
2011 Large-scale, real-time logo recognition in broadcast videos
abstract
Robust, real-time, multi-class logo detection in high resolution broadcast videos presents several difficult challenges. For most logos we only have a few training samples, which makes training robust classifiers hard. Also, logos could potentially occur anywhere in the image, and traditional sliding window approaches for logo/object detection are computationally intensive. We present a system that addresses these issues by first identifying a small set of possible logo locations in a frame, based on temporal continuity and multi-resolution search, and then successively pruning these locations for each logo template, using a cascade of color and edge based features. We present experimental results that demonstrate our system for detecting a total of 270 different logo classes in broadcast video from 5 different languages (English, Indonesian, Malay, Simplified and Traditional Chinese).
Pradeep Natarajan, Yue Wu 0001, Shirin Saleem, Ehry MacRostie, Fred Bernardin, Rohit Prasad, Premkumar Natarajan
ICME2
2011 A novel information entropy based randomness test for image encryption
abstract
In this paper, a new information entropy based randomness test for image encryption about is proposed. Unlike the conventional information entropy test working on the global image, the new proposed measurement focuses on local image blocks and calculates the sample mean of information entropy over a number of random selected image blocks within a test encrypted image. Reference values are mathematical derived from the ideally encrypted image model. As a result, the randomness of an encrypted image can be easily told by comparing the test values with theoretical ones. Ultimately, a hypothesis test to accept or reject an encrypted image is also derived to accept or reject the null hypothesis that ciphertext image is ideally encrypted/random-like with a significance level α. In such a way, the proposed block entropy test provides both quantitative and qualitative results for image encryption. Experimental results on existing image cipher show the effectiveness of the proposed test. The same idea can be also to other digital data, like videos or audios.
Yue Wu 0001, Joseph P. Noonan, Sos S. Agaian
SMC1
2011 Dynamic and implicit latin square doubly stochastic S-boxes with reversibility
abstract
S-Boxes play a vital role in cipher designs and have been researched for years. In this article a new way of dynamically designing S-boxes using Latin Square doubly stochastic matrix is proposed. And it is demonstrated that the enciphering/deciphering process is a mimic of a Markov chain Monte Carlo simulation. Unlike conventional dynamic S-boxes, the proposed S-boxes are not directly defined by keys, but contained in the key dependent doubly stochastic matrix. The created matrix has desired properties including: 1) it implicitly contains S-boxes and thus complicates the internal structure of S-boxes; 2) it naturally defines S-boxes with reversibility and thus any S-boxes used for encryption can be directly used for decryption; 3) it satisfies the Strict Avalanche Criterion (SAC) for S-Boxes and thus it has good resistance to differential or linear cryptanalysis; 4) it guarantees the independence of the ciphertext distribution from plaintext one; and 5) it ensures that the expected ciphertext distribution is uniform and thus attains excellent confusion properties when use these S-Boxes iteratively. Theoretical and experimental results show that the proposed S-box has a high security level and is suitable to design cryptosystem for data encryption. We also extend our S-boxes to a simple image cipher. Experimental results show that the proposed dynamical flexible structure S-boxes has cryptographic properties comparable or better than some existing image encryption methods.
Yue Wu 0001, Joseph P. Noonan, Sos S. Agaian
SMC1
2010 Binary data encryption using the Sudoku block cipher
abstract
This paper presents a novel block cipher based on the Sudoku matrix. The offered block cipher combines many advantages of chaos-based encryption and traditional transform-based encryption techniques. Computer simulations show that a) the encrypted data have very random-like properties under many statistical metrics, b) unlike most chaos-based encryption methods generating unpredictable output, our new method is robust and effective for generating uniform-like encrypted data; c) it has high sensitivity to the KEY. The offered scheme can be applied to many different data types, such as audio, image and video.
Yue Wu 0001, Joseph P. Noonan, Sos S. Agaian
SMC1