VLDB 2026 Research / reviewers in the wild / expert
Vishnu Naresh Boddeti
dblp:55/6988 · also Vishnu Boddeti
· DBLP profile ↗
61ranked-venue papers
5as first author
36since 2021 · last 2026
0000-0002-8918-9385ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 4 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 3 first-author · 17 since 2021Security and privacy · 5 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Show Me: Unifying Instructional Image and Video Generation with Diffusion ModelsabstractGenerating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image manipulation or video prediction, these tasks are typically treated in isolation. This separation reveals a fundamental issue: image manipulation methods overlook how actions unfold over time, while video prediction models often ignore the intended outcomes. To this end, we propose ShowMe, a unified framework that enables both tasks by selectively activating the spatial and temporal components of video diffusion models. In addition, we introduce structure and motion consistency rewards to improve structural fidelity and temporal coherence. Notably, this unification brings dual benefits: the spatial knowledge gained through video pretraining enhances contextual consistency and realism in non-rigid image edits, while the instruction-guided manipulation stage equips the model with stronger goal-oriented reasoning for video prediction. Experiments on diverse benchmarks demonstrate that our method outperforms expert models in both instructional image and video generation, highlighting the strength of video diffusion models as a unified action-object state transformer. Our code will be available at https://yujiangpu20.github.io/showme/. Yujiang Pu, Zhanbo Huang, Vishnu Naresh Boddeti, Yu Kong 0001 |
WACV | 3 |
| 2025 | CryptoFace: End-to-End Encrypted Face RecognitionabstractFace recognition is central to many authentication, security, and personalized applications. Yet, it suffers from significant privacy risks, particularly arising from unauthorized access to sensitive biometric data. This paper introduces CryptoFace, the first end-to-end encrypted face recognition system with fully homomorphic encryption (FHE). It enables secure processing of facial data across all stages of a face-recognition process—feature extraction, storage, and matching—without exposing raw images or features. We introduce a mixture of shallow patch convolutional networks to support higher-dimensional tensors via patch-based processing while reducing the multiplicative depth and thus inference latency. Parallel FHE evaluation of these networks ensures near-resolution-independent latency. On standard face recognition benchmarks, CryptoFace significantly accelerates inference and increases verification accuracy compared to the state-of-the-art FHE neural networks adapted for face recognition. CryptoFace will facilitate secure face recognition systems requiring robust and provable security. The code is available at https://github.com/human-analysis/CryptoFace. Vishnu Naresh Boddeti |
CVPR | 2 |
| 2025 | DiverseFlow: Sample-Efficient Diverse Mode Coverage in FlowsabstractMany real-world applications of flow-based generative models desire a diverse set of samples that cover multiple modes of the target distribution. However, the predominant approach for obtaining diverse sets is not sample-efficient, as it involves independently obtaining many samples from the source distribution and mapping them through the flow until the desired mode coverage is achieved. As an alternative to repeated sampling, we introduce DiverseFlow: a training-free approach to improve the diversity of flow models. Our key idea is to employ a determinantal point process to induce a coupling between the samples that drives diversity under a fixed sampling budget. In essence, DiverseFlow allows exploration of more variations in a learned flow model with fewer samples. We demonstrate the efficacy of our method for tasks where sample-efficient diversity is desirable, such as text-guided image generation with polysemous words, inverse problems like large-hole inpainting, and class-conditional image synthesis. Mashrur Mahmud Morshed, Vishnu Naresh Boddeti |
CVPR | 2 |
| 2025 | SEAL: Semantic Attention Learning for Long Video RepresentationabstractLong video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must efficiently process such redundancy while preserving essential contents for downstream tasks. This paper introduces SEmantic Attention Learning (SEAL), a novel unified representation for long videos. To reduce computational complexity, long videos are decomposed into three distinct types of semantic entities: scenes, objects, and actions, allowing models to operate on a compact set of entities rather than a large number of frames or pixels. To further address redundancy, we propose an attention learning module that balances token relevance with diversity, formulated as a subset selection optimization problem. Our representation is versatile and applicable across various long video understanding tasks. Extensive experiments demonstrate that SEAL significantly outperforms state-of-the-art methods in video question answering and temporal grounding tasks across diverse benchmarks, including LVBench, MovieChat-1K, and Ego4D. Yujia Chen 0001, Du Tran, Vishnu Naresh Boddeti, Wen-Sheng Chu |
CVPR | 4 |
| 2025 | Generative Zero-Shot Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a vision-language task utilizing queries comprising images and textual descriptions to achieve precise image retrieval. This task seeks to find images that are visually similar to a reference image while incorporating specific changes or features described textually (visual delta). CIR enables a more flexible and user-specific retrieval by bridging visual data with verbal instructions. This paper introduces a novel generative method that augments Composed Image Retrieval by Composed Image Generation (CIG) to provide pseudo-target images. CIG utilizes a textual inversion network to map reference images into semantic word space, which generates pseudo-target images in combination with textual descriptions. These images serve as additional visual information, significantly improving the accuracy and relevance of retrieved images when integrated into existing retrieval frameworks. Experiments conducted across multiple CIR datasets and several baseline methods demonstrate improvements in retrieval performance, which shows the potential of our approach as an effective add-on for existing composed image retrieval. Project Page: https://lan-lw.github.io/CIG/ Vishnu Naresh Boddeti, Ser-Nam Lim |
CVPR | 3 |
| 2025 | Shielding Latent Face Representations From Privacy AttacksabstractIn today’s data-driven analytics landscape, deep learning has become a powerful tool, with latent representations, known as embeddings, playing a central role in several applications. In the face analytics domain, such embeddings are commonly used for biometric recognition (e.g., face identification). However, these embeddings, or templates, can inadvertently expose sensitive attributes such as age, gender, and ethnicity. Leaking such information can compromise personal privacy and affect civil liberty and human rights. To address these concerns, we introduce a multi-layer protection framework for embeddings. It consists of a sequence of operations: (a) encrypting embeddings using Fully Homomorphic Encryption (FHE), and (b) hashing them using irreversible feature manifold hashing. Unlike conventional encryption methods, FHE enables computations directly on encrypted data, allowing downstream analytics while maintaining strong privacy guarantees. To reduce the overhead of encrypted processing, we employ embedding compression. Our proposed method shields latent representations of sensitive data from leaking private attributes (such as age and gender) while retaining essential functional capabilities (such as face identification). Extensive experiments on two datasets using two face encoders demonstrate that our approach outperforms several state-of-the-art privacy protection methods. Arjun Ramesh Kaushik, Bharat Yalavarthi, Arun Ross, Vishnu Naresh Boddeti, Nalini K. Ratha |
FG | 4 |
| 2025 | OASIS Uncovers: High-Quality T2I Models, Same Old StereotypesabstractImages generated by text-to-image (T2I) models often exhibit visual biases and stereotypes of concepts such as culture and profession. Existing quantitative measures of stereotypes are based on statistical parity that does not align with the sociological definition of stereotypes and, therefore, incorrectly categorizes biases as stereotypes. Instead of oversimplifying stereotypes as biases, we propose a quantitative measure of stereotypes that aligns with its sociological definition. We then propose OASIS to measure the stereotypes in a generated dataset and understand their origins within the T2I model. OASIS includes two scores to measure stereotypes from a generated image dataset: **(M1)** Stereotype Score to measure the distributional violation of stereotypical attributes, and **(M2)** WALS to measure spectral variance in the images along a stereotypical attribute. OASIS
also includes two methods to understand the origins of stereotypes in T2I models: **(U1)** StOP to discover attributes that the T2I model internally associates with a given concept, and **(U2)** SPI to quantify the emergence of stereotypical attributes in the latent space of the T2I model during image generation. Despite the considerable progress in image fidelity, using OASIS, we conclude that newer T2I models such as FLUX.1 and SDv3 contain strong stereotypical predispositions about concepts and still generate images with widespread stereotypical attributes. Additionally, the quantity of stereotypes worsens for nationalities with lower Internet footprints. Sepehr Dehdashtian, Gautam Sreekumar, Vishnu Naresh Boddeti |
ICLR | 3 |
| 2025 | CoInD: Enabling Logical Compositions in Diffusion ModelsabstractHow can we learn generative models to sample data with arbitrary logical compositions of statistically independent attributes? The prevailing solution is to sample from distributions expressed as a composition of attributes' conditional marginal distributions under the assumption that they are statistically independent. This paper shows that standard conditional diffusion models violate this assumption, even when all attribute compositions are observed during training. And, this violation is significantly more severe when only a subset of the compositions is observed. We propose CoInD to address this problem. It explicitly enforces statistical independence between the conditional marginal distributions by minimizing Fisher’s divergence between the joint and marginal distributions. The theoretical advantages of CoInD are reflected in both qualitative and quantitative experiments, demonstrating a significantly more faithful and controlled generation of samples for arbitrary logical compositions of attributes. The benefit is more pronounced for scenarios that current solutions relying on the assumption of conditionally independent marginals struggle with, namely, logical compositions involving the NOT operation and when only a subset of compositions are observed during training. Sachit Gaudi, Gautam Sreekumar, Vishnu Naresh Boddeti |
ICLR | 3 |
| 2025 | Obliviator Reveals the Cost of Nonlinear Guardedness in Concept ErasureabstractConcept erasure aims to remove unwanted attributes, such as social or demographic factors, from learned representations, while preserving their task-relevant utility. While the goal of concept erasure is protection against all adversaries, existing methods remain vulnerable to nonlinear ones. This vulnerability arises from their failure to fully capture the complex, nonlinear statistical dependencies between learned representations and unwanted attributes. Moreover, although the existence of a trade-off between utility and erasure is expected, its progression during the erasure process, i.e., the cost of erasure, remains unstudied. In this work, we introduce Obliviator, a post-hoc erasure method designed to fully capture nonlinear statistical dependencies. We formulate erasure from a functional perspective, leading to an optimization problem involving a composition of kernels that lacks a closed-form solution. Instead of solving this problem in a single shot, we adopt an iterative approach that gradually morphs the feature space to achieve a more utility-preserving erasure. Unlike prior methods, Obliviator guards unwanted attribute against nonlinear adversaries. Our gradual approach quantifies the cost of nonlinear guardedness and reveals the dynamics between attribute protection and utility-preservation over the course of erasure. The utility-erasure trade-off curves obtained by Obliviator outperform the baselines and demonstrate its strong generalizability: its erasure becomes more utility-preserving when applied to the better-disentangled representations learned by more capable models. Ramin Akbari, Milad Afshari, Vishnu Naresh Boddeti |
NeurIPS | 3 |
| 2025 | PolyJuice Makes It Real: Black-Box, Universal Red Teaming for Synthetic Image DetectorsabstractSynthetic image detectors (SIDs) are a key defense against the risks posed by the growing realism of images from text-to-image (T2I) models. Red teaming improves SID’s effectiveness by identifying and exploiting their failure modes via misclassified synthetic images. However, existing red-teaming solutions (i) require white-box access to SIDs, which is infeasible for proprietary state-of-the-art detectors, and (ii) generate image-specific attacks through expensive online optimization. To address these limitations, we propose PolyJuice, the first black-box, image-agnostic red-teaming method for SIDs, based on an observed distribution shift in the T2I latent space between samples correctly and incorrectly classified by the SID. PolyJuice generates attacks by (i) identifying the direction of this shift through a lightweight offline process that only requires black-box access to the SID, and (ii) exploiting this direction by universally steering all generated images towards the SID’s failure modes. PolyJuice-steered T2I models are significantly more effective at deceiving SIDs (up to 84%) compared to their unsteered counterparts. We also show that the steering directions can be estimated efficiently at lower resolutions and transferred to higher resolutions using simple interpolation, reducing computational overhead. Finally, tuning SID models on PolyJuice-augmented datasets notably enhances the performance of the detectors (up to 30%). Sepehr Dehdashtian, Mashrur Mahmud Morshed, Jacob H. Seidman, Gaurav Bharaj, Vishnu Naresh Boddeti |
NeurIPS | 5 |
| 2024 | Utility-Fairness Trade-Offs and how to Find ThemabstractWhen building classification systems with demographic fairness considerations, there are two objectives to satisfy: 1) maximizing utility for the specific task and 2) ensuring fairness w.r.t. a known demographic attribute. These objectives often compete, so optimizing both can lead to a trade-off between utility and fairness. While existing works acknowledge the trade-offs and study their limits, two questions remain unanswered: 1) What are the optimal trade-offs between utility and fairness? and 2) How can we nu-merically quantify these trade-offs from data for a desired prediction task and demographic attribute of interest? This paper addresses these questions. We introduce two utility-fairness trade-offs: the Data-Space and Label-Space Trade-off. The trade-offs reveal three regions within the utility-fairness plane, delineating what is fully and partially possible and impossible. We propose U-FaTE, a method to nu-merically quantify the trade-offs for a given prediction task and group fairness definition from data samples. Based on the trade-offs, we introduce a new scheme for evaluating representations. An extensive evaluation of fair representation learning methods and representations from over 1000 pre-trained models revealed that most current approaches are far from the estimated and achievable fairness-utility trade-offs across multiple datasets and prediction tasks. Sepehr Dehdashtian, Bashir Sadeghi, Vishnu Naresh Boddeti |
CVPR | 3 |
| 2024 | Enhancing Privacy in Face Analytics Using Fully Homomorphic EncryptionabstractModern face recognition systems utilize deep neural networks to extract salient features from a face. These features denote embeddings in latent space and are often stored as templates in a face recognition system. These embeddings are susceptible to data leakage and, in some cases, can even be used to reconstruct the original face image. To prevent compromising identities, template protection schemes are commonly employed. However, these schemes may still not prevent the leakage of soft biometric information such as age, gender and race. To alleviate this issue, we propose a novel technique that combines Fully Homomorphic Encryption (FHE) with an existing template protection scheme known as PolyProtect. We show that the embeddings can be compressed and encrypted using FHE and transformed into a secure PolyProtect template using polynomial transformation, for additional protection. We demonstrate the efficacy of the proposed approach through extensive experiments on multiple datasets. Our proposed approach ensures irreversibility and unlinkability, effectively preventing the leakage of soft biometric attributes from face embeddings without compromising recognition accuracy. Bharat Yalavarthi, Arjun Ramesh Kaushik, Arun Ross, Vishnu Naresh Boddeti, Nalini K. Ratha |
FG | 4 |
| 2024 | FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSsabstractLarge pre-trained vision-language models such as CLIP provide compact and general-purpose representations of text and images that are demonstrably effective across multiple downstream zero-shot prediction tasks. However, owing to the nature of their training process, these models have the potential to 1) propagate or amplify societal biases in the training data and 2) learn to rely on spurious features. This paper proposes FairerCLIP, a general approach for making zero-shot predictions of CLIP more fair and robust to spurious correlations. We formulate the problem of jointly debiasing CLIP’s image and text representations in reproducing kernel Hilbert spaces (RKHSs), which affords multiple benefits: 1) Flexibility: Unlike existing approaches, which are specialized to either learn with or without ground-truth labels, FairerCLIP is adaptable to learning in both scenarios. 2) Ease of Optimization: FairerCLIP lends itself to an iterative optimization involving closed-form solvers, which leads to 4×-10× faster training than the existing methods. 3) Sample Efficiency: Under sample-limited conditions, FairerCLIP significantly outperforms baselines when they fail entirely. And, 4) Performance: Empirically, FairerCLIP achieves appreciable accuracy gains on benchmark fairness and spurious correlation datasets over their respective baselines. Sepehr Dehdashtian, Vishnu Naresh Boddeti |
ICLR | 3 |
| 2024 | AutoFHE: Automated Adaption of CNNs for Efficient Evaluation over FHE
Vishnu Naresh Boddeti |
USENIX Security Symposium | 2 |
| 2023 | Mitigating Task Interference in Multi-Task Learning via Explicit Task Routing with Non-Learnable PrimitivesabstractMulti-task learning (MTL) seeks to learn a single model to accomplish multiple tasks by leveraging shared information among the tasks. Existing MTL models, however, have been known to suffer from negative interference among tasks. Efforts to mitigate task interference have focused on either loss/gradient balancing or implicit parameter partitioning with partial overlaps among the tasks. In this paper, we propose ETR-NLP to mitigate task interference through a synergistic combination of non-learnable primitives (NLPs) and explicit task routing (ETR). Our key idea is to employ non-learnable primitives to extract a diverse set of task-agnostic features and recombine them into a shared branch common to all tasks and explicit task-specific branches reserved for each task. The non-learnable primitives and the explicit decoupling of learnable parameters into shared and task-specific ones afford the flexibility needed for minimizing task interference. We evaluate the efficacy of ETR-NLP networks for both image-level classification and pixel-level dense prediction MTL problems. Experimental results indicate that ETR-NLP significantly outperforms state-of-the-art baselines with fewer learnable parameters and similar FLOPs across all datasets. Code is available at this URL. Chuntao Ding, Zhichao Lu, Shangguang Wang, Ran Cheng 0004, Vishnu Naresh Boddeti |
CVPR | 5 |
| 2023 | Revisiting Residual Networks for Adversarial RobustnessabstractEfforts to improve the adversarial robustness of convolutional neural networks have primarily focused on developing more effective adversarial training methods. In contrast, little attention was devoted to analyzing the role of architectural elements (e.g., topology, depth, and width) on adversarial robustness. This paper seeks to bridge this gap and present a holistic study on the impact of architectural design on adversarial robustness. We focus on residual networks and consider architecture design at the block level as well as at the network scaling level. In both cases, we first derive insights through systematic experiments. Then we design a robust residual block, dubbed RobustResBlock, and a compound scaling rule, dubbed RobustScaling, to distribute depth and width at the desired FLOP count. Finally, we combine RobustResBlock and RobustScaling and present a portfolio of adversarially robust residual networks, RobustResNets, spanning a broad spectrum of model capacities. Experimental validation across multiple datasets and adversarial attacks demonstrate that RobustResNets consistently outperform both the standard WRNs and other existing robust architectures, achieving state-of-the-art AutoAttack robust accuracy 63.7% with 500K external data while being 2× more compact in terms of parameters. Code is available at this URL. Shihua Huang, Zhichao Lu, Kalyanmoy Deb, Vishnu Naresh Boddeti |
CVPR | 4 |
| 2023 | ProTéGé: Untrimmed Pretraining for Video Temporal Grounding by Video Temporal GroundingabstractVideo temporal grounding (VTG) is the task of localizing a given natural language text query in an arbitrarily long untrimmed video. While the task involves untrimmed videos, all existing VTG methods leverage features from video backbones pretrained on trimmed videos. This is largely due to the lack of large-scale well-annotated VTG dataset to perform pretraining. As a result, the pretrained features lack a notion of temporal boundaries leading to the video-text alignment being less distinguishable between correct and incorrect locations. We present ProTéGé as the first method to perform VTG-based untrimmed pretraining to bridge the gap between trimmed pretrained backbones and downstream VTG tasks. ProTéGé reconfigures the HowTo100M dataset, with noisily correlated video-text pairs, into a VTG dataset and introduces a novel Video-Text Similarity-based Grounding Module and a pretraining objective to make pretraining robust to noise in HowTo100M. Extensive experiments on multiple datasets across downstream tasks with all variations of supervision validate that pretrained features from ProTéGé can significantly outperform features from trimmed pretrained backbones on VTG. Gaurav Mittal, Sandra Sajeev, Ye Yu 0003, Vishnu Naresh Boddeti |
CVPR | 6 |
| 2023 | MOAZ: A Multi-Objective AutoML-Zero FrameworkabstractAutomated machine learning (AutoML) greatly eases human efforts in architecture engineering. However, mainstream AutoML methods like neural architecture search (NAS) are customized for well-designed search spaces wherein promising architectures are densely distributed. In contrast, AutoML-Zero builds machine-learning algorithms using basic primitives and can explore novel architectures beyond human knowledge. AutoML-Zero shows the potential to deploy machine learning systems by not taking advantage of either feature engineering or architectural engineering. In its current form, it only optimizes a single objective like accuracy and has no mechanism to ensure that the constraints of real-world applications are satisfied. We propose a multi-objective variant of AutoML-Zero called MOAZ, that distributes solutions on a Pareto front by trading off accuracy against the computational complexity of the machine learning algorithm. In addition to generating different Pareto-optimal solutions, MOAZ can effectively explore the sparse search space to improve search efficiency. Experimental results on linear regression tasks show MOAZ reduces the median complexity by 87.4% compared to AutoML-Zero while accelerating the median target performance achievement speed by 82%. In addition, our preliminary results on non-linear regression tasks show the potential for further improvements in search accuracy and for reducing the need for human intervention in AutoML. Ritam Guha, Vishnu Naresh Boddeti, Erik D. Goodman, Wolfgang Banzhaf, Kalyanmoy Deb |
GECCO | 4 |
| 2023 | On the Biometric Capacity of Generative Face ModelsabstractThere has been tremendous progress in generating realistic faces with high fidelity over the past few years. Despite this progress, a crucial question remains unanswered: “Given a generative face model, how many unique identities can it generate?” In other words, what is the biometric capacity of the generative face model? A scientific basis for answering this question will benefit evaluating and comparing different generative face models and establish an upper bound on their scalability. This paper proposes a statistical approach to estimate the biometric capacity of generated face images in a hyperspherical feature space. We employ our approach on multiple generative models, including unconditional generators like StyleGAN, Latent Diffusion Model, and “Generated Photos,” as well as DCFace, a class-conditional generator. We also estimate capacity w.r.t. demographic attributes such as gender and age. Our capacity estimates indicate that (a) under ArcFace representation at a false acceptance rate (FAR) of 0.1%, StyleGAN3 and DCFace have a capacity upper bound of $1.43 \times 10^{6}$ and $1.190 \times 10^{4}$, respectively; (b) the capacity reduces drastically as we lower the desired FAR with an estimate of $1.796 \times 10^{4}$ and 562 at FAR of 1% and 10%, respectively, for StyleGAN3; (c) there is no discernible disparity in the capacity w.r.t gender; and (d) for some generative models, there is an appreciable disparity in the capacity w.r.t age. Code is available at https://github.com/humananalysis/capacity-generative-face-models. Vishnu Naresh Boddeti, Gautam Sreekumar, Arun Ross |
IJCB | 1 |
| 2023 | Seed Feature Maps-based CNN Models for LEO Satellite Remote Sensing ServicesabstractDeploying high-performance convolutional neural network (CNN) models on low-earth orbit (LEO) satellites for rapid remote sensing image processing has attracted significant interest from industry and academia. However, the limited resources available on LEO satellites contrast with the demands of resource-intensive CNN models, necessitating the adoption of ground-station server assistance for training and updating these models. Existing approaches often require large floating-point operations (FLOPs) and substantial model parameter transmissions, presenting considerable challenges. To address these issues, this paper introduces a ground-station server-assisted framework. With the proposed framework, each layer of the CNN model contains only one learnable feature map (called the seed feature map) from which other feature maps are generated based on specific rules. The hyperparameters of these rules are randomly generated instead of being trained, thus enabling the generation of multiple feature maps from the seed feature map and significantly reducing FLOPs. Furthermore, since the random hyperparameters can be saved using a few random seeds, the ground station server assistance can be facilitated in updating the CNN model deployed on the LEO satellite. Experimental results on the ISPRS Vaihingen, ISPRS Potsdam, UAVid, and LoveDA datasets for semantic segmentation services demonstrate that the proposed framework outperforms existing state-of-the-art approaches. In particular, the SineFM-based model achieves a higher mIoU than the UNetFormer on the UAVid dataset, with 3.3 × fewer parameters and 2.2 × fewer FLOPs. Zhichao Lu, Chuntao Ding, Shangguang Wang, Ran Cheng 0004, Felix Juefei-Xu, Vishnu Naresh Boddeti |
ICWS | 6 |
| 2023 | Discovering Adaptable Symbolic Algorithms from ScratchabstractAutonomous robots deployed in the real world will need control policies that rapidly adapt to environmental changes. To this end, we propose AutoRobotics-Zero (ARZ), a method based on AutoML-Zero that discovers zero-shot adaptable policies from scratch. In contrast to neural network adaption policies, where only model parameters are optimized, ARZ can build control algorithms with the full expressive power of a linear register machine. We evolve modular policies that tune their model parameters and alter their inference algorithm on-the-fly to adapt to sudden environmental changes. We demonstrate our method on a realistic simulated quadruped robot, for which we evolve safe control policies that avoid falling when individual limbs suddenly break. This is a challenging task in which two popular neural network baselines fail. Finally, we conduct a detailed analysis of our method on a novel and challenging non-stationary control task dubbed Cataclysmic Cartpole. Results confirm our findings that ARZ is significantly more robust to sudden environmental changes and can build simple, interpretable control policies. Daniel S. Park, Xingyou Song, Mitchell McIntire, Pranav Nashikkar, Ritam Guha, Wolfgang Banzhaf, Kalyanmoy Deb, Vishnu Naresh Boddeti, Jie Tan 0001, Esteban Real |
IROS | 9 |
| 2023 | International Workshop on Federated Learning for Distributed Data MiningabstractThe past decade has witnessed wide applications of machine learning to various domains for decision-making, including crime detection, urban planning, drug discovery, and health monitoring, which benefited from surging data resources. As data collection in real-world applications is often done in different locations, being able to mine and discover knowledge from distributed data sources is an essential requirement for building powerful predictive models. However, directly uploading all data sources to an untrustworthy centralized data server for learning will lead to risks of privacy leakage. Federated Learning (FL) emerges as a decentralized learning framework that aggregates knowledge from distributed data without centralizing them, hence mitigating privacy risks. By hosting this workshop, we aim to attract a broad spectrum of audiences, including researchers and practitioners from academia and industry interested in the latest advances in FL. As an effort to advance the fundamental development of FL in data mining, this workshop will encourage ideas exchange on the trustworthiness, scalability, robustness, and broad applications of FL. Junyuan Hong, Zhuangdi Zhu, Lingjuan Lyu, Yang Zhou 0001, Vishnu Naresh Boddeti |
KDD | 5 |
| 2023 | Into the LAION's Den: Investigating Hate in Multimodal Datasetsabstract`Scale the model, scale the data, scale the compute' is the reigning sentiment in the world of generative AI today. While the impact of model scaling has been extensively studied, we are only beginning to scratch the surface of data scaling and its consequences. This is especially of critical importance in the context of vision-language datasets such as LAION. These datasets are continually growing in size and are built based on large-scale internet dumps such as the Common Crawl, which is known to have numerous drawbacks ranging from quality, legality, and content. The datasets then serve as the backbone for large generative models, contributing to the operationalization and perpetuation of harmful societal and historical biases and stereotypes. In this paper, we investigate the effect of scaling datasets on hateful content through a comparative audit of two datasets: LAION-400M and LAION-2B. Our results show that hate content increased by nearly 12% with dataset scale, measured both qualitatively and quantitatively using a metric that we term as Hate Content Rate (HCR). We also found that filtering dataset contents based on Not Safe For Work (NSFW) values calculated based on images alone does not exclude all the harmful content in alt-text. Instead, we found that trace amounts of hateful, targeted, and aggressive text remain even when carrying out conservative filtering. We end with a reflection and a discussion of the significance of our results for dataset curation and usage in the AI community.Code and the meta-data assets curated in this paper are publicly available at https://github.com/vinayprabhu/hate_scaling. Content warning: This paper contains examples of hateful text that might be disturbing, distressing, and/or offensive. Abeba Birhane, Vinay Uday Prabhu, Vishnu Naresh Boddeti, Sasha Luccioni |
NeurIPS | 4 |
| 2023 | Physics informed neural network for dynamic stress prediction
Hamed Bolandi, Gautam Sreekumar, Nizar Lajnef, Vishnu Naresh Boddeti |
Appl. Intell. | 5 |
| 2023 | Towards Transmission-Friendly and Robust CNN Models over Cloud and DeviceabstractDeploying deep convolutional neural network (CNN) models on ubiquitous Internet of Things (IoT) devices has attracted much attention from industry and academia since it greatly facilitates our lives by providing various rapid-response services. Due to the limited resources of IoT devices, cloud-assisted training of CNN models has become the mainstream. However, most existing related works suffer froma large amount of model parameter transmission and weak model robustness. To this end, this paper proposes a cloud-assisted CNN training framework with low model parameter transmission and strong model robustness. In the proposed framework, we first introduce MonoCNN, which contains only a few learnable filters, and other filters are nonlearnable. These nonlearnable filter parameters are generated according to certain rules, i.e., the filter generation function (FGF), and can be saved and reproduced by a few random seeds. Thus, the cloud server only needs to send these learnable filters and a few seeds to the IoT device. Compared to transmitting all model parameters, sending several learnable filter parameters and seeds can significantly reduce parameter transmission. Then, we investigate multiple FGFs and enable the IoT device to use the FGF to generate multiple filters and combine them into MonoCNN. Thus, MonoCNN is affected not only by the training data but also by the FGF. The rules of the FGF play a role in regularizing the MonoCNN, thereby improving its robustness. Experimental results show that compared to state-of-the-art methods, our proposed framework can reduce a large amount of model parameter transfer between the cloud server and the IoT device while improving the performance by approximately 2.2% when dealing with corrupted data. Chuntao Ding, Zhichao Lu, Felix Juefei-Xu, Vishnu Naresh Boddeti, Yidong Li, Jiannong Cao 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | TFormer: A Transmission-Friendly ViT Model for IoT DevicesabstractDeploying high-performance vision transformer (ViT) models on ubiquitous Internet of Things (IoT) devices to provide high-quality vision services will revolutionize the way we live, work, and interact with the world. Due to the contradiction between the limited resources of IoT devices and resource-intensive ViT models, the use of cloud servers to assist ViT model training has become mainstream. However, due to the larger number of parameters and floating-point operations (FLOPs) of the existing ViT models, the model parameters transmitted by cloud servers are large and difficult to run on resource-constrained IoT devices. To this end, this article proposes a transmission-friendly ViT model, TFormer, for deployment on resource-constrained IoT devices with the assistance of a cloud server. The high performance and small number of model parameters and FLOPs of TFormer are attributed to the proposed hybrid layer and the proposed partially connected feed-forward network (PCS-FFN). The hybrid layer consists of nonlearnable modules and a pointwise convolution, which can obtain multitype and multiscale features with only a few parameters and FLOPs to improve the TFormer performance. The PCS-FFN adopts group convolution to reduce the number of parameters. The key idea of this article is to propose TFormer with few model parameters and FLOPs to facilitate applications running on resource-constrained IoT devices to benefit from the high performance of the ViT models. Experimental results on the ImageNet-1K, MS COCO, and ADE20K datasets for image classification, object detection, and semantic segmentation tasks demonstrate that the proposed model outperforms other state-of-the-art models. Specifically, TFormer-S achieves 5% higher accuracy on ImageNet-1K than ResNet18 with 1.4× fewer parameters and FLOPs. Zhichao Lu, Chuntao Ding, Felix Juefei-Xu, Vishnu Naresh Boddeti, Shangguang Wang, Yun Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Generating Diverse 3D Reconstructions from a Single Occluded Face ImageabstractOcclusions are a common occurrence in unconstrained face images. Single image 3D reconstruction from such face images often suffers from corruption due to the pres-ence of occlusions. Furthermore, while a plurality of 3D reconstructions is plausible in the occluded regions, existing approaches are limited to generating only a single so-lution. To address both of these challenges, we present Diverse3DFace, which is specifically designed to simulta-neously generate a diverse and realistic set of 3D reconstructions from a single occluded face image. It comprises three components; a global+local shape fitting process, a graph neural network-based mesh VAE, and a determinan-tal point process based diversity-promoting iterative opti-mization procedure. Quantitative and qualitative comparisons of 3D reconstruction on occluded faces show that Di-verse3DFace can estimate 3D shapes that are consistent with the visible regions in the target image while exhibiting high, yet realistic, levels of diversity in the occluded regions. On face images occluded by masks, glasses, and other random objects, Diverse3DFace generates a distri-bution of 3D shapes having ~50% higher diversity on the occluded regions compared to the baselines. Moreover, our closest sample to the ground truth has ~40% lower MSE than the singular reconstructions by existing approaches. Code and data available at: https://github.com/human-analysis/diverse3dface Rahul Dey, Vishnu Naresh Boddeti |
CVPR | 2 |
| 2022 | Do learned representations respect causal relationships?abstractData often has many semantic attributes that are causally associated with each other. But do attribute-specific learned representations of data also respect the same causal relations? We answer this question in three steps. First, we introduce NCINet, an approach for obser-vational causal discovery from high-dimensional data. It is trained purely on synthetically generated representations and can be applied to real representations, and is specif-ically designed to mitigate the domain gap between the two. Second, we apply NCINet to identify the causal relations between image representations of different pairs of at-tributes with known and unknown causal relations between the labels. For this purpose, we consider image represen-tations learned for predicting attributes on the 3D Shapes, CelebA, and the CASIA-WebFace datasets, which we an-notate with multiple multi-class attributes. Third, we an-alyze the effect on the underlying causal relation between learned representations induced by various design choices in representation learning. Our experiments indicate that (1) NCINet significantly outperforms existing observational causal discovery approaches for estimating the causal relation between pairs of random samples, both in the presence and absence of an unobserved confounder, (2) under controlled scenarios, learned representations can indeed satisfy the underlying causal relations between their respective labels, and (3) the causal relations are positively correlated with the predictive capability of the representations. Code and annotations are available at: https://github.com/human-analysis/causal-relations-between-representations. Vishnu Naresh Boddeti |
CVPR | 2 |
| 2022 | HEFT: Homomorphically Encrypted Fusion of Biometric TemplatesabstractThis paper proposes a non-interactive end-to-end solution for secure fusion and matching of biometric templates using fully homomorphic encryption (FHE). Given a pair of encrypted feature vectors, we perform the following ciphertext operations, i) feature concatenation, ii) fusion and dimensionality reduction through a learned linear projection, iii) scale normalization to unit ℓ2-norm, and iv) match score computation. Our method, dubbed HEFT (Homomorphi-cally Encrypted Fusion of biometric Templates), is custom-designed to overcome the unique constraint imposed by FHE, namely the lack of support for non-arithmetic operations. From an inference perspective, we systematically explore different data packing schemes for computationally efficient linear projection and introduce a polynomial approximation for scale normalization. From a training perspective, we introduce an FHE-aware algorithm for learning the linear projection matrix to mitigate errors induced by approximate normalization. Experimental evaluation for template fusion and matching of face and voice biometrics shows that HEFT (i) improves biometric verification performance by 11.07% and 9.58% AUROC compared to the respective unibiometric representations while compressing the feature vectors by a factor of 16 (512D to 32D), and (ii) fuses a pair of encrypted feature vectors and computes its match score against a gallery of size 1024 in 884 ms. Code and data are available at https://github.com/humananalysis/encrypted-biometric-fusion Luke Sperling, Nalini K. Ratha, Arun Ross, Vishnu Naresh Boddeti |
IJCB | 4 |
| 2022 | 3DFaceFill: An Analysis-By-Synthesis Approach to Face CompletionabstractExisting face completion solutions are primarily driven by end-to-end models that directly generate 2D completions of 2D masked faces. By having to implicitly account for geometric and photometric variations in facial shape and appearance, such approaches result in unrealistic completions, especially under large variations in pose, shape, illumination and mask sizes. To alleviate these limitations, we introduce 3DFaceFill, an analysis-by-synthesis approach for face completion that explicitly considers the image formation process. It comprises three components, (1) an encoder that disentangles the face into its constituent 3D mesh, 3D pose, illumination and albedo factors, (2) an autoencoder that inpaints the UV representation of facial albedo, and (3) a renderer that resynthesizes the completed face. By operating on the UV representation, 3DFaceFill affords the power of correspondence and allows us to naturally enforce geometrical priors (e.g. facial symmetry) more effectively. Quantitatively, 3DFaceFill improves the state-of-the-art by up to 4dB higher PSNR and 25% better LPIPS for large masks. And, qualitatively, it leads to demonstrably more photorealistic face completions over a range of masks and occlusions while preserving consistency in global and component-wise shape, pose, illumination and eye-gaze. Rahul Dey, Vishnu Naresh Boddeti |
WACV | 2 |
| 2021 | Multi-objective Coevolution and Decision-making for Cooperative and Competitive EnvironmentsabstractCo-evolutionary algorithms involve two co-evolving populations, each having its own set of objectives and constraints, that interact with each other during function evaluation. Co-evolutionary algorithms are of great interest in cooperative and competing games and search tasks in which multiple agents having different interests are in play. Despite a number of single-objective co-evolutionary studies, there has been limited interest in multi-objective co-evolutionary algorithms. A recent study has revealed that in addition to the challenges associated with the development of an efficient algorithm, a proper understanding of the conflicting objectives within a single population and their interaction among objectives of the second population becomes extremely difficult to comprehend. In this paper, we extend the previous proof-of-principle multi-objective co-evolutionary (MOCoEv) study in three important directions. First, we enhance MOCoEv's ability to handle mixed cooperating and conflicting scenarios among different players. Second, we propose an iterative multi-criterion decision-making (MCDM) approach to demonstrate how, in an arms-race type scenario, the most appropriate solution can be selected from the obtained Pareto-optimal solution set iteratively. Third, we extend the previous MOCoEv algorithm with a many-objective evolutionary algorithm (NSGA-III) to make them applicable to three or more objectives for each player. These three developments reveal better insights about the intricate issues related to multiple objectives and decision-making for co-evolutionary optimization and take MOCoEv a step closer to solving more complex multi-player problems. Anirudh Suresh, Jaturong Kongmanee, Kalyanmoy Deb, Vishnu Naresh Boddeti |
CEC | 4 |
| 2021 | Towards Multi-objective Co-evolutionary Problem Solving
Anirudh Suresh, Kalyanmoy Deb, Vishnu Naresh Boddeti |
EMO | 3 |
| 2021 | Spatially-Adaptive Image Restoration using Distortion-Guided NetworksabstractWe present a general learning-based solution for restoring images suffering from spatially-varying degradations. Prior approaches are typically degradation-specific and employ the same processing across different images and different pixels within. However, we hypothesize that such spatially rigid processing is suboptimal for simultaneously restoring the degraded pixels as well as reconstructing the clean regions of the image. To overcome this limitation, we propose SPAIR, a network design that harnesses distortion-localization information and dynamically adjusts computation to difficult regions in the image. SPAIR comprises of two components, (1) a localization network that identifies degraded pixels, and (2) a restoration network that exploits knowledge from the localization network in filter and feature domain to selectively and adaptively restore degraded pixels. Our key idea is to exploit the non-uniformity of heavy degradations in spatial-domain and suitably embed this knowledge within distortion-guided modules performing sparse normalization, feature extraction and attention. Our architecture is agnostic to physical formation model and generalizes across several types of spatially-varying degradations. We demonstrate the efficacy of SPAIR individually on four restoration tasks- removal of rain-streaks, raindrops, shadows and motion blur. Extensive qualitative and quantitative comparisons with prior art on 11 benchmark datasets demonstrate that our degradation-agnostic network design offers significant performance gains over state-of-the-art degradation-specific architectures. Code available at https://github.com/humananalysis/spatially-adaptive-image-restoration. Kuldeep Purohit, Maitreya Suin, A. N. Rajagopalan 0001, Vishnu Naresh Boddeti |
ICCV | 4 |
| 2021 | Adversarial Representation Learning with Closed-Form SolversabstractAdversarial representation learning aims to learn data representations for a target task while removing unwanted sensitive information at the same time. Existing methods learn model parameters iteratively through stochastic gradient descent-ascent, which is often unstable and unreliable in practice. To overcome this challenge, we adopt closed-form solvers for the adversary and target task. We model them as kernel ridge regressors and analytically determine an upper-bound on the optimal dimensionality of representation. Our solution, dubbed OptNet-ARL, reduces to a stable one one-shot optimization problem that can be solved reliably and efficiently. OptNet-ARL can be easily generalized to the case of multiple target tasks and sensitive attributes. Numerical experiments, on both small and large scale datasets, show that, from an optimization perspective, OptNet-ARL is stable and exhibits three to five times faster convergence. Performance wise, when the target and sensitive attributes are dependent, OptNet-ARL learns representations that offer a better trade-off front between (a) utility and bias for fair classification and (b) utility and privacy by mitigating leakage of private information than existing solutions.Code is available at https://github.com/human-analysis. Bashir Sadeghi, Vishnu Naresh Boddeti |
ECML/PKDD (2) | 3 |
| 2021 | Neural Architecture TransferabstractNeural architecture search (NAS) has emerged as a promising avenue for automatically designing task-specific neural networks. Existing NAS approaches require one complete search for each deployment specification of hardware or objective. This is a computationally impractical endeavor given the potentially large number of application scenarios. In this paper, we propose Neural Architecture Transfer (NAT) to overcome this limitation. NAT is designed to efficiently generate task-specific custom models that are competitive under multiple conflicting objectives. To realize this goal we learn task-specific supernets from which specialized subnets can be sampled without any additional training. The key to our approach is an integrated online transfer learning and many-objective evolutionary search procedure. A pre-trained supernet is iteratively adapted while simultaneously searching for task-specific subnets. We demonstrate the efficacy of NAT on 11 benchmark image classification tasks ranging from large-scale multi-class to small-scale fine-grained datasets. In all cases, including ImageNet, NATNets improve upon the state-of-the-art under mobile settings ( ≤ 600M Multiply-Adds). Surprisingly, small-scale fine-grained datasets benefit the most from NAT. At the same time, the architecture search and transfer is orders of magnitude more efficient than existing NAS methods. Overall, experimental evaluation indicates that, across diverse image classification tasks and computational objectives, NAT is an appreciably more effective alternative to conventional transfer learning of fine-tuning weights of an existing network architecture learned on standard datasets. Code is available at https://github.com/human-analysis/neural-architecture-transfer. Zhichao Lu, Gautam Sreekumar, Erik D. Goodman, Wolfgang Banzhaf, Kalyanmoy Deb, Vishnu Naresh Boddeti |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Multiobjective Evolutionary Design of Deep Convolutional Neural Networks for Image ClassificationabstractConvolutional neural networks (CNNs) are the backbones of deep learning paradigms for numerous vision tasks. Early advancements in CNN architectures are primarily driven by human expertise and by elaborate design processes. Recently, neural architecture search was proposed with the aim of automating the network design process and generating task-dependent architectures. While existing approaches have achieved competitive performance in image classification, they are not well suited to problems where the computational budget is limited for two reasons: 1) the obtained architectures are either solely optimized for classification performance, or only for one deployment scenario and 2) the search process requires vast computational resources in most approaches. To overcome these limitations, we propose an evolutionary algorithm for searching neural architectures under multiple objectives, such as classification performance and floating point operations (FLOPs). The proposed method addresses the first shortcoming by populating a set of architectures to approximate the entire Pareto frontier through genetic operations that recombine and modify architectural components progressively. Our approach improves computational efficiency by carefully down-scaling the architectures during the search as well as reinforcing the patterns commonly shared among past successful architectures through Bayesian model learning. The integration of these two main contributions allows an efficient design of architectures that are competitive and in most cases outperform both manually and automatically designed architectures on benchmark image classification datasets: CIFAR, ImageNet, and human chest X-ray. The flexibility provided from simultaneously obtaining multiple architecture choices for different compute requirements further differentiates our approach from other methods in the literature. Zhichao Lu, Ian Whalen, Yashesh D. Dhebar, Kalyanmoy Deb, Erik D. Goodman, Wolfgang Banzhaf, Vishnu Naresh Boddeti |
IEEE Trans. Evol. Comput. | 7 |
| 2020 | MUXConv: Information Multiplexing in Convolutional Neural NetworksabstractConvolutional neural networks have witnessed remarkable improvements in computational efficiency in recent years. A key driving force has been the idea of trading-off model expressivity and efficiency through a combination of 1x1 and depth-wise separable convolutions in lieu of a standard convolutional layer. The price of the efficiency, however, is the sub-optimal flow of information across space and channels in the network. To overcome this limitation, we present MUXConv, a layer that is designed to increase the flow of information by progressively multiplexing channel and spatial information in the network, while mitigating computational complexity. Furthermore, to demonstrate the effectiveness of MUXConv, we integrate it within an efficient multi-objective evolutionary algorithm to search for the optimal model hyper-parameters while simultaneously optimizing accuracy, compactness, and computational efficiency. On ImageNet, the resulting models, dubbed MUXNets, match the performance (75.3% top-1 accuracy) and multiply-add operations (218M) of MobileNetV3 while being 1.6x more compact, and outperform other mobile models in all the three criteria. MUXNet also performs well under transfer learning and when adapted to object detection. On the ChestX-Ray 14 benchmark, its accuracy is comparable to the state-of-the-art while being 3.3x more compact and 14x more efficient. Similarly, detection on PASCAL VOC 2007 is 1.2% more accurate, 28% faster and 6% more compact compared to MobileNetV2. Zhichao Lu, Kalyanmoy Deb, Vishnu Naresh Boddeti |
CVPR | 3 |
| 2020 | NSGANetV2: Evolutionary Multi-objective Surrogate-Assisted Neural Architecture Search
Zhichao Lu, Kalyanmoy Deb, Erik D. Goodman, Wolfgang Banzhaf, Vishnu Naresh Boddeti |
ECCV (1) | 5 |
| 2020 | NSGA-Net: Neural Architecture Search using Multi-Objective Genetic Algorithm (Extended Abstract)abstractConvolutional neural networks (CNNs) are the backbones of deep learning paradigms for numerous vision tasks. Early advancements in CNN architectures are primarily driven by human expertise and elaborate design. Recently, neural architecture search (NAS) was proposed with the aim of automating the network design process and generating task-dependent architectures. This paper introduces NSGA-Net -- an evolutionary search algorithm that explores a space of potential neural network architectures in three steps, namely, a population initialization step that is based on prior-knowledge from hand-crafted architectures, an exploration step comprising crossover and mutation of architectures, and finally an exploitation step that utilizes the hidden useful knowledge stored in the entire history of evaluated neural architectures in the form of a Bayesian Network. The integration of these components allows an efficient design of architectures that are competitive and in many cases outperform both manually and automatically designed architectures on CIFAR-10 classification task. The flexibility provided from simultaneously obtaining multiple architecture choices for different compute requirements further differentiates our approach from other methods in the literature. Zhichao Lu, Ian Whalen, Yashesh D. Dhebar, Kalyanmoy Deb, Erik D. Goodman, Wolfgang Banzhaf, Vishnu Naresh Boddeti |
IJCAI | 7 |
| 2019 | On the Intrinsic Dimensionality of Image RepresentationsabstractThis paper addresses the following questions pertaining to the intrinsic dimensionality of any given image representation: (i) estimate its intrinsic dimensionality, (ii) develop a deep neural network based non-linear mapping, dubbed DeepMDS, that transforms the ambient representation to the minimal intrinsic space, and (iii) validate the veracity of the mapping through image matching in the intrinsic space. Experiments on benchmark image datasets (LFW, IJB-C and ImageNet-100) reveal that the intrinsic dimensionality of deep neural network representations is significantly lower than the dimensionality of the ambient features. For instance, SphereFace's 512-dim face representation and ResNet's 512-dim image representation have an intrinsic dimensionality of 16 and 19 respectively. Further, the DeepMDS mapping is able to obtain a representation of significantly lower dimensionality while maintaining discriminative ability to a large extent, 59.75% TAR @ 0.1% FAR in 16-dim vs 71.26% TAR in 512-dim on IJB-C and a Top-1 accuracy of 77.0% at 19-dim vs 83.4% at 512-dim on ImageNet-100. Sixue Gong, Vishnu Naresh Boddeti, Anil K. Jain 0001 |
CVPR | 2 |
| 2019 | Mitigating Information Leakage in Image Representations: A Maximum Entropy ApproachabstractImage recognition systems have demonstrated tremendous progress over the past few decades thanks, in part, to our ability of learning compact and robust representations of images. As we witness the wide spread adoption of these systems, it is imperative to consider the problem of unintended leakage of information from an image representation, which might compromise the privacy of the data owner. This paper investigates the problem of learning an image representation that minimizes such leakage of user information. We formulate the problem as an adversarial non-zero sum game of finding a good embedding function with two competing goals: to retain as much task dependent discriminative image information as possible, while simultaneously minimizing the amount of information, as measured by entropy, about other sensitive attributes of the user. We analyze the stability and convergence dynamics of the proposed formulation using tools from non-linear systems theory and compare to that of the corresponding adversarial zero-sum game formulation that optimizes likelihood as a measure of information content. Numerical experiments on UCI, Extended Yale B, CIFAR-10 and CIFAR-100 datasets indicate that our proposed approach is able to learn image representations that exhibit high task performance while mitigating leakage of predefined sensitive information. Proteek Chandan Roy, Vishnu Naresh Boddeti |
CVPR | 2 |
| 2019 | NSGA-Net: neural architecture search using multi-objective genetic algorithmabstractThis paper introduces NSGA-Net --- an evolutionary approach for neural architecture search (NAS). NSGA-Net is designed with three goals in mind: (1) a procedure considering multiple and conflicting objectives, (2) an efficient procedure balancing exploration and exploitation of the space of potential neural network architectures, and (3) a procedure finding a diverse set of trade-off network architectures achieved in a single run. NSGA-Net is a population-based search algorithm that explores a space of potential neural network architectures in three steps, namely, a population initialization step that is based on prior-knowledge from hand-crafted architectures, an exploration step comprising crossover and mutation of architectures, and finally an exploitation step that utilizes the hidden useful knowledge stored in the entire history of evaluated neural architectures in the form of a Bayesian Network. Experimental results suggest that combining the dual objectives of minimizing an error metric and computational complexity, as measured by FLOPs, allows NSGA-Net to find competitive neural architectures. Moreover, NSGA-Net achieves error rate on the CIFAR-10 dataset on par with other state-of-the-art NAS methods while using orders of magnitude less computational resources. These results are encouraging and shows the promise to further use of EC methods in various deep-learning paradigms. Zhichao Lu, Ian Whalen, Vishnu Naresh Boddeti, Yashesh D. Dhebar, Kalyanmoy Deb, Erik D. Goodman, Wolfgang Banzhaf |
GECCO | 3 |
| 2019 | On the Global Optima of Kernelized Adversarial Representation LearningabstractAdversarial representation learning is a promising paradigm for obtaining data representations that are invariant to certain sensitive attributes while retaining the information necessary for predicting target attributes. Existing approaches solve this problem through iterative adversarial minimax optimization and lack theoretical guarantees. In this paper, we first study the “linear” form of this problem i.e., the setting where all the players are linear functions. We show that the resulting optimization problem is both non-convex and non-differentiable. We obtain an exact closed-form expression for its global optima through spectral learning and provide performance guarantees in terms of analytical bounds on the achievable utility and invariance. We then extend this solution and analysis to non-linear functions through kernel representation. Numerical experiments on UCI, Extended Yale B and CIFAR100 datasets indicate that, (a) practically, our solution is ideal for “imparting” provable invariance to any biased pre-trained data representation, and (b) the global optima of the “kernel” form can provide a comparable trade-off between utility and invariance in comparison to iterative minimax optimization of existing deep neural network based approaches, but with provable guarantees. Bashir Sadeghi, Runyi Yu 0001, Vishnu Naresh Boddeti |
ICCV | 3 |
| 2018 | Efficient K-Shot Learning With Regularized Deep NetworksabstractFeature representations from pre-trained deep neural networks have been known to exhibit excellent generalization and utility across a variety of related tasks. Fine-tuning is by far the simplest and most widely used approach that seeks to exploit and adapt these feature representations to novel tasks with limited data. Despite the effectiveness of fine-tuning, it is often sub-optimal and requires very careful optimization to prevent severe over-fitting to small datasets. The problem of sub-optimality and overfitting, is due in part to the large number of parameters used in a typical deep convolutional neural network. To address these problems, we propose a simple yet effective regularization method for fine-tuning pre-trained deep networks for the task of k-shot learning. To prevent overfitting, our key strategy is to cluster the model parameters while ensuring intra-cluster similarity and inter-cluster diversity of the parameters, effectively regularizing the dimensionality of the parameter search space. In particular, we identify groups of neurons within each layer of a deep network that shares similar activation patterns. When the network is to be fine-tuned for a classification task using only k examples, we propagate a single gradient to all of the neuron parameters that belong to the same group. The grouping of neurons is non-trivial as neuron activations depend on the distribution of the input data. To efficiently search for optimal groupings conditioned on the input data, we propose a reinforcement learning search strategy using recurrent networks to learn the optimal group assignments for each network layer. Experimental results show that our method can be easily applied to several popular convolutional neural networks and improve upon other state-of-the-art fine-tuning based k-shot learning strategies by more than 10%. Donghyun Yoo, Haoqi Fan 0001, Vishnu Naresh Boddeti, Kris Makoto Kitani |
AAAI | 3 |
| 2018 | RankGAN: A Maximum Margin Ranking GAN for Generating Faces
Felix Juefei-Xu, Rahul Dey, Vishnu Naresh Boddeti, Marios Savvides |
ACCV (3) | 3 |
| 2018 | Perturbative Neural NetworksabstractConvolutional neural networks are witnessing wide adoption in computer vision systems with numerous applications across a range of visual recognition tasks. Much of this progress is fueled through advances in convolutional neural network architectures and learning algorithms even as the basic premise of a convolutional layer has remained unchanged. In this paper, we seek to revisit the convolutional layer that has been the workhorse of state-of-the-art visual recognition models. We introduce a very simple, yet effective, module called a perturbation layer as an alternative to a convolutional layer. The perturbation layer does away with convolution in the traditional sense and instead computes its response as a weighted linear combination of non-linearly activated additive noise perturbed inputs. We demonstrate both analytically and empirically that this perturbation layer can be an effective replacement for a standard convolutional layer. Empirically, deep neural networks with perturbation layers, called Perturbative Neural Networks (PNNs), in lieu of convolutional layers perform comparably with standard CNNs on a range of visual datasets (MNIST, CIFAR-10, PASCAL VOC, and ImageNet) with fewer parameters. Felix Juefei-Xu, Vishnu Naresh Boddeti, Marios Savvides |
CVPR | 2 |
| 2018 | Synthesizing a Scene-Specific Pedestrian Detector and Pose Estimator for Static Video Surveillance - Can We Learn Pedestrian Detectors and Pose Estimators Without Real Data?
Hironori Hattori, Namhoon Lee, Vishnu Naresh Boddeti, Fares Beainy, Kris Makoto Kitani, Takeo Kanade |
Int. J. Comput. Vis. | 3 |
| 2017 | Local Binary Convolutional Neural Networks
Felix Juefei-Xu, Vishnu Naresh Boddeti, Marios Savvides |
CVPR | 2 |
| 2017 | Privacy-Preserving Visual Learning Using Doubly Permuted Homomorphic EncryptionabstractWe propose a privacy-preserving framework for learning visual classifiers by leveraging distributed private image data. This framework is designed to aggregate multiple classifiers updated locally using private data and to ensure that no private information about the data is exposed during and after its learning procedure. We utilize a homomorphic cryptosystem that can aggregate the local classifiers while they are encrypted and thus kept secret. To overcome the high computational cost of homomorphic encryption of high-dimensional classifiers, we (1) impose sparsity constraints on local classifier updates and (2) propose a novel efficient encryption scheme named doublypermuted homomorphic encryption (DPHE) which is tailored to sparse high-dimensional data. DPHE (i) decomposes sparse data into its constituent non-zero values and their corresponding support indices, (ii) applies homomorphic encryption only to the non-zero values, and (iii) employs double permutations on the support indices to make them secret. Our experimental evaluation on several public datasets shows that the proposed approach achieves comparable performance against state-of-the-art visual recognition methods while preserving privacy and significantly outperforms other privacy-preserving methods. Ryo Yonetani, Vishnu Naresh Boddeti, Kris Makoto Kitani, Yoichi Sato 0001 |
ICCV | 2 |
| 2016 | Stacked correlation filters for biometric verificationabstractCorrelation filters (CFs) are a well-known pattern classification approach used in biometrics. A CF is a spatial-frequency array that is specifically synthesized from a set of training patterns to produce a sharp correlation output peak at the location of the best match for an authentic image comparison and no such peak for an impostor image comparison. The underlying premise when using CFs is that this correlation output peak behavior on training data ideally extends to testing data. Yet in 1:1 verification scenarios, where there is limited training data available to represent pattern distortions, the correlation output from an authentic comparison can be difficult to discern from the correlation output from an impostor. In this paper we introduce Stacked Correlation Filters (SCFs), a simple and powerful approach to address this problem by training an additional set of classifiers which learn to differentiate correlation outputs from authentic and impostor match pairs. This is done by training a series of stacked modular CFs with each layer refining the output of the previous layer. Our basic premise is that since correlation outputs have an expected shape, an additional CF can be trained to recognize such shape and refine the final output. As previous works with CFs have only focused on individual filter design or application, which assumes the CF to provide a sharp peak, this is a new CF paradigm that can benefit many existing CF designs and applications. Jonathon M. Smereka, Vishnu Naresh Boddeti, B. V. K. Vijaya Kumar, Andres Rodriguez 0001 |
ICASSP | 2 |
| 2015 | Learning scene-specific pedestrian detectors without real dataabstractWe consider the problem of designing a scene-specific pedestrian detector in a scenario where we have zero instances of real pedestrian data (i.e., no labeled real data or unsupervised real data). This scenario may arise when a new surveillance system is installed in a novel location and a scene-specific pedestrian detector must be trained prior to any observations of pedestrians. The key idea of our approach is to infer the potential appearance of pedestrians using geometric scene data and a customizable database of virtual simulations of pedestrian motion. We propose an efficient discriminative learning method that generates a spatially-varying pedestrian appearance model that takes into the account the perspective geometry of the scene. As a result, our method is able to learn a unique pedestrian classifier customized for every possible location in the scene. Our experimental results show that our proposed approach outperforms classical pedestrian detection models and hybrid synthetic-real models. Our results also yield a surprising result, that our method using purely synthetic data is able to outperform models trained on real scene-specific data when data is limited. Hironori Hattori, Vishnu Naresh Boddeti, Kris Makoto Kitani, Takeo Kanade |
CVPR | 2 |
| 2015 | Face Alignment RefinementabstractAchieving sub-pixel accuracy with face alignment algorithms is a difficult task given the diversity of appearance in real world facial profiles. To capture variations in perspective, occlusion, and illumination with adequate precision, current face alignment approaches rely on detecting facial landmarks and iteratively adjusting deformable models that encode prior knowledge of facial structure. However, these methods involve optimization in latent sub-spaces, where user-specific face shape information is easily lost after dimensionality reduction. Attempting to retain this information to capture this wide range of variation requires a large training distribution, which is difficult to obtain without high computational complexity. Subsequently, many face alignment methods lack the pixel-level accuracy necessary to satisfy the aesthetic requirements of tasks such as face deidentification, face swapping, and face modeling. In many such applications, the primary source of aesthetic inadequacy is a misaligned jaw line or facial contour. In this work, we explore the idea of an image-based refinement method to fix the landmark points of a misaligned facial contour. We propose an efficient two stage process - an intuitively constructed edge detection based algorithm to actively adjust facial contour landmark points, and a data driven validation system to filter out erroneous adjustments. Experimental results show that state-of-the-art face alignment combined with our proposed post-processing method yields improved overall performance over multiple face image datasets. Andy Zeng 0001, Vishnu Naresh Boddeti, Kris Makoto Kitani, Takeo Kanade |
WACV | 2 |
| 2015 | Zero-Aliasing Correlation Filters for Object RecognitionabstractCorrelation filters (CFs) are a class of classifiers that are attractive for object localization and tracking applications. Traditionally, CFs have been designed in the frequency domain using the discrete Fourier transform (DFT), where correlation is efficiently implemented. However, existing CF designs do not account for the fact that the multiplication of two DFTs in the frequency domain corresponds to a circular correlation in the time/spatial domain. Because this was previously unaccounted for, prior CF designs are not truly optimal, as their optimization criteria do not accurately quantify their optimization intention. In this paper, we introduce new zero-aliasing constraints that completely eliminate this aliasing problem by ensuring that the optimization criterion for a given CF corresponds to a linear correlation rather than a circular correlation. This means that previous CF designs can be significantly improved by this reformulation. We demonstrate the benefits of this new CF design approach with several important CFs. We present experimental results on diverse data sets and present solutions to the computational challenges associated with computing these CFs. Code for the CFs described in this paper and their respective zero-aliasing versions is available at http://vishnu.boddeti.net/projects/correlation-filters.html. Joseph A. Fernandez, Vishnu Naresh Boddeti, Andres Rodriguez 0001, B. V. K. Vijaya Kumar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Probabilistic Deformation Models for Challenging Periocular Image VerificationabstractThe periocular region as a biometric trait has recently gained considerable traction, especially under challenging scenarios where reliable iris information is not available for human authentication. In this paper, we consider the problem of one-to-one (1 : 1) matching of highly nonideal periocular images captured in-the-wild under unconstrained imaging conditions. Such images exhibit considerable appearance variations, including nonuniform illumination variations, motion and defocus blur, off-axis gaze, and nonstationary pattern deformations. To address these challenges, we propose periocular probabilistic deformation models (PPDMs) that: 1) reduce the image matching problem to matching local image regions and 2) approximate the periocular distortions by local patch level spatial translations whose relationships are modeled by a Gaussian Markov random field. Given a periocular image pair, we determine the distortion-tolerant similarity metric by regularizing local match scores by the maximum aposteriori probability estimate of the relative local deformations between them. Unlike the existing global periocular image matching techniques, by accounting for local image deformations in the periocular matching process, PPDM exhibits greater tolerance to pattern variations. We demonstrate the effectiveness of our model via extensive evaluation on a large number of in-the-wild periocular images. We find that PPDMs outperform many benchmark 1 : 1 image matching techniques (improving verification rates at 0.1% false accept rate by ~30% over previous work and ~40% when compared with the best baseline) in challenging scenarios leading to state-of-the-art verification performance on multiple real-world periocular data sets. Jonathon M. Smereka, Vishnu Naresh Boddeti, B. V. K. Vijaya Kumar |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | 3D Pose-by-Detection of Vehicles via Discriminatively Reduced Ensembles of Correlation Filters
Yair Movshovitz-Attias, Yaser Sheikh, Vishnu Naresh Boddeti, Zijun Wei |
BMVC | 3 |
| 2013 | Correlation Filters for Object AlignmentabstractAlignment of 3D objects from 2D images is one of the most important and well studied problems in computer vision. A typical object alignment system consists of a landmark appearance model which is used to obtain an initial shape and a shape model which refines this initial shape by correcting the initialization errors. Since errors in landmark initialization from the appearance model propagate through the shape model, it is critical to have a robust landmark appearance model. While there has been much progress in designing sophisticated and robust shape models, there has been relatively less progress in designing robust landmark detection models. In this paper we present an efficient and robust landmark detection model which is designed specifically to minimize localization errors thereby leading to state-of-the-art object alignment performance. We demonstrate the efficacy and speed of the proposed approach on the challenging task of multi-view car alignment. Vishnu Naresh Boddeti, Takeo Kanade, B. V. K. Vijaya Kumar |
CVPR | 1 |
| 2013 | A Framework for Binding and Retrieving Class-Specific Information to and from Image Patterns Using Correlation FiltersabstractWe describe a template-based framework to bind class-specific information to a set of image patterns and retrieve that information by matching the template to a query pattern of the same class. This is done by mapping the class-specific information to a set of spatial translations which are applied to the set of image patterns from which a template is designed, taking advantage of the properties of correlation filters. The bound information is retrieved during matching with an authentic query by estimating the spatial translations applied to the images that were used to design the template. In this paper, we focus on the problem of binding information to biometric signatures as an application of our framework. Our framework is flexible enough to allow spreading the information to be bound over multiple pattern classes which, in the context of biometric key-binding, enables multiclass and multimodal biometric key-binding. We demonstrate the effectiveness of the proposed scheme via extensive numerical results on multiple biometric databases. Vishnu Naresh Boddeti, B. V. K. Vijaya Kumar |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Maximum Margin Correlation Filter: A New Approach for Localization and ClassificationabstractSupport vector machine (SVM) classifiers are popular in many computer vision tasks. In most of them, the SVM classifier assumes that the object to be classified is centered in the query image, which might not always be valid, e.g., when locating and classifying a particular class of vehicles in a large scene. In this paper, we introduce a new classifier called Maximum Margin Correlation Filter (MMCF), which, while exhibiting the good generalization capabilities of SVM classifiers, is also capable of localizing objects of interest, thereby avoiding the need for image centering as is usually required in SVM classifiers. In other words, MMCF can simultaneously localize and classify objects of interest. We test the efficacy of the proposed classifier on three different tasks: vehicle recognition, eye localization, and face classification. We demonstrate that MMCF outperforms SVM classifiers as well as well known correlation filters. Andres Rodriguez 0001, Vishnu Naresh Boddeti, B. V. K. Vijaya Kumar, Abhijit Mahalanobis |
IEEE Trans. Image Process. | 2 |
| 2012 | RainMon: an integrated approach to mining bursty timeseries monitoring dataabstractMetrics like disk activity and network traffic are widespread sources of diagnosis and monitoring information in datacenters and networks. However, as the scale of these systems increases, examining the raw data yields diminishing insight. We present RainMon, a novel end-to-end approach for mining timeseries monitoring data designed to handle its size and unique characteristics. Our system is able to (a) mine large, bursty, real-world monitoring data, (b) find significant trends and anomalies in the data, (c) compress the raw data effectively, and (d) estimate trends to make forecasts. Furthermore, RainMon integrates the full analysis process from data storage to the user interface to provide accessible long-term diagnosis. We apply RainMon to three real-world datasets from production systems and show its utility in discovering anomalous machines and time periods. Ilari Shafer, Kai Ren 0001, Vishnu Naresh Boddeti, Yoshihisa Abe, Gregory R. Ganger, Christos Faloutsos |
KDD | 3 |
| 2011 | A comparative evaluation of iris and ocular recognition methods on challenging ocular imagesabstractIris recognition is believed to offer excellent recognition rates for iris images acquired under controlled conditions. However, recognition rates degrade considerably when images exhibit impairments such as off-axis gaze, partial occlusions, specular reflections and out-of-focus and motion-induced blur. In this paper, we use the recently-available face and ocular challenge set (FOCS) to investigate the comparative recognition performance gains of using ocular images (i.e., iris regions as well as the surrounding peri-ocular regions) instead of just the iris regions. A new method for ocular recognition is presented and it is shown that use of ocular regions leads to better recognition rates than iris recognition on FOCS dataset. Another advantage of using ocular images for recognition is that it avoids the need for segmenting the iris images from their surrounding regions. Vishnu Naresh Boddeti, Jonathon M. Smereka, B. V. K. Vijaya Kumar |
IJCB | 1 |
| 2010 | Extended-Depth-of-Field Iris Recognition Using Unrestored Wavefront-Coded ImageryabstractIris recognition can offer high-accuracy person recognition, particularly when the acquired iris image is well focused. However, in some practical scenarios, user cooperation may not be sufficient to acquire iris images in focus; therefore, iris recognition using camera systems with a large depth of field is very desirable. One approach to achieve extended depth of field is to use a wavefront-coding system as proposed by Dowski and Cathey, which uses a phase modulation mask. The conventional approach when using a camera system with such a phase mask is to restore the raw images acquired from the camera before feeding them into the iris recognition module. In this paper, we investigate the feasibility of skipping the image restoration step with minimal degradation in recognition performance while still increasing the depth of field of the whole system compared to an imaging system without a phase mask. By using a simulated wavefront-coded imagery, we present the results of two different iris recognition algorithms, namely, Daugman's iriscode and correlation-filter-based iris recognition, using more than 1000 iris images taken from the iris challenge evaluation database. We carefully study the effect of an off-the-shelf phase mask on iris segmentation and iris matching, and finally, to better enable the use of unrestored wavefront-coded images, we design a custom phase mask by formulating an optimization problem. Our results suggest that, in exchange for some degradation in recognition performance at the best focus, we can increase the depth of field by a factor of about four (over a conventional camera system without a phase mask) by carefully designing the phase masks. Vishnu Naresh Boddeti, B. V. K. Vijaya Kumar |
IEEE Trans. Syst. Man Cybern. Part A | 1 |