VLDB 2026 Research / reviewers in the wild / expert
Sachin Mehta
dblp:34/11140
· DBLP profile ↗
32ranked-venue papers
17as first author
17since 2021 · last 2025
0000-0002-5420-4725ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 10 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TiC-LM: A Web-Scale Benchmark for Time-Continual LLM PretrainingabstractJeffrey Li, Mohammadreza Armandpour, Seyed Iman Mirzadeh, Sachin Mehta, Vaishaal Shankar, Raviteja Vemulapalli, Samy Bengio, Oncel Tuzel, Mehrdad Farajtabar, Hadi Pouransari, Fartash Faghri. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jeffrey Li, Mohammadreza Armandpour, Iman Mirzadeh, Sachin Mehta, Vaishaal Shankar, Raviteja Vemulapalli, Samy Bengio, Oncel Tuzel, Mehrdad Farajtabar, Hadi Pouransari, Fartash Faghri |
ACL (1) | 4 |
| 2025 | SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random GeneratorsabstractLarge Language Models (LLMs) have transformed natural language processing, but face significant challenges in widespread deployment due to their high runtime cost. In this paper, we introduce SeedLM, a novel post-training compression method that uses seeds of a pseudo-random generator to encode and compress model weights. Specifically, for each block of weights, we find a seed that is fed into a Linear Feedback Shift Register (LFSR) during inference to efficiently generate a random matrix. This matrix is then linearly combined with compressed coefficients to reconstruct the weight block. SeedLM reduces memory access and leverages idle compute cycles during inference, effectively speeding up memory-bound tasks by trading compute for fewer memory accesses. Unlike state-of-the-art methods that rely on calibration data, our approach is data-free and generalizes well across diverse tasks. Our experiments with Llama3 70B, which is particularly challenging, show zero-shot accuracy retention at 4- and 3-bit compression to be on par with or better than state-of-the-art methods, while maintaining performance comparable to FP16 baselines. Additionally, FPGA-based tests demonstrate that 4-bit SeedLM, as model size increases, approaches a 4x speed-up over an FP16 Llama 2/3 baseline. Rasoul Shafipour, David Harrison, Maxwell Horton, Jeffrey Marker, Houman Bedayat, Sachin Mehta, Mohammad Rastegari, Mahyar Najibi, Saman Naderiparizi |
ICLR | 6 |
| 2025 | CtrlSynth: Controllable Image Text Synthesis for Data-Efficient Multimodal LearningabstractPretraining robust vision or multimodal foundation models (e.g., CLIP) relies on large-scale datasets that may be noisy, potentially misaligned, and have long-tail distributions. Previous works have shown promising results in augmenting datasets by generating synthetic samples. However, they only support domain-specific ad hoc use cases (e.g., either image or text only, but not both), and are limited in data diversity due to a lack of fine-grained control over the synthesis process. In this paper, we design a controllable image-text synthesis pipeline, CtrlSynth, for data-efficient and robust multimodal learning. The key idea is to decompose the visual semantics of an image into basic elements, apply user-specified control policies (e.g., remove, add, or replace operations), and recompose them to synthesize images or texts. The decompose and recompose feature in CtrlSynth allows users to control data synthesis in a fine-grained manner by defining customized control policies to manipulate the basic elements. CtrlSynth leverages the capabilities of pretrained foundation models such as large language models or diffusion models to reason and recompose basic elements such that synthetic samples are natural and composed in diverse ways. CtrlSynth is a closed-loop, training-free, and modular framework, making it easy to support different pretrained models. With extensive experiments on 31 datasets spanning different vision and vision-language tasks, we show that CtrlSynth substantially improves zero-shot classification, image-text retrieval, and compositional reasoning performance of CLIP models. Mahyar Najibi, Sachin Mehta |
ICML | 3 |
| 2024 | TiC-CLIP: Continual Training of CLIP ModelsabstractKeeping large foundation models up to date on latest data is inherently expensive. To avoid the prohibitive costs of constantly retraining, it is imperative to continually train these models. This problem is exacerbated by the lack of any large scale continual learning benchmarks or baselines. We introduce the first set of web-scale Time-Continual (TiC) benchmarks for training vision-language models: TiC-DataComp, TiC-YFCC, and TiC-Redcaps. TiC-DataComp, our largest dataset, contains over 12.7B timestamped image-text pairs spanning 9 years (2014-2022). We first use our benchmarks to curate various dynamic evaluations to measure temporal robustness of existing models. We show OpenAI's CLIP (trained on data up to 2020) loses $\approx 8\%$ zero-shot accuracy on our curated retrieval task from 2021-2022 compared with more recently trained models in OpenCLIP repository. We then study how to efficiently train models on time-continuous data. We demonstrate that a simple rehearsal-based approach that continues training from the last checkpoint and replays old data reduces compute by $2.5\times$ when compared to the standard practice of retraining from scratch. Code is available at https://github.com/apple/ml-tic-clip. Mehrdad Farajtabar, Hadi Pouransari, Raviteja Vemulapalli, Sachin Mehta, Oncel Tuzel, Vaishaal Shankar, Fartash Faghri |
ICLR | 5 |
| 2024 | ReLU Strikes Back: Exploiting Activation Sparsity in Large Language ModelsabstractLarge Language Models (LLMs) with billions of parameters have drastically transformed AI applications. However, their demanding computation during inference has raised significant challenges for deployment on resource-constrained devices. Despite recent trends favoring alternative activation functions such as GELU or SiLU, known for increased computation, this study strongly advocates for reinstating ReLU activation in LLMs. We demonstrate that using the ReLU activation function has a negligible impact on convergence and performance while significantly reducing computation and weight transfer. This reduction is particularly valuable during the memory-bound inference step, where efficiency is paramount. Exploring sparsity patterns in ReLU-based LLMs, we unveil the reutilization of activated neurons for generating new tokens and leveraging these insights, we propose practical strategies to substantially reduce LLM inference computation up to three times, using ReLU activations with minimal performance trade-offs. Iman Mirzadeh, Keivan Alizadeh-Vahid, Sachin Mehta, Carlo C. del Mundo, Oncel Tuzel, Golnoosh Samei, Mohammad Rastegari, Mehrdad Farajtabar |
ICLR | 3 |
| 2024 | Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific ModelsabstractVision Foundation Models (VFMs) pretrained on massive datasets exhibit impressive performance on various downstream tasks, especially with limited labeled target data. However, due to their high inference compute cost, these models cannot be deployed for many real-world applications. Motivated by this, we ask the following important question, "How can we leverage the knowledge from a large VFM to train a small task-specific model for a new target task with limited labeled training data?", and propose a simple task-oriented knowledge transfer approach as a highly effective solution to this problem. Our experimental results on five target tasks show that the proposed approach outperforms task-agnostic VFM distillation, web-scale CLIP pretraining, supervised ImageNet pretraining, and self-supervised DINO pretraining by up to 11.6%, 22.1%, 13.7%, and 29.8%, respectively. Furthermore, the proposed approach also demonstrates up to 9x, 4x and 15x reduction in pretraining compute cost when compared to task-agnostic VFM distillation, ImageNet pretraining and DINO pretraining, respectively, while outperforming them. We also show that the dataset used for transferring knowledge has a significant effect on the final target task performance, and introduce a retrieval-augmented knowledge transfer strategy that uses web-scale image retrieval to curate effective transfer sets. Raviteja Vemulapalli, Hadi Pouransari, Fartash Faghri, Sachin Mehta, Mehrdad Farajtabar, Mohammad Rastegari, Oncel Tuzel |
ICML | 4 |
| 2023 | Reinforce Data, Multiply Impact: Improved Model Accuracy and Robustness with Dataset ReinforcementabstractWe propose Dataset Reinforcement, a strategy to improve a dataset once such that the accuracy of any model architecture trained on the reinforced dataset is improved at no additional training cost for users. We propose a Dataset Reinforcement strategy based on data augmentation and knowledge distillation. Our generic strategy is designed based on extensive analysis across CNN- and transformer-based models and performing large-scale study of distillation with state-of-the-art models with various data augmentations. We create a reinforced version of the ImageNet training dataset, called ImageNet+, as well as reinforced datasets CIFAR-100+, Flowers-102+, and Food-101+. Models trained with ImageNet+are more accurate, robust, and calibrated, and transfer well to downstream tasks (e.g., segmentation and detection). As an example, the accuracy of ResNet-50 improves by 1.7% on the ImageNet validation set, 3.5% on ImageNetV2, and 10.0% on ImageNet-R. Expected Calibration Error (ECE) on the ImageNet validation set is also reduced by 9.9%. Using this backbone with Mask-RCNN for object detection on MS-COCO, the mean average precision improves by 0.8%. We reach similar gains for MobileNets, ViTs, and Swin-Transformers. For MobileNetV3 and Swin-Tiny, we observe significant improvements on ImageNet-R/A/C of up to 20% improved robustness. Models pretrained on ImageNet+and fine-tuned on CIFAR-100+, Flowers-102+, and Food-101+, reach up to 3.4% improved accuracy. The code, datasets, and pretrained models are available at https://github.com/apple/ml-dr. Fartash Faghri, Hadi Pouransari, Sachin Mehta, Mehrdad Farajtabar, Ali Farhadi, Mohammad Rastegari, Oncel Tuzel |
ICCV | 3 |
| 2022 | SPIN: An Empirical Evaluation on Sharing Parameters of Isotropic Networks
Chien-Yu Lin, Anish Prabhu, Thomas Merth, Sachin Mehta, Anurag Ranjan, Maxwell Horton, Mohammad Rastegari |
ECCV (11) | 4 |
| 2022 | Calibration Error Prediction: Ensuring High-Quality Mobile Eye-TrackingabstractGaze calibration is common in traditional infrared oculographic eye tracking. However, it is not well studied in visible-light mobile/remote eye tracking. We developed a lightweight real-time gaze error estimator and analyzed calibration errors from two perspectives: facial feature-based and Monte Carlo-based. Both methods correlated with gaze estimation errors, but the Monte Carlo method associated more strongly. Facial feature associations with gaze error were interpretable, relating movements of the face to the visibility of the eye. We highlight the degradation of gaze estimation quality in a sample of children with autism spectrum disorder (as compared to typical adults), and note that calibration methods may improve Euclidean error by 10%. Beibin Li, James C. Snider, Quan Wang 0003, Sachin Mehta, Claire E. Foster, Erin Barney, Linda G. Shapiro, Pamela Ventola, Frédérick Shic |
ETRA | 4 |
| 2022 | MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer
Sachin Mehta, Mohammad Rastegari |
ICLR | 1 |
| 2022 | CVNets: High Performance Library for Computer VisionabstractWe introduce CVNets, a high-performance open-source library for training deep neural networks for visual recognition tasks, including classification, detection, and segmentation. CVNets supports image and video understanding tools, including data loading, data transformations, novel data sampling methods, and implementations of several standard networks with similar or better performance than previous studies. Our source code is available at: https://github.com/apple/ml-cvnets. Sachin Mehta, Farzad Abdolhosseini, Mohammad Rastegari |
ACM Multimedia | 1 |
| 2022 | End-to-End diagnosis of breast biopsy images with transformers
Sachin Mehta, Ximing Lu, Donald L. Weaver, Hannaneh Hajishirzi, Joann G. Elmore, Linda G. Shapiro |
Medical Image Anal. | 1 |
| 2022 | DiCENet: Dimension-Wise Convolutions for Efficient NetworksabstractWe introduce a novel and generic convolutional unit, DiCE unit, that is built using dimension-wise convolutions and dimension-wise fusion. The dimension-wise convolutions apply light-weight convolutional filtering across each dimension of the input tensor while dimension-wise fusion efficiently combines these dimension-wise representations; allowing the DiCE unit to efficiently encode spatial and channel-wise information contained in the input tensor. The DiCE unit is simple and can be seamlessly integrated with any architecture to improve its efficiency and performance. Compared to depth-wise separable convolutions, the DiCE unit shows significant improvements across different architectures. When DiCE units are stacked to build the DiCENet model, we observe significant improvements over state-of-the-art models across various computer vision tasks including image classification, object detection, and semantic segmentation. On the ImageNet dataset, the DiCENet delivers 2-4 percent higher accuracy than state-of-the-art manually designed models (e.g., MobileNetv2 and ShuffleNetv2). Also, DiCENet generalizes better to tasks (e.g., object detection) that are often used in resource-constrained devices in comparison to state-of-the-art separable convolution-based efficient networks, including neural search-based methods (e.g., MobileNetv3 and MixNet). Sachin Mehta, Hannaneh Hajishirzi, Mohammad Rastegari |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Collecting Sidewalk Network Data at Scale for Accessible Pedestrian TravelabstractSidewalks are central to an accessible transportation network, as they connect all other transportation modes. The street-side environment, especially the location and connectivity of the sidewalks, has not been widely integrated into information systems used to report accessibility and walkability in wayfinding applications. Typical sidewalk mapping methods rely on surveyor collections, which are non-standardized, laborious, costly, difficult to maintain, and do not scale well. In this work, we introduce a working proof-of-concept system for automated mapping of sidewalk networks on portable computing devices. Our system utilizes efficient neural networks, image sensing, GPS, and compact hardware to perform sidewalk mapping on portable devices. We discuss future opportunities for cities and transportation agencies to advance their knowledge of the transportation network they own and manage in order to improve accessibility for all travelers. Yuxiang Zhang 0010, Sachin Mehta, Anat Caspi |
ASSETS | 2 |
| 2021 | Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and TextabstractChristopher Clark, Jordi Salvador, Dustin Schwenk, Derrick Bonafilia, Mark Yatskar, Eric Kolve, Alvaro Herrasti, Jonghyun Choi, Sachin Mehta, Sam Skjonsberg, Carissa Schoenick, Aaron Sarnat, Hannaneh Hajishirzi, Aniruddha Kembhavi, Oren Etzioni, Ali Farhadi. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Jordi Salvador, Dustin Schwenk, Derrick Bonafilia, Mark Yatskar, Eric Kolve, Alvaro Herrasti, Sachin Mehta, Sam Skjonsberg, Carissa Schoenick, Aaron Sarnat, Hannaneh Hajishirzi, Aniruddha Kembhavi, Oren Etzioni, Ali Farhadi |
EMNLP (1) | 9 |
| 2021 | DeLighT: Deep and Light-weight Transformer
Sachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer 0001, Luke Zettlemoyer, Hannaneh Hajishirzi |
ICLR | 1 |
| 2021 | EVRNet: Efficient Video Restoration on Edge DevicesabstractIn video transmission applications, video signals are transmitted over lossy channels, resulting in low-quality received signals. To re- store videos on recipient edge devices in real-time, we introduce an efficient video restoration network, EVRNet. EVRNet efficiently allocates parameters inside the network using alignment, differential, and fusion modules. With extensive experiments on different video restoration tasks (deblocking, denoising, and super-resolution), we demonstrate that EVRNet delivers competitive performance to existing methods with significantly fewer parameters and MACs. For example, EVRNet has 260× fewer parameters and 958× fewer MACs than enhanced deformable convolution-based video restoration net- work (EDVR) for 4× video super-resolution while its SSIM score is 0.018 less than EDVR. We also evaluated the performance of EVR-Net under multiple distortions on unseen dataset to demonstrate its ability in modeling variable-length sequences under both camera and object motion. Sachin Mehta, Amit Kumar 0013, Fitsum A. Reda, Varun Nasery, Vikram Mulukutla, Vikas Chandra |
ACM Multimedia | 1 |
| 2020 | DeFINE: Deep Factorized Input Token Embeddings for Neural Sequence Modeling
Sachin Mehta, Rik Koncel-Kedziorski, Mohammad Rastegari, Hannaneh Hajishirzi |
ICLR | 1 |
| 2020 | Classifying Breast Histopathology Images with a Ductal Instance-Oriented PipelineabstractIn this study, we propose the Ductal Instance-Oriented Pipeline (DIOP) that contains a duct-level instance segmentation model, a tissue-level semantic segmentation model, and three-levels of features for diagnostic classification. Based on recent advancements in instance segmentation and the Mask RCNN model, our duct-level segmenter tries to identify each ductal individual inside a microscopic image; then, it extracts tissue-level information from the identified ductal instances. Leveraging three levels of information obtained from these ductal instances and also the histopathology image, the proposed DIOP outperforms previous approaches (both feature-based and CNN-based) in all diagnostic tasks; for the four-way classification task, the DIOP achieves comparable performance to general pathologists in this unique dataset. The proposed DIOP only takes a few seconds to run in the inference time, which could be used interactively on most modern computers. More clinical explorations are needed to study the robustness and generalizability of this system in the future. Beibin Li, Ezgi Mercan, Sachin Mehta, Stevan Knezevich, Corey W. Arnold, Donald L. Weaver, Joann G. Elmore, Linda G. Shapiro |
ICPR | 3 |
| 2020 | Leveraging Unlabeled Data for Glioma Molecular Subtype and Survival PredictionabstractIn this paper, we address two long-standing radio-genomic challenges in glioma subtype and survival prediction: (1) how to leverage large amounts of unlabeled magnetic resonance (MR) imaging data and (2) how to unite MR data and genomic data. We propose a novel application of multi-task learning (MTL) that leverages unlabeled MR data by jointly learning an auxiliary tumor segmentation task with glioma subtype prediction and that can learn from patients with and without genomic data. We analyze multi-parametric MR data from 542 patients in the combined training, validation, and testing sets of the 2018 Multimodal Brain Tumor Segmentation Challenge and somatic copy number alteration (SCNA) data from 1090 patients in The Cancer Genome Atlas' (TCGA) lower-grade glioma and glioblastoma projects. Our MTL model significantly outperforms comparable classification models trained only on labeled MR data for both IDH1/2 mutation and 1p/19q co-deletion subtype prediction tasks. We also show that embeddings produced by our MTL models improve survival predictions beyond MR or SCNA on their own. Our code is available at https://github.com/nknuecht/glioma_mtl. Nicholas Nuechterlein, Beibin Li, Mehmet Saygin Seyfioglu, Sachin Mehta, Patrick J. Cimino, Linda G. Shapiro |
ICPR | 4 |
| 2019 | ESPNetv2: A Light-Weight, Power Efficient, and General Purpose Convolutional Neural NetworkabstractWe introduce a light-weight, power efficient, and general purpose convolutional neural network, ESPNetv2, for modeling visual and sequential data. Our network uses group point-wise and depth-wise dilated separable convolutions to learn representations from a large effective receptive field with fewer FLOPs and parameters. The performance of our network is evaluated on four different tasks: (1) object classification, (2) semantic segmentation, (3) object detection, and (4) language modeling. Experiments on these tasks, including image classification on the ImageNet and language modeling on the PenTree bank dataset, demonstrate the superior performance of our method over the state-of-the-art methods. Our network outperforms ESPNet by 4-5% and has 2-4x fewer FLOPs on the PASCAL VOC and the Cityscapes dataset. Compared to YOLOv2 on the MS-COCO object detection, ESPNetv2 delivers 4.4% higher accuracy with 6x fewer FLOPs. Our experiments show that ESPNetv2 is much more power efficient than existing state-of-the-art efficient methods including ShuffleNets and MobileNets. Our code is open-source and available at https://github.com/sacmehta/ESPNetv2. Sachin Mehta, Mohammad Rastegari, Linda G. Shapiro, Hannaneh Hajishirzi |
CVPR | 1 |
| 2019 | A Facial Affect Analysis System for Autism Spectrum DisorderabstractIn this paper, we introduce an end-to-end machine learning-based system for classifying autism spectrum disorder (ASD) using facial attributes such as expressions, action units, arousal, and valence. Our system classifies ASD using representations of different facial attributes from convolutional neural networks, which are trained on images in the wild. Our experimental results show that different facial attributes used in our system are statistically significant and improve sensitivity, specificity, and F1 score of ASD classification by a large margin. In particular, the addition of different facial attributes improves the performance of ASD classification by about 7% which achieves a F1 score of 76%. Beibin Li, Sachin Mehta, Deepali Aneja, Claire E. Foster, Pamela Ventola, Frédérick Shic, Linda G. Shapiro |
ICIP | 2 |
| 2018 | ESPNet: Efficient Spatial Pyramid of Dilated Convolutions for Semantic Segmentation
Sachin Mehta, Mohammad Rastegari, Anat Caspi, Linda G. Shapiro, Hannaneh Hajishirzi |
ECCV (10) | 1 |
| 2018 | Pyramidal Recurrent Unit for Language ModelingabstractLSTMs are powerful tools for modeling contextual information, as evidenced by their success at the task of language modeling.However, modeling contexts in very high dimensional space can lead to poor generalizability.We introduce the Pyramidal Recurrent Unit (PRU), which enables learning representations in high dimensional space with more generalization power and fewer parameters.PRUs replace the linear transformation in LSTMs with more sophisticated interactions including pyramidal and grouped linear transformations.This architecture gives strong results on wordlevel language modeling while reducing the number of parameters significantly.In particular, PRU improves the perplexity of a recent state-of-the-art language model Merity et al. (2018) by up to 1.3 points while learning 15-20% fewer parameters.For similar number of model parameters, PRU outperforms all previous RNN models that exploit different gating mechanisms and transformations.We provide a detailed examination of the PRU and its behavior on the language modeling tasks. Sachin Mehta, Rik Koncel-Kedziorski, Mohammad Rastegari, Hannaneh Hajishirzi |
EMNLP | 1 |
| 2018 | Automated Diagnosis of Breast Cancer and Pre-invasive Lesions on Digital Whole Slide Images
Ezgi Mercan, Sachin Mehta, Jamen Bartlett, Donald L. Weaver, Joann G. Elmore, Linda G. Shapiro |
ICPRAM | 2 |
| 2018 | Y-Net: Joint Segmentation and Classification for Diagnosis of Breast Biopsy Images
Sachin Mehta, Ezgi Mercan, Jamen Bartlett, Donald L. Weaver, Joann G. Elmore, Linda G. Shapiro |
MICCAI (2) | 1 |
| 2018 | DeepSolarEye: Power Loss Prediction and Weakly Supervised Soiling Localization via Fully Convolutional Networks for Solar PanelsabstractThe impact of soiling on solar panels is an important and well-studied problem in renewable energy sector. In this paper, we present the first convolutional neural network (CNN) based approach for solar panel soiling and defect analysis. Our approach takes an RGB image of solar panel and environmental factors as inputs to predict power loss, soiling localization, and soiling type. In computer vision, localization is a complex task which typically requires manually labeled training data such as bounding boxes or segmentation masks. Our proposed approach consists of specialized four stages which completely avoids localization ground truth and only needs panel images with power loss labels for training. The region of impact area obtained from the predicted localization masks are classified into soiling types using the webly supervised learning. For improving localization capabilities of CNNs, we introduce a novel bi-directional input-aware fusion (BiDIAF) block that reinforces the input at different levels of CNN to learn input-specific feature maps. Our empirical study shows that BiDIAF improves the power loss prediction accuracy by about 3% and localization accuracy by about 4%. Our end-to-end model yields further improvement of about 24% on localization when learned in a weakly supervised manner. Our approach is generalizable and showed promising results on web crawled solar panel images. Our system has a frame rate of 22 fps (including all steps) on a NVIDIA TitanX GPU. Additionally, we collected first of it's kind dataset for solar panel image analysis consisting 45,000+ images. Sachin Mehta, Amar Prakash Azad, Saneem A. Chemmengath, Vikas C. Raykar, Shivkumar Kalyanaraman |
WACV | 1 |
| 2018 | Learning to Segment Breast Biopsy Whole Slide ImagesabstractWe trained and applied an encoder-decoder model to semantically segment breast biopsy images into biologically meaningful tissue labels. Since conventional encoderdecoder networks cannot be applied directly on large biopsy images and the different sized structures in biopsies present novel challenges, we propose four modifications: (1) an input-aware encoding block to compensate for information loss, (2) a new dense connection pattern between encoder and decoder, (3) dense and sparse decoders to combine multi-level features, (4) a multi-resolution network that fuses the results of encoder-decoders run on different resolutions. Our model outperforms a feature-based approach and conventional encoder-decoders from the literature. We use semantic segmentations produced with our model in an automated diagnosis task and obtain higher accuracies than a baseline approach that employs an SVM for featurebased segmentation, both using the same segmentationbased diagnostic features. Sachin Mehta, Ezgi Mercan, Jamen Bartlett, Donald L. Weaver, Joann G. Elmore, Linda G. Shapiro |
WACV | 1 |
| 2016 | Region graph based method for multi-object detection and tracking using depth camerasabstractIn this paper, we propose a multi-object detection and tracking method using depth cameras. Depth maps are very noisy and obscure in object detection. We first propose a region-based method to suppress high magnitude noise which cannot be filtered using spatial filters. Second, the proposed method detect Region of Interests by temporal learning which are then tracked using weighted graph-based approach. We demonstrate the performance of the proposed method on standard depth camera datasets with and without object occlusions. Experimental results show that the proposed, method is able to suppress high magnitude noise in depth maps and detect/track the objects (with and without occlusion). Sachin Mehta, B. Prabhakaran 0001 |
WACV | 1 |
| 2016 | Scene-based fingerprinting method for traitor tracing
Sachin Mehta, Rajarathnam Nallusamy, B. Prabhakaran 0001 |
Multim. Syst. | 1 |
| 2014 | 3D content fingerprintingabstractFingerprint is a set of features that uniquely characterizes a video. The aim of content fingerprinting is to determine the duplicate videos over the Internet. In this paper, a method for content fingerprinting of Depth-Image-Based-Rendering (DIBR) 3D videos is proposed. The proposed method is two pronged approach: (i) histogram based global fingerprint and (ii) keypoint based local fingerprint. Though global fingerprint is fast and robust towards DIBR 3D pre-processing, it is not robust against severe distortions such as change in brightness. To make the proposed method robust against such severe distortions, we have complemented the global fingerprint with widely used keypoint based local fingerprint. Experimental results show that the proposed two pass method as robust as keypoint based local fingerprinting method. Additionally, the proposed method improves the video matching time of keypoint based local fingerprint method by 60% to 100%. Sachin Mehta, B. Prabhakaran 0001 |
ICIP | 1 |
| 2012 | Tampering resistant self recoverable watermarking method using error correction codesabstractRecent advances in digital technologies have made alteration of digital multimedia content simple and unnoticeable by human beings. Although digital multimedia content offers numerous advantages such as easy to store, easy to transmit, etc., it has made illegal replication/distribution of digital content easier and highlights the need for digital rights management (DRM). Digital watermarking has emerged as a solution for the copyright protection of digital multimedia content. Digital watermarking hides the owner’s information inside the digital multimedia content which can be extracted in future to prove the owner’s copyright over the content. In this paper, self-recoverable digital watermarking method is presented to protect the digital images from tampering attacks. Experimental results show that the proposed method is robust against attacks such as cropping, blurring, noising, etc. Sachin Mehta, Vijayaraghavan Varadharajan, Rajarathnam Nallusamy |
Int. J. Inf. Comput. Secur. | 1 |