EDBT 2026 Demo / reviewers in the wild / expert
Gaurav Mittal
dblp:98/6246
· DBLP profile ↗
38ranked-venue papers
14as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 11 since 2021Systems, architecture and hardware · 12 · 4 first-author · 2 since 2021Theory of computation · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Convergence analysis of a frozen stochastic iteratively regularized Gauss-Newton method
Gaurav Mittal |
J. Complex. | 1 |
| 2026 | A new and efficient image authenticated encryption scheme based on modified two-square cipher and adaptive Arnold map
Kanchan Bhagtani, Gaurav Mittal |
Multim. Tools Appl. | 2 |
| 2026 | On the Convergence of the Iterative Regularization Method Assisted by the Graph Laplacian with Early StoppingabstractAbstract. We present a data-assisted iterative regularization method for solving ill-posed inverse problems. The proposed approach, termed IRMGL+[Formula: see text], integrates classical iterative techniques with a data-driven regularization term realized through an iteratively updated graph Laplacian. Our method commences by computing a preliminary solution using any suitable reconstruction method, which then serves as the basis for constructing the initial graph Laplacian. The solution is subsequently refined through an iterative process, where the graph Laplacian is simultaneously recalibrated at each step to effectively capture the evolving structure of the solution. A key innovation of this work lies in the formulation of this iterative scheme and the rigorous justification of the classical discrepancy principle as a reliable early stopping criterion specifically tailored to the proposed method. Under standard assumptions, we establish stability and convergence results for the scheme when the discrepancy principle is applied. Furthermore, we demonstrate the robustness and effectiveness of our method through numerical experiments utilizing four distinct initial reconstructors [Formula: see text]: the adjoint operator (Adj), filtered back projection, total variation denoising, and standard Tikhonov regularization. It is observed that IRMGL + Adj demonstrates a distinct advantage over the other initializers, producing a robust and stable approximate solution directly from a basic initial reconstruction. Harshit Bajpai, Gaurav Mittal, Ankik Kumar Giri |
SIAM J. Imaging Sci. | 2 |
| 2026 | Towards quantum-resistant and anonymous authenticated key establishment in multi-specialty IoT healthcare environments
Gaurav Mittal, Arvind Yadav |
J. Supercomput. | 2 |
| 2025 | DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long VideosabstractLong Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing video into clips and processing each clip via a full-scale expert encoder is challenging to scale due to prohibitive computational costs of processing a large number of clips in long videos. To address this issue, we introduce DeCafNet, an approach employing "delegate-and-conquer" strategy to achieve computation efficiency without sacrificing grounding performance. DeCafNet introduces a sidekick encoder that performs dense feature extraction over all video clips in a resource-efficient manner, while generating a saliency map to identify the most relevant clips for full processing by the expert encoder. To effectively leverage features from sidekick and expert encoders that exist at different temporal resolutions, we introduce DeCaf-Grounder, which unifies and refines them via query-aware temporal aggregation and multi-scale temporal refinement for accurate grounding. Experiments on two LTVG benchmark datasets demonstrate that DeCafNet reduces computation by up to 47% while still outperforming existing methods, establishing a new state-of-the-art for LTVG in terms of both efficiency and performance. Zijia Lu, A S. M. Iftekhar, Gaurav Mittal, Tianjian Meng, Xiawei Wang, Rohith Kukkala, Ehsan Elhamifar |
CVPR | 3 |
| 2025 | Hummingbird: High Fidelity Image Generation via Multimodal Context AlignmentabstractWhile diffusion models are powerful in generating high-quality, diverse synthetic data for object-centric tasks, existing methods struggle with scene-aware tasks such as Visual Question Answering (VQA) and Human-Object Interaction (HOI) Reasoning, where it is critical to preserve scene attributes in generated images consistent with a multimodal context, i.e. a reference image with accompanying text guidance query. To address this, we introduce **Hummingbird**, the first diffusion-based image generator which, given a multimodal context, generates highly diverse images w.r.t. the reference image while ensuring high fidelity by accurately preserving scene attributes, such as object interactions and spatial relationships from the text guidance. Hummingbird employs a novel Multimodal Context Evaluator that simultaneously optimizes our formulated Global Semantic and Fine-grained Consistency Rewards to ensure generated images preserve the scene attributes of reference images in relation to the text guidance while maintaining diversity. As the first model to address the task of maintaining both diversity and fidelity given a multimodal context, we introduce a new benchmark formulation incorporating MME Perception and Bongard HOI datasets. Benchmark experiments show Hummingbird outperforms all existing methods by achieving superior fidelity while maintaining diversity, validating Hummingbird's potential as a robust multimodal context-aligned image generator in complex visual tasks. Project page: https://roar-ai.github.io/hummingbird Minh-Quan Le, Gaurav Mittal, Tianjian Meng, A S. M. Iftekhar, Vishwas Suryanarayanan, Barun Patra, Dimitris Samaras |
ICLR | 2 |
| 2025 | LoSA: Long-Short-Range Adapter for Scaling End-to-End Temporal Action LocalizationabstractTemporal Action Localization (TAL) involves localizing and classifying action snippets in an untrimmed video. The emergence of large video foundation models has led RGB-only video backbones to outperform previous methods needing both RGB and optical flow modalities. Leveraging these large models is often limited to training only the TAL head due to the prohibitively large GPU memory required to adapt the video backbone for TAL. To overcome this limitation, we introduce LoSA, the first memory-and-parameter-efficient backbone adapter designed specifically for TAL to handle untrimmed videos. LoSA specializes for TAL by introducing Long-Short-range Adapters that adapt the intermediate layers of the video backbone over different temporal ranges. These adapters run parallel to the video backbone to significantly reduce memory footprint. LoSA also includes Long-Short-range Gated Fusion that strategically combines the output of these adapters from the video backbone layers to enhance the video features provided to the TAL head. Experiments show that LoSA significantly outperforms all existing methods on standard TAL benchmarks, THUMOS-14 and ActivityNet-v1.3, by scaling end-to-end backbone adaptation to billion-parameter-plus models like VideoMAEv2 (ViT-g) and leveraging them beyond head-only transfer learning. Gaurav Mittal, Ahmed Magooda, Ye Yu 0003, Graham W. Taylor |
WACV | 2 |
| 2025 | Three Party Post Quantum Secure Lattice Based Construction of Authenticated Key Establishment Protocol for Mobile CommunicationabstractABSTRACT A three‐party post‐quantum key agreement protocol involves server with two communicating parties securely agreeing on a shared secret key in a way that is resistant to quantum attacks. Once the shared secret key is shared using authenticated key agreement protocol, then user (A), and user (B) can use it for securing communication channel using symmetric‐key encryption AES‐256 algorithm. Although there are few third‐party post‐quantum authenticated and key agreement schemes exist, but the recent studies in this paper illustrates that they are not satisfying properties like unlinkability, anonymity, perfect forward secrecy, and signal leakage attacks. Therefore, the proposed protocol ensures anonymity, unlinkablity, perfect forward secrecy, and resistant against signal leakage attacks. The proposed protocol uses different random numbers for each of sessions and ensures freshness of the session key to maintain forward secrecy. In this protocol, the user (A) only communicates with server, and establish an authenticated session key with user (B) which avoids server overheads. The use of ring learning with errors (RLWE) instead of the simpler learning with errors (LWE) is primarily motivated by the need for efficiency, compactness, and scalability in cryptographic applications. A comparative study, including both performance and security assessments, demonstrates that the proposed design is more secure and efficient. Gaurav Mittal, Arvind Yadav |
Concurr. Comput. Pract. Exp. | 2 |
| 2025 | Convergence analysis of iteratively regularized Landweber iteration with uniformly convex constraints in Banach spaces
Gaurav Mittal, Harshit Bajpai, Ankik Kumar Giri |
J. Complex. | 1 |
| 2024 | Rethinking Multimodal Content Moderation from an Asymmetric Angle with Mixed-modalityabstractThere is a rapidly growing need for multimodal content moderation (CM) as more and more content on social media is multimodal in nature. Existing unimodal CM systems may fail to catch harmful content that crosses modalities (e.g., memes or videos), which may lead to severe consequences. In this paper, we present a novel CM model, Asymmetric Mixed-Modal Moderation (AM3), to target multimodal and unimodal CM tasks. Specifically, to address the asymmetry in semantics between vision and language, AM3 has a novel asymmetric fusion architecture that is designed to not only fuse the common knowledge in both modalities but also to exploit the unique information in each modality. Unlike previous works that focus on representing the two modalities into a similar feature space while overlooking the intrinsic difference between the information conveyed in multimodality and in unimodality (asymmetry in modalities), we propose a novel cross-modality contrastive loss to learn the unique knowledge that only appears in multimodality. This is critical as some harmful intent may only be conveyed through the intersection of both modalities. With extensive experiments, we show that AM3 outperforms all existing state-of-the-art methods on both multimodal and unimodal CM benchmarks. Jialin Yuan, Ye Yu 0003, Gaurav Mittal, Sandra Sajeev |
WACV | 3 |
| 2024 | HyperSTAR: Task-Aware Hyperparameter Recommendation for Training and Compression
Chang Liu 0022, Gaurav Mittal, Nikolaos Karianakis, Victor Fragoso, Ye Yu 0003, Yun Fu 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | A novel and efficient undeniable signature scheme based on group ring
Gaurav Mittal, Shubham Mittal 0004 |
Soft Comput. | 1 |
| 2023 | Rule By Example: Harnessing Logical Rules for Explainable Hate Speech DetectionabstractChristopher Clarke, Matthew Hall, Gaurav Mittal, Ye Yu, Sandra Sajeev, Jason Mars, Mei Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Christopher Clarke, Gaurav Mittal, Ye Yu 0003, Sandra Sajeev, Jason Mars |
ACL (1) | 3 |
| 2023 | PivoTAL: Prior-Driven Supervision for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised Temporal Action Localization (WTAL) attempts to localize the actions in untrimmed videos using only video-level supervision. Most recent works approach WTAL from a localization-by-classification perspective where these methods try to classify each video frame followed by a manually-designed post-processing pipeline to aggregate these per-frame action predictions into action snippets. Due to this perspective, the model lacks any explicit understanding of action boundaries and tends to focus only on the most discriminative parts of the video resulting in incomplete action localization. To address this, we present PivoTAL, Prior-driven Supervision for Weakly-supervised Temporal Action Localization, to approach WTAL from a localization-by-localization perspective by learning to localize the action snippets directly. To this end, PivoTAL leverages the underlying spatio-temporal regularities in videos in the form of action-specific scene prior, action snippet generation prior, and learnable Gaussian prior to supervise the localization-based training. PivoTAL shows significant improvement (of at least 3% avg mAP) over all existing methods on the benchmark datasets, THUMOS-14 and ActivitNet-v1.3. Mamshad Nayeem Rizve, Gaurav Mittal, Ye Yu 0003, Sandra Sajeev, Mubarak Shah |
CVPR | 2 |
| 2023 | ProTéGé: Untrimmed Pretraining for Video Temporal Grounding by Video Temporal GroundingabstractVideo temporal grounding (VTG) is the task of localizing a given natural language text query in an arbitrarily long untrimmed video. While the task involves untrimmed videos, all existing VTG methods leverage features from video backbones pretrained on trimmed videos. This is largely due to the lack of large-scale well-annotated VTG dataset to perform pretraining. As a result, the pretrained features lack a notion of temporal boundaries leading to the video-text alignment being less distinguishable between correct and incorrect locations. We present ProTéGé as the first method to perform VTG-based untrimmed pretraining to bridge the gap between trimmed pretrained backbones and downstream VTG tasks. ProTéGé reconfigures the HowTo100M dataset, with noisily correlated video-text pairs, into a VTG dataset and introduces a novel Video-Text Similarity-based Grounding Module and a pretraining objective to make pretraining robust to noise in HowTo100M. Extensive experiments on multiple datasets across downstream tasks with all variations of supervision validate that pretrained features from ProTéGé can significantly outperform features from trimmed pretrained backbones on VTG. Gaurav Mittal, Sandra Sajeev, Ye Yu 0003, Vishnu Naresh Boddeti |
CVPR | 2 |
| 2023 | Convergence analysis of an optimally accurate frozen multi-level projected steepest descent iteration for solving inverse problems
Gaurav Mittal, Ankik Kumar Giri |
J. Complex. | 1 |
| 2022 | GateHUB: Gated History Unit with Background Suppression for Online Action DetectionabstractOnline action detection is the task of predicting the action as soon as it happens in a streaming video. A major challenge is that the model does not have access to the future and has to solely rely on the history, i.e., the frames observed so far, to make predictions. It is therefore important to accentuate parts of the history that are more informative to the prediction of the current frame. We present GateHUB, Gated History Unit with Background Suppression, that comprises a novel position-guided gated cross attention mechanism to enhance or suppress parts of the history as per how informative they are for current frame prediction. GateHUB further proposes Future-augmented History (FaH) to make history features more informative by using subsequently observed frames when available. In a single unified framework, GateHUB integrates the transformer's ability of long-range temporal modeling and the recurrent model's capacity to selectively encode relevant information. GateHUB also introduces a background suppression objective to further mitigate false positive background frames that closely resemble the action frames. Extensive validation on three benchmark datasets, THUMOS, TVSeries, and HDD, demonstrates that GateHUB significantly outperforms all existing methods and is also more efficient than the existing best work. Furthermore, a flow free version of GateHUB is able to achieve higher or close accuracy at 2.8× higher frame rate compared to all existing methods that require both RGB and optical flow information for prediction. Junwen Chen 0001, Gaurav Mittal, Ye Yu 0003, Yu Kong 0001 |
CVPR | 2 |
| 2022 | BATMAN: Bilateral Attention Transformer in Motion-Appearance Neighboring Space for Video Object Segmentation
Ye Yu 0003, Jialing Yuan, Gaurav Mittal, Fuxin Li |
ECCV (29) | 3 |
| 2021 | MUSE: Feature Self-Distillation with Mutual Information and Self-Information
Ye Yu 0003, Gaurav Mittal, Greg Mori |
BMVC | 3 |
| 2021 | Unsupervised Few-Shot Action Recognition via Action-Appearance Aligned Meta-AdaptationabstractWe present MetaUVFS as the first Unsupervised Meta-learning algorithm for Video Few-Shot action recognition. MetaUVFS leverages over 550K unlabeled videos to train a two-stream 2D and 3D CNN architecture via contrastive learning to capture the appearance-specific spatial and action-specific spatio-temporal video features respectively. MetaUVFS comprises a novel Action-Appearance Aligned Meta-adaptation (A3M) module that learns to focus on the action-oriented video features in relation to the appearance features via explicit few-shot episodic meta-learning over unsupervised hard-mined episodes. Our action-appearance alignment and explicit few-shot learner conditions the unsupervised training to mimic the downstream few-shot task, enabling MetaUVFS to significantly outperform all state-of-the-art unsupervised methods on few-shot benchmarks. Moreover, unlike previous few-shot action recognition methods that are supervised, MetaUVFS needs neither base-class labels nor a supervised pretrained backbone. Thus, we need to train MetaUVFS just once to perform competitively or sometimes even outperform state-of-the-art supervised methods on popular HMDB51, UCF101, and Kinetics100 few-shot datasets. Jay Patravali, Gaurav Mittal, Ye Yu 0003, Fuxin Li |
ICCV | 2 |
| 2020 | BLT: Balancing Long-Tailed Datasets with Adversarially-Perturbed Images
Jedrzej Kozerawski, Victor Fragoso, Nikolaos Karianakis, Gaurav Mittal, Matthew Turk 0001 |
ACCV (3) | 4 |
| 2020 | HyperSTAR: Task-Aware Hyperparameters for Deep NetworksabstractWhile deep neural networks excel in solving visual recognition tasks, they require significant effort to find hyperparameters that make them work optimally. Hyperparameter Optimization (HPO) approaches have automated the process of finding good hyperparameters but they do not adapt to a given task (task-agnostic), making them computationally inefficient. To reduce HPO time, we present HyperSTAR (System for Task Aware Hyperparameter Recommendation), a task-aware method to warm-start HPO for deep neural networks. HyperSTAR ranks and recommends hyperparameters by predicting their performance conditioned on a joint dataset-hyperparameter space. It learns a dataset (task) representation along with the performance predictor directly from raw images in an end-to-end fashion. The recommendations, when integrated with an existing HPO method, make it task-aware and significantly reduce the time to achieve optimal performance. We conduct extensive experiments on 10 publicly available large-scale image classification datasets over two different network architectures, validating that HyperSTAR evaluates 50% less configurations to achieve the best performance compared to existing methods. We further demonstrate that HyperSTAR makes Hyperband (HB) task-aware, achieving the optimal accuracy in just 25% of the budget required by both vanilla HB and Bayesian Optimized HB (BOHB). Gaurav Mittal, Chang Liu 0022, Nikolaos Karianakis, Victor Fragoso, Yun Fu 0001 |
CVPR | 1 |
| 2020 | Animating Face using Disentangled Audio RepresentationsabstractPrevious methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone. As we show empirically, one can easily break these systems by simply adding certain background noise to the utterance or changing its emotional tone (to for example, sad). To make talking head generation robust to such variations, we propose an explicit audio representation learning framework that disentangles audio sequences into various factors such as phonetic content, emotional tone, background noise and others. We conduct experiments to validate that when conditioned on disentangled content representation, the generated mouth movement by our model is significantly more accurate than previous approaches (without disentangled learning) in the presence of noise and emotional variations. We further demonstrate that our framework is compatible with current state-of-the-art approaches by replacing their original component to learn audio based representation with ours. To the best of our knowledge, this is the first work which improves the performance of talking head generation through a disentangled audio representation perspective, which is important for many real-world applications. Gaurav Mittal, Baoyuan Wang |
WACV | 1 |
| 2017 | Attentive Semantic Video Generation Using CaptionsabstractThis paper proposes a network architecture to perform variable length semantic video generation using captions. We adopt a new perspective towards video generation where we allow the captions to be combined with the long-term and short-term dependencies between video frames and thus generate a video in an incremental manner. Our experiments demonstrate our network architecture’s ability to distinguish between objects, actions and interactions in a video and combine them to generate videos for unseen captions. The network also exhibits the capability to perform spatio-temporal style transfer when asked to generate videos for a sequence of captions. We also show that the network’s ability to learn a latent representation allows it generate videos in an unsupervised manner and perform other tasks such as action recognition. Tanya Marwah, Gaurav Mittal, Vineeth N. Balasubramanian |
ICCV | 2 |
| 2017 | Sync-DRAW: Automatic Video Generation using Deep Recurrent Attentive ArchitecturesabstractThis paper introduces a novel approach for generating videos called Synchronized Deep Recurrent Attentive Writer (Sync-DRAW). Sync-DRAW can also perform text-to-video generation which, to the best of our knowledge, makes it the first approach of its kind. It combines a Variational Autoencoder(VAE) with a Recurrent Attention Mechanism in a novel manner to create a temporally dependent sequence of frames that are gradually formed over time. The recurrent attention mechanism in Sync-DRAW attends to each individual frame of the video in sychronization, while the VAE learns a latent distribution for the entire video at the global level. Our experiments with Bouncing MNIST, KTH and UCF-101 suggest that Sync-DRAW is efficient in learning the spatial and temporal information of the videos and generates frames with high structural integrity, and can generate videos from simple captions on these datasets. Gaurav Mittal, Tanya Marwah, Vineeth N. Balasubramanian |
ACM Multimedia | 1 |
| 2016 | SpotGarbage: smartphone app to detect garbage using deep learningabstractMaintaining a clean and hygienic civic environment is an indispensable yet formidable task, especially in developing countries. With the aim of engaging citizens to track and report on their neighborhoods, this paper presents a novel smartphone app, called SpotGarbage, which detects and coarsely segments garbage regions in a user-clicked geo-tagged image. The app utilizes the proposed deep architecture of fully convolutional networks for detecting garbage in images. The model has been trained on a newly introduced Garbage In Images (GINI) dataset, achieving a mean accuracy of 87.69%. The paper also proposes optimizations in the network architecture resulting in a reduction of 87.9% in memory usage and 96.8% in prediction time with no loss in accuracy, facilitating its usage in resource constrained smartphones. Gaurav Mittal, Kaushal B. Yagnik, Narayanan Chatapuram Krishnan |
UbiComp | 1 |
| 2014 | Symmetrical predictor structure based integrated lossy, near lossless/lossless coding of imagesabstractPrediction based algorithms reported in the literature are not able to integrate lossy and near-lossless/lossless coding and uses only causal pixels (non-symmetrical predictor structure) for prediction. A non-symmetrical predictor structure, however, is not able to efficiently adapt near the intensity varying areas, which results into poor prediction. Hence, we propose a novel two-stage algorithm for lossy, near lossless/lossless compression using a symmetrical predictor structure is proposed. In the first stage, the proposed algorithm encodes and decodes the given image using the JPEG-2000 standard algorithm (lossy coding). This JPEG-2000 decoded image in the first stage, enables us to use the symmetrical predictor (using both causal and non-causal pixels) for prediction in the second stage. A performance evaluation shows that our algorithm is significantly better in terms of compression performance as compared to some of the computationally complex methods. Vinit Jakhetiya, Oscar C. Au, Sunil Prasad Jaiswal, Luheng Jia, Gaurav Mittal |
ISCAS | 5 |
| 2012 | Bit-depth expansion using Minimum Risk Based ClassificationabstractBit-depth expansion is an art of converting low bit-depth image into high bit-depth image. Bit-depth of an image represents the number of bits required to represent an intensity value of the image. Bit-depth expansion is an important field since it directly affects the display quality. In this paper, we propose a novel method for bit-depth expansion which uses Minimum Risk Based Classification to create high bit-depth image. Blurring and other annoying artifacts are lowered in this method. Our method gives better objective (PSNR) and superior visual quality as compared to recently developed bit-depth expansion algorithms. Gaurav Mittal, Vinit Jakhetiya, Sunil Prasad Jaiswal, Oscar C. Au, Anil Kumar Tiwari, Dai Wei |
VCIP | 1 |
| 2010 | Automatic Generation of Stream Descriptors for Streaming ArchitecturesabstractWe describe a novel approach for automatically generating streaming architectures from software programs. While existing systems require user-defined stream models, our method automatically identifies producer-consumer streaming relationships and translates them into streaming architectures. Data streams between producer-consumer kernels are represented using a combination of stream descriptors and CFGs, which are categorized into four stream types. A bridge module is generated based on the stream type in the streaming architecture to facilitate data streaming between each producer-consumer pair. Several optimizations are also developed to improve throughput and parallelism. We demonstrate our results on a FPGA based platform. The automatically generated streaming architectures show 1.5-3x speedups over the non-streaming designs by employing spatial and temporal data independence to increase parallelism. David Zaretsky, Gaurav Mittal, Dan Schonfeld, Prithviraj Banerjee |
ICPP | 3 |
| 2009 | Streaming implementation of a sequential decompression algorithm on an FPGAabstractThis paper describes an FPGA based implementation of a real time compression algorithm used in transactions between financial institutions such as exchanges and trading houses. FIX is a protocol that has gained widespread popularity for exchanging financial information such as stock prices and purchases over the Internet. If a financial trader can speed up the processing of these protocols, he can make significant financial profits by buying or selling stocks when there is a lot of variability in the share prices. Our methodology tries to recognize and exploit streaming characteristics of the software design in order to implement a pipelined parallel processing system in reconfigurable hardware. It introduces the concept of caches to keep stream pipelines filled more often. The system implemented on a Xilinx Virtex5 LX110T FPGA shows a 17x speedup in throughput over a software implementation running on a dual core Intel Pentium workstation. These techniques are being developed as part of commercial compiler project to automatically translate software binaries to streaming RTL VHDL systems. Gaurav Mittal, David Zaretsky, Prithviraj Banerjee |
FPGA | 1 |
| 2009 | An Automated Algorithm to Generate Stream ProgramsabstractWith the proliferation of reconfigurable systems and flexible memory architectures, there has been intense interest in stream systems. While the existing stream systems require the programs to be written using special models, this paper demonstrates an approach to automatically generate stream programs from existing applications written for non-stream scalar processors. As a part of this approach, we provide a new comparison methodology and an algorithm to automatically generate stream descriptions. A second algorithm identifies processing kernels that can be pipelined. We demonstrate our results on an FPGA based platform. Gaurav Mittal, David Zaretsky, Dan Schonfeld, Prithviraj Banerjee |
ISCAS | 2 |
| 2009 | Streaming Implementation of the ZLIB Decoder Algorithm on an FPGAabstractMany new real-time system require high-speed compression and decompression solutions that provide low latency links between systems over a network interface. We describe a methodology for implementing an optimized streaming ZLIB decoder system on a Xilinx Virtex-5 FPGA board, which exploits the fine-grain parallelism in the software architecture to improve the performance. We describe a ZLIB decoder system in hardware and concrete examples of how to transform the sequential software algorithm into a highly optimized hardware implementation in RTL VHDL. Experimental results show 50times speedup in terms of cycles and 2.83times speedup in terms of time in the FPGA over the software. The ZLIB decoder was shown to operate at a rate of 1 GBit/s. David Zaretsky, Gaurav Mittal, Prithviraj Banerjee |
ISCAS | 2 |
| 2007 | An Overview of a Compiler for Mapping Software Binaries to HardwareabstractAs new applications in embedded communications and control systems push the computational limits of digital signal processing (DSP) functions, there will be an increasing need for software applications to be migrated to hardware in the form of a hardware-software codesign system. In many cases, access to the high-level source code may not be available. It is thus desirable to have a technology to translate the software binaries intended for processors to hardware implementations. This paper provides details on the retargetable FREEDOM compiler. The compiler automatically translates DSP software binaries to register-transfer level (RTL) VHDL and Verilog for implementation on field-programmable gate arrays (FPGAs) as standalone or system-on-chip implementations. We describe the underlying optimizations and some novel algorithms for alias analysis, data dependency analysis, memory optimizations, procedure call recovery, and back-end code scheduling. Experimental results on resource usage and performance are shown for several program binaries intended for the Texas Instruments C 6211 DSP (VLIW) and the ARM 922 T reduced instruction set computer (RISC) processors. Implementation results for four kernels from the Simulink demo library and others from commonly used DSP applications, such as MPEG-4, Viterbi, and JPEG are also discussed. The compiler generated RTL code is mapped to Xilinx Virtex II and Altera Stratix FPGAs. We record overall performance gains of 1.5-26.9 for the hardware implementations of the kernels. Comparisons with the power aware compiler techniques (PACT) high-level synthesis compiler are used to show that software binaries can be used as intermediate representations from any high-level language and generate efficient hardware implementations. Gaurav Mittal, David Zaretsky, Xiaoyong Tang, Prithviraj Banerjee |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2006 | An Efficient Video Enhancement Method Using LA*B* AnalysisabstractIn this paper, an efficient technique is proposed for real-time enhancement of video containing inconsistent and complex conditions like nonuniform and insufficient lighting. Beyond providing digital tools for video enhancement, this method provides a better approach to enhance a video in low lighting conditions without any loss of color information. This approach is based on histogram manipulations on La*b* color model of the video frame. This algorithm provides an effective way for video enhancement with simple computational procedures, which makes real-time enhancement for homeland security application successfully realized. Simulation is done on videos taken under bad lighting conditions and results are presented. Gaurav Mittal, Sushrutha Locharam, Sreela Sasi, Glenn R. Shaffer, Ajith K. Kumar |
AVSS | 1 |
| 2005 | Automatic extraction of function bodies from software binariesabstractThis paper describes a method for automatically extracting function bodies from linked software binaries. It utilizes procedure-calling conventions along with limited control and data flow information. It has been tested with the TI C6000 DSP processor platform. Results are reported on eight benchmarks for which our algorithm successfully identifies all functions. It identifies 198% more functions than by the use procedure calling conventions alone. Gaurav Mittal, David Zaretsky, Gokhan Memik, Prithviraj Banerjee |
ASP-DAC | 1 |
| 2004 | Automatic translation of software binaries onto FPGAsabstractThe introduction of advanced FPGA architectures, with built-in DSP support, has given DSP designers a new hardware alternative. By exploiting its inherent parallelism, it is expected that FPGAs can outperform DSP processors. This paper describes the process and considerations for automatically translating binaries targeted for general DSP processors into Register Transfer Level (RTL) VHDL or Verilog code to be mapped onto commercial FPGAs. The Texas Instruments C6000 DSP processor architecture is chosen as the DSP processor platform, and the Xilinx Virtex II as a target FPGA. Various optimizations are discussed, including data dependency analysis, procedure extraction, induction variable analysis, memory optimizations, and scheduling. Experimental results on resource usage and performance are shown for ten software binary benchmarks. Results show performance gains of 3-20X in the FPGA designs over that of the DSP processors in terms of reductions of execution cycles. Gaurav Mittal, David Zaretsky, Xiaoyong Tang, Prithviraj Banerjee |
DAC | 1 |
| 2004 | Overview of the FREEDOM Compiler for Mapping DSP Software to FPGAsabstractApplications that require digital signal processing (DSP) functions are typically mapped onto general purpose DSP processors. With the introduction of advanced FPGA architectures with built-in DSP support, a new hardware alternative is available for DSP designers. By exploiting its inherent parallelism, it is expected that FPGAs can outperform DSP processors. However, the migration of assembly code to hardware is typically a very arduous process. This paper describes the process and considerations for automatically translating software assembly and binary codes targeted for general DSP processors into register transfer level (RTL) VHDL or Verilog code to be mapped onto commercial FPGAs. The Texas instruments C6000 DSP processor architecture has been used as the DSP processor platform, and the Xilinx Virtex II as the target FPGA. Various optimizations are discussed, including loop unrolling, induction variable analysis, memory and register optimizations, scheduling and resource binding. Experimental results on resource usage and performance are shown for ten software binary benchmarks in the signal processing and image processing domains. Results show performance gains of 3-20x in terms of reductions in execution cycles and 1.3-5x in terms of reductions in execution times for the FPGA designs over that of the DSP processors in terms of reductions in execution cycles. David Zaretsky, Gaurav Mittal, Xiaoyong Tang, Prithviraj Banerjee |
FCCM | 2 |
| 2004 | Evaluation of scheduling and allocation algorithms while mapping assembly code onto FPGAsabstractMigration of software from older general purpose embedded processors onto newer mixed hardware/software Systems-On-Chip (SOC) platforms is becoming an increasingly important topic. Automatic translation of general purpose software binaries and assembly code onto hardware implementations using FPGAs require sophisticated scheduling and allocation algorithms to maximize the resource utilization of such hardware devices. This paper describes the effects of scheduling and chaining of node operations in a CDFG onto an FPGA. The effects of register allocation on scheduled nodes are also discussed. The Texas Instruments C6000 DSP processor architecture was chosen as the DSP processor platform and assembly code, and the Xilinx Virtex II XC2V250 was chosen as the target FPGA. Results are reported on ten benchmarks, which show that scheduling with chaining operations produces the best results on FPGAs, while the addition of register allocation in fact generates poorer designs in terms of area and frequency. David Zaretsky, Gaurav Mittal, Xiaoyong Tang, Prithviraj Banerjee |
ACM Great Lakes Symposium on VLSI | 2 |