Suren Jayasuriya

dblp:153/9770 · DBLP profile ↗
← Back
38ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0001-7143-4429ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 12 since 2021Systems, architecture and hardware · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 SH-SAS: An Implicit Neural Representation for Complex Spherical-Harmonic Scattering Fields for 3D Synthetic Aperture Sonar
abstract
Synthetic aperture sonar (SAS) reconstruction requires recovering both the spatial distribution of acoustic scatterers and their direction-dependent response. Time-domain backprojection is the most common 3D SAS reconstruction algorithm, but it does not model directionality and can suffer from sampling limitations, aliasing and occlusion. Prior neural volumetric methods applied to synthetic aperture sonar, e.g. Reed et al. [43], treat each voxel as an isotropic scattering density, not modeling anisotropic returns. We introduce SH-SAS, an implicit neural representation that expresses the complex acoustic scattering field as a set of spherical harmonic (SH) coefficients. A multi-resolution hash encoder feeds a lightweight MLP that outputs complex SH coefficients up to a specified degree L. The zerothorder coefficient acts as an isotropic scattering field, which also serves as the density term, while higher orders compactly capture directional scattering with minimal parameter overhead. Because the model predicts the complex amplitude for any transmit-receive baseline, training is performed directly from 1-D time-of-flight (ToF) signals without the need to beamform intermediate images for supervision. Across synthetic and real SAS (both in-air and underwater) benchmarks, results show that SH-SAS performs better in terms of 3D reconstruction quality and geometric metrics than previous methods such as time-domain backprojection and Reed et al. [43].
Omkar Vengurlekar, Adithya Kumar Pediredla, Suren Jayasuriya
3DV3
2025 VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
abstract
Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly across multiple images remains a significant challenge. To address this, we introduce VOILA, a large-scale, open-ended, dynamic benchmark designed to evaluate MLLMs' perceptual understanding and abstract relational reasoning. VOILA employs an analogical mapping approach in the visual domain, requiring models to generate an image that completes an analogy between two given image pairs, reference and application, without relying on predefined choices. Our experiments demonstrate that the analogical reasoning tasks in VOILA present a challenge to MLLMs. Through multi-step analysis, we reveal that current MLLMs struggle to comprehend inter-image relationships and exhibit limited capabilities in high-level relational reasoning. Notably, we observe that performance improves when following a multi-step strategy of least-to-most prompting. Comprehensive evaluations on open-source models and GPT-4o show that on text-based answers, the best accuracy for challenging scenarios is 13% (LLaMa 3.2) and even for simpler tasks is only 29% (GPT-4o), while human performance is significantly higher at 70% across both difficulty levels.
Nilay Yilmaz, Maitreya Patel, Yiran Luo 0001, Tejas Gokhale, Chitta Baral, Suren Jayasuriya, Yezhou Yang
ICLR6
2025 MetaVIn: Meteorological and Visual Integration for Atmospheric Turbulence Strength Estimation
abstract
Long-range image understanding is a challenging task for computer vision due to the presence of atmospheric turbulence. Turbulence can degrade image quality (blur and geometric distortion) due to the medium's spatio-temporal varying index of refraction bending light rays. The strength of atmospheric turbulence is quantified by the refractive index structure parameter$C_n^2$, and estimating it is important both as an indicator of image degradation and is useful for downstream tasks including video restoration and estimating true shape and range/depth. However, traditional methods for estimating$C_{n}^{2}$involve expensive and complex optical equipment, limiting their practicality. In this paper, we propose MetaVIn: a Meteorological and Visual Integration system to predict atmospheric turbulence strength. Our method leverages image quality metrics to capture sharpness and blur, combined with meteorological information within a Kolmogorov Arnold Network (KAN). We demonstrate that this approach provides a more accurate and generalizable estimation of$C_{n}^{2}$, outperforming previous state-of-the-art methods in both blind image quality assessment and passive video-based turbulence strength estimation on a large dataset of 35,364 image samples with accompanying ground truth scintillometer measurements for$C_n^2$. Our method enables better prediction and mitigation of atmospheric image degradation while being useful in applications such as shape and range estimation, enhancing the practical utility of our approach.
Ripon K. Saha, Scott McCloskey, Suren Jayasuriya
WACV3
2025 DAATSim: Depth-Aware Atmospheric Turbulence Simulation for Fast Image Rendering
abstract
Abstract Simulating the effects of atmospheric turbulence for imaging systems operating over long distances is a significant challenge for optical and computer graphics models. Physically‐based ray tracing over kilometers of distance is difficult due to the need to define a spatio‐temporal volume of varying refractive index. Even if such a volume can be defined, Monte Carlo rendering approximations for light refraction through the environment would not yield real‐time solutions needed for video game engines or online dataset augmentation for machine learning. While existing simulators based on procedurally‐generated noise or textures have been proposed in these settings, these simulators often neglect the significant impact of scene depth, leading to unrealistic degradations for scenes with substantial foreground‐background separation. This paper introduces a novel, physically‐based atmospheric turbulence simulator that explicitly models depth‐dependent effects while rendering frames at interactive/near real‐time (> 10 FPS) rates for image resolutions up to 1024 × 1024 (real‐time 35 FPS at 256× 256 resolution with depth or 512×512 at 33 FPS without depth). Our hybrid approach combines spatially‐varying wavefront aberrations using Zernike polynomials with pixel‐wise depth modulation of both blur (via Point Spread Function interpolation) and geometric distortion or tilt. Our approach includes a novel fusion technique that integrates complementary strengths of leading monocular depth estimators to generate metrically accurate depth maps with enhanced edge fidelity. DAATSim is implemented efficiently on GPUs using Py‐Torch incorporating optimizations like mixed‐precision computation and caching to achieve efficient performance. We present quantitative and qualitative validation demonstrating the simulator's physical plausibility for generating turbulent video. DAAT‐Sim is made publicly available and open‐source to the community: https://github.com/Riponcs/DAATSim .
Ripon K. Saha, Yufan Zhang 0001, Jinwei Ye, Suren Jayasuriya
Comput. Graph. Forum4
2025 Z-Splat: Z-Axis Gaussian Splatting for Camera-Sonar Fusion
abstract
Differentiable 3D-Gaussian splatting (GS) is emerging as a prominent technique in computer vision and graphics for reconstructing 3D scenes. GS represents a scene as a set of 3D Gaussians with varying opacities and employs a computationally efficient splatting operation along with analytical derivatives to compute the 3D Gaussian parameters given scene images captured from various viewpoints. Unfortunately, capturing surround view ($\text{360}^{\circ }$360∘ viewpoint) images is impossible or impractical in many real-world imaging scenarios, including underwater imaging, rooms inside a building, and autonomous navigation. In these restricted baseline imaging scenarios, the GS algorithm suffers from a well-known 'missing cone' problem, which results in poor reconstruction along the depth axis. In this paper, we demonstrate that using transient data (from sonars) allows us to address the missing cone problem by sampling high-frequency data along the depth axis. We extend the Gaussian splatting algorithms for two commonly used sonars and propose fusion algorithms that simultaneously utilize RGB camera data and sonar data. Through simulations, emulations, and hardware experiments across various imaging scenarios, we show that the proposed fusion algorithms lead to significantly better novel view synthesis (5 dB improvement in PSNR) and 3D geometry reconstruction (60% lower Chamfer distance).
Ziyuan Qu, Omkar Vengurlekar, Mohamad Qadri, Kevin Zhang 0003, Michael Kaess, Christopher A. Metzler, Suren Jayasuriya, Adithya Kumar Pediredla
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 ImageSTEAM: Teacher Professional Development for Integrating Visual Computing into Middle School Lessons
abstract
Artificial intelligence (AI) and its teaching in the K-12 grades has been championed as a vital need for the United States due to the technology's future prominence in the 21st century. However, there remain several barriers to effective AI lessons at these age groups including the broad range of interdisciplinary knowledge needed and the lack of formal training or preparation for teachers to implement these lessons. In this experience report, we present ImageSTEAM, a teacher professional development for creating lessons surrounding computer vision, machine learning, and computational photography/cameras targeted for middle school grades 6-8 classes. Teacher professional development workshops were conducted in the states of Arizona and Georgia from 2021-2023 where lessons were co-created with teachers to introduce various specific visual computing concepts while aligning to state and national standards. In addition, the use of a variety of computer vision and image processing software including custom designed Python notebooks were created as technology activities and demonstrations to be used in the classroom. Educational research showed that teachers improved their self-efficacy and outcomes for concepts in computer vision, machine learning, and artificial intelligence when participating in the program. Results from the professional development workshops highlight key opportunities and challenges in integrating this content into the standard curriculum, the benefits of a co-creation pedagogy, and the positive impact on teacher and student's learning experiences. The open-source program curriculum is available at www.imagesteam.org.
Suren Jayasuriya, Kimberlee Swisher, Joshua D. Rego, Sreenithy Chandran, John M. Mativo, Terri Kurz, Cerenity E. Collins, Dawn T. Robinson, Ramana Pidaparti
AAAI1
2024 Turb-Seg-Res: A Segment-then-Restore Pipeline for Dynamic Videos with Atmospheric Turbulence
abstract
Tackling image degradation due to atmospheric turbu-lence, particularly in dynamic environments, remains a challenge for long-range imaging systems. Existing techniques have been primarily designed for static scenes or scenes with small motion. This paper presents the first segment-then-restore pipeline for restoring the videos of dy-namic scenes in turbulent environments. We leverage mean optical flow with an unsupervised motion segmentation method to separate dynamic and static scene components prior to restoration. After camera shake compensation and segmentation, we introduce foreground/background en-hancement leveraging the statistics of turbulence strength and a transformer model trained on a novel noise-based procedural turbulence generator for fast dataset augmen-tation. Benchmarked against existing restoration meth-ods, our approach restores most of the geometric distortion and enhances the sharpness of videos. We make our code, simulator, and data publicly available to ad-vance the field of video restoration from turbulence: riponcs.github.io/TurbSegRes
Ripon K. Saha, Dehao Qin, Nianyi Li, Jinwei Ye, Suren Jayasuriya
CVPR5
2024 Unsupervised Moving Object Segmentation with Atmospheric Turbulence
Dehao Qin, Ripon K. Saha, Woojeh Chung, Suren Jayasuriya, Jinwei Ye, Nianyi Li
ECCV (6)4
2024 PathFinder: Attention-Driven Dynamic Non-Line-of-Sight Tracking with a Mobile Robot
abstract
The study of non-line-of-sight (NLOS) imaging is growing due to its many potential applications, including rescue operations and pedestrian detection by self-driving cars. However, implementing NLOS imaging on a moving camera remains an open area of research. Existing NLOS imaging methods rely on time-resolved detectors and laser configurations that require precise optical alignment, making it difficult to deploy them in dynamic environments. This work proposes a data-driven approach to NLOS imaging, PathFinder, that can be used with a standard RGB camera mounted on a small, power-constrained mobile robot, such as an aerial drone. Our experimental pipeline is designed to accurately estimate the 2D trajectory of a person who moves in a Manhattan-world environment while remaining hidden from the camera’s field-of-view. We introduce a novel approach to process a sequence of dynamic successive frames in a line-of-sight (LOS) video using an attention-based neural network that performs inference in real-time. The method also includes a preprocessing selection metric that analyzes images from a moving camera which contain multiple vertical planar surfaces, such as walls and building facades, and extracts planes that return maximum NLOS information. We validate the approach on in-the-wild scenes using a drone for video capture, thus demonstrating low-cost NLOS imaging in dynamic capture environments. The real-world dataset that we collected and used to train the network can be found at https://srchandr.github.io/DynamicNLOS/.
Shenbagaraj Kannapiran, Sreenithy Chandran, Suren Jayasuriya, Spring Berman
IROS3
2024 NeRF-enabled Analysis-Through-Synthesis for ISAR Imaging of Small Everyday Objects with Sparse and Noisy UWB Radar Data
abstract
Inverse Synthetic Aperture Radar (ISAR) imaging presents a formidable challenge when it comes to small everyday objects due to their limited Radar Cross-Section (RCS) and the inherent resolution constraints of radar systems. Existing ISAR reconstruction methods including backprojection (BP) often require complex setups and controlled environments, rendering them impractical for many real-world noisy scenarios. In this paper, we propose a novel Analysis-through-Synthesis (ATS) framework enabled by Neural Radiance Fields (NeRF) for high-resolution coherent ISAR imaging of small objects using sparse and noisy Ultra-Wideband (UWB) radar data with an inexpensive and portable setup. Our end-to-end framework integrates ultra-wideband radar wave propagation, reflection characteristics, and scene priors, enabling efficient 2D scene reconstruction without the need for costly anechoic chambers or complex measurement test beds. With qualitative and quantitative comparisons, we demonstrate that the proposed method outperforms traditional techniques and generates ISAR images of complex scenes with multiple targets and complex structures in Non-Line-of-Sight (NLOS) and noisy scenarios, particularly with limited number of views and sparse UWB radar scans. This work represents a significant step towards practical, cost-effective ISAR imaging of small everyday objects, with broad implications for robotics and mobile sensing applications.
Md. Farhan Tasnim Oshim, Albert W. Reed, Suren Jayasuriya, Tauhidur Rahman
IROS3
2024 Learning-based Spotlight Position Optimization for Non-Line-of-Sight Human Localization and Posture Classification
abstract
Non-line-of-sight imaging (NLOS) is the process of estimating information about a scene that is hidden from the direct line of sight of the camera. NLOS imaging typically requires time-resolved detectors and a laser source for illumination, which are both expensive and computationally intensive to handle. In this paper, we propose an NLOS-based localization and posture classification technique that uses an off-the-shelf projector and camera system. We leverage a message-passing neural network to learn a visible scene geometry and predict the best position to be spotlighted by the projector that can maximize the NLOS signal. The neural network is trained end-to-end and the network parameters are optimized to maximize the NLOS performance. Unlike prior deep-learning-based NLOS techniques that assume planar relay walls, our system allows us to handle line-of-sight scenes where scene geometries are more arbitrary. Our method demonstrates state-of-the-art performance in object localization and position classification using both synthetic and real scenes.
Sreenithy Chandran, Tatsuya Yatagawa, Hiroyuki Kubo, Suren Jayasuriya
WACV4
2023 Learning-Based Tone Mapping to Improve 3D SAS ATR
abstract
Automatic target recognition (ATR) for 3D synthetic aperture sonar (SAS) imagery is an intrinsic challenge in highly cluttered ocean environments, especially for objects partially or completely buried in the sediment. Conventional dynamic range compression (DRC) techniques such as log-compression, which is a type of tone mapping intended to appeal to the human visual system, can further obscure the sonar signatures of these already physically occluded objects and lead to suboptimal downstream ATR performance, particularly for convolutional neural networks (CNNs). In this paper, we present a novel machine learning-based approach for tone mapping sub-bottom SAS imagery as a pre-processing stage in the 3D SAS ATR pipeline. This learned tone mapping function can be jointly optimized with a CNN-based ATR algorithm. We train and validate our method on measured volumetric SAS data captured by the Sediment Volume Search Sonar (SVSS) system.
Gregory D. Vetaw, Benjamin Cowen, Daniel C. Brown, Suren Jayasuriya
IGARSS5
2023 Learning Repeatable Speech Embeddings Using An Intra-class Correlation Regularizer
abstract
A good supervised embedding for a specific machine learning task is only sensitive to changes in the label of interest and is invariant to other confounding factors. We leverage the concept of repeatability from measurement theory to describe this property and propose to use the intra-class correlation coefficient (ICC) to evaluate the repeatability of embeddings. We then propose a novel regularizer, the ICC regularizer, as a complementary component for contrastive losses to guide deep neural networks to produce embeddings with higher repeatability. We use simulated data to explain why the ICC regularizer works better on minimizing the intra-class variance than the contrastive loss alone. We implement the ICC regularizer and apply it to three speech tasks: speaker verification, voice style conversion, and a clinical application for detecting dysphonic voice. The experimental results demonstrate that adding an ICC regularizer can improve the repeatability of learned embeddings compared to only using the contrastive loss; further, these embeddings lead to improved performance in these downstream tasks.
Suren Jayasuriya, Visar Berisha
NeurIPS2
2023 Software-Defined Imaging: A Survey
abstract
Huge advancements have been made over the years in terms of modern image-sensing hardware and visual computing algorithms (e.g., computer vision, image processing, and computational photography). However, to this day, there still exists a current gap between the hardware and software design in an imaging system, which silos one research domain from another. Bridging this gap is the key to unlocking new visual computing capabilities for end applications in commercial photography, industrial inspection, and robotics. In this survey, we explore existing works in the literature that can be leveraged to replace conventional hardware components in an imaging system with software for enhanced reconfigurability. As a result, the user can program the image sensor in a way best suited to the end application. We refer to this as software-defined imaging (SDI), where image sensor behavior can be altered by the system software depending on the user’s needs. The scope of our survey covers imaging systems for single-image capture, multi-image, and burst photography, as well as video. We review works related to the sensor primitives, image signal processor (ISP) pipeline, computer architecture, and operating system elements of the SDI stack. Finally, we outline the infrastructure and resources for SDI systems, and we also discuss possible future research directions for the field.
Suren Jayasuriya, Odrika Iqbal, Venkatesh Kodukula, Victor Isaac Torres Muro, Robert LiKamWa, Andreas Spanias
Proc. IEEE1
2023 Robust Vocal Quality Feature Embeddings for Dysphonic Voice Detection
abstract
Approximately 1.2% of the world's population has impaired voice production. As a result, automatic dysphonic voice detection has attracted considerable academic and clinical interest. However, existing methods for automated voice assessment often fail to generalize outside the training conditions or to other related applications. In this paper, we propose a deep learning framework for generating acoustic feature embeddings sensitive to vocal quality and robust across different corpora. A contrastive loss is combined with a classification loss to train our deep learning model jointly. Data warping methods are used on input voice samples to improve the robustness of our method. Empirical results demonstrate that our method not only achieves high in-corpus and cross-corpus classification accuracy but also generates good embeddings sensitive to voice quality and robust across different corpora. We also compare our results against three baseline methods on clean and three variations of deteriorated in-corpus and cross-corpus datasets and demonstrate that the proposed model consistently outperforms the baseline methods.
Julie M. Liss, Suren Jayasuriya, Visar Berisha
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Neural Volumetric Reconstruction for Coherent Synthetic Aperture Sonar
abstract
Synthetic aperture sonar (SAS) measures a scene from multiple views in order to increase the resolution of reconstructed imagery. Image reconstruction methods for SAS coherently combine measurements to focus acoustic energy onto the scene. However, image formation is typically under-constrained due to a limited number of measurements and bandlimited hardware, which limits the capabilities of existing reconstruction methods. To help meet these challenges, we design an analysis-by-synthesis optimization that leverages recent advances in neural rendering to perform coherent SAS imaging. Our optimization enables us to incorporate physics-based constraints and scene priors into the image formation process. We validate our method on simulation and experimental results captured in both air and water. We demonstrate both quantitatively and qualitatively that our method typically produces superior reconstructions than existing approaches. We share code and data for reproducibility.
Albert W. Reed, Thomas E. Blanford, Adithya Kumar Pediredla, Daniel C. Brown, Suren Jayasuriya
ACM Trans. Graph.6
2021 Zen and the Art of STEAM: Student Knowledge and Experiences in Interdisciplinary and Traditional Engineering Capstone Experiences
abstract
This Research Full Paper examines the concept of flow, derived from Zen philosophy and positive psychology, and how interdisciplinary STEAM (science, technology, engineering, arts, and mathematics) and disciplinary electrical engineering students find flow within their coursework and their capstone design experiences. STEAM education incorporates the arts and humanities into the traditional disciplines of STEM. However, students involved in this interdisciplinary space often struggle to find a balance in applying both creative and logical knowledge in their work. The theoretical framework for this study leverages the concept of pure experience from Zen philosophy to analyze flow states in students' interdisciplinary experiences. This theory focuses on the unity of subject/object and rejection of purely logical, positivist thinking for more integrative knowledge acquisition while in flow states. In this secondary analysis, we analyzed interviews conducted with electrical engineering and STEAM students. STEAM students from an interdisciplinary program were found to approach their coursework differently than engineering students, likely because of a difference in assignment guidelines. The engineering students in the study had more restrictive guidelines, while the STEAM students were given more freedom to move between disciplines. Alternatively, students from both disciplines shared many similar values about education and knowledge including the need for enjoyment and personal interest within the coursework as well as finding a balance between logical thought and the desire for creation that a student's program did not determine whether they reached a state of pure experience, or flow, in their work. However, rigid adherence to either the arts or engineering seemed to create disharmony and very few students find cohesion between their values and their approach to knowledge. This paper points to new insights into the design of capstone experiences for STEAM education.
Dominique Dredd, Nadia N. Kellam, Suren Jayasuriya
FIE3
2021 Use of AI-Generated Visual Media in Interviews to Understand Power Differentials in Gender, Romantic, and Sexual Minority Students
abstract
This work-in-progress briefly outlines the theoretical background, methods, and preliminary results of a qualitative study conducted with gender, romantic, and sexual minority (GRSM) students immersed in higher education spaces. We elaborate on the efficacy of our innovative qualitative methodologies through the use of AI-human art-making interactions during our interviews, which helped to produce richer qualitative data from our participants. Our methodology was constructed using a Foucauldian theoretical framework to inform the framework of this study, focusing explicitly on GRSM students' experiences with power in higher education and when using technology, as well as the ways in which they resist power through the use of technology and AI -generated visual media.
Madeleine Jennings, Jorge Sandoval, Jeanne Sanders, Mirka Koro, Nadia N. Kellam, Suren Jayasuriya
FIE6
2021 Unsupervised Non-Rigid Image Distortion Removal via Grid Deformation
abstract
Many computer vision problems face difficulties when imaging through turbulent refractive media (e.g., air and water) due to the refraction and scattering of light. These effects cause geometric distortion that requires either handcrafted physical priors or supervised learning methods to remove. In this paper, we present a novel unsupervised network to recover the latent distortion-free image. The key idea is to model non-rigid distortions as deformable grids. Our network consists of a grid deformer that estimates the distortion field and an image generator that outputs the distortion-free image. By leveraging the positional encoding operator, we can simplify the network structure while maintaining fine spatial details in the recovered images. Our method doesn't need to be trained on labeled data and has good transferability across various turbulent image datasets with different types of distortions. Extensive experiments on both simulated and real-captured turbulent images demonstrate that our method can remove both air and water distortions without much customization.
Nianyi Li, Simron Thapa, Cameron Whyte, Albert W. Reed, Suren Jayasuriya, Jinwei Ye
ICCV5
2021 Dynamic CT Reconstruction from Limited Views with Implicit Neural Representations and Parametric Motion Fields
abstract
Reconstructing dynamic, time-varying scenes with computed tomography (4D-CT) is a challenging and ill-posed problem common to industrial and medical settings. Existing 4D-CT reconstructions are designed for sparse sampling schemes that require fast CT scanners to capture multiple, rapid revolutions around the scene in order to generate high quality results. However, if the scene is moving too fast, then the sampling occurs along a limited view and is difficult to reconstruct due to spatiotemporal ambiguities. In this work, we design a reconstruction pipeline using implicit neural representations coupled with a novel parametric motion field warping to perform limited view 4D-CT reconstruction of rapidly deforming scenes. Importantly, we utilize a differentiable analysis-bysynthesis approach to compare with captured x-ray sinogram data in a self-supervised fashion. Thus, our resulting optimization method requires no training data to reconstruct the scene. We demonstrate that our proposed system robustly reconstructs scenes containing deformable and periodic motion and validate against state-of-the-art baselines. Further, we demonstrate an ability to reconstruct continuous spatiotemporal representations of our scenes and upsample them to arbitrary volumes and frame rates post-optimization. This research opens a new avenue for implicit neural representations in computed tomography reconstruction in general. Code is available at https://github.com/awreed/DynamicCTReconstruction.
Albert W. Reed, Hyojin Kim 0001, Rushil Anirudh, K. Aditya Mohan, Kyle Champley, Suren Jayasuriya
ICCV7
2021 Restoring Degraded Speech via a Modified Diffusion Model
abstract
There are many deterministic mathematical operations (e.g. compression, clipping, downsampling) that degrade speech quality considerably. In this paper we introduce a neural network architecture, based on a modification of the DiffWave model, that aims to restore the original speech signal. DiffWave, a recently published diffusion-based vocoder, has shown state-of-the-art synthesized speech quality and relatively shorter waveform generation times, with only a small set of parameters. We replace the mel-spectrum upsampler in DiffWave with a deep CNN upsampler, which is trained to alter the degraded speech mel-spectrum to match that of the original speech. The model is trained using the original speech waveform, but conditioned on the degraded speech mel-spectrum. Post-training, only the degraded mel-spectrum is used as input and the model generates an estimate of the original speech. Our model results in improved speech quality (original DiffWave model as baseline) on several different experiments. These include improving the quality of speech degraded by LPC-10 compression, AMR-NB compression, and signal clipping. Compared to the original DiffWave architecture, our scheme achieves better performance on several objective perceptual metrics and in subjective comparisons. Improvements over baseline are further amplified in a out-of-corpus evaluation setting.
Suren Jayasuriya, Visar Berisha
Interspeech2
2021 Robust Lensless Image Reconstruction via PSF Estimation
abstract
Lensless imaging is a new, emerging modality where image sensors utilize optical elements in front of the sensor to perform multiplexed imaging. There have been several recent papers to reconstruct images from lensless imagers, including methods that utilize deep learning for state-of-the-art performance. However, many of these methods require explicit knowledge of the optical element, such as the point spread function, or learn the reconstruction mapping for a single fixed PSF. In this paper, we explore a neural network architecture that performs joint image reconstruction and PSF estimation to robustly recover images captured with multiple PSFs from different cameras. Using adversarial learning, this approach achieves improved reconstruction results that do not require explicit knowledge of the PSF at test-time and shows an added improvement in the reconstruction model's ability to generalize to variations in the camera's PSF. This allows lensless cameras to be utilized in a wider range of applications that require multiple cameras without the need to explicitly train a separate model for each new camera.
Joshua D. Rego, Karthik Kulkarni, Suren Jayasuriya
WACV3
2021 Programmable Non-Epipolar Indirect Light Transport: Capture and Analysis
abstract
The decomposition of light transport into direct and global components, diffuse and specular interreflections, and subsurface scattering allows for new visualizations of light in everyday scenes. In particular, indirect light contains a myriad of information about the complex appearance of materials useful for computer vision and inverse rendering applications. In this paper, we present a new imaging technique that captures and analyzes components of indirect light via light transport using a synchronized projector-camera system. The rectified system illuminates the scene with epipolar planes corresponding to projector rows, and we vary two key parameters to capture plane-to-ray light transport between projector row and camera pixel: (1) the offset between projector row and camera row in the rolling shutter (implemented as synchronization delay), and (2) the exposure of the camera row. We describe how this synchronized rolling shutter performs illumination multiplexing, and develop a nonlinear optimization algorithm to demultiplex the resulting 3D light transport operator. Using our system, we are able to capture live short and long-range non-epipolar indirect light transport, disambiguate subsurface scattering, diffuse and specular interreflections, and distinguish materials according to their subsurface scattering properties. In particular, we show the utility of indirect imaging for capturing and analyzing the hidden structure of veins in human skin.
Hiroyuki Kubo, Suren Jayasuriya, Takafumi Iwaguchi, Takuya Funatomi, Yasuhiro Mukaigawa, Srinivasa G. Narasimhan
IEEE Trans. Vis. Comput. Graph.2
2020 Differentiable Programming for Hyperspectral Unmixing Using a Physics-Based Dispersion Model
John Janiczek, Parth Thaker, Gautam Dasarathy, Christopher S. Edwards, Philip Christensen, Suren Jayasuriya
ECCV (27)6
2020 Analyzing Sensor Quantization Of Raw Images For Visual Slam
abstract
Visual simultaneous localization and mapping (SLAM) is an emerging technology that enables low-power devices with a single camera to perform robotic navigation. However, most visual SLAM algorithms are tuned for images produced through the image sensor processing (ISP) pipeline optimized for highly aesthetic photography. In this paper, we investigate the feasibility of varying sensor quantization on RAW images directly from the sensor to save energy for visual SLAM. In particular, we compare linear and logarithmic image quantization and show visual SLAM is robust to the latter. Further, we introduce a new gradient-based image quantization scheme that outperforms logarithmic quantization's energy savings while preserving accuracy for feature-based visual SLAM algorithms. This work opens a new direction in energy-efficient image sensing for SLAM in the future.
Olivia Christie, Joshua D. Rego, Suren Jayasuriya
ICIP3
2020 Design and FPGA Implementation of an Adaptive video Subsampling Algorithm for Energy-Efficient Single Object Tracking
abstract
Image sensors with programmable region-of-interest (ROI) readout are a new sensing technology important for energyefficient embedded computer vision. In particular, ROIs can subsample the number of pixels being readout while performing single object tracking in a video. In this paper, we develop adaptive sampling algorithms which perform joint object tracking and predictive video subsampling. We utilize an object detection consisting of either mean shift tracking or a neural network, coupled with a Kalman filter for prediction. We show that our algorithms achieve mean average precision of 0.70 or higher on a dataset of 20 videos in software. Further, we implement hardware acceleration of mean shift tracking with Kalman filter adaptive subsampling on an FPGA. Hardware results show a 23 × improvement in clock cycles and latency as compared to baseline methods and achieves 38FPS real-time performance. This research points to a new domain of hardware-software co-design for adaptive video subsampling in embedded computer vision.
Odrika Iqbal, Saquib Siddiqui, Joshua Martin, Sameeksha Katoch, Andreas Spanias, Daniel W. Bliss, Suren Jayasuriya
ICIP7
2019 Adaptive Lighting for Data-Driven Non-Line-of-Sight 3D Localization and Object Identification
Sreenithy Chandran, Suren Jayasuriya
BMVC2
2019 An REU Experience in Machine Learning and Computational Cameras
abstract
In this work in progress paper, we describe an REU summer experience on imaging sensors that involved a female junior level Electrical Engineering student, a graduate student advisor, and three faculty. A research plan was designed to embed the student in a sensor and machine learning research with specific emphasis on energy-efficient cameras. The motivation for submitting this paper is the unique planning and the quality of the overall student experience which resulted in continuous engagement of the REU student with the faculty after the REU summer program completed. The program resulted in a major presentation at an international event, an NSF I/UCRC poster presentation, a research conference submission which is remarkable for an undergraduate student, and finally a new research direction for the graduate mentor and faculty. This paper describes successful strategies for research engagement for undergraduates in state-of-the-art research fields which yield positive outcomes for all participants, and is grounded in contemporary educational methodology and theoretical frameworks.
Divya Mohan, Sameeksha Katoch, Suren Jayasuriya, Pavan Turaga, Andreas Spanias
FIE3
2019 Slope Disparity Gating using a Synchronized Projector-Camera System
abstract
Active illumination systems which perform disparity gating, or the ability to selectively image photons that arrive from a specified surface geometry some distance away, have recently shown usefulness for robotics, autonomous vehicles, and surveillance applications. In this paper, we present a new technique for sloped disparity gating, capturing a particular set of sloped planar surfaces in a scene, implemented using the synchronization between a raster-scanning projector and the rolling shutter of a camera. We demonstrate how to control the slope and thickness of these planar surfaces using hardware parameters of pixel clock, synchronization delay, and exposure. Finally, we perform applications including real-time image masking and imaging in scattering media with a real hardware prototype in the lab. This work showcases the potential for energy-efficient, geometry-aware disparity gating in the future.
Tomoki Ueda, Hiroyuki Kubo, Suren Jayasuriya, Takuya Funatomi, Yasuhiro Mukaigawa
ICCP3
2019 Non-Parametric Priors For Generative Adversarial Networks
abstract
The advent of generative adversarial networks (GAN) has enabled new capabilities in synthesis, interpolation, and data augmentation heretofore considered very challenging. However, one of the common assumptions in most GAN architectures is the assumption of simple parametric latent-space distributions. While easy to implement, a simple latent-space distribution can be problematic for uses such as interpolation. This is due to distributional mismatches when samples are interpolated in the latent space. We present a straightforward formalization of this problem; using basic results from probability theory and off-the-shelf-optimization tools, we develop ways to arrive at appropriate non-parametric priors. The obtained prior exhibits unusual qualitative properties in terms of its shape, and quantitative benefits in terms of lower divergence with its mid-point distribution. We demonstrate that our designed prior helps improve image generation along any Euclidean straight line during interpolation, both qualitatively and quantitatively, without any additional training or architectural modifications. The proposed formulation is quite flexible, paving the way to impose newer constraints on the latent-space statistics.
Rajhans Singh, Pavan Turaga, Suren Jayasuriya, Ravi Garg, Martin W. Braun
ICML3
2018 Acquiring and characterizing plane-to-ray indirect light transport
abstract
Separation of light transport into direct and indirect paths has enabled new visualizations of light in everyday scenes. However, indirect light itself contains a variety of components from subsurface scattering to diffuse and specular interreflections, all of which contribute to complex visual appearance. In this paper, we present a new imaging technique that captures and analyzes these components of indirect light via light transport between epipolar planes of illumination and rays of received light. This plane-to-ray light transport is captured using a rectified projector-camera system where we vary the offset between projector and camera rows (implemented as synchronization delay) as well as the exposure of each camera row. The resulting delay-exposure stack of images can capture live short and long-range indirect light transport, disambiguate subsurface scattering, diffuse and specular interreflections, and distinguish materials according to their subsurface scattering properties.
Hiroyuki Kubo, Suren Jayasuriya, Takafumi Iwaguchi, Takuya Funatomi, Yasuhiro Mukaigawa, Srinivasa G. Narasimhan
ICCP2
2018 CS-VQA: Visual Question Answering with Compressively Sensed Images
abstract
Visual Question Answering (VQA) is a complex semantic task requiring both natural language processing and visual recognition. In this paper, we explore whether VQA is solvable when images are captured in a sub-Nyquist compressive paradigm. We develop a series of deep-network architectures that exploit available compressive data to increasing degrees of accuracy, and show that VQA is indeed solvable in the compressed domain. Our results show that there is nominal degradation in VQA performance when using compressive measurements, but that accuracy can be recovered when VQA pipelines are used in conjunction with state-of-the-art deep neural networks for CS reconstruction. The results presented yield important implications for resource-constrained VQA applications.
Li-Chi Huang, Kuldeep Kulkarni, Anik Jha, Suhas Lohit, Suren Jayasuriya, Pavan Turaga
ICIP5
2018 EVA2: Exploiting Temporal Redundancy in Live Computer Vision
abstract
Hardware support for deep convolutional neural networks (CNNs) is critical to advanced computer vision in mobile and embedded devices. Current designs, however, accelerate generic CNNs; they do not exploit the unique characteristics of real-time vision. We propose to use the temporal redundancy in natural video to avoid unnecessary computation on most frames. A new algorithm, activation motion compensation, detects changes in the visual input and incrementally updates a previously-computed activation. The technique takes inspiration from video compression and applies well-known motion estimation techniques to adapt to visual changes. We use an adaptive key frame rate to control the trade-off between efficiency and vision quality as the input changes. We implement the technique in hardware as an extension to state-of-the-art CNN accelerator designs. The new unit reduces the average energy per frame by 54%, 62%, and 87% for three CNNs with less than 1% loss in vision accuracy.
Mark Buckler, Philip Bedoukian, Suren Jayasuriya, Adrian Sampson
ISCA3
2017 Reconfiguring the Imaging Pipeline for Computer Vision
abstract
Advancements in deep learning have ignited an explosion of research on efficient hardware for embedded computer vision. Hardware vision acceleration, however, does not address the cost of capturing and processing the image data that feeds these algorithms. We examine the role of the image signal processing (ISP) pipeline in computer vision to identify opportunities to reduce computation and save energy. The key insight is that imaging pipelines should be be configurable: to switch between a traditional photography mode and a low-power vision mode that produces lower-quality image data suitable only for computer vision. We use eight computer vision algorithms and a reversible pipeline simulation tool to study the imaging system's impact on vision performance. For both CNN-based and classical vision algorithms, we observe that only two ISP stages, demosaicing and gamma compression, are critical for task performance. We propose a new image sensor design that can compensate for these stages. The sensor design features an adjustable resolution and tunable analog-to-digital converters (ADCs). Our proposed imaging system's vision mode disables the ISP entirely and configures the sensor to produce subsampled, lower-precision image data. This vision mode can save ~75% of the average energy of a baseline photography mode with only a small impact on vision task accuracy.
Mark Buckler, Suren Jayasuriya, Adrian Sampson
ICCV2
2016 ASP Vision: Optically Computing the First Layer of Convolutional Neural Networks Using Angle Sensitive Pixels
abstract
Deep learning using convolutional neural networks (CNNs) is quickly becoming the state-of-the-art for challenging computer vision applications. However, deep learning's power consumption and bandwidth requirements currently limit its application in embedded and mobile systems with tight energy budgets. In this paper, we explore the energy savings of optically computing the first layer of CNNs. To do so, we utilize bio-inspired Angle Sensitive Pixels (ASPs), custom CMOS diffractive image sensors which act similar to Gabor filter banks in the V1 layer of the human visual cortex. ASPs replace both image sensing and the first layer of a conventional CNN by directly performing optical edge filtering, saving sensing energy, data bandwidth, and CNN FLOPS to compute. Our experimental results (both on synthetic data and a hardware prototype) for a variety of vision tasks such as digit recognition, object recognition, and face identification demonstrate 97% reduction in image sensor power consumption and 90% reduction in data bandwidth from sensor to CPU, while achieving similar performance compared to traditional deep learning pipelines.
Huaijin G. Chen, Suren Jayasuriya, Jiyue Yang, Judy Stephen, Sriram Sivaramakrishnan, Ashok Veeraraghavan, Alyosha C. Molnar
CVPR2
2016 Experiences using a novel Python-based hardware modeling framework for computer architecture test chips
abstract
This poster will describe a taped-out 2×2mm 1.3 M-transistor test chip in IBM 130 nm designed using our new Python-based hardware modeling framework. The goal of our tapeout was to demonstrate the ability of this framework to enable Agile hardware design flows.
Christopher Torng, Moyang Wang, Bharath Sudheendra, Nagaraj Murali, Suren Jayasuriya, Shreesha Srinath, Taylor Pritchard, Robin Ying, Christopher Batten
Hot Chips Symposium5
2015 Depth Fields: Extending Light Field Techniques to Time-of-Flight Imaging
abstract
A variety of techniques such as light field, structured illumination, and time-of-flight (TOF) are commonly used for depth acquisition in consumer imaging, robotics and many other applications. Unfortunately, each technique suffers from its individual limitations preventing robust depth sensing. In this paper, we explore the strengths and weaknesses of combining light field and time-of-flight imaging, particularly the feasibility of an on-chip implementation as a single hybrid depth sensor. We refer to this combination as depth field imaging. Depth fields combine light field advantages such as synthetic aperture refocusing with TOF imaging advantages such as high depth resolution and coded signal processing to resolve multipath interference. We show applications including synthesizing virtual apertures for TOF imaging, improved depth mapping through partial and scattering occluders, and single frequency TOF phase unwrapping. Utilizing space, angle, and temporal coding, depth fields can improve depth sensing in the wild and generate new insights into the dimensions of light's plenoptic function.
Suren Jayasuriya, Adithya Kumar Pediredla, Sriram Sivaramakrishnan, Alyosha C. Molnar, Ashok Veeraraghavan
3DV1
2014 A switchable light field camera architecture with Angle Sensitive Pixels and dictionary-based sparse coding
abstract
We propose a flexible light field camera architecture that is at the convergence of optics, sensor electronics, and applied mathematics. Through the co-design of a sensor that comprises tailored, Angle Sensitive Pixels and advanced reconstruction algorithms, we show that—contrary to light field cameras today—our system can use the same measurements captured in a single sensor image to recover either a high-resolution 2D image, a low-resolution 4D light field using fast, linear processing, or a high-resolution light field using sparsity-constrained optimization.
Matthew Hirsch, Sriram Sivaramakrishnan, Suren Jayasuriya, Albert Wang 0004, Alyosha C. Molnar, Ramesh Raskar, Gordon Wetzstein
ICCP3