Julien N. P. Martel

dblp:150/2876 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
8since 2021 · last 2022
0000-0002-4928-5487ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Systems, architecture and hardware · 9 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 CryoAI: Amortized Inference of Poses for Ab Initio Reconstruction of 3D Molecular Volumes from Real Cryo-EM Images
Axel Levy, Frédéric Poitevin, Julien N. P. Martel, Youssef S. G. Nashed, Ariana Peck, Nina Miolane, Daniel Ratner, Mike Dunne, Gordon Wetzstein
ECCV (21)3
2022 Learning Spatially Varying Pixel Exposures for Motion Deblurring
abstract
Computationally removing the motion blur introduced by camera shake or object motion in a captured image remains a challenging task in computational photography. Deblurring methods are often limited by the fixed global exposure time of the image capture process. The post-processing algorithm either must deblur a longer exposure that contains relatively little noise or denoise a short exposure that intentionally removes the opportunity for blur at the cost of increased noise. We present a novel approach of leveraging spatially varying pixel exposures for motion deblurring using next-generation focal-plane sensor-processors along with an end-to-end design of these exposures and a machine learning-based motion-deblurring framework. We demonstrate in simulation and a physical prototype that learned spatially varying pixel exposures (L-SVPE) can successfully deblur scenes while recovering high frequency detail. Our work illustrates the promising role that focal-plane sensor-processors can play in the future of computational imaging.
Cindy M. Nguyen, Julien N. P. Martel, Gordon Wetzstein
ICCP2
2022 MantissaCam: Learning Snapshot High-dynamic-range Imaging with Perceptually-based In-pixel Irradiance Encoding
abstract
The ability to image high-dynamic-range (HDR) scenes is crucial in many computer vision applications. The dynamic range of conventional sensors, however, is fundamentally limited by their well capacity, resulting in saturation of bright scene parts. To overcome this limitation, emerging sensors offer in-pixel processing capabilities to encode the incident irradiance. Among the most promising encoding schemes is modulo wrapping, which results in a computational photography problem where the HDR scene is computed by an irradiance unwrapping algorithm from the wrapped low-dynamic-range (LDR) sensor image. Here, we design a neural network-based algorithm that outperforms previous irradiance unwrapping methods and we design a perceptually inspired “mantissa,” or log-modulo, encoding scheme that more efficiently wraps an HDR scene into an LDR sensor. Combined with our reconstruction framework, MantissaCam achieves state-of-the-art results among modulo-type snapshot HDR imaging approaches. We demonstrate the efficacy of our method in simulation and show benefits of our algorithm on modulo images captured with a prototype implemented with a programmable sensor.
Haley M. So, Julien N. P. Martel, Gordon Wetzstein, Piotr Dudek
ICCP2
2022 Amortized Inference for Heterogeneous Reconstruction in Cryo-EM
abstract
Cryo-electron microscopy (cryo-EM) is an imaging modality that provides unique insights into the dynamics of proteins and other building blocks of life. The algorithmic challenge of jointly estimating the poses, 3D structure, and conformational heterogeneity of a biomolecule from millions of noisy and randomly oriented 2D projections in a computationally efficient manner, however, remains unsolved. Our method, cryoFIRE, performs ab initio heterogeneous reconstruction with unknown poses in an amortized framework, thereby avoiding the computationally expensive step of pose search while enabling the analysis of conformational heterogeneity. Poses and conformation are jointly estimated by an encoder while a physics-based decoder aggregates the images into an implicit neural representation of the conformational space. We show that our method can provide one order of magnitude speedup on datasets containing millions of images, without any loss of accuracy. We validate that the joint estimation of poses and conformations can be amortized over the size of the dataset. For the first time, we prove that an amortized method can extract interpretable dynamic information from experimental datasets.
Axel Levy, Gordon Wetzstein, Julien N. P. Martel, Frédéric Poitevin, Ellen D. Zhong
NeurIPS3
2021 AutoInt: Automatic Integration for Fast Neural Volume Rendering
abstract
Numerical integration is a foundational technique in scientific computing and is at the core of many computer vision applications. Among these applications, neural volume rendering has recently been proposed as a new paradigm for view synthesis, achieving photorealistic image quality. However, a fundamental obstacle to making these methods practical is the extreme computational and memory requirements caused by the required volume integrations along the rendered rays during training and inference. Millions of rays, each requiring hundreds of forward passes through a neural network are needed to approximate those integrations with Monte Carlo sampling. Here, we propose automatic integration, a new framework for learning efficient, closed-form solutions to integrals using coordinate-based neural networks. For training, we instantiate the computational graph corresponding to the derivative of the coordinate-based network. The graph is fitted to the signal to integrate. After optimization, we reassemble the graph to obtain a network that represents the antiderivative. By the fundamental theorem of calculus, this enables the calculation of any definite integral in two evaluations of the network. Applying this approach to neural rendering, we improve a tradeoff between rendering speed and image quality: improving render times by greater than 10× with a tradeoff of reduced image quality.
David B. Lindell, Julien N. P. Martel, Gordon Wetzstein
CVPR2
2021 Time-Multiplexed Coded Aperture Imaging: Learned Coded Aperture and Pixel Exposures for Compressive Imaging Systems
abstract
Compressive imaging using coded apertures (CA) is a powerful technique that can be used to recover depth, light fields, hyperspectral images and other quantities from a single snapshot. The performance of compressive imaging systems based on CAs mostly depends on two factors: the properties of the mask's attenuation pattern, that we refer to as "codification", and the computational techniques used to recover the quantity of interest from the coded snapshot. In this work, we introduce the idea of using time-varying CAs synchronized with spatially varying pixel shutters. We divide the exposure of a sensor into sub-exposures at the beginning of which the CA mask changes and at which the sensor's pixels are simultaneously and individually switched "on" or "off". This is a practically appealing codification as it does not introduce additional optical components other than the already present CA but uses a change in the pixel shutter that can be easily realized electronically. We show that our proposed time-multiplexed coded aperture (TMCA) can be optimized end to end and induces better coded snapshots enabling superior reconstructions in two different applications: compressive light field imaging and hyperspectral imaging. We demonstrate both in simulation and with real captures (taken with prototypes we built) that this codification outperforms the state-of-the-art compressive imaging systems by a large margin in those applications.
Edwin Vargas, Julien N. P. Martel, Gordon Wetzstein, Henry Arguello
ICCV2
2021 Acorn: adaptive coordinate networks for neural scene representation
abstract
Neural representations have emerged as a new paradigm for applications in rendering, imaging, geometric modeling, and simulation. Compared to traditional representations such as meshes, point clouds, or volumes they can be flexibly incorporated into differentiable learning-based pipelines. While recent improvements to neural representations now make it possible to represent signals with fine details at moderate resolutions (e.g., for images and 3D shapes), adequately representing large-scale or complex scenes has proven a challenge. Current neural representations fail to accurately represent images at resolutions greater than a megapixel or 3D scenes with more than a few hundred thousand polygons. Here, we introduce a new hybrid implicit-explicit network architecture and training strategy that adaptively allocates resources during training and inference based on the local complexity of a signal of interest. Our approach uses a multiscale block-coordinate decomposition, similar to a quadtree or octree, that is optimized during training. The network architecture operates in two stages: using the bulk of the network parameters, a coordinate encoder generates a feature grid in a single forward pass. Then, hundreds or thousands of samples within each block can be efficiently evaluated using a lightweight feature decoder. With this hybrid implicit-explicit network architecture, we demonstrate the first experiments that fit gigapixel images to nearly 40 dB peak signal-to-noise ratio. Notably this represents an increase in scale of over 1000X compared to the resolution of previously demonstrated image-fitting experiments. Moreover, our approach is able to represent 3D shapes significantly faster and better than previous techniques; it reduces training times from days to hours or minutes and memory requirements by over an order of magnitude.
Julien N. P. Martel, David B. Lindell, Connor Z. Lin, Eric R. Chan, Marco Monteiro, Gordon Wetzstein
ACM Trans. Graph.1
2021 Event-Based Near-Eye Gaze Tracking Beyond 10, 000 Hz
abstract
The cameras in modern gaze-tracking systems suffer from fundamental bandwidth and power limitations, constraining data acquisition speed to 300 Hz realistically. This obstructs the use of mobile eye trackers to perform, e.g., low latency predictive rendering, or to study quick and subtle eye motions like microsaccades using head-mounted devices in the wild. Here, we propose a hybrid frame-event-based near-eye gaze tracking system offering update rates beyond 10,000 Hz with an accuracy that matches that of high-end desktop-mounted commercial trackers when evaluated in the same conditions. Our system, previewed in Figure 1, builds on emerging event cameras that simultaneously acquire regularly sampled frames and adaptively sampled events. We develop an online 2D pupil fitting method that updates a parametric model every one or few events. Moreover, we propose a polynomial regressor for estimating the point of gaze from the parametric pupil model in real time. Using the first event-based gaze dataset, we demonstrate that our system achieves accuracies of 0.45°-1.75° for fields of view from 45° to 98°. With this technology, we hope to enable a new generation of ultra-low-latency gaze-contingent rendering and display techniques for virtual and augmented reality.
Anastasios Angelopoulos, Julien N. P. Martel, Amit P. S. Kohli, Jörg Conradt, Gordon Wetzstein
IEEE Trans. Vis. Comput. Graph.2
2020 Implicit Neural Representations with Periodic Activation Functions
abstract
Implicitly defined, continuous, differentiable signal representations parameterized by neural networks have emerged as a powerful paradigm, offering many possible benefits over conventional representations. However, current network architectures for such implicit neural representations are incapable of modeling signals with fine detail, and fail to represent a signal's spatial and temporal derivatives, despite the fact that these are essential to many physical signals defined implicitly as the solution to partial differential equations. We propose to leverage periodic activation functions for implicit neural representations and demonstrate that these networks, dubbed sinusoidal representation networks or SIRENs, are ideally suited for representing complex natural signals and their derivatives. We analyze SIREN activation statistics to propose a principled initialization scheme and demonstrate the representation of images, wavefields, video, sound, and their derivatives. Further, we show how SIRENs can be leveraged to solve challenging boundary value problems, such as particular Eikonal equations (yielding signed distance functions), the Poisson equation, and the Helmholtz and wave equations. Lastly, we combine SIRENs with hypernetworks to learn priors over the space of SIREN functions.
Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell, Gordon Wetzstein
NeurIPS2
2020 Neural Sensors: Learning Pixel Exposures for HDR Imaging and Video Compressive Sensing With Programmable Sensors
abstract
Camera sensors rely on global or rolling shutter functions to expose an image. This fixed function approach severely limits the sensors' ability to capture high-dynamic-range (HDR) scenes and resolve high-speed dynamics. Spatially varying pixel exposures have been introduced as a powerful computational photography approach to optically encode irradiance on a sensor and computationally recover additional information of a scene, but existing approaches rely on heuristic coding schemes and bulky spatial light modulators to optically implement these exposure functions. Here, we introduce neural sensors as a methodology to optimize per-pixel shutter functions jointly with a differentiable image processing method, such as a neural network, in an end-to-end fashion. Moreover, we demonstrate how to leverage emerging programmable and re-configurable sensor-processors to implement the optimized exposure functions directly on the sensor. Our system takes specific limitations of the sensor into account to optimize physically feasible optical codes and we evaluate its performance for snapshot HDR and high-speed compressive imaging both in simulation and experimentally with real scenes.
Julien N. P. Martel, Lorenz K. Müller, Stephen J. Carey, Piotr Dudek, Gordon Wetzstein
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Adaptive motor control and learning in a spiking neural network realised on a mixed-signal neuromorphic processor
abstract
Neuromorphic computing is a new paradigm for design of both the computing hardware and algorithms inspired by biological neural networks. The event-based nature and the inherent parallelism make neuromorphic computing a promising paradigm for building efficient neural network based architectures for control of fast and agile robots. In this paper, we present a spiking neural network architecture that uses sensory feedback to control rotational velocity of a robotic vehicle. When the velocity reaches the target value, the mapping from the target velocity of the vehicle to the correct motor command, both represented in the spiking neural network on the neuromorphic device, is autonomously stored on the device using on-chip plastic synaptic weights. We validate the controller using a wheel motor of a miniature mobile vehicle and inertia measurement unit as the sensory feedback and demonstrate online learning of a simple “inverse model” in a two-layer spiking neural network on the neuromorphic chip. The prototype neuromorphic device that features 256 spiking neurons allows us to realise a simple proof of concept architecture for the purely neuromorphic motor control and learning. The architecture can be easily scaled-up if a larger neuromorphic device is available.
Sebastian Glatz, Julien N. P. Martel, Raphaela Kreiser, Yulia Sandamirskaya
ICRA2
2018 Kernelized Synaptic Weight Matrices
abstract
In this paper we introduce a novel neural network architecture, in which weight matrices are re-parametrized in terms of low-dimensional vectors, interacting through kernel functions. A layer of our network can be interpreted as introducing a (potentially infinitely wide) linear layer between input and output. We describe the theory underpinning this model and validate it with concrete examples, exploring how it can be used to impose structure on neural networks in diverse applications ranging from data visualization to recommender systems. We achieve state-of-the-art performance in a collaborative filtering task (MovieLens).
Lorenz K. Müller, Julien N. P. Martel, Giacomo Indiveri
ICML2
2018 A Neuromorphic Approach to Path Integration: A Head-Direction Spiking Neural Network with Vision-driven Reset
abstract
Simultaneous localization and mapping (SLAM) is one of the core tasks of mobile autonomous robots. Looking for power efficient and embedded solutions for SLAM is an important challenge when building controllers for small and agile robots. Biological neural systems of even simple animals are until now unprecedented in their ability to localize themselves in an unknown environment. Neuromorphic engineering offers ultra low-power and compact computing hardware, in which biologically inspired neuronal architectures for SLAM can be realised. In this paper, we propose an on chip approach for one of the components of SLAM: path integration. Our solution takes inspiration from biology and uses motor command information to estimate the orientation of an agent solely in a spiking neural network. We realise this network on a neuromorphic device that implements artificial neurons and synapses with analog electronics. The neural network receives visual input from an event-based camera and uses this information to correct the on-chip spiking neurons estimate of the robot's orientation. This system can be easily integrated with other localization and mapping components on chip and is a step towards a fully neuromorphic SLAM.
Raphaela Kreiser, Matteo Cartiglia, Julien N. P. Martel, Jörg Conradt, Yulia Sandamirskaya
ISCAS3
2018 An Active Approach to Solving the Stereo Matching Problem using Event-Based Sensors
abstract
The problem of inferring distances from a visual sensor to objects in a scene - referred to as depth estimation - can be solved in various ways. Among those, stereo vision is a method in which two sensors observe the same scene from different viewpoints. To recover the three-dimensional coordinates of a point, its two projections - one in each view - can be used for triangulation. However, the pair of points in the two views that correspond to each other has to be found first. This is known as stereo-matching and is usually a computationally expensive operation. Traditionally, this is performed by describing a point in the first view with some information from its surrounding, e.g. in a feature vector, and then searching for a match with a point described in a similar way in the other view. In this work, we propose a simple idea that alleviates this stereo-matching problem using an active component: a mirror-galvanometer driven laser. The laser beam is deflected by actuating two mirrors, thus creating a sequence of "light spots" in the scene. At these spots, contrast changes quickly. We capture those contrast changes by two Dynamic Vision Sensors (DVS). The high time-resolution of these sensors enables the detection of the laser-induced events in time and their matching using lightweight computation. This method enables event-based depth estimation at a high speed, low computational cost, and without exact sensor synchronization.
Julien N. P. Martel, Jonathan Müller, Jörg Conradt, Yulia Sandamirskaya
ISCAS1
2018 Live Demonstration: An Active System for Depth Reconstruction using Event-Based Sensors
abstract
We demonstrate a system that can reconstruct the three-dimensional structure of a scene using two Dynamic Vision Sensors (DVS) and an active component: a Mirror Galvanometer-driven Laser (MGDL). Our system uses two concurrent methods to estimate depth: 1) triangulation from the two DVS cameras calibrated as stereo rig, where stereo matching is greatly simplified by actively generating events in the scene with the laser beam, and 2) a technique derived from structured light, where we send the laser beam in a known direction and detect its projection in each of the two cameras. For the demonstration, we add two interactive components: the user can trigger (re)scanning of a specific location in the scene by pointing a hand-held laser pointer to it, and virtual objects can be rendered with correct scaling and distance on a screen using the reconstructed depth.
Julien N. P. Martel, Jonathan Müller, Jörg Conradt, Yulia Sandamirskaya
ISCAS1
2017 High-speed depth from focus on a programmable vision chip using a focus tunable lens
abstract
In this paper, we present a 3D imaging system providing a semi-dense depth map, using a passive, low-power, compact, static, monocular camera. The demonstrated depth estimation system reconstructs 32 depth-levels in real-time at 25FPS drawing less than 1.9W of power. This is achieved by performing computation on an analog focal-plane processor that analyses frames captured through a vibrating liquid focus-tunable lens. The optical system provides shallow depth of focus images and fast sweeps of optical power, while the use of pixel-level processing removes the sensor-processor bandwidth limitations of depth-from-focus systems built using conventional imaging and processor technologies. All-in-focus images are also obtained.
Julien N. P. Martel, Lorenz K. Müller, Stephen J. Carey, Piotr Dudek
ISCAS1
2017 Live demonstration: Depth from focus on a focal plane processor using a focus tunable liquid lens
abstract
We demonstrate a 3D imaging system that produces sparse depth maps. It consists in a liquid focus-tunable lens whose focal power can be changed at high speed, placed in front of a SCAMP5 vision-chip embedding processing capabilities in each pixel. The focus-tunable lens performs focal sweeps with shallow depth of fields. These are sampled by the vision chip taking multiple images at different focus and analyzed on-chip to produce a single depth frame. The combination of the focus tunable-lens with the vision-chip, enabling near-focal plane processing, allows us to present a compact passive system that is static, monocular, real-time (> 25FPS) and low-power (<; 1.6W).
Julien N. P. Martel, Lorenz K. Müller, Stephen J. Carey, Jonathan Müller, Yulia Sandamirskaya, Piotr Dudek
ISCAS1
2017 Color temporal contrast sensitivity in dynamic vision sensors
abstract
This paper introduces the first simulations and measurements of event data obtained from the first Dynamic and Active Vision Sensors (DAVIS) with RGBW color filters. The absolute quantum efficiency spectral responses of the RGBW photodiodes were measured, the behavior of the color-sensitive DVS pixels were simulated and measured, and reconstruction through color events interpolation was developed.
Diederik Paul Moeys, Cheng-Han Li, Julien N. P. Martel, Simeon A. Bamford, Luca Longinotti, Vasyl Motsnyi, David San Segundo Bello, Tobi Delbruck
ISCAS3
2016 Parallel HDR tone mapping and auto-focus on a cellular processor array vision chip
abstract
To improve computational efficiency, it may be advantageous to transfer part of the intelligence lying in the core of a system to its sensors. Vision sensors equipped with small programmable processors at each pixel allow us to follow this principle in so-called near-focal plane processing, which is performed on-chip directly where light is being collected. Such devices need then only to communicate relevant pre-processed visual information to other parts of the system. In this work, we demonstrate how two classical problems, namely high dynamic range imaging and auto-focus, can be solved efficiently using two simple parallel algorithms implemented on such a chip. We illustrate with these two examples that embedding uncomplicated algorithms on-chip, directly where information acquisition takes place can replace more complex dedicated post-processing. Adapting data acquisition by bringing processing at the sensor level allows us to explore solutions that would not be feasible in a conventional sensor-ADC-processor pipeline.
Julien N. P. Martel, Lorenz K. Müller, Stephen J. Carey, Piotr Dudek
ISCAS1
2015 Toward joint approximate inference of visual quantities on cellular processor arrays
abstract
The interacting visual maps (IVM) algorithm introduced in [1] is able to perform the joint approximate inference of several visual quantities such as optic-flow, gray-level intensities and ego-motion, using a sparse input coming from a neuromorphic dynamic vision sensor (DVS). We show that features of the model such as the intrinsic parallelism and distributed nature of its computation make it a natural candidate to benefit from the cellular processor array (CPA) hardware architecture. We have now implemented the IVM algorithm on a general-purpose CPA simulator, and here we present results of our simulations and demonstrate that the IVM algorithm indeed naturally fits the CPA architecture. Our work indicates that extended versions of the IVM algorithm could benefit greatly from a dedicated hardware implementation, eventually yielding a high speed, low power visual odometry chip.
Julien N. P. Martel, Miguel Chau, Piotr Dudek, Matthew Cook 0001
ISCAS1
2014 Candidate Sampling for Neuron Reconstruction from Anisotropic Electron Microscopy Volumes
Jan Funke, Julien N. P. Martel, Stephan Gerhard, Bjoern Andres, Dan C. Ciresan, Alessandro Giusti, Luca Maria Gambardella, Jürgen Schmidhuber, Hanspeter Pfister, Albert Cardona, Matthew Cook 0001
MICCAI (1)2
2013 A Combination of Hand-Crafted and Hierarchical High-Level Learnt Feature Extraction for Music Genre Classification
Julien N. P. Martel, Toru Nakashika, Christophe Garcia, Khalid Idrissi
ICANN1