Stanley J. Osher

dblp:24/5359 · DBLP profile ↗
← Back
81ranked-venue papers
2as first author
20since 2021 · last 2025
0000-0002-7900-4658ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 35 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 5Computer networks · 4 · 3 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
abstract
We study Transformers through the perspective of optimal control theory, using tools from continuous-time formulations to derive actionable insights into training and architecture design. This framework improves the performance of existing Transformer models while providing desirable theoretical guarantees, including generalization and robustness. Our framework is designed to be plug-and-play, enabling seamless integration with established Transformer models and requiring only slight changes to the implementation. We conduct seven extensive experiments on tasks motivated by text generation, sentiment analysis, image classification, and point cloud classification. Experimental results show that the framework improves the test performance of the baselines, while being more parameter-efficient. On character-level text generation with nanoGPT, our framework achieves a 46\% reduction in final test loss while using 42\% fewer parameters. On GPT-2, our framework achieves a 9.3\% reduction in final test loss, demonstrating scalability to larger models. To the best of our knowledge, this is the first work that applies optimal control theory to both the training and architecture of Transformers. It offers a new foundation for systematic, theory-driven improvements and moves beyond costly trial-and-error approaches.
Kelvin Kan, Xingjian Li 0005, Benjamin J. Zhang, Tuhin Sahai, Stanley J. Osher, Markos A. Katsoulakis
NeurIPS5
2025 Laplace Meets Moreau: Smooth Approximation to Infimal Convolutions Using Laplace's Method
abstract
We study approximations to the Moreau envelope---and infimal convolutions more broadly---based on Laplace's method, a classical tool in analysis which ties certain integrals to suprema of their integrands. We believe the connection between Laplace's method and infimal convolutions is generally deserving of more attention in the study of optimization and partial differential equations, since it bears numerous potentially important applications, from proximal-type algorithms to Hamilton-Jacobi equations.
Ryan J. Tibshirani, Samy Wu Fung, Howard Heaton, Stanley J. Osher
J. Mach. Learn. Res.4
2025 Fine-tune language models as multi-modal differential equation solvers
abstract
In the growing domain of scientific machine learning, in-context operator learning has shown notable potential in building foundation models, as in this framework the model is trained to learn operators and solve differential equations using prompted data, during the inference stage without weight updates. However, the current model's overdependence on function data overlooks the invaluable human insight into the operator. To address this, we present a transformation of in-context operator learning into a multi-modal paradigm. In particular, we take inspiration from the recent success of large language models, and propose using "captions" to integrate human knowledge about the operator, expressed through natural language descriptions and equations. Also, we introduce a novel approach to train a language-model-like architecture, or directly fine-tune existing language models, for in-context operator learning. We beat the baseline on single-modal learning tasks, and also demonstrated the effectiveness of multi-modal learning in enhancing performance and reducing function data requirements. The proposed method not only significantly enhanced the development of the in-context operator learning paradigm, but also created a new path for the application of language models.
Liu Yang 0027, Siting Liu 0003, Stanley J. Osher
Neural Networks3
2024 Low-Light Phase Retrieval With Implicit Generative Priors
abstract
Phase retrieval (PR) is fundamentally important in scientific imaging and is crucial for nanoscale techniques like coherent diffractive imaging (CDI). Low radiation dose imaging is essential for applications involving radiation-sensitive samples. However, most PR methods struggle in low-dose scenarios due to high shot noise. Recent advancements in optical data acquisition setups, such as in-situ CDI, have shown promise for low-dose imaging, but they rely on a time series of measurements, making them unsuitable for single-image applications. Similarly, data-driven phase retrieval techniques are not easily adaptable to data-scarce situations. Zero-shot deep learning methods based on pre-trained and implicit generative priors have been effective in various imaging tasks but have shown limited success in PR. In this work, we propose low-dose deep image prior (LoDIP), which combines in-situ CDI with the power of implicit generative priors to address single-image low-dose phase retrieval. Quantitative evaluations demonstrate LoDIP's superior performance in this task and its applicability to real experimental scenarios.
Raunak Manekar, Elisa Negrini, Minh Pham 0003, Daniel Jacobs, Jaideep Srivastava, Stanley J. Osher, Jianwei Miao
IEEE Trans. Image Process.6
2023 A Probabilistic Framework for Pruning Transformers Via a Finite Admixture of Keys
abstract
Pairwise dot product-based self-attention is key to the success of transformers which achieve state-of-the-art performance across a variety of applications in language and vision, but are costly to compute. It has been shown that most attention scores and keys in transformers are redundant and can be removed without loss of accuracy. In this paper, we develop a novel probabilistic framework for pruning attention scores and keys in transformers. We first formulate an admixture model of attention keys whose input data to be clustered are attention queries. We show that attention scores in self-attention correspond to the posterior distribution of this model when attention keys admit a uniform prior distribution. We then relax this uniform prior constraint and let the model learn these priors from data, resulting in a new Finite Admixture of Keys (FiAK). The learned priors are used for pruning away redundant attention scores and keys in the baseline transformers, improving the diversity of attention patterns that the models capture. We corroborate the efficiency of transformers pruned with FiAK on the ImageNet object classification and WikiText-103 language modeling tasks. Our experiments demonstrate that transformers pruned with FiAK yield similar or better accuracy than the baseline dense transformers while being much more efficient in terms of memory and computational cost.
Tan M. Nguyen, Long Bui, Hai Do, Duy Khuong Nguyen, Dung D. Le, Hung Tran-The, Nhat Ho, Stanley J. Osher, Richard G. Baraniuk
ICASSP9
2023 A Primal-Dual Framework for Transformers and Neural Networks
Tan M. Nguyen, Tam Minh Nguyen, Nhat Ho, Andrea L. Bertozzi, Richard G. Baraniuk, Stanley J. Osher
ICLR6
2023 Neural Collapse in Deep Linear Networks: From Balanced to Imbalanced Data
abstract
Modern deep neural networks have achieved impressive performance on tasks from image classification to natural language processing. Surprisingly, these complex systems with massive amounts of parameters exhibit the same structural properties in their last-layer features and classifiers across canonical datasets when training until convergence. In particular, it has been observed that the last-layer features collapse to their class-means, and those class-means are the vertices of a simplex Equiangular Tight Frame (ETF). This phenomenon is known as Neural Collapse (NC). Recent papers have theoretically shown that NC emerges in the global minimizers of training problems with the simplified ``unconstrained feature model''. In this context, we take a step further and prove the NC occurrences in deep linear networks for the popular mean squared error (MSE) and cross entropy (CE) losses, showing that global solutions exhibit NC properties across the linear layers. Furthermore, we extend our study to imbalanced data for MSE loss and present the first geometric analysis of NC under bias-free setting. Our results demonstrate the convergence of the last-layer features and classifiers to a geometry consisting of orthogonal vectors, whose lengths depend on the amount of data in their corresponding classes. Finally, we empirically validate our theoretical analyses on synthetic and practical network architectures with both balanced and imbalanced scenarios.
Hien Dang 0003, Tho Tran, Stanley J. Osher, Hung Tran-The, Nhat Ho, Tan M. Nguyen
ICML3
2023 Revisiting Over-smoothing and Over-squashing Using Ollivier-Ricci Curvature
abstract
Graph Neural Networks (GNNs) had been demonstrated to be inherently susceptible to the problems of over-smoothing and over-squashing. These issues prohibit the ability of GNNs to model complex graph interactions by limiting their effectiveness in taking into account distant information. Our study reveals the key connection between the local graph geometry and the occurrence of both of these issues, thereby providing a unified framework for studying them at a local scale using the Ollivier-Ricci curvature. Specifically, we demonstrate that over-smoothing is linked to positive graph curvature while over-squashing is linked to negative graph curvature. Based on our theory, we propose the Batch Ollivier-Ricci Flow, a novel rewiring algorithm capable of simultaneously addressing both over-smoothing and over-squashing.
Khang Nguyen 0002, Nong Minh Hieu, Vinh Duc Nguyen, Nhat Ho, Stanley J. Osher, Tan M. Nguyen
ICML5
2022 JFB: Jacobian-Free Backpropagation for Implicit Networks
abstract
A promising trend in deep learning replaces traditional feedforward networks with implicit networks. Unlike traditional networks, implicit networks solve a fixed point equation to compute inferences. Solving for the fixed point varies in complexity, depending on provided data and an error tolerance. Importantly, implicit networks may be trained with fixed memory costs in stark contrast to feedforward networks, whose memory requirements scale linearly with depth. However, there is no free lunch --- backpropagation through implicit networks often requires solving a costly Jacobian-based equation arising from the implicit function theorem. We propose Jacobian-Free Backpropagation (JFB), a fixed-memory approach that circumvents the need to solve Jacobian-based equations. JFB makes implicit networks faster to train and significantly easier to implement, without sacrificing test accuracy. Our experiments show implicit networks trained with JFB are competitive with feedforward networks and prior implicit networks given the same number of parameters.
Samy Wu Fung, Howard Heaton, Qiuwei Li, Daniel McKenzie, Stanley J. Osher, Wotao Yin
AAAI5
2022 GRAND++: Graph Neural Diffusion with A Source Term
Matthew Thorpe, Tan M. Nguyen, Hedi Xia, Thomas Strohmer, Andrea L. Bertozzi, Stanley J. Osher, Bao Wang 0001
ICLR6
2022 Improving Transformers with Probabilistic Attention Keys
abstract
Multi-head attention is a driving force behind state-of-the-art transformers, which achieve remarkable performance across a variety of natural language processing (NLP) and computer vision tasks. It has been observed that for many applications, those attention heads learn redundant embedding, and most of them can be removed without degrading the performance of the model. Inspired by this observation, we propose Transformer with a Mixture of Gaussian Keys (Transformer-MGK), a novel transformer architecture that replaces redundant heads in transformers with a mixture of keys at each head. These mixtures of keys follow a Gaussian mixture model and allow each attention head to focus on different parts of the input sequence efficiently. Compared to its conventional transformer counterpart, Transformer-MGK accelerates training and inference, has fewer parameters, and requires fewer FLOPs to compute while achieving comparable or better accuracy across tasks. Transformer-MGK can also be easily extended to use with linear attention. We empirically demonstrate the advantage of Transformer-MGK in a range of practical applications, including language modeling and tasks that involve very long sequences. On the Wikitext-103 and Long Range Arena benchmark, Transformer-MGKs with 4 heads attain comparable or better performance to the baseline transformers with 8 heads.
Tam Minh Nguyen, Tan M. Nguyen, Dung D. Le, Duy Khuong Nguyen, Viet-Anh Tran, Richard G. Baraniuk, Nhat Ho, Stanley J. Osher
ICML8
2022 FourierFormer: Transformer Meets Generalized Fourier Integral Theorem
abstract
Multi-head attention empowers the recent success of transformers, the state-of-the-art models that have achieved remarkable success in sequence modeling and beyond. These attention mechanisms compute the pairwise dot products between the queries and keys, which results from the use of unnormalized Gaussian kernels with the assumption that the queries follow a mixture of Gaussian distribution. There is no guarantee that this assumption is valid in practice. In response, we first interpret attention in transformers as a nonparametric kernel regression. We then propose the FourierFormer, a new class of transformers in which the dot-product kernels are replaced by the novel generalized Fourier integral kernels. Different from the dot-product kernels, where we need to choose a good covariance matrix to capture the dependency of the features of data, the generalized Fourier integral kernels can automatically capture such dependency and remove the need to tune the covariance matrix. We theoretically prove that our proposed Fourier integral kernels can efficiently approximate any key and query distributions. Compared to the conventional transformers with dot-product attention, FourierFormers attain better accuracy and reduce the redundancy between attention heads. We empirically corroborate the advantages of FourierFormers over the baseline transformers in a variety of practical applications including language modeling and image classification.
Minh Pham 0003, Stanley J. Osher, Nhat Ho
NeurIPS5
2022 Improving Transformer with an Admixture of Attention Heads
abstract
Transformers with multi-head self-attention have achieved remarkable success in sequence modeling and beyond. However, they suffer from high computational and memory complexities for computing the attention matrix at each head. Recently, it has been shown that those attention matrices lie on a low-dimensional manifold and, thus, are redundant. We propose the Transformer with a Finite Admixture of Shared Heads (FiSHformers), a novel class of efficient and flexible transformers that allow the sharing of attention matrices between attention heads. At the core of FiSHformer is a novel finite admixture model of shared heads (FiSH) that samples attention matrices from a set of global attention matrices. The number of global attention matrices is much smaller than the number of local attention matrices generated. FiSHformers directly learn these global attention matrices rather than the local ones as in other transformers, thus significantly improving the computational and memory efficiency of the model. We empirically verify the advantages of the FiSHformer over the baseline transformers in a wide range of practical applications including language modeling, machine translation, and image classification. On the WikiText-103, IWSLT'14 De-En and WMT'14 En-De, FiSHformers use much fewer floating-point operations per second (FLOPs), memory, and parameters compared to the baseline transformers.
Hai Do, Vishwanath Saragadam, Minh Pham 0003, Duy Khuong Nguyen, Nhat Ho, Stanley J. Osher
NeurIPS9
2022 Improving Neural Ordinary Differential Equations with Nesterov's Accelerated Gradient Method
abstract
We propose the Nesterov neural ordinary differential equations (NesterovNODEs), whose layers solve the second-order ordinary differential equations (ODEs) limit of Nesterov's accelerated gradient (NAG) method, and a generalization called GNesterovNODEs. Taking the advantage of the convergence rate $\mathcal{O}(1/k^{2})$ of the NAG scheme, GNesterovNODEs speed up training and inference by reducing the number of function evaluations (NFEs) needed to solve the ODEs. We also prove that the adjoint state of a GNesterovNODEs also satisfies a GNesterovNODEs, thus accelerating both forward and backward ODE solvers and allowing the model to be scaled up for large-scale tasks. We empirically corroborate the advantage of GNesterovNODEs on a wide range of practical applications, including point cloud separation, image classification, and sequence modeling. Compared to NODEs, GNesterovNODEs require a significantly smaller number of NFEs while achieving better accuracy across our experiments.
Ho Huu Nghia Nguyen, Huyen Vo, Stanley J. Osher, Thieu Vo
NeurIPS4
2022 Scheduled Restart Momentum for Accelerated Stochastic Gradient Descent
abstract
Stochastic gradient descent (SGD) algorithms, with constant momentum and its variants such as Adam, are the optimization methods of choice for training deep neural networks (DNNs). There is great interest in speeding up the convergence of these methods due to their high computational expense. Nesterov accelerated gradient with a time-varying momentum (NAG) improves the convergence rate of gradient descent for convex optimization using a specially designed momentum; however, it accumulates error when the stochastic gradient is used, slowing convergence at best and diverging at worst. In this paper, we propose scheduled restart SGD (SRSGD), a new NAG-style scheme for training DNNs. SRSGD replaces the constant momentum in SGD by the increasing momentum in NAG but stabilizes the iterations by resetting the momentum to zero according to a schedule. Using a variety of models and benchmarks for image classification, we demonstrate that, in training DNNs, SRSGD significantly improves convergence and generalization; for instance, in training ResNet-200 for ImageNet classification, SRSGD achieves an error rate of 20.93% versus the benchmark of 22.13%. These improvements become more significant as the network grows deeper. Furthermore, on both CIFAR and ImageNet, SRSGD reaches similar or even better error rates with significantly fewer training epochs compared to the SGD baseline. Our implementation of SRSGD is available at https://github.com/minhtannguyen/SRSGD.
Bao Wang 0001, Tan M. Nguyen, Tao Sun 0005, Andrea L. Bertozzi, Richard G. Baraniuk, Stanley J. Osher
SIAM J. Imaging Sci.6
2021 Belief and Opinion Evolution in Social Networks: A High-Dimensional Mean Field Game Approach
abstract
Belief and opinion evolution in social networks (SNs) can aid in understanding how people influence others’ decisions through social relationships as well as provide a solid foundation for many valuable social applications. As large numbers of users are involved in SNs, the complexity of traditional optimization techniques is high as they deal with the interactions between users separately. Moreover, the state variable (opinion) is high-dimensional because a person usually has opinions about many different social issues. To overcome those challenges, we formulate the opinion evolution in SNs as a high-dimensional stochastic mean field game (MFG). Numerical methods for high-dimensional MFGs are practically non-existent because of the need for grid-based spatial discretization. Thus, we propose a machine-learning based method, where we use an alternating population and agent control neural network (APAC-net), to tractably solve high-dimensional stochastic MFGs. Through APAC-net, solving MFGs can be regarded as a special case of training a generative adversarial network (GAN). To the best of our knowledge, the APAC-Net is the first model that can solve high-dimensional stochastic MFGs. The simulation results affirm the efficiency of the APAC-net.
Hao Gao 0008, Alex Tong Lin, Reginald Banez, Wuchen Li, Zhu Han 0001, Stanley J. Osher, H. Vincent Poor
ICC6
2021 Task Selection and Route Planning for Mobile Crowd Sensing Using Multi-Population Mean-Field Games
abstract
With the increasing deployment of mobile vehicles, such as mobile robots and unmanned aerial vehicles (UAVs), it is foreseen that they will play an important role in mobile crowd sensing (MCS). Specifically, mobile vehicles equipped with sensors and computing devices are able to collect massive data due to their fast and flexible mobility in MCS systems. In this paper, we consider a mobile vehicle-based MCS system where vehicles owned by different operators or individuals compete against others for limited sensing resources. We investigate the joint task selection and route planning problem for such an MCS system. However, since the structural complexity and computational complexity of the original problem is very high, we propose a multi-population Mean-Field Game (MFG) problem by simplifying the interaction between vehicles as a distribution over their strategy space, known as the mean-field term. To solve the multi-population MFG problem efficiently, we propose a G-prox primal-dual hybrid gradient method (PDHG) algorithm whose computational complexity is independent of the number of vehicles. Numerical results show that the proposed multi-population MFG scheme and algorithm are of effectiveness and efficiency.
Yuhan Kang, Siting Liu 0003, Hongliang Zhang 0001, Zhu Han 0001, Stanley J. Osher, H. Vincent Poor
ICC5
2021 FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
abstract
We propose FMMformers, a class of efficient and flexible transformers inspired by the celebrated fast multipole method (FMM) for accelerating interacting particle simulation. FMM decomposes particle-particle interaction into near-field and far-field components and then performs direct and coarse-grained computation, respectively. Similarly, FMMformers decompose the attention into near-field and far-field attention, modeling the near-field attention by a banded matrix and the far-field attention by a low-rank matrix. Computing the attention matrix for FMMformers requires linear complexity in computational time and memory footprint with respect to the sequence length. In contrast, standard transformers suffer from quadratic complexity. We analyze and validate the advantage of FMMformers over the standard transformer on the Long Range Arena and language modeling benchmarks. FMMformers can even outperform the standard transformer in terms of accuracy by a significant margin. For instance, FMMformers achieve an average classification accuracy of $60.74\%$ over the five Long Range Arena tasks, which is significantly better than the standard transformer's average accuracy of $58.70\%$.
Tan M. Nguyen, Vai Suliafu, Stanley J. Osher, Bao Wang 0001
NeurIPS3
2021 Heavy Ball Neural Ordinary Differential Equations
abstract
We propose heavy ball neural ordinary differential equations (HBNODEs), leveraging the continuous limit of the classical momentum accelerated gradient descent, to improve neural ODEs (NODEs) training and inference. HBNODEs have two properties that imply practical advantages over NODEs: (i) The adjoint state of an HBNODE also satisfies an HBNODE, accelerating both forward and backward ODE solvers, thus significantly reducing the number of function evaluations (NFEs) and improving the utility of the trained models. (ii) The spectrum of HBNODEs is well structured, enabling effective learning of long-term dependencies from complex sequential data. We verify the advantages of HBNODEs over NODEs on benchmark tasks, including image classification, learning complex dynamics, and sequential modeling. Our method requires remarkably fewer forward and backward NFEs, is more accurate, and learns long-term dependencies more effectively than the other ODE-based neural network models. Code is available at \url{https://github.com/hedixia/HeavyBallNODE}.
Hedi Xia, Vai Suliafu, Hangjie Ji, Tan M. Nguyen, Andrea L. Bertozzi, Stanley J. Osher, Bao Wang 0001
NeurIPS6
2021 Joint Sensing Task Assignment and Collision-Free Trajectory Optimization for Mobile Vehicle Networks Using Mean-Field Games
abstract
With the increasing popularity of mobile vehicles, such as unmanned aerial vehicles (UAVs) and mobile robots, it is foreseen that they will play an important role in Internet-of-Things (IoT) networks due to their high mobility and rapid deployment. Specifically, mobile vehicles equipped with sensors act as IoT devices and can be dispatched to several sensing regions to perform sensing tasks. In this article, we consider mobile vehicles for sensing applications and investigate the corresponding joint task assignment and collision-free trajectory optimization problem. This problem is challenging as the number of involved vehicles can be very large, and to tackle the problem efficiently, we reformulate the original optimization problem into a mean-field-game (MFG) problem by simplifying the interaction between vehicles as a distribution over their state space, known as the mean-field term. To solve the MFG problem efficiently, we propose a G-prox primal-dual hybrid gradient (PDHG) algorithm that transforms the MFG problem into a saddle-point problem by defining a Lagrangian functional with a proximal operator. The complexity of this algorithm is shown to be linear with the total number of grid points in the proposed MFG problem. We provide a comprehensive theoretical analysis of the proposed model and algorithm. Numerical results together with the practical implementation on real mobile robots show that our proposed system model and algorithm are of significant effectiveness and efficiency.
Yuhan Kang, Siting Liu 0003, Hongliang Zhang 0001, Wuchen Li, Zhu Han 0001, Stanley J. Osher, H. Vincent Poor
IEEE Internet Things J.6
2020 Energy-efficient Velocity Control for Massive Numbers of Rotary-Wing UAVs: A Mean Field Game Approach
abstract
When a disaster happens in a metropolitan area, wireless communication systems in the area are highly affected, degrading the efficiency of the search and rescue (SAR) mission. An emergency wireless network must be deployed quickly and efficiently to preserve human lives. Teams of low-altitude rotarywing unmanned aerial vehicles (UAVs) are useful as on-demand temporal wireless networks because they are generally faster to deploy, flexible to reconfigure, and able to provide good communication services with short line-of-sight links. However, rotary-wing UAVs' limited on-board batteries require that they need to recharge and reconFigure frequently during a mission. Therefore, we formulate the velocity control problem for massive numbers of rotary-wing UAVs as a Schrödinger bridge problem which can describe the frequent reconfiguration of UAVs. Then we transform it into a mean field game and solve it with the Gprox primal dual hybrid gradient (PDHG) method. Finally, we show the efficiency of our algorithm and analyze the influence of wind dynamics with numerical results.
Hao Gao 0008, Wonjun Lee 0004, Wuchen Li, Zhu Han 0001, Stanley J. Osher, H. Vincent Poor
GLOBECOM5
2020 MomentumRNN: Integrating Momentum into Recurrent Neural Networks
abstract
Designing deep neural networks is an art that often involves an expensive search over candidate architectures. To overcome this for recurrent neural nets (RNNs), we establish a connection between the hidden state dynamics in an RNN and gradient descent (GD). We then integrate momentum into this framework and propose a new family of RNNs, called {\em MomentumRNNs}. We theoretically prove and numerically demonstrate that MomentumRNNs alleviate the vanishing gradient issue in training RNNs. We study the momentum long-short term memory (MomentumLSTM) and verify its advantages in convergence speed and accuracy over its LSTM counterpart across a variety of benchmarks. We also demonstrate that MomentumRNN is applicable to many types of recurrent cells, including those in the state-of-the-art orthogonal RNNs. Finally, we show that other advanced momentum-based optimization methods, such as Adam and Nesterov accelerated gradients with a restart, can be easily incorporated into the MomentumRNN framework for designing new recurrent cells with even better performance.
Tan M. Nguyen, Richard G. Baraniuk, Andrea L. Bertozzi, Stanley J. Osher, Bao Wang 0001
NeurIPS4
2019 Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets
Penghang Yin, Jiancheng Lyu, Shuai Zhang 0009, Stanley J. Osher, Yingyong Qi, Jack Xin
ICLR (Poster)4
2019 ResNets Ensemble via the Feynman-Kac Formalism to Improve Natural and Robust Accuracies
abstract
We unify the theory of optimal control of transport equations with the practice of training and testing of ResNets. Based on this unified viewpoint, we propose a simple yet effective ResNets ensemble algorithm to boost the accuracy of the robustly trained model on both clean and adversarial images. The proposed algorithm consists of two components: First, we modify the base ResNets by injecting a variance specified Gaussian noise to the output of each residual mapping. Second, we average over the production of multiple jointly trained modified ResNets to get the final prediction. These two steps give an approximation to the Feynman-Kac formula for representing the solution of a convection-diffusion equation. For the CIFAR10 benchmark, this simple algorithm leads to a robust model with a natural accuracy of {\bf 85.62}\% on clean images and a robust accuracy of ${\bf 57.94 \%}$ under the 20 iterations of the IFGSM attack, which outperforms the current state-of-the-art in defending against IFGSM attack on the CIFAR10.
Bao Wang 0001, Zuoqiang Shi, Stanley J. Osher
NeurIPS3
2018 Deep Neural Nets with Interpolating Function as Output Activation
abstract
We replace the output layer of deep neural nets, typically the softmax function, by a novel interpolating function. And we propose end-to-end training and testing algorithms for this new architecture. Compared to classical neural nets with softmax function as output activation, the surrogate with interpolating function as output activation combines advantages of both deep and manifold learning. The new framework demonstrates the following major advantages: First, it is better applicable to the case with insufficient training data. Second, it significantly improves the generalization accuracy on a wide variety of networks. The algorithm is implemented in PyTorch, and the code is available at https://github.com/ BaoWangMath/DNN-DataDependentActivation.
Bao Wang 0001, Xiyang Luo, Zhen Li 0029, Wei Zhu 0007, Zuoqiang Shi, Stanley J. Osher
NeurIPS6
2018 BinaryRelax: A Relaxation Approach for Training Deep Neural Networks with Quantized Weights
abstract
We propose BinaryRelax, a simple two-phase algorithm, for training deep neural networks with quantized weights. The set constraint that characterizes the quantization of weights is not imposed until the late stage of training, and a sequence of pseudo quantized weights is maintained. Specifically, we relax the hard constraint into a continuous regularizer via a Moreau envelope, which turns out to be the squared Euclidean distance to the set of quantized weights. The pseudo quantized weights are obtained by linearly interpolating between the float weights and their quantizations. A continuation strategy is adopted to push the weights toward the quantized state by gradually increasing the regularization parameter. In the second phase, an exact quantization scheme with a small learning rate is invoked to guarantee fully quantized weights. We test BinaryRelax on the benchmark CIFAR and ImageNet color image datasets to demonstrate the superiority of the relaxed quantization approach and the improved accuracy over the state-of-the-art training methods. Finally, we prove the convergence of BinaryRelax under an approximate orthogonality condition.
Penghang Yin, Shuai Zhang 0009, Jiancheng Lyu, Stanley J. Osher, Yingyong Qi, Jack Xin
SIAM J. Imaging Sci.4
2017 Pre-processing and classification of hyperspectral imagery via selective inpainting
abstract
We propose a semi-supervised algorithm for processing and classification of hyperspectral imagery. For initialization, we keep 20% of the data intact, and use Principal Component Analysis to discard voxels from noisier bands and pixels. Then, we use either an Accelerated Proximal Gradient algorithm (APGL), or a modified APGL algorithm with a penalty term for distance between inpainted pixels and endmembers (APGL Hyp), on the initialized datacube to inpaint the missing data. APGL and APGL Hyp are distinguished by performance on datasets with full pixels removed or extreme noise. This inpainting technique results in band-by-band datacube sharpening and removal of noise from individual spectral signatures. We can also classify the inpainted cube by assigning each pixel to its nearest endmember via Euclidean distance. We demonstrate improved accuracy in classification over data-mining techniques like k-means, unmixing techniques like Hierarchical Non-Negative Matrix Factorization, and graph-based methods like Non-Local Total Variation.
Victoria Chayes, Rasika Bhalerao, Wei Zhu 0007, Andrea L. Bertozzi, Wenzi Liao, Stanley J. Osher
ICASSP8
2017 Low Dimensional Manifold Model for Image Processing
abstract
In this paper, we propose a novel low dimensional manifold model (LDMM) and apply it to some image processing problems. LDMM is based on the fact that the patch manifolds of many natural images have low dimensional structure. Based on this fact, the dimension of the patch manifold is used as a regularization to recover the image. The key step in LDMM is to solve a Laplace--Beltrami equation over a point cloud which is solved by the point integral method. The point integral method enforces the sample point constraints correctly and gives better results than the standard graph Laplacian. Numerical simulations in image denoising, inpainting, and superresolution problems show that LDMM is a powerful method in image processing.
Stanley J. Osher, Zuoqiang Shi, Wei Zhu 0007
SIAM J. Imaging Sci.1
2017 Unsupervised Classification in Hyperspectral Imagery With Nonlocal Total Variation and Primal-Dual Hybrid Gradient Algorithm
abstract
In this paper, a graph-based nonlocal total variation method is proposed for unsupervised classification of hyperspectral images (HSI). The variational problem is solved by the primal-dual hybrid gradient algorithm. By squaring the labeling function and using a stable simplex clustering routine, an unsupervised clustering method with random initialization can be implemented. The effectiveness of this proposed algorithm is illustrated on both synthetic and real-world HSI, and numerical results show that the proposed algorithm outperforms other standard unsupervised clustering methods, such as spherical K -means, nonnegative matrix factorization, and the graph-based Merriman-Bence-Osher scheme.
Wei Zhu 0007, Victoria Chayes, Alexandre Tiard, Stephanie Sánchez, Devin Dahlberg, Andrea L. Bertozzi, Stanley J. Osher, Dominique Zosso, Da Kuang
IEEE Trans. Geosci. Remote. Sens.7
2015 Normal Estimation of a Transparent Object Using a Video
abstract
Reconstructing transparent objects is a challenging problem. While producing reasonable results for quite complex objects, existing approaches require custom calibration or somewhat expensive labor to achieve high precision. When an overall shape preserving salient and fine details is sufficient, we show in this paper a significant step toward solving the problem when the object's silhouette is available and simple user interaction is allowed, by using a video of a transparent object shot under varying illumination. Specifically, we estimate the normal map of the exterior surface of a given solid transparent object, from which the surface depth can be integrated. Our technical contribution lies in relating this normal estimation problem to one of graph-cut segmentation. Unlike conventional formulations, however, our graph is dual-layered, since we can see a transparent object's foreground as well as the background behind it. Quantitative and qualitative evaluation are performed to verify the efficacy of this practical solution.
Sai-Kit Yeung, Tai-Pang Wu, Chi-Keung Tang, Tony F. Chan, Stanley J. Osher
IEEE Trans. Pattern Anal. Mach. Intell.5
2015 A Weighted Difference of Anisotropic and Isotropic Total Variation Model for Image Processing
abstract
We propose a weighted difference of anisotropic and isotropic total variation (TV) as a regularization for image processing tasks, based on the well-known TV model and natural image statistics. Due to the form of our model, it is natural to compute via a difference of convex algorithm (DCA). We draw its connection to the Bregman iteration for convex problems and prove that the iteration generated from our algorithm converges to a stationary point with the objective function values decreasing monotonically. A stopping strategy based on the stable oscillatory pattern of the iteration error from the ground truth is introduced. In numerical experiments on image denoising, image deblurring, and magnetic resonance imaging (MRI) reconstruction, our method improves on the classical TV model consistently and is on par with representative state-of-the-art methods.
Yifei Lou, Tieyong Zeng, Stanley J. Osher, Jack Xin
SIAM J. Imaging Sci.3
2015 Space-Time Regularization for Video Decompression
abstract
We consider the problem of reconstructing frames from a video which has been compressed using the video compressive sensing (VCS) method. In VCS data, each frame comes from first subsampling the original video data in space and then averaging the subsampled sequence in time. This results in a large linear system of equations whose inversion is ill-posed. We introduce a convex regularizer to invert the system, where the spatial component is regularized by the total variation seminorm, and the temporal component is regularized by enforcing sparsity on the difference between the spatial gradients of each frame. Since the regularizers are $L^1$-like norms, the model can be written in the form of an easy-to-solve saddle point problem. The saddle point problem is solved by the primal-dual algorithm, whose implementation calls for nearly pointwise operations (i.e., no direct linear inversion) and has a simple parallel version. Results show that our model decompresses videos more accurately than other popular models, with PSNR gains of several dB.
Hayden Schaeffer, Yi Yang 0010, Stanley J. Osher
SIAM J. Imaging Sci.3
2015 Non-Local Retinex - A Unifying Framework and Beyond
abstract
In this paper, we provide a short review of Retinex and then present a unifying framework. The fundamental assumption of all Retinex models is that the observed image is a multiplication between the illumination and the true underlying reflectance of the object. Starting from Morel's 2010 PDE model, where illumination is supposed to vary smoothly and where the reflectance is thus recovered from a hard-thresholded Laplacian of the observed image in a Poisson equation, we define our unifying Retinex model in two similar, but more general, steps. We reinterpret the gradient thresholding model as variational models with sparsity constraints. First, we look for a filtered gradient that is the solution of an optimization problem consisting of two terms: a sparsity prior of the reflectance and a fidelity prior of the reflectance gradient to the observed image gradient. Second, since this filtered gradient almost certainly is not a consistent image gradient, we then fit an actual reflectance gradient to it, subject to further sparsity and fidelity priors. This generalized formulation allows making connections with other variational or kernel-based Retinex implementations. We provide simple algorithms for the optimization problems resulting from our framework. In particular, in the quadratic case, we can link our model to a plausible neural mechanism through Wilson--Cowan equations. Beyond unifying existing models, we derive entirely novel Retinex flavors by using more interesting non-local versions for the sparsity and fidelity priors. Eventually, we define within a single framework new Retinex applications to shadow detection and removal, nonuniformity correction, cartoon-texture decomposition, as well as color and hyperspectral image enhancement.
Dominique Zosso, Giang Tran, Stanley J. Osher
SIAM J. Imaging Sci.3
2014 Optimal data collection for informative rankings expose well-connected graphs
Braxton Osting, Christoph Brune, Stanley J. Osher
J. Mach. Learn. Res.3
2014 2D Empirical Transforms. Wavelets, Ridgelets, and Curvelets Revisited
abstract
A recently developed approach, called “empirical wavelet transform,” aims to build one-dimensional (1D) adaptive wavelet frames accordingly to the analyzed signal. In this paper, we present several extensions of this approach to two-dimensional (2D) signals (images). We revisit some well-known transforms (tensor wavelets, Littlewood--Paley wavelets, ridgelets, and curvelets) and show that it is possible to build their empirical counterparts. We prove that such constructions lead to different adaptive frames which show some promising properties for image analysis and processing.
Jérôme Gilles, Giang Tran, Stanley J. Osher
SIAM J. Imaging Sci.3
2013 Enhanced statistical rankings via targeted data collection
abstract
Given a graph where vertices represent alternatives and pairwise comparison data, y_ij, is given on the edges, the statistical ranking problem is to find a potential function, defined on the vertices, such that the gradient of the potential function agrees with pairwise comparisons. We study the dependence of the statistical ranking problem on the available pairwise data, i.e., pairs (i,j) for which the pairwise comparison data y_ij is known, and propose a framework to identify data which, when augmented with the current dataset, maximally increases the Fisher information of the ranking. Under certain assumptions, the data collection problem decouples, reducing to a problem of finding an edge set on the graph (with a fixed number of edges) such that the second eigenvalue of the graph Laplacian is maximal. This reduction of the data collection problem to a spectral graph-theoretic question is one of the primary contributions of this work. As an application, we study the Yahoo! Movie user rating dataset and demonstrate that the addition of a small number of well-chosen pairwise comparisons can significantly increase the Fisher informativeness of the ranking.
Braxton Osting, Christoph Brune, Stanley J. Osher
ICML (1)3
2013 A Low Patch-Rank Interpretation of Texture
abstract
We propose a novel cartoon-texture separation model using a sparse low-rank decomposition. Our texture model connects the separate ideas of robust principal component analysis (PCA) [E. J. Candès, X. Li, Y. Ma, and J. Wright, J. ACM, 58 (2011), 11], nonlocal methods [A. Buades, B. Coll, and J.-M. Morel, Multiscale Model. Simul., 4 (2005), pp. 490--530], [A. Buades, B. Coll, and J.-M. Morel, Numer. Math., 105 (2006), pp. 1--34], [G. Gilboa and S. Osher, Multiscale Model. Simul., 6 (2007), pp. 595--630], [G. Gilboa and S. Osher, Multiscale Model. Simul., 7 (2008), pp. 1005--1028], and cartoon-texture decompositions in an interesting way, taking advantage of each of these methodologies. We define our texture norm using the nuclear norm applied to patches in the image, interpreting the texture patches to be low-rank. In particular, this norm is easier to implement than many of the weak function space norms in the literature and is computationally faster than nonlocal methods since there is no explicit weight function to compute. This norm is used as an additional regularizer in several image recovery models. Using total variation as the cartoon norm and our new texture norm, we solve the proposed variational problems using the split Bregman algorithm [T. Goldstein and S. Osher, SIAM J. Imaging Sci., 2 (2009), pp. 323--343]. Since both of our regularizers are of $L^1$ type, a double splitting provides a fast algorithm that is simple to implement. Based on experimental results, we demonstrate our algorithm's success on a wide range of textures. Also, our particular cartoon-texture decomposition model has the advantage of separating noise from texture. Our proposed texture norm is shown to better reconstruct texture for other applications such as denoising, deblurring, sparse reconstruction, and pattern regularization.
Hayden Schaeffer, Stanley J. Osher
SIAM J. Imaging Sci.2
2012 Multi-Channel l1 Regularized Convex Speech Enhancement Model and Fast Computation by the Split Bregman Method
abstract
A convex speech enhancement (CSE) method is presented based on convex optimization and pause detection of the speech sources. Channel spatial difference is identified for enhancing each speech source individually while suppressing other interfering sources. Sparse unmixing filters indicating channel spatial differences are sought byl1norm regularization and the split Bregman method. A subdivided split Bregman method is developed for efficiently solving the problem in severely reverberant environments. The speech pause detection is based on a binary mask source separation method. The CSE method is evaluated objectively and subjectively, and found to outperform a list of existing blind speech separation approaches on both synthetic and room recorded speech mixtures in terms of the overall computational speed and separation quality.
Meng Yu 0003, Wenye Ma, Jack Xin, Stanley J. Osher
IEEE Trans. Speech Audio Process.4
2012 Multiple Aerosol Unmixing by the Split Bregman Algorithm
abstract
For more than a decade, the U.S. government has been developing laser-based sensors for detecting, locating, and classifying aerosols in the atmosphere at safe standoff ranges. The motivation for this work is the need to discriminate aerosols of biological origin from interferent materials such as smoke and dust using the backscatter from multiple wavelengths in the long wave infrared (LWIR) spectral region. Through previous work, algorithms have been developed for estimating the aerosol spectral dependence and concentration range dependence from these data. The range dependence is required for locating and tracking the aerosol plumes, and the backscatter spectral dependence is used for discrimination by a support vector machine classifier. Substantial progress has been made in these algorithms for the case of a single aerosol present in the lidar line-of-sight (LOS). Often, however, mixtures of aerosols are present along the same LOS overlapped in range and time. Analysis of these mixtures of aerosols presents a difficult inverse problem that cannot be successfully treated by the methods used for single aerosols. Fortunately, recent advances have been made in the analysis of inverse problems using shrinkage-basedL1-regularization techniques. Of the severalL1-regularization methods currently known, the split Bregman algorithm is straightforward to implement, converges rapidly, and is applicable to a broad range of inverse problems including our aerosol unmixing. In this paper, we show how the split Bregman algorithm can successfully resolve LWIR lidar data containing mixtures of bioaerosol simulants and interferents into their separate components. The individual components then can be classified as bio- or nonbioaerosol by our SVM classifier. We illustrate the approach through data collected in field tests over the past several years using the U.S. Army FAL sensor in testing at Dugway Proving Ground, UT.
Russell E. Warren, Stanley J. Osher, Richard G. Vanderbeek
IEEE Trans. Geosci. Remote. Sens.2
2012 A Convex Model for Nonnegative Matrix Factorization and Dimensionality Reduction on Physical Space
abstract
A collaborative convex framework for factoring a data matrix X into a nonnegative product AS , with a sparse coefficient matrix S, is proposed. We restrict the columns of the dictionary matrix A to coincide with certain columns of the data matrix X, thereby guaranteeing a physically meaningful dictionary and dimensionality reduction. We use l(1, ∞) regularization to select the dictionary from the data and show that this leads to an exact convex relaxation of l(0) in the case of distinct noise-free data. We also show how to relax the restriction-to- X constraint by initializing an alternating minimization approach with the solution of the convex model, obtaining a dictionary close to but not necessarily in X. We focus on applications of the proposed framework to hyperspectral endmember and abundance identification and also show an application to blind source separation of nuclear magnetic resonance data.
Ernie Esser, Michael Möller 0001, Stanley J. Osher, Guillermo Sapiro, Jack Xin
IEEE Trans. Image Process.3
2012 Efficient Algorithm for Level Set Method Preserving Distance Function
abstract
The level set method is a popular technique for tracking moving interfaces in several disciplines, including computer vision and fluid dynamics. However, despite its high flexibility, the original level set method is limited by two important numerical issues. First, the level set method does not implicitly preserve the level set function as a distance function, which is necessary to estimate accurately geometric features, s.a. the curvature or the contour normal. Second, the level set algorithm is slow because the time step is limited by the standard Courant-Friedrichs-Lewy (CFL) condition, which is also essential to the numerical stability of the iterative scheme. Recent advances with graph cut methods and continuous convex relaxation methods provide powerful alternatives to the level set method for image processing problems because they are fast, accurate, and guaranteed to find the global minimizer independently to the initialization. These recent techniques use binary functions to represent the contour rather than distance functions, which are usually considered for the level set method. However, the binary function cannot provide the distance information, which can be essential for some applications, s.a. the surface reconstruction problem from scattered points and the cortex segmentation problem in medical imaging. In this paper, we propose a fast algorithm to preserve distance functions in level set methods. Our algorithm is inspired by recent efficient l(1) optimization techniques, which will provide an efficient and easy to implement algorithm. It is interesting to note that our algorithm is not limited by the CFL condition and it naturally preserves the level set function as a distance function during the evolution, which avoids the classical re-distancing problem in level set methods. We apply the proposed algorithm to carry out image segmentation, where our methods prove to be 5-6 times faster than standard distance preserving level set techniques. We also present two applications where preserving a distance function is essential. Nonetheless, our method stays generic and can be applied to any level set methods that require the distance information.
Virginia Estellers, Dominique Zosso, Rongjie Lai, Stanley J. Osher, Jean-Philippe Thiran, Xavier Bresson
IEEE Trans. Image Process.4
2011 An L1-based variational model for Retinex theory and its application to medical images
abstract
Human visual system (HVS) can perceive constant color under varying illumination conditions while digital images record information of both reflectance (physical color) of objects and illumination. Retinex theory, formulated by Edwin H. Land, aimed to simulate and explain this feature of HVS. However, to recover the reflectance from a given image is in general an ill-posed problem. In this paper, we establish an L1-based variational model for Retinex theory that can be solved by a fast computational approach based on Bregman iteration. Compared with previous works, our L1-Retinex method is more accurate for recovering the reflectance, which is illustrated by examples and statistics. In medical images such as magnetic resonance imaging (MRI), intensity inhomogeneity is often encountered due to bias fields. This is a similar formulation to Retinex theory while the MRI has some specific properties. We then modify the L1-Retinex method and develop a new algorithm for MRI data. We demonstrate the performance of our method by comparison with previous work on simulated and real data.
Wenye Ma, Jean-Michel Morel, Stanley J. Osher, Aichi Chien
CVPR3
2011 Adequate reconstruction of transparent objects on a shoestring budget
abstract
Reconstructing transparent objects is a challenging problem. While producing reasonable results for quite complex objects, existing approaches require custom calibration or somewhat expensive labor to achieve high precision. On the other hand, when an overall shape preserving salient and fine details is sufficient, we show in this paper a significant step toward solving the problem on a shoestring budget, by using only a video camera, a moving spotlight, and a small chrome sphere. Specifically, the problem we address is to estimate the normal map of the exterior surface of a given solid transparent object, from which the surface depth can be integrated. Our technical contribution lies in relating this normal reconstruction problem to one of graph-cut segmentation. Unlike conventional formulations, however, our graph is dual-layered, since we can see a transparent object's foreground as well as the background behind it. Quantitative and qualitative evaluation are performed to verify the efficacy of this practical solution.
Sai-Kit Yeung, Tai-Pang Wu, Chi-Keung Tang, Tony F. Chan, Stanley J. Osher
CVPR5
2011 Fast edge-filtered image upsampling
abstract
We present a novel edge preserved interpolation scheme for fast upsampling of natural images. The proposed piecewise hyperbolic operator uses a slope-limiter function that conveniently lends itself to higher-order approximations and is responsible for restricting spatial oscillations arising due to the edges and sharp details in the image. As a consequence the upsampled image not only exhibits enhanced edges, and discontinuities across boundaries, but also preserves smoothly varying features in images. Experimental results show an improvement in the PSNR compared to typical cubic, and spline-based interpolation approaches.
Shantanu H. Joshi, Antonio Marquina, Stanley J. Osher, Ivo D. Dinov, Arthur W. Toga, John D. Van Horn
ICIP3
2011 Make it home: automatic optimization of furniture arrangement
abstract
We present a system that automatically synthesizes indoor scenes realistically populated by a variety of furniture objects. Given examples of sensibly furnished indoor scenes, our system extracts, in advance, hierarchical and spatial relationships for various furniture objects, encoding them into priors associated with ergonomic factors, such as visibility and accessibility, which are assembled into a cost function whose optimization yields realistic furniture arrangements. To deal with the prohibitively large search space, the cost function is optimized by simulated annealing using a Metropolis-Hastings state search step. We demonstrate that our system can synthesize multiple realistic furniture arrangements and, through a perceptual study, investigate whether there is a significant difference in the perceived functionality of the automatically synthesized results relative to furniture arrangements produced by human designers.
Lap-Fai Yu, Sai-Kit Yeung, Chi-Keung Tang, Demetri Terzopoulos, Tony F. Chan, Stanley J. Osher
ACM Trans. Graph.6
2010 A split Bregman method for non-negative sparsity penalized least squares with applications to hyperspectral demixing
abstract
We will describe an alternating direction (aka split Bregman) method for solving problems of the form minu∥Au - f∥2+ η∥u∥1such that u ≥ 0, where A is an m×n matrix, and η is a nonnegative parameter. The algorithm works especially well for solving large numbers of small to medium overdetermined problems (i.e. m > n) with a fixed A. We will demonstrate applications in the analysis of hyperspectral images.
Arthur Szlam, Zhaohui Guo, Stanley J. Osher
ICIP3
2010 Reducing musical noise in blind source separation by time-domain sparse filters and split bregman method
abstract
Musical noise often arises in the outputs of time-frequency binary mask based blind source separation approaches. Postprocessing is desired to enhance the separation quality. An efficient musical noise reduction method by time-domain sparse filters is presented using convex optimization. The sparse filters are sought by l1 regularization and the split Bregman method. The proposed musical noise reduction method is evaluated by both synthetic and room recorded speech and music data, and found to outperform existing musical noise reduction methods in terms of the objective and subjective measures. Index Terms: Musical noise, time-frequency mask, timedomain sparse filters, split Bregman method.
Wenye Ma, Meng Yu 0003, Jack Xin, Stanley J. Osher
INTERSPEECH4
2010 Convexity and fast speech extraction by split bregman method
abstract
A fast speech extraction (FSE) method is presented using convex optimization made possible by pause detection of the speech sources. Sparse unmixing filters are sought by l1 regularization and the split Bregman method. A subdivided split Bregman method is developed for efficiently estimating long reverberations in real room recordings. The speech pause detection is based on a binary mask source separation method. The FSE method is evaluated and found to outperform existing blind speech separation approaches on both synthetic and room recorded data in terms of the overall computational speed and separation quality. Index Terms: convexity, sparse filters, split Bregman method, fast blind speech extraction.
Meng Yu 0003, Wenye Ma, Jack Xin, Stanley J. Osher
INTERSPEECH4
2010 A Fast Hybrid Algorithm for Large-Scale l1-Regularized Logistic Regression
Jianing Shi, Wotao Yin, Stanley J. Osher, Paul Sajda
J. Mach. Learn. Res.3
2010 Bregmanized Nonlocal Regularization for Deconvolution and Sparse Reconstruction
abstract
Bregman methods introduced in [S. Osher, M. Burger, D. Goldfarb, J. Xu, and W. Yin, Multiscale Model. Simul., 4 (2005), pp. 460–489] to image processing are demonstrated to be an efficient optimization method for solving sparse reconstruction with convex functionals, such as the $\ell^1$ norm and total variation [W. Yin, S. Osher, D. Goldfarb, and J. Darbon, SIAM J. Imaging Sci., 1 (2008), pp. 143–168; T. Goldstein and S. Osher, SIAM J. Imaging Sci., 2 (2009), pp. 323–343]. In particular, the efficiency of this method relies on the performance of inner solvers for the resulting subproblems. In this paper, we propose a general algorithm framework for inverse problem regularization with a single forward-backward operator splitting step [P. L. Combettes and V. R. Wajs, Multiscale Model. Simul., 4 (2005), pp. 1168–1200], which is used to solve the subproblems of the Bregman iteration. We prove that the proposed algorithm, namely, Bregmanized operator splitting (BOS), converges without fully solving the subproblems. Furthermore, we apply the BOS algorithm and a preconditioned one for solving inverse problems with nonlocal functionals. Our numerical results on deconvolution and compressive sensing illustrate the performance of nonlocal total variation regularization under the proposed algorithm framework, compared to other regularization techniques such as the standard total variation method and the wavelet-based regularization method.
Xiaoqun Zhang, Martin Burger 0001, Xavier Bresson, Stanley J. Osher
SIAM J. Imaging Sci.4
2010 Development and Optimization of Regularized Tomographic Reconstruction Algorithms Utilizing Equally-Sloped Tomography
abstract
We develop two new algorithms for tomographic reconstruction which incorporate the technique of equally-sloped tomography (EST) and allow for the optimized and flexible implementation of regularization schemes, such as total variation constraints, and the incorporation of arbitrary physical constraints. The founding structure of the developed algorithms is EST, a technique of tomographic acquisition and reconstruction first proposed by Miao in 2005 for performing tomographic image reconstructions from a limited number of noisy projections in an accurate manner by avoiding direct interpolations. EST has recently been successfully applied to coherent diffraction microscopy, electron microscopy, and computed tomography for image enhancement and radiation dose reduction. However, the bottleneck of EST lies in its slow speed due to its higher computation requirements. In this paper, we formulate the EST approach as a constrained problem and subsequently transform it into a series of linear problems, which can be accurately solved by the operator splitting method. Based on these mathematical formulations, we develop two iterative algorithms for tomographic image reconstructions through EST, which incorporate Bregman and continuative regularization. Our numerical experiment results indicate that the new tomographic image reconstruction algorithms not only significantly reduce the computational time, but also improve the image quality. We anticipate that EST coupled with the novel iterative algorithms will find broad applications in X-ray tomography, electron microscopy, coherent diffraction microscopy, and other tomography fields.
Benjamin P. Fahimian, Stanley J. Osher, Jianwei Miao
IEEE Trans. Image Process.3
2009 A note on the Bregmanized Total Variation and dual forms
abstract
This paper considers two approaches to perform image restoration while preserving the contrast. The first one is the Total Variation-based Bregman iterations while the second consists in the minimization of an energy that involves robust edge preserving regularization. We show that these two approaches can be derived form a common framework. This allows us to deduce new properties and to extend and generalize these two previous approaches.
Jérôme Darbon, Igor Ciril, Antonio Marquina, Tony F. Chan, Stanley J. Osher
ICIP5
2009 Comparing registration methods for mapping brain change using tensor-based morphometry
Igor Yanovsky, Alex D. Leow, Suh Lee, Stanley J. Osher, Paul M. Thompson
Medical Image Anal.4
2009 Linearized Bregman Iterations for Frame-Based Image Deblurring
abstract
Real images usually have sparse approximations under some tight frame systems derived from framelets, an oversampled discrete (window) cosine, or a Fourier transform. In this paper, we propose a method for image deblurring in tight frame domains. It is reduced to finding a sparse solution of a system of linear equations whose coefficient matrix is rectangular. Then, a modified version of the linearized Bregman iteration proposed and analyzed in [J.-F. Cai, S. Osher, and Z. Shen, Math. Comp., to appear, UCLA CAM Report (08-52), 2008; J.-F. Cai, S. Osher, and Z. Shen, Math. Comp., to appear, UCLA CAM Report (08-06), 2008; S. Osher et al., UCLA CAM Report (08-37), 2008; W. Yin et al., SIAM J. Imaging Sci., 1 (2008), pp. 143–168] can be applied. Numerical examples show that the method is very simple to implement, robust to noise, and effective for image deblurring.
Jian-Feng Cai 0001, Stanley J. Osher, Zuowei Shen
SIAM J. Imaging Sci.2
2009 The Split Bregman Method for L1-Regularized Problems
abstract
The class of L1-regularized optimization problems has received much attention recently because of the introduction of “compressed sensing,” which allows images and signals to be reconstructed from small amounts of data. Despite this recent attention, many L1-regularized problems still remain difficult to solve, or require techniques that are very problem-specific. In this paper, we show that Bregman iteration can be used to solve a wide variety of constrained optimization problems. Using this technique, we propose a “split Bregman” method, which can solve a very broad class of L1-regularized problems. We apply this technique to the Rudin–Osher–Fatemi functional for image denoising and to a compressed sensing problem that arises in magnetic resonance imaging.
Tom Goldstein, Stanley J. Osher
SIAM J. Imaging Sci.2
2008 Level Set Based Surface Capturing in 3D Medical Images
Bin Dong 0001, Aichi Chien, Stanley J. Osher
MICCAI (1)5
2008 Shape from Defocus via Diffusion
abstract
Defocus can be modeled as a diffusion process and represented mathematically using the heat equation, where image blur corresponds to the diffusion of heat. This analogy can be extended to non-planar scenes by allowing a space-varying diffusion coefficient. The inverse problem of reconstructing 3-D structure from blurred images corresponds to an "inverse diffusion" that is notoriously ill-posed. We show how to bypass this problem by using the notion of relative blur. Given two images, within each neighborhood, the amount of diffusion necessary to transform the sharper image into the blurrier one depends on the depth of the scene. This can be used to devise a global algorithm to estimate the depth profile of the scene without recovering the deblurred image, using only forward diffusion.
Paolo Favaro, Stefano Soatto, Martin Burger 0001, Stanley J. Osher
IEEE Trans. Pattern Anal. Mach. Intell.4
2008 Topology Preserving Linear Filtering Applied to Medical Imaging
abstract
One of the central problems of medical imaging is the three-dimensional (3D) visualization of body parts. The 3D volume can be viewed in slices, but the extraction of a part requires a segmentation process. Inasmuch as body parts are distinguishable by their various densities, a widely accepted method for extracting an organ is to extract isodensity surfaces by a simple threshold. Unfortunately, the density of organs, arteries, etc. varies spatially due to morphology, and no unique threshold allows one to extract the organs boundaries. The snake or active contour methods have attempted to capture these boundaries as smooth and overall contrasted surfaces. The snake method suffers, however, from severe drawbacks. The contour has to be initialized near the boundary. In addition, many body parts have too complex a topology. In this paper we focus on another idea, which is to preprocess the image before thresholding. The preprocessing aims at the homogeneity of the different parts while preserving small features. Starting from a recent seminal work by Grady and Funka-Lea [in Computer Vision and Mathematical Methods in Medical and Biomedical Image Analysis: ECCV 2004 Workshops CVAMIA and MMBIA, Prague, Czech Republic, May 2004, Revised Selected Papers, Springer, Berlin, 2004, pp. 230–245], several linear heat equations on images will be compared. They stem from nonlinear partial differential equations or from their associated nonlinear filters. By linearizing these processes one obtains more accurate topology preserving methods. These linear filters will be tested comparatively to visualize challenging angiography images of arteries. A salient fact of the method will emerge. By a concentration phenomenon, peaks in the image histogram become much more concentrated under the linear heat equations, thus permitting us to fix the thresholds defining the surfaces without supervision. Automatic extraction can be performed in this way for angiography images taken at a one-year or longer delay.
Antoni Buades, Aichi Chien, Jean-Michel Morel, Stanley J. Osher
SIAM J. Imaging Sci.4
2008 A Nonlinear Inverse Scale Space Method for a Convex Multiplicative Noise Model
abstract
We are motivated by a recently developed nonlinear inverse scale space method for image denoising [M. Burger, G. Gilboa, S. Osher, and J. Xu, Commun. Math. Sci., 4 (2006), pp. 179–212; M. Burger, S. Osher, J. Xu, and G. Gilboa, in Variational, Geometric, and Level Set Methods in Computer Vision, Lecture Notes in Comput. Sci. 3752, Springer, Berlin, 2005, pp. 25–36], whereby noise can be removed with minimal degradation. The additive noise model has been studied extensively, using the Rudin–Osher–Fatemi model [L. I. Rudin, S. Osher, and E. Fatemi, Phys. D, 60 (1992), pp. 259–268], an iterative regularization method [S. Osher, M. Burger, D. Goldfarb, J. Xu, and W. Yin, Multiscale Model. Simul., 4 (2005), pp. 460–489], and the inverse scale space flow [M. Burger, G. Gilboa, S. Osher, and J. Xu, Commun. Math. Sci., 4 (2006), pp. 179–212; M. Burger, S. Osher, J. Xu, and G. Gilboa, in Variational, Geometric, and Level Set Methods in Computer Vision, Lecture Notes in Comput. Sci. 3752, Springer, Berlin, 2005, pp. 25–36]. However, the multiplicative noise model has not yet been studied thoroughly. Earlier total variation models for the multiplicative noise cannot easily be extended to the inverse scale space, due to the lack of global convexity. In this paper, we review existing multiplicative models and present a new total variation framework for the multiplicative noise model, which is globally strictly convex. We extend this convex model to the nonlinear inverse scale space flow and its corresponding relaxed inverse scale space flow. We demonstrate the convergence of the flow for the multiplicative noise model, as well as its regularization effect and its relation to the Bregman distance. We investigate the properties of the flow and study the dependence on flow parameters. The numerical results show an excellent denoising effect and significant improvement over earlier multiplicative models.
Jianing Shi, Stanley J. Osher
SIAM J. Imaging Sci.2
2008 Bregman Iterative Algorithms for \ell1-Minimization with Applications to Compressed Sensing
abstract
We propose simple and extremely efficient methods for solving the basis pursuit problem $\min\{\|u\|_1 : Au = f, u\in\mathbb{R}^n\},$ which is used in compressed sensing. Our methods are based on Bregman iterative regularization, and they give a very accurate solution after solving only a very small number of instances of the unconstrained problem $\min_{u\in\mathbb{R}^n} \mu\|u\|_1+\frac{1}{2}\|Au-f^k\|_2^2$ for given matrix A and vector $f^k$. We show analytically that this iterative approach yields exact solutions in a finite number of steps and present numerical results that demonstrate that as few as two to six iterations are sufficient in most cases. Our approach is especially useful for many compressed sensing applications where matrix-vector operations involving A and $A^\top$ can be computed by fast transforms. Utilizing a fast fixed-point continuation solver that is based solely on such operations for solving the above unconstrained subproblem, we were able to quickly solve huge instances of compressed sensing problems on a standard PC.
Wotao Yin, Stanley J. Osher, Donald Goldfarb, Jérôme Darbon
SIAM J. Imaging Sci.2
2007 Topology Preserving Log-Unbiased Nonlinear Image Registration: Theory and Implementation
abstract
In this paper, we present a novel framework for constructing large deformation log-unbiased image registration models that generate theoretically and intuitively correct deformation maps. Such registration models do not rely on regridding and are inherently topology preserving. We apply information theory to quantify the magnitude of deformations and examine the statistical distributions of Jacobian maps in the logarithmic space. To demonstrate the power of the proposed framework, we generalize the well known viscous fluid registration model to compute log-unbiased deformations. We tested the proposed method using a pair of binary corpus callosum images, a pair of two-dimensional serial MRI images, and a set of three-dimensional serial MRI brain images. We compared our results to those computed using the viscous fluid registration method, and demonstrated that the proposed method is advantageous when recovering voxel-wise maps of local tissue change.
Igor Yanovsky, Paul M. Thompson, Stanley J. Osher, Alex D. Leow
CVPR3
2007 Multiphase Segmentation of Deformation using Logarithmic Priors
abstract
In [8], the authors proposed the large deformation log-unbiased diffeomorphic nonlinear image registration model which has been successfully used to obtain theoretically and intuitively correct deformation maps. In this paper, we extend this idea to simultaneously registering and tracking deforming objects in a sequence of two or more images. We generalize a level set based Chan-Vese multiphase segmentation model to consider Jacobian fields while segmenting regions of growth and shrinkage in deformations. Deforming objects are thus classified based on magnitude of homogeneous deformation. Numerical experiments demonstrating our results include a pair of two-dimensional synthetic images and pairs of two-dimensional and three-dimensional serial MRI images.
Igor Yanovsky, Paul M. Thompson, Stanley J. Osher, Luminita A. Vese, Alex D. Leow
CVPR3
2007 A comparison of three total variation based texture extraction models
Wotao Yin, Donald Goldfarb, Stanley J. Osher
J. Vis. Commun. Image Represent.3
2007 Direct cortical mapping via solving partial differential equations on implicit surfaces
Yonggang Shi, Paul M. Thompson, Ivo D. Dinov, Stanley J. Osher, Arthur W. Toga
Medical Image Anal.4
2007 Iterative Regularization and Nonlinear Inverse Scale Space Applied to Wavelet-Based Denoising
abstract
In this paper, we generalize the iterative regularization method and the inverse scale space method, recently developed for total-variation (TV) based image restoration, to wavelet-based image restoration. This continues our earlier joint work with others where we applied these techniques to variational-based image restoration, obtaining significant improvement over the Rudin-Osher-Fatemi TV-based restoration. Here, we apply these techniques to soft shrinkage and obtain the somewhat surprising result that a) the iterative procedure applied to soft shrinkage gives firm shrinkage and converges to hard shrinkage and b) that these procedures enhance the noise-removal capability both theoretically, in the sense of generalized Bregman distance, and for some examples, experimentally, in terms of the signal-to-noise ratio, leaving less signal in the residual.
Jinjun Xu, Stanley J. Osher
IEEE Trans. Image Process.2
2006 Structure-Texture Image Decomposition - Modeling, Algorithms, and Parameter Selection
Jean-François Aujol, Guy Gilboa, Tony F. Chan, Stanley J. Osher
Int. J. Comput. Vis.4
2006 Kernel Density Estimation and Intrinsic Alignment for Shape Priors in Level Set Segmentation
Daniel Cremers, Stanley J. Osher, Stefano Soatto
Int. J. Comput. Vis.2
2004 Noise removal using smoothed normals and surface fitting
abstract
In this work, we use partial differential equation techniques to remove noise from digital images. The removal is done in two steps. We first use a total-variation filter to smooth the normal vectors of the level curves of a noise image. After this, we try to find a surface to fit the smoothed normal vectors. For each of these two stages, the problem is reduced to a nonlinear partial differential equation. Finite difference schemes are used to solve these equations. A broad range of numerical examples are given in the paper.
Ola Marius Lysaker, Stanley J. Osher, Xue-Cheng Tai
IEEE Trans. Image Process.2
2003 Simultaneous Structure and Texture Image Inpainting
abstract
An algorithm for the simultaneous filling-in of texture and structure in regions of missing image information is presented. The basic idea is to first decompose the image into the sum of two functions with different basic characteristics, and then reconstruct each one of these functions separately with structure and texture filling-in algorithms. The first function used in the decomposition is of bounded variation, representing the underlying image structure, while the second function captures the texture and possible noise. The region of missing information in the bounded variation image is reconstructed using image inpainting algorithms, while the same region in the texture image is filled-in with texture synthesis techniques. The original image is then reconstructed adding back these two sub-images. The novel contribution of the paper is then in the combination of these three previously developed components: image decomposition with inpainting and texture synthesis, which permits the simultaneous use of filling-in algorithms that are suited for different image characteristics. Examples on real images show the advantages of this proposed approach.
Marcelo Bertalmío, Luminita A. Vese, Guillermo Sapiro, Stanley J. Osher
CVPR (2)4
2003 3D Shape from Anisotropic Diffusion
abstract
We cast the problem of inferring the 3D shape of a scene from a collection of defocused images in the framework of anisotropic diffusion. We propose an algorithm that can estimate the shape of a scene by inferring the diffusion coefficient of a heat equation. The method is optimal, as we pose it as the minimization of a certain cost functional based on the input images, and fast. Furthermore, we also extend our algorithm to the case of multiple images, and derive a 3D scene segmentation algorithm that can work in the presence of pictorial camouflage.
Paolo Favaro, Stanley J. Osher, Stefano Soatto, Luminita A. Vese
CVPR (1)2
2003 Image filling-in in a decomposition space
abstract
An algorithm for the simultaneous filling-in of texture and structure in regions of missing image information is presented in this paper. The basic idea is to first decompose the image into the sum of two functions with different basic characteristics, and then reconstruct each one of these functions separately with structure and texture filling-in algorithms. The first function used in the decomposition is of bounded variation, representing the underlying image structure, while the second function captures the texture and possible noise. The region of missing information in the bounded variation image is reconstructed using image inpainting algorithms, while the same region in the texture image is filled-in with texture synthesis techniques. The original image is then reconstructed adding back these two subimages. The novel contribution of this paper is then in the combination of these three previously developed components, image decomposition with in-painting and texture synthesis, which permits the simultaneous use of filling-in algorithms that are suited for different image characteristics. The novelty in the approach is to perform filling-in in a domain different from the original given image space. Examples on real images show the advantages of this proposed approach.
Marcelo Bertalmío, Luminita A. Vese, Guillermo Sapiro, Stanley J. Osher
ICIP (1)4
2003 New framework for object warping: semi-Lagrangian level set approach
abstract
In this paper, we present a new framework for object matching between two images. This method could handle multiple pairs of overlapping and non-overlapping shapes, open curves, and landmarks. When implemented in 3-D, the same framework could be used to warp 3-D objects with minimal modification. Our approach is to use the level set formulation to represent the objects to be matched. Using this representation, the problem becomes an energy minimization problem. Cost functions for warping overlapping, non-overlapping, open curves, and landmarks are proposed. Euler-Lagrange equations are applied and gradient descent is used to solve the corresponding partial differential equations. Moreover, a general framework for linking the level set approach and the infinite dimensional group actions is discussed.
Wei-Hsun Liao, Hillary Protas, Luminita A. Vese, Sung-Cheng Huang, Marvin Bergsneider, Stanley J. Osher
ICIP (3)6
2003 Image decomposition, image restoration, and texture modeling using total variation minimization and the H/sup-1/ norm
abstract
We propose a new model for image restoration and decomposition, based on the total variation minimization of Rudin-Osher-Fatemi (1992), and on some new techniques by Y. Meyer (2002) for oscillatory functions. An initial image f is decomposed into a cartoon part u and a texture or noise part v. The u component is modeled by a function of bounded variation, while the v component by an oscillatory function, with bounded H/sup -1/ norm. After some transformation, the resulting PDE is of fourth order. The proposed model continues the ideas and techniques previously introduced by the authors in L Vese et al., (2002). Image decomposition and denoising numerical results will be shown by the proposed new fourth order nonlinear partial differential equation.
Stanley J. Osher, Andres Fco. Solé, Luminita A. Vese
ICIP (1)1
2003 Level set methods in image science
abstract
In this article, we discuss the question "what level set methods can do for image science". We examine the presence of these and related methods in image science and introduce some relevant level set techniques that are potentially useful for this class of applications. We will show that image science demands multidisciplinary knowledge and flexible but still robust methods, and that is why the level set method has become a thriving technique in this field.
Richard Tsai 0001, Stanley J. Osher
ICIP (2)2
2003 Simultaneous structure and texture image inpainting
abstract
An algorithm for the simultaneous filling-in of texture and structure in regions of missing image information is presented in this paper. The basic idea is to first decompose the image into the sum of two functions with different basic characteristics, and then reconstruct each one of these functions separately with structure and texture filling-in algorithms. The first function used in the decomposition is of bounded variation, representing the underlying image structure, while the second function captures the texture and possible noise. The region of missing information in the bounded variation image is reconstructed using image inpainting algorithms, while the same region in the texture image is filled-in with texture synthesis techniques. The original image is then reconstructed adding back these two sub-images. The novel contribution of this paper is then in the combination of these three previously developed components, image decomposition with inpainting and texture synthesis, which permits the simultaneous use of filling-in algorithms that are suited for different image characteristics. Examples on real images show the advantages of this proposed approach.
Marcelo Bertalmío, Luminita A. Vese, Guillermo Sapiro, Stanley J. Osher
IEEE Trans. Image Process.4
2003 Geometric surface processing via normal maps
abstract
We propose that the generalization of signal and image processing to surfaces entails filtering the normals of the surface, rather than filtering the positions of points on a mesh. Using a variational strategy, penalty functions on the surface geometry can be formulated as penalty functions on the surface normals, which are computed using geometry-based shape metrics and minimized using fourth-order gradient descent partial differential equations (PDEs). In this paper, we introduce a two-step approach to implementing geometric processing tools for surfaces: (i) operating on the normal map of a surface, and (ii) manipulating the surface to fit the processed normals. Iterating this two-step process, we efficiently can implement geometric fourth-order flows by solving a set of coupled second-order PDEs. The computational approach uses level set surface models; therefore, the processing does not depend on any underlying parameterization. This paper will demonstrate that the proposed strategy provides for a wide range of surface processing operations, including edge-preserving smoothing and high-boost filtering. Furthermore, the generality of the implementation makes it appropriate for very complex surface models, for example, those constructed directly from measured data.
Tolga Tasdizen, Ross T. Whitaker, Paul Burchard, Stanley J. Osher
ACM Trans. Graph.4
2002 Geometric Surface Smoothing via Anisotropic Diffusion of Normals
abstract
This paper introduces a method for smoothing complex, noisy surfaces, while preserving (and enhancing) sharp, geometric features. It has two main advantages over previous approaches to feature preserving surface smoothing. First is the use of level set surface models, which allows us to process very complex shapes of arbitrary and changing topology. This generality makes it well suited for processing surfaces that are derived directly from measured data. The second advantage is that the proposed method derives from a well-founded formulation, which is a natural generalization of anisotropic diffusion, as used in image processing. This formulation is based on the proposition that the generalization of image filtering entails filtering the normals of the surface, rather than processing the positions of points on a mesh.
Tolga Tasdizen, Ross T. Whitaker, Paul Burchard, Stanley J. Osher
IEEE Visualization4
2001 The digital TV filter and nonlinear denoising
abstract
Motivated by the classical TV (total variation) restoration model, we propose a new nonlinear filter-the digital TV filter for denoising and enhancing digital images, or more generally, data living on graphs. The digital TV filter is a data dependent lowpass filter, capable of denoising data without blurring jumps or edges. In iterations, it solves a global total variational (or L(1)) optimization problem, which differs from most statistical filters. Applications are given in the denoising of one dimensional (1-D) signals, two-dimensional (2-D) data with irregular structures, gray scale and color images, and nonflat image features such as chromaticity.
Tony F. Chan, Stanley J. Osher, Jianhong Shen
IEEE Trans. Image Process.2
2000 Implicit and Nonparametric Shape Reconstruction from Unorganized Data Using a Variational Level Set Method
Hongkai Zhao, Stanley J. Osher, Barry Merriman, Myungjoo Kang
Comput. Vis. Image Underst.2
1994 Total Variation Based Image Restoration with Free Local Constraints
abstract
A new total variation based approach was developed by Rudin, Osher and Fatemi (see Physica D., vol.60, p.259, 1992) to overcome the basic limitations of all smooth regularization algorithms. The TV-based technique use the L/sup 1/ norm of the magnitude of a gradient, thus making discontinuous and nonsmooth solutions possible. In TV image restoration, the solution is obtained by solving a time-dependent, nonlinear PDE on a manifold that satisfies the degradation constraints. In practical applications, one assumes a space-varying blurring kernel and signal-dependent (e.g. multiplicative) noise. The evolution part of the TV-based PDE turned out to be related to the curve shortening equation, but scaled by an inverse.>
Leonid I. Rudin, Stanley J. Osher
ICIP (1)2
1991 Solution of the hydrodynamic device model using high-order nonoscillatory shock capturing algorithms
abstract
Simulation results for the hydrodynamic model are presented for an n/sup +/-n-n/sup +/ diode by use of shock-capturing numerical algorithms applied to the transient model with subsequent passage to the steady state. The numerical method is first order in time, but of high spatial order in regions of smoothness. Implementation typically requires a few thousand time steps. These algorithms, termed essentially nonoscillatory, have been successfully applied in other contexts to model the flow in gas dynamics, magnetohydrodynamics, and other physical situations involving the conservation laws of fluid mechanics. The presented semiconductor simulations reveal temporal and spatial velocity overshot, as well as overshoot relative to an electric field induced by the Poisson equation. Shocks are observed in the transient simulations for certain low-temperature parameter regimes.>
Emad Fatemi, Joseph W. Jerome, Stanley J. Osher
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3