Raja Bala

dblp:63/5464 · also Raja Balasubramanian · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0002-0142-9859ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
YearPublicationVenuePosition
2024 MRC-Net: 6-DoF Pose Estimation with MultiScale Residual Correlation
abstract
We propose a single-shot approach to determining 6-DoF pose of an object with available 3D computer-aided design (CAD) model from a single RGB image. Our method, dubbed MRC-Net, comprises two stages. The first performs pose classification and renders the 3D object in the classified pose. The second stage performs regression to predict fine-grained residual pose within class. Connecting the two stages is a novel multi-scale residual correlation (MRC) layer that captures high-and-low level correspondences between the input image and rendering from first stage. MRC-Net employs a Siamese network with shared weights between both stages to learn embeddings for input and rendered images. To mitigate ambiguity when predicting discrete pose class labels on symmetric objects, we use soft probabilistic labels to define pose class in the first stage. We demonstrate state-of-the-art accuracy, outperforming all competing RGB-based methods on four challenging BOP benchmark datasets: T-LESS, LM-O, YCB-V, and ITODD. Our method is non-iterative and requires no complex post-processing. Our code and pretrained models are available at https://github.com/amzn/mrc-net-6d-pose.
Yafei Mao, Raja Bala, Sunil Hadap
CVPR3
2022 Human Body Measurement Estimation with Adversarial Augmentation
abstract
We present a Body Measurement network (BMnet) for estimating 3D anthropomorphic measurements of the human body shape from silhouette images. Training of BMnet is performed on data from real human subjects, and augmented with a novel adversarial body simulator (ABS) that finds and synthesizes challenging body shapes. ABS is based on the skinned multiperson linear (SMPL) body model, and aims to maximize BMnet measurement prediction error with respect to latent SMPL shape parameters. ABS is fully differentiable with respect to these parameters, and trained end-to-end via backpropagation with BMnet in the loop. Experiments show that ABS effectively discovers adversarial examples, such as bodies with extreme body mass indices (BMI), consistent with the rarity of extreme-BMI bodies in BMnet's training set. Thus ABS is able to reveal gaps in training data and potential failures in predicting under-represented body shapes. Results show that training BMnet with ABS improves measurement prediction accuracy on real bodies by up to 10%, when compared to no augmentation or random body shape sampling. Furthermore, our method significantly outperforms SOTA measurement estimation methods by as much as 3x. Finally, we release BodyM, the first challenging, large-scale dataset of photo silhouettes and body measurements of real human subjects, to further promote research in this area. Project website: https://adversarialbodysim.github.io.
Nataniel Ruiz, Miriam Bellver, Timo Bolkart, Ambuj Arora, Ming C. Lin, Javier Romero 0002, Raja Bala
3DV7
2022 Structural Prior Models for 3-D Deep Vessel Segmentation
abstract
We address the problem of 3-D blood vessel segmentation with a deep learning method that incorporates domain information via priors and regularizers on vessel structure and morphology. Inspired by the observation that 3-D vessel structures project onto 2-D image slices with distinctive edges that can aid 3-D vessel segmentation, we propose a novel multi-task learning architecture comprising a shared encoder and two decoders that respectively predict vessel segmentation maps and edge profiles. 3-D features from the two branches are concatenated to facilitate edge-guidance when learning segmentation maps. We introduce new regularization terms that encourage local homogeneity of 3-D blood vessel volumes brought about by biomarkers, as well as sparsity of edge pixels. Experiments on benchmark datasets demonstrate superior performance of our method over the state-of-the-art, especially when training data is limited.
Xuelu Li, Raja Bala, Vishal Monga
ICASSP2
2022 Robust Deep 3D Blood Vessel Segmentation Using Structural Priors
abstract
Deep learning has enabled significant improvements in the accuracy of 3D blood vessel segmentation. Open challenges remain in scenarios where labeled 3D segmentation maps for training are severely limited, as is often the case in practice, and in ensuring robustness to noise. Inspired by the observation that 3D vessel structures project onto 2D image slices with informative and unique edge profiles, we propose a novel deep 3D vessel segmentation network guided by edge profiles. Our network architecture comprises a shared encoder and two decoders that learn segmentation maps and edge profiles jointly. 3D context is mined in both the segmentation and edge prediction branches by employing bidirectional convolutional long-short term memory (BCLSTM) modules. 3D features from the two branches are concatenated to facilitate learning of the segmentation map. As a key contribution, we introduce new regularization terms that: a) capture the local homogeneity of 3D blood vessel volumes in the presence of biomarkers; and b) ensure performance robustness to domain-specific noise by suppressing false positive responses. Experiments on benchmark datasets with ground truth labels reveal that the proposed approach outperforms state-of-the-art techniques on standard measures such as DICE overlap and mean Intersection-over-Union. The performance gains of our method are even more pronounced when training is limited. Furthermore, the computational cost of our network inference is among the lowest compared with state-of-the-art.
Xuelu Li, Raja Bala, Vishal Monga
IEEE Trans. Image Process.2
2021 STRIVE: Scene Text Replacement In Videos
abstract
We propose replacing scene text in videos using deep style transfer and learned photometric transformations. Building on recent progress on still image text replacement, we present extensions that alter text while preserving the appearance and motion characteristics of the original video. Compared to the problem of still image text replacement, our method addresses additional challenges introduced by video, namely effects induced by changing lighting, motion blur, diverse variations in camera-object pose over time, and preservation of temporal consistency. We parse the problem into three steps. First, the text in all frames is normalized to a frontal pose using a spatio-temporal transformer network. Second, the text is replaced in a single reference frame using a state-of-art still-image text replacement method. Finally, the new text is transferred from the reference to remaining frames using a novel learned image transformation network that captures lighting and blur effects in a temporally consistent manner. Results on synthetic and challenging real videos show realistic text transfer, competitive quantitative and qualitative performance, and superior inference speed relative to alternatives. We introduce new synthetic and real-world datasets with paired text objects. To the best of our knowledge this is the first attempt at deep video text replacement.
Jeyasri Subramanian, Varnith Chordia, Eugene Bart, Shaobo Fang, Kelly Guan, Raja Bala
ICCV7
2020 Editing in Style: Uncovering the Local Semantics of GANs
abstract
While the quality of GAN image synthesis has improved tremendously in recent years, our ability to control and condition the output is still limited. Focusing on StyleGAN, we introduce a simple and effective method for making local, semantically-aware edits to a target output image. This is accomplished by borrowing elements from a source image, also a GAN output, via a novel manipulation of style vectors. Our method requires neither supervision from an external model, nor involves complex spatial morphing operations. Instead, it relies on the emergent disentanglement of semantic objects that is learned by StyleGAN during its training. Semantic editing is demonstrated on GANs producing human faces, indoor scenes, cats, and cars. We measure the locality and photorealism of the edits produced by our method, and find that it accomplishes both.
Edo Collins, Raja Bala, Bob Price, Sabine Süsstrunk
CVPR2
2020 Deep Retinal Image Segmentation With Regularization Under Geometric Priors
abstract
Vessel segmentation of retinal images is a key diagnostic capability in ophthalmology. This problem faces several challenges including low contrast, variable vessel size and thickness, and presence of interfering pathology such as micro-aneurysms and hemorrhages. Early approaches addressing this problem employed hand-crafted filters to capture vessel structures, accompanied by morphological post-processing. More recently, deep learning techniques have been employed with significantly enhanced segmentation accuracy. We propose a novel domain enriched deep network that consists of two components: 1) a representation network that learns geometric features specific to retinal images, and 2) a custom designed computationally efficient residual task network that utilizes the features obtained from the representation layer to perform pixel-level segmentation. The representation and task networks are jointly learned for any given training set. To obtain physically meaningful and practically effective representation filters, we propose two new constraints that are inspired by expected prior structure on these filters: 1) orientation constraint that promotes geometric diversity of curvilinear features, and 2) a data adaptive noise regularizer that penalizes false positives. Multi-scale extensions are developed to enable accurate detection of thin vessels. Experiments performed on three challenging benchmark databases under a variety of training scenarios show that the proposed prior guided deep network outperforms state of the art alternatives as measured by common evaluation metrics, while being more economical in network size and inference time.
Venkateswararao Cherukuri, Raja Bala, Vishal Monga
IEEE Trans. Image Process.3
2019 Multi-Scale Regularized Deep Network for Retinal Vessel Segmentation
abstract
Vessel segmentation of retinal images is a key diagnostic capability in ophthalmology. Early approaches addressing this problem employed hand-crafted filters to capture vessel structures, accompanied by morphological processing. More recently, deep learning techniques have been employed to significantly enhance segmentation accuracy. We propose a novel domain enriched deep network that consists of two components: 1) a representation network which learns geometric (specifically curvilinear) features that are tailored to retinal images, followed by 2) a task network that utilizes the features obtained from the representation layer to perform pixel-level segmentation. The representation and task networks are learned jointly for any given training set. To obtain effective representation filters, we develop a new orientation constraint that enables geometric diversity of curvilinear features. A multi-scale extension is further developed to enhance segmentation of thin vessels. Experiments performed on two challenging benchmark databases reveal that the proposed regularized deep network can outperform state of the art alternatives as measured by common evaluation metrics. Further, the proposed method exhibits a more graceful decay in performance as training data is reduced.
Venkateswararao Cherukuri, Raja Bala, Vishal Monga
ICIP3
2018 Deep Temporal Multimodal Fusion for Medical Procedure Monitoring Using Wearable Sensors
abstract
Process monitoring and verification have a wide range of uses in the medical and healthcare fields. Currently, such tasks are often carried out by a trained specialist, which makes them expensive, inefficient, and time-consuming. Recent advances in automated video- and multimodal-data-based action and activity recognition have made it possible to reduce the extent of manual intervention required to effectively carry out process supervision tasks. In this paper, we propose algorithms for automated egocentric human action and activity recognition from multimodal data, with a target application of monitoring and assisting a user perform a multistep medical procedure. We propose a supervised deep multimodal fusion framework that relies on concurrent processing of motion data acquired with wearable sensors and video data acquired with an egocentric or body-mounted camera. We demonstrate the effectiveness of the algorithm on a public multimodal dataset and conclude that automated process monitoring via the use of multiple heterogeneous sensors is a viable alternative to its manual counterpart. Furthermore, we demonstrate that the application of previously proposed adaptive sampling schemes to the video processing branch of the multimodal framework results in significant performance improvements.
Edgar A. Bernal, Xitong Yang, Qun Li 0003, Jayant Kumar, Sriganesh Madhvanath, Palghat Ramesh, Raja Bala
IEEE Trans. Multim.7
2016 A study on the discriminability of facs from spontaneous facial expressions
abstract
This paper investigates the discriminative capabilities of facial action units (AUs) exhibited by an individual while performing a task on a tablet computer in a semi-unconstrained environment. To that end, AUs are measured on a frame-by-frame basis from videos of 96 different subjects participating in a game-show-like quiz game that included a prize incentive. We propose a method that leverages the activation characteristics, as well as the temporal dynamics of facial behavior. In order to demonstrate the discriminative capabilities of the proposed approach, we perform identity matching across all subject pairs. Overall, the rank-1 matching performance of our algorithm ranges from 55% and up to 85%, on scenarios where the emotional disparity between the reference and query samples is largest and smallest, respectively. We believe these results represent a significant improvement relative to existing work relying on the use of AUs for human identification, in particular because the experimental settings guarantee that the facial expressions involved are spontaneous.
Matthew Shreve, Edgar A. Bernal, Qun Li 0003, Jayant Kumar, Raja Bala
ICIP5
2014 Low rank sparsity prior for robust video anomaly detection
abstract
Recently, sparsity based classification has been applied to video anomaly detection. A linear model is assumed over video features (e.g. trajectories) such that the feature representation of a new event is written as a sparse linear combination of existing feature representations in the dictionary. Sparsity based video anomaly detection shows promise but open challenges remain in that existing methods assume object specific and class specific event dictionaries making them applicable mostly in highly structured scenarios. Second, using conventional sparsity models on matrices/vectors, the computational burden is often high. In this work, we advocate a more general and practical sparsity model using a low-rank structure on the matrix of sparse coefficients. We find that enforcing a low-rank structure can ease the rigidity of traditional row-sparse constraints on sparse coefficient vectors/matrices. Because low-rank matrices are of course not always sparse, an additional l1regularization term is added. Further, if rank is substituted by its convex nuclear norm alternative, then significant computational benefits can be obtained over existing methods in sparsity based video anomaly detection. Experimental evaluation on benchmark video datasets reveal, our method is competitive with state-of-the art while providing robustness benefits under occlusion.
Xuan Mo, Vishal Monga, Raja Bala, Zhigang Fan 0001, Aaron M. Burry
ICASSP3
2014 Flash/no-flash fusion for mobile document image binarization
abstract
We propose a novel algorithm for improving the quality of binarized document images captured with a mobile device under low light conditions. In such scenarios, images captured without a flash often result in blur artifacts and poor signal-to-noise ratio, while images taken with the flash may produce information loss due to specular reflection in a localized region termed a “flash spot”. Our algorithm automatically triggers the capture of a pair of images, one with and one without flash, in rapid succession. The flash spot region (FSR) is first detected. The two images are then accurately aligned within the FSR using a multiscale alignment technique. Finally the images are binarized and fused in the vicinity of the FSR using an intelligent technique that minimizes fusion boundary artifacts. The result is a binary image that is largely identical to the binarized flash image, except within the FSR where content from the no-flash image is selectively incorporated. To our knowledge this is the first attempt to employ flash/no-flash fusion to improve binarization of mobile document images. Results show superior qualitative and quantitative performance of the proposed algorithm when compared with standard binarization applied to either the flash or no-flash image.
Jayant Kumar, Martin S. Maltz, Raja Bala
ICIP3
2014 Simultaneous sparsity model for multi-perspective video anomaly detection
abstract
Recently, sparsity based classification has been applied to video anomaly detection. A linear model is assumed over video features (e.g. trajectories) such that the feature representation of a new event is written as a sparse linear combination of existing feature representations in the dictionary. Sparsity based video anomaly detection has shown promise over alternate video anomaly detection methods in that the sparse representations exhibit excellent robustness under noise (common in surveillance videos) and missing or corrupted features, e.g. vehicle occlusion in transportation videos. One limitation of existing sparsity based video anomaly detection techniques is that they are based on only a single feature representation (known formally as video event encoding). One can easily envision that different event representations such as object trajectories and spatio-temporal volumes often contain correlated yet complementary information. In this paper, we propose to extend sparsity models based on single feature representations to simultaneous sparse representations of multiple feature representations. In this model, the matrix of sparse coefficients does not confirm to the commonly seen row-sparsity and a modified greedy heuristic approach that extends simultaneous orthogonal matching pursuit (SOMP) is needed to solve the resulting optimization problem. Experiments on two benchmark video datasets reveal that our method significantly outperforms state-of-the art approaches that utilize only a single-perspective or event encoding.
Xuan Mo, Vishal Monga, Raja Bala
ICIP3
2014 Adaptive Sparse Representations for Video Anomaly Detection
abstract
Video anomaly detection can be used in the transportation domain to identify unusual patterns such as traffic violations, accidents, unsafe driver behavior, street crime, and other suspicious activities. A common class of approaches relies on object tracking and trajectory analysis. Very recently, sparse reconstruction techniques have been employed in video anomaly detection. The fundamental underlying assumption of these methods is that any new feature representation of a normal/anomalous event can be approximately modeled as a (sparse) linear combination prelabeled feature representations (of previously observed events) in a training dictionary. Sparsity can be a powerful prior on model coefficients but challenges remain in the detection of anomalies involving multiple objects and the ability of the linear sparsity model to effectively allow for class separation. The proposed research addresses both these issues. First, we develop a new joint sparsity model for anomaly detection that enables the detection of joint anomalies involving multiple objects. This extension is highly nontrivial since it leads to a new simultaneous sparsity problem that we solve using a greedy pursuit technique. Second, we introduce nonlinearity into, that is, kernelize. The linear sparsity model to enable superior class separability and hence anomaly detection. We extensively test on several real world video datasets involving both single and multiple object anomalies. Results show marked improvements in detection of anomalies in both supervised and unsupervised scenarios when using the proposed sparsity models.
Xuan Mo, Vishal Monga, Raja Bala, Zhigang Fan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2012 Design and Optimization of Color Lookup Tables on a Simplex Topology
abstract
An important computational problem in color imaging is the design of color transforms that map color between devices or from a device-dependent space (e.g., RGB/CMYK) to a device-independent space (e.g., CIELAB) and vice versa. Real-time processing constraints entail that such nonlinear color transforms be implemented using multidimensional lookup tables (LUTs). Furthermore, relatively sparse LUTs (with efficient interpolation) are employed in practice because of storage and memory constraints. This paper presents a principled design methodology rooted in constrained convex optimization to design color LUTs on a simplex topology. The use of n simplexes, i.e., simplexes in n dimensions, as opposed to traditional lattices, recently has been of great interest in color LUT design for simplex topologies that allow both more analytically tractable formulations and greater efficiency in the LUT. In this framework of n-simplex interpolation, our central contribution is to develop an elegant iterative algorithm that jointly optimizes the placement of nodes of the color LUT and the output values at those nodes to minimize interpolation error in an expected sense. This is in contrast to existing work, which exclusively designs either node locations or the output values. We also develop new analytical results for the problem of node location optimization, which reduces to constrained optimization of a large but sparse interpolation matrix in our framework. We evaluate our n -simplex color LUTs against the state-of-the-art lattice (e.g., International Color Consortium profiles) and simplex-based techniques for approximating two representative multidimensional color transforms that characterize a CMYK xerographic printer and an RGB scanner, respectively. The results show that color LUTs designed on simplexes offer very significant benefits over traditional lattice-based alternatives in improving color transform accuracy even with a much smaller number of nodes.
Vishal Monga, Raja Bala, Xuan Mo
IEEE Trans. Image Process.2
2010 Algorithms for color look-up-table (LUT) design via joint optimization of node locations and output values
abstract
Real-time processing constraints entail that non-linear color transforms be implemented using multi-dimensional look-up-tables (LUT). Further, relatively sparse LUTs (with efficient interpolation) are employed in practice because of storage and memory constraints. Much research has been devoted towards optimizing “nodes” (or equivalently partitioning the input color space) of this color LUT based on the curvature of the color transform to be processed through the LUT. Likewise, for a given LUT structure, the optimization of transform output values has been suggested so as to minimize interpolation error in an expected sense even if the values stored in the LUT do not agree with true transform output values. This paper presents a principled algorithmic approach to combine the merits of these two complementary techniques. The error (cost) function does not exhibit joint convexity over the multidimensional variable sets of node locations and corresponding output values which makes this optimization particularly challenging. The paper makes two significant contributions: 1.) for the case of simplex interpolation, a cost function is formulated that exhibits separable convexity in its arguments and enables an efficient alternating convex optimization algorithm, and 2.) in the aforementioned framework, for fixed node outputs, the optimization of node locations is split into a primary and an auxiliary optimization, which greatly improves the quality of the solution over traditional alternatives where node locations are directly optimized. Preliminary experiments show remarkable improvements in color transform accuracy over what is obtained by individually optimizing just the node locations or output values.
Vishal Monga, Raja Bala
ICASSP2
2005 Two-dimensional transforms for device color correction and calibration
abstract
Color device calibration is traditionally performed using one-dimensional (1-D) per-channel tone-response corrections (TRCs). While 1-D TRCs are attractive in view of their low implementation complexity and efficient real-time processing of color images, their use severely restricts the degree of control that can be exercised along various device axes. A typical example is that per separation (or per-channel), TRCs in a printer can be used to either ensure gray balance along the C = M = Y axis or to provide a linear response in delta-E units along each of the individual (C, M, and Y) axis, but not both. This paper proposes a novel two-dimensional color correction architecture that enables much greater control over the device color gamut with a modest increase in implementation cost. Results show significant improvement in calibration accuracy and stability when compared to traditional 1-D calibration. Superior cost quality tradeoffs (over 1-D methods) are also achieved for emulation of one color device on another.
Raja Bala, Gaurav Sharma 0001, Vishal Monga, Jean-Pierre Van de Capelle
IEEE Trans. Image Process.1
2002 Detection and segmentation of sweeps in color graphics images
abstract
Business graphics are an important class of digital imagery. Such images are computer-generated, and comprise synthetic elements such as solid fills, line art, and color sweeps. Often these images are first printed and then scanned for further electronic reuse. The printing and scanning process destroys the synthetic structure of a graphics image, and furthermore introduces distortions due to halftoning and other forms of printer and scanner noise. Subsequent reproductions usually amplify these distortions thus resulting in rapid degradation of image quality. It would thus be desirable to detect and reconstruct the original synthetic structure from the scanned image. This paper presents an effort in this direction, namely a method to detect color sweeps in scanned images. Once detected, the synthetic signature of the sweep is derived, namely its starting and ending color. This information can be used to optimize subsequent image processing operations such as rendering to an output device, or image compression. This work represents a novel application of known image processing techniques to extract semantic information from graphics images.
Salil Prabhakar, Raja Bala, John C. Handley, Ying-wei Lin
ICIP (3)3
1995 Sequential scalar quantization of vectors: an analysis
abstract
Proposes an efficient vector quantization (VQ) technique called sequential scalar quantization (SSQ). The scalar components of the vector are individually quantized in a sequence, with the quantization of each component utilizing conditional information from the quantization of previous components. Unlike conventional independent scalar quantization (ISQ), SSQ has the ability to exploit intercomponent correlation. At the same time, since quantization is performed on scalar rather than vector variables, SSQ offers a significant computational advantage over conventional VQ techniques and is easily amenable to a hardware implementation. In order to analyze the performance of SSQ, the authors appeal to asymptotic quantization theory, where the codebook size is assumed to be large. Closed-form expressions are derived for the quantizer mean squared error (MSE). These expressions are used to compare the asymptotic performance of SSQ with other VQ techniques. The authors also demonstrate the use of asymptotic theory in designing SSQ for a practical application (color image quantization), where the codebook size is typically small. Theoretical and experimental results show that SSQ far outperforms ISQ with respect to MSE while offering a considerable reduction in computation over conventional VQ at the expense of a moderate increase in MSE.
Raja Bala, Charles A. Bouman, Jan P. Allebach
IEEE Trans. Image Process.1