Sundaram Muthu

dblp:264/6819 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0003-3023-5372ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 53% Transfer learning and domain adaptation · 12% Segmentation and scene understanding · 10%
Computer graphics and multimedia
2 papers
Rendering · 82% Geometric modeling and processing · 18%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Rendering › gaussian splatting
2d gaussian splatting
0.912025
GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction · CVPR 2025
Geometric modeling and processing
3d reconstruction
0.912025
GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction · CVPR 2025
Rendering
gaussian splatting
0.912025
GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction · CVPR 2025
Rendering › inverse rendering
reflective object reconstruction
0.912025
GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction · CVPR 2025
Computer vision › 3D vision › visual localization
cross-view geo-localization
0.812024
View from Above: Orthogonal-View Aware Cross-View Localization · CVPR 2024
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.812024
View from Above: Orthogonal-View Aware Cross-View Localization · CVPR 2024
Computer vision › 3D vision
visual localization
0.812024
View from Above: Orthogonal-View Aware Cross-View Localization · CVPR 2024
Computer vision › 3D vision › 3d reconstruction
neural reconstruction
0.712023
Seeing Through the Glass: Neural 3D Reconstruction of Object Inside a Transparent Container · CVPR 2023
Robotics › Autonomous driving › perception › environment perception
perception for self-driving vehicles
0.712023
Homography Guided Temporal Fusion for Road Line and Marking Segmentation · ICCV 2023
Computer vision › 3D vision › 3d reconstruction
transparent object reconstruction
0.712023
Seeing Through the Glass: Neural 3D Reconstruction of Object Inside a Transparent Container · CVPR 2023
Rendering
hybrid rendering
0.712023
Seeing Through the Glass: Neural 3D Reconstruction of Object Inside a Transparent Container · CVPR 2023
Rendering
volume rendering
0.712023
Seeing Through the Glass: Neural 3D Reconstruction of Object Inside a Transparent Container · CVPR 2023
Computer vision › Video understanding and tracking
motion segmentation
0.412020
Motion Segmentation of RGB-D Sequences: Combining Semantic and Motion Information Using Statistical Inference · IEEE Trans. Image Process. 2020
Computer vision › 3D vision › 3d reconstruction
surface reconstruction
0.312025
GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction · CVPR 2025
Computer vision › 3D vision › object pose estimation
vehicle pose estimation
0.212024
View from Above: Orthogonal-View Aware Cross-View Localization · CVPR 2024
Computer vision › Video understanding and tracking › temporal modeling
temporal fusion
0.212023
Homography Guided Temporal Fusion for Road Line and Marking Segmentation · ICCV 2023
Robotics › Robot navigation and mapping › SLAM › visual SLAM
RGB-D SLAM
0.112020
Motion Segmentation of RGB-D Sequences: Combining Semantic and Motion Information Using Statistical Inference · IEEE Trans. Image Process. 2020
Robotics › Robot navigation and mapping
SLAM
0.112020
Motion Segmentation of RGB-D Sequences: Combining Semantic and Motion Information Using Statistical Inference · IEEE Trans. Image Process. 2020

Methods — techniques the papers use, named apart from their topics

signed distance function · 1.7gaussian splatting · 1.7foundation model geometric supervision · 1.7volume rendering · 1.3ray tracing · 1.3neural implicit surface reconstruction · 1.3top-to-ground aggregation · 0.8equidistant re-projection loss · 0.8pixel-to-pixel attention · 0.7homography estimation · 0.7
YearPublicationVenuePosition
2026 MFGS: Mask-free Gaussian separation for 3D object reconstruction
abstract
Accurate 3D reconstruction from multi-view images is a fundamental problem in computer vision. A common acquisition strategy involves placing an object on a rotating turntable while moving the camera to capture it from various viewpoints. In such scenarios, object moves relative to the background, many existing reconstruction methods rely on object masks to separate the foreground from the background. The quality of these masks significantly affects the final reconstruction, yet obtaining high-quality and consistent masks is a challenging and laborious process, especially when controlled environments like green screens are unavailable. To address this limitation, we introduce Mask-free Gaussian Separation (MFGS), a novel method that performs simultaneous object reconstruction and segmentation without requiring any input masks. Our approach builds on Gaussian Splatting and automatically disentangles the scene by extending each Gaussian primitive with a learnable parameter that represents its probability of belonging to the dynamic foreground object. This separation is optimized in a self-supervised manner, optimized by the object and camera transformation constraints. We evaluated MFGS on new synthetic and real-world datasets designed to reflect this challenging capture scenario. Experimental results demonstrate that our mask-free approach significantly outperforms existing methods. Notably, MFGS surpasses the performance of the state-of-the-art method(2DGS) that relies on high-quality segmentation masks, achieving a 27% improvement in novel view synthesis and a 7% improvement in geometry reconstruction.
Jinguang Tong, Xuesong Li 0001, Sundaram Muthu, Fahira A. Maken, Lars Petersson, Hongdong Li
Pattern Recognit.3
2025 GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction
abstract
3D modeling of highly reflective objects remains challenging due to strong view-dependent appearances. While previous SDF-based methods can recover high-quality meshes, they are often time-consuming and tend to produce over-smoothed surfaces. In contrast, 3D Gaussian Splatting (3DGS) offers the advantage of high speed and detailed real-time rendering, but extracting surfaces from the Gaussians can be noisy due to the lack of geometric constraints. To bridge the gap between these approaches, we propose a novel reconstruction method called GS-2DGS for reflective objects based on 2D Gaussian Splatting (2DGS). Our approach combines the rapid rendering capabilities of Gaussian Splatting with additional geometric information from foundation models. Experimental results on synthetic and real datasets demonstrate that our method significantly outperforms Gaussian-based techniques in terms of reconstruction and relighting and achieves performance comparable to SDF-based methods while being an order of magnitude faster. Code is available at https://github.com/hirotong/GS2DGS
Jinguang Tong, Xuesong Li 0001, Fahira A. Maken, Sundaram Muthu, Lars Petersson, Hongdong Li
CVPR4
2024 View from Above: Orthogonal-View Aware Cross-View Localization
abstract
This paper presents a novel aerial-to-ground feature ag-gregation strategy, tailored for the task of cross- view image-based geo-localization. Conventional vision-based methods heavily rely on matching ground-view image features with a pre-recorded image database, often through establishing planar homography correspondences via a planar ground assumption. As such, they tend to ignore features that are off-ground and not suited for handling visual occlusions, leading to unreliable localization in challenging scenarios. We propose a Top-to-Ground Aggregation (T2GA) module that capitalizes aerial orthographic views to aggregate features down to the ground level, leveraging reliable off-ground information to improve feature alignment. Furthermore, we introduce a Cycle Domain Adaptation (CycDA) loss that ensures feature extraction robustness across do-main changes. Additionally, an Equidistant Re-projection (ERP) loss is introduced to equalize the impact of all key-points on orientation error, leading to a more extended distribution of keypoints which benefits orientation estimation. On both KITTI and Ford Multi-AV datasets, our method consistently achieves the lowest mean longitudinal and lateral translations across different settings and obtains the smallest orientation error when the initial pose is less ac-curate, a more challenging setting. Further, it can complete an entire route through continual vehicle pose estimation with initial vehicle pose given only at the starting point.11Code is available at https://github.com/ShanWang-Shan/View FromAbove.
Shan Wang 0010, Jiawei Liu 0005, Yanhao Zhang 0003, Sundaram Muthu, Fahira A. Maken, Kaihao Zhang, Hongdong Li
CVPR5
2023 Seeing Through the Glass: Neural 3D Reconstruction of Object Inside a Transparent Container
abstract
In this paper, we define a new problem of recovering the 3D geometry of an object confined in a transparent enclosure. We also propose a novel method for solving this challenging problem. Transparent enclosures pose challenges of multiple light reflections and refractions at the interface between different propagation media e.g. air or glass. These multiple reflections and refractions cause serious image distortions which invalidate the single viewpoint assumption. Hence the 3D geometry of such objects cannot be reliably reconstructed using existing methods, such as traditional structure from motion or modern neural reconstruction methods. We solve this problem by explicitly modeling the scene as two distinct sub-spaces, inside and outside the transparent enclosure. We use an existing neural reconstruction method (NeuS) that implicitly represents the geometry and appearance of the inner subspace. In order to account for complex light interactions, we develop a hybrid rendering strategy that combines volume rendering with ray tracing. We then recover the underlying geometry and appearance of the model by minimizing the difference between the real and rendered images. We evaluate our method on both synthetic and real data. Experiment results show that our method outperforms the state-of-the-art (SOTA) methods. Codes and data will be available at https://github.com/hirotong/ReNeuS
Jinguang Tong, Sundaram Muthu, Fahira A. Maken, Hongdong Li
CVPR2
2023 Homography Guided Temporal Fusion for Road Line and Marking Segmentation
abstract
Reliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occluded in the presence of moving vehicles, shadow, and glare and (2) highly structured with low intra-class shape variance and overall high appearance consistency. To solve these issues, we propose a Homography Guided Fusion (HomoFusion) module to exploit temporally-adjacent video frames for complementary cues facilitating the correct classification of the partially occluded road lines or markings. To reduce computational complexity, a novel surface normal estimator is proposed to establish spatial correspondences between the sampled frames, allowing the HomoFusion module to perform a pixel-to-pixel attention mechanism in updating the representation of the occluded road lines or markings. Experiments on ApolloScape, a large-scale lane mark segmentation dataset, and ApolloScape Night with artificial simulated night-time road conditions, demonstrate that our method outperforms other existing SOTA lane mark segmentation models with less than 9% of their parameters and computational complexity. We show that exploiting available camera intrinsic data and ground plane assumption for cross-frame correspondence can lead to a light-weight network with significantly improved performances in speed and accuracy. We also prove the versatility of our HomoFusion approach by applying it to the problem of water puddle segmentation and achieving SOTA performance1.
Shan Wang 0010, Jiawei Liu 0005, Kaihao Zhang, Wenhan Luo, Yanhao Zhang 0003, Sundaram Muthu, Fahira A. Maken, Hongdong Li
ICCV7
2023 Generalized framework for image and video object segmentation using affinity learning and message passing GNNS
abstract
Despite significant amount of work reported in the computer vision literature, segmenting images or videos based on multiple cues such as objectness, texture and motion, is still a challenge. This is particularly true when the number of objects to be segmented is not known or there are objects that are not classified in the training data (unknown objects). A possible remedy to this problem is to utiize graph-based clustering techniques such as Correlation Clustering. It is known that using long range affinities (Lifted multicut), makes correlation clustering more accurate than using only adjacent affinities (Multicut). However, the former is computationally expensive and hard to use. In this paper, we introduce a new framework to perform image/motion segmentation using an affinity learning module and a Message Passing Graph Neural Network (MPGNN). The affinity learning module uses a permutation invariant affinity representation to overcome the multi-object problem. The paper shows, both theoretically and empirically, that the proposed MPGNN aggregates higher order information and thereby converts the Lifted Multicut Problem (LMP) to a Multicut Problem (MP), which is easier and faster to solve. Importantly, the proposed method can be generalized to deal with different clustering problems with the same MPGNN architecture. For instance, our method produces competitive results for single image segmentation (on BSDS dataset) as well as unsupervised video object segmentation (on DAVIS17 dataset), by only changing the feature extraction part. In addition, using an ablation study on the proposed MPGNN architecture, we show that the way we update the parameterized affinities directly contributes to the accuracy of the results.
Sundaram Muthu, Ruwan B. Tennakoon, Tharindu Rathnayake, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
Comput. Vis. Image Underst.1
2020 Motion Segmentation of RGB-D Sequences: Combining Semantic and Motion Information Using Statistical Inference
abstract
This paper presents an innovative method for motion segmentation in RGB-D dynamic videos with multiple moving objects. The focus is on finding static, small or slow moving objects (often overlooked by other methods) that their inclusion can improve the motion segmentation results. In our approach, semantic object based segmentation and motion cues are combined to estimate the number of moving objects, their motion parameters and perform segmentation. Selective object-based sampling and correspondence matching are used to estimate object specific motion parameters. The main issue with such an approach is the over segmentation of moving parts due to the fact that different objects can have the same motion (e.g. background objects). To resolve this issue, we propose to identify objects with similar motions by characterizing each motion by a distribution of a simple metric and using a statistical inference theory to assess their similarities. To demonstrate the significance of the proposed statistical inference, we present an ablation study, with and without static objects inclusion, on SLAM accuracy using the TUM-RGBD dataset. To test the effectiveness of the proposed method for finding small or slow moving objects, we applied the method to RGB-D MultiBody and SBM-RGBD motion segmentation datasets. The results showed that we can improve the accuracy of motion segmentation for small objects while remaining competitive on overall measures.
Sundaram Muthu, Ruwan B. Tennakoon, Tharindu Rathnayake, Reza Hoseinnezhad, David Suter, Alireza Bab-Hadiashar
IEEE Trans. Image Process.1