Sk Aziz Ali

dblp:191/4692 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-6396-8436ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 76% Generative modeling · 21% Image recognition and object detection · 3%
Computer graphics and multimedia
3 papers
Geometric modeling and processing · 72% Visual content generation and editing · 28%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › pose estimation
hand pose and shape estimation
1.022022
HandVoxNet++: 3D Hand Shape and Pose Estimation Using Voxel-Based Neural Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2022
HandVoxNet: Deep Voxel-Based Network for 3D Hand Shape and Pose Estimation From a Single Depth Map · CVPR 2020
Visual content generation and editing › 3d content generation
text-to-3d generation
0.912025
MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation · CVPR 2025
Machine learning › Generative modeling
autoregressive model
0.812024
Text2CAD: Generating Sequential CAD Designs from Beginner-to-Expert Level Text Prompts · NeurIPS 2024
Geometric modeling and processing › reverse engineering
CAD model reconstruction
0.812024
CAD-SIGNet: CAD Language Inference from Point Clouds Using Layer-Wise Sketch Instance Guided Attention · CVPR 2024
Geometric modeling and processing
computer-aided design
0.812024
Text2CAD: Generating Sequential CAD Designs from Beginner-to-Expert Level Text Prompts · NeurIPS 2024
Geometric modeling and processing
reverse engineering
0.812024
CAD-SIGNet: CAD Language Inference from Point Clouds Using Layer-Wise Sketch Instance Guided Attention · CVPR 2024
Computer vision › 3D vision
point cloud registration
0.512021
RPSRNet: End-to-End Trainable Rigid Point Set Registration Network Using Barnes-Hut 2D-Tree Representation · CVPR 2021
Computer vision › 3D vision › geometric estimation › registration › rigid registration
rigid point set registration
0.512021
RPSRNet: End-to-End Trainable Rigid Point Set Registration Network Using Barnes-Hut 2D-Tree Representation · CVPR 2021
Computer vision › 3D vision › point cloud registration
point set registration
0.212016
Gravitational Approach for Point Set Registration · CVPR 2016
Computer vision › 3D vision › geometric estimation › registration
rigid registration
0.212016
Gravitational Approach for Point Set Registration · CVPR 2016
Machine learning › Generative modeling › cross-modal generation
text-conditioned generation
0.212024
Text2CAD: Generating Sequential CAD Designs from Beginner-to-Expert Level Text Prompts · NeurIPS 2024
Computer vision › 3D vision › geometric estimation › 3d registration
mesh registration
0.212022
HandVoxNet++: 3D Hand Shape and Pose Estimation Using Voxel-Based Neural Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Computer vision › Image recognition and object detection › point set representation
point cloud representation
0.112021
RPSRNet: End-to-End Trainable Rigid Point Set Registration Network Using Barnes-Hut 2D-Tree Representation · CVPR 2021

Methods — techniques the papers use, named apart from their topics

vision-language model · 1.7stable diffusion · 1.7large language model · 1.7image-to-3d · 1.7transformer · 1.5data annotation pipeline · 1.5autoregressive generation · 1.5point cloud · 0.8cross-attention · 0.8autoregressive model · 0.8non-rigid registration · 0.6graph convolution · 0.63d convolution · 0.6
YearPublicationVenuePosition
2025 MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation
abstract
Generating high-fidelity 3D content from text prompts remains a significant challenge in computer vision due to the limited size, diversity, and annotation depth of the existing datasets. To address this, we introduce MARVEL-40M+, an extensive dataset with 40 million text annotations for over 8.9 million 3D assets aggregated from seven major 3D datasets. Our contribution is a novel multi-stage annotation pipeline that integrates open-source pretrained multi-view VLMs and LLMs to automatically produce multi-level descriptions, ranging from detailed (150-200 words) to concise semantic tags (10-20 words). This structure supports both fine-grained 3D reconstruction and rapid prototyping. Furthermore, we incorporate human metadata from source datasets into our annotation pipeline to add domain-specific information in our annotation and reduce VLM hallucinations. Additionally, we develop MARVEL-FX3D, a two-stage text-to-3D pipeline. We fine-tune Stable Diffusion with our annotations and use a pretrained image-to-3D network to generate 3D textured meshes within 15s. Extensive evaluations show that MARVEL-40M+ significantly outperforms existing datasets in annotation quality and linguistic diversity, achieving win rates of 72.41% by GPT-4 and 73.40% by human evaluators. Project page is available at https://sankalpsinha-cmos.github.io/MARVEL/.
Sankalp Sinha, Mohammad Sadil Khan, Shino Sam, Didier Stricker, Sk Aziz Ali, Muhammad Zeshan Afzal
CVPR6
2024 CAD-SIGNet: CAD Language Inference from Point Clouds Using Layer-Wise Sketch Instance Guided Attention
abstract
Reverse engineering in the realm of Computer-Aided Design (CAD) has been a longstanding aspiration, though not yet entirely realized. Its primary aim is to uncover the CAD process behind a physical object given its 3D scan. We propose CAD-SIGNet, an end-to-end trainable and aetoregressive architecture to recover the design history of a CAD model represented as a sequence of sketch-and-extrusion from an input point cloud. Our model learns CAD visual-language representations by layer-wise crossattention between point cloud and CAD language embedding. In particular, a new Sketch instance Guided Attention (SGA) module is proposed in order to reconstruct the finegrained details of the sketches. Thanks to its auto-regressive nature, CAD-SIGNet not only reconstructs a unique full design history of the corresponding CAD model given an input point cloud but also provides multiple plausible design choices. This allows for an interactive reverse engineering scenario by providing designers with multiple next step choices along with the design process. Extensive experiments on publicly available CAD datasets showcase the effectiveness of our approach against existing baseline models in two settings, namely, full design history recovery and conditional auto-completion from point clouds.
Mohammad Sadil Khan, Elona Dupont, Sk Aziz Ali, Kseniya Cherenkova, Anis Kacem 0001, Djamila Aouada
CVPR3
2024 Text2CAD: Generating Sequential CAD Designs from Beginner-to-Expert Level Text Prompts
abstract
Prototyping complex computer-aided design (CAD) models in modern softwares can be very time-consuming. This is due to the lack of intelligent systems that can quickly generate simpler intermediate parts. We propose Text2CAD, the first AI framework for generating text-to-parametric CAD models using designer-friendly instructions for all skill levels. Furthermore, we introduce a data annotation pipeline for generating text prompts based on natural language instructions for the DeepCAD dataset using Mistral and LLaVA-NeXT. The dataset contains $\sim170$K models and $\sim660$K text annotations, from abstract CAD descriptions (e.g., _generate two concentric cylinders_) to detailed specifications (e.g., _draw two circles with center_ $(x,y)$ and _radius_ $r_{1}$, $r_{2}$, \textit{and extrude along the normal by} $d$...). Within the Text2CAD framework, we propose an end-to-end transformer-based auto-regressive network to generate parametric CAD models from input texts. We evaluate the performance of our model through a mixture of metrics, including visual quality, parametric precision, and geometrical accuracy. Our proposed framework shows great potential in AI-aided design applications. Project page is available at https://sadilkhan.github.io/text2cad-project/.
Mohammad Sadil Khan, Sankalp Sinha, Talha Uddin Sheikh, Didier Stricker, Sk Aziz Ali, Muhammad Zeshan Afzal
NeurIPS5
2022 CADOps-Net: Jointly Learning CAD Operation Types and Steps from Boundary-Representations
abstract
3D reverse engineering is a long sought-after, yet not completely achieved goal in the Computer-Aided Design (CAD) industry. The objective is to recover the construction history of a CAD model. Starting from a Boundary Representation (B-Rep) of a CAD model, this paper proposes a new deep neural network, CADOps-Net, that jointly learns the CAD operation types and the decomposition into different CAD operation steps. This joint learning allows to divide a B-Rep into parts that were created by various types of CAD operations at the same construction step; therefore providing relevant information for further recovery of the design history. Furthermore, we propose the novel CC3D-Ops dataset that includes over 37k CAD models annotated with CAD operation type labels and step labels. Compared to existing datasets, the complexity and variety of CC3D-Ops models are closer to those used for industrial purposes. Our experiments, conducted on the proposed CC3D-Ops and the publicly available Fusion360 datasets, demonstrate the competitive performance of CADOps-Net with respect to state-of-the-art, and confirm the importance of the joint learning of CAD operation types and steps.
Elona Dupont, Kseniya Cherenkova, Anis Kacem 0001, Sk Aziz Ali, Ilya Arzhannikov, Gleb Gusev, Djamila Aouada
3DV4
2022 HandVoxNet++: 3D Hand Shape and Pose Estimation Using Voxel-Based Neural Networks
abstract
3D hand shape and pose estimation from a single depth map is a new and challenging computer vision problem with many applications. Existing methods addressing it directly regress hand meshes via 2D convolutional neural networks, which leads to artifacts due to perspective distortions in the images. To address the limitations of the existing methods, we develop HandVoxNet++, i.e., a voxel-based deep network with 3D and graph convolutions trained in a fully supervised manner. The input to our network is a 3D voxelized-depth-map-based on the truncated signed distance function (TSDF). HandVoxNet++ relies on two hand shape representations. The first one is the 3D voxelized grid of hand shape, which does not preserve the mesh topology and which is the most accurate representation. The second representation is the hand surface that preserves the mesh topology. We combine the advantages of both representations by aligning the hand surface to the voxelized hand shape either with a new neural Graph-Convolutions-based Mesh Registration (GCN-MeshReg) or classical segment-wise Non-Rigid Gravitational Approach (NRGA++) which does not rely on training data. In extensive evaluations on three public benchmarks, i.e., SynHand5M, depth-based HANDS19 challenge and HO-3D, the proposed HandVoxNet++ achieves the state-of-the-art performance. In this journal extension of our previous approach presented at CVPR 2020, we gain 41.09% and 13.7% higher shape alignment accuracy on SynHand5M and HANDS19 datasets, respectively. Our method is ranked first on the HANDS19 challenge dataset (Task 1: Depth-Based 3D Hand Pose Estimation) at the moment of the submission of our results to the portal in August 2020.
Jameel Malik, Soshi Shimada, Ahmed Elhayek, Sk Aziz Ali, Christian Theobalt, Vladislav Golyanik, Didier Stricker
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 RPSRNet: End-to-End Trainable Rigid Point Set Registration Network Using Barnes-Hut 2D-Tree Representation
abstract
We propose RPSRNet - a novel end-to-end trainable deep neural network for rigid point set registration. For this task, we use a novel 2D-tree representation for the input point sets and a hierarchical deep feature embedding in the neural network. An iterative transformation refinement module of our network boosts the feature matching accuracy in the intermediate stages. We achieve an inference speed of ∼12-15 ms to register a pair of input point clouds as large as ∼250K. Extensive evaluations on (i) KITTI LiDAR-odometry and (ii) ModelNet-40 datasets show that our method outperforms prior state-of-the-art methods – e.g., on the KITTI dataset, DCP-v2 by 1.3 and 1.5 times, and PointNetLK by 1.8 and 1.9 times better rotational and translational accuracy respectively. Evaluation on ModelNet40 shows that RPSRNet is more robust than other benchmark methods when the samples contain a significant amount of noise and disturbance. RPSRNet accurately registers point clouds with non-uniform sampling densities, e.g., LiDAR data, which cannot be processed by many existing deep-learning-based registration methods.
Sk Aziz Ali, Kerem Kahraman, Gerd Reis, Didier Stricker
CVPR1
2020 HandVoxNet: Deep Voxel-Based Network for 3D Hand Shape and Pose Estimation From a Single Depth Map
abstract
3D hand shape and pose estimation from a single depth map is a new and challenging computer vision problem with many applications. The state-of-the-art methods directly regress 3D hand meshes from 2D depth images via 2D convolutional neural networks, which leads to artefacts in the estimations due to perspective distortions in the images. In contrast, we propose a novel architecture with 3D convolutions trained in a weakly-supervised manner. The input to our method is a 3D voxelized depth map, and we rely on two hand shape representations. The first one is the 3D voxelized grid of the shape which is accurate but does not preserve the mesh topology and the number of mesh vertices. The second representation is the 3D hand surface which is less accurate but does not suffer from the limitations of the first representation. We combine the advantages of these two representations by registering the hand surface to the voxelized hand shape. In the extensive experiments, the proposed approach improves over the state of the art by47.8% on the SynHand5M dataset. Moreover, our augmentation policy for voxelized depth maps further enhances the accuracy of 3D hand pose estimation on real data. Our method produces visually more reasonable and realistic hand shapes on NYU and BigHand2.2M datasets compared to the existing approaches.
Jameel Malik, Ibrahim Abdelaziz, Ahmed Elhayek, Soshi Shimada, Sk Aziz Ali, Vladislav Golyanik, Christian Theobalt, Didier Stricker
CVPR5
2020 Foldmatch: Accurate and High Fidelity Garment Fitting Onto 3D Scans
abstract
In this paper, we propose a new template fitting method that can capture fine details of garments in target 3D scans of dressed human bodies. Matching the high fidelity details of such loose/tight-fit garments is a challenging task as they express intricate folds, creases, wrinkle patterns, and other high fidelity surface details. Our proposed method of non-rigid shape fitting - FoldMatch - uses physics-based particle dynamics to explicitly model the deformation of loose-fit garments and wrinkle vector fields for capturing clothing details. The 3D scan point cloud behaves as a collection of astrophysical particles, which attracts the points in template mesh and defines the template motion model. We use this point-based motion model to derive regularized deformation gradients for the template mesh. We show the parameterization of the wrinkle vector fields helps in the accurate shape fitting. Our method shows better performance than the state-of-the-art methods. We define several deformation and shape matching quality measurement metrics to evaluate FoldMatch on synthetic and real data sets.
Sk Aziz Ali, Sikang Yan, Wolfgang Dornisch, Didier Stricker
ICIP1
2018 NRGA: Gravitational Approach for Non-rigid Point Set Registration
abstract
Recovery of correspondences between point sets which differ by some non-rigid transformation is an ill-posed problem. Many existing methods underperform on noisy or corrupted input data. In this study, a novel physics-based approach -- Non-Rigid Gravitational Approach (NRGA) -- for non-rigid point set registration is introduced which is robust to the mentioned artifacts. Thereafter, a distributed N-body simulation and iterative Procrustes alignment non-rigidly transform and register the template point set. Furthermore, in the force field evolution, per-point Gaussian curvature serves as a shape matching descriptor whereas the displacement fields are regularized by coherent collective motion. The optimal alignment is referred to as the state of minimum gravitational potential energy between the point sets. A thorough experimental evaluation and comparison are provided with widely used state-of-the-art methods on 2D and 3D data sets. Experiments show NRGA's robustness against uniform outliers and missing data.
Sk Aziz Ali, Vladislav Golyanik, Didier Stricker
3DV1
2016 Gravitational Approach for Point Set Registration
abstract
In this paper a new astrodynamics inspired rigid point set registration algorithm is introduced-the Gravitational Approach (GA). We formulate point set registration as a modified N-body problem with additional constraints and obtain an algorithm with unique properties which is fully scalable with the number of processing cores. In GA, a template point set moves in a viscous medium under gravitational forces induced by a reference point set. Pose updates are completed by numerically solving the differential equations of Newtonian mechanics. We discuss techniques for efficient implementation of the new algorithm and evaluate it on several synthetic and real-world scenarios. GA is compared with the widely used Iterative Closest Point and the state of the art rigid Coherent Point Drift algorithms. Experiments evidence that the new approach is robust against noise and can handle challenging scenarios with structured outliers.
Vladislav Golyanik, Sk Aziz Ali, Didier Stricker
CVPR2