Ali Mahdavi-Amiri

dblp:33/10499 · also Ali Mahdavi Amiri · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0002-4693-3565ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 17 · 16 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MultiCOIN: Multi-Modal COntrollable Inbetweening
abstract
Abstract Video inbetweening creates smooth transitions between two frames making it an indispensable tool for video editing and longform video synthesis. Existing methods struggle with large or complex motion and offer limited control over intermediate frames, often misaligning with user intent. We introduce MultiCOIN, a video inbetweening framework supporting multi‐modal controls, including depth transitions and layering, motion trajectories, text prompts, and target regions for movement localization. It balances flexibility, usability, and fine‐grained precision. Built on a Diffusion Transformer (DiT), due to its proven capability to generate high‐quality long video, our model maps all motion controls into a unified sparse point‐based representation compatible with the denoising process. Further, to respect the variety of controls which operate at varying levels of granularity and influence, we separate content and motion into two branches, enabling dedicated generators for each. A stage‐wise training strategy ensures stable learning of multi‐modal controls. Extensive experiments show improved motion complexity, controllability, and narrative consistency. Project Page: MultiCOIN.
Maham Tanveer, Yang Zhou 0007, Simon Niklaus, Ali Mahdavi-Amiri, Hao (Richard) Zhang, Krishna Kumar Singh, Nanxuan Zhao
Comput. Graph. Forum4
2025 SMITE: Segment Me In TimE
abstract
Segmenting an object in a video presents significant challenges. Each pixel must be accurately labeled, and these labels must remain consistent across frames. The difficulty increases when the segmentation is with arbitrary granularity, meaning the number of segments can vary arbitrarily, and masks are defined based on only one or a few sample images. In this paper, we address this issue by employing a pre-trained text to image diffusion model supplemented with an additional tracking mechanism. We demonstrate that our approach can effectively manage various segmentation scenarios and outperforms state-of-the-art alternatives. The project page is available at https://segment-me-in-time.github.io/
Amirhossein Alimohammadi, Sauradip Nag, Saeid Asgari Taghanaki, Andrea Tagliasacchi, Ghassan Hamarneh, Ali Mahdavi-Amiri
ICLR6
2025 SINGAPO: Single Image Controlled Generation of Articulated Parts in Objects
abstract
We address the challenge of creating 3D assets for household articulated objects from a single image. Prior work on articulated object creation either requires multi-view multi-state input, or only allows coarse control over the generation process. These limitations hinder the scalability and practicality for articulated object modeling. In this work, we propose a method to generate articulated objects from a single image. Observing the object in a resting state from an arbitrary view, our method generates an articulated object that is visually consistent with the input image. To capture the ambiguity in part shape and motion posed by a single view of the object, we design a diffusion model that learns the plausible variations of objects in terms of geometry and kinematics. To tackle the complexity of generating structured data with attributes in multiple domains, we design a pipeline that produces articulated objects from high-level structure to geometric details in a coarse-to-fine manner, where we use a part connectivity graph and part abstraction as proxies. Our experiments show that our method outperforms the state-of-the-art in articulated object creation by a large margin in terms of the generated object realism, resemblance to the input image, and reconstruction quality.
Jiayi Liu 0006, Denys Iliash, Angel X. Chang, Manolis Savva, Ali Mahdavi-Amiri
ICLR5
2025 GALA: Geometry-Aware Local Adaptive Grids for Detailed 3D Generation
abstract
We propose GALA, a novel representation of 3D shapes that (i) excels at capturing and reproducing complex geometry and surface details, (ii) is computationally efficient, and (iii) lends itself to 3D generative modelling with modern, diffusion-based schemes. The key idea of GALA is to exploit both the global sparsity of surfaces within a 3D volume and their local surface properties. *Sparsity* is promoted by covering only the 3D object boundaries, not empty space, with an ensemble of tree root voxels. Each voxel contains an octree to further limit storage and compute to regions that contain surfaces. *Adaptivity* is achieved by fitting one local and geometry-aware coordinate frame in each non-empty leaf node. Adjusting the orientation of the local grid, as well as the anisotropic scales of its axes, to the local surface shape greatly increases the amount of detail that can be stored in a given amount of memory, which in turn allows for quantization without loss of quality. With our optimized C++/CUDA implementation, GALA can be fitted to an object in less than 10 seconds. Moreover, the representation can efficiently be flattened and manipulated with transformer networks. We provide a cascaded generation pipeline capable of generating 3D shapes with great geometric detail. For more information, please visit our [project page](https://santisy.github.io/GALA/).
Dingdong Yang, Yizhi Wang 0006, Konrad Schindler, Ali Mahdavi-Amiri, Hao (Richard) Zhang
ICLR4
2025 In-2-4D: Inbetweening from Two Single-View Images to 4D Generation
abstract
We pose a new problem, In-2-4D, for generative 4D (i.e., 3D + motion) inbetweening to interpolate two single-view images. In contrast to video/4D generation from only text or a single image, our interpolative task can leverage more precise motion control to better constrain the generation. Given two monocular RGB images representing the start and end states of an object in motion, our goal is to generate and reconstruct the motion in 4D, without making assumptions on the object category, motion type, length, or complexity. To handle such arbitrary and diverse motions, we utilize a foundational video interpolation model for motion prediction. However, large frame-to-frame motion gaps can lead to ambiguous interpretations. To this end, we employ a hierarchical approach to identify keyframes that are visually close to the input states while exhibiting significant motions, then generate smooth fragments between them. For each fragment, we construct a 3D representation of the keyframe using Gaussian Splatting (3DGS). The temporal frames within the fragment guide the motion, enabling their transformation into dynamic 3DGS through a deformation field. To improve temporal consistency and refine the 3D motion, we expand the self-attention of multi-view diffusion across timesteps and apply rigid transformation regularization. Finally, we merge the independently generated 3D motion segments by interpolating boundary deformation fields and optimizing them to align with the guiding video, ensuring smooth and flicker-free transitions. Through extensive qualitative and quantitive experiments as well as a user study, we demonstrate the effectiveness of our method and design choices.
Sauradip Nag, Daniel Cohen-Or, Hao (Richard) Zhang, Ali Mahdavi-Amiri
SIGGRAPH Asia4
2025 ASIA: Adaptive 3D Segmentation using Few Image Annotations
abstract
We introduce ASIA (Adaptive 3D Segmentation using few Image Annotations), a novel framework that enables segmentation of possibly non-semantic and non-text describable "parts" in 3D. Our segmentation is controllable through a few user-annotated in-the-wild images, which are easier to collect than multi-view images, less demanding to annotate than 3D models, and more precise than potentially ambiguous text descriptions. Our method leverages the rich priors of text-to-image diffusion models, such as Stable Diffusion, to transfer segmentations from image space to 3D, even when the annotated and target objects differ significantly in geometry or structure. During training, we optimize a text token for each segment and fine-tune our model with a novel cross-view part correspondence loss. At inference, we segment multi-view renderings of the 3D mesh, fuse the labels in UV-space via voting, refine them with our novel Noise Optimization technique, and finally map the UV-labels back onto the mesh. ASIA provides a practical and generalizable solution for both semantic and non-semantic 3D segmentation tasks, outperforming existing methods by a noticeable margin in both quantitative and qualitative evaluations.
Perla Sai Raj Kishore, Aditya Vora, Sauradip Nag, Ali Mahdavi-Amiri, Hao (Richard) Zhang
SIGGRAPH Asia4
2025 Survey on Modeling of Human-made Articulated Objects
abstract
Abstract 3D modeling of articulated objects is a research problem within computer vision, graphics, and robotics. Its objective is to understand the shape and motion of the articulated components, represent the geometry and mobility of object parts, and create realistic models that reflect articulated objects in the real world. This survey provides a comprehensive overview of the current state‐of‐the‐art in 3D modeling of articulated objects, with a specific focus on the task of articulated part perception and articulated object creation (reconstruction and generation). We systematically review and discuss the relevant literature from two perspectives: geometry modeling (i.e., structure and shape of articulated parts) and articulation modeling (i.e., dynamics and motion of parts). Through this survey, we highlight the substantial progress made in these areas, outline the ongoing challenges, and identify gaps for future research. Our survey aims to serve as a foundational reference for researchers and practitioners in computer vision and graphics, offering insights into the complexities of articulated object modeling.
Jiayi Liu 0006, Manolis Savva, Ali Mahdavi-Amiri
Comput. Graph. Forum3
2024 CAGE: Controllable Articulation GEneration
abstract
We address the challenge of generating 3D articulated objects in a controllable fashion. Currently, modeling artic-ulated 3D objects is either achieved through laborious manual authoring, or using methods from prior work that are hard to scale and control directly. We leverage the interplay between part shape, connectivity, and motion using a de-noising diffusion-based method with attention modules de-signed to extract correlations between part attributes. Our method takes an object category label and a part connectivity graph as input and generates an object's geometry and motion parameters. The generated objects conform to user-specified constraints on the object category, part shape, and part articulation. Our experiments show that our method outperforms the state-of-the-art in articulated object generation, producing more realistic objects while conforming better to user constraints.
Jiayi Liu 0006, Hou In Ivan Tam, Ali Mahdavi-Amiri, Manolis Savva
CVPR3
2024 CLiC: Concept Learning in Context
abstract
This paper addresses the challenge of learning a local visual pattern of an object from one image, and generating images depicting objects with that pattern. Learning a localized concept and placing it on an object in a target image is a nontrivial task, as the objects may have different orientations and shapes. Our approach builds upon recent advancements in visual concept learning. It involves ac-quiring a visual concept (e.g., an ornament) from a source image and subsequently applying it to an object (e.g., a chair) in a target image. Our key idea is to perform in-context concept learning, acquiring the local visual concept within the broader context of the objects they belong to. To localize the concept learning, we employ soft masks that contain both the concept within the mask and the surrounding image area. We demonstrate our approach through object generation within an image, showcasing plausible embedding of in-context learned concepts. We also introduce methods for directing acquired concepts to specific locations within target images, employing cross-attention mechanisms, and establishing correspondences between source and target objects. The effectiveness of our method is demonstrated through quantitative and qualitative experiments, along with comparisons against baseline techniques.
Mehdi Safaee, Aryan Mikaeili, Or Patashnik, Daniel Cohen-Or, Ali Mahdavi-Amiri
CVPR5
2024 Slice3D: Multi-Slice, Occlusion-Revealing, Single View 3D Reconstruction
abstract
We introduce multi-slice reasoning, a new notion for single-view 3D reconstruction which challenges the current and prevailing belief that multi-view synthesis is the most natural conduit between single-view and 3D. Our key ob-servation is that object slicing is a more direct, and hence more advantageous, means to reveal occluded structures than altering camera views. Specifically, slicing can peel through any occluder without obstruction, and in the limit (i.e., with infinitely many slices), it is guaranteed to unveil all hidden object parts. We realize our idea by developing Slice3D, a novel method for single-view 3D reconstruction which first predicts multi-slice images from a single RGB input image and then integrates the slices into a 3D model using a coordinate-based transformer network to product a signed distance function. The slice images can be regressed or generated, both through a U-Net based network. For the former, we inject a learnable slice indicator code to desig-nate each decoded image into a spatial slice location, while the slice generator is a denoising diffusion model operating on the entirety of slice images stacked on the input channels. We conduct extensive evaluation against state-of-the-art alternatives to demonstrate superiority of our method, especially in recovering complex and severely occluded shape structures, amid ambiguities. All Slice3D results were produced by networks trained on a single Nvidia A40 GPU, with an inference time of less than 20 seconds.
Yizhi Wang 0006, Wallace P. Lira, Wenqi Wang 0003, Ali Mahdavi-Amiri, Hao (Richard) Zhang
CVPR4
2024 AFreeCA: Annotation-Free Counting for All
Adriano C. D'Alessandro, Ali Mahdavi-Amiri, Ghassan Hamarneh
ECCV (4)2
2024 SweepNet: Unsupervised Learning Shape Abstraction via Neural Sweepers
Mingrui Zhao, Yizhi Wang 0006, Fenggen Yu, Changqing Zou, Ali Mahdavi-Amiri
ECCV (37)5
2024 SLiMe: Segment Like Me
abstract
Significant strides have been made using large vision-language models, like Stable Diffusion (SD), for a variety of downstream tasks, including image generation, image editing, and 3D shape generation. Inspired by these advancements, we explore leveraging these vision-language models for segmenting images at any desired granularity using as few as one annotated sample. We propose SLiMe, which frames this problem as an optimization task. Specifically, given a single image and its segmentation mask, we first extract our novel “weighted accumulated self-attention map” along with cross-attention map from the SD prior. Then, using these extracted maps, the text embeddings of SD are optimized to highlight the segmented region in these attention maps, which in turn can be used to derive new segmentation results. Moreover, leveraging additional training data when available, i.e. few-shot, improves the performance of SLiMe. We performed comprehensive experiments examining various design factors and showed that SLiMe outperforms other existing one-shot and few-shot segmentation methods.
Aliasghar Khani, Saeid Asgari Taghanaki, Aditya Sanghi, Ali Mahdavi-Amiri, Ghassan Hamarneh
ICLR4
2024 EASI-Tex: Edge-Aware Mesh Texturing from Single Image
abstract
We present a novel approach for single-image mesh texturing , which employs a diffusion model with judicious conditioning to seamlessly transfer an object's texture from a single RGB image to a given 3D mesh object. We do not assume that the two objects belong to the same category, and even if they do, there can be significant discrepancies in their geometry and part proportions. Our method aims to rectify the discrepancies by conditioning a pre-trained Stable Diffusion generator with edges describing the mesh through ControlNet, and features extracted from the input image using IP-Adapter to generate textures that respect the underlying geometry of the mesh and the input texture without any optimization or training. We also introduce Image Inversion , a novel technique to quickly personalize the diffusion model for a single concept using a single image , for cases where the pre-trained IP-Adapter falls short in capturing all the details from the input image faithfully. Experimental results demonstrate the efficiency and effectiveness of our edge-aware single-image mesh texturing approach, coined EASI-Tex, in preserving the details of the input texture on diverse 3D objects, while respecting their geometry. Code : https://github.com/sairajk/easi-tex
Perla Sai Raj Kishore, Yizhi Wang 0006, Ali Mahdavi-Amiri, Hao (Richard) Zhang
ACM Trans. Graph.3
2024 SAC-GAN: Structure-Aware Image Composition
abstract
We introduce an end-to-end learning framework for image-to-image composition, aiming to plausibly compose an object represented as a cropped patch from an object image into a background scene image. As our approach emphasizes more on semantic and structural coherence of the composed images, rather than their pixel-level RGB accuracies, we tailor the input and output of our network with structure-aware features and design our network losses accordingly, with ground truth established in a self-supervised setting through the object cropping. Specifically, our network takes the semantic layout features from the input scene image, features encoded from the edges and silhouette in the input object patch, as well as a latent code as inputs, and generates a 2D spatial affine transform defining the translation and scaling of the object patch. The learned parameters are further fed into a differentiable spatial transformer network to transform the object patch into the target image, where our model is trained adversarially using an affine transform discriminator and a layout discriminator. We evaluate our network, coined SAC-GAN, for various image composition scenarios in terms of quality, composability, and generalizability of the composite images. Comparisons are made to state-of-the-art alternatives, including Instance Insertion, ST-GAN, CompGAN and PlaceNet, confirming superiority of our method.
Hang Zhou 0007, Rui Ma 0011, Ling-Xiao Zhang, Lin Gao 0004, Ali Mahdavi-Amiri, Hao (Richard) Zhang
IEEE Trans. Vis. Comput. Graph.5
2023 PARIS: Part-level Reconstruction and Motion Analysis for Articulated Objects
abstract
We address the task of simultaneous part-level reconstruction and motion parameter estimation for articulated objects. Given two sets of multi-view images of an object in two static articulation states, we decouple the movable part from the static part and reconstruct shape and appearance while predicting the motion parameters. To tackle this problem, we present PARIS: a self-supervised, end-to-end architecture that learns part-level implicit shape and appearance models and optimizes motion parameters jointly without any 3D supervision, motion, or semantic annotation. Our experiments show that our method generalizes better across object categories, and outperforms baselines and prior work that are given 3D point clouds as input. Our approach improves reconstruction relative to state-of-the-art baselines with a Chamfer-L1 distance reduction of 3.94 (45.2%) for objects and 26.79 (84.5%) for parts, and achieves 5% error rate for motion estimation across 10 object categories.
Jiayi Liu 0006, Ali Mahdavi-Amiri, Manolis Savva
ICCV2
2023 SKED: Sketch-guided Text-based 3D Editing
abstract
Text-to-image diffusion models are gradually introduced into computer graphics, recently enabling the development of Text-to-3D pipelines in an open domain. However, for interactive editing purposes, local manipulations of content through a simplistic textual interface can be arduous. Incorporating user guided sketches with Text-to-image pipelines offers users more intuitive control. Still, as state-of-the-art Text-to-3D pipelines rely on optimizing Neural Radiance Fields (NeRF) through gradients from arbitrary rendering views, conditioning on sketches is not straightforward. In this paper, we present SKED, a technique for editing 3D shapes represented by NeRFs. Our technique utilizes as few as two guiding sketches from different views to alter an existing neural field. The edited region respects the prompt semantics through a pre-trained diffusion model. To ensure the generated output adheres to the provided sketches, we propose novel loss functions to generate the desired edits while preserving the density and radiance of the base instance. We demonstrate the effectiveness of our proposed method through several qualitative and quantitative experiments. https://sked-paper.github.io/
Aryan Mikaeili, Or Perel, Mehdi Safaee, Daniel Cohen-Or, Ali Mahdavi-Amiri
ICCV5
2023 DS-Fusion: Artistic Typography via Discriminated and Stylized Diffusion
abstract
We introduce a novel method to automatically generate an artistic typography by stylizing one or more letter fonts to visually convey the semantics of an input word, while ensuring that the output remains readable. To address an assortment of challenges with our task at hand including conflicting goals (artistic stylization vs. legibility), lack of ground truth, and immense search space, our approach utilizes large language models to bridge texts and visual images for stylization and build an unsupervised generative model with a diffusion model backbone. Specifically, we employ the denoising generator in Latent Diffusion Model (LDM), with the key addition of a CNN-based discriminator to adapt the input style onto the input text. The discriminator uses rasterized images of a given letter/word font as real samples and the output of the denoising generator as fake samples. Our model is coined DS-Fusion for discriminated and stylized diffusion. We showcase the quality and versatility of our method through numerous examples, qualitative and quantitative evaluation, and ablation studies. User studies comparing to strong baselines including CLIPDraw, DALL-E 2, Stable Diffusion, as well as artist-crafted typographies, demonstrate strong performance of DS-Fusion. Code is available at https://ds-fusion.github.io/.
Maham Tanveer, Yizhi Wang 0006, Ali Mahdavi-Amiri, Hao (Richard) Zhang
ICCV3
2023 D2CSG: Unsupervised Learning of Compact CSG Trees with Dual Complements and Dropouts
abstract
We present D$^2$CSG, a neural model composed of two dual and complementary network branches, with dropouts, for unsupervised learning of compact constructive solid geometry (CSG) representations of 3D CAD shapes. Our network is trained to reconstruct a 3D shape by a fixed-order assembly of quadric primitives, with both branches producing a union of primitive intersections or inverses. A key difference between D$^2$CSG and all prior neural CSG models is its dedicated residual branch to assemble the potentially complex shape complement, which is subtracted from an overall shape modeled by the cover branch. With the shape complements, our network is provably general, while the weight dropout further improves compactness of the CSG tree by removing redundant primitives. We demonstrate both quantitatively and qualitatively that D$^2$CSG produces compact CSG reconstructions with superior quality and more natural primitives than all existing alternatives, especially over complex and high-genus CAD shapes.
Fenggen Yu, Qimin Chen, Maham Tanveer, Ali Mahdavi-Amiri, Hao (Richard) Zhang
NeurIPS4
2022 UNIST: Unpaired Neural Implicit Shape Translation Network
abstract
We introduce UNIST, the first deep neural implicit model for general-purpose, unpaired shape-to-shape translation, in both 2D and 3D domains. Our model is built on autoencoding implicit fields, rather than point clouds which represents the state of the art. Furthermore, our translation network is trained to perform the task over a latent grid representation which combines the merits of both latent-space processing and position awareness, to not only enable drastic shape transforms but also well preserve spatial features and fine local details for natural shape translations. With the same network architecture and only dictated by the input domain pairs, our model can learn both style-preserving content alteration and content-preserving style transfer. We demonstrate the generality and quality of the translation results, and compare them to well-known baselines. Code is available at https://qiminchen.github.io/unist/.
Qimin Chen, Johannes Merz, Aditya Sanghi, Hooman Shayani, Ali Mahdavi-Amiri, Hao (Richard) Zhang
CVPR5
2022 CAPRI-Net: Learning Compact CAD Shapes with Adaptive Primitive Assembly
abstract
We introduce CAPRI-Net, a self-supervised neural network for learning compact and interpretable implicit representations of 3D computer-aided design (CAD) models, in the form of adaptive primitive assemblies. Given an input 3D shape, our network reconstructs it by an assembly of quadric surface primitives via constructive solid geometry (CSG) operations. Without any ground-truth shape assemblies, our self-supervised network is trained with a reconstruction loss, leading to faithful 3D reconstructions with sharp edges and plausible CSG trees. While the parametric nature of CAD models does make them more predictable locally, at the shape level, there is much structural and topological variation, which presents a significant generalizability challenge to state-of-the-art neural models for 3D shapes. Our network addresses this challenge by adaptive training with respect to each test shape, with which we fine-tune the network that was pre-trained on a model collection. We evaluate our learning framework on both ShapeNet and ABC, the largest and most diverse CAD dataset to date, in terms of reconstruction quality, sharp edges, compactness, and interpretability, to demonstrate superiority over current alternatives for neural CAD reconstruction.
Fenggen Yu, Manyi Li, Aditya Sanghi, Hooman Shayani, Ali Mahdavi-Amiri, Hao (Richard) Zhang
CVPR6
2022 MaskTune: Mitigating Spurious Correlations by Forcing to Explore
abstract
A fundamental challenge of over-parameterized deep learning models is learning meaningful data representations that yield good performance on a downstream task without over-fitting spurious input features. This work proposes MaskTune, a masking strategy that prevents over-reliance on spurious (or a limited number of) features. MaskTune forces the trained model to explore new features during a single epoch finetuning by masking previously discovered features. MaskTune, unlike earlier approaches for mitigating shortcut learning, does not require any supervision, such as annotating spurious features or labels for subgroup samples in a dataset. Our empirical results on biased MNIST, CelebA, Waterbirds, and ImagenNet-9L datasets show that MaskTune is effective on tasks that often suffer from the existence of spurious correlations. Finally, we show that \method{} outperforms or achieves similar performance to the competing methods when applied to the selective classification (classification with rejection option) task. Code for MaskTune is available at https://github.com/aliasgharkhani/Masktune.
Saeid Asgari Taghanaki, Aliasghar Khani, Fereshte Khani, Ali Mahdavi-Amiri, Ghassan Hamarneh
NeurIPS6
2021 Data to Physicalization: A Survey of the Physical Rendering Process
abstract
Abstract Physical representations of data offer physical and spatial ways of looking at, navigating, and interacting with data. While digital fabrication has facilitated the creation of objects with data‐driven geometry, rendering data as a physically fabricated object is still a daunting leap for many physicalization designers. Rendering in the scope of this research refers to the back‐and‐forth process from digital design to digital fabrication and its specific challenges. We developed a corpus of example data physicalizations from research literature and physicalization practice. This survey then unpacks the “rendering” phase of the extended InfoVis pipeline in greater detail through these examples, with the aim of identifying ways that researchers, artists, and industry practitioners “render” physicalizations using digital design and fabrication tools.
Hessam Djavaherpour, Faramarz F. Samavati, Ali Mahdavi-Amiri, Fatemeh Yazdanbakhsh, Samuel Huron, Richard Levy, Yvonne Jansen, Lora Oehlberg
Comput. Graph. Forum3
2021 RIAS: Repeated Invertible Averaging for Surface Multiresolution of Arbitrary Degree
abstract
In this article, we introduce two local surface averaging operators with local inverses and use them to devise a method for surface multiresolution (subdivision and reverse subdivision) of arbitrary degree. Similar to previous works by Stam, Zorin, and Schröder that achieved forward subdivision only, our averaging operators involve only direct neighbours of a vertex, and can be configured to generalize B-Spline multiresolution to arbitrary topology surfaces. Our subdivision surfaces are hence able to exhibit Cdcontinuity at regular vertices (for arbitrary values ofd) and appear to exhibit C1continuity at extraordinary vertices. Smooth reverse and non-uniform subdivisions are additionally supported.
Troy F. Alderson, Ali Mahdavi-Amiri, Faramarz F. Samavati
IEEE Trans. Vis. Comput. Graph.2
2020 PIE-NET: Parametric Inference of Point Cloud Edges
abstract
We introduce an end-to-end learnable technique to robustly identify feature edges in 3D point cloud data. We represent these edges as a collection of parametric curves (i.e.,~lines, circles, and B-splines). Accordingly, our deep neural network, coined PIE-NET, is trained for parametric inference of edges. The network relies on a "region proposal" architecture, where a first module proposes an over-complete collection of edge and corner points, and a second module ranks each proposal to decide whether it should be considered. We train and evaluate our method on the ABC dataset, a large dataset of CAD models, and compare our results to those produced by traditional (non-learning) processing pipelines, as well as a recent deep learning based edge detector (EC-NET). Our results significantly improve over the state-of-the-art from both a quantitative and qualitative standpoint.
Xiaogang Wang 0005, Yuelang Xu, Kai Xu 0004, Andrea Tagliasacchi, Ali Mahdavi-Amiri, Hao (Richard) Zhang
NeurIPS6
2020 VDAC: volume decompose-and-carve for subtractive manufacturing
abstract
We introduce carvable volume decomposition for efficient 3-axis CNC machining of 3D freeform objects, where our goal is to develop a fully automatic method to jointly optimize setup and path planning. We formulate our joint optimization as a volume decomposition problem which prioritizes minimizing the number of setup directions while striving for a minimum number of continuously carvable volumes, where a 3D volume is continuously carvable, or simply carvable, if it can be carved with the machine cutter traversing a single continuous path. Geometrically, carvability combines visibility and monotonicity and presents a new shape property which had not been studied before. Given a target 3D shape and the initial material block, our algorithm first finds the minimum number of carving directions by solving a set cover problem. Specifically, we analyze cutter accessibility and select the carving directions based on an assessment of how likely they would lead to a small carvable volume decomposition. Next, to obtain a minimum decomposition based on the selected carving directions efficiently, we narrow down the solution search by focusing on a special kind of points in the residual volume, single access or SA points, which are points that can be accessed from one and only one of the selected carving directions. Candidate carvable volumes are grown starting from the SA points. Finally, we devise an energy term to evaluate the carvable volumes and their combinations, leading to the final decomposition. We demonstrate the performance of our decomposition algorithm on a variety of 2D and 3D examples and evaluate it against the ground truth, where possible, and solutions provided by human experts. Physically machined models are produced where each carvable volume is continuously carved following a connected Fermat spiral toolpath.
Ali Mahdavi-Amiri, Fenggen Yu, Haisen Zhao, Adriana Schulz, Hao (Richard) Zhang
ACM Trans. Graph.1
2018 Landscaper: A Modeling System for 3D Printing Scale Models of Landscapes
abstract
Abstract Landscape models of geospatial regions provide an intuitive mechanism for exploring complex geospatial information. However, the methods currently used to create these scale models require a large amount of resources, which restricts the availability of these models to a limited number of popular public places, such as museums and airports. In this paper, we have proposed a system for creating these physical models using an affordable 3D printer in order to make the creation of these models more widely accessible. Our system retrieves GIS relevant to creating a physical model of a geospatial region and then addresses the two major limitations of affordable 3D printers, namely the limited number of materials and available printing volume. This is accomplished by separating features into distinct extruded layers and splitting large models into smaller pieces, allowing us to employ different methods for the visualization of different geospatial features, like vegetation and residential areas, in a 3D printing context. We confirm the functionality of our system by printing two large physical models of relatively complex landscape regions.
K. Allahverdi, Hessam Djavaherpour, Ali Mahdavi-Amiri, Faramarz F. Samavati
Comput. Graph. Forum3
2018 Construction and fabrication of reversible shape transforms
abstract
We study a new and elegant instance of geometric dissection of 2D shapes: reversible hinged dissection, which corresponds to a dual transform between two shapes where one of them can be dissected in its interior and then inverted inside-out , with hinges on the shape boundary, to reproduce the other shape, and vice versa. We call such a transform reversible inside-out transform or RIOT. Since it is rare for two shapes to possess even a rough RIOT, let alone an exact one, we develop both a RIOT construction algorithm and a quick filtering mechanism to pick, from a shape collection, potential shape pairs that are likely to possess the transform. Our construction algorithm is fully automatic. It computes an approximate RIOT between two given input 2D shapes, whose boundaries can undergo slight deformations, while the filtering scheme picks good inputs for the construction. Furthermore, we add properly designed hinges and connectors to the shape pieces and fabricate them using a 3D printer so that they can be played as an assembly puzzle. With many interesting and fun RIOT pairs constructed from shapes found online, we demonstrate that our method significantly expands the range of shapes to be considered for RIOT, a seemingly impossible shape transform, and offers a practical way to construct and physically realize these transforms.
Ali Mahdavi-Amiri, Ruizhen Hu, Han Liu 0003, Changqing Zou, Oliver van Kaick, Xiuping Liu, Hui Huang 0004, Hao (Richard) Zhang
ACM Trans. Graph.2
2018 Semi-Supervised Co-Analysis of 3D Shape Styles from Projected Lines
abstract
We present a semi-supervised co-analysis method for learning 3D shape styles from projected feature lines , achieving style patch localization with only weak supervision. Given a collection of 3D shapes spanning multiple object categories and styles, we perform style co-analysis over projected feature lines of each 3D shape and then back-project the learned style features onto the 3D shapes. Our core analysis pipeline starts with mid-level patch sampling and pre-selection of candidate style patches. Projective features are then encoded via patch convolution. Multi-view feature integration and style clustering are carried out under the framework of partially shared latent factor (PSLF) learning, a multi-view feature learning scheme. PSLF achieves effective multi-view feature fusion by distilling and exploiting consistent and complementary feature information from multiple views, while also selecting style patches from the candidates. Our style analysis approach supports both unsupervised and semi-supervised analysis. For the latter, our method accepts both user-specified shape labels and style-ranked triplets as clustering constraints. We demonstrate results from 3D shape style analysis and patch localization as well as improvements over state-of-the-art methods. We also present several applications enabled by our style analysis.
Fenggen Yu, Yan Zhang 0057, Kai Xu 0004, Ali Mahdavi-Amiri, Hao (Richard) Zhang
ACM Trans. Graph.4
2018 Offsetting spherical curves in vector and raster form
Troy F. Alderson, Ali Mahdavi-Amiri, Faramarz F. Samavati
Vis. Comput.2
2017 Diagrammatic approach for constructing multiresolution of primal subdivisions
Richard H. Bartels, Ali Mahdavi-Amiri, Faramarz F. Samavati, Nezam Mahdavi-Amiri
Comput. Aided Geom. Des.2
2016 Hierarchical grid conversion
abstract
Hierarchical grids appear in various applications in computer graphics such as subdivision and multiresolution surfaces, and terrain models. Since the different grid types perform better at different tasks, it is desired to switch between regular grids to take advantages of these grids. Based on a 2D domain obtained from the connectivity information of a mesh, we can define simple conversions to switch between regular grids. In this paper, we introduce a general framework that can be used to convert a given grid to another and we discuss the properties of these refinements such as their transformations. This framework is hierarchical meaning that it provides conversions between meshes at different level of refinement. To describe the use of this framework, we define new regular and near-regular refinements with good properties such as small factors. We also describe how grid conversion enables us to use patch-based data structures for hexagonal cells and near-regular refinements. To do so, meshes are converted to a set of quadrilateral patches that can be stored in simple structures. Near-regular refinements are also supported by defining two sets of neighborhood vectors that connect a vertex to its neighbors and are useful to address connectivity queries.
Ali Mahdavi-Amiri, Erika Harrison, Faramarz F. Samavati
Comput. Aided Des.1
2016 Multiresolution on spherical curves
abstract
In this paper, we present an approximating multiresolution framework of arbitrary degree for curves on the surface of a sphere. Multiresolution by subdivision and reverse subdivision allows one to decrease and restore the resolution of a curve, and is typically defined by affine combinations of points in Euclidean space . While translating such combinations to spherical space is possible, ensuring perfect reconstruction of the curve remains challenging. Hence, current spherical multiresolution schemes tend to be interpolating or midpoint-interpolating, as achieving perfect reconstruction in these cases is more straightforward. We use a simple geometric construction for a non-interpolating and non-midpoint-interpolating multiresolution scheme on the sphere, which is made up of easily generalized components and based on a modified Lane–Riesenfeld algorithm.
Troy F. Alderson, Ali Mahdavi-Amiri, Faramarz F. Samavati
Graph. Model.2
2015 Cover-it: an interactive system for covering 3d prints
Ali Mahdavi-Amiri, Philip Whittingham, Faramarz F. Samavati
Graphics Interface1
2015 A Survey of Digital Earth
Ali Mahdavi-Amiri, Troy F. Alderson, Faramarz F. Samavati
Comput. Graph.1
2014 Atlas of connectivity maps
Ali Mahdavi-Amiri, Faramarz F. Samavati
Comput. Graph.1
2013 ACM: atlas of connectivity maps for semiregular models
Ali Mahdavi-Amiri, Faramarz F. Samavati
Graphics Interface1
2011 Optimization of Inverse Snyder Polyhedral Projection
abstract
Modern techniques in area preserving projections used by cartographers and other glossarial researchers have closed forms when projecting from the sphere to the plane, as based on their initial derivations. Inversions, from the planar map to the spherical approximation of the Earth which are important for modern 3D analysis and visualizations, are slower, requiring iterative root finding approaches, or not determined at all. We introduce optimization techniques for Snyder's inverse polyhedral projection by reducing iterations, and using polynomial approximations for avoiding them entirely. Results including speed up, iteration reduction, and error analysis are provided.
Erika Harrison, Ali Mahdavi-Amiri, Faramarz F. Samavati
CW2