EDBT 2026 Demo / reviewers in the wild / expert
Ariel Shamir
dblp:66/2852
· DBLP profile ↗
139ranked-venue papers
12as first author
52since 2021 · last 2026
0000-0001-7082-7845ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 124 · 10 first-author · 44 since 2021Artificial intelligence and machine learning · 17 · 13 since 2021Human-computer interaction and ubiquitous computing · 13 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Low-Vision Chart Accessibility via On-Cursor Visual ContextabstractDespite widespread use, charts remain largely inaccessible for Low-Vision Individuals (LVI). Reading charts requires viewing data points within a global context, which is difficult for LVI who may rely on magnification or experience a partial field of vision. We aim to improve exploration by providing visual access to critical context. To inform this, we conducted a formative study with five LVI. We identified four fundamental contextual elements common across chart types: axes, legend, grid lines, and the overview. We propose two pointer-based interaction methods to provide this context: Dynamic Context, a novel focus+context interaction, and Mini-map, which adapts overview+detail principles for LVI. In a study with N=22 LVI, we compared both methods and evaluated their integration to current tools. Our results show that Dynamic Context had significant positive impact on access, usability, and effort reduction; however, worsened visual load. Mini-map strengthened spatial understanding, but was less preferred for this task. We offer design insights to guide the development of future systems that support LVI with visual context while balancing visual load. Yotam Sechayk, Hennes Rave, Max Rädler, Mark Colley, Zhongyi Zhou, Ariel Shamir, Takeo Igarashi |
CHI | 6 |
| 2026 | CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question AnsweringabstractVision-Language Models (VLMs) demonstrate remarkable capabilities in visual understanding and reasoning, such as in Visual Question Answering (VQA), where the model is asked a question related to a visual input. Still, these models can make distinctly unnatural errors, for example, providing (wrong) answers to unanswerable VQA questions, such as questions asking about objects that do not appear in the image.To address this issue, we propose CLIP-UP: CLIP-based Unanswerable Problem detection, a novel lightweight method for equipping VLMs with the ability to withhold answers to unanswerable questions. CLIP-UP leverages CLIP-based similarity measures to extract question-image alignment information to detect unanswerability, requiring efficient training of only a few additional layers, while keeping the original VLMs’ weights unchanged.Tested across several models, CLIP-UP achieves significant improvements on benchmarks assessing unanswerability in both multiple-choice and open-ended VQA, surpassing other methods, while preserving original performance on other tasks. Ben Vardi, Oron Nir, Ariel Shamir |
WACV | 3 |
| 2026 | Palette Aligned Image DiffusionabstractAbstract We introduce the Palette‐Adapter , a novel method for conditioning text‐to‐image diffusion models on a user‐specified color palette. While palettes are a compact and intuitive tool widely used in creative workflows, they introduce significant ambiguity and instability when used for conditioning image generation. Our approach addresses this challenge by interpreting palettes as sparse histograms and introducing two scalar control parameters: histogram entropy and palette‐to‐histogram distance , which allow flexible control over the degree of palette adherence and color variation. We further introduce a negative histogram mechanism that allows users to suppress specific undesired hues, improving adherence to the intended palette under the standard classifier‐free guidance mechanism. To ensure broad generalization across the color space, we train on a carefully curated dataset with balanced coverage of rare and common colors. Our method enables stable, semantically coherent generation across a wide range of palettes and prompts. We evaluate our method qualitatively, quantitatively, and through a human evaluation, and show that it consistently outperforms existing approaches in achieving both strong palette adherence and high image quality. Elad Aharoni, Noy Porat, Dani Lischinski, Ariel Shamir |
Comput. Graph. Forum | 4 |
| 2026 | Conversational Gesture Model (CGM): Extending Speaker-Centric Audio-Driven Motion Generation to Full Conversation Gestures
Tomer Koren, Adi Rosenthal, Doron Friedman, Ariel Shamir |
Comput. Graph. Forum | 4 |
| 2026 | Conversational Gesture Model (CGM): Extending Speaker-Centric Audio-Driven Motion Generation to Full Conversation GesturesabstractAbstract In this work we extend speaker‐centric audio‐driven gesture synthesis toward a unified conversational model that jointly captures both speaking and listening behaviors. Existing speaker‐centric models effectively generate gestures aligned with speech but overlook the bidirectional dynamics that characterize natural dialogue. To address this limitation, we propose the Conversational Gesture Model (CGM), a cross‐attention‐based model capable of synthesizing gestures conditioned on interlocutor conversational cues such as gestures, tone, and textual semantics. By leveraging cross‐attention mechanisms, the model fuses interlocutor audio and text features with character gesture encodings, enabling a single system to seamlessly alternate between speaking and listening roles of the same character. Hence, our model enables a single system to act as both speaker and listener, capturing the fluid role shifts and mutual influence inherent in conversation. Experiments demonstrate that this approach preserves the quality of speaker‐driven gestures while significantly improving the realism, coherence, and responsiveness of full conversational interactions. Tomer Koren, Adi Rosenthal, Doron Friedman, Ariel Shamir |
Comput. Graph. Forum | 4 |
| 2026 | SiGnature: Explicit Motion Diffusion for Stylized Semantic Gesture GenerationabstractAbstract While recent advances in co‐speech gesture generation have achieved impressive rhythmic synchronization, synthesizing gestures that are both semantically meaningful and faithful to a speaker's unique non‐verbal style remains an open challenge. Semantic gestures, such as iconic shapes or deictic pointing, are statistically sparse, making them difficult to learn effectively within standard generative models. We present SiGnature, a framework for Stylized and Semantic Gesture generation that reconciles precise semantic control with high‐fidelity style preservation. Unlike prevalent methods that rely on entangled latent representations, SiGnature operates in an explicit joint‐rotation space. This design enables our core contribution, Joint Motion Integration (JMI), a training‐free inference mechanism capable of injecting any external motion sequence, particularly in‐the‐wild semantic gestures, directly into the diffusion process. JMI automatically identifies the specific “active joints” conveying a semantic action and injects them into the generation, while relying on the diffusion backbone to synthesize the remaining body dynamics, including posture and flow, in accordance with the pre‐learned style of the target speaker. This allows for the plug‐and‐play integration of arbitrary motions, including complex semantic gestures, without retraining or introducing the “Frankenstein” artifacts typical of cut‐and‐paste methods. Extensive experiments and perceptual studies demonstrate that SiGnature offers superior semantic motion control while maintaining smooth and natural co‐speech gesture generation and preserving the distinct characteristics of the speaker, thereby outperforming state‐of‐the‐art baselines. Adi Rosenthal, Tomer Koren, Nadav Shaked, Doron Friedman, Ariel Shamir |
Comput. Graph. Forum | 5 |
| 2026 | Gait-Synced Translation Gain for Naturalistic VR MotionabstractTranslation gain is a key Redirected Walking (RDW) technique in Virtual Reality (VR) that enables users to navigate virtual environments (VEs) larger than the available physical space. The technique was originally developed to scale users' walking distance in the VE and is typically applied continuously, regardless of the user's motion state. We introduce Gait-Synced Translation Gain (GSTG), a novel approach that adapts translation gain by synchronizing it with the user's gait cycle. GSTG leverages the single-limb support phase of walking-when users are less stable and thus less sensitive to external disturbances-to apply higher levels of gain. This approach allows greater manipulation while preserving natural walking sensations and avoiding additional cybersickness. A user study comparing GSTG with continuous translation gain demonstrates significant improvements in perceived naturalness and comfort. Our results highlight the potential of gait-synchronized gain to enhance immersion, offering new possibilities for more realistic and comfortable VR locomotion. Fiona Xiao Yu Chen, Sen-Zhe Xu 0001, Kui Huang, Ariel Shamir, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | VeasyGuide: Personalized Visual Guidance for Low-vision Learners on Instructor Actions in Presentation VideosabstractToggle (a.1) Zoom settings panel (a.2) Highlight settings panel (a) The VeasyGuide player (b) Zoomed-in view (c) VeasyGuide's personalization settings panelsFigure 1: VeasyGuide's interface includes: (a) a video player with highlighted areas (red rectangle and hand pointer), (a.1) zoom settings toggle, (a.2) highlight settings toggle, (b) a zoomed window (toggled with the Z key), and (c) zoom and highlight settings panels.While watching a video, VeasyGuide auto-highlights pointing, marking, and sketching activities.Users can zoom into highlighted portions with the Z key, adjust zoom with arrow keys, and customize highlight and zoom appearance in real time. Yotam Sechayk, Ariel Shamir, Amy Pavel, Takeo Igarashi |
ASSETS | 2 |
| 2025 | FontCraft: Multimodal Font Design Using Interactive Bayesian OptimizationabstractInternational audience Yuki Tatsukawa, I-Chao Shen, Mustafa Doga Dogan, Anran Qi, Yuki Koyama 0001, Ariel Shamir, Takeo Igarashi |
CHI | 6 |
| 2025 | Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image GenerationabstractBalancing content fidelity and artistic style is a pivotal challenge in image generation. While traditional style transfer methods and modern Denoising Diffusion Probabilistic Models (DDPMs) strive to achieve this balance, they often struggle to do so without sacrificing either style, content, or sometimes both. This work addresses this challenge by analyzing the ability of DDPMs to maintain content and style equilibrium. We introduce a novel method to identify sensitivities within the DDPM attention layers, identifying specific layers that correspond to different stylistic aspects. By directing conditional inputs only to these sensitive layers, our approach enables fine-grained control over style and content, significantly reducing issues arising from over-constrained inputs. Our findings demonstrate that this method enhances recent stylization techniques by better aligning style and content, ultimately improving the quality of generated visual content. Nadav Z. Cohen, Oron Nir, Ariel Shamir |
CVPR | 3 |
| 2025 | Generative PhotomontageabstractText-to-image models are powerful tools for image creation. However, the generation process is akin to a dice roll and makes it difficult to achieve a single image that captures everything a user wants. In this paper, we propose a framework for creating the desired image by compositing it from various parts of generated images, in essence forming a Generative Photomontage. Given a stack of images generated by ControlNet using the same input condition and different seeds, we let users select desired parts from the generated results using a brush stroke interface. We introduce a novel technique that takes in the user’s brush strokes, segments the generated images using a graph-based optimization in diffusion feature space, and then composites the segmented regions via a new feature-space blending method. Our method faithfully preserves the user-selected regions while compositing them harmoniously. We demonstrate that our flexible framework can be used for many applications, including generating new appearance combinations, fixing incorrect shapes and artifacts, and improving prompt alignment. We show compelling results for each application and demonstrate that our method outperforms existing image blending methods and various baselines. Sean J. Liu, Nupur Kumari, Ariel Shamir, Jun-Yan Zhu |
CVPR | 3 |
| 2025 | Playing Along: Building AI Agents for Co-Creation of Improvised Stories
Idan Dov Vidra, Gal Kimron, Lior Noy, Ariel Shamir |
ICCC | 4 |
| 2025 | NeuFrameQ: Neural Frame Fields for Scalable and Generalizable Anisotropic Quadrangulation
Ying-Tian Liu, Xin Yu 0004, Yan-Pei Cao 0001, Ding Liang, Ariel Shamir, Song-Hai Zhang |
ICCV | 8 |
| 2025 | Unimodal Strategies in Density-Based Clustering
Oron Nir, Jay Tenenbaum, Ariel Shamir |
ECML/PKDD (1) | 3 |
| 2025 | AutoSketch: VLM-assisted Style-Aware Vector Sketch CompletionabstractSketches are an important medium of expression and recently many works concentrate on automatic sketch creations. One such ability very useful for amateurs is text-based completion of a partial sketch to create a complex scene, while preserving the style of the partial sketch. Existing methods focus solely on generating sketch that match the content in the input prompt in a predefined style, ignoring the styles of the input partial sketches, e.g., the global abstraction level and local stroke styles. To address this challenge, we introduce AutoSketch, a style-aware vector sketch completion method that accommodates diverse sketch styles and supports iterative sketch completion. AutoSketch completes the input sketch in a style-consistent manner using a two-stage method. In the first stage, we initially optimize the strokes to match an input prompt augmented by style descriptions extracted from a vision-language model (VLM). Such style descriptions lead to non-photorealistic guidance images which enable more content to be depicted through new strokes. In the second stage, we utilize the VLM to adjust the strokes from the previous stage to adhere to the style present in the input partial sketch through an iterative style adjustment process. In each iteration, the VLM identifies a list of style differences between the input sketch and the strokes generated in the previous stage, translating these differences into adjustment codes to modify the strokes. We compare our method with existing methods using various sketch styles and prompts, perform extensive ablation studies and qualitative and quantitative evaluations, and demonstrate that AutoSketch can support diverse sketching scenarios. Hsiao-Yuan Chin, I-Chao Shen, Yi-Ting Chiu, Ariel Shamir, Bing-Yu Chen 0004 |
SIGGRAPH Asia | 4 |
| 2025 | Duet dancing from a solo dance videoabstractThis paper introduces a method to generate duet dance videos from an input solo dancer’s video performance. Addressing this novel problem, our system tackles a set of sub-tasks, including dancer segmentation, camera motion handling, stage reconstruction and the intricate management of geometric constraints such as dancer scale preservation and dancers collision prevention. The proposed approach leverages existing methodologies and new solutions. Notably, we address collisions in 2D space–time directly, departing from traditional 3D approaches. We modify the initial location of the dancer to avoid long-time collisions globally, while also modulating the pace of the dance by deliberate slowing down or accelerating motion to avoid short collisions locally. Experimental results attest to the efficacy of our approach. The system not only successfully synthesizes engaging duet dance sequences but also upholds the authenticity of individual performances, as shown by a user study. Yarin Moshe, Yael Moses, Ariel Shamir |
Comput. Graph. | 3 |
| 2025 | Sketch2Data: Recovering data from hand-drawn infographics
Anran Qi, Theophanis Tsandilas, Ariel Shamir, Adrien Bousseau |
Comput. Graph. | 3 |
| 2025 | REED-VAE: RE-Encode Decode Training for Iterative Image Editing with Diffusion ModelsabstractAbstract While latent diffusion models achieve impressive image editing results, their application to iterative editing of the same image is severely restricted. When trying to apply consecutive edit operations using current models, they accumulate artifacts and noise due to repeated transitions between pixel and latent spaces. Some methods have attempted to address this limitation by performing the entire edit chain within the latent space, sacrificing flexibility by supporting only a limited, predetermined set of diffusion editing operations. We present a re‐encode decode (REED) training scheme for variational autoencoders (VAEs), which promotes image quality preservation even after many iterations. Our work enables multi‐method iterative image editing: users can perform a variety of iterative edit operations, with each operation building on the output of the previous one using both diffusion based operations and conventional editing techniques. We demonstrate the advantage of REED‐VAE across a range of image editing scenarios, including text‐based and mask‐based editing frameworks. In addition, we show how REED‐VAE enhances the overall editability of images, increasing the likelihood of successful and precise edit operations. We hope that this work will serve as a benchmark for the newly introduced task of multi‐method image editing. Gal Almog, Ariel Shamir, Ohad Fried |
Comput. Graph. Forum | 2 |
| 2025 | LayoutRectifier: An Optimization-based Post-processing for Graphic Design Layout GenerationabstractAbstract Recent deep learning methods can generate diverse graphic design layouts efficiently. However, these methods often create layouts with flaws, such as misalignment, unwanted overlaps, and unsatisfied containment. To tackle this issue, we propose an optimization‐based method called LayoutRectifier, which gracefully rectifies auto‐generated graphic design layouts to reduce these flaws while minimizing deviation from the generated layout. The core of our method is a two‐stage optimization. First, we utilize grid systems, which professional designers commonly use to organize elements, to mitigate misalignments through discrete search. Second, we introduce a novel box containment function designed to adjust the positions and sizes of the layout elements, preventing unwanted overlapping and promoting desired containment. We evaluate our method on content‐agnostic and content‐aware layout generation tasks and achieve better‐quality layouts that are more suitable for downstream graphic design tasks. Our method complements learning‐based layout generation methods and does not require additional training. I-Chao Shen, Ariel Shamir, Takeo Igarashi |
Comput. Graph. Forum | 2 |
| 2025 | Prediction of Scene PlausibilityabstractUnderstanding the 3D world from 2D images involves more than detection and segmentation of the objects within the scene. It also includes the interpretation of the structure and arrangement of the scene elements. Such understanding is often rooted in recognizing the physical world and its limitations, and in prior knowledge as to how similar typical scenes are arranged. In this research, we pose a new challenge for neural networks, or other scene understanding algorithms—can they distinguish plausible scenes from implausible ones? Plausibility is defined both in terms of physical properties and in terms of functional and typical arrangements. Loosely, it can be seen as the probability of encountering a given scene in the real physical world. We have built a dataset of synthetic images containing both plausible and implausible scenes, and test the success of various vision models for the task of classifying and recognizing different types of implausibility. Our code and dataset is publicly available at https://github.com/ornachmias/scene_plusibility_prediction. Or Nachmias, Ohad Fried, Ariel Shamir |
Comput. Vis. Media | 3 |
| 2025 | Neural Image abstraction using long smoothing B-splinesabstractWe integrate smoothing B-splines into a standard differentiable vector graphics (DiffVG) pipeline through linear mapping, and show how this can be used to generate smooth and arbitrarily long paths within image-based deep learning systems. We take advantage of derivative-based smoothing costs for parametric control of fidelity vs. simplicity tradeoffs, while also enabling stylization control in geometric and image spaces. The proposed pipeline is compatible with recent vector graphics generation and vectorization methods. We demonstrate the versatility of our approach with four applications aimed at the generation of stylized vector graphics: stylized space-filling path generation, stroke-based image abstraction, closed-area image abstraction, and stylized text generation. Daniel Berio, Michael Stroh, Sylvain Calinon, Frederic Fol Leymarie, Oliver Deussen, Ariel Shamir |
ACM Trans. Graph. | 6 |
| 2025 | MiGumi: Making Tightly Coupled Integral Joints MillableabstractTraditional integral wood joints, despite their strength, durability, and elegance, remain rare in modern workflows due to the cost and difficulty of manual fabrication. CNC milling offers a scalable alternative, but directly milling traditional joints often fails to produce functional results because milling induces geometric deviations—such as rounded inner corners—that alter the target geometries of the parts. Since joints rely on tightly fitting surfaces, such deviations introduce gaps or overlaps that undermine fit or block assembly. We propose to overcome this problem by (1) designing a language that represent millable geometry, and (2) co-optimizing part geometries to restore coupling. We introduce Millable Extrusion Geometry (MXG), a language for representing geometry as the outcome of milling operations performed with flat-end drill bits. MXG represents each operation as a subtractive extrusion volume defined by a tool direction and drill radius. This parameterization enables the modeling of artifact-free geometry under an idealized zero-radius drill bit, matching traditional joint designs. Increasing the radius then reveals milling-induced deviations, which compromise the integrity of the joint. To restore coupling, we formalize tight coupling in terms of both surface proximity and proximity constraints on the mill-bit paths associated with mating surfaces. We then derive two tractable, differentiable losses that enable efficient optimization of joint geometry. We evaluate our method on 30 traditional joint designs, demonstrating that it produces CNC-compatible, tightly fitting joints that approximates the original geometry. By reinterpreting traditional joints for CNC workflows, we continue the evolution of this heritage craft and help ensure its relevance in future making practices. Aditya Ganeshan, Kurt Fleischer, Wenzel Jakob, Ariel Shamir, Daniel Ritchie 0001, Takeo Igarashi, Maria Larsson |
ACM Trans. Graph. | 4 |
| 2025 | The Mokume Dataset and Inverse Modeling of Solid Wood TexturesabstractWe present the Mokume dataset for solid wood texturing consisting of 190 cube-shaped samples of various hard and softwood species documented by high-resolution exterior photographs, annual ring annotations, and volumetric computed tomography (CT) scans. A subset of samples further includes photographs along slanted cuts through the cube for validation purposes. Using this dataset, we propose a three-stage inverse modeling pipeline to infer solid wood textures using only exterior photographs. Our method begins by evaluating a neural model to localize year rings on the cube face photographs. We then extend these exterior 2D observations into a globally consistent 3D representation by optimizing a procedural growth field using a novel iso-contour loss. Finally, we synthesize a detailed volumetric color texture from the growth field. For this last step, we propose two methods with different efficiency and quality characteristics: a fast inverse procedural texture method, and a neural cellular automaton (NCA). We demonstrate the synergy between the Mokume dataset and the proposed algorithms through comprehensive comparisons with unseen captured data. We also present experiments demonstrating the efficiency of our pipeline's components against ablations and baselines. Our code, the dataset, and reconstructions are available via https://mokumeproject.github.io/. Maria Larsson, Hodaka Yamaguchi, Ehsan Pajouheshgar, I-Chao Shen, Kenji Tojo, Chia-Ming Chang 0003, Lars Hansson, Olof Broman, Takashi Ijiri, Ariel Shamir, Wenzel Jakob, Takeo Igarashi |
ACM Trans. Graph. | 10 |
| 2024 | iPose: Interactive Human Pose Reconstruction from VideoabstractReconstructing 3D human poses from video has wide applications, such as character animation and sports analysis. Automatic 3D pose reconstruction methods have demonstrated promising results, but failure cases can still appear due to the diversity of human actions, capturing conditions, and depth ambiguities. Thus, manual intervention remains indispensable, which can be time-consuming and require professional skills. We thus present iPose, an interactive tool that facilitates intuitive human pose reconstruction from a given video. Our tool incorporates both human perception in specifying pose appearance to achieve controllability, and video frame processing algorithms to achieve precision and automation. A user manipulates the projection of a 3D pose via 2D operations on top of video frames, and the 3D poses are updated correspondingly while satisfying both kinematic and video frame constraints. The pose updates are propagated temporally to reduce user workload. We evaluate the effectiveness of iPose with a user study on the 3DPW dataset and expert interviews. Li-Yi Wei, Ariel Shamir, Takeo Igarashi |
CHI | 3 |
| 2024 | Breathing Life Into Sketches Using Text-to-Video PriorsabstractA sketch is one of the most intuitive and versatile tools humans use to convey their ideas visually. An animated sketch opens another dimension to the expression of ideas and is widely used by designers for a variety of purposes. Animating sketches is a laborious process, requiring extensive experience and professional design skills. In this work, we present a method that automatically adds motion to a single-subject sketch (hence, “breathing life into it”), merely by providing a text prompt indicating the desired motion. The output is a short animation provided in vector representation, which can be easily edited. Our method does not require extensive training, but instead leverages the motion prior of a large pretrained text-to- video diffusion model using a score-distillation loss to guide the placement of strokes. To promote natural and smooth motion and to better preserve the sketch's appearance, we model the learned motion through two components. The first governs small local deformations and the second controls global affine transformations. Surprisingly, wefind that even models that struggle to generate sketch videos on their own can still serve as a useful backbone for animating abstract representations. Rinon Gal, Yael Vinker, Yuval Alaluf, Amit Bermano, Daniel Cohen-Or, Ariel Shamir, Gal Chechik |
CVPR | 6 |
| 2024 | Implicit Style-Content Separation Using B-LoRA
Yarden Frenkel, Yael Vinker, Ariel Shamir, Daniel Cohen-Or |
ECCV (10) | 3 |
| 2024 | PALP: Prompt Aligned Personalization of Text-to-Image Models
Moab Arar, Andrey Voynov, Amir Hertz, Omri Avrahami, Shlomi Fruchter, Yael Pritch, Daniel Cohen-Or, Ariel Shamir |
SIGGRAPH Asia | 8 |
| 2024 | PromptonomyViT: Multi-Task Prompt Learning Improves Video Transformers using Synthetic Scene DataabstractAction recognition models have achieved impressive results by incorporating scene-level annotations, such as objects, their relations, 3D structure, and more. However, obtaining annotations of scene structure for videos requires a significant amount of effort to gather and annotate, making these methods expensive to train. In contrast, synthetic datasets generated by graphics engines provide powerful alternatives for generating scene-level annotations across multiple tasks. In this work, we propose an approach to leverage synthetic scene data for improving video understanding. We present a multi-task prompt learning approach for video transformers, where a shared video transformer backbone is enhanced by a small set of specialized parameters for each task. Specifically, we add a set of "task prompts", each corresponding to a different task, and let each prompt predict task-related annotations. This design allows the model to capture information shared among synthetic scene tasks as well as information shared between synthetic scene tasks and a real video downstream task throughout the entire network. We refer to this approach as "Promptonomy", since the prompts model task-related structure. We propose the PromptonomyViT model (PViT), a video transformer that incorporates various types of scene-level information from synthetic data using the "Promptonomy" approach. PViT shows strong performance improvements on multiple video understanding tasks and datasets. Project page: https://ofir1080.github.io/PromptonomyViT Roei Herzig, Ofir Abramovich, Elad Ben-Avraham, Assaf Arbelle, Leonid Karlinsky, Ariel Shamir, Trevor Darrell, Amir Globerson |
WACV | 6 |
| 2024 | Learned Inference of Annual Ring Pattern of Solid WoodabstractAbstract We propose a method for inferring the internal anisotropic volumetric texture of a given wood block from annotated photographs of its external surfaces. The global structure of the annual ring pattern is represented using a continuous spatial scalar field referred to as the growth time field (GTF). First, we train a generic neural model that can represent various GTFs using procedurally generated training data. Next, we fit the generic model to the GTF of a given wood block based on surface annotations. Finally, we convert the GTF to an annual ring field (ARF) revealing the layered pattern and apply neural style transfer to render orientation‐dependent small‐scale features and colors on a cut surface. We show rendered results of various physically cut real wood samples. Our method has physical and virtual applications such as cut‐preview before subtractive fabricating solid wood artifacts and simulating object breaking. Maria Larsson, Takashi Ijiri, I-Chao Shen, Hironori Yoshida, Ariel Shamir, Takeo Igarashi |
Comput. Graph. Forum | 5 |
| 2024 | FontCLIP: A Semantic Typography Visual-Language Model for Multilingual Font ApplicationsabstractAbstract Acquiring the desired font for various design tasks can be challenging and requires professional typographic knowledge. While previous font retrieval or generation works have alleviated some of these difficulties, they often lack support for multiple languages and semantic attributes beyond the training data domains. To solve this problem, we present FontCLIP – a model that connects the semantic understanding of a large vision‐language model with typographical knowledge. We integrate typography‐specific knowledge into the comprehensive vision‐language knowledge of a pretrained CLIP model through a novel finetuning approach. We propose to use a compound descriptive prompt that encapsulates adaptively sampled attributes from a font attribute dataset focusing on Roman alphabet characters. FontCLIP's semantic typographic latent space demonstrates two unprecedented generalization abilities. First, FontCLIP generalizes to different languages including Chinese, Japanese, and Korean (CJK), capturing the typographical features of fonts across different languages, even though it was only finetuned using fonts of Roman characters. Second, FontCLIP can recognize the semantic attributes that are not presented in the training data. FontCLIP's dual‐modality and generalization abilities enable multilingual and cross‐lingual font retrieval and letter shape optimization, reducing the burden of obtaining desired fonts. Yuki Tatsukawa, I-Chao Shen, Anran Qi, Yuki Koyama 0001, Takeo Igarashi, Ariel Shamir |
Comput. Graph. Forum | 6 |
| 2023 | Evaluating machine comprehension of sketch meaning at different levels of abstraction
Kushin Mukherjee, Xuanchen Lu, Holly Huey, Yael Vinker, Rio Aguina-Kang, Ariel Shamir, Judith E. Fan |
CogSci | 6 |
| 2023 | ARO-Net: Learning Implicit Fields from Anchored Radial ObservationsabstractWe introduce anchored radial observations (ARO), a novel shape encoding for learning implicit field representation of 3D shapes that is category-agnostic and generalizable amid significant shape variations. The main idea behind our work is to reason about shapes through partial observations from a set of viewpoints, called anchors. We develop a general and unified shape representation by employing a fixed set of anchors, via Fibonacci sampling, and designing a coordinate-based deep neural network to predict the occupancy value of a query point in space. Differently from prior neural implicit models that use global shape feature, our shape encoder operates on contextual, query-specific features. To predict point occupancy, locally observed shape information from the perspective of the anchors surrounding the input query point are encoded and aggregated through an attention module, before implicit decoding is performed. We demonstrate the quality and generality of our network, coined ARO-Net, on surface reconstruction from sparse point clouds, with tests on novel and unseen object categories, “one-shape” training, and comparisons to state-of-the-art neural and classical methods for reconstruction and tessellation. Yizhi Wang 0006, Ariel Shamir, Hui Huang 0004, Hao (Richard) Zhang, Ruizhen Hu |
CVPR | 3 |
| 2023 | Semantify: Simplifying the Control of 3D Morphable Models using CLIPabstractWe present Semantify: a self-supervised method that utilizes the semantic power of CLIP language-vision foundation model [32] to simplify the control of 3D morphable models. Given a parametric model, training data is created by randomly sampling the model’s parameters, creating various shapes and rendering them. The similarity between the output images and a set of word descriptors is calculated in CLIP’s latent space. Our key idea is first to choose a small set of semantically meaningful and disentangled descriptors that characterize the 3DMM, and then learn a non-linear mapping from scores across this set to the parametric coefficients of the given 3DMM. The nonlinear mapping is defined by training a neural network without a human-in-the-loop. We present results on numerous 3DMMs: body shape models, face shape and expression models, as well as animal shapes. We demonstrate how our method defines a simple slider interface for intuitive modeling, and show how the mapping can be used to instantly fit a 3D parametric body shape to in-the-wild images. See our project page at https://omergral.github.io/Semantify/ Omer Gralnik, Guy Gafni, Ariel Shamir |
ICCV | 3 |
| 2023 | CLIPascene: Scene Sketching with Different Types and Levels of AbstractionabstractIn this paper, we present a method for converting a given scene image into a sketch using different types and multiple levels of abstraction. We distinguish between two types of abstraction. The first considers the fidelity of the sketch, varying its representation from a more precise portrayal of the input to a looser depiction. The second is defined by the visual simplicity of the sketch, moving from a detailed depiction to a sparse sketch. Using an explicit disentanglement into two abstraction axes — and multiple levels for each one — provides users additional control over selecting the desired sketch based on their personal goals and preferences. To form a sketch at a given level of fidelity and simplification, we train two MLP networks. The first network learns the desired placement of strokes, while the second network learns to gradually remove strokes from the sketch without harming its recognizability and semantics. Our approach is able to generate sketches of complex scenes including those with complex backgrounds (e.g. natural and urban settings) and subjects (e.g. animals and people) while depicting gradual abstractions of the input scene in terms of fidelity and simplicity. https://clipascene.github.io/CLIPascene/ Yael Vinker, Yuval Alaluf, Daniel Cohen-Or, Ariel Shamir |
ICCV | 4 |
| 2023 | SEVA: Leveraging sketches to evaluate alignment between human and machine visual abstractionabstractSketching is a powerful tool for creating abstract images that are sparse but meaningful. Sketch understanding poses fundamental challenges for general-purpose vision algorithms because it requires robustness to the sparsity of sketches relative to natural visual inputs and because it demands tolerance for semantic ambiguity, as sketches can reliably evoke multiple meanings. While current vision algorithms have achieved high performance on a variety of visual tasks, it remains unclear to what extent they understand sketches in a human-like way. Here we introduce $\texttt{SEVA}$, a new benchmark dataset containing approximately 90K human-generated sketches of 128 object concepts produced under different time constraints, and thus systematically varying in sparsity. We evaluated a suite of state-of-the-art vision algorithms on their ability to correctly identify the target concept depicted in these sketches and to generate responses that are strongly aligned with human response patterns on the same sketch recognition task. We found that vision algorithms that better predicted human sketch recognition performance also better approximated human uncertainty about sketch meaning, but there remains a sizable gap between model and human response patterns. To explore the potential of models that emulate human visual abstraction in generative tasks, we conducted further evaluations of a recently developed sketch generation algorithm (Vinker et al., 2022) capable of generating sketches that vary in sparsity. We hope that public release of this dataset and evaluation protocol will catalyze progress towards algorithms with enhanced capacities for human-like visual abstraction. Kushin Mukherjee, Holly Huey, Xuanchen Lu, Yael Vinker, Rio Aguina-Kang, Ariel Shamir, Judith E. Fan |
NeurIPS | 6 |
| 2023 | Domain-Agnostic Tuning-Encoder for Fast Personalization of Text-To-Image ModelsabstractText-to-image (T2I) personalization allows users to guide the creative image generation process by combining their own visual concepts in natural language prompts. Recently, encoder-based techniques have emerged as a new effective approach for T2I personalization, reducing the need for multiple images and long training times. However, most existing encoders are limited to a single-class domain, which hinders their ability to handle diverse concepts. In this work, we propose a domain-agnostic method that does not require any specialized dataset or prior information about the personalized concepts. We introduce a novel contrastive-based regularization technique to maintain high fidelity to the target concept characteristics while keeping the predicted embeddings close to editable regions of the latent space, by pushing the predicted tokens toward their nearest existing CLIP tokens. Our experimental results demonstrate the effectiveness of our approach and show how the learned tokens are more semantic than tokens predicted by unregularized models. This leads to a better representation that achieves state-of-the-art performance while being more flexible than previous methods. Moab Arar, Rinon Gal, Yuval Atzmon, Gal Chechik, Daniel Cohen-Or, Ariel Shamir, Amit Bermano |
SIGGRAPH Asia | 6 |
| 2023 | DeepPortraitDrawing: Generating human body images from freehand sketches
Xian Wu 0004, Chen Wang 0049, Hongbo Fu 0001, Ariel Shamir, Song-Hai Zhang |
Comput. Graph. | 4 |
| 2023 | HoughLaneNet: Lane detection with deep hough transform and dynamic convolution
Jia-Qi Zhang, Hao-Bin Duan, Ariel Shamir, Miao Wang 0004 |
Comput. Graph. | 4 |
| 2023 | Word-As-Image for Semantic TypographyabstractA word-as-image is a semantic typography technique where a word illustration presents a visualization of the meaning of the word, while also preserving its readability. We present a method to create word-as-image illustrations automatically. This task is highly challenging as it requires semantic understanding of the word and a creative idea of where and how to depict these semantics in a visually pleasing and legible manner. We rely on the remarkable ability of recent large pretrained language-vision models to distill textual concepts visually. We target simple, concise, black-and-white designs that convey the semantics clearly. We deliberately do not change the color or texture of the letters and do not use embellishments. Our method optimizes the outline of each letter to convey the desired concept, guided by a pretrained Stable Diffusion model. We incorporate additional loss terms to ensure the legibility of the text and the preservation of the style of the font. We show high quality and engaging results on numerous examples and compare to alternative techniques. Code and demo will be available at our project page. Shir Iluz, Yael Vinker, Amir Hertz, Daniel Berio, Daniel Cohen-Or, Ariel Shamir |
ACM Trans. Graph. | 6 |
| 2023 | Concept Decomposition for Visual Exploration and InspirationabstractA creative idea is often born from transforming, combining, and modifying ideas from existing visual examples capturing various concepts. However, one cannot simply copy the concept as a whole, and inspiration is achieved by examining certain aspects of the concept. Hence, it is often necessary to separate a concept into different aspects to provide new perspectives. In this paper, we propose a method to decompose a visual concept, represented as a set of images, into different visual aspects encoded in a hierarchical tree structure. We utilize large vision-language models and their rich latent space for concept decomposition and generation. Each node in the tree represents a sub-concept using a learned vector embedding injected into the latent space of a pretrained text-to-image model. We use a set of regularizations to guide the optimization of the embedding vectors encoded in the nodes to follow the hierarchical structure of the tree. Our method allows to explore and discover new concepts derived from the original one. The tree provides the possibility of endless visual sampling at each node, allowing the user to explore the hidden sub-concepts of the object of interest. The learned aspects in each node can be combined within and across trees to create new visual ideas, and can be used in natural language sentences to apply such aspects to new designs. Project page: https://inspirationtree.github.io/inspirationtree/ Yael Vinker, Andrey Voynov, Daniel Cohen-Or, Ariel Shamir |
ACM Trans. Graph. | 4 |
| 2023 | Rhythm is a Dancer: Music-Driven Motion Synthesis With Global StructureabstractSynthesizing human motion with a global structure, such as a choreography, is a challenging task. Existing methods tend to concentrate on local smooth pose transitions and neglect the global context or the theme of the motion. In this work, we present a music-driven motion synthesis framework that generates long-term sequences of human motions which are synchronized with the input beats, and jointly form a global structure that respects a specific dance genre. In addition, our framework enables generation of diverse motions that are controlled by the content of the music, and not only by the beat. Our music-driven dance synthesis framework is a hierarchical system that consists of three levels: pose, motif, and choreography. The pose level consists of an LSTM component that generates temporally coherent sequences of poses. The motif level guides sets of consecutive poses to form a movement that belongs to a specific distribution using a novel motion perceptual-loss. And the choreography level selects the order of the performed movements and drives the system to follow the global structure of a dance genre. Our results demonstrate the effectiveness of our music-driven framework to generate natural and consistent movements on various dance types, having control over the content of the synthesized motions, and respecting the overall structure of the dance. Andreas Aristidou, Anastasios Yiannakidis, Kfir Aberman, Daniel Cohen-Or, Ariel Shamir, Yiorgos Chrysanthou |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Learned Queries for Efficient Local AttentionabstractVision Transformers (ViT) serve as powerful vision models. Unlike convolutional neural networks, which dominated vision research in previous years, vision transformers enjoy the ability to capture long-range dependencies in the data. Nonetheless, an integral part of any transformer architecture, the self-attention mechanism, suffers from high latency and inefficient memory utilization, making it less suitable for high-resolution input images. To alleviate these shortcomings, hierarchical vision models locally employ self-attention on non-interleaving windows. This relaxation reduces the complexity to be linear in the input size; however, it limits the cross-window interaction, hurting the model performance. In this paper, we propose a new shift-invariant local attention layer, called query and attend (QnA), that aggregates the input locally in an overlapping manner, much like convolutions. The key idea behind QnA is to introduce learned queries, which allow fast and efficient implementation. We verify the effectiveness of our layer by incorporating it into a hierarchical vision transformer model. We show improvements in speed and memory complexity while achieving comparable accuracy with state-of-the-art models. Finally, our layer scales especially well with window size, requiring up to x10 less memory while being up to x5 faster than existing methods. The code is publicly available at https://github.com/moabarar/qna. Moab Arar, Ariel Shamir, Amit Bermano |
CVPR | 2 |
| 2022 | Ordered Attention for Coherent Visual StorytellingabstractWe address the problem of visual storytelling, i.e., generating a story for a given sequence of images. While each story sentence should describe a corresponding image, a coherent story also needs to be consistent and relate to both future and past images. Current approaches encode images independently, disregarding relations between images. Our approach learns to encode images with different interactions based on the story position (i.e., past image or future image). To this end, we develop a novel message-passing-like algorithm for ordered image attention (OIA) that collects interactions across all the images in the sequence. Finally, to generate the story's sentences, a second attention mechanism picks the important image attention vectors with an Image-Sentence Attention (ISA). The obtained results improve the METEOR score on the VIST dataset by 1%. Furthermore, a thorough human study confirms improvements and demonstrates that order-based interactions significantly improve coherency (64.20% \vs 28.70%). Source code available at \urlhttps://github.com/tomateb/OIAVist.git Tom Braude, Idan Schwartz, Alexander G. Schwing, Ariel Shamir |
ACM Multimedia | 4 |
| 2022 | Semantic Segmentation in Art PaintingsabstractAbstract Semantic segmentation is a difficult task even when trained in a supervised manner on photographs. In this paper, we tackle the problem of semantic segmentation of artistic paintings, an even more challenging task because of a much larger diversity in colors, textures, and shapes and because there are no ground truth annotations available for segmentation. We propose an unsupervised method for semantic segmentation of paintings using domain adaptation. Our approach creates a training set of pseudo‐paintings in specific artistic styles by using style‐transfer on the PASCAL VOC 2012 dataset, and then applies domain confusion between PASCAL VOC 2012 and real paintings. These two steps build on a new dataset we gathered called DRAM (Diverse Realism in Art Movements) composed of figurative art paintings from four movements, which are highly diverse in pattern, color, and geometry. To segment new paintings, we present a composite multi‐domain adaptation method that trains on each sub‐domain separately and composes their solutions during inference time. Our method provides better segmentation results not only on the specific artistic movements of DRAM, but also on other, unseen ones. We compare our approach to alternative methods and show applications of semantic segmentation in art paintings. The code and models for our approach are publicly available at: https://github.com/Nadavc220/SemanticSegmentationInArtPaintings . Nadav Cohen 0003, Yael Newman, Ariel Shamir |
Comput. Graph. Forum | 3 |
| 2022 | CAST: Character labeling in Animation using Self-supervision by TrackingabstractAbstract Cartoons and animation domain videos have very different characteristics compared to real‐life images and videos. In addition, this domain carries a large variability in styles. Current computer vision and deep‐learning solutions often fail on animated content because they were trained on natural images. In this paper we present a method to refine a semantic representation suitable for specific animated content. We first train a neural network on a large‐scale set of animation videos and use the mapping to deep features as an embedding space. Next, we use self‐supervision to refine the representation for any specific animation style by gathering many examples of animated characters in this style, using a multi‐object tracking. These examples are used to define triplets for contrastive loss training. The refined semantic space allows better clustering of animated characters even when they have diverse manifestations. Using this space we can build dictionaries of characters in an animation videos, and define specialized classifiers for specific stylistic content (e.g., characters in a specific animation series) with very little user effort. These classifiers are the basis for automatically labeling characters in animation videos. We present results on a collection of characters in a variety of animation styles. Code and resources are available at: https://github.com/oronnir/CAST . Oron Nir, Gal Rapoport, Ariel Shamir |
Comput. Graph. Forum | 3 |
| 2022 | NPRportrait 1.0: A three-level benchmark for non-photorealistic rendering of portraitsabstractRecently, there has been an upsurge of activity in image-based non-photorealistic rendering (NPR), and in particular portrait image stylisation, due to the advent of neural style transfer (NST). However, the state of performance evaluation in this field is poor, especially compared to the norms in the computer vision and machine learning communities. Unfortunately, the task of evaluating image stylisation is thus far not well defined, since it involves subjective, perceptual, and aesthetic aspects. To make progress towards a solution, this paper proposes a new structured, three-level, benchmark dataset for the evaluation of stylised portrait images. Rigorous criteria were used for its construction, and its consistency was validated by user studies. Moreover, a new methodology has been developed for evaluating portrait stylisation algorithms, which makes use of the different benchmark levels as well as annotations provided by user studies regarding the characteristics of the faces. We perform evaluation for a wide variety of image stylisation methods (both portrait-specific and general purpose, and also both traditional NPR approaches and NST) using the new benchmark dataset. Paul L. Rosin, Yukun Lai, David Mould, Ran Yi 0002, Itamar Berger, Lars Doyle, Seungyong Lee 0001, Chuan Li 0001, Yong-Jin Liu 0001, Amir Semmo, Ariel Shamir, Minjung Son 0001, Holger Winnemöller |
Comput. Vis. Media | 11 |
| 2022 | User-Guided Deep Human Image Matting Using Arbitrary TrimapsabstractImage matting is widely studied for accurate foreground extraction. Most algorithms, including deep-learning based solutions, require a carefully edited trimap. Recent works attempt to combine the segmentation stage and matting stage in one CNN model, but errors occurring at the segmentation stage lead to unsatisfactory matte. We propose a user-guided approach for practical human matting. More precisely, we provide a good automatic initial matting and a natural way of interaction that reduces the workload of drawing trimaps and allows users to guide the matting in ambiguous situation. We also combine the segmentation and matting stage in an end-to-end CNN architecture and introduce a residual-learning module to support convenient stroke-based interaction. The proposed model learns to propagate the input trimap and modify the deep image features, which can efficiently correct the segmentation errors. Our model supports arbitrary forms of trimaps from carefully edited to totally unknown maps. Our model also allows users to choose from different foreground estimations according to their preference. We collected a large human matting dataset consisting of 12K real-world human images with complex background and human-object relations. The proposed model is trained on the new dataset with a novel trimap generation strategy that enables the model to tackle different test situations and highly improves the interaction efficiency. Our method outperforms other state-of-the-art automatic methods and achieve competitive accuracy when high-quality trimaps are provided. Experiments indicate that our interactive matting strategy is superior to separately estimating the trimap and alpha matte using two models. Xiaonan Fang 0001, Song-Hai Zhang, Tao Chen 0015, Xian Wu 0004, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | CLIPasso: semantically-aware object sketchingabstractAbstraction is at the heart of sketching due to the simple and minimal nature of line drawings. Abstraction entails identifying the essential visual properties of an object or scene, which requires semantic understanding and prior knowledge of high-level concepts. Abstract depictions are therefore challenging for artists, and even more so for machines. We present CLIPasso, an object sketching method that can achieve different levels of abstraction, guided by geometric and semantic simplifications. While sketch generation methods often rely on explicit sketch datasets for training, we utilize the remarkable ability of CLIP (Contrastive-Language-Image-Pretraining) to distill semantic concepts from sketches and images alike. We define a sketch as a set of Bézier curves and use a differentiable rasterizer to optimize the parameters of the curves directly with respect to a CLIP-based perceptual loss. The abstraction degree is controlled by varying the number of strokes. The generated sketches demonstrate multiple levels of abstraction while maintaining recognizability, underlying structure, and essential visual components of the subject drawn. Yael Vinker, Ehsan Pajouheshgar, Jessica Y. Bo, Roman Bachmann 0001, Amit Bermano, Daniel Cohen-Or, Amir Zamir, Ariel Shamir |
ACM Trans. Graph. | 8 |
| 2021 | Deep Symmetric Network for Underexposed Image Enhancement with Recurrent Attentional LearningabstractUnderexposed image enhancement is of importance in many research domains. In this paper, we take this problem as image feature transformation between the underexposed image and its paired enhanced version, and we propose a deep symmetric network for the issue. Our symmetric network adapts invertible neural networks (INN) for bidirectional feature learning between images, and to ensure the mutual propagation invertible we specifically construct two pairs of encoder-decoder with the same pretrained parameters. This invertible mechanism with bidirectional feature transformations enable us to both avoid colour bias and recover the content effectively for image enhancement. In addition, we propose a new recurrent residual-attention module (RRAM), where the recurrent learning network is designed to gradually perform the desired colour adjustments. Ablation experiments are executed to show the role of each component of our new architecture. We conduct a large number of experiments on two datasets to demonstrate that our method achieves the state-of-the-art effect in underexposed image enhancement. Code is available at https://www.shaopinglu.net/proj-iccv21/ImageEnhancement.html. Shao-Ping Lu, Tao Chen 0015, Zhenglu Yang, Ariel Shamir |
ICCV | 5 |
| 2021 | Image resizing by reconstruction from deep featuresabstractTraditional image resizing methods usually work in pixel space and use various saliency measures. The challenge is to adjust the image shape while trying to preserve important content. In this paper we perform image resizing in feature space using the deep layers of a neural network containing rich important semantic information. We directly adjust the image feature maps, extracted from a pre-trained classification network, and reconstruct the resized image using neural-network based optimization. This novel approach leverages the hierarchical encoding of the network, and in particular, the high-level discriminative power of its deeper layers, that can recognize semantic regions and objects, thereby allowing maintenance of their aspect ratios. Our use of reconstruction from deep features results in less noticeable artifacts than use of imagespace resizing operators. We evaluate our method on benchmarks, compare it to alternative approaches, and demonstrate its strengths on challenging images. Dov Danon, Moab Arar, Daniel Cohen-Or, Ariel Shamir |
Comput. Vis. Media | 4 |
| 2021 | On the role of geometry in geo-localizationabstractConsider the geo-localization task of finding the pose of a camera in a large 3D scene from a single image. Most existing CNN-based methods use as input textured images. We aim to experimentally explore whether texture and correlation between nearby images are necessary in a CNN-based solution for the geo-localization task. To do so, we consider lean images , textureless projections of a simple 3D model of a city. They only contain information related to the geometry of the scene viewed (edges, faces, and relative depth). The main contributions of this paper are: (i) to demonstrate the ability of CNNs to recover camera pose using lean images; and (ii) to provide insight into the role of geometry in the CNN learning process. Moti Kadosh, Yael Moses, Ariel Shamir |
Comput. Vis. Media | 3 |
| 2021 | Prominent Structures for Video Analysis and EditingabstractWe present prominent structures in video, a representation of visually strong, spatially sparse and temporally stable structural units, for use in video analysis and editing. With a novel quality measurement of prominent structures in video, we develop a general framework for prominent structure computation, and an efficient hierarchical structure alignment algorithm between a pair of videos. The prominent structural unit map is proposed to encode both binary prominence guidance and numerical strength and geometry details for each video frame. Even though the detailed appearance of videos could be visually different, the proposed alignment algorithm can find matched prominent structure sub-volumes. Prominent structures in video support a wide range of video analysis and editing applications including graphic match-cut between successive videos, instant cut editing, finding transition portals from a video collection, structure-aware video re-ranking, visualizing human action differences, etc. Miao Wang 0004, Xiaonan Fang 0001, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | Adult2child: Motion Style Transfer using CycleGANsabstractChild characters are commonly seen in leading roles in top-selling video games. Previous studies have shown that child motions are perceptually and stylistically different from those of adults. Creating motion for these characters by motion capturing children is uniquely challenging because of confusion, lack of patience and regulations. Retargeting adult motion, which is much easier to record, onto child skeletons, does not capture the stylistic differences. In this paper, we propose that style translation is an effective way to transform adult motion capture data to the style of child motion. Our method is based on CycleGAN, which allows training on a relatively small number of sequences of child and adult motions that do not even need to be temporally aligned. Our adult2child network converts short sequences of motions called motion words from one domain to the other. The network was trained using a motion capture database collected by our team containing 23 locomotion and exercise motions. We conducted a perception study to evaluate the success of style translation algorithms, including our algorithm and recently presented style translation neural networks. Results show that the translated adult motions are recognized as child motions significantly more often than adult motions. Yuzhu Dong, Andreas Aristidou, Ariel Shamir, Moshe Mahler, Eakta Jain |
MIG | 3 |
| 2020 | Deep Portrait Image Completion and ExtrapolationabstractGeneral image completion and extrapolation methods often fail on portrait images where parts of the human body need to be recovered -a task that requires accurate human body structure and appearance synthesis. We present a twostage deep learning framework for tackling this problem. In the first stage, given a portrait image with an incomplete human body, we extract a complete, coherent human body structure through a human parsing network, which focuses on structure recovery inside the unknown region with the help of full-body pose estimation. In the second stage, we use an image completion network to fill the unknown region, guided by the structure map recovered in the first stage. For realistic synthesis the completion network is trained with both perceptual loss and conditional adversarial loss.We further propose a face refinement network to improve the fidelity of the synthesized face region. We evaluate our method on publicly-available portrait image datasets, and show that it outperforms other state-of-the-art general image completion methods. Our method enables new portrait image editing applications such as occlusion removal and portrait extrapolation. We further show that the proposed general learning framework can be applied to other types of images, e.g. animal images. Xian Wu 0004, Ruilong Li, Jian-Cheng Liu, Jue Wang 0001, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 6 |
| 2019 | Can Children Understand Machine Learning Concepts?: The Effect of Uncovering Black BoxesabstractMachine Learning services are integrated into various aspects of everyday life. Their underlying processes are typically black-boxed to increase ease-of-use. Consequently, children lack the opportunity to explore such processes and develop essential mental models. We present a gesture recognition research platform, designed to support learning from experience by uncovering Machine Learning building blocks: Data Labeling and Evaluation. Children used the platform to perform physical gestures, iterating between sampling and evaluation. Their understanding was tested in a pre/post experimental design, in three conditions: learning activity uncovering Data Labeling only, Evaluation only, or both. Our findings show that both building blocks are imperative to enhance children's understanding of basic Machine Learning concepts. Children were able to apply their new knowledge to everyday life context, including personally meaningful applications. We conclude that children's interaction with uncovered black boxes of Machine Learning contributes to a better understanding of the world around them. Tom Hitron, Yoav Orlev, Iddo Wald, Ariel Shamir, Hadas Erel, Oren Zuckerman |
CHI | 4 |
| 2019 | Learning Explicit Smoothing Kernels for Joint Image FilteringabstractAbstract Smoothing noises while preserving strong edges in images is an important problem in image processing. Image smoothing filters can be either explicit (based on local weighted average) or implicit (based on global optimization). Implicit methods are usually time‐consuming and cannot be applied to joint image filtering tasks, i.e., leveraging the structural information of a guidance image to filter a target image. Previous deep learning based image smoothing filters are all implicit and unavailable for joint filtering. In this paper, we propose to learn explicit guidance feature maps as well as offset maps from the guidance image and smoothing parameter that can be utilized to smooth the input itself or to filter images in other target domains. We design a deep convolutional neural network consisting of a fully‐convolution block for guidance and offset maps extraction together with a stacked spatially varying deformable convolution block for joint image filtering. Our models can approximate several representative image smoothing filters with high accuracy comparable to state‐of‐the‐art methods, and serve as general tools for other joint image filtering tasks, such as color interpolation, depth map upsampling, saliency map upsampling, flash/non‐flash image denoising and RGB/NIR image denoising. Xiaonan Fang 0001, Miao Wang 0004, Ariel Shamir, Shi-Min Hu 0001 |
Comput. Graph. Forum | 3 |
| 2019 | Deep Online Video Stabilization With Multi-Grid Warping Transformation LearningabstractVideo stabilization techniques are essential for most hand-held captured videos due to high-frequency shakes. Several 2D, 2.5D and 3D-based stabilization techniques have been presented previously, but to our knowledge, no solutions based on deep neural networks had been proposed to date. The main reason for this omission is shortage in training data as well as the challenge of modeling the problem using neural networks. In this paper, we present a video stabilization technique using a convolutional neural network. Previous works usually propose an offline algorithm that smoothes a holistic camera path based on feature matching. Instead, we focus on low-latency, real-time camera path smoothing, that does not explicitly represent the camera path, and does not use future frames. Our neural network model, called StabNet, learns a set of mesh-grid transformations progressively for each input frame from the previous set of stabalized camera frames, and creates stable corresponding latent camera paths implicitly. To train the network, we collect a dataset of synchronized steady and unsteady video pairs via a specially designed hand-held hardware. Experimental results show that our proposed online method performs comparatively to traditional offline video stabilization methods without using future frames, while running about 10× faster. More importantly, our proposed StabNet is able to handle low-quality videos such as night-scene videos, watermarked videos, blurry videos and noisy videos, where existing methods fail in feature extraction or matching. Miao Wang 0004, Guo-Ye Yang, Jin-Kun Lin, Song-Hai Zhang, Ariel Shamir, Shao-Ping Lu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | GRAINS: Generative Recursive Autoencoders for INdoor ScenesabstractWe present a generative neural network that enables us to generate plausible 3D indoor scenes in large quantities and varieties, easily and highly efficiently. Our key observation is that indoor scene structures are inherently hierarchical . Hence, our network is not convolutional; it is a recursive neural network, or RvNN. Using a dataset of annotated scene hierarchies, we train a variational recursive autoencoder , or RvNN-VAE, which performs scene object grouping during its encoding phase and scene generation during decoding. Specifically, a set of encoders are recursively applied to group 3D objects based on support, surround, and co-occurrence relations in a scene, encoding information about objects’ spatial properties, semantics , and relative positioning with respect to other objects in the hierarchy. By training a variational autoencoder (VAE), the resulting fixed-length codes roughly follow a Gaussian distribution. A novel 3D scene can be generated hierarchically by the decoder from a randomly sampled code from the learned distribution. We coin our method GRAINS, for Generative Recursive Autoencoders for INdoor Scenes. We demonstrate the capability of GRAINS to generate plausible and diverse 3D indoor scenes and compare with existing methods for 3D scene synthesis. We show applications of GRAINS including 3D scene modeling from 2D layouts, scene editing, and semantic scene segmentation via PointNet whose performance is boosted by the large quantity and variety of 3D scenes generated by our method. Manyi Li, Akshay Gadi Patil, Kai Xu 0004, Siddhartha Chaudhuri, Owais Khan, Ariel Shamir, Changhe Tu, Baoquan Chen, Daniel Cohen-Or, Hao (Richard) Zhang |
ACM Trans. Graph. | 6 |
| 2019 | Write-a-video: computational video montage from themed textabstractWe present Write-A-Video , a tool for the creation of video montage using mostly text-editing. Given an input themed text and a related video repository either from online websites or personal albums, the tool allows novice users to generate a video montage much more easily than current video editing tools. The resulting video illustrates the given narrative, provides diverse visual content, and follows cinematographic guidelines. The process involves three simple steps: (1) the user provides input, mostly in the form of editing the text, (2) the tool automatically searches for semantically matching candidate shots from the video repository, and (3) an optimization method assembles the video montage. Visual-semantic matching between segmented text and shots is performed by cascaded keyword matching and visual-semantic embedding, that have better accuracy than alternative solutions. The video assembly is formulated as a hybrid optimization problem over a graph of shots, considering temporal constraints, cinematography metrics such as camera movement and tone, and user-specified cinematography idioms. Using our system, users without video editing experience are able to generate appealing videos. Miao Wang 0004, Shi-Min Hu 0001, Shing-Tung Yau, Ariel Shamir |
ACM Trans. Graph. | 5 |
| 2019 | The face of art: landmark detection and geometric style in portraitsabstractFacial Landmark detection in natural images is a very active research domain. Impressive progress has been made in recent years, with the rise of neural-network based methods and large-scale datasets. However, it is still a challenging and largely unexplored problem in the artistic portraits domain. Compared to natural face images, artistic portraits are much more diverse. They contain a much wider style variation in both geometry and texture and are more complex to analyze. Moreover, datasets that are necessary to train neural networks are unavailable. We propose a method for artistic augmentation of natural face images that enables training deep neural networks for landmark detection in artistic portraits. We utilize conventional facial landmarks datasets, and transform their content from natural images into "artistic face" images. In addition, we use a feature-based landmark correction step, to reduce the dependency between the different facial features, which is necessary due to position and shape variations of facial landmarks in artworks. To evaluate our landmark detection framework, we created an "Artistic-Faces" dataset, containing 160 artworks of various art genres, artists and styles, with a large variation in both geometry and texture. Using our method, we can detect facial features in artistic portraits and analyze their geometric style. This allows the definition of signatures for artistic styles of artworks and artists, that encode both the geometry and the texture style. It also allows us to present a geometric-aware style transfer method for portraits. Jordan Yaniv, Yael Newman, Ariel Shamir |
ACM Trans. Graph. | 3 |
| 2018 | Computer-aided design of resistance micro-fluidic circuits for 3D printing
Elishai Ezra Tsur, Ariel Shamir |
Comput. Aided Des. | 2 |
| 2018 | Self-similarity Analysis for Motion Capture CleaningabstractAbstract Motion capture sequences may contain erroneous data, especially when the motion is complex or performers are interacting closely and occlusions are frequent. Common practice is to have specialists visually detect the abnormalities and fix them manually. In this paper, we present a method to automatically analyze and fix motion capture sequences by using self‐similarity analysis. The premise of this work is that human motion data has a high‐degree of self‐similarity. Therefore, given enough motion data, erroneous motions are distinct when compared to other motions. We utilizemotion‐wordsthat consist of short sequences of transformations of groups of joints around a given motion frame. We search for the K‐nearest neighbors (KNN) set of each word using dynamic time warping and use it to detect and fix erroneous motions automatically. We demonstrate the effectiveness of our method in various examples, and evaluate by comparing to alternative methods and to manual cleaning. Andreas Aristidou, Daniel Cohen-Or, Jessica K. Hodgins, Ariel Shamir |
Comput. Graph. Forum | 4 |
| 2018 | Inverse Kinematics Techniques in Computer Graphics: A SurveyabstractAbstract Inverse kinematics (IK) is the use of kinematic equations to determine the joint parameters of a manipulator so that the end effector moves to a desired position; IK can be applied in many areas, including robotics, engineering, computer graphics and video games. In this survey, we present a comprehensive review of the IK problem and the solutions developed over the years from the computer graphics point of view. The paper starts with the definition of forward and IK, their mathematical formulations and explains how to distinguish the unsolvable cases, indicating when a solution is available. The IK literature in this report is divided into four main categories: the analytical , the numerical , the data‐driven and the hybrid methods. A timeline illustrating key methods is presented, explaining how the IK approaches have progressed over the years. The most popular IK methods are discussed with regard to their performance, computational cost and the smoothness of their resulting postures, while we suggest which IK family of solvers is best suited for particular problems. Finally, we indicate the limitations of the current IK methodologies and propose future research directions. Andreas Aristidou, Joan Lasenby, Yiorgos Chrysanthou, Ariel Shamir |
Comput. Graph. Forum | 4 |
| 2018 | Preface
Shi-Min Hu 0001, Cewu Lu, Ariel Shamir |
J. Comput. Sci. Technol. | 3 |
| 2018 | Hyper-Lapse From Multiple Spatially-Overlapping VideosabstractHyper-lapse video with high speed-up rate is an efficient way to overview long videos, such as a human activity in first-person view. Existing hyper-lapse video creation methods produce a fast-forward video effect using only one video source. In this paper, we present a novel hyper-lapse video creation approach based on multiple spatially-overlapping videos. We assume the videos share a common view or location, and find transition points where jumps from one video to another may occur. We represent the collection of videos using a hyper-lapse transition graph; the edges between nodes represent possible hyper-lapse frame transitions. To create a hyper-lapse video, a shortest path search is performed on this digraph to optimize frame sampling and assembly simultaneously. Finally, we render the hyper-lapse results using video stabilization and appearance smoothing techniques on the selected frames. Our technique can synthesize novel virtual hyper-lapse routes, which may not exist originally. We show various application results on both indoor and outdoor video collections with static scenes, moving objects, and crowds. Miao Wang 0004, Jun-Bang Liang, Song-Hai Zhang, Shao-Ping Lu, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | BiggerSelfie: Selfie Video Expansion With Hand-Held CameraabstractSelfie photography from the hand-held camera is becoming a popular media type. Although being convenient and flexible, it suffers from low camera motion stability, small field of view, and limited background content. These limitations can annoy users, especially, when touring a place of interest and taking selfie videos. In this paper, we present a novel method to create what we call a BiggerSelfie that deals with these shortcomings. Using a video of the environment that has partial content overlap with the selfie video, we stitch plausible frames selected from the environment video to the original selfie frames and stabilize the composed video content with a portrait-preserving constraint. Using the proposed method, one can easily obtain a stable selfie video with expanded background content by merely capturing some background shots. We show various results and several evaluations to demonstrate the applicability of our method. Miao Wang 0004, Ariel Shamir, Guo-Ye Yang, Jin-Kun Lin, Shao-Ping Lu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Deep motifs and motion signaturesabstractMany analysis tasks for human motion rely on high-level similarity between sequences of motions, that are not an exact matches in joint angles, timing, or ordering of actions. Even the same movements performed by the same person can vary in duration and speed. Similar motions are characterized by similar sets of actions that appear frequently. In this paper we introduce motion motifs and motion signatures that are a succinct but descriptive representation of motion sequences. We first break the motion sequences to short-term movements called motion words, and then cluster the words in a high-dimensional feature space to find motifs. Hence, motifs are words that are both common and descriptive, and their distribution represents the motion sequence. To cluster words and find motifs, the challenge is to define an effective feature space, where the distances among motion words are semantically meaningful, and where variations in speed and duration are handled. To this end, we use a deep neural network to embed the motion words into feature space using a triplet loss function. To define a signature, we choose a finite set of motion-motifs, creating a bag-of-motifs representation for the sequence. Motion signatures are agnostic to movement order, speed or duration variations, and can distinguish fine-grained differences between motions of the same class. We illustrate examples of characterizing motion sequences by motifs, and for the use of motion signatures in a number of applications. Andreas Aristidou, Daniel Cohen-Or, Jessica K. Hodgins, Yiorgos Chrysanthou, Ariel Shamir |
ACM Trans. Graph. | 5 |
| 2018 | Predictive and generative neural networks for object functionalityabstractHumans can predict the functionality of an object even without any surroundings, since their knowledge and experience would allow them to "hallucinate" the interaction or usage scenarios involving the object. We develop predictive and generative deep convolutional neural networks to replicate this feat. Specifically, our work focuses on functionalities of man-made 3D objects characterized by human-object or object-object interactions. Our networks are trained on a database of scene contexts, called interaction contexts , each consisting of a central object and one or more surrounding objects, that represent object functionalities. Given a 3D object in isolation , our functional similarity network (fSIM-NET), a variation of the triplet network, is trained to predict the functionality of the object by inferring functionality-revealing interaction contexts. fSIM-NET is complemented by a generative network (iGEN-NET) and a segmentation network (iSEG-NET). iGEN-NET takes a single voxelized 3D object with a functionality label and synthesizes a voxelized surround, i.e., the interaction context which visually demonstrates the corresponding functionality. iSEG-NET further separates the interacting objects into different groups according to their interaction types. Ruizhen Hu, Zihao Yan, Oliver van Kaick, Ariel Shamir, Hao (Richard) Zhang, Hui Huang 0004 |
ACM Trans. Graph. | 5 |
| 2017 | Learn How to Choose: Independent Detectors Versus Composite Visual PhrasesabstractMost approaches for scene parsing, recognition or retrieval use detectors that are either (i) independently trained or (ii) jointly trained for conjunctions of object-object or object-attribute phrases. We posit that neither of these two extremes is uniformly optimal, in terms of performance, across all categories and conjunctions. The choice of whether one should train an independent or composite detector should be made for each possible conjunction separately, and depends on the statistics of the dataset as well. For example, person holding phone may be more accurately modeled using a single composite detector, while tall person may be more accurately modeled as combination of two detectors. We extensively study this issue in the context of multiple problems and datasets. Further, for e ciency, we propose a predictor that is based on a number of category speci c features (e.g., sample size, entropy, etc.) for whether independent or joint composite detector may be more accurate for a given conjunction. We show that our prediction and selection mechanism generalizes and leads to improved performance on a number of large-scale datasets and vision tasks. Guy Rosenthal, Ariel Shamir, Leonid Sigal |
WACV | 2 |
| 2017 | Learning to predict part mobility from a single static snapshotabstractWe introduce a method for learning a model for the mobility of parts in 3D objects. Our method allows not only to understand the dynamic functionalities of one or more parts in a 3D object, but also to apply the mobility functions to static 3D models. Specifically, the learned part mobility model can predict mobilities for parts of a 3D object given in the form of a single static snapshot reflecting the spatial configuration of the object parts in 3D space, and transfer the mobility from relevant units in the training data. The training data consists of a set of mobility units of different motion types. Each unit is composed of a pair of 3D object parts (one moving and one reference part), along with usage examples consisting of a few snapshots capturing different motion states of the unit. Taking advantage of a linearity characteristic exhibited by most part motions in everyday objects, and utilizing a set of part-relation descriptors, we define a mapping from static snapshots to dynamic units. This mapping employs a motion-dependent snapshot-to-unit distance obtained via metric learning. We show that our learning scheme leads to accurate motion prediction from single static snapshots and allows proper motion transfer. We also demonstrate other applications such as motion-driven object detection and motion hierarchy construction. Ruizhen Hu, Wenchao Li 0005, Oliver van Kaick, Ariel Shamir, Hao (Richard) Zhang, Hui Huang 0004 |
ACM Trans. Graph. | 4 |
| 2017 | Retrieval on Parametric Shape CollectionsabstractWhile collections of parametric shapes are growing in size and use, little progress has been made on the fundamental problem of shape-based matching and retrieval for parametric shapes in a collection. The search space for such collections is both discrete (number of shapes) and continuous (parameter values). In this work, we propose representing this space using descriptors that have shown to be effective for single shape retrieval. While single shapes can be represented as points in a descriptor space, parametric shapes are mapped into larger continuous regions. For smooth descriptors, we can assume that these regions are bounded low-dimensional manifolds where the dimensionality is given by the number of shape parameters. We propose representing these manifolds with a set of primitives, namely, points and bounded tangent spaces. Our algorithm describes how to define these primitives and how to use them to construct a manifold approximation that allows accurate and fast retrieval. We perform an analysis based on curvature, boundary evaluation, and the allowed approximation error to select between primitive types. We show how to compute decision variables with no need for empirical parameter adjustments and discuss theoretical guarantees on retrieval accuracy. We validate our approach with experiments that use different types of descriptors on a collection of shapes from multiple categories. Adriana Schulz, Ariel Shamir, Ilya Baran, David I. W. Levin, Pitchaya Sitthi-amorn, Wojciech Matusik |
ACM Trans. Graph. | 2 |
| 2016 | Line-Drawing Video StylizationabstractAbstract We present a method to automatically convert videos and CG animations to stylized animated line drawings. Using a data‐driven approach, the animated drawings can follow the sketching style of a specific artist. Given an input video, we first extract edges from the video frames and vectorize them to curves. The curves are matched to strokes from an artist's library, while following the artist's stroke distribution and characteristics. The key challenge in this process is to match the large number of curves in the frames over time, despite topological and geometric changes, allowing to maintain temporal coherence in the output animation. We solve this problem using constrained optimization to build correspondences between tracked points and create smooth sheets over time. These sheets are then replaced with strokes from the artist's database to render the final animation. We evaluate our tracking algorithm on various examples and show stylized animation results based on various artists. N. Ben-Zvi, José Bento 0001, Moshe Mahler, Jessica K. Hodgins, Ariel Shamir |
Comput. Graph. Forum | 5 |
| 2016 | Learning how objects function via co-analysis of interactionsabstractWe introduce a co-analysis method which learns a functionality model for an object category, e.g., strollers or backpacks. Like previous works on functionality, we analyze object-to-object interactions and intra-object properties and relations. Differently from previous works, our model goes beyond providing a functionality-oriented descriptor for a single object; it prototypes the functionality of a category of 3D objects by co-analyzing typical interactions involving objects from the category. Furthermore, our co-analysis localizes the studied properties to the specific locations, or surface patches, that support specific functionalities, and then integrates the patch-level properties into a category functionality model. Thus our model focuses on the how , via common interactions, and where , via patch localization, of functionality analysis. Given a collection of 3D objects belonging to the same category, with each object provided within a scene context, our co-analysis yields a set of proto-patches , each of which is a patch prototype supporting a specific type of interaction, e.g., stroller handle held by hand. The learned category functionality model is composed of proto-patches, along with their pairwise relations, which together summarize the functional properties of all the patches that appear in the input object category. With the learned functionality models for various object categories serving as a knowledge base, we are able to form a functional understanding of an individual 3D object, without a scene context. With patch localization in the model, functionality-aware modeling , e.g, functional object enhancement and the creation of functional object hybrids, is made possible. Ruizhen Hu, Oliver van Kaick, Bojian Wu, Hui Huang 0004, Ariel Shamir, Hao (Richard) Zhang |
ACM Trans. Graph. | 5 |
| 2016 | Stochastic structural analysis for context-aware design and fabricationabstractIn this paper we propose failure probabilities as a semantically and mechanically meaningful measure of object fragility. We present a stochastic finite element method which exploits fast rigid body simulation and reduced-space approaches to compute spatially varying failure probabilities. We use an explicit rigid body simulation to emulate the real-world loading conditions an object might experience, including persistent and transient frictional contact, while allowing us to combine several such scenarios together. Thus, our estimates better reflect real-world failure modes than previous methods. We validate our results using a series of real-world tests. Finally, we show how to embed failure probabilities into a stress constrained topology optimization which we use to design objects such as weight bearing brackets and robust 3D printable objects. Timothy R. Langlois, Ariel Shamir, Daniel Dror, Wojciech Matusik, David I. W. Levin |
ACM Trans. Graph. | 2 |
| 2015 | Mirror Puppeteering: Animating Toy Robots in Front of a WebcamabstractMirror Puppeteering is a system for easily creating gestures ("animations") for robotic toys, custom robots, and virtual characters. Lay users can record animations by simply moving a robot's limbs in front of a webcam. Makers and hobbyists can use the system to easily set up their custom-built robots for animation. Gamers and amateur animators can real-time control or save animations for virtual characters. Our system works by tracking circular markers on the robot's surface and translating these into motor commands, using a calibration map between marker locations in camera space and motor angles. New robots can be quickly set up for Mirror Puppeteering without knowledge of the robot's 3D structure, as we demonstrate on several robots. In a user study, participants found our method more enjoyable, usable, easy to learn, and successful than traditional animation methods. Ronit Slyper, Guy Hoffman, Ariel Shamir |
TEI | 3 |
| 2015 | A Survey on Data-Driven Video CompletionabstractAbstract Image completion techniques aim to complete selected regions of an image in a natural looking manner with little or no user interaction. Video Completion, the space–time equivalent of the image completion problem, inherits and extends both the difficulties and the solutions of the original 2D problem, but also imposes new ones—mainly temporal coherency and space complexity (videos contain significantly more information than images). Data‐driven approaches to completion have been established as a favoured choice, especially when large regions have to be filled. In this survey, we present the current state of the art in data‐driven video completion techniques. For unacquainted researchers, we aim to provide a broad yet easy to follow introduction to the subject (including an extensive review of the image completion foundations) and early guidance to the challenges ahead. For a versed reader, we offer a comprehensive review of the contemporary techniques, sectioned out by their approaches to key aspects of the problem. Shachar Ilan, Ariel Shamir |
Comput. Graph. Forum | 2 |
| 2015 | Interaction context (ICON): towards a geometric functionality descriptorabstractWe introduce a contextual descriptor which aims to provide a geometric description of the functionality of a 3D object in the context of a given scene. Differently from previous works, we do not regard functionality as an abstract label or represent it implicitly through an agent. Our descriptor, called interaction context or ICON for short, explicitly represents the geometry of object-to-object interactions. Our approach to object functionality analysis is based on the key premise that functionality should mainly be derived from interactions between objects and not objects in isolation. Specifically, ICON collects geometric and structural features to encode interactions between a central object in a 3D scene and its surrounding objects. These interactions are then grouped based on feature similarity, leading to a hierarchical structure. By focusing on interactions and their organization, ICON is insensitive to the numbers of objects that appear in a scene, the specific disposition of objects around the central object, or the objects' fine-grained geometry. With a series of experiments, we demonstrate the potential of ICON in functionality-oriented shape processing, including shape retrieval (either directly or by complementing existing shape descriptors), segmentation, and synthesis. Ruizhen Hu, Chenyang Zhu 0002, Oliver van Kaick, Ligang Liu 0001, Ariel Shamir, Hao (Richard) Zhang |
ACM Trans. Graph. | 5 |
| 2015 | Gaze-Driven Video Re-EditingabstractGiven the current profusion of devices for viewing media, video content created at one aspect ratio is often viewed on displays with different aspect ratios. Many previous solutions address this problem by retargeting or resizing the video, but a more general solution would re-edit the video for the new display. Our method employs the three primary editing operations: pan, cut, and zoom. We let viewers implicitly reveal what is important in a video by tracking their gaze as they watch the video. We present an algorithm that optimizes the path of a cropping window based on the collected eyetracking data, finds places to cut, and computes the size of the cropping window. We present results on a variety of video clips, including close-up and distant shots, and stationary and moving cameras. We conduct two experiments to evaluate our results. First, we eyetrack viewers on the result videos generated by our algorithm, and second, we perform a subjective assessment of viewer preference. These experiments show that viewer gaze patterns are similar on our result videos and on the original video clips, and that viewers prefer our results to an optimized crop-and-warp algorithm. Eakta Jain, Yaser Sheikh, Ariel Shamir, Jessica K. Hodgins |
ACM Trans. Graph. | 3 |
| 2015 | AutoConnect: computational design of 3D-printable connectorsabstractWe present AutoConnect, an automatic method that creates customized, 3D-printable connectors attaching two physical objects together. Users simply position and orient virtual models of the two objects that they want to connect and indicate some auxiliary information such as weight and dimensions. Then, AutoConnect creates several alternative designs that users can choose from for 3D printing. The design of the connector is created by combining two holders, one for each object. We categorize the holders into two types. The first type holds standard objects such as pipes and planes. We utilize a database of parameterized mechanical holders and optimize the holder shape based on the grip strength and material consumption. The second type holds free-form objects. These are procedurally generated shell-gripper designs created based on geometric analysis of the object. We illustrate the use of our method by demonstrating many examples of connectors and practical use cases. Yuki Koyama 0001, Shinjiro Sueda, Emma Steinhardt, Takeo Igarashi, Ariel Shamir, Wojciech Matusik |
ACM Trans. Graph. | 5 |
| 2015 | Fab forms: customizable objects for fabrication with validity and geometry cachingabstractWe address the problem of allowing casual users to customize parametric models while maintaining their valid state as 3D-printable functional objects. We define Fab Form as any design representation that lends itself to interactive customization by a novice user, while remaining valid and manufacturable. We propose a method to achieve these Fab Form requirements for general parametric designs tagged with a general set of automated validity tests and a small number of parameters exposed to the casual user. Our solution separates Fab Form evaluation into a precomputation stage and a runtime stage. Parts of the geometry and design validity (such as manufacturability) are evaluated and stored in the precomputation stage by adaptively sampling the design space. At runtime the remainder of the evaluation is performed. This allows interactive navigation in the valid regions of the design space using an automatically generated Web user interface (UI). We evaluate our approach by converting several parametric models into corresponding Fab Forms. Maria Shugrina, Ariel Shamir, Wojciech Matusik |
ACM Trans. Graph. | 2 |
| 2014 | Organizing heterogeneous scene collections through contextual focal pointsabstractWe introduce focal points for characterizing, comparing, and organizing collections of complex and heterogeneous data and apply the concepts and algorithms developed to collections of 3D indoor scenes. We represent each scene by a graph of its constituent objects and define focal points as representative substructures in a scene collection. To organize a heterogeneous scene collection, we cluster the scenes based on a set of extracted focal points: scenes in a cluster are closely connected when viewed from the perspective of the representative focal points of that cluster. The key concept of representativity requires that the focal points occur frequently in the cluster and that they result in a compact cluster. Hence, the problem of focal point extraction is intermixed with the problem of clustering groups of scenes based on their representative focal points. We present a co-analysis algorithm which interleaves frequent pattern mining and subspace clustering to extract a set of contextual focal points which guide the clustering of the scene collection. We demonstrate advantages of focal-centric scene comparison and organization over existing approaches, particularly in dealing with hybrid scenes, scenes consisting of elements which suggest membership in different semantic categories. Kai Xu 0004, Rui Ma 0011, Hao (Richard) Zhang, Chenyang Zhu 0002, Ariel Shamir, Daniel Cohen-Or, Hui Huang 0004 |
ACM Trans. Graph. | 5 |
| 2014 | Automatic editing of footage from multiple social camerasabstractWe present an approach that takes multiple videos captured by social cameras---cameras that are carried or worn by members of the group involved in an activity---and produces a coherent "cut" video of the activity. Footage from social cameras contains an intimate, personalized view that reflects the part of an event that was of importance to the camera operator (or wearer). We leverage the insight that social cameras share the focus of attention of the people carrying them. We use this insight to determine where the important "content" in a scene is taking place, and use it in conjunction with cinematographic guidelines to select which cameras to cut to and to determine the timing of those cuts. A trellis graph representation is used to optimize an objective function that maximizes coverage of the important content in the scene, while respecting cinematographic guidelines such as the 180-degree rule and avoiding jump cuts. We demonstrate cuts of the videos in various styles and lengths for a number of scenarios, including sports games, street performances, family activities, and social get-togethers. We evaluate our results through an in-depth analysis of the cuts in the resulting videos and through comparison with videos produced by a professional editor and existing commercial solutions. Ido Arev, Hyun Soo Park, Yaser Sheikh, Jessica K. Hodgins, Ariel Shamir |
ACM Trans. Graph. | 5 |
| 2014 | Boxelization: folding 3D objects into boxesabstractWe present a method for transforming a 3D object into a cube or a box using a continuous folding sequence. Our method produces a single, connected object that can be physically fabricated and folded from one shape to the other. We segment the object into voxels and search for a voxel-tree that can fold from the input shape to the target shape. This involves three major steps: finding a good voxelization, finding the tree structure that can form the input and target shapes' configurations, and finding a non-intersecting folding sequence. We demonstrate our results on several input 3D objects and also physically fabricate some using a 3D printer. Yahan Zhou, Shinjiro Sueda, Wojciech Matusik, Ariel Shamir |
ACM Trans. Graph. | 4 |
| 2014 | Design and fabrication by exampleabstractWe propose a data-driven method for designing 3D models that can be fabricated. First, our approach converts a collection of expert-created designs to a dataset of parameterized design templates that includes all information necessary for fabrication. The templates are then used in an interactive design system to create new fabri-cable models in a design-by-example manner. A simple interface allows novice users to choose template parts from the database, change their parameters, and combine them to create new models. Using the information in the template database, the system can automatically position, align, and connect parts: the system accomplishes this by adjusting parameters, adding appropriate constraints, and assigning connectors. This process ensures that the created models can be fabricated, saves the user from many tedious but necessary tasks, and makes it possible for non-experts to design and create actual physical objects. To demonstrate our data-driven method, we present several examples of complex functional objects that we designed and manufactured using our system. Adriana Schulz, Ariel Shamir, David I. W. Levin, Pitchaya Sitthi-amorn, Wojciech Matusik |
ACM Trans. Graph. | 2 |
| 2014 | Filling Your Shelves: Synthesizing Diverse Style-Preserving Artifact ArrangementsabstractOur homes and workspaces are filled with collections of dozens of artifacts laid out on surfaces such as shelves, counters, and mantles. The content and layout of these arrangements reflect both context, e.g., kitchen or living room, and style, e.g., neat or messy. Manually assembling such arrangements in virtual scenes is highly time consuming, especially when one needs to generate multiple diverse arrangements for numerous support surfaces and living spaces. We present a data-driven method especially designed for artifact arrangement which automatically populates empty surfaces with diverse believable arrangements of artifacts in a given style. The input to our method is an annotated photograph or a 3D model of an exemplar arrangement, that reflects the desired context and style. Our method leverages this exemplar to generate diverse arrangements reflecting the exemplar style for arbitrary furniture setups and layout dimensions. To simultaneously achieve scalability, diversity and style preservation, we define a valid solution space of arrangements that reflect the input style. We obtain solutions within this space using barrier functions and stochastic optimization. Lucas Majerowicz, Ariel Shamir, Alla Sheffer, Holger H. Hoos |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2013 | Curve Style Analysis in a Set of ShapesabstractAbstract The word ‘style’ can be interpreted in so many different ways in so many different contexts. To provide a general analysis and understanding of styles is a highly challenging problem. We pose the open question ‘how to extract styles from geometric shapes?’ and address one instance of the problem. Specifically, we present an unsupervised algorithm for identifying curve styles in a set of shapes. In our setting, a curve style is explicitly represented by a mode of curve features appearing along the 2D silhouettes of the shapes in the set. Unlike previous attempts, we do not rely on any preconceived conceptual characterisations, for example, via specific shape descriptors, to define what is or is not a style. Our definition of styles is data‐dependent; it depends on the input set but we do not require computing a shape correspondence across the set. We provide an operational definition of curve styles which focuses on separating curve features that represent styles from curve features that are content revealing. To this end, we develop a novel formulation and associated algorithm for style‐content separation. The analysis is based on a feature‐shape association matrix (FSM) whose rows correspond to modes of curve features, columns to shapes in the set, and each entry expresses the extent a feature mode is present in a shape. We make several assumptions to drive style‐content separation which only involve properties of, and relations between, rows of the FSM. Computationally, our algorithm only requires row‐wise correlation analysis in the FSM and a heuristic solution of an instance of the set cover problem. Results are demonstrated on several data sets showing the identification of curve styles. We also develop and demonstrate several style‐related applications including style exaggeration, removal, blending, and style transfer for 2D shape synthesis. Honghua Li, Hao (Richard) Zhang, Junjie Cao 0001, Ariel Shamir, Daniel Cohen-Or |
Comput. Graph. Forum | 5 |
| 2013 | Geosemantic Snapping for Sketch-Based ModelingabstractAbstract Modeling 3D objects from sketches is a process that requires several challenging problems including segmentation, recognition and reconstruction. Some of these tasks are harder for humans and some are harder for the machine. At the core of the problem lies the need for semantic understanding of the shape's geometry from the sketch. In this paper we propose a method to model 3D objects from sketches by utilizing humans specifically for semantic tasks that are very simple for humans and extremely difficult for the machine, while utilizing the machine for tasks that are harder for humans. The user assists recognition and segmentation by choosing and placing specific geometric primitives on the relevant parts of the sketch. The machine first snaps the primitive to the sketch by fitting its projection to the sketch lines, and then improves the model globally by inferringgeosemanticconstraints that link the different parts. The fitting occurs in real‐time, allowing the user to be only as precise as needed to have a good starting configuration for this non‐convex optimization problem. We evaluate the accessibility of our approach with a user study. Alex Shtof, Alexander Agathos, Yotam I. Gingold, Ariel Shamir, Daniel Cohen-Or |
Comput. Graph. Forum | 4 |
| 2013 | Motion-Aware Gradient Domain Video CompositionabstractFor images, gradient domain composition methods like Poisson blending offer practical solutions for uncertain object boundaries and differences in illumination conditions. However, adapting Poisson image blending to video presents new challenges due to the added temporal dimension. In video, the human eye is sensitive to small changes in blending boundaries across frames and slight differences in motions of the source patch and target video. We present a novel video blending approach that tackles these problems by merging the gradient of source and target videos and optimizing a consistent blending boundary based on a user-provided blending trimap for the source video. Our approach extends mean-value coordinates interpolation to support hybrid blending with a dynamic boundary while maintaining interactive performance. We also provide a user interface and source object positioning method that can efficiently deal with complex video sequences beyond the capabilities of alpha blending. Tao Chen 0015, Jun-Yan Zhu, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 3 |
| 2013 | Style and abstraction in portrait sketchingabstractWe use a data-driven approach to study both style and abstraction in sketching of a human face. We gather and analyze data from a number of artists as they sketch a human face from a reference photograph. To achieve different levels of abstraction in the sketches, decreasing time limits were imposed -- from four and a half minutes to fifteen seconds. We analyzed the data at two levels: strokes and geometric shape. In each, we create a model that captures both the style of the different artists and the process of abstraction. These models are then used for a portrait sketch synthesis application. Starting from a novel face photograph, we can synthesize a sketch in the various artistic styles and in different levels of abstraction. Itamar Berger, Ariel Shamir, Moshe Mahler, Elizabeth J. Carter, Jessica K. Hodgins |
ACM Trans. Graph. | 2 |
| 2013 | 3-Sweep: extracting editable objects from a single photoabstractWe introduce an interactive technique for manipulating simple 3D shapes based on extracting them from a single photograph. Such extraction requires understanding of the components of the shape, their projections, and relations. These simple cognitive tasks for humans are particularly difficult for automatic algorithms. Thus, our approach combines the cognitive abilities of humans with the computational accuracy of the machine to solve this problem. Our technique provides the user the means to quickly create editable 3D parts---human assistance implicitly segments a complex object into its components, and positions them in space. In our interface, three strokes are used to generate a 3D component that snaps to the shape's outline in the photograph, where each stroke defines one dimension of the component. The computer reshapes the component to fit the image of the object in the photograph as well as to satisfy various inferred geometric constraints imposed by its global 3D structure. We show that with this intelligent interactive modeling tool, the daunting task of object extraction is made simple. Once the 3D object has been extracted, it can be quickly edited and placed back into photos or 3D scenes, permitting object-driven photo editing tasks which are impossible to perform in image-space. We show several examples and present a user study illustrating the usefulness of our technique. Tao Chen 0015, Zhe Zhu, Ariel Shamir, Shi-Min Hu 0001, Daniel Cohen-Or |
ACM Trans. Graph. | 3 |
| 2013 | Qualitative organization of collections of shapes via quartet analysisabstractWe present a method for organizing a heterogeneous collection of 3D shapes for overview and exploration. Instead of relying on quantitative distances, which may become unreliable between dissimilar shapes, we introduce aqualitativeanalysis which utilizes multiple distance measures but only in cases where the measures can be reliably compared. Our analysis is based on the notion ofquartets, each defined by two pairs of shapes, where the shapes in each pair are close to each other, but far apart from the shapes in the other pair. Combining the information from many quartets computed across a shape collection using several distance measures, we create a hierarchical structure we callcategorization treeof the shape collection. This tree satisfies the topological (qualitative) constraints imposed by the quartets creating an effective organization of the shapes. We present categorization trees computed on various collections of shapes and compare them to ground truth data from human categorization. We further introduce the concept ofdegree of separationchart for every shape in the collection and show the effectiveness of using it for interactive shapes exploration. Shi-Sheng Huang, Ariel Shamir, Chao-Hui Shen, Hao (Richard) Zhang, Alla Sheffer, Shi-Min Hu 0001, Daniel Cohen-Or |
ACM Trans. Graph. | 2 |
| 2013 | Co-hierarchical analysis of shape structuresabstractWe introduce an unsupervised co-hierarchical analysis of a set of shapes, aimed at discovering their hierarchical part structures and revealing relations between geometrically dissimilar yet functionally equivalent shape parts across the set. The core problem is that of representative co-selection . For each shape in the set, one representative hierarchy (tree) is selected from among many possible interpretations of the hierarchical structure of the shape. Collectively, the selected tree representatives maximize the within-cluster structural similarity among them. We develop an iterative algorithm for representative co-selection. At each step, a novel cluster-and-select scheme is applied to a set of candidate trees for all the shapes. The tree-to-tree distance for clustering caters to structural shape analysis by focusing on spatial arrangement of shape parts, rather than their geometric details. The final set of representative trees are unified to form a structural co-hierarchy. We demonstrate co-hierarchical analysis on families of man-made shapes exhibiting high degrees of geometric and finer-scale structural variabilities. Oliver van Kaick, Kai Xu 0004, Hao (Richard) Zhang, Shuyang Sun, Ariel Shamir, Daniel Cohen-Or |
ACM Trans. Graph. | 6 |
| 2013 | Content-adaptive image downscalingabstractThis paper introduces a novel content-adaptive image downscaling method. The key idea is to optimize the shape and locations of the downsampling kernels to better align with local image features. Our content-adaptive kernels are formed as a bilateral combination of two Gaussian kernels defined over space and color, respectively. This yields a continuum ranging from smoothing to edge/detail preserving kernels driven by image content. We optimize these kernels to represent the input image well, by finding an output image from which the input can be well reconstructed. This is technically realized as an iterative maximum-likelihood optimization using a constrained variation of the Expectation-Maximization algorithm. In comparison to previous downscaling algorithms, our results remain crisper without suffering from ringing artifacts. Besides natural images, our algorithm is also effective for creating pixel art images from vector graphics inputs, due to its ability to keep linear features sharp and connected. Johannes Kopf 0001, Ariel Shamir, Pieter Peers |
ACM Trans. Graph. | 2 |
| 2013 | PoseShop: Human Image Database Construction and Personalized Content SynthesisabstractWe present PoseShop--a pipeline to construct segmented human image database with minimal manual intervention. By downloading, analyzing, and filtering massive amounts of human images from the Internet, we achieve a database which contains 400 thousands human figures that are segmented out of their background. The human figures are organized based on action semantic, clothes attributes, and indexed by the shape of their poses. They can be queried using either silhouette sketch or a skeleton to find a given pose. We demonstrate applications for this database for multiframe personalized content synthesis in the form of comic-strips, where the main character is the user or his/her friends. We address the two challenges of such synthesis, namely personalization and consistency over a set of frames, by introducing head swapping and clothes swapping techniques. We also demonstrate an action correlation analysis application to show the usefulness of the database for vision application. Tao Chen 0015, Li-Qian Ma, Ming-Ming Cheng, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2012 | Interactive visual queries for multivariate graphs exploration
Ariel Shamir, Alla Stolpnik |
Comput. Graph. | 1 |
| 2012 | Data-Driven Object Manipulation in ImagesabstractAbstract We present a framework for interactively manipulating objects in a photograph using related objects obtained from internet images. Given an image, the user selects an object to modify, and provides keywords to describe it. Objects with a similar shape are retrieved and segmented from online images matching the keywords, and deformed to correspond with the selected object. By matching the candidate object and adjusting manipulation parameters, our method appropriately modifies candidate objects and composites them into the scene. Supported manipulations include transferring texture, color and shape from the matched object to the target in a seamless manner. We demonstrate the versatility of our framework using several inputs of varying complexity, for object completion, augmentation, replacement and revealing. Our results are evaluated using a user study. Chen Goldberg, Tao Chen 0015, Ariel Shamir, Shi-Min Hu 0001 |
Comput. Graph. Forum | 4 |
| 2012 | Micro perceptual human computation for visual tasksabstractHuman Computation (HC) utilizes humans to solve problems or carry out tasks that are hard for pure computational algorithms. Many graphics and vision problems have such tasks. Previous HC approaches mainly focus on generating data in batch, to gather benchmarks, or perform surveys demanding nontrivial interactions. We advocate a tighter integration of human computation into online, interactive algorithms. We aim to distill the differences between humans and computers and maximize the advantages of both in one algorithm. Our key idea is to decompose such a problem into a massive number of very simple, carefully designed, human micro-tasks that are based on perception , and whose answers can be combined algorithmically to solve the original problem. Our approach is inspired by previous work on micro-tasks and perception experiments. We present three specific examples for the design of micro perceptual human computation algorithms to extract depth layers and image normals from a single photograph, and to augment an image with high-level semantic information such as symmetry. Yotam I. Gingold, Ariel Shamir, Daniel Cohen-Or |
ACM Trans. Graph. | 2 |
| 2012 | StackabilizationabstractWe introduce the geometric problem of stackabilization : how to geometrically modify a 3D object so that it is more amenable to stacking. Given a 3D object and a stacking direction, we define a measure of stackability, which is derived from the gap between the lower and upper envelopes of the object in a stacking configuration along the stacking direction. The main challenge in stackabilization lies in the desire to modify the object's geometry only subtly so that the intended functionality and aesthetic appearance of the original object are not significantly affected. We present an automatic algorithm to deform a 3D object to meet a target stackability score using energy minimization. The optimized energy accounts for both the scales of the deformation parameters as well as the preservation of pre-existing geometric and structural properties in the object, e. g., symmetry, as a means of maintaining its functionality. We also present an intelligent editing tool that assists a modeler when modifying a given 3D object to improve its stackability. Finally, we explore a few fun variations of the stackabilization problem. Honghua Li, Ibraheem Alhashim, Hao (Richard) Zhang, Ariel Shamir, Daniel Cohen-Or |
ACM Trans. Graph. | 4 |
| 2012 | Combining color and depth for enhanced image segmentation and retargeting
Meir Johnathan Dahan, Nir Chen, Ariel Shamir, Daniel Cohen-Or |
Vis. Comput. | 3 |
| 2011 | Symmetry Hierarchy of Man-Made ObjectsabstractAbstract We introduce symmetry hierarchy of man‐made objects, a high‐level structural representation of a 3D model providing a symmetry‐induced, hierarchical organization of the model's constituent parts. Given an input mesh, we segment it into primitive parts and build an initial graph which encodes inter‐part symmetries and connectivity relations, as well as self‐symmetries in individual parts. The symmetry hierarchy is constructed from the initial graph via recursive graph contraction which either groups parts by symmetry or assembles connected sets of parts. The order of graph contraction is dictated by a set of precedence rules designed primarily to respect the law of symmetry in perceptual grouping and the principle of compactness of representation. We show that symmetry hierarchy naturally implies a hierarchical segmentation that is more meaningful than those produced by local geometric considerations. We also develop an application of symmetry hierarchies for structural shape editing. Kai Xu 0004, Jun Li 0042, Hao (Richard) Zhang, Ariel Shamir, Ligang Liu 0001, Zhi-Quan Cheng, Yueshan Xiong |
Comput. Graph. Forum | 5 |
| 2011 | Digital micrographyabstractWe present an algorithm for creating digital micrography images, or micrograms , a special type of calligrams created from minuscule text. These attractive text-art works successfully combine beautiful images with readable meaningful text. Traditional micrograms are created by highly skilled artists and involve a huge amount of tedious manual work. We aim to simplify this process by providing a computerized digital micrography design tool. The main challenge in creating digital micrograms is designing textual layouts that simultaneously convey the input image, are readable and appealing. To generate such layout we use the streamlines of singularity free, low curvature, smooth vector fields, especially designed for our needs. The vector fields are computed using a new approach which controls field properties via a priori boundary condition design that balances the different requirements we aim to satisfy. The optimal boundary conditions are computed using a graph-cut approach balancing local and global design considerations. The generated layouts are further processed to obtain the final micrograms. Our method automatically generates engaging, readable micrograms starting from a vector image and an input text while providing a variety of optional high-level controls to the user. Ron Maharik, Mikhail Bessmeltsev, Alla Sheffer, Ariel Shamir, Nathan Carr 0001 |
ACM Trans. Graph. | 4 |
| 2010 | Motion fields to predict play evolution in dynamic sport scenesabstractVideos of multi-player team sports provide a challenging domain for dynamic scene analysis. Player actions and interactions are complex as they are driven by many factors, such as the short-term goals of the individual player, the overall team strategy, the rules of the sport, and the current context of the game. We show that constrained multi-agent events can be analyzed and even predicted from video. Such analysis requires estimating the global movements of all players in the scene at any time, and is needed for modeling and predicting how the multi-agent play evolves over time on the field. To this end, we propose a novel approach to detect the locations of where the play evolution will proceed, e.g. where interesting events will occur, by tracking player positions and movements over time. We start by extracting the ground level sparse movement of players in each time-step, and then generate a dense motion field. Using this field we detect locations where the motion converges, implying positions towards which the play is evolving. We evaluate our approach by analyzing videos of a variety of complex soccer plays. Matthias Grundmann 0002, Ariel Shamir, Iain A. Matthews, Jessica K. Hodgins, Irfan A. Essa |
CVPR | 3 |
| 2010 | Context-Dependent Crowd EvaluationabstractAbstract Many times, even if a crowd simulation looks good in general, there could be some specific individual behaviors which do not seem correct. Spotting such problems manually can become tedious, but ignoring them may harm the simulation's credibility. In this paper we present a data‐driven approach for evaluating the behaviors of individuals within a simulated crowd. Based on video‐footage of a real crowd, a database of behavior examples is generated. Given a simulation of a crowd, an analog analysis is performed on it, defining a set of queries, which are matched by a similarity function to the database examples. The results offer a possible objective answer to the question of how similar are the simulated individual behaviors to real observed behaviors. Moreover, by changing the video input one can change the context of evaluation. We show several examples of evaluating simulated crowds produced using different techniques and comprising of dense crowds, sparse crowds and flocks. Alon Lerner, Yiorgos Chrysanthou, Ariel Shamir, Daniel Cohen-Or |
Comput. Graph. Forum | 3 |
| 2010 | Contextual Part Analogies in 3D Objects
Lior Shapira, Shy Shalom, Ariel Shamir, Daniel Cohen-Or, Hao (Richard) Zhang |
Int. J. Comput. Vis. | 3 |
| 2010 | A comparative study of image retargetingabstractThe numerous works on media retargeting call for a methodological approach for evaluating retargeting results. We present the first comprehensive perceptual study and analysis of image retargeting. First, we create a benchmark of images and conduct a large scale user study to compare a representative number of state-of-the-art retargeting methods. Second, we present analysis of the users' responses, where we find that humans in general agree on the evaluation of the results and show that some retargeting methods are consistently more favorable than others. Third, we examine whether computational image distance metrics can predict human retargeting perception. We show that current measures used in this context are not necessarily consistent with human rankings, and demonstrate that better results can be achieved using image features that were not previously considered for this task. We also reveal specific qualities in retargeted media that are more important for viewers. The importance of our work lies in promoting better measures to assess and guide retargeting algorithms in the future. The full benchmark we collected, including all images, retargeted results, and the collected user data, are available to the research community for further investigation at http://people.csail.mit.edu/mrub/retargetme. Michael Rubinstein, Diego Gutierrez, Olga Sorkine-Hornung, Ariel Shamir |
ACM Trans. Graph. | 4 |
| 2010 | Cone carving for surface reconstructionabstractWe present cone carving, a novel space carving technique supporting topologically correct surface reconstruction from an incomplete scanned point cloud. The technique utilizes the point samples not only for local surface position estimation but also to obtain global visibility information under the assumption that each acquired point is visible from a point lying outside the shape. This enables associating each point with a generalized cone, called the visibility cone , that carves a portion of the outside ambient space of the shape from the inside out. These cones collectively provide a means to better approximate the signed distances to the shape specifically near regions containing large holes in the scan, allowing one to infer the correct surface topology. Combining the new distance measure with conventional RBF, we define an implicit function whose zero level set defines the surface of the shape. We demonstrate the utility of cone carving in coping with significant missing data and raw scans from a commercial 3D scanner as well as synthetic input. Shy Shalom, Ariel Shamir, Hao (Richard) Zhang, Daniel Cohen-Or |
ACM Trans. Graph. | 2 |
| 2009 | Mode-detection via median-shiftabstractMedian-shift is a mode seeking algorithm that relies on computing the median of local neighborhoods, instead of the mean. We further combine median-shift with Locality Sensitive Hashing (LSH) and show that the combined algorithm is suitable for clustering large scale, high dimensional data sets. In particular, we propose a new mode detection step that greatly accelerates performance. In the past, LSH was used in conjunction with mean shift only to accelerate nearest neighbor queries. Here we show that we can analyze the density of the LSH bins to quickly detect potential mode candidates and use only them to initialize the median-shift procedure. We use the median, instead of the mean (or its discrete counterpart - the medoid) because the median is more robust and because the median of a set is a point in the set. A median is well defined for scalars but there is no single agreed upon extension of the median to high dimensional data. We adopt a particular extension, known as the Tukey median, and show that it can be computed efficiently using random projections of the high dimensional data onto 1D lines, just like LSH, leading to a tightly integrated and efficient algorithm. Lior Shapira, Shai Avidan, Ariel Shamir |
ICCV | 3 |
| 2009 | A Part-aware Surface Metric for Shape AnalysisabstractAbstract The notion of parts in a shape plays an important role in many geometry problems, including segmentation, correspondence, recognition, editing, and animation. As the fundamental geometric representation of 3D objects in computer graphics is surface‐based, solutions of many such problems utilize a surface metric, a distance function defined over pairs of points on the surface, to assist shape analysis and understanding. The main contribution of our work is to bring together these two fundamental concepts: shape parts and surface metric. Specifically, we develop a surface metric that is part‐aware. To encode part information at a point on a shape, we model its volumetric context – called the volumetric shape image (VSI) – inside the shape's enclosed volume, to capture relevant visibility information. We then define the part‐aware metric by combining an appropriate VSI distance with geodesic distance and normal variation. We show how the volumetric view on part separation addresses certain limitations of the surface view, which relies on concavity measures over a surface as implied by the well‐known minima rule. We demonstrate how the new metric can be effectively utilized in various applications including mesh segmentation, shape registration, part‐aware sampling and shape retrieval. Hao (Richard) Zhang, Ariel Shamir, Daniel Cohen-Or |
Comput. Graph. Forum | 3 |
| 2009 | Image Appearance Exploration by Model-Based NavigationabstractAbstract Changing the appearance of an image can be a complex and non‐intuitive task. Many times the target image colors and look are only known vaguely and many trials are needed to reach the desired results. Moreover, the effect of a specific change on an image is difficult to envision, since one must take into account spatial image considerations along with the color constraints. Tools provided today by image processing applications can become highly technical and non‐intuitive including various gauges and knobs. In this paper we introduce a method for changing image appearance by navigation, focusing on recoloring images. The user visually navigates a high dimensional space of possible color manipulations of an image. He can either explore in it for inspiration or refine his choices by navigating into sub regions of this space to a specific goal. This navigation is enabled by modeling the chroma channels of an image's colors using a Gaussian Mixture Model (GMM). The Gaussians model both color and spatial image coordinates, and provide a high dimensional parameterization space of a rich variety of color manipulations. The user's actions are translated into transformations of the parameters of the model, which recolor the image. This approach provides both inspiration and intuitive navigation in the complex space of image color manipulations. Lior Shapira, Ariel Shamir, Daniel Cohen-Or |
Comput. Graph. Forum | 2 |
| 2009 | Sketch2Photo: internet image montageabstractWe present a system that composes a realistic picture from a simple freehand sketch annotated with text labels. The composed picture is generated by seamlessly stitching several photographs in agreement with the sketch and text labels; these are found by searching the Internet. Although online image search generates many inappropriate results, our system is able to automatically select suitable photographs to generate a high quality composition, using a filtering scheme to exclude undesirable images. We also provide a novel image blending algorithm to allow seamless image composition. Each blending result is given a numeric score, allowing us to find an optimal combination of discovered images. Experimental results show the method is very successful; we also evaluate our system using the results from two user studies. Tao Chen 0015, Ming-Ming Cheng, Ariel Shamir, Shi-Min Hu 0001 |
ACM Trans. Graph. | 4 |
| 2009 | Multi-operator media retargetingabstractContent aware resizing gained popularity lately and users can now choose from a battery of methods to retarget their media. However, no single retargeting operator performs well on all images and all target sizes. In a user study we conducted, we found that users prefer to combine seam carving with cropping and scaling to produce results they are satisfied with. This inspires us to propose an algorithm that combines different operators in an optimal manner. We define a resizing space as a conceptual multi-dimensional space combining several resizing operators, and show how a path in this space defines a sequence of operations to retarget media. We define a new image similarity measure, which we term Bi-Directional Warping (BDW), and use it with a dynamic programming algorithm to find an optimal path in the resizing space. In addition, we show a simple and intuitive user interface allowing users to explore the resizing space of various image sizes interactively. Using key-frames and interpolation we also extend our technique to retarget video, providing the flexibility to use the best combination of operators at different times in the sequence. Michael Rubinstein, Ariel Shamir, Shai Avidan |
ACM Trans. Graph. | 2 |
| 2009 | Relief analysis and extractionabstractWe present an approach for extracting reliefs and details from relief surfaces. We consider a relief surface as a surface composed of two components: a base surface and a height function which is defined over this base. However, since the base surface is unknown, the decoupling of these components is a challenge. We show how to estimate a robust height function over the base, without explicitly extracting the base surface. This height function is utilized to separate the relief from the base. Several applications benefiting from this extraction are demonstrated, including relief segmentation, detail exaggeration and dampening, copying of details from one object to another, and curve drawing on meshes. Rony Zatzarinni, Ayellet Tal, Ariel Shamir |
ACM Trans. Graph. | 3 |
| 2008 | A survey on Mesh Segmentation TechniquesabstractAbstract We present a review of the state of the art of segmentation and partitioning techniques of boundary meshes. Recently, these have become a part of many mesh and object manipulation algorithms in computer graphics, geometric modelling and computer aided design. We formulate the segmentation problem as an optimization problem and identify two primarily distinct types of mesh segmentation, namelypartsegmentation andsurface‐patchsegmentation. We classify previous segmentation solutions according to the different segmentation goals, the optimization criteria and features used, and the various algorithmic techniques employed. We also present some generic algorithms for the major segmentation techniques. Ariel Shamir |
Comput. Graph. Forum | 1 |
| 2008 | Non-homogeneous resizing of complex modelsabstractResizing of 3D models can be very useful when creating new models or placing models inside different scenes. However, uniform scaling is limited in its applicability while straightforward non-uniform scaling can destroy features and lead to serious visual artifacts. Our goal is to define a method that protects model features and structures during resizing. We observe that typically, during scaling some parts of the models are more vulnerable than others, undergoing undesirable deformation. We automatically detect vulnerable regions and carry this information to a protective grid defined around the object, defining a vulnerability map. The 3D model is then resized by a space-deformation technique which scales the grid non-homogeneously while respecting this map. Using space-deformation allows processing of common models of man-made objects that consist of multiple components and contain non-manifold structures. We show that our technique resizes models while suppressing undesirable distortion, creating models that preserve the structure and features of the original ones. Vladislav Kraevoy, Alla Sheffer, Ariel Shamir, Daniel Cohen-Or |
ACM Trans. Graph. | 3 |
| 2008 | Improved seam carving for video retargetingabstractVideo, like images, should support content aware resizing. We present video retargeting using an improved seam carving operator. Instead of removing 1D seams from 2D images we remove 2D seam manifolds from 3D space-time volumes. To achieve this we replace the dynamic programming method of seam carving with graph cuts that are suitable for 3D volumes. In the new formulation, a seam is given by a minimal cut in the graph and we show how to construct a graph such that the resulting cut is a valid seam. That is, the cut is monotonic and connected. In addition, we present a novel energy criterion that improves the visual quality of the retargeted images and videos. The original seam carving operator is focused on removing seams with the least amount of energy, ignoring energy that is introduced into the images and video by applying the operator. To counter this, the new criterion is looking forward in time - removing seams that introduce the least amount of energy into the retargeted result. We show how to encode the improved criterion into graph cuts (for images and video) as well as dynamic programming (for images). We apply our technique to images and videos and present results of various applications. Michael Rubinstein, Ariel Shamir, Shai Avidan |
ACM Trans. Graph. | 2 |
| 2008 | Consistent mesh partitioning and skeletonisation using the shape diameter function
Lior Shapira, Ariel Shamir, Daniel Cohen-Or |
Vis. Comput. | 2 |
| 2007 | Surface reconstruction using local shape priors
Ran Gal, Ariel Shamir, Tal Hassner, Mark Pauly, Daniel Cohen-Or |
Symposium on Geometry Processing | 2 |
| 2007 | On-the-fly Curve-skeleton Computation for 3D ShapesabstractAbstract The curve‐skeleton of a 3D object is an abstract geometrical and topological representation of its 3D shape. It maps the spatial relation of geometrically meaningful parts to a graph structure. Each arc of this graph represents a part of the object with roughly constant diameter or thickness, and approximates its centerline. This makes the curve‐skeleton suitable to describe and handle articulated objects such as characters for animation. We present an algorithm to extract such a skeleton on‐the‐fly, both from point clouds and polygonal meshes. The algorithm is based on a deformable model evolution that captures the object's volumetric shape. The deformable model involves multiple competing fronts which evolve inside the object in a coarse‐to‐fine manner. We first track these fronts' centers, and then merge and filter the resulting arcs to obtain a curve‐skeleton of the object. The process inherits the robustness of the reconstruction technique, being able to cope with noisy input, intricate geometry and complex topology. It creates a natural segmentation of the object and computes a center curve for each segment while maintaining a full correspondence between the skeleton and the boundary of the object. Andrei Sharf, Thomas Lewiner, Ariel Shamir, Leif Kobbelt |
Comput. Graph. Forum | 3 |
| 2007 | Seam carving for content-aware image resizingabstractEffective resizing of images should not only use geometric constraints, but consider the image content as well. We present a simple image operator called seam carving that supports content-aware image resizing for both reduction and expansion. A seam is an optimal 8-connected path of pixels on a single image from top to bottom, or left to right, where optimality is defined by an image energy function. By repeatedly carving out or inserting seams in one direction we can change the aspect ratio of an image. By applying these operators in both directions we can retarget the image to a new size. The selection and order of seams protect the content of the image, as defined by the energy function. Seam carving can also be used for image content enhancement and object removal. We support various visual saliency measures for defining the energy of an image, and can also include user input to guide the process. By storing the order of seams in an image we create multi-size images, that are able to continuously change in real time to fit a given size. Shai Avidan, Ariel Shamir |
ACM Trans. Graph. | 2 |
| 2007 | Pose-Oblivious Shape SignatureabstractA 3D shape signature is a compact representation for some essence of a shape. Shape signatures are commonly utilized as a fast indexing mechanism for shape retrieval. Effective shape signatures capture some global geometric properties which are scale, translation, and rotation invariant. In this paper, we introduce an effective shape signature which is also pose-oblivious. This means that the signature is also insensitive to transformations which change the pose of a 3D shape such as skeletal articulations. Although some topology-based matching methods can be considered pose-oblivious as well, our new signature retains the simplicity and speed of signature indexing. Moreover, contrary to topology-based methods, the new signature is also insensitive to the topology change of the shape, allowing us to match similar shapes with different genus. Our shape signature is a 2D histogram which is a combination of the distribution of two scalar functions defined on the boundary surface of the 3D shape. The first is a definition of a novel function called the local-diameter function. This function measures the diameter of the 3D shape in the neighborhood of each vertex. The histogram of this function is an informative measure of the shape which is insensitive to pose changes. The second is the centricity function that measures the average geodesic distance from one vertex to all other vertices on the mesh. We evaluate and compare a number of methods for measuring the similarity between two signatures, and demonstrate the effectiveness of our pose-oblivious shape signature within a 3D search engine application for different databases containing hundreds of models. Ran Gal, Ariel Shamir, Daniel Cohen-Or |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2006 | Editorial
Ayellet Tal, Thomas A. Funkhouser, Ariel Shamir |
Comput. Graph. | 3 |
| 2006 | Competing Fronts for Coarse-to-Fine Surface ReconstructionabstractAbstract We present a deformable model to reconstruct a surface from a point cloud. The model is based on an explicit mesh representation composed of multiple competing evolving fronts. These fronts adapt to the local feature size of the target shape in a coarse–to–fine manner. Hence, they approach towards the finer (local) features of the target shape only after the reconstruction of the coarse (global) features has been completed. This conservative approach leads to a better control and interpretation of the reconstructed topology. The use of an explicit representation for the deformable model guarantees water‐tightness and simple tracking of topological events. Furthermore, the coarse–to–fine nature of reconstruction enables adaptive handling of non‐homogenous sample density, including robustness to missing data in defected areas. Categories and Subject Descriptors (according to ACM CCS): I.3.3 [Computer Graphics]: Digitizing and scanning. Keywords: surface reconstruction, deformable models Andrei Sharf, Thomas Lewiner, Ariel Shamir, Leif Kobbelt, Daniel Cohen-Or |
Comput. Graph. Forum | 3 |
| 2006 | Skeleton based solid representation with topology preservation
Ariel Shamir, Amir Shaham |
Graph. Model. | 1 |
| 2006 | Mesh analysis using geodesic mean-shift
Ariel Shamir, Lior Shapira, Daniel Cohen-Or |
Vis. Comput. | 1 |
| 2006 | SnapPaste: an interactive technique for easy mesh composition
Andrei Sharf, Marina Blumenkrants, Ariel Shamir, Daniel Cohen-Or |
Vis. Comput. | 3 |
| 2005 | Mesh scissoring with minima rule and part salience
Yunjin Lee, Seungyong Lee 0001, Ariel Shamir, Daniel Cohen-Or, Hans-Peter Seidel |
Comput. Aided Geom. Des. | 3 |
| 2004 | Feature-Sensitive 3D Shape MatchingabstractThree dimensional shape matching plays an important role in many of today's applications. Nevertheless, shape matching is a difficult problem since there is no unique measure that defines shape similarity and since computing shape distance using various measures is an elaborate task. We present a new framework for matching shapes, represented by union of spheres hierarchies. Our approach is "feature-sensitive" since it extends the usual geometry-based matching by adding sensitivity to shape topology, shape features (sharp angles, chemical attributes) and their relative positioning on the shape. Our method can be used in conjecture with other geometric matching methods as a pre or post-processing filtering stage, or it can be used as a stand-alone feature-sensitive matching. Andrei Sharf, Ariel Shamir |
Computer Graphics International | 2 |
| 2004 | Intelligent Mesh Scissoring Using 3D SnakesabstractMesh partitioning and parts extraction have become key ingredients for many mesh manipulation applications both manual and automatic. In this paper, we present an intelligent scissoring operator for meshes which supports both automatic segmentation and manual cutting. Instead of segmenting the mesh by clustering, our approach concentrates on finding and defining the contours for cutting. This approach is based on the minima rule, which states that human perception usually divides a surface into parts along the contours of concave discontinuity of the tangent plane. The technique uses feature extraction to find such candidate feature contours. Subsequently, such a contour can be selected either automatically or manually, or the user may draw a 2D line to start the scissoring process. The given open contour is completed to form a loop around a specific part of the mesh, and this loop is used as the initial position of a 3D geometric snake. The snake moves by relaxation until it settles to define the final scissoring position. This process uses several fundamental geometric mesh attributes, such as curvature and centricity, and enables both automatic segmentation and an easy-to-use intelligent-scissoring operator. Yunjin Lee, Seungyong Lee 0001, Ariel Shamir, Daniel Cohen-Or, Hans-Peter Seidel |
PG | 3 |
| 2004 | Automated Creation of Movie Summaries in Interactive Virtual Environments
Doron Friedman, Yishai A. Feldman, Ariel Shamir, Tsvi Dagan |
VR | 3 |
| 2004 | Colorplate: Automated Creation of Movie Summaries in Interactive Virtual Environments
Doron Friedman, Yishai A. Feldman, Ariel Shamir, Tsvi Dagan |
VR | 3 |
| 2004 | Special Section on the Fourth Israel-Korea Bi-National Conference on Geometric Modeling and Computer Graphics
Ariel Shamir, Seungyong Lee 0001 |
Vis. Comput. | 1 |
| 2003 | Feature Space Analysis of Unstructured MeshesabstractUnstructured meshes are often used in simulations and imaging applications. They provide advanced flexibility in modeling abilities but are more difficult to manipulate and analyze than regular data. This work provides a novel approach for the analysis of unstructured meshes using feature-space clustering and feature-detection. Analyzing and revealing underlying structures in data involve operators on both spatial and functional domains. Slicing concentrates more on the spatial domain, while iso-surfacing or volume rendering concentrate more on the functional domain. Nevertheless, many times it is the combination of the two domains which provides real insight on the structure of the data. In this work, a combined feature-space is defined on top of unstructured meshes in order to search for structure in the data. A point in feature-space includes the spatial coordinates of the point in the mesh domain and all chosen attributes defined on the mesh. A distance measures between points in feature-space is defined enabling the utilization of clustering using the mean shift procedure (previously used for images) on unstructured meshes. Feature space analysis is shown to be useful for feature-extraction, for data exploration and partitioning. Ariel Shamir |
IEEE Visualization | 1 |
| 2003 | Dynamic maintenance and visualization of molecular surfaces
Chandrajit L. Bajaj, Valerio Pascucci, Ariel Shamir, Robert J. Holt, Arun N. Netravali |
Discret. Appl. Math. | 3 |
| 2003 | Constraint-based approach for automatic hinting of digital typefacesabstractThe rasterization process of characters from digital outline fonts to bitmaps on displays must include additional information in the form of hints beside the shape of characters in order to produce high quality bitmaps. Hints describe constraints on sizes and shapes inside characters and across the font that should be preserved during rasterization. We describe a novel, fast and fully automatic method for adding those hints to characters. The method is based on identifying hinting situations inside characters. It includes gathering global font information and linking it to characters, defining a set of constraints, sorting them, and converting them to hints in any known hinting technology (PostScript, TrueType or other). Our scheme is general enough to be applied on any language and on complex scripts such as Chinese Japanese and Korean. Although still inferior to expert manual hinting, our method produces high quality bitmaps which approach this goal. The method can also be used as a solid base for further hinting refinements done manually. Ariel Shamir |
ACM Trans. Graph. | 1 |
| 2001 | Temporal and spatial level of details for dynamic meshesabstractMulti-resolution techniques enhance the ability of graphics and visual systems to overcome limitations in time, space and transmission costs. Numerous techniques have been presented which concentrate on creating level of detail models for static meshes. Time-dependent deformable meshes impose even greater difficulties on such systems. In this paper we describe a solution for using level of details for time dependent meshes. Our solution allows for both temporal and spatial level of details to be combined in an efficient manner. By separating low and high frequency temporal information, we gain the ability to create very fast coarse updates in the temporal dimension, which can be adaptively refined for greater details. Ariel Shamir, Valerio Pascucci |
VRST | 1 |
| 2000 | Multi-resolution dynamic meshes with arbitrary deformationsabstractMulti-resolution techniques and models have been shown to be effective for the display and transmission of large static geometric object. Dynamic environments with internally deforming models and scientific simulations using dynamic meshes pose greater challenges in terms of time and space, and need the development of similar solutions. We introduce the T-DAG, an adaptive multi-resolution representation for dynamic meshes with arbitrary deformations including attribute, position, connectivity and topology changes. T-DAG stands for time-dependent directed acyclic graph which defines the structure supporting this representation. We also provide an incremental algorithm (in time) for constructing the T-DAG representation of a given input mesh. This enables the traversal and use of the multi-resolution dynamic model for partial playback while still constructing new time-steps. Ariel Shamir, Chandrajit L. Bajaj, Valerio Pascucci |
IEEE Visualization | 1 |
| 1999 | Compacting oriental fonts by optimizing parametric elements
Ariel Shamir, Ari Rappoport |
Vis. Comput. | 1 |
| 1997 | Quality enhancements of digital outline fonts
Ariel Shamir, Ari Rappoport |
Comput. Graph. | 1 |
| 1996 | Extraction of Typographic Elements from Outline Representations of FontsabstractAbstract Digital typefaces for computer graphics and multimedia applications should be capable of supporting operations such as font variations, transformations. deformations and blending. A powerful implementation of such operations must rely on the inherent typographic attributes of the typeface. However, even today's most advanced typeface representations support only geometric outline representations and basic font variations. In this paper we discuss high‐level typeface representations which we term Parametric Typographic Representations (PTRs). We present an algorithm for automatically extracting typographic elements of typefaces from their outline representation, which, is an essential initial step in converting typefaces from outline representations to PTRs. The extracted typographic elements include serifs, bars. sterns, slants, bows, arcs, curve stems and curve bars. Most notable is the treatment of serifs, which are represented by finite‐automata. The algorithm only needs to learn a serif type once, and is then capable of automatically recognizing it in different typefaces. We show an application of a PTR for automatic high‐quality hinting of fonts, which is one of the most important stages in, digital font production. Our system was used to generate hints for dozens of thousands of Kanji, Roman and Hebrew characters. Ariel Shamir, Ari Rappoport |
Comput. Graph. Forum | 1 |