VLDB 2026 Research / reviewers in the wild / expert
Jose Echevarria
dblp:245/4272
· DBLP profile ↗
23ranked-venue papers
0as first author
15since 2021 · last 2025
0000-0001-6802-0911ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 13 since 2021Artificial intelligence and machine learning · 10 · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MaDCoW: Marginal Distortion Correction for Wide-Angle Photography with Arbitrary ObjectsabstractWe introduce MaDCoW, a method for correcting marginal distortion of arbitrary objects in wide-angle photography. People often use wide-angle photography—it is the default in smartphone cameras—but very-wide-fields-of-view produce distorted object appearance in image margins. In our system, a user annotates straight lines and regions of interest. MaDCoW solves for a separate linear perspective projection for each region and then jointly solves for a distortion-minimizing projection for the whole photograph. We show that MaDCoW can produce good results in cases where previous methods yield visible distortions. Kevin Zhang 0003, Jia-Bin Huang 0001, Jose Echevarria, Stephen DiVerdi, Aaron Hertzmann |
CVPR | 3 |
| 2025 | Palette-Based Color HarmonizationabstractWe present a palette-based framework for color composition for visual applications and three large-scale, wide-ranging perceptual studies on the perception of color harmonization. We abstract relationships between palette colors as a compact set of axes describing harmonic templates over perceptually uniform color wheels. Our framework provides a basis for interactive color-aware operations such as color harmonization of images and videos. Because our approach to harmonization is palette-based, we are able to conduct the first controlled perceptual experiments evaluating preferences for harmonized images and color palettes. In a third study, we compare preference for archetypical harmonic palettes. In total, our studies involved over 1000 participants. We found that participants do not prefer harmonized images and that some archetypal palettes are reliably viewed as less harmonious than random palettes. These studies raise important questions for research and artistic practice. Jianchao Tan, Jose Echevarria, Yotam I. Gingold |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Palette-Based Recolouring of Gradient MeshesabstractAbstract Gradient meshes are a vector graphics primitive formed by a regular grid of bicubic quad patches. They allow for the creation of complex geometries and colour gradients, with recent extensions supporting features such as local refinement and sharp colour transitions. While many methods exist for recolouring raster images, often achieved by modifying an automatically detected palette of the image, gradient meshes have not received the same amount of attention when it comes to global colour editing. We present a novel method that allows for real‐time palette‐based recolouring of gradient meshes, including gradient meshes constructed using local refinement and containing sharp colour transitions. We demonstrate the utility of our method on synthetic illustrative examples as well as on complex gradient meshes. Willard A. Verschoore de la Houssaije, Jose Echevarria, Jirí Kosinka |
Comput. Graph. Forum | 2 |
| 2024 | LF2MV: Learning an Editable Meta-View Towards Light Field RepresentationabstractLight fields are 4D scene representations that are typically structured as arrays of views or several directional samples per pixel in a single view. However, this highly correlated structure is not very efficient to transmit and manipulate, especially for editing. To tackle this issue, we propose a novel representation learning framework that can encode the light field into a single meta-view that is both compact and editable. Specifically, the meta-view composes of three visual channels and a complementary meta channel that is embedded with geometric and residual appearance information. The visual channels can be edited using existing 2D image editing tools, before reconstructing the whole edited light field. To facilitate edit propagation against occlusion, we design a special editing-aware decoding network that consistently propagates the visual edits to the whole light field upon reconstruction. Extensive experiments show that our proposed method achieves competitive representation accuracy and meanwhile enables consistent edit propagation. Menghan Xia, Jose Echevarria, Minshan Xie, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | SVGformer: Representation Learning for Continuous Vector Graphics using TransformersabstractAdvances in representation learning have led to great success in understanding and generating data in various domains. However, in modeling vector graphics data, the pure data-driven approach often yields unsatisfactory results in downstream tasks as existing deep learning methods often require the quantization of SVG parameters and cannot exploit the geometric properties explicitly. In this paper, we propose a transformer-based representation learning model (SVG-former) that directly operates on continuous input values and manipulates the geometric information of SVG to encode outline details and long-distance dependencies. SVGfomer can be used for various downstream tasks: reconstruction, classification, interpolation, retrieval, etc. We have conducted extensive experiments on vector font and icon datasets to show that our model can capture high-quality representation information and outperform the previous state-of-the-art on downstream tasks significantly. Defu Cao, Jose Echevarria, Yan Liu 0002 |
CVPR | 3 |
| 2023 | Cross-modal Latent Space Alignment for Image to Avatar TranslationabstractWe present a novel method for automatic vectorized avatar generation from a single portrait image. Most existing approaches that create avatars rely on image-to-image translation methods, which present some limitations when applied to 3D rendering, animation, or video. Instead, we leverage modality-specific autoencoders trained on large-scale unpaired portraits and parametric avatars, and then learn a mapping between both modalities via an alignment module trained on a significantly smaller amount of data. The resulting cross-modal latent space preserves facial identity, producing more visually appealing and higher fidelity avatars than previous methods, as supported by our quantitative and qualitative evaluations. Moreover, our method’s virtue of being resolution-independent makes it highly versatile and applicable in a wide range of settings. Manuel Ladron de Guevara, Yannick Hold-Geoffroy, Jose Echevarria, Cameron Smith, Daichi Ito |
ICCV | 3 |
| 2023 | Interactive Portrait Harmonization
Jeya Maria Jose Valanarasu, He Zhang 0004, Jianming Zhang 0001, Yilin Wang 0002, Zhe Lin 0001, Jose Echevarria, Yinglan Ma, Zijun Wei, Kalyan Sunkavalli, Vishal M. Patel |
ICLR | 6 |
| 2023 | LoCoPalettes: Local Control for Palette-based Image EditingabstractAbstract Palette‐based image editing takes advantage of the fact that color palettes are intuitive abstractions of images. They allow users to make global edits to an image by adjusting a small set of colors. Many algorithms have been proposed to compute color palettes and corresponding mixing weights. However, in many cases, especially in complex scenes, a single global palette may not adequately represent all potential objects of interest. Edits made using a single palette cannot be localized to specific semantic regions. We introduce an adaptive solution to the usability problem based on optimizing RGB palette colors to achieve arbitrary image‐space constraints and automatically splitting the image into semantic sub‐regions with more representative local palettes when the constraints cannot be satisfied. Our algorithm automatically decomposes a given image into a semantic hierarchy of soft segments. Difficult‐to‐achieve edits become straightforward with our method. Our results show the flexibility, control, and generality of our method. Cheng-Kang Ted Chao, Jason Klein, Jianchao Tan, Jose Echevarria, Yotam I. Gingold |
Comput. Graph. Forum | 4 |
| 2023 | ColorfulCurves: Palette-Aware Lightness Control and Color Editing via Sparse OptimizationabstractColor editing in images often consists of two main tasks: changing hue and saturation, and editing lightness or tone curves. State-of-the-art palette-based recoloring approaches entangle these two tasks. A user's only lightness control is changing the lightness of individual palette colors. This is inferior to state-of-the-art commercial software, where lightness editing is based on flexible tone curves that remap lightness. However, tone curves are only provided globally or per color channel (e.g., RGB). They are unrelated to the image content. Neither tone curves nor palette-based approaches support direct image-space edits---changing a specific pixel to a desired hue, saturation, and lightness. ColorfulCurves solves both of these problems by uniting palette-based and tone curve editing. In ColorfulCurves , users directly edit palette colors' hue and saturation, per-palette tone curves, or image pixels (hue, saturation, and lightness). ColorfulCurves solves an L 2,1 optimization problem in real-time to find a sparse edit that satisfies all user constraints. Our expert study found overwhelming support for ColorfulCurves over experts' preferred tools. Cheng-Kang Ted Chao, Jason Klein, Jianchao Tan, Jose Echevarria, Yotam I. Gingold |
ACM Trans. Graph. | 4 |
| 2022 | Intelli-Paint: Towards Developing More Human-Intelligible Painting Agents
Cameron Smith, Jose Echevarria, Liang Zheng 0001 |
ECCV (16) | 3 |
| 2022 | Paint2Pix: Interactive Painting Based Progressive Image Synthesis and Editing
Liang Zheng 0001, Cameron Smith, Jose Echevarria |
ECCV (14) | 4 |
| 2022 | Adaptive image vectorisation and brushing using mesh coloursabstractWe propose the use of curved triangles and mesh colours as a vector primitive for image vectorisation. We show that our representation has clear benefits for rendering performance, texture detail, as well as further editing of the resulting vector images. The proposed method focuses on efficiency, but it still leads to results that compare favourably with those from previous work. We show results over a variety of input images ranging from photos, drawings, paintings, all the way to designs and cartoons. We implemented several editing workflows facilitated by our representation: interactive user-guided vectorisation, and novel raster-style feature-aware brushing capabilities. Gerben J. Hettinga, Jose Echevarria, Jirí Kosinka |
Comput. Graph. | 2 |
| 2022 | StrokeStyles: Stroke-based Segmentation and Stylization of FontsabstractWe develop a method to automatically segment a font’s glyphs into a set of overlapping and intersecting strokes with the aim of generating artistic stylizations. The segmentation method relies on a geometric analysis of the glyph’s outline, its interior, and the surrounding areas and is grounded in perceptually informed principles and measures. Our method does not require training data or templates and applies to glyphs in a large variety of input languages, writing systems, and styles. It uses the medial axis, curvilinear shape features that specify convex and concave outline parts, links that connect concavities, and seven junction types. We show that the resulting decomposition in strokes can be used to create variations, stylizations, and animations in different artistic or design-oriented styles while remaining recognizably similar to the input font. Daniel Berio, Frederic Fol Leymarie, Paul Asente, Jose Echevarria |
ACM Trans. Graph. | 4 |
| 2022 | Instant Reality: Gaze-Contingent Perceptual Optimization for 3D Virtual Reality StreamingabstractMedia streaming, with an edge-cloud setting, has been adopted for a variety of applications such as entertainment, visualization, and design. Unlike video/audio streaming where the content is usually consumed passively, virtual reality applications require 3D assets stored on the edge to facilitate frequent edge-side interactions such as object manipulation and viewpoint movement. Compared to audio and video streaming, 3D asset streaming often requires larger data sizes and yet lower latency to ensure sufficient rendering quality, resolution, and latency for perceptual comfort. Thus, streaming 3D assets faces remarkably additional than streaming audios/videos, and existing solutions often suffer from long loading time or limited quality. To address this challenge, we propose a perceptually-optimized progressive 3D streaming method for spatial quality and temporal consistency in immersive interactions. On the cloud-side, our main idea is to estimate perceptual importance in 2D image space based on user gaze behaviors, including where they are looking and how their eyes move. The estimated importance is then mapped to 3D object space for scheduling the streaming priorities for edge-side rendering. Since this computational pipeline could be heavy, we also develop a simple neural network to accelerate the cloud-side scheduling process. We evaluate our method via subjective studies and objective analysis under varying network conditions (from 3G to 5G) and edge devices (HMD and traditional displays), and demonstrate better visual quality and temporal consistency than alternative solutions. Shaoyu Chen, Budmonde Duinkharjav, Xin Sun 0014, Li-Yi Wei, Stefano Petrangeli, Jose Echevarria, Cláudio T. Silva, Qi Sun 0003 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Dynamic Guidance for Decluttering Photographic CompositionsabstractUnwanted clutter in a photo can be incredibly distracting. However in the moment, photographers have so many things to simultaneously consider, it can be hard to catch every detail. Designers have long known the benefits of abstraction for seeing a more holistic view of their design. We wondered if, similarly, some form of image abstraction might be helpful for photographers as an alternative perspective or “lens” with which to see their image. Specifically, we wondered if such abstraction might draw the photographer’s attention away from details in the subject to noticing objects in the background, such as unwanted clutter. We present our process for designing such a camera overlay, based on the idea of using abstraction to recognize clutter. Our final design uses object-based saliency and edge detection to highlight contrast along subject and image borders, outlining potential distractors in these regions. We describe the implementation and evaluation of a capture-time tool that interactively displays these overlays and find that the tool is helpful for making users more confident in their ability to take decluttered photos that clearly convey their intended story. Jane L., Kevin Y. Zhai, Jose Echevarria, Ohad Fried, Pat Hanrahan, James A. Landay |
UIST | 3 |
| 2020 | Let Me Choose: From Verbal Context to Font SelectionabstractIn this paper, we aim to learn associations between visual attributes of fonts and the verbal context of the texts they are typically applied to. Compared to related work leveraging the surrounding visual context, we choose to focus only on the input text as this can enable new applications for which the text is the only visual element in the document. We introduce a new dataset, containing examples of different topics in social media posts and ads, labeled through crowd-sourcing. Due to the subjective nature of the task, multiple fonts might be perceived as acceptable for an input text, which makes this problem challenging. To this end, we investigate different end-to-end models to learn label distributions on crowd-sourced data and capture inter-subjectivity across all annotations. Amirreza Shirani, Franck Dernoncourt, Jose Echevarria, Paul Asente, Nedim Lipka, Thamar Solorio |
ACL | 3 |
| 2020 | Adaptive Photographic Composition GuidanceabstractPhotographic composition is often taught as alignment with composition grids-most commonly, the rule of thirds. Professional photographers use more complex grids, like the harmonic armature, to achieve more diverse dynamic compositions. We are interested in understanding whether these complex grids are helpful to amateurs. Jane E, Ohad Fried, Jingwan Lu, Jianming Zhang 0001, Radomír Mech, Jose Echevarria, Pat Hanrahan, James A. Landay |
CHI | 6 |
| 2020 | ICONATE: Automatic Compound Icon Generation and IdeationabstractCompound icons are prevalent on signs, webpages, and infographics, effectively conveying complex and abstract concepts, such as "no smoking" and "health insurance", with simple graphical representations. However, designing such icons requires experience and creativity, in order to efficiently navigate the semantics, space, and style features of icons. In this paper, we aim to automate the process of generating icons given compound concepts, to facilitate rapid compound icon creation and ideation. Informed by ethnographic interviews with professional icon designers, we have developed ICONATE, a novel system that automatically generates compound icons based on textual queries and allows users to explore and customize the generated icons. At the core of ICONATE is a computational pipeline that automatically finds commonly used icons for sub-concepts and arranges them according to inferred conventions. To enable the pipeline, we collected a new dataset, Compicon1k, consisting of 1000 compound icons annotated with semantic labels (i.e., concepts). Through user studies, we have demonstrated that our tool is able to automate or accelerate the compound icon design process for both novices and professionals. Nanxuan Zhao, Laura Mariah Herman, Hanspeter Pfister, Rynson W. H. Lau, Jose Echevarria, Zoya Bylinskii |
CHI | 6 |
| 2020 | Intuitive, Interactive Beard and Hair Synthesis With Generative ModelsabstractWe present an interactive approach to synthesizing realistic variations in facial hair in images, ranging from subtle edits to existing hair to the addition of complex and challenging hair in images of clean-shaven subjects. To circumvent the tedious and computationally expensive tasks of modeling, rendering and compositing the 3D geometry of the target hairstyle using the traditional graphics pipeline, we employ a neural network pipeline that synthesizes realistic and detailed images of facial hair directly in the target image in under one second. The synthesis is controlled by simple and sparse guide strokes from the user defining the general structural and color properties of the target hairstyle. We qualitatively and quantitatively evaluate our chosen method compared to several alternative approaches. We show compelling interactive editing results with a prototype user interface that allows novice users to progressively refine the generated image to match their desired hairstyle, and demonstrate that our approach also allows for flexible and high-fidelity scalp hair synthesis. Kyle Olszewski, Duygu Ceylan, Jun Xing, Jose Echevarria, Weikai Chen 0001, Hao Li 0015 |
CVPR | 4 |
| 2020 | Texture Hallucination for Large-Factor Painting Super-Resolution
Yulun Zhang 0001, Stephen DiVerdi, Jose Echevarria, Yun Fu 0001 |
ECCV (7) | 5 |
| 2020 | MakeltTalk: speaker-aware talking-head animationabstractWe present a method that generates expressive talking-head videos from a single facial image with audio as the only input. In contrast to previous attempts to learn direct mappings from audio to raw pixels for creating talking faces, our method first disentangles the content and speaker information in the input audio signal. The audio content robustly controls the motion of lips and nearby facial regions, while the speaker information determines the specifics of facial expressions and the rest of the talking-head dynamics. Another key component of our method is the prediction of facial landmarks reflecting the speaker-aware dynamics. Based on this intermediate representation, our method works with many portrait images in a single unified framework, including artistic paintings, sketches, 2D cartoon characters, Japanese mangas, and stylized caricatures. In addition, our method generalizes well for faces and characters that were not observed during training. We present extensive quantitative and qualitative evaluation of our method, in addition to user studies, demonstrating generated talking-heads of significantly higher quality compared to prior state-of-the-art methods. Yang Zhou 0009, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, Dingzeyu Li |
ACM Trans. Graph. | 4 |
| 2019 | Learning Emphasis Selection for Written Text in Visual Media from Crowd-Sourced Label DistributionsabstractAmirreza Shirani, Franck Dernoncourt, Paul Asente, Nedim Lipka, Seokhwan Kim, Jose Echevarria, Thamar Solorio. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Amirreza Shirani, Franck Dernoncourt, Paul Asente, Nedim Lipka, Seokhwan Kim, Jose Echevarria, Thamar Solorio |
ACL (1) | 6 |
| 2018 | Efficient palette-based decomposition and recoloring of images via RGBXY-space geometryabstractWe introduce an extremely scalable and efficient yet simple palette-based image decomposition algorithm. Given an RGB image and set of palette colors, our algorithm decomposes the image into a set of additive mixing layers, each of which corresponds to a palette color applied with varying weight. Our approach is based on the geometry of images in RGBXY-space. This new geometric approach is orders of magnitude more efficient than previous work and requires no numerical optimization. We provide an implementation of the algorithm in 48 lines of Python code. We demonstrate a real-time layer decomposition tool in which users can interactively edit the palette to adjust the layers. After preprocessing, our algorithm can decompose 6 MP images into layers in 20 milliseconds. Jianchao Tan, Jose Echevarria, Yotam I. Gingold |
ACM Trans. Graph. | 2 |