Cusuh Ham

dblp:182/9376 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0002-2686-052XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 46% 3D vision · 36% Robot manipulation · 18%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 67% Multimedia analysis and retrieval · 33%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
Personalized Residuals for Concept-Driven Text-to-Image Generation · CVPR 2024
Visual content generation and editing › image-to-image translation
sketch-to-image generation
0.612022
CoGS: Controllable Generation and Search from Sketch and Style · ECCV (16) 2022
Computer vision › 3D vision › 3d shape analysis
3d shape understanding
0.412019
ContactDB: Analyzing and Predicting Grasp Contact via Thermal Imaging · CVPR 2019
Computer vision › 3D vision › human body modeling
contact map prediction
0.412019
ContactDB: Analyzing and Predicting Grasp Contact via Thermal Imaging · CVPR 2019
Robotics › Robot manipulation
grasping
0.412019
ContactDB: Analyzing and Predicting Grasp Contact via Thermal Imaging · CVPR 2019
Multimedia analysis and retrieval › multimedia retrieval › content-based retrieval
cross-domain retrieval
0.212016
The sketchy database: learning to retrieve badly drawn bunnies · ACM Trans. Graph. 2016
Multimedia analysis and retrieval › image retrieval
sketch-based image retrieval
0.212016
The sketchy database: learning to retrieve badly drawn bunnies · ACM Trans. Graph. 2016
Machine learning › Generative modeling
diffusion model
0.212024
Personalized Residuals for Concept-Driven Text-to-Image Generation · CVPR 2024
Wearable and physiological sensing › camera-based sensing
thermal imaging
0.112019
ContactDB: Analyzing and Predicting Grasp Contact via Thermal Imaging · CVPR 2019
Multimedia analysis and retrieval › image analysis
image understanding
0.112016
The sketchy database: learning to retrieve badly drawn bunnies · ACM Trans. Graph. 2016

Methods — techniques the papers use, named apart from their topics

low-rank residual · 0.8image translation · 0.8cross-attention localization · 0.83d convolution · 0.8cross-domain embedding · 0.2convolutional neural network · 0.2
YearPublicationVenuePosition
2026 CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition
abstract
Multi-shot generation requires preserving the identity of characters and settings across frames. Cinematic scene composition goes beyond standard multi-shot generation, introducing additional challenges such as expressing complex interactions among multiple characters and visual effects to convey creative narratives—challenges existing datasets cannot fully address. We present CineVerse, a large-scale dataset of diverse movie scenes labeled with shot-level annotations tailored for filmmaking. CineVerse includes refined scene descriptions, shot-type information, and newly extracted shot, character, setting descriptions. We validate our dataset by developing a baseline framework that first generates a scene plan containing detailed information for the overall scene and each individual shot, then produces a set of coherent keyframes. Our results show significant improvements in controlling and synthesizing cinematic content through the added context provided by CineVerse.
Quynh Phung, Long Mai, Fabian Caba Heilbron, Feng Liu 0015, Jia-Bin Huang 0001, Cusuh Ham
WACV6
2024 Personalized Residuals for Concept-Driven Text-to-Image Generation
abstract
We present personalized residuals and localized attention-guided sampling for efficient concept-driven generation using text-to-image diffusion models. Our method first represents concepts by freezing the weights of a pretrained text-conditioned diffusion model and learning low-rank residuals for a small subset of the model's layers. The residual-based approach then directly enables application of our proposed sampling technique, which applies the learned residuals only in areas where the concept is localized via cross-attention and applies the original diffusion weights in all other regions. Localized sampling therefore combines the learned identity of the concept with the existing generative prior of the underlying diffusion model. We show that personalized residuals effectively capture the identity of a concept in$\sim$3 minutes on a single GPU without the use of regularization images and with fewer parameters than previous models, and localized sampling allows using the original model as strong prior for large parts of the image.
Cusuh Ham, Matthew Fisher, James Hays, Nicholas I. Kolkin, Yuchen Liu 0002, Richard Zhang 0001, Tobias Hinz
CVPR1
2022 CoGS: Controllable Generation and Search from Sketch and Style
Cusuh Ham, Gemma Canet Tarrés, Tu Bui, James Hays, Zhe Lin 0001, John P. Collomosse
ECCV (16)1
2021 Density of States Estimation for Out of Distribution Detection
abstract
Perhaps surprisingly, recent studies have shown probabilistic model likelihoods have poor specificity for out-of-distribution (OOD) detection and often assign higher likelihoods to OOD data than in-distribution data. To ameliorate this issue we propose DoSE, the density of states estimator. Drawing on the statistical physics notion of “density of states,” the DoSE decision rule avoids direct comparison of model probabilities, and instead utilizes the “probability of the model probability,” or indeed the frequency of any reasonable statistic. The frequency is calculated using nonparametric density estimators (e.g., KDE and one-class SVM) which measure the typicality of various model statistics given the training data and from which we can flag test points with low typicality as anomalous. Unlike many other methods, DoSE requires neither labeled data nor OOD examples. DoSE is modular and can be trivially applied to any existing, trained model. We demonstrate DoSE’s state-of-the-art performance against other unsupervised OOD detectors on previously established “hard” benchmarks.
Warren R. Morningstar, Cusuh Ham, Andrew G. Gallagher, Balaji Lakshminarayanan, Alexander A. Alemi, Joshua V. Dillon
AISTATS2
2021 Automatic Differentiation Variational Inference with Mixtures
abstract
Automatic Differentiation Variational Inference (ADVI) is a useful tool for efficiently learning probabilistic models in machine learning. Generally approximate posteriors learned by ADVI are forced to be unimodal in order to facilitate use of the reparameterization trick. In this paper, we show how stratified sampling may be used to enable mixture distributions as the approximate posterior, and derive a new lower bound on the evidence analogous to the importance weighted autoencoder (IWAE). We show that this "SIWAE" is a tighter bound than both IWAE and the traditional ELBO, both of which are special instances of this bound. We verify empirically that the traditional ELBO objective disfavors the presence of multimodal posterior distributions and may therefore not be able to fully capture structure in the latent space. Our experiments show that using the SIWAE objective allows the encoder to learn more complex distributions which regularly contain multimodality, resulting in higher accuracy and better calibration in the presence of incomplete, limited, or corrupted data.
Warren R. Morningstar, Sharad M. Vikram, Cusuh Ham, Andrew G. Gallagher, Joshua V. Dillon
AISTATS3
2019 ContactDB: Analyzing and Predicting Grasp Contact via Thermal Imaging
abstract
Grasping and manipulating objects is an important human skill. Since hand-object contact is fundamental to grasping, capturing it can lead to important insights. However, observing contact through external sensors is challenging because of occlusion and the complexity of the human hand. We present ContactDB, a novel dataset of contact maps for household objects that captures the rich hand-object contact that occurs during grasping, enabled by use of a thermal camera. Participants in our study grasped 3D printed objects with a post-grasp functional intent. ContactDB includes 3750 3D meshes of 50 household objects textured with contact maps and 375K frames of synchronized RGB-D+thermal images. To the best of our knowledge, this is the first large-scale dataset that records detailed contact maps for human grasps. Analysis of this data shows the influence of functional intent and object size on grasping, the tendency to touch/avoid `active areas', and the high frequency of palm and proximal finger contact. Finally, we train state-of-the art image translation and 3D convolution algorithms to predict diverse contact patterns from object shape. Data, code and models are available at https://contactdb.cc.gatech.edu.
Samarth Brahmbhatt, Cusuh Ham, Charles C. Kemp, James Hays
CVPR2
2016 The sketchy database: learning to retrieve badly drawn bunnies
abstract
We present the Sketchy database , the first large-scale collection of sketch-photo pairs. We ask crowd workers to sketch particular photographic objects sampled from 125 categories and acquire 75,471 sketches of 12,500 objects. The Sketchy database gives us fine-grained associations between particular photos and sketches, and we use this to train cross-domain convolutional networks which embed sketches and photographs in a common feature space. We use our database as a benchmark for fine-grained retrieval and show that our learned representation significantly outperforms both hand-crafted features as well as deep features trained for sketch or photo classification. Beyond image retrieval, we believe the Sketchy database opens up new opportunities for sketch and image understanding and synthesis.
Patsorn Sangkloy, Nathan Burnell, Cusuh Ham, James Hays
ACM Trans. Graph.3