EDBT 2026 Demo / reviewers in the wild / expert
Jihun Hamm
dblp:69/7426 · also Jihun Ham
· DBLP profile ↗
40ranked-venue papers
14as first author
17since 2021 · last 2025
0000-0002-0680-0901ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 9 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3 · 2 first-authorSecurity and privacy · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DPCore: Dynamic Prompt Coreset for Continual Test-Time AdaptationabstractContinual Test-Time Adaptation (CTTA) seeks to adapt source pre-trained models to continually changing, unseen target domains. While existing CTTA methods assume structured domain changes with uniform durations, real-world environments often exhibit dynamic patterns where domains recur with varying frequencies and durations. Current approaches, which adapt the same parameters across different domains, struggle in such dynamic conditions—they face convergence issues with brief domain exposures, risk forgetting previously learned knowledge, or misapplying it to irrelevant domains. To remedy this, we propose **DPCore**, a method designed for robust performance across diverse domain change patterns while ensuring computational efficiency. DPCore integrates three key components: Visual Prompt Adaptation for efficient domain alignment, a Prompt Coreset for knowledge preservation, and a Dynamic Update mechanism that intelligently adjusts existing prompts for similar domains while creating new ones for substantially different domains. Extensive experiments on four benchmarks demonstrate that DPCore consistently outperforms various CTTA methods, achieving state-of-the-art performance in both structured and dynamic settings while reducing trainable parameters by 99% and computation time by 64% compared to previous approaches. Yunbei Zhang, Akshay Mehra, Shuaicheng Niu, Jihun Hamm |
ICML | 4 |
| 2025 | SOFA: Deep Learning Framework for Simulating and Optimizing Atrial Fibrillation Ablation
Yunsung Chung, Chanho Lim, Ghassan Bidaoui, Christian Massad, Nassir Marrouche, Jihun Hamm |
MICCAI (4) | 6 |
| 2025 | eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin DiseasesabstractSkin Neglected Tropical Diseases (NTDs) impose severe health and socioeconomic burdens in impoverished tropical communities. Yet, advancements in AI-driven diagnostic support are hindered by data scarcity, particularly for underrepresented populations and rare manifestations of NTDs. Existing dermatological datasets often lack the demographic and disease spectrum crucial for developing reliable recognition models of NTDs. To address this, we introduce eSkinHealth, a novel dermatological dataset collected on-site in Côte d'Ivoire and Ghana. Specifically, eSkinHealth contains 5,623 images from 1,639 cases and encompasses 47 skin diseases, focusing uniquely on skin NTDs and rare conditions among West African populations. We further propose an AI-expert collaboration paradigm to implement foundation language and segmentation models for efficient generation of multimodal annotations, under dermatologists' guidance. In addition to patient metadata and diagnosis labels, eSkinHealth also includes semantic lesion masks, instance-specific visual captions, and clinical concepts. Overall, our work provides a valuable new resource and a scalable annotation framework, aiming to catalyze the development of more equitable, accurate, and interpretable AI tools for global dermatology. Janet Wang, Yunbei Zhang, Diabate Almamy, Vagamon Bamba, Konan Amos Sébastien Koffi, Koffi Aubin Yao, Zhengming Ding, Jihun Hamm, Rie Roselyne Yotsu |
ACM Multimedia | 9 |
| 2025 | Visual Instance-aware Prompt TuningabstractVisual Prompt Tuning (VPT) has emerged as a parameter-efficient fine-tuning paradigm for vision transformers, with conventional approaches utilizing dataset-level prompts that remain the same across all input instances. We observe that this strategy results in sub-optimal performance due to high variance in downstream datasets. To address this challenge, we propose Visual Instance-aware Prompt Tuning (ViaPT), which generates instance-aware prompts based on each individual input and fuses them with dataset-level prompts, leveraging Principal Component Analysis (PCA) to retain important prompting information. Moreover, we reveal that VPT-Deep and VPT-Shallow represent two corner cases based on a conceptual understanding, in which they fail to effectively capture instance-specific information, while random dimension reduction on prompts only yields performance between the two extremes. Instead, ViaPT overcomes these limitations by balancing dataset-level and instance-level knowledge, while reducing the amount of learnable parameters compared to VPT-Deep. Extensive experiments across 34 diverse datasets demonstrate that our method consistently outperforms state-of-the-art baselines, establishing a new paradigm for analyzing and optimizing visual prompts for vision transformers. Xi Xiao 0003, Yunbei Zhang, Xingjian Li 0002, Tianyang Wang 0004, Xiao Wang 0004, Yuxiang Wei 0004, Jihun Hamm, Min Xu 0009 |
ACM Multimedia | 7 |
| 2025 | Doctor Approved: Generating Medically Accurate Skin Disease Images through AI-Expert FeedbackabstractPaucity of medical data severely limits the generalizability of diagnostic ML models, as the full spectrum of disease variability can not be represented by a small clinical dataset. To address this, diffusion models (DMs) have been considered as a promising avenue for synthetic image generation and augmentation. However, they frequently produce _medically inaccurate_ images, deteriorating the model performance. Expert domain knowledge is critical for synthesizing images that correctly encode clinical information, especially when data is scarce and quality outweighs quantity. Existing approaches for incorporating human feedback, such as reinforcement learning (RL) and Direct Preference Optimization (DPO), rely on robust reward functions or demand labor-intensive expert evaluations. Recent progress in Multimodal Large Language Models (MLLMs) reveals their strong visual reasoning capabilities, making them adept candidates as evaluators. In this work, we propose a novel framework, coined MAGIC (**M**edically **A**ccurate **G**eneration of **I**mages through AI-Expert **C**ollaboration), that synthesizes clinically accurate skin disease images for data augmentation. Our method creatively translates expert-defined criteria into actionable feedback for image synthesis of DMs, significantly improving clinical accuracy while reducing the direct human workload. Experiments demonstrate that our method greatly improves the clinical quality of synthesized skin disease images, with outputs aligning with dermatologist assessments. Additionally, augmenting training data with these synthesized images improves diagnostic accuracy by +9.02% on a challenging 20-condition skin disease classification task, and by +13.89% in the few-shot setting. Beyond image synthesis, MAGIC illustrates a task-centric alignment paradigm: instead of adapting MLLMs to niche medical tasks, it adapts tasks to the evaluative strengths of general-purpose MLLMs by decomposing domain knowledge into attribute-level checklists. This design offers a scalable and reliable path for leveraging foundation models in specialized domains. Janet Wang, Yunbei Zhang, Zhengming Ding, Jihun Hamm |
NeurIPS | 4 |
| 2025 | SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation
Yunsung Chung, Yunbei Zhang, Nassir Marrouche, Jihun Hamm |
USENIX Security Symposium | 4 |
| 2025 | Enhancing Skin Disease Diagnosis: Interpretable Visual Concept Discovery with SAMabstractCurrent AI-assisted skin image diagnosis has achieved dermatologist-level performance in classifying skin cancer, driven by rapid advancements in deep learning architectures. However, unlike traditional vision tasks, skin images in general present unique challenges due to the limited availability of well-annotated datasets, complex variations in conditions, and the necessity for detailed interpretations to ensure patient safety. Previous segmentation methods have sought to reduce image noise and enhance diagnostic performance, but these techniques require fine-grained, pixel-level ground truth masks for training. In contrast, with the rise of foundation models, the Segment Anything Model (SAM) has been introduced to facilitate promptable segmentation, enabling the automation of the segmentation process with simple yet effective prompts. Efforts applying SAM pre-dominantly focus on dermatoscopy images, which present more easily identifiable lesion boundaries than clinical photos taken with smartphones. This limitation constrains the practicality of these approaches to real-world applications. To overcome the challenges posed by noisy clinical photos acquired via non-standardized protocols and to improve diagnostic accessibility, we propose a novel Cross-Attentive Fusion framework for interpretable skin lesion diagnosis. Our method leverages SAM to generate visual concepts for skin diseases using prompts, integrating local visual concepts with global image features to enhance model performance. Extensive evaluation on two skin disease datasets demonstrates our proposed method's effectiveness on lesion diagnosis and interpretability. Janet Wang, Jihun Hamm, Rie Roselyne Yotsu, Zhengming Ding |
WACV | 3 |
| 2025 | OT-VP: Optimal Transport-Guided Visual Prompting for Test-Time AdaptationabstractVision Transformers (ViTs) have demonstrated remarkable capabilities in learning representations, but their performance is compromised when applied to unseen domains. Previous methods either engage in prompt learning during the training phase or modify model parameters at test time through entropy minimization. The former often over-looks unlabeled target data, while the latter doesn't fully address domain shifts. In this work, our approach, Optimal Transport-guided Test-Time Visual Prompting (OT-VP), handles these problems by leveraging prompt learning at test time to align the target and source domains without accessing the training process or altering pretrained model parameters. This method involves learning a universal visual prompt for the target domain by optimizing the Optimal Transport distance. With only four learned prompt tokens, OT-VP exceeds state-of-the-art performance across three stylistic datasets—PACS, VLCS, OfficeHome, and one corrupted dataset ImageNet-C. Additionally, OT-VP operates efficiently, both in terms of memory and computation, and is adaptable for extension to online settings. The code is available at https://github.com/zybeich/OT-VP. Yunbei Zhang, Akshay Mehra, Jihun Hamm |
WACV | 3 |
| 2024 | Understanding the Transferability of Representations via Task-RelatednessabstractThe growing popularity of transfer learning due to the availability of models pre-trained on vast amounts of data, makes it imperative to understand when the knowledge of these pre-trained models can be transferred to obtain high-performing models on downstream target tasks. However, the exact conditions under which transfer learning succeeds in a cross-domain cross-task setting are still poorly understood. To bridge this gap, we propose a novel analysis that analyzes the transferability of the representations of pre-trained models to downstream tasks in terms of their relatedness to a given reference task. Our analysis leads to an upper bound on transferability in terms of task-relatedness, quantified using the difference between the class priors, label sets, and features of the two tasks.Our experiments using state-of-the-art pre-trained models show the effectiveness of task-relatedness in explaining transferability on various vision and language tasks. The efficient computability of task-relatedness even without labels of the target task and its high correlation with the model's accuracy after end-to-end fine-tuning on the target task makes it a useful metric for transferability estimation. Our empirical results of using task-relatedness on the problem of selecting the best pre-trained model from a model zoo for a target task highlight its utility for practical problems. Akshay Mehra, Yunbei Zhang, Jihun Hamm |
NeurIPS | 3 |
| 2024 | On the Fly Neural Style Smoothing for Risk-Averse Domain GeneralizationabstractAchieving high accuracy on data from domains unseen during training is a fundamental challenge in domain generalization (DG). While state-of-the-art (SOTA) DG classifiers have demonstrated impressive performance across various tasks, they have shown a bias towards domain-dependent information, such as image styles, rather than domain-invariant information, such as image content. This bias renders them unreliable for deployment in risk-sensitive scenarios such as autonomous driving where a misclassification could have catastrophic consequences. To enable risk-averse predictions from a DG classifier, we propose a novel inference procedure, Test-Time Neural Style Smoothing (TT-NSS), that uses a "style-smoothed" version of the DG classifier for prediction at test time. Specifically, the style-smoothed classifier classifies a test image as the most probable class predicted by the DG classifier on random re-stylizations of the test image. TT-NSS uses a neural style transfer module to stylize a test image on the fly, requires only black-box access to the DG classifier, and crucially, abstains when predictions of the DG classifier on the stylized test images lack consensus. Additionally, we propose a neural style smoothing (NSS) based training procedure that can be seamlessly integrated with existing DG methods. This procedure enhances the prediction consistency of DG classifiers, improving the performance of TT-NSS on non-abstained samples. Our empirical results demonstrate the effectiveness of TT-NSS and NSS at producing and improving risk-averse predictions on unseen domains from DG classifiers trained with SOTA training methods on various benchmark datasets and their variations. Akshay Mehra, Yunbei Zhang, Bhavya Kailkhura, Jihun Hamm |
WACV | 4 |
| 2024 | Marginalized Augmented Few-Shot Domain AdaptationabstractDomain adaptation (DA) has recently drawn a lot of attention, as it facilitates unlabeled target learning by borrowing knowledge from an external source domain. Most existing DA solutions seek to align feature representations between the labeled source and unlabeled target data. However, the scarcity of target data easily results in negative transfer, as it misleads the cross DA to the dominance of the source. To address the challenging few-shot domain adaptation (FSDA) problem, in this article, we propose a novel marginalized augmented FSDA (MAF) approach to address the cross-domain distribution disparity and insufficiency of target data simultaneously. On the one hand, cross-domain continuity augmentation (CCA) synthesizes abundant intermediate patterns across domains leading to a continuous domain-invariant latent space. On the other hand, sufficient source-supervised semantic augmentation (SSA) is explored to progressively diversify the conditional distribution within and across domains. Moreover, the proposed augmentation strategies are implemented efficiently via an expected transferable cross-entropy (CE) loss over the augmented distribution instead of explicit data synthesis, and minimizing the upper bound of the expected loss introduces negligible extra computing cost. Experimentally, our method outperforms the state of the art in various FSDA benchmarks, which demonstrates the effectiveness and contribution of our work. Our source code is provided at https://github.com/scottjingtt/MAF.git. Taotao Jing, Haifeng Xia, Jihun Hamm, Zhengming Ding |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | A Spectral View of Randomized Smoothing Under Common Corruptions: Benchmarking and Improving Certified Robustness
Akshay Mehra, Bhavya Kailkhura, Dan Hendrycks, Jihun Hamm, Z. Morley Mao |
ECCV (4) | 6 |
| 2022 | Online Evasion Attacks on Recurrent Models: The Power of Hallucinating the FutureabstractRecurrent models are frequently being used in online tasks such as autonomous driving, and a comprehensive study of their vulnerability is called for. Existing research is limited in generality only addressing application-specific vulnerability or making implausible assumptions such as the knowledge of future input. In this paper, we present a general attack framework for online tasks incorporating the unique constraints of the online setting different from offline tasks. Our framework is versatile in that it covers time-varying adversarial objectives and various optimization constraints, allowing for a comprehensive study of robustness. Using the framework, we also present a novel white-box attack called Predictive Attack that `hallucinates' the future. The attack achieves 98 percent of the performance of the ideal but infeasible clairvoyant attack on average. We validate the effectiveness of the proposed framework and attacks through various experiments. Byunggill Joe, Insik Shin, Jihun Hamm |
IJCAI | 3 |
| 2022 | Augmented Multimodality Fusion for Generalized Zero-Shot Sketch-Based Visual RetrievalabstractZero-shot sketch-based image retrieval (ZS-SBIR) has attracted great attention recently, due to the potential application of sketch-based retrieval under zero-shot scenarios, where the categories of query sketches and gallery photos are not observed in the training stage. However, it is still under insufficient exploration for the general and practical scenario when the query sketches and gallery photos contain both seen and unseen categories. Such a problem is defined as generalized zero-shot sketch-based image retrieval (GZS-SBIR), which is the focus of this work. To this end, we propose a novel Augmented Multi-modality Fusion (AMF) framework to generalize seen concepts to unobserved ones efficiently. Specifically, a novel knowledge discovery module named cross-domain augmentation is designed in both visual and semantic space to mimic novel knowledge unseen from the training stage, which is the key to handling the GZS-SBIR challenge. Moreover, a triplet domain alignment module is proposed to couple the cross-domain distribution between photo and sketch in visual space. To enhance the robustness of our model, we explore embedding propagation to refine both visual and semantic features by removing undesired noise. Eventually, visual-semantic fusion representations are concatenated for further domain discrimination and task-specific recognition, which tend to trigger the cross-domain alignment in both visual and semantic feature space. Experimental evaluations are conducted on popular ZS-SBIR benchmarks as well as a new evaluation protocol designed for GZS-SBIR from DomainNet dataset with more diverse sub-domains, and the promising results demonstrate the superiority of the proposed solution over other baselines. The source code is available at https://github.com/scottjingtt/AMF_GZS_SBIR.git. Taotao Jing, Haifeng Xia, Jihun Hamm, Zhengming Ding |
IEEE Trans. Image Process. | 3 |
| 2021 | Penalty Method for Inversion-Free Deep Bilevel OptimizationabstractSolving a bilevel optimization problem is at the core of several machine learning problems such as hyperparameter tuning, data denoising, meta- and few-shot learning, and trainingdata poisoning. Different from simultaneous or multi-objective optimization, the steepest descent direction for minimizing the upper-level cost in a bilevel problem requires the inverse of the Hessian of the lower-level cost. In this work, we propose a novel algorithm for solving bilevel optimization problems based on the classical penalty function approach. Our method avoids computing the Hessian inverse and can handle constrained bilevel problems easily. We prove the convergence of the method under mild conditions and show that the exact hypergradient is obtained asymptotically. Our method’s simplicity and small space and time complexities enable us to effectively solve large-scale bilevel problems involving deep neural networks. We present results on data denoising, few-shot learning, and training-data poisoning problems in a large-scale setting. Our results show that our approach outperforms or is comparable to previously proposed methods based on automatic differentiation and approximate inversion in terms of accuracy, run-time, and convergence speed Akshay Mehra, Jihun Hamm |
ACML | 2 |
| 2021 | How Robust Are Randomized Smoothing Based Defenses to Data Poisoning?abstractPredictions of certifiably robust classifiers remain constant in a neighborhood of a point, making them resilient to test-time attacks with a guarantee. In this work, we present a previously unrecognized threat to robust machine learning models that highlights the importance of training-data quality in achieving high certified adversarial robustness. Specifically, we propose a novel bilevel optimization based data poisoning attack that degrades the robustness guarantees of certifiably robust classifiers. Unlike other poisoning attacks that reduce the accuracy of the poisoned models on a small set of target points, our attack reduces the average certified radius (ACR) of an entire target class in the dataset. Moreover, our attack is effective even when the victim trains the models from scratch using state-of-the-art robust training methods such as Gaussian data augmentation[8], MACER[36], and SmoothAdv[29] that achieve high certified adversarial robustness. To make the attack harder to detect, we use clean-label poisoning points with imperceptible distortions. The effectiveness of the proposed method is evaluated by poisoning MNIST and CIFAR10 datasets and training deep neural networks using previously mentioned training methods and certifying the robustness with randomized smoothing. The ACR of the target class, for models trained on generated poison data, can be reduced by more than 30%. Moreover, the poisoned data is transferable to models trained with different training methods and models with different architectures. Akshay Mehra, Bhavya Kailkhura, Jihun Hamm |
CVPR | 4 |
| 2021 | Understanding the Limits of Unsupervised Domain Adaptation via Data PoisoningabstractUnsupervised domain adaptation (UDA) enables cross-domain learning without target domain labels by transferring knowledge from a labeled source domain whose distribution differs from that of the target. However, UDA is not always successful and several accounts of `negative transfer' have been reported in the literature. In this work, we prove a simple lower bound on the target domain error that complements the existing upper bound. Our bound shows the insufficiency of minimizing source domain error and marginal distribution mismatch for a guaranteed reduction in the target domain error, due to the possible increase of induced labeling function mismatch. This insufficiency is further illustrated through simple distributions for which the same UDA approach succeeds, fails, and may succeed or fail with an equal chance. Motivated from this, we propose novel data poisoning attacks to fool UDA methods into learning representations that produce large target domain errors. We evaluate the effect of these attacks on popular UDA methods using benchmark datasets where they have been previously shown to be successful. Our results show that poisoning can significantly decrease the target domain accuracy, dropping it to almost 0% in some cases, with the addition of only 10% poisoned data in the source domain. The failure of these UDA methods demonstrates their limitations at guaranteeing cross-domain generalization consistent with our lower bound. Thus, evaluating UDA methods in adversarial settings such as data poisoning provides a better sense of their robustness to data distributions unfavorable for UDA. Akshay Mehra, Bhavya Kailkhura, Jihun Hamm |
NeurIPS | 4 |
| 2019 | Statistical Privacy for Streaming Traffic
Xiaokuan Zhang, Jihun Hamm, Michael K. Reiter, Yinqian Zhang |
NDSS | 2 |
| 2018 | K-Beam Minimax: Efficient Optimization for Deep Adversarial LearningabstractMinimax optimization plays a key role in adversarial training of machine learning algorithms, such as learning generative models, domain adaptation, privacy preservation, and robust learning. In this paper, we demonstrate the failure of alternating gradient descent in minimax optimization problems due to the discontinuity of solutions of the inner maximization. To address this, we propose a new $\epsilon$-subgradient descent algorithm that addresses this problem by simultaneously tracking $K$ candidate solutions. Practically, the algorithm can find solutions that previous saddle-point algorithms cannot find, with only a sublinear increase of complexity in $K$. We analyze the conditions under which the algorithm converges to the true solution in detail. A significant improvement in stability and convergence speed of the algorithm is observed in simple representative problems, GAN training, and domain-adaptation problems. Jihun Hamm, Yung-Kyun Noh |
ICML | 1 |
| 2018 | Fluid Dynamic Models for Bhattacharyya-Based Discriminant AnalysisabstractClassical discriminant analysis attempts to discover a low-dimensional subspace where class label information is maximally preserved under projection. Canonical methods for estimating the subspace optimize an information-theoretic criterion that measures the separation between the class-conditional distributions. Unfortunately, direct optimization of the information-theoretic criteria is generally non-convex and intractable in high-dimensional spaces. In this work, we propose a novel, tractable algorithm for discriminant analysis that considers the class-conditional densities as interacting fluids in the high-dimensional embedding space. We use the Bhattacharyya criterion as a potential function that generates forces between the interacting fluids, and derive a computationally tractable method for finding the low-dimensional subspace that optimally constrains the resulting fluid flow. We show that this model properly reduces to the optimal solution for homoscedastic data as well as for heteroscedastic Gaussian distributions with equal means. We also extend this model to discover optimal filters for discriminating Gaussian processes and provide experimental results and comparisons on a number of datasets. Yung-Kyun Noh, Jihun Hamm, Frank C. Park 0001, Byoung-Tak Zhang, Daniel D. Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Enhancing utility and privacy with noisy minimax filtersabstractPreserving privacy of continuous and/or high-dimensional data such as images, videos and audios is challenging. Syntactic anonymization methods were proposed typically for discrete data types and can be unsuitable. Differential privacy, which provides a stricter type of privacy, has shown more success in sanitizing continuous data. However, both syntactic and differential privacy are susceptible to inference attacks, i.e., an adversary can accurately guess sensitive attributes from insensitive attributes. On the other hand, minimax filters were proposed previously to minimize the accuracy of inference while maximizing utility at the same time. The paper presents noisy minimax filter that combines minimax filter and differentially private mechanism, which can attain high average utility and protection against inference attacks and a formal worst-case privacy guarantee. The proposed algorithm is demonstrated with real databases of faces, voices, and motion data. Jihun Hamm |
ICASSP | 1 |
| 2017 | Crowd-ML: A library for privacy-preserving machine learning on smart devicesabstractWhen user-generated data such as audio and video signals are used to train machine learning algorithms, users' privacy must be considered before the learned model is released. In this work, we present an open-source library for privacy-preserving machine learning framework on smart devices. The library allows Android and iOS devices to collectively learn a common classifier/regression model from distributed data with differential privacy, using a variant of minibatch stochastic gradient descent method. The library allows researchers and developers to easily implement and deploy customized tasks that use on-device sensors to collect sensitive data for machine learning. Jihun Hamm, Jackson Luken, Yani Xie |
ICASSP | 1 |
| 2017 | Minimax Filter: Learning to Preserve Privacy from Inference AttacksabstractPreserving privacy of continuous and/or high-dimensional data such as images, videos and audios, can be challenging with syntactic anonymization methods which are designed for discrete attributes. Differentially privacy, which uses a more rigorous definition of privacy loss, has shown more success in sanitizing continuous data. However, both syntactic and differential privacy are susceptible to inference attacks, i.e., an adversary can accurately infer sensitive attributes from sanitized data. The paper proposes a novel filter-based mechanism which preserves privacy of continuous and high-dimensional attributes against inference attacks. Finding the optimal utility-privacy tradeoff is formulated as a min-diff-max optimization problem. The paper provides an ERM-like analysis of the generalization error and also a practical algorithm to perform minimax optimization. In addition, the paper proposes a noisy minimax filter which combines minimax filter and differentially-private mechanism. Advantages of the method over purely noisy mechanisms is explained and demonstrated with examples. Experiments with several real-world tasks including facial expression classification, speech emotion classification, and activity classification from motion, show that the minimax filter can simultaneously achieve similar or higher target task accuracy and lower inference accuracy, often significantly lower than previous methods. Jihun Hamm |
J. Mach. Learn. Res. | 1 |
| 2016 | Learning privately from multiparty dataabstractLearning a classifier from private data distributed across multiple parties is an important problem that has many potential applications. How can we build an accurate and differentially private global classifier by combining locally-trained classifiers from different parties, without access to any party’s private data? We propose to transfer the “knowledge” of the local classifier ensemble by first creating labeled data from auxiliary unlabeled data, and then train a global differentially private classifier. We show that majority voting is too sensitive and therefore propose a new risk weighted by class probabilities estimated from the ensemble. Relative to a non-private solution, our private solution has a generalization error bounded by O(ε^-2 M^-2). This allows strong privacy without performance loss when the number of participating parties M is large, such as in crowdsensing applications. We demonstrate the performance of our framework with realistic tasks of activity recognition, network intrusion detection, and malicious URL detection. Jihun Hamm, Yingjun Cao, Mikhail Belkin |
ICML | 1 |
| 2015 | Preserving Privacy of Continuous High-dimensional Data with Minimax FiltersabstractPreserving privacy of high-dimensional and continuous data such as images or biometric data is a challenging problem. This paper formulates this problem as a learning game between three parties: 1) data contributors using a filter to sanitize data samples, 2) a cooperative data aggregator learning a target task using the filtered samples, and 3) an adversary learning to identify contributors using the same filtered samples. Minimax filters that achieve the optimal privacy-utility trade-off from broad families of filters and loss/classifiers are defined, and algorithms for learning the filers in batch or distributed settings are presented. Experiments with several real-world tasks including facial expression recognition, speech emotion recognition, and activity recognition from motion, show that the minimax filter can simultaneously achieve similar or better target task accuracy and lower privacy risk, often significantly lower than previous methods. Jihun Hamm |
AISTATS | 1 |
| 2015 | Crowd-ML: A Privacy-Preserving Learning Framework for a Crowd of Smart DevicesabstractSmart devices with built-in sensors, computational capabilities, and network connectivity have become increasingly pervasive. Crowds of smart devices offer opportunities to collectively sense and perform computing tasks at an unprecedented scale. This paper presents Crowd-ML, a privacy-preserving machine learning framework for a crowd of smart devices, which can solve a wide range of learning problems for crowd sensing data with differential privacy guarantees. Crowd-ML endows a crowd sensing system with the ability to learn classifiers or predictors online from crowd sensing data privately with minimal computational overhead on devices and servers, suitable for practical large-scale use of the framework. We analyze the performance and scalability of Crowd-ML and implement the system with off-the-shelf smartphones as a proof of concept. We demonstrate the advantages of Crowd-ML with real and simulated experiments under various conditions. Jihun Hamm, Adam C. Champion, Guoxing Chen, Mikhail Belkin, Dong Xuan |
ICDCS | 1 |
| 2015 | Qualitative Tracking Performance Evaluation without Ground-TruthabstractWe present a qualitative tracking performance evaluation algorithm without ground-truth, where several representative frames are automatically selected and visualized in a principled way. Although tracking algorithms are typically evaluated by quantitative scores based on predefined measures, qualitative evaluation is also useful especially when the ground-truth of a target state is unavailable or unreliable. However, there is no prior study on how to present frames for better qualitative evaluation of tracking algorithms. Motivated by this fact, we propose an unbiased frame selection technique, where salient and unique features in tracking results are captured effectively. Our method identifies a set of representative frames by 1) analyzing the sequence structure using manifold learning, and 2) selecting frames by formulating the task as a facility location problem. By presenting the manifold and the selected frames, one can understand the sequence structure as well as the characteristics of tracking results. The effectiveness of our method is illustrated with single and multiple tracking results for sequences without ground-truth. Bohyung Han, Jihun Hamm |
WACV | 2 |
| 2014 | Regional Manifold Learning for Disease ClassificationabstractWhile manifold learning from images itself has become widely used in medical image analysis, the accuracy of existing implementations suffers from viewing each image as a single data point. To address this issue, we parcellate images into regions and then separately learn the manifold for each region. We use the regional manifolds as low-dimensional descriptors of high-dimensional morphological image features, which are then fed into a classifier to identify regions affected by disease. We produce a single ensemble decision for each scan by the weighted combination of these regional classification results. Each weight is determined by the regional accuracy of detecting the disease. When applied to cardiac magnetic resonance imaging of 50 normal controls and 50 patients with reconstructive surgery of Tetralogy of Fallot, our method achieves significantly better classification accuracy than approaches learning a single manifold across the entire image domain. Dong Hye Ye, Benoit Desjardins, Jihun Hamm, Harold Litt, Kilian M. Pohl |
IEEE Trans. Medical Imaging | 3 |
| 2013 | FLOOR: Fusing Locally Optimal Registrations
Dong Hye Ye, Jihun Hamm, Benoit Desjardins, Kilian M. Pohl |
MICCAI (3) | 2 |
| 2012 | Regional Manifold Learning for Deformable Registration of Brain MR Images
Dong Hye Ye, Jihun Hamm, Dongjin Kwon, Christos Davatzikos, Kilian M. Pohl |
MICCAI (3) | 2 |
| 2011 | Personalized video summarization with human in the loopabstractIn automatic video summarization, visual summary is constructed typically based on the analysis of low-level features with little consideration of video semantics. However, the contextual and semantic information of a video is marginally related to low-level features in practice although they are useful to compute visual similarity between frames. Therefore, we propose a novel video summarization technique, where the semantically important information is extracted from a set of keyframes given by human and the summary of a video is constructed based on the automatic temporal segmentation using the analysis of inter-frame similarity to the keyframes. Toward this goal, we model a video sequence with a dissimilarity matrix based on bidirectional similarity measure between every pair of frames, and subsequently characterize the structure of the video by a nonlinear manifold embedding. Then, we formulate video summarization as a variant of the 0-1 knapsack problem, which is solved by dynamic programming efficiently. The effectiveness of our algorithm is illustrated quantitatively and qualitatively using realistic videos collected from YouTube. Bohyung Han, Jihun Hamm, Jack Sim |
WACV | 2 |
| 2010 | GRAM: A framework for geodesic registration on anatomical manifolds
Jihun Hamm, Dong Hye Ye, Ragini Verma, Christos Davatzikos |
Medical Image Anal. | 1 |
| 2009 | Efficient Large Deformation Registration via Geodesics on a Learned Manifold of Images
Jihun Hamm, Christos Davatzikos, Ragini Verma |
MICCAI (1) | 1 |
| 2008 | Grassmann discriminant analysis: a unifying view on subspace-based learningabstractIn this paper we propose a discriminant learning framework for problems in which data consist of linear subspaces instead of vectors. By treating subspaces as basic elements, we can make learning algorithms adapt naturally to the problems with linear invariant structures. We propose a unifying view on the subspace-based learning method by formulating the problems on the Grassmann manifold, which is the set of fixed-dimensional linear subspaces of a Euclidean space. Previous methods on the problem typically adopt an inconsistent strategy: feature extraction is performed in the Euclidean space while non-Euclidean distances are used. In our approach, we treat each sub-space as a point in the Grassmann space, and perform feature extraction and classification in the same space. We show feasibility of the approach by using the Grassmann kernel functions such as the Projection kernel and the Binet-Cauchy kernel. Experiments with real image databases show that the proposed method performs well compared with state-of-the-art algorithms. Jihun Hamm, Daniel D. Lee |
ICML | 1 |
| 2008 | Regularized discriminant analysis for transformation-invariant object recognitionabstractWe present a novel method for incorporating prior knowledge about invariances in object recognition for discriminant analysis. In contrast to conventional isotropic regularization approaches, our approach shows how to incorporate known transformation invariances in the geometry of the problem to better regularize discriminant analysis. In particular, we show how to incorporate group invariance and tangent vector structure with multiple parameters and derive special covariance terms that are used to regularize discriminant analysis. We apply this method to Fisher discriminant analysis, as well as its kernelized version, and show that this invariant regularization improves recognition performance over conventional regularization techniques. Yung-Kyun Noh, Jihun Hamm, Daniel D. Lee |
ICPR | 2 |
| 2008 | Extended Grassmann Kernels for Subspace-Based LearningabstractSubspace-based learning problems involve data whose elements are linear subspaces of a vector space. To handle such data structures, Grassmann kernels have been proposed and used previously. In this paper, we analyze the relationship between Grassmann kernels and probabilistic similarity measures. Firstly, we show that the KL distance in the limit yields the Projection kernel on the Grassmann manifold, whereas the Bhattacharyya kernel becomes trivial in the limit and is suboptimal for subspace-based problems. Secondly, based on our analysis of the KL distance, we propose extensions of the Projection kernel which can be extended to the set of affine as well as scaled subspaces. We demonstrate the advantages of these extended kernels for classification and recognition tasks with Support Vector Machines and Kernel Discriminant Analysis using synthetic and real image databases. Jihun Hamm, Daniel D. Lee |
NIPS | 1 |
| 2006 | Learning a manifold-constrained map between image sets: applications to matching and pose estimationabstractThis paper proposes a method for matching two sets of images given a small number of training examples by exploiting the underlying structure of the image manifolds. A nonlinear map from one manifold to another is constructed by combining linear maps locally defined on the tangent spaces of the manifolds. This construction imposes strong constraints on the choice of the maps, and makes possible good generalization of correspondences between all of the image sets. This map is flexible enough to approximate an arbitrary diffeomorphism between manifolds and can serve many purposes for applications. The underlying algorithm is a non-iterative efficient procedure whose complexity mainly depends on the number of matched training examples and the dimensionality of the manifold, and not on the number of samples nor on the dimensionality of the images. Several experiments were performed to demonstrate the potential of our method in image analysis and pose estimation. The first example demonstrates how images from a rotating camera can be mapped to the underlying pose manifold. Second, computer generated images from articulating toy figures are matched using the underlying 4 dimensional manifold to generate image-driven animations. Finally, two sets of actual lip images during speech are matched by their appearance manifold. In all these cases, our algorithm is able to obtain reasonable matches between thousands of large-dimensional images, with a minimum of computation. Jihun Hamm, Ikkjin Ahn, Daniel D. Lee |
CVPR (1) | 1 |
| 2005 | Learning nonlinear appearance manifolds for robot localizationabstractWe propose a nonlinear method for learning the low-dimensional pose of a robot from high-dimensional panoramic images. The panoramic images are assumed to lie on a nonlinear low-dimensional appearance manifold that is embedded in a high-dimensional image space. We demonstrate that the local geometry of a point and its nearest neighbors on this manifold can be used to project the point onto a low-dimensional coordinate space. Using this embedding, the unknown camera position can be estimated from a novel panoramic image. We show how the image-based position measurements can be integrated with odometry information in a Bayesian framework to yield an online estimate of a robot's position. Results from simulated data show that the proposed method outperforms other appearance-based models based upon principal components analysis and kernel density estimation. Jihun Hamm, Yuanqing Lin, Daniel D. Lee |
IROS | 1 |
| 2005 | Cooperative relative robot localization with audible acoustic sensingabstractWe describe a method for estimating the relative poses of a team of mobile robots using only acoustic sensing. The relative distances and bearing angles of the robots are estimated using the time of arrival of audible sound signals on stereo microphones. The robots emit specially designed sound waveforms that simultaneously enable robot identification and time of arrival estimation. These acoustic observations are then combined with odometry to update a belief state describing the positions and heading angles of all the robots. To efficiently resolve the ambiguity in the heading angle of the observing robot as well as the back-front ambiguity of the observed robot, we employ a Rao-Blackwellised particle filter (RBPF) where the distribution over heading angles is represented by a discrete set of particles, and the uncertainty in the translational positions conditioned on each of these particles is described by a Gaussian. This approach combines the representational accuracy of conventional particle filters with the efficiency of Kalman filter updates in modeling the pose distribution over a number of robots. We demonstrate how the RBPF can quickly resolve uncertainties in the binaural acoustic measurements and yield a globally consistent pose estimate. Simulations as well as an experimental implementation on robots with generic sound hardware illustrate the accuracy and the convergence of the resulting pose estimates. Yuanqing Lin, Paul Vernaza, Jihun Hamm, Daniel D. Lee |
IROS | 3 |
| 2004 | A kernel view of the dimensionality reduction of manifoldsabstractWe interpret several well-known algorithms for dimensionality reduction of manifolds as kernel methods. Isomap, graph Laplacian eigenmap, and locally linear embedding (LLE) all utilize local neighborhood information to construct a global embedding of the manifold. We show how all three algorithms can be described as kernel PCA on specially constructed Gram matrices, and illustrate the similarities and differences between the algorithms with representative examples. Jihun Hamm, Daniel D. Lee, Sebastian Mika, Bernhard Schölkopf |
ICML | 1 |