EDBT 2026 Demo / reviewers in the wild / expert
Jong Chul Ye
dblp:15/5613
· DBLP profile ↗
164ranked-venue papers
18as first author
107since 2021 · last 2026
0000-0001-9763-9609ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 75 · 2 first-author · 72 since 2021Graphics, computer vision, multimedia, augmented reality and games · 75 · 14 first-author · 43 since 2021Applied, interdisciplinary, general and emerging computing · 47 · 28 since 2021Theory of computation · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DreamMakeup: Face Makeup Customization using Latent Diffusion ModelsabstractThe exponential growth of the global makeup market has paralleled advancements in virtual makeup simulation technology. Despite the progress led by GANs, their application still encounters significant challenges, including training instability and limited customization capabilities. Addressing these challenges, we introduce DreamMakup – a novel training-free Diffusion model based Makeup Customization method, leveraging the inherent advantages of diffusion models for superior controllability and precise real-image editing. DreamMakeup employs early-stopped DDIM inversion to preserve the facial structure and identity while enabling extensive customization through various conditioning inputs such as reference images, specific RGB colors, and textual descriptions. Our model demonstrates notable improvements over existing GAN-based and recent diffusion-based frameworks – improved customization, color-matching capabil ities, identity preservation and compatibility with textual descriptions or LLMs with affordable computational costs. Geon Yeong Park, Inhwa Han, Serin Yang, Yeobin Hong, Seongmin Jeong, Heechan Jeon, Myeongjin Goh, Sung Won Yi, Jin Nam, Jong Chul Ye |
WACV | 10 |
| 2026 | Read like a radiologist: Efficient vision-language model for 3D medical imaging interpretation
Changsun Lee, Sangjoon Park, Cheong-Il Shin, Woo Hee Choi, Hyun Jeong Park, Jeong Eun Lee, Jong Chul Ye |
Medical Image Anal. | 7 |
| 2025 | Spectral Motion Alignment for Video Motion Transfer Using Diffusion ModelsabstractDiffusion models have significantly facilitated the customization of input video with target appearance while maintaining its motion patterns. To distill the motion information from video frames, existing works often estimate motion representations as frame difference or correlation in pixel-/feature-space. Despite its simplicity, these methods have unexplored limitations, including lack of understanding of global motion context, and the introduction of motion-independent spatial distortions. To address this, we present Spectral Motion Alignment (SMA), a novel framework that refines and aligns motion representations in the spectral domain. Specifically, SMA learns spectral motion representations, facilitating the learning of whole-frame global motion dynamics, and effectively mitigating motion-independent artifacts. Extensive experiments demonstrate SMA's efficacy in improving motion transfer while maintaining computational efficiency and compatibility across various video customization frameworks. Geon Yeong Park, Hyeonho Jeong, Sang Wan Lee, Jong Chul Ye |
AAAI | 4 |
| 2025 | Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video GenerationabstractWhile recent foundational video generators produce visually rich output, they still struggle with appearance drift, where objects gradually degrade or change inconsistently across frames, breaking visual coherence. We hypothesize that this is because there is no explicit supervision in terms of spatial tracking at the feature level. We propose Track4Gen, a spatially aware video generator that combines video diffusion loss with point tracking across frames, providing enhanced spatial supervision on the diffusion features. Track4Gen merges the video generation and point tracking tasks into a single network by making minimal changes to existing video generation architectures. Using Stable Video Diffusion [4] as a backbone, Track4Gen demonstrates that it is possible to unify video generation and point tracking, which are typically handled as separate tasks. Our extensive evaluations show that Track4Gen effectively reduces appearance drift, resulting in temporally stable and visually coherent video generation. Project page: hyeonho99.github.io/track4gen Hyeonho Jeong, Chun-Hao P. Huang, Jong Chul Ye, Niloy J. Mitra, Duygu Ceylan |
CVPR | 3 |
| 2025 | Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAIabstractGradient-based methods are a prototypical family of "explainability for AI" (XAI) techniques, especially for image-based models. However, they (1) require white-box access to models, (2) are vulnerable to adversarial attacks, and (3) produce attributions that lie off the image manifold, leading to explanations that are not amenable to human perception. To overcome these challenges, we introduce Derivative-Free Diffusion Manifold-Contrained Gradients (FreeMCG): by leveraging ensemble Kalman filters and diffusion models, we derive a derivative-free approximation of the model’s gradient projected onto the data manifold, requiring access only to the model’s outputs (i.e., black-box setting). We demonstrate the effectiveness of FreeMCG by applying it to both counterfactual generation and feature attribution, which have traditionally been treated as different tasks requiring distinct methods. Through comprehensive evaluation on both counter-factual explanation and feature attribution we show that our method yields state-of-the-art results for both tasks while preserving the essential properties expected of XAI tools. Code: https://github.com/one-june/FreeMCG. Won Jun Kim, Hyungjin Chung, Byeongsu Sim, Jong Chul Ye |
CVPR | 6 |
| 2025 | VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's GuideabstractText-to-image (T2I) diffusion models have revolutionized visual content creation, but extending these capabilities to text-to-video (T2V) generation remains a challenge, particularly in preserving temporal consistency. Existing methods that aim to improve consistency often cause trade-offs such as reduced imaging quality and impractical computational time. To address these issues we introduce VideoGuide, a novel framework that enhances the temporal consistency of pretrained T2V models without the need for additional training or fine-tuning. Instead, VideoGuide leverages any pretrained video diffusion model (VDM) or itself as a guide during the early stages of inference, improving temporal quality by interpolating the guiding model’s denoised samples into the sampling model’s denoising process. The proposed method brings about significant improvement in temporal consistency and image fidelity, providing a cost-effective and practical solution that synergizes the strengths of various video diffusion models. Furthermore, we demonstrate prior distillation, revealing that base models can achieve enhanced text coherence by utilizing the superior data prior of the guiding model through the proposed method. Project Page: https://dohunlee1.github.io/videoguide.github.io/ Bryan Sangwoo Kim, Geon Yeong Park, Jong Chul Ye |
CVPR | 4 |
| 2025 | Optical-Flow Guided Prompt Optimization for Coherent Video GenerationabstractWhile text-to-video diffusion models have made significant strides, many still face challenges in generating videos with temporal consistency. Within diffusion frameworks, guidance techniques have proven effective in enhancing output quality during inference; however, applying these methods to video diffusion models introduces additional complexity of handling computations across entire sequences.To address this, we propose a novel framework called MotionPrompt that guides the video generation process via optical flow. Specifically, we train a discriminator to distinguish optical flow between random pairs of frames from real videos and generated ones. Given that prompts can influence the entire video, we optimize learnable token embeddings during reverse sampling steps by using gradients from a trained discriminator applied to random frame pairs. This approach allows our method to generate visually coherent video sequences that closely reflect natural motion dynamics, without compromising the fidelity of the generated content. We demonstrate the effectiveness of our approach across various models. Project Page: https://motionprompt.github.io/ Hyelin Nam, Jong Chul Ye |
CVPR | 4 |
| 2025 | Minority-Focused Text-to-Image Generation via Prompt OptimizationabstractWe investigate the generation of minority samples using pretrained text-to-image (T2I) latent diffusion models. Minority instances, in the context of T2I generation, can be defined as ones living on low-density regions of text-conditional data distributions. They are valuable for various applications of modern T2I generators, such as data augmentation and creative AI. Unfortunately, existing pretrained T2I diffusion models primarily focus on high-density regions, largely due to the influence of guided samplers (like CFG) that are essential for high-quality generation. To address this, we present a novel framework to counter the high-density-focus of T2I diffusion models. Specifically, we first develop an online prompt optimization framework that encourages emergence of desired properties during inference while preserving semantic contents of user-provided prompts. We subsequently tailor this generic prompt optimizer into a specialized solver that promotes generation of minority features by incorporating a carefully-crafted likelihood objective. Extensive experiments conducted across various types of T2I models demonstrate that our approach significantly enhances the capability to produce high-quality minority instances compared to existing samplers. Code is available at https://github.com/soobin-um/MinorityPrompt. Soobin Um, Jong Chul Ye |
CVPR | 2 |
| 2025 | Reangle-A-Video: 4D Video Generation as Video-to-Video TranslationabstractWe introduce Reangle-A-Video, a unified framework for generating synchronized multi-view videos from a single input video. Unlike mainstream approaches that train multi-view video diffusion models on large-scale 4D datasets, our method reframes the multi-view video generation task as video-to-videos translation, leveraging publicly available image and video diffusion priors. In essence, Reangle-A-Video operates in two stages. (1) Multi-View Motion Learning: An image-to-video diffusion transformer is synchronously fine-tuned in a self-supervised manner to distill view-invariant motion from a set of warped videos. (2) Multi-View Consistent Image-to-Images Translation: The first frame of the input video is warped and inpainted into various camera perspectives under an inference-time cross-view consistency guidance using DUSt3R, generating multi-view consistent starting images. Extensive experiments on static view transport and dynamic camera control show that Reangle-A-Video surpasses existing methods, establishing a new solution for multi-view video generation. We will publicly release our code and data. Project page: https://hyeonho99.github.io/reangle-a-video/ Hyeonho Jeong, Suhyeon Lee 0004, Jong Chul Ye |
ICCV | 3 |
| 2025 | Free2 Guide: Training-Free Text-to-Video Alignment Using Image LVLM
Bryan Sangwoo Kim, Jong Chul Ye |
ICCV | 3 |
| 2025 | FlowDPS: Flow-Driven Posterior Sampling for Inverse ProblemsabstractFlow matching is a recent state-of-the-art framework for generative modeling based on ordinary differential equations (ODEs). While closely related to diffusion models, it provides a more general perspective on generative modeling. Although inverse problem solving has been extensively explored using diffusion models, it has not been rigorously examined within the broader context of flow models. Therefore, here we extend the diffusion inverse solvers (DIS) - which perform posterior sampling by combining a denoising diffusion prior with an likelihood gradient - into the flow framework. Specifically, by driving the flow-version of Tweedie's formula, we decompose the flow ODE into two components: one for clean image estimation and the other for noise estimation. By integrating the likelihood gradient and stochastic noise into each component, respectively, we demonstrate that posterior sampling for inverse problem solving can be effectively achieved using flows. Our proposed solver, Flow-Driven Posterior Sampling (FlowDPS), can also be seamlessly integrated into a latent flow model with a transformer architecture. Across four linear inverse problems, we confirm that FlowDPS outperforms state-of-the-art alternatives, all without requiring additional training. Jeongsol Kim, Bryan Sangwoo Kim, Jong Chul Ye |
ICCV | 3 |
| 2025 | VISION-XL: High Definition Video Inverse Problem Solver using Latent Image Diffusion ModelsabstractIn this paper, we propose a novel framework for solving high-definition video inverse problems using latent image diffusion models. Building on recent advancements in spatio-temporal optimization for video inverse problems using image diffusion models, our approach leverages latent-space diffusion models to achieve enhanced video quality and resolution. To address the high computational demands of processing high-resolution frames, we introduce a pseudo-batch consistent sampling strategy, allowing efficient operation on a single GPU. Additionally, to improve temporal consistency, we present pseudo-batch inversion, an initialization technique that incorporates informative latents from the measurement. By integrating with SDXL, our framework achieves state-of-the-art video reconstruction across a wide range of spatio-temporal inverse problems, including complex combinations of frame averaging and various spatial degradations, such as deblurring, super-resolution, and inpainting. Unlike previous methods, our approach supports multiple aspect ratios (landscape, vertical, and square) and delivers HD-resolution reconstructions (exceeding 1280x720) in under 6 seconds per frame on a single NVIDIA 4090 GPU. Taesung Kwon, Jong Chul Ye |
ICCV | 2 |
| 2025 | Inference-Time Diffusion Model DistillationabstractDiffusion distillation models effectively accelerate reverse sampling by compressing the process into fewer steps. However, these models still exhibit a performance gap compared to their pre-trained diffusion model counterparts, exacerbated by distribution shifts and accumulated errors during multi-step sampling. To address this, we introduce Distillation++, a novel inference-time distillation framework that reduces this gap by incorporating teacher-guided refinement during sampling. Inspired by recent advances in conditional sampling, our approach recasts student model sampling as a proximal optimization problem with a score distillation sampling loss (SDS). To this end, we integrate distillation optimization during reverse sampling, which can be viewed as teacher guidance that drives student sampling trajectory towards the clean manifold using pre-trained diffusion models. Thus, Distillation++ improves the denoising process in real-time without additional source data or fine-tuning. Distillation++ demonstrates substantial improvements over state-of-the-art distillation baselines, particularly in early sampling stages, positioning itself as a robust guided sampling process crafted for diffusion distillation models. Code: https://github.com/geonyeong-park/inference_distillation. Geon Yeong Park, Sang Wan Lee, Jong Chul Ye |
ICCV | 3 |
| 2025 | CFG++: Manifold-constrained Classifier Free Guidance for Diffusion ModelsabstractClassifier-free guidance (CFG) is a fundamental tool in modern diffusion models for text-guided generation. Although effective, CFG has notable drawbacks. For instance, DDIM with CFG lacks invertibility, complicating image editing; furthermore, high guidance scales, essential for high-quality outputs, frequently result in issues like mode collapse. Contrary to the widespread belief that these are inherent limitations of diffusion models, this paper reveals that the problems actually stem from the off-manifold phenomenon associated with CFG, rather than the diffusion models themselves. More specifically, inspired by the recent advancements of diffusion model-based inverse problem solvers (DIS), we reformulate text-guidance as an inverse problem with a text-conditioned score matching loss and develop CFG++, a novel approach that tackles the off-manifold challenges inherent in traditional CFG. CFG++ features a surprisingly simple fix to CFG, yet it offers significant improvements, including better sample quality for text-to-image generation, invertibility, smaller guidance scales, reduced etc. Furthermore, CFG++ enables seamless interpolation between unconditional and conditional sampling at lower guidance scales, consistently outperforming traditional CFG at all scales. Moreover, CFG++ can be easily integrated into the high-order diffusion solvers and naturally extends to distilled diffusion models. Experimental results confirm that our method significantly enhances performance in text-to-image generation, DDIM inversion, editing, and solving inverse problems, suggesting a wide-ranging impact and potential applications in various fields that utilize text guidance. Project Page: https://cfgpp-diffusion.github.io/anon Hyungjin Chung, Jeongsol Kim, Geon Yeong Park, Hyelin Nam, Jong Chul Ye |
ICLR | 5 |
| 2025 | Simple ReFlow: Improved Techniques for Fast Flow ModelsabstractDiffusion and flow-matching models achieve remarkable generative performance but at the cost of many neural function evaluations (NFE), which slows inference and limits applicability to time-critical tasks. The ReFlow procedure can accelerate sampling by straightening generation trajectories. But it is an iterative procedure, typically requiring training on simulated data, and results in reduced sample quality. To mitigate sample deterioration, we examine the design space of ReFlow and highlight potential pitfalls in prior heuristic practices. We then propose seven improvements for training dynamics, learning and inference, which are verified with thorough ablation studies on CIFAR10 $32 \times 32$, AFHQv2 $64 \times 64$, and FFHQ $64 \times 64$. Combining all our techniques, we achieve state-of-the-art FID scores (without / with guidance, resp.) for fast generation via neural ODEs: $2.23$ / $1.98$ on CIFAR10, $2.30$ / $1.91$ on AFHQv2, $2.84$ / $2.67$ on FFHQ, and $3.49$ / $1.74$ on ImageNet-64, all with merely $9$ NFEs. Yu-Guan Hsieh, Michal Klein, Marco Cuturi, Jong Chul Ye, Bahjat Kawar, James Thornton |
ICLR | 5 |
| 2025 | Generalized Consistency Trajectory Models for Image ManipulationabstractDiffusion-based generative models excel in unconditional generation, as well as on applied tasks such as image editing and restoration. The success of diffusion models lies in the iterative nature of diffusion: diffusion breaks down the complex process of mapping noise to data into a sequence of simple denoising tasks. Moreover, we are able to exert fine-grained control over the generation process by injecting guidance terms into each denoising step. However, the iterative process is also computationally intensive, often taking from tens up to thousands of function evaluations. Although consistency trajectory models (CTMs) enable traversal between any time points along the probability flow ODE (PFODE) and score inference with a single function evaluation, CTMs only allow translation from Gaussian noise to data. Thus, this work aims to unlock the full potential of CTMs by proposing generalized CTMs (GCTMs), which translate between arbitrary distributions via ODEs. We discuss the design space of GCTMs and demonstrate their efficacy in various image manipulation tasks such as image-to-image translation, restoration, and editing. Code is available at https://github.com/1202kbs/GCTM. Jeongsol Kim, Jong Chul Ye |
ICLR | 4 |
| 2025 | Regularization by Texts for Latent Diffusion Inverse SolversabstractThe recent development of diffusion models has led to significant progress in solving inverse problems by leveraging these models as powerful generative priors. However, challenges persist due to the ill-posed nature of such problems, often arising from ambiguities in measurements or intrinsic system symmetries. To address this, we introduce a novel latent diffusion inverse solver, regularization by text (TReg), inspired by the human ability to resolve visual ambiguities through perceptual biases. TReg integrates textual descriptions of preconceptions about the solution during reverse diffusion sampling, dynamically reinforcing these descriptions through null-text optimization, which we refer to as adaptive negation. Our comprehensive experimental results demonstrate that TReg effectively mitigates ambiguity in inverse problems, improving both accuracy and efficiency. Jeongsol Kim, Geon Yeong Park, Hyungjin Chung, Jong Chul Ye |
ICLR | 4 |
| 2025 | TweedieMix: Improving Multi-Concept Fusion for Diffusion-based Image/Video GenerationabstractDespite significant advancements in customizing text-to-image and video generation models, generating images and videos that effectively integrate multiple personalized concepts remains challenging. To address this, we present TweedieMix, a novel method for composing customized diffusion models during the inference phase. By analyzing the properties of reverse diffusion sampling, our approach divides the sampling process into two stages. During the initial steps, we apply a multiple object-aware sampling technique to ensure the inclusion of the desired target objects. In the later steps, we blend the appearances of the custom concepts in the de-noised image space using Tweedie's formula. Our results demonstrate that TweedieMix can generate multiple personalized concepts with higher fidelity than existing methods. Moreover, our framework can be effortlessly extended to image-to-video diffusion models by extending the residual layer's features across frames}, enabling the generation of videos that feature multiple personalized concepts. Gihyun Kwon, Jong Chul Ye |
ICLR | 2 |
| 2025 | Solving Video Inverse Problems Using Image Diffusion ModelsabstractRecently, diffusion model-based inverse problem solvers (DIS) have emerged as state-of-the-art approaches for addressing inverse problems, including image super-resolution, deblurring, inpainting, etc.
However, their application to video inverse problems arising from spatio-temporal degradation remains largely unexplored due to the challenges in training video diffusion models.
To address this issue, here we introduce an innovative video inverse solver that leverages only image diffusion models.
Specifically, by
drawing inspiration from the success of the recent decomposed diffusion sampler (DDS),
our method treats the time dimension of a video as the batch dimension of image diffusion models and solves spatio-temporal optimization problems within denoised spatio-temporal batches derived from each image diffusion model.
Moreover, we introduce a batch-consistent diffusion sampling strategy that encourages consistency across batches by synchronizing the stochastic noise components in image diffusion models.
Our approach synergistically combines batch-consistent sampling with simultaneous optimization of denoised spatio-temporal batches at each reverse diffusion step, resulting in a novel and efficient diffusion sampling strategy for video inverse problems.
Experimental results demonstrate that our method effectively addresses various spatio-temporal degradations in video inverse problems, achieving state-of-the-art reconstructions.
Project page: https://svi-diffusion.github.io/ Taesung Kwon, Jong Chul Ye |
ICLR | 2 |
| 2025 | ViBiDSampler: Enhancing Video Interpolation Using Bidirectional Diffusion SamplerabstractRecent progress in large-scale text-to-video (T2V) and image-to-video (I2V) diffusion models has greatly enhanced video generation, especially in terms of keyframe interpolation. However, current image-to-video diffusion models, while powerful in generating videos from a single conditioning frame, need adaptation for two-frame (start \& end) conditioned generation, which is essential for effective bounded interpolation. Unfortunately, existing approaches that fuse temporally forward and backward paths in parallel often suffer from off-manifold issues, leading to artifacts or requiring multiple iterative re-noising steps. In this work, we introduce a novel, bidirectional sampling strategy to address these off-manifold issues without requiring extensive re-noising or fine-tuning. Our method employs sequential sampling along both forward and backward paths, conditioned on the start and end frames, respectively, ensuring more coherent and on-manifold generation of intermediate frames. Additionally, we incorporate advanced guidance techniques, CFG++ and DDS, to further enhance the interpolation process. By integrating these, our method achieves state-of-the-art performance, efficiently generating high-quality, smooth videos between keyframes. On a single 3090 GPU, our method can interpolate 25 frames at 1024$\times$576 resolution in just 195 seconds, establishing it as a leading solution for keyframe interpolation.
Project page: https://vibidsampler.github.io/ Serin Yang, Taesung Kwon, Jong Chul Ye |
ICLR | 3 |
| 2025 | LDMol: A Text-to-Molecule Diffusion Model with Structurally Informative Latent Space Surpasses AR ModelsabstractWith the emergence of diffusion models as a frontline generative model, many researchers have proposed molecule generation techniques with conditional diffusion models. However, the unavoidable discreteness of a molecule makes it difficult for a diffusion model to connect raw data with highly complex conditions like natural language. To address this, here we present a novel latent diffusion model dubbed LDMol for text-conditioned molecule generation.
By recognizing that the suitable latent space design is the key to the diffusion model performance, we employ a contrastive learning strategy to extract novel feature space from text data that embeds the unique characteristics of the molecule structure.
Experiments show that LDMol outperforms the existing autoregressive baselines on the text-to-molecule generation benchmark, being one of the first diffusion models that outperforms autoregressive models in textual data generation with a better choice of the latent domain.
Furthermore, we show that LDMol can be applied to downstream tasks such as molecule-to-text retrieval and text-guided molecule editing, demonstrating its versatility as a diffusion model. Jinho Chang, Jong Chul Ye |
ICML | 2 |
| 2025 | Boost-and-Skip: A Simple Guidance-Free Diffusion for Minority GenerationabstractMinority samples are underrepresented instances located in low-density regions of a data manifold, and are valuable in many generative AI applications, such as data augmentation, creative content generation, etc. Unfortunately, existing diffusion-based minority generators often rely on computationally expensive guidance dedicated for minority generation. To address this, here we present a simple yet powerful guidance-free approach called Boost-and-Skip for generating minority samples using diffusion models. The key advantage of our framework requires only two minimal changes to standard generative processes: (i) variance-boosted initialization and (ii) timestep skipping. We highlight that these seemingly-trivial modifications are supported by solid theoretical and empirical evidence, thereby effectively promoting emergence of underrepresented minority features. Our comprehensive experiments demonstrate that Boost-and-Skip greatly enhances the capability of generating minority samples, even rivaling guidance-based state-of-the-art approaches while requiring significantly fewer computations. Code is available at https://github.com/soobin-um/BnS. Soobin Um, Jong Chul Ye |
ICML | 3 |
| 2025 | InvFusion: Bridging Supervised and Zero-shot Diffusion for Inverse ProblemsabstractDiffusion Models have demonstrated remarkable capabilities in handling inverse problems, offering high-quality posterior-sampling-based solutions. Despite significant advances, a fundamental trade-off persists regarding the way the conditioned synthesis is employed: Zero-shot approaches can accommodate any linear degradation but rely on approximations that reduce accuracy. In contrast, training-based methods model the posterior correctly, but cannot adapt to the degradation at test-time. Here we introduce InvFusion, the first training-based degradation-aware posterior sampler. InvFusion combines the best of both worlds - the strong performance of supervised approaches and the flexibility of zero-shot methods. This is achieved through a novel architectural design that seamlessly integrates the degradation operator directly into the diffusion denoiser. We compare InvFusion against existing general-purpose posterior samplers, both degradation-aware zero-shot techniques and blind training-based methods. Experiments on the FFHQ and ImageNet datasets demonstrate state-of-the-art performance. Beyond posterior sampling, we further demonstrate the applicability of our architecture, operating as a general Minimum Mean Square Error predictor, and as a Neural Posterior Principal Component estimator. Noam Elata, Hyungjin Chung, Jong Chul Ye, Tomer Michaeli, Miki Elad |
NeurIPS | 3 |
| 2025 | Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference AlignmentabstractModern single-image super-resolution (SISR) models deliver photo-realistic results at the scale factors on which they are trained, but collapse when asked to magnify far beyond that regime. We address this scalability bottleneck with Chain-of-Zoom (CoZ), a model-agnostic framework that factorizes SISR into an autoregressive chain of intermediate scale-states with multi-scale-aware prompts. CoZ repeatedly re-uses a backbone SR model, decomposing the conditional probability into tractable sub-problems to achieve extreme resolutions without additional training. Because visual cues diminish at high magnifications, we augment each zoom step with multi-scale-aware text prompts generated by a vision-language model (VLM). The prompt extractor itself is fine-tuned using Generalized Reward Policy Optimization (GRPO) with a critic VLM, aligning text guidance towards human preference. Experiments show that a standard $4\times$ diffusion SR model wrapped in CoZ attains beyond $256\times$ enlargement with high perceptual quality and fidelity. Bryan Sangwoo Kim, Jeongsol Kim, Jong Chul Ye |
NeurIPS | 3 |
| 2025 | Aligning Text to Image in Diffusion Models is Easier Than You ThinkabstractWhile recent advancements in generative modeling have significantly improved text-image alignment, some residual misalignment between text and image representations still remains. Some approaches address this issue by fine-tuning models in terms of preference optimization, etc., which require tailored datasets. Orthogonal to these methods, we revisit the challenge from the perspective of representation alignment—an approach that has gained popularity with the success of REPresentation Alignment (REPA). We first argue that conventional text-to-image (T2I) diffusion models, typically trained on paired image and text data (i.e., positive pairs) by minimizing score matching or flow matching losses, is suboptimal from the standpoint of representation alignment. Instead, a better alignment can be achieved through contrastive learning that leverages existing dataset as both positive and negative pairs. To enable efficient alignment with pretrained models, we propose SoftREPA—a lightweight contrastive fine-tuning strategy that leverages soft text tokens for representation alignment. This approach improves alignment with minimal computational overhead by adding fewer than 1M trainable parameters to the pretrained model. Our theoretical analysis demonstrates that our method explicitly increases the mutual information between text and image representations, leading to enhanced semantic consistency. Experimental results across text-to-image generation and text-guided image editing tasks validate the effectiveness of our approach in improving the semantic consistency of T2I generative models. Jaayeon Lee, Byunghee Cha, Jeongsol Kim, Jong Chul Ye |
NeurIPS | 4 |
| 2025 | Guided Diffusion Sampling on Function Spaces with Applications to PDEsabstractWe propose a general framework for conditional sampling in PDE-based inverse problems, targeting the recovery of whole solutions from extremely sparse or noisy measurements.
This is accomplished by a function-space diffusion model and plug-and-play guidance for conditioning.
Our method first trains an unconditional discretization-agnostic denoising model using neural operator architectures.
At inference, we refine the samples to satisfy sparse observation data via a gradient-based guidance mechanism.
Through rigorous mathematical analysis, we extend Tweedie's formula to infinite-dimensional Banach spaces, providing the theoretical foundation for our posterior sampling approach.
Our method (FunDPS) accurately captures posterior distribution in function spaces under minimal supervision and severe data scarcity. Across five PDE tasks with only 3\% observation, our method achieves an average 32\% accuracy improvement over state-of-the-art fixed-resolution diffusion baselines while reducing sampling steps by 4x. Furthermore, multi-resolution fine-tuning ensures strong cross-resolution generalizability and speedup. To the best of our knowledge, this is the first diffusion-based framework to operate independently of discretization, offering a practical and flexible solution for forward and inverse problems in the context of PDEs. Code is available at https://github.com/neuraloperator/FunDPS. Jiachen Yao, Abbas Mammadov, Julius Berner, Gavin Kerrigan, Jong Chul Ye, Kamyar Azizzadenesheli, Anima Anandkumar |
NeurIPS | 5 |
| 2025 | Text-to-Image Synthesis for Domain Generalization in Face Anti-SpoofingabstractThis paper addresses the challenge of developing robust Face Anti-Spoofing (FAS) models for face recognition systems. Traditional FAS protocols are limited by a lack of diversity in subject identities and environmental conditions, restricting generalization to real-world scenarios. Recent advancements in spoof image synthesis have mitigated data scarcity but still fail to capture the full range of facial attributes and environmental variability needed for effective domain generalization. To address this, we propose a novel framework capable of generating diverse, realistic facial images with text-guided control. We fine-tune Stable Diffusion to extract real facial features and specifically train LoRA layers to capture detailed spoof patterns. Addition-ally, the text-guided control of attributes helps overcome the lack of diversity seen in previous methods. Extensive experiments demonstrate that our text-to-image-based syn-thetic data generation significantly enhances the robustness of FAS models, establishing a new benchmark for domain-independent and reliable anti-spoofing systems. Naeun Ko, Yonghyun Jeong, Jong Chul Ye |
WACV | 3 |
| 2025 | End-to-end breast cancer radiotherapy planning via LMMs with consistency embeddingabstractRecent advances in AI foundation models have significant potential for lightening the clinical workload by mimicking the comprehensive and multi-faceted approaches used by medical professionals. In the field of radiation oncology, the integration of multiple modalities holds great importance, so the opportunity of foundational model is abundant. Inspired by this, here we present RO-LMM, a multi-purpose, comprehensive large multimodal model (LMM) tailored for the field of radiation oncology. This model effectively manages a series of tasks within the clinical workflow, including clinical context summarization, radiotherapy strategy suggestion, and plan-guided target volume segmentation by leveraging the capabilities of LMM. In particular, to perform consecutive clinical tasks without error accumulation, we present a novel Consistency Embedding Fine-Tuning (CEFTune) technique, which boosts LMM's robustness to noisy inputs while preserving the consistency of handling clean inputs. We further extend this concept to LMM-driven segmentation framework, leading to a novel Consistency Embedding Segmentation (CESEG) techniques. Experimental results including multi-center validation confirm that our RO-LMM with CEFTune and CESEG results in promising performance for multiple clinical tasks with generalization capabilities. Kwan-Young Kim, Yujin Oh, Sangjoon Park, Hwa Kyung Byun, Joongyo Lee, Yong Bae Kim, Jong Chul Ye |
Medical Image Anal. | 8 |
| 2025 | Video Diffusion Posterior Sampling for Seeing Beyond Dynamic Scattering LayersabstractImaging through scattering is challenging, as even a thin layer can randomly perturb light propagation and obscure hidden objects. Accurate closed-form modeling of forward scattering remains difficult, particularly for dynamically varying or thick layers. Here, we introduce a plug-and-play inverse solver based on video diffusion models with a physically grounded forward model tailored to dynamic scattering layers. Our method extends Diffusion Posterior Sampling (DPS) to the spatio-temporal domain, thereby capturing statistical correlations between video frames and scattered signals more effectively. Leveraging these temporal correlations, our approach recovers high-resolution spatial details that spatial-only methods typically fail to reconstruct. We also propose an inference-time optimization with a lightweight mapping network, enabling joint estimation of low-dimensional forward-model parameters without additional training. This joint optimization significantly enhances adaptability to unknown, time-varying degradations, making our method suitable for blind inverse scattering problems. We validate across diverse conditions, including different scene types, layer thicknesses, and scene-layer distances. And real-world experiments using multiple datasets confirm the robustness and effectiveness of our approach, even under real noise and forward-model approximation mismatches. Finally, we validate our method as a general video-restoration framework across dehazing, deblurring, inpainting, and blind restoration under complex optical aberrations. Taesung Kwon, Gookho Song, Yoosun Kim, Jeongsol Kim, Jong Chul Ye, Mooseok Jang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Text optimization with latent inversion for non-rigid image editing
Yunji Jung, Seokju Lee, Tair Djanibekov, Jong Chul Ye, Hyunjung Shim |
Pattern Recognit. Lett. | 4 |
| 2025 | Steerable Conditional Diffusion for Out-of-Distribution Adaptation in Medical Image ReconstructionabstractDenoising diffusion models have emerged as the go-to generative framework for solving inverse problems in imaging. A critical concern regarding these models is their performance on out-of-distribution tasks, which remains an under-explored challenge. Using a diffusion model on an out-of-distribution dataset, realistic reconstructions can be generated, but with hallucinating image features that are uniquely present in the training dataset. To address this discrepancy and improve reconstruction accuracy, we introduce a novel test-time adaptation sampling framework called Steerable Conditional Diffusion. Specifically, this framework adapts the diffusion model, concurrently with image reconstruction, based solely on the information provided by the available measurement. Utilising the proposed method, we achieve substantial enhancements in out-of-distribution performance across diverse imaging modalities, advancing the robust deployment of denoising diffusion models in real-world applications. Riccardo Barbano, Alexander Denker, Hyungjin Chung, Tae-Hoon Roh, Simon R. Arridge, Peter Maass, Bangti Jin, Jong Chul Ye |
IEEE Trans. Medical Imaging | 8 |
| 2025 | Guest Editorial Special Issue on Advancements in Foundation Models for Medical ImagingabstractPretrained on massive datasets, Foundation Models (FMs) are revolutionizing medical imaging by offering scalable and generalizable solutions to longstanding challenges. This Special Issue on Advancements in Foundation Models for Medical Imaging presents FM-related works that explore the potential of FMs to address data scarcity, domain shifts, and multimodal integration across a wide range of medical imaging tasks, including segmentation, diagnosis, reconstruction, and prognosis. The included papers also examine critical concerns such as interpretability, efficiency, benchmarking, and ethics in the adoption of FMs for medical imaging. Collectively, the articles in this Special Issue mark a significant step toward establishing FMs as a cornerstone of next-generation medical imaging AI. Tianming Liu 0001, Dinggang Shen, Jong Chul Ye, Marleen de Bruijne |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Patch-Wise Graph Contrastive Learning for Image TranslationabstractRecently, patch-wise contrastive learning is drawing attention for the image translation by exploring the semantic correspondence between the input image and the output image. To further explore the patch-wise topology for high-level semantic understanding, here we exploit the graph neural network to capture the topology-aware features. Specifically, we construct the graph based on the patch-wise similarity from a pretrained encoder, whose adjacency matrix is shared to enhance the consistency of patch-wise relation between the input and the output. Then, we obtain the node feature from the graph neural network, and enhance the correspondence between the nodes by increasing mutual information using the contrastive loss. In order to capture the hierarchical semantic structure, we further propose the graph pooling. Experimental results demonstrate the state-of-art results for the image translation thanks to the semantic encoding by the constructed graphs. Chanyong Jung, Gihyun Kwon, Jong Chul Ye |
AAAI | 3 |
| 2024 | VMC: Video Motion Customization Using Temporal Attention Adaption for Text-to-Video Diffusion ModelsabstractText-to-video diffusion models have advanced video generation significantly. However, customizing these models to generate videos with tailored motions presents a substantial challenge. In specific, they encounter hurdles in (a) accurately reproducing motion from a target video, and (b) creating diverse visual variations. For example, straight-forward extensions of static image customization methods to video often lead to intricate entanglements of appearance and motion data. To tackle this, here we present the Video Motion Customization (VMC) framework, a novel one-shot tuning approach crafted to adapt temporal attention layers within video diffusion models. Our approach introduces a novel motion distillation objective using residual vectors between consecutive noisy latent frames as a motion reference. The diffusion process then preserve low-frequency motion trajectories while mitigating high-frequency motion-unrelated noise in image space. We validate our method against state-of-the-art video generative models across diverse real-world motions and contexts. Our code and data can be found at https://video-motion-customization.github.io/. Hyeonho Jeong, Geon Yeong Park, Jong Chul Ye |
CVPR | 3 |
| 2024 | Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image ModelsabstractWhile there has been significant progress in customizing text-to-image generation models, generating images that combine multiple personalized concepts remains challenging. In this work, we introduce Concept Weaver, a method for composing customized text-to-image diffusion models at inference time. Specifically, the method breaks the process into two steps: creating a template image aligned with the semantics of input prompts, and then personalizing the template using a concept fusion strategy. The fusion strategy incorporates the appearance of the target concepts into the template image while retaining its structural details. The results indicate that our method can generate multiple custom concepts with higher identity fidelity compared to alternative approaches. Furthermore, the method is shown to seamlessly handle more than two concepts and closely follow the semantic meaning of the input prompt without blending appearances across different subjects. Gihyun Kwon, Simon Jenni, Dingzeyu Li, Joon-Young Lee, Jong Chul Ye, Fabian Caba Heilbron |
CVPR | 5 |
| 2024 | Contrastive Denoising Score for Text-Guided Latent Diffusion Image EditingabstractWith the remarkable advent of text-to-image diffusion models, image editing methods have become more diverse and continue to evolve. A promising recent approach in this realm is Delta Denoising Score (DDS) - an image editing technique based on Score Distillation Sampling (SDS) framework that leverages the rich generative prior of text-to-image diffusion models. However, relying solely on the difference between scoring functions is insufficient for preserving specific structural elements from the original image, a crucial aspect of image editing. To address this, here we present an embarrassingly simple yet very powerful modification of DDS, called Contrastive Denoising Score (CDS), for latent diffusion models (LDM). Inspired by the similarities and differences between DDS and the contrastive learning for unpaired image-to-image translation(CUT), we introduce a straightforward approach using CUT loss within the DDS framework. Rather than employing auxiliary networks as in the original CUT approach, we leverage the intermediate features of LDM, specifically those from the self-attention layers, which possesses rich spatial information. Our approach enables zero-shot image-to-image translation and neural radiance field (NeRF) editing, achieving structural correspondence between the input and output while maintaining content controllability. Qualitative results and comparisons demonstrates the effectiveness of our proposed method. Project page: https://hyelinnam.github.io/CDS/ Hyelin Nam, Gihyun Kwon, Geon Yeong Park, Jong Chul Ye |
CVPR | 4 |
| 2024 | Self-Supervised Debiasing Using Low Rank RegularizationabstractSpurious correlations can cause strong biases in deep neural networks, impairing generalization ability. While most existing debiasing methods require full supervision on either spurious attributes or target labels, training a debiased model from a limited amount of both annotations is still an open question. To address this issue, we investigate an interesting phenomenon using the spectral analysis of latent representations: spuriously correlated attributes make neural networks inductively biased towards encoding lower effective rank representations. We also show that a rank regularization can amplify this bias in a way that encourages highly correlated features. Leveraging these findings, we propose a self-supervised debiasing framework potentially compatible with unlabeled samples. Specifically, we first pretrain a biased encoder in a self-supervised manner with the rank regularization, serving as a semantic bottleneck to enforce the encoder to learn the spuriously correlated attributes. This biased encoder is then used to discover and upweight bias-conflicting samples in a downstream task, serving as a boosting to effectively debias the main model. Remarkably, the proposed debiasing framework significantly improves the generalization performance of self-supervised learning baselines and, in some cases, even outperforms state-of-the-art supervised debiasing approaches. Geon Yeong Park, Chanyong Jung, Sangmin Lee 0017, Jong Chul Ye, Sang Wan Lee |
CVPR | 4 |
| 2024 | Deep Diffusion Image Prior for Efficient OOD Adaptation in 3D Inverse Problems
Hyungjin Chung, Jong Chul Ye |
ECCV (75) | 2 |
| 2024 | DreamMotion: Space-Time Self-similar Score Distillation for Zero-Shot Video Editing
Hyeonho Jeong, Jinho Chang, Geon Yeong Park, Jong Chul Ye |
ECCV (30) | 4 |
| 2024 | OTSeg: Multi-Prompt Sinkhorn Attention for Zero-Shot Semantic Segmentation
Kwan-Young Kim, Yujin Oh, Jong Chul Ye |
ECCV (77) | 3 |
| 2024 | DreamSampler: Unifying Diffusion Sampling and Score Distillation for Image Manipulation
Jeongsol Kim, Geon Yeong Park, Jong Chul Ye |
ECCV (82) | 3 |
| 2024 | Self-Guided Generation of Minority Samples Using Diffusion Models
Soobin Um, Jong Chul Ye |
ECCV (68) | 2 |
| 2024 | Noise2one: One-Shot Image Denoising with Local Implicit LearningabstractRecently, the self-supervised learning paradigm, involving pretraining and fine-tuning large-scale models for downstream tasks, has shown promise in computer vision. Inspired by this, here we introduce Noise2One, a simple and effective image denoising method building upon this paradigm. Noise2One leverages self-supervised learning to pre-train a model on noisy images by learning a score function. Subsequently, fine-tuning on a single clean image enables denoising noisy images. Unlike Noise2Score, which relies on Tweedie’s formula, our method introduces a lightweight decoder based on Local Implicit Image Function (LIIF) for per-pixel noise adjustment and clean image reconstruction. This approach is more versatile, accommodating various noise models, including real-world noise. Our extensive experiments on benchmark datasets demonstrate that Noise2One achieves state-of-the-art denoising performance with only a 2.5% (+0.04M) increase in parameters compared to existing methods. Kwan-Young Kim, Jong Chul Ye |
ICASSP | 2 |
| 2024 | LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and GenerationabstractFollowing the impressive development of LLMs, vision-language alignment in LLMs is actively being researched to enable multimodal reasoning and visual input/output. This direction of research is particularly relevant to medical imaging because accurate medical image analysis and generation consist of a combination of reasoning based on visual features and prior knowledge. Many recent works have focused on training adapter networks that serve as an information bridge between image processing (encoding or generating) networks and LLMs; but presumably, in order to achieve maximum reasoning potential of LLMs on visual information as well, visual and language features should be allowed to interact more freely. This is especially important in the medical domain because understanding and generating medical images such as chest X-rays (CXR) require not only accurate visual and language-based reasoning but also a more intimate mapping between the two modalities. Thus, taking inspiration from previous work on the transformer and VQ-GAN combination for bidirectional image and text generation, we build upon this approach and develop a method for instruction-tuning an LLM pre-trained only on text to gain vision-language capabilities for medical images. Specifically, we leverage a pretrained LLM’s existing question-answering and instruction-following abilities to teach it to understand visual inputs by instructing it to answer questions about image inputs and, symmetrically, output both text and image responses appropriate to a given query by tuning the LLM with diverse tasks that encompass image-based text-generation and text-based image-generation. We show that our LLM-CXR trained in this approach shows better image-text alignment in both CXR understanding and generation tasks while being smaller in size compared to previously developed models that perform a narrower range of tasks. Suhyeon Lee 0004, Won Jun Kim, Jinho Chang, Jong Chul Ye |
ICLR | 4 |
| 2024 | Decomposed Diffusion Sampler for Accelerating Large-Scale Inverse ProblemsabstractKrylov subspace, which is generated by multiplying a given vector by the matrix of a linear transformation and its successive powers, has been extensively studied in classical optimization literature to design algorithms that converge quickly for large linear inverse problems. For example, the conjugate gradient method (CG), one of the most popular Krylov subspace methods, is based on the idea of minimizing the residual error in the Krylov subspace. However, with the recent advancement of high-performance diffusion solvers for inverse problems, it is not clear how classical wisdom can be synergistically combined with modern diffusion models. In this study, we propose a novel and efficient diffusion sampling strategy that synergistically combines the diffusion sampling and Krylov subspace methods. Specifically, we prove that if the tangent space at a denoised sample by Tweedie's formula forms a Krylov subspace, then the CG initialized with the denoised data ensures the data consistency update to remain in the tangent space. This negates the need to compute the manifold-constrained gradient (MCG), leading to a more efficient diffusion sampling method. Our method is applicable regardless of the parametrization and setting (i.e., VE, VP). Notably, we achieve state-of-the-art reconstruction quality on challenging real-world medical inverse imaging problems, including multi-coil MRI reconstruction and 3D CT reconstruction. Moreover, our proposed method achieves more than 80 times faster inference time than the previous state-of-the-art method. Code is available at https://github.com/HJ-harry/DDS Hyungjin Chung, Suhyeon Lee 0004, Jong Chul Ye |
ICLR | 3 |
| 2024 | Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion ModelsabstractThis paper introduces a novel grounding-guided video-to-video translation framework called Ground-A-Video for multi-attribute video editing.
Recent endeavors in video editing have showcased promising results in single-attribute editing or style transfer tasks, either by training T2V models on text-video data or adopting training-free methods.
However, when confronted with the complexities of multi-attribute editing scenarios, they exhibit shortcomings such as omitting or overlooking intended attribute changes, modifying the wrong elements of the input video, and failing to preserve regions of the input video that should remain intact.
Ground-A-Video attains temporally consistent multi-attribute editing of input videos in a training-free manner without aforementioned shortcomings.
Central to our method is the introduction of cross-frame gated attention which incorporates groundings information into the latent representations in a temporally consistent fashion, along with Modulated Cross-Attention and optical flow guided inverted latents smoothing.
Extensive experiments and applications demonstrate that Ground-A-Video's zero-shot capacity outperforms other baseline methods in terms of edit-accuracy and frame consistency.
Further results and code are available at our project page ( http://ground-a-video.github.io ) Hyeonho Jeong, Jong Chul Ye |
ICLR | 2 |
| 2024 | Unpaired Image-to-Image Translation via Neural Schrödinger BridgeabstractDiffusion models are a powerful class of generative models which simulate stochastic differential equations (SDEs) to generate data from noise. While diffusion models have achieved remarkable progress, they have limitations in unpaired image-to-image (I2I) translation tasks due to the Gaussian prior assumption. Schrödinger Bridge (SB), which learns an SDE to translate between two arbitrary distributions, have risen as an attractive solution to this problem. Yet, to our best knowledge, none of SB models so far have been successful at unpaired translation between high-resolution images. In this work, we propose Unpaired Neural Schrödinger Bridge (UNSB), which expresses the SB problem as a sequence of adversarial learning problems. This allows us to incorporate advanced discriminators and regularization to learn a SB between unpaired data. We show that UNSB is scalable and successfully solves various unpaired I2I translation tasks. Code: \url{https://github.com/cyclomon/UNSB} Gihyun Kwon, Kwan-Young Kim, Jong Chul Ye |
ICLR | 4 |
| 2024 | ED-NeRF: Efficient Text-Guided Editing of 3D Scene With Latent Space NeRFabstractRecently, there has been a significant advancement in text-to-image diffusion models, leading to groundbreaking performance in 2D image generation. These advancements have been extended to 3D models, enabling the generation of novel 3D objects from textual descriptions. This has evolved into NeRF editing methods, which allow the manipulation of existing 3D objects through textual conditioning. However, existing NeRF editing techniques have faced limitations in their performance due to slow training speeds and the use of loss functions that do not adequately consider editing. To address this, here we present a novel 3D NeRF editing approach dubbed ED-NeRF by successfully embedding real-world scenes into the latent space of the latent diffusion model (LDM) through a unique refinement layer. This approach enables us to obtain a NeRF backbone that is not only faster but also more amenable to editing compared to traditional image space NeRF editing. Furthermore, we propose an improved loss function tailored for editing by migrating the delta denoising score (DDS) distillation loss, originally used in 2D image editing to the three-dimensional domain. This novel loss function surpasses the well-known score distillation sampling (SDS) loss in terms of suitability for editing purposes. Our experimental results demonstrate that ED-NeRF achieves faster editing speed while producing improved output quality compared to state-of-the-art 3D editing models. Jangho Park, Gihyun Kwon, Jong Chul Ye |
ICLR | 3 |
| 2024 | Don't Play Favorites: Minority Guidance for Diffusion ModelsabstractWe explore the problem of generating minority samples using diffusion models. The minority samples are instances that lie on low-density regions of a data manifold. Generating a sufficient number of such minority instances is important, since they often contain some unique attributes of the data. However, the conventional generation process of the diffusion models mostly yields majority samples (that lie on high-density regions of the manifold) due to their high likelihoods, making themselves ineffective and time-consuming for the minority generating task. In this work, we present a novel framework that can make the generation process of the diffusion models focus on the minority samples. We first highlight that Tweedie's denoising formula yields favorable results for majority samples. The observation motivates us to introduce a metric that describes the uniqueness of a given sample. To address the inherent preference of the diffusion models w.r.t. the majority samples, we further develop *minority guidance*, a sampling technique that can guide the generation process toward regions with desired likelihood levels. Experiments on benchmark real datasets demonstrate that our minority guidance can greatly improve the capability of generating high-quality minority samples over existing generative samplers. We showcase that the performance benefit of our framework persists even in demanding real-world scenarios such as medical imaging, further underscoring the practical significance of our work. Code is available at https://github.com/soobin-um/minority-guidance. Soobin Um, Suhyeon Lee 0004, Jong Chul Ye |
ICLR | 3 |
| 2024 | Defining Neural Network Architecture through Polytope Structures of DatasetsabstractCurrent theoretical and empirical research in neural networks suggests that complex datasets require large network architectures for thorough classification, yet the precise nature of this relationship remains unclear. This paper tackles this issue by defining upper and lower bounds for neural network widths, which are informed by the polytope structure of the dataset in question. We also delve into the application of these principles to simplicial complexes and specific manifold shapes, explaining how the requirement for network width varies in accordance with the geometric complexity of the dataset. Moreover, we develop an algorithm to investigate a converse situation where the polytope structure of a dataset can be inferred from its corresponding trained neural networks. Through our algorithm, it is established that popular datasets such as MNIST, Fashion-MNIST, and CIFAR10 can be efficiently encapsulated using no more than two polytopes with a small number of faces. Sangmin Lee 0017, Abbas Mammadov, Jong Chul Ye |
ICML | 3 |
| 2024 | Prompt-tuning Latent Diffusion Models for Inverse ProblemsabstractWe propose a new method for solving imaging inverse problems using text-to-image latent diffusion models as general priors. Existing methods using latent diffusion models for inverse problems typically rely on simple null text prompts, which can lead to suboptimal performance. To improve upon this, we introduce a method for prompt tuning, which jointly optimizes the text embedding on-the-fly while running the reverse diffusion. This allows us to generate images that are more faithful to the diffusion prior. Specifically, our approach involves a unified optimization framework that simultaneously considers the prompt, latent, and pixel values through alternating minimization. This significantly diminishes image artifacts - a major problem when using latent diffusion models instead of pixel-based diffusion ones. Our method, called P2L, outperforms both pixel- and latent-diffusion model-based inverse problem solvers on a variety of tasks, such as super-resolution, deblurring, and inpainting. Furthermore, P2L demonstrates remarkable scalability to higher resolutions without artifacts. Hyungjin Chung, Jong Chul Ye, Peyman Milanfar, Mauricio Delbracio |
ICML | 2 |
| 2024 | C-DARL: Contrastive diffusion adversarial representation learning for label-free blood vessel segmentation
Boah Kim, Yujin Oh, Bradford J. Wood, Ronald M. Summers, Jong Chul Ye |
Medical Image Anal. | 5 |
| 2024 | Self-supervised multi-modal training from uncurated images and reports enables monitoring AI in radiology
Sangjoon Park, Eun Sun Lee, Kyung Sook Shin, Jeong Eun Lee, Jong Chul Ye |
Medical Image Anal. | 5 |
| 2024 | Magnitude and angle dynamics in training single ReLU neurons
Sangmin Lee 0017, Byeongsu Sim, Jong Chul Ye |
Neural Networks | 3 |
| 2024 | Improving Medical Speech-to-Text Accuracy using Vision-Language Pre-training ModelsabstractAutomatic Speech Recognition (ASR) is a technology that converts spoken words into text, facilitating interaction between humans and machines. One of the most common applications of ASR is Speech-To-Text (STT) technology, which simplifies user workflows by transcribing spoken words into text. In the medical field, STT has the potential to significantly reduce the workload of clinicians who rely on typists to transcribe their voice recordings. However, developing an STT model for the medical domain is challenging due to the lack of sufficient speech and text datasets. To address this issue, we propose a medical-domain text correction method that modifies the output text of a general STT system using the Vision Language Pre-training (VLP) method. VLP combines textual and visual information to correct text based on image knowledge. Our extensive experiments demonstrate that the proposed method offers quantitatively and clinically significant improvements in STT performance in the medical field. We further show that multi-modal understanding of image and text information outperforms single-modal understanding using only text information. Jaeyoung Huh, Sangjoon Park, Jeong Eun Lee, Jong Chul Ye |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Fundus Image Enhancement Through Direct Diffusion BridgesabstractWe propose FD3, a fundus image enhancement method based on direct diffusion bridges, which can cope with a wide range of complex degradations, including haze, blur, noise, and shadow. We first propose a synthetic forward model through a human feedback loop with board-certified ophthalmologists for maximal quality improvement of low-quality in-vivo images. Using the proposed forward model, we train a robust and flexible diffusion-based image enhancement network that is highly effective as a stand-alone method, unlike previous diffusion model-based approaches which act only as a refiner on top of pre-trained models. Through extensive experiments, we show that FD3 establishes superior quality not only on synthetic degradations but also on in vivo studies with low-quality fundus photos taken from patients with cataracts or small pupils. Sehui Kim, Hyungjin Chung, Se Hie Park, Eui-Sang Chung, Kayoung Yi, Jong Chul Ye |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | MS-DINO: Masked Self-Supervised Distributed Learning Using Vision TransformerabstractDespite promising advancements in deep learning in medical domains, challenges still remain owing to data scarcity, compounded by privacy concerns and data ownership disputes. Recent explorations of distributed-learning paradigms, particularly federated learning, have aimed to mitigate these challenges. However, these approaches are often encumbered by substantial communication and computational overhead, and potential vulnerabilities in privacy safeguards. Therefore, we propose a self-supervised masked sampling distillation technique called MS-DINO, tailored to the vision transformer architecture. This approach removes the need for incessant communication and strengthens privacy using a modified encryption mechanism inherent to the vision transformer while minimizing the computational burden on client-side devices. Rigorous evaluations across various tasks confirmed that our method outperforms existing self-supervised distributed learning strategies and fine-tuned baselines. Sangjoon Park, Ik-Jae Lee, Jun Won Kim, Jong Chul Ye |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Parallel Diffusion Models of Operator and Image for Blind Inverse ProblemsabstractDiffusion model-based inverse problem solvers have demonstrated state-of-the-art performance in cases where the forward operator is known (i.e. non-blind). However, the applicability of the method to blind inverse problems has yet to be explored. In this work, we show that we can indeed solve a family of blind inverse problems by constructing another diffusion prior for the forward operator. Specifically, parallel reverse diffusion guided by gradients from the intermediate stages enables joint optimization of both the forward operator parameters as well as the image, such that both are jointly estimated at the end of the parallel reverse diffusion procedure. We show the efficacy of our method on two representative tasks - blind deblurring, and imaging through turbulence - and show that our method yields state-of-the-art performance, while also being flexible to be applicable to general blind inverse problems when we know the functional forms. Code available: https://github.com/BlindDPS/blind-dps Hyungjin Chung, Jeongsol Kim, Sehui Kim, Jong Chul Ye |
CVPR | 4 |
| 2023 | Solving 3D Inverse Problems Using Pre-Trained 2D Diffusion ModelsabstractDiffusion models have emerged as the new state-of-the-art generative model with high quality samples, with intriguing properties such as mode coverage and high flexibility. They have also been shown to be effective inverse problem solvers, acting as the prior of the distribution, while the information of the forward model can be granted at the sampling stage. Nonetheless, as the generative process remains in the same high dimensional (i.e. identical to data dimension) space, the models have not been extended to 3D inverse problems due to the extremely high memory and computational cost. In this paper, we combine the ideas from the conventional model-based iterative reconstruction with the modern diffusion models, which leads to a highly effective method for solving 3D medical image reconstruction tasks such as sparse-view tomography, limited angle tomography, compressed sensing MRI from pre-trained 2D diffusion models. In essence, we propose to augment the 2D diffusion prior with a model-based prior in the remaining direction at test time, such that one can achieve coherent reconstructions across all dimensions. Our method can be run in a single commodity GPU, and establishes the new state-of-the-art, showing that the proposed method can perform reconstructions of high fidelity and accuracy even in the most extreme cases (e.g. 2-view 3D tomography). We further reveal that the generalization capacity of the proposed method is surprisingly high, and can be used to reconstruct volumes that are entirely different from the training dataset. Code available: https://github.com/HJ-harry/DiffusionMBIR Hyungjin Chung, Dohoon Ryu, Michael T. McCann, Marc Louis Klasky, Jong Chul Ye |
CVPR | 5 |
| 2023 | Training Debiased Subnetworks with Contrastive Weight PruningabstractNeural networks are often biased to spuriously correlated features that provide misleading statistical evidence that does not generalize. This raises an interesting question: “Does an optimal unbiased functional subnetwork exist in a severely biased network? If so, how to extract such subnetwork?” While empirical evidence has been accumulated about the existence of such unbiased subnetworks, these observations are mainly based on the guidance of ground-truth unbiased samples. Thus, it is unexplored how to discover the optimal subnetworks with biased training datasets in practice. To address this, here we first present our theoretical insight that alerts potential limitations of existing algorithms in exploring unbiased subnetworks in the presence of strong spurious correlations. We then further elucidate the importance of bias-conflicting samples on structure learning. Motivated by these observations, we propose a Debiased Contrastive Weight Pruning (DCWP) algorithm, which probes unbiased subnetworks without expensive group annotations. Experimental results demonstrate that our approach significantly outperforms state-of-the-art debiasing methods despite its considerable reduction in the number of parameters. Geon Yeong Park, Sangmin Lee 0017, Sang Wan Lee, Jong Chul Ye |
CVPR | 4 |
| 2023 | Ultrasound Image Quality Control Using Speech-Assisted Switchable CycleGANabstractUnlike computed tomography (CT) and magnetic resonance imaging (MRI) in which the image quality (IQ) is controlled by predefined acquisition setups, the IQ of ultrasound (US) is heavily dependent upon operators. In particular, an operator often adjusts the system parameters in a real-time manner based on his/her preference. Unfortunately, there are many cases where such real-time control of IQ is difficult, especially in the intensive care unit (ICU) or operating room (OR), since the operator should simultaneously treat patients in sterile status and adjust the system parameters. To address this, inspired by the recent success of Switchable CycleGAN using Adaptive Instance Normalization (AdaIN) layers, here we propose a novel speech-assisted Switchable CycleGAN architecture that can be controlled by operator’s verbal commands. Specifically, we employ a Speech Recognition Module (SRM) to generate AdaIN codes for Switchable CycleGAN. In particular, our SRM is based on the pre-trained model so that it can be less effected by the operator’s voice and robustly generate AdaIN codes. Various experimental results confirm that the proposed method can successfully control IQ of US by speech. Jaeyoung Huh, Shujaat Khan, Eun Sun Lee, Jong Chul Ye |
ICASSP | 4 |
| 2023 | Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion ModelsabstractDiffusion models have become a popular approach for image generation and reconstruction due to their numerous advantages. However, most diffusion-based inverse problem-solving methods only deal with 2D images, and even recently published 3D methods do not fully exploit the 3D distribution prior. To address this, we propose a novel approach using two perpendicular pre-trained 2D diffusion models to solve the 3D inverse problem. By modeling the 3D data distribution as a product of 2D distributions sliced in different directions, our method effectively addresses the curse of dimensionality. Our experimental results demonstrate that our method is highly effective for 3D medical image reconstruction tasks, including MRI Z-axis super-resolution, compressed sensing MRI, and sparse-view CT. Our method can generate high-quality voxel volumes suitable for medical applications. The code is available at https://github.com/hyn2028/tpdm Suhyeon Lee 0004, Hyungjin Chung, Jonghyuk Park 0006, Wi-Sun Ryu, Jong Chul Ye |
ICCV | 6 |
| 2023 | Zero-Shot Contrastive Loss for Text-Guided Diffusion Image Style TransferabstractDiffusion models have shown great promise in text-guided image style transfer, but there is a trade-off between style transformation and content preservation due to their stochastic nature. Existing methods require computationally expensive fine-tuning of diffusion models or additional neural network. To address this, here we propose a zero-shot contrastive loss for diffusion models that doesn’t require additional fine-tuning or auxiliary networks. By leveraging patch-wise contrastive loss between generated samples and original image embeddings in the pre-trained diffusion model, our method can generate images with the same semantic content as the source image in a zero-shot manner. Our approach outperforms existing methods while preserving content and requiring no additional training, not only for image style transfer but also for image-to-image translation and manipulation. Our experimental results validate the effectiveness of our proposed method. Code is available at https://github.com/YSerin/ZeCon. Serin Yang, Hyunmin Hwang, Jong Chul Ye |
ICCV | 3 |
| 2023 | Score-Based Diffusion Models for Bayesian Image ReconstructionabstractThis paper explores the use of score-based diffusion models for Bayesian image reconstruction. Diffusion models are an efficient tool for generative modeling. Diffusion models can also be used for solving image reconstruction problems. We present a simple and flexible algorithm for training a diffusion model and using it for maximum a posteriori reconstruction, minimum mean square error reconstruction, and posterior sampling. We present experiments on both a linear and a nonlinear reconstruction problem that highlight the strengths and limitations of the approach. Michael T. McCann, Hyungjin Chung, Jong Chul Ye, Marc Louis Klasky |
ICIP | 3 |
| 2023 | Diffusion Posterior Sampling for General Noisy Inverse Problems
Hyungjin Chung, Jeongsol Kim, Michael T. McCann, Marc Louis Klasky, Jong Chul Ye |
ICLR | 5 |
| 2023 | Diffusion Adversarial Representation Learning for Self-supervised Vessel Segmentation
Boah Kim, Yujin Oh, Jong Chul Ye |
ICLR | 3 |
| 2023 | Diffusion-based Image Translation using disentangled style and content representation
Gihyun Kwon, Jong Chul Ye |
ICLR | 2 |
| 2023 | Denoising MCMC for Accelerating Diffusion-Based Generative ModelsabstractThe sampling process of diffusion models can be interpreted as solving the reverse stochastic differential equation (SDE) or the ordinary differential equation (ODE) of the diffusion process, which often requires up to thousands of discretization steps to generate a single image. This has sparked a great interest in developing efficient integration techniques for reverse-S/ODEs. Here, we propose an orthogonal approach to accelerating score-based sampling: Denoising MCMC (DMCMC). DMCMC first uses MCMC to produce initialization points for reverse-S/ODE in the product space of data and diffusion time. Then, a reverse-S/ODE integrator is used to denoise the initialization points. Since MCMC traverses close to the data manifold, the cost of producing a clean sample for DMCMC is much less than that of producing a clean sample from noise. Denoising Langevin Gibbs, an instance of DMCMC, successfully accelerates all six reverse-S/ODE integrators considered in this work, and achieves state-of-the-art results: in the limited number of score function evaluation (NFE) setting on CIFAR10, we have $3.25$ FID with $\approx 10$ NFE and $2.49$ FID with $\approx 16$ NFE. On CelebA-HQ-256, we have $6.99$ FID with $\approx 160$ NFE, which beats the current best record of Kim et al. (2022) among score-based models, $7.16$ FID with $4000$ NFE. Code: https://github.com/1202kbs/DMCMC Jong Chul Ye |
ICML | 2 |
| 2023 | Minimizing Trajectory Curvature of ODE-based Generative ModelsabstractRecent ODE/SDE-based generative models, such as diffusion models, rectified flows, and flow matching, define a generative process as a time reversal of a fixed forward process. Even though these models show impressive performance on large-scale datasets, numerical simulation requires multiple evaluations of a neural network, leading to a slow sampling speed. We attribute the reason to the high curvature of the learned generative trajectories, as it is directly related to the truncation error of a numerical solver. Based on the relationship between the forward process and the curvature, here we present an efficient method of training the forward process to minimize the curvature of generative trajectories without any ODE/SDE simulation. Experiments show that our method achieves a lower curvature than previous models and, therefore, decreased sampling costs while maintaining competitive performance. Code is available at https://github.com/sangyun884/fast-ode. Jong Chul Ye |
ICML | 3 |
| 2023 | Direct Diffusion Bridge using Data Consistency for Inverse ProblemsabstractDiffusion model-based inverse problem solvers have shown impressive performance, but are limited in speed, mostly as they require reverse diffusion sampling starting from noise. Several recent works have tried to alleviate this problem by building a diffusion process, directly bridging the clean and the corrupted for specific inverse problems. In this paper, we first unify these existing works under the name Direct Diffusion Bridges (DDB), showing that while motivated by different theories, the resulting algorithms only differ in the choice of parameters. Then, we highlight a critical limitation of the current DDB framework, namely that it does not ensure data consistency. To address this problem, we propose a modified inference procedure that imposes data consistency without the need for fine-tuning. We term the resulting method data Consistent DDB (CDDB), which outperforms its inconsistent counterpart in terms of both perception and distortion metrics, thereby effectively pushing the Pareto-frontier toward the optimum. Our proposed method achieves state-of-the-art results on both evaluation criteria, showcasing its superiority over existing methods. Code is open-sourced [here](https://github.com/HJ-harry/CDDB). Hyungjin Chung, Jeongsol Kim, Jong Chul Ye |
NeurIPS | 3 |
| 2023 | Energy-Based Cross Attention for Bayesian Context Update in Text-to-Image Diffusion ModelsabstractDespite the remarkable performance of text-to-image diffusion models in image generation tasks, recent studies have raised the issue that generated images sometimes cannot capture the intended semantic contents of the text prompts, which phenomenon is often called semantic misalignment. To address this, here we present a novel energy-based model (EBM) framework for adaptive context control by modeling the posterior of context vectors. Specifically, we first formulate EBMs of latent image representations and text embeddings in each cross-attention layer of the denoising autoencoder. Then, we obtain the gradient of the log posterior of context vectors, which can be updated and transferred to the subsequent cross-attention layer, thereby implicitly minimizing a nested hierarchy of energy functions.
Our latent EBMs further allow zero-shot compositional generation as a linear combination of cross-attention outputs from different contexts.
Using extensive experiments, we demonstrate that the proposed method is highly effective in handling various image generation tasks, including multi-concept generation, text-guided image inpainting, and real and synthetic image editing. Code: https://github.com/EnergyAttention/Energy-Based-CrossAttention. Geon Yeong Park, Jeongsol Kim, Sang Wan Lee, Jong Chul Ye |
NeurIPS | 5 |
| 2023 | Tunable image quality control of 3-D ultrasound using switchable CycleGAN
Jaeyoung Huh, Shujaat Khan, Sungjin Choi, Dongkuk Shin, Jeong Eun Lee, Eun Sun Lee, Jong Chul Ye |
Medical Image Anal. | 7 |
| 2023 | One-Shot Adaptation of GAN in Just One CLIPabstractThere are many recent research efforts to fine-tune a pre-trained generator with a few target images to generate images of a novel domain. Unfortunately, these methods often suffer from overfitting or under-fitting when fine-tuned with a single target image. To address this, here we present a novel single-shot GAN adaptation method through unified CLIP space manipulations. Specifically, our model employs a two-step training strategy: reference image search in the source generator using a CLIP-guided latent optimization, followed by generator fine-tuning with a novel loss function that imposes CLIP space consistency between the source and adapted generators. To further improve the adapted model to produce spatially consistent samples with respect to the source generator, we also propose contrastive regularization for patchwise relationships in the CLIP space. Experimental results show that our model generates diverse outputs with the target texture and outperforms the baseline models both qualitatively and quantitatively. Furthermore, we show that our CLIP space manipulation strategy allows more effective attribute editing. Gihyun Kwon, Jong Chul Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Task-Agnostic Vision Transformer for Distributed Learning of Image ProcessingabstractRecently, distributed learning approaches have been studied for using data from multiple sources without sharing them, but they are not usually suitable in applications where each client carries out different tasks. Meanwhile, Transformer has been widely explored in computer vision area due to its capability to learn the common representation through global attention. By leveraging the advantages of Transformer, here we present a new distributed learning framework for multiple image processing tasks, allowing clients to learn distinct tasks with their local data. This arises from a disentangled representation of local and non-local features using a task-specific head/tail and a task-agnostic Vision Transformer. Each client learns a translation from its own task to a common representation using the task-specific networks, while the Transformer body on the server learns global attention between the features embedded in the representation. To enable decomposition between the task-specific and common representations, we propose an alternating training strategy between clients and server. Experimental results on distributed learning for various tasks show that our method synergistically improves the performance of each client with its own data. Boah Kim, Jeongsol Kim, Jong Chul Ye |
IEEE Trans. Image Process. | 3 |
| 2023 | Multi-Scale Hybrid Vision Transformer for Learning Gastric Histology: AI-Based Decision Support System for Gastric Cancer TreatmentabstractGastric endoscopic screening is an effective way to decide appropriate gastric cancer treatment at an early stage, reducing gastric cancer-associated mortality rate. Although artificial intelligence has brought a great promise to assist pathologist to screen digitalized endoscopic biopsies, existing artificial intelligence systems are limited to be utilized in planning gastric cancer treatment. We propose a practical artificial intelligence-based decision support system that enables five subclassifications of gastric cancer pathology, which can be directly matched to general gastric cancer treatment guidance. The proposed framework is designed to efficiently differentiate multi-classes of gastric cancer through multiscale self-attention mechanism using 2-stage hybrid vision transformer networks, by mimicking the way how human pathologists understand histology. The proposed system demonstrates its reliable diagnostic performance by achieving class-average sensitivity of above 0.85 for multicentric cohort tests. Moreover, the proposed system demonstrates its great generalization capability on gastrointestinal track organ cancer by achieving the best class-average sensitivity among contemporary networks. Furthermore, in the observational study, artificial intelligence-assisted pathologists show significantly improved diagnostic sensitivity within saved screening time compared to human pathologists. Our results demonstrate that the proposed artificial intelligence system has a great potential for providing presumptive pathologic opinion and supporting decision of appropriate gastric cancer treatment in practical clinical settings. Yujin Oh, Go Eun Bae, Kyung-Hee Kim, Min-Kyung Yeo, Jong Chul Ye |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | MR Image Denoising and Super-Resolution Using Regularized Reverse DiffusionabstractPatient scans from MRI often suffer from noise, which hampers the diagnostic capability of such images. As a method to mitigate such artifacts, denoising is largely studied both within the medical imaging community and beyond the community as a general subject. However, recent deep neural network-based approaches mostly rely on the minimum mean squared error (MMSE) estimates, which tend to produce a blurred output. Moreover, such models suffer when deployed in real-world situations: out-of-distribution data, and complex noise distributions that deviate from the usual parametric noise models. In this work, we propose a new denoising method based on score-based reverse diffusion sampling, which overcomes all the aforementioned drawbacks. Our network, trained only with coronal knee scans, excels even on out-of-distribution in vivo liver MRI data, contaminated with a complex mixture of noise. Even more, we propose a method to enhance the resolution of the denoised image with the same network. With extensive experiments, we show that our method establishes state-of-the-art performance while having desirable properties which prior MMSE denoisers did not have: flexibly choosing the extent of denoising, and quantifying uncertainty. Hyungjin Chung, Eun Sun Lee, Jong Chul Ye |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Multi-Task Distributed Learning Using Vision Transformer With Random Patch PermutationabstractThe widespread application of artificial intelligence in health research is currently hampered by limitations in data availability. Distributed learning methods such as federated learning (FL) and split learning (SL) are introduced to solve this problem as well as data management and ownership issues with their different strengths and weaknesses. The recent proposal of federated split task-agnostic (F eSTA) learning tries to reconcile the distinct merits of FL and SL by enabling the multi-task collaboration between participants through Vision Transformer (ViT) architecture, but they suffer from higher communication overhead. To address this, here we present a multi-task distributed learning using ViT with random patch permutation, dubbed p -F eSTA. Instead of using a CNN-based head as in F eSTA, p -F eSTA adopts a simple patch embedder with random permutation, improving the multi-task learning performance without sacrificing privacy. Experimental results confirm that the proposed method significantly enhances the benefit of multi-task collaboration, communication efficiency, and privacy preservation, shedding light on practical multi-task distributed learning in the field of medical imaging. Sangjoon Park, Jong Chul Ye |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Come-Closer-Diffuse-Faster: Accelerating Conditional Diffusion Models for Inverse Problems through Stochastic ContractionabstractDiffusion models have recently attained significant interest within the community owing to their strong performance as generative models. Furthermore, its application to inverse problems have demonstrated state-of-the-art performance. Unfortunately, diffusion models have a critical downside - they are inherently slow to sample from, needing few thousand steps of iteration to generate images from pure Gaussian noise. In this work, we show that starting from Gaussian noise is unnecessary. Instead, starting from a single forward diffusion with better initialization significantly reduces the number of sampling steps in the reverse conditional diffusion. This phenomenon is formally explained by the contraction theory of the stochastic difference equations like our conditional diffusion strategy - the alternating applications of reverse diffusion followed by a non-expansive data consistency step. The new sampling strategy, dubbed Come-Closer-Diffuse-Faster (CCDF), also reveals a new insight on how the existing feed-forward neural network approaches for inverse problems can be synergistically combined with the diffusion models. Experimental results with super-resolution, image inpainting, and compressed sensing MRI demonstrate that our method can achieve state-of-the-art reconstruction performance at significantly reduced sampling steps. Hyungjin Chung, Byeongsu Sim, Jong Chul Ye |
CVPR | 3 |
| 2022 | Exploring Patch-wise Semantic Relation for Contrastive Learning in Image-to-Image Translation TasksabstractRecently, contrastive learning-based image translation methods have been proposed, which contrasts different spatial locations to enhance the spatial correspondence. However, the methods often ignore the diverse semantic relation within the images. To address this, here we propose a novel semantic relation consistency (SRC) regularization along with the decoupled contrastive learning, which utilize the diverse semantics by focusing on the heterogeneous semantics between the image patches of a single image. To further improve the performance, we present a hard negative mining by exploiting the semantic relation. We verified our method for three tasks: single-modal and multi-modal image translations, and GAN compression task for image translation. Experimental results confirmed the state-of-art performance of our method in all the three tasks. Chanyong Jung, Gihyun Kwon, Jong Chul Ye |
CVPR | 3 |
| 2022 | Noise Distribution Adaptive Self-Supervised Image Denoising using Tweedie Distribution and Score MatchingabstractTweedie distributions are a special case of exponential dispersion models, which are often used in classical statistics as distributions for generalized linear models. Here, we show that Tweedie distributions also play key roles in modern deep learning era, leading to a distribution adaptive self-supervised image denoising formula without clean reference images. Specifically, by combining with the recent Noise2Score self-supervised image denoising approach and the saddle point approximation of Tweedie distribution, we provide a general closed-form denoising formula that can be used for large classes of noise distributions without ever knowing the underlying noise distribution. Similar to the original Noise2Score, the new approach is composed of two successive steps: score matching using perturbed noisy images, followed by a closed form image denoising formula via distribution-independent Tweedie's formula. In addition, we reveal a systematic algorithm to estimate the noise model and noise parameters for a given noisy image data set. Through extensive experiments, we demonstrate that the proposed method can accurately estimate noise models and parameters, and provide the state-of-the-art self-supervised image denoising performance in the benchmark dataset and real-world dataset. Kwan-Young Kim, Taesung Kwon, Jong Chul Ye |
CVPR | 3 |
| 2022 | DiffusionCLIP: Text-Guided Diffusion Models for Robust Image ManipulationabstractRecently, GAN inversion methods combined with Contrastive Language-Image Pretraining (CLIP) enables zeroshot image manipulation guided by text prompts. However, their applications to diverse real images are still difficult due to the limited GAN inversion capability. Specifically, these approaches often have difficulties in reconstructing images with novel poses, views, and highly variable contents compared to the training data, altering object identity, or producing unwanted image artifacts. To mitigate these problems and enable faithful manipulation of real images, we propose a novel method, dubbed DiffusionCLIP, that performs textdriven image manipulation using diffusion models. Based on full inversion capability and high-quality image generation power of recent diffusion models, our method performs zeroshot image manipulation successfully even between unseen domains and takes another step towards general application by manipulating images from a widely varying ImageNet dataset. Furthermore, we propose a novel noise combination method that allows straightforward multi-attribute manipulation. Extensive experiments and human evaluation confirmed robust and superior manipulation performance of our methods compared to the existing baselines. Code is available at https://github.com/gwang-kim/DiffusionCLIP.git Gwanghyun Kim, Taesung Kwon, Jong Chul Ye |
CVPR | 3 |
| 2022 | CLIPstyler: Image Style Transfer with a Single Text ConditionabstractExisting neural style transfer methods require reference style images to transfer texture information of style images to content images. However, in many practical situations, users may not have reference style images but still be inter-ested in transferring styles by just imagining them. In order to deal with such applications, we propose a new framework that enables a style transfer ‘without’ a style image, but only with a text description of the desired style. Using the pre-trained text-image embedding model of CLIP, we demonstrate the modulation of the style of content images only with a single text condition. Specifically, we propose a patch-wise text-image matching loss with multiview augmentations for realistic texture transfer. Extensive experimental results confirmed the successful image style transfer with realistic textures that reflect semantic query texts. Gihyun Kwon, Jong Chul Ye |
CVPR | 2 |
| 2022 | DiffuseMorph: Unsupervised Deformable Image Registration Using Diffusion Model
Boah Kim, Inhwa Han, Jong Chul Ye |
ECCV (31) | 3 |
| 2022 | CXR Segmentation by AdaIN-Based Domain Adaptation and Knowledge Distillation
Yujin Oh, Jong Chul Ye |
ECCV (21) | 2 |
| 2022 | Multi-Domain Unpaired Ultrasound Image Artifact Removal Using a Single Convolutional Neural NetworkabstractUltrasound imaging (US) often suffers from distinct image artifacts from various sources. Classic approaches for solving these problems are usually model-based iterative approaches that have been developed specifically for each type of artifact, which are often computationally intensive. Recently, deep learning approaches have been proposed as computationally efficient and high performance alternatives. Unfortunately, in the current deep learning approaches, a dedicated neural network should be trained with matched training data for each artifact type. This poses a fundamental limitation in the practical use of deep learning for US, since large number of paired data is required for supervised training of multiple models to deal with various US image artifacts. Inspired by the recent success of multi-domain image transfer, herein, we propose a novel unpaired deep learning approach where a single neural network can deal with different types of US artifacts simply by changing a mask vector that switches between different target domains. The proposed method can generate high quality images by removing distinct artifacts, which are comparable to those obtained by separately trained multiple neural networks. Jaeyoung Huh, Shujaat Khan, Jong Chul Ye |
ICASSP | 3 |
| 2022 | Patch-Wise Deep Metric Learning for Unsupervised Low-Dose CT Denoising
Chanyong Jung, Joonhyung Lee, Sunkyoung You, Jong Chul Ye |
MICCAI (6) | 4 |
| 2022 | Diffusion Deformable Model for 4D Temporal Medical Image Generation
Boah Kim, Jong Chul Ye |
MICCAI (1) | 2 |
| 2022 | Improving Diffusion Models for Inverse Problems using Manifold ConstraintsabstractRecently, diffusion models have been used to solve various inverse problems in an unsupervised manner with appropriate modifications to the sampling process. However, the current solvers, which recursively apply a reverse diffusion step followed by a projection-based measurement consistency step, often produce sub-optimal results. By studying the generative sampling path, here we show that current solvers throw the sample path off the data manifold, and hence the error accumulates. To address this, we propose an additional correction term inspired by the manifold constraint, which can be used synergistically with the previous solvers to make the iterations close to the manifold. The proposed manifold constraint is straightforward to implement within a few lines of code, yet boosts the performance by a surprisingly large margin. With extensive experiments, we show that our method is superior to the previous methods both theoretically and empirically, producing promising results in many applications such as image inpainting, colorization, and sparse-view computed tomography. Code available https://github.com/HJ-harry/MCG_diffusion Hyungjin Chung, Byeongsu Sim, Dohoon Ryu, Jong Chul Ye |
NeurIPS | 4 |
| 2022 | Energy-Based Contrastive Learning of Visual RepresentationsabstractContrastive learning is a method of learning visual representations by training Deep Neural Networks (DNNs) to increase the similarity between representations of positive pairs (transformations of the same image) and reduce the similarity between representations of negative pairs (transformations of different images). Here we explore Energy-Based Contrastive Learning (EBCLR) that leverages the power of generative learning by combining contrastive learning with Energy-Based Models (EBMs). EBCLR can be theoretically interpreted as learning the joint distribution of positive pairs, and it shows promising results on small and medium-scale datasets such as MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100. Specifically, we find EBCLR demonstrates from $\times 4$ up to $\times 20$ acceleration compared to SimCLR and MoCo v2 in terms of training epochs. Furthermore, in contrast to SimCLR, we observe EBCLR achieves nearly the same performance with $254$ negative pairs (batch size $128$) and $30$ negative pairs (batch size $16$) per positive pair, demonstrating the robustness of EBCLR to small numbers of negative pairs. Hence, EBCLR provides a novel avenue for improving contrastive learning methods that usually require large datasets with a significant number of negative pairs per iteration to achieve reasonable performance on downstream tasks. Code: https://github.com/1202kbs/EBCLR Jong Chul Ye |
NeurIPS | 2 |
| 2022 | Score-based diffusion models for accelerated MRI
Hyungjin Chung, Jong Chul Ye |
Medical Image Anal. | 2 |
| 2022 | Unsupervised resolution-agnostic quantitative susceptibility mapping using adaptive instance normalization
Gyutaek Oh, Hyokyoung Bae, Hyun-Seo Ahn, Sung-Hong Park, Won-Jin Moon, Jong Chul Ye |
Medical Image Anal. | 6 |
| 2022 | Multi-task vision transformer using low-level chest X-ray feature corpus for COVID-19 diagnosis and severity quantification
Sangjoon Park, Gwanghyun Kim, Yujin Oh, Joon Beom Seo, Sangmin Lee 0017, Jin Hwan Kim, Sungjun Moon, Jae-Kwang Lim, Jong Chul Ye |
Medical Image Anal. | 9 |
| 2022 | DeepPhaseCut: Deep Relaxation in Phase for Unsupervised Fourier Phase RetrievalabstractFourier phase retrieval is a classical problem of restoring a signal only from the measured magnitude of its Fourier transform. Although Fienup-type algorithms, which use prior knowledge in both spatial and Fourier domains, have been widely used in practice, they can often stall in local minima. Convex relaxation methods such as PhaseLift and PhaseCut may offer performance guarantees, but these algorithms are usually computationally expensive for practical use. To address this problem, here we propose a novel unsupervised feed-forward neural network for Fourier phase retrieval which generates high quality reconstruction immediately. Unlike the existing deep learning approaches that use a neural network as a regularization term or an end-to-end blackbox model for supervised training, our algorithm is a feed-forward neural network implementation of physics-driven constraints in an unsupervised learning framework. Specifically, our network is composed of two generators: one for the phase estimation using PhaseCut loss, followed by another generator for image reconstruction, all of which are trained simultaneously without matched data. The link to the classical Fienup-type algorithms and the recent symmetry-breaking learning approach is also revealed. Extensive experiments demonstrate that the proposed method outperforms all existing approaches in Fourier phase retrieval problems. Eun Ju Cha, Chanseok Lee, Mooseok Jang, Jong Chul Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Switchable and Tunable Deep Beamformer Using Adaptive Instance Normalization for Medical UltrasoundabstractRecent proposals of deep learning-based beamformers for ultrasound imaging (US) have attracted significant attention as computational efficient alternatives to adaptive and compressive beamformers. Moreover, deep beamformers are versatile in that image post-processing algorithms can be readily combined. Unfortunately, with the existing technology, a large number of beamformers need to be trained and stored for different probes, organs, depth ranges, operating frequency, and desired target 'styles', demanding significant resources such as training data, etc. To address this problem, here we propose a switchable and tunable deep beamformer that can switch between various types of outputs such as DAS, MVBF, DMAS, GCF, etc., and also adjust noise removal levels at the inference phase, by using a simple switch or tunable nozzle. This novel mechanism is implemented through Adaptive Instance Normalization (AdaIN) layers, so that distinct outputs can be generated using a single generator by merely changing the AdaIN codes. Experimental results using B-mode focused ultrasound confirm the flexibility and efficacy of the proposed method for various applications. Shujaat Khan, Jaeyoung Huh, Jong Chul Ye |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Diagonal Attention and Style-based GAN for Content-Style Disentanglement in Image Generation and TranslationabstractOne of the important research topics in image generative models is to disentangle the spatial contents and styles for their separate control. Although StyleGAN can generate content feature vectors from random noises, the resulting spatial content control is primarily intended for minor spatial variations, and the disentanglement of global content and styles is by no means complete. Inspired by a mathematical understanding of normalization and attention, here we present a novel hierarchical adaptive Diagonal spatial ATtention (DAT) layers to separately manipulate the spatial contents from styles in a hierarchical manner. Using DAT and AdaIN, our method enables coarse-to-fine level disentanglement of spatial contents and styles. In addition, our generator can be easily integrated into the GAN inversion framework so that the content and style of translated images from multi-domain image translation tasks can be flexibly controlled. By using various datasets, we confirm that the proposed method not only outperforms the existing models in disentanglement scores, but also provides more flexible control over spatial features in the generated images. Gihyun Kwon, Jong Chul Ye |
ICCV | 2 |
| 2021 | Noise2Score: Tweedie's Approach to Self-Supervised Image Denoising without Clean ImagesabstractRecently, there has been extensive research interest in training deep networks to denoise images without clean reference.However, the representative approaches such as Noise2Noise, Noise2Void, Stein's unbiased risk estimator (SURE), etc. seem to differ from one another and it is difficult to find the coherent mathematical structure. To address this, here we present a novel approach, called Noise2Score, which reveals a missing link in order to unite these seemingly different approaches.Specifically, we show that image denoising problems without clean images can be addressed by finding the mode of the posterior distribution and that the Tweedie's formula offers an explicit solution through the score function (i.e. the gradient of loglikelihood). Our method then uses the recent finding that the score function can be stably estimated from the noisy images using the amortized residual denoising autoencoder, the method of which is closely related to Noise2Noise or Nose2Void. Our Noise2Score approach is so universal that the same network training can be used to remove noises from images that are corrupted by any exponential family distributions and noise parameters. Using extensive experiments with Gaussian, Poisson, and Gamma noises, we show that Noise2Score significantly outperforms the state-of-the-art self-supervised denoising methods in the benchmark data set such as (C)BSD68, Set12, and Kodak, etc. Kwan-Young Kim, Jong Chul Ye |
NeurIPS | 2 |
| 2021 | Learning Dynamic Graph Representation of Brain Connectome with Spatio-Temporal AttentionabstractFunctional connectivity (FC) between regions of the brain can be assessed by the degree of temporal correlation measured with functional neuroimaging modalities. Based on the fact that these connectivities build a network, graph-based approaches for analyzing the brain connectome have provided insights into the functions of the human brain. The development of graph neural networks (GNNs) capable of learning representation from graph structured data has led to increased interest in learning the graph representation of the brain connectome. Although recent attempts to apply GNN to the FC network have shown promising results, there is still a common limitation that they usually do not incorporate the dynamic characteristics of the FC network which fluctuates over time. In addition, a few studies that have attempted to use dynamic FC as an input for the GNN reported a reduction in performance compared to static FC methods, and did not provide temporal explainability. Here, we propose STAGIN, a method for learning dynamic graph representation of the brain connectome with spatio-temporal attention. Specifically, a temporal sequence of brain graphs is input to the STAGIN to obtain the dynamic graph representation, while novel READOUT functions and the Transformer encoder provide spatial and temporal explainability with attention, respectively. Experiments on the HCP-Rest and the HCP-Task datasets demonstrate exceptional performance of our proposed method. Analysis of the spatio-temporal attention also provide concurrent interpretation with the neuroscientific knowledge, which further validates our method. Code is available at https://github.com/egyptdj/stagin Byung-Hoon Kim, Jong Chul Ye, Jae-Jin Kim |
NeurIPS | 2 |
| 2021 | Federated Split Task-Agnostic Vision Transformer for COVID-19 CXR DiagnosisabstractFederated learning, which shares the weights of the neural network across clients, is gaining attention in the healthcare sector as it enables training on a large corpus of decentralized data while maintaining data privacy. For example, this enables neural network training for COVID-19 diagnosis on chest X-ray (CXR) images without collecting patient CXR data across multiple hospitals. Unfortunately, the exchange of the weights quickly consumes the network bandwidth if highly expressive network architecture is employed. So-called split learning partially solves this problem by dividing a neural network into a client and a server part, so that the client part of the network takes up less extensive computation resources and bandwidth. However, it is not clear how to find the optimal split without sacrificing the overall network performance. To amalgamate these methods and thereby maximize their distinct strengths, here we show that the Vision Transformer, a recently developed deep learning architecture with straightforward decomposable configuration, is ideally suitable for split learning without sacrificing performance. Even under the non-independent and identically distributed data distribution which emulates a real collaboration between hospitals using CXR datasets from multiple sources, the proposed framework was able to attain performance comparable to data-centralized training. In addition, the proposed framework along with heterogeneous multi-task clients also improves individual task performances including the diagnosis of COVID-19, eliminating the need for sharing large weights with innumerable parameters. Our results affirm the suitability of Transformer for collaborative learning in medical imaging and pave the way forward for future real-world implementations. Sangjoon Park, Gwanghyun Kim, Jeongsol Kim, Boah Kim, Jong Chul Ye |
NeurIPS | 5 |
| 2021 | Two-stage deep learning for accelerated 3D time-of-flight MRA without matched training data
Hyungjin Chung, Eun Ju Cha, Leonard Sunwoo, Jong Chul Ye |
Medical Image Anal. | 4 |
| 2021 | CycleGAN denoising of extreme low-dose cardiac CT using wavelet-assisted noise disentanglement
Jawook Gu, Tae Seong Yang, Jong Chul Ye, Dong Hyun Yang |
Medical Image Anal. | 3 |
| 2021 | CycleMorph: Cycle consistent unsupervised deformable image registration
Boah Kim, Dong Hwan Kim, Seong Ho Park, June-Goo Lee, Jong Chul Ye |
Medical Image Anal. | 6 |
| 2021 | Unsupervised Denoising for Satellite Imagery Using Wavelet Directional CycleGANabstractMultispectral satellite imaging sensors acquire various spectral band images and have a unique spectroscopic property in each band. Unfortunately, image artifacts from imaging sensor noise often affect the quality of scenes and have a negative impact on applications for satellite imagery. Recently, deep learning approaches have been extensively explored to remove noise in satellite imagery. Most deep learning denoising methods, however, follow a supervised learning scheme, which requires matched noisy image and clean image pairs that are difficult to collect in real situations. In this article, we propose a novel unsupervised multispectral denoising method for satellite imagery using a wavelet directional cycle-consistent adversarial network (WavCycleGAN). The proposed method is based on an unsupervised learning scheme using adversarial loss and cycle-consistency loss to overcome the lack of paired data. Moreover, in contrast to the standard image-domain cycleGAN, we introduce a wavelet directional learning scheme for effective denoising without sacrificing high-frequency components such as edges and detailed information. Experimental results for the removal of vertical stripes and wave noise in satellite imaging sensors demonstrate that the proposed method effectively removes noise and preserves important high-frequency features of satellite images. Joonyoung Song, Dae-Soon Park, Hyun-Ho Kim, Jong Chul Ye |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | Unpaired Training of Deep Learning tMRA for Flexible Spatio-Temporal ResolutionabstractTime-resolved MR angiography (tMRA) has been widely used for dynamic contrast enhanced MRI (DCE-MRI) due to its highly accelerated acquisition. In tMRA, the periphery of the k -space data are sparsely sampled so that neighbouring frames can be merged to construct one temporal frame. However, this view-sharing scheme fundamentally limits the temporal resolution, and it is not possible to change the view-sharing number to achieve different spatio-temporal resolution trade-offs. Although many deep learning approaches have been recently proposed for MR reconstruction from sparse samples, the existing approaches usually require matched fully sampled k -space reference data for supervised training, which is not suitable for tMRA due to the lack of high spatio-temporal resolution ground-truth images. To address this problem, here we propose a novel unpaired training scheme for deep learning using optimal transport driven cycle-consistent generative adversarial network (cycleGAN). In contrast to the conventional cycleGAN with two pairs of generator and discriminator, the new architecture requires just a single pair of generator and discriminator, which makes the training much simpler but still improves the performance. Reconstruction results using in vivo tMRA and simulation data set confirm that the proposed method can immediately generate high quality reconstruction results at various choices of view-sharing numbers, allowing us to exploit better trade-off between spatial and temporal resolution in time-resolved MR angiography. Eun Ju Cha, Hyungjin Chung, Eung-Yeop Kim, Jong Chul Ye |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Unsupervised CT Metal Artifact Learning Using Attention-Guided β-CycleGANabstractMetal artifact reduction (MAR) is one of the most important research topics in computed tomography (CT). With the advance of deep learning approaches for image reconstruction, various deep learning methods have been suggested for metal artifact reduction, among which supervised learning methods are most popular. However, matched metal-artifact-free and metal artifact corrupted image pairs are difficult to obtain in real CT acquisition. Recently, a promising unsupervised learning for MAR was proposed using feature disentanglement, but the resulting network architecture is so complicated that it is difficult to handle large size clinical images. To address this, here we propose a simple and effective unsupervised learning method for MAR. The proposed method is based on a novel β -cycleGAN architecture derived from the optimal transport theory for appropriate feature space disentanglement. Moreover, by adding the convolutional block attention module (CBAM) layers in the generator, we show that the metal artifacts can be more focused so that it can be effectively removed. Experimental results confirm that we can achieve improved metal artifact reduction that preserves the detailed texture of the original image. Jawook Gu, Jong Chul Ye |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Unpaired MR Motion Artifact Deep Learning Using Outlier-Rejecting Bootstrap AggregationabstractRecently, deep learning approaches for MR motion artifact correction have been extensively studied. Although these approaches have shown high performance and lower computational complexity compared to classical methods, most of them require supervised training using paired artifact-free and artifact-corrupted images, which may prohibit its use in many important clinical applications. For example, transient severe motion (TSM) due to acute transient dyspnea in Gd-EOB-DTPA-enhanced MR is difficult to control and model for paired data generation. To address this issue, here we propose a novel unpaired deep learning scheme that does not require matched motion-free and motion artifact images. Specifically, the first step of our method is k -space random subsampling along the phase encoding direction that can remove some outliers probabilistically. In the second step, the neural network reconstructs fully sampled resolution image from a downsampled k -space data, and motion artifacts can be reduced in this step. Last, the aggregation step through averaging can further improve the results from the reconstruction network. We verify that our method can be applied for artifact correction from simulated motion as well as real motion from TSM successfully from both single and multi-coil data with and without k -space raw data, outperforming existing state-of-the-art deep learning methods. Gyutaek Oh, Jeong Eun Lee, Jong Chul Ye |
IEEE Trans. Medical Imaging | 3 |
| 2021 | DeepRegularizer: Rapid Resolution Enhancement of Tomographic Imaging Using Deep LearningabstractOptical diffraction tomography measures the three-dimensional refractive index map of a specimen and visualizes biochemical phenomena at the nanoscale in a non-destructive manner. One major drawback of optical diffraction tomography is poor axial resolution due to limited access to the three-dimensional optical transfer function. This missing cone problem has been addressed through regularization algorithms that use a priori information, such as non-negativity and sample smoothness. However, the iterative nature of these algorithms and their parameter dependency make real-time visualization impossible. In this article, we propose and experimentally demonstrate a deep neural network, which we term DeepRegularizer, that rapidly improves the resolution of a three-dimensional refractive index map. Trained with pairs of datasets (a raw refractive index tomogram and a resolution-enhanced refractive index tomogram via the iterative total variation algorithm), the three-dimensional U-net-based convolutional neural network learns a transformation between the two tomogram domains. The feasibility and generalizability of our network are demonstrated using bacterial cells and a human leukaemic cell line, and by validating the model across different samples. DeepRegularizer offers more than an order of magnitude faster regularization performance compared to the conventional iterative method. We envision that the proposed data-driven approach can bypass the high time complexity of various image reconstructions in other imaging modalities. DongHun Ryu, Dongmin Ryu, YoonSeok Baek, Hyungjoo Cho, Young Seo Kim, Yongki Lee, Yoosik Kim, Jong Chul Ye, Hyunseok Min, YongKeun Park |
IEEE Trans. Medical Imaging | 9 |
| 2021 | Continuous Conversion of CT Kernel Using Switchable CycleGAN With AdaINabstractX-ray computed tomography (CT) uses different filter kernels to highlight different structures. Since the raw sinogram data is usually removed after the reconstruction, in case there is additional need for other types of kernel images that were not previously generated, the patient may need to be scanned again. Accordingly, there exists increasing demand for post-hoc image domain conversion from one kernel to another without sacrificing the image quality. In this paper, we propose a novel unsupervised continuous kernel conversion method using cycle-consistent generative adversarial network (cycleGAN) with adaptive instance normalization (AdaIN). Even without paired training data, not only can our network translate the images between two different kernels, but it can also convert images along the interpolation path between the two kernel domains. We also show that the quality of generated images can be further improved if intermediate kernel domain images are available. Experimental results confirm that our method not only enables accurate kernel conversion that is comparable to supervised learning methods, but also generates intermediate kernel images in the unseen domain that are useful for hypopharyngeal cancer diagnosis. Serin Yang, Eung-Yeop Kim, Jong Chul Ye |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Optimal Transport Structure of CycleGAN for Unsupervised Learning for Inverse ProblemsabstractOptimal transport (OT) is a mathematical theory that can provide a tool how to transfer one measure to another measure at minimal cost, thus serve another framework for computer vision tasks of image processing without reference. Cycleconsistent generative adversarial network (cycleGAN) is a recent extension of GAN to learn target distributions with less mode collapsing behavior, and also does not need matched data during training. In this article, we explain the link between these two framework by mathematical formula and experimental results. We prove that cycleGAN architecture can be derived from optimal transport problem, and this implies that cycleGAN is a plausible way to learn target distribution when it comes to handling data far from the training set. Using accelerated MR imaging experiments, we confirmed the flexibility and efficacy of our theoretical framework. Byeongsu Sim, Gyutaek Oh, Jong Chul Ye |
ICASSP | 3 |
| 2020 | Image Reconstruction: From Sparsity to Data-Adaptive Methods and Machine LearningabstractThe field of medical image reconstruction has seen roughly four types of methods. The first type tended to be analytical methods, such as filtered backprojection (FBP) for X-ray computed tomography (CT) and the inverse Fourier transform for magnetic resonance imaging (MRI), based on simple mathematical models for the imaging systems. These methods are typically fast, but have suboptimal properties such as poor resolution-noise tradeoff for CT. A second type is iterative reconstruction methods based on more complete models for the imaging system physics and, where appropriate, models for the sensor statistics. These iterative methods improved image quality by reducing noise and artifacts. The U.S. Food and Drug Administration (FDA)-approved methods among these have been based on relatively simple regularization models. A third type of methods has been designed to accommodate modified data acquisition methods, such as reduced sampling in MRI and CT to reduce scan time or radiation dose. These methods typically involve mathematical image models involving assumptions such as sparsity or low rank. A fourth type of methods replaces mathematically designed models of signals and systems with data-driven or adaptive models inspired by the field of machine learning. This article focuses on the two most recent trends in medical image reconstruction: methods based on sparsity or low-rank models and data-driven methods based on machine learning techniques. Saiprasad Ravishankar, Jong Chul Ye, Jeffrey A. Fessler |
Proc. IEEE | 2 |
| 2020 | Optimal Transport Driven CycleGAN for Unsupervised Learning in Inverse ProblemsabstractTo improve the performance of classical generative adversarial networks (GANs), Wasserstein generative adversarial networks (WGANs) were developed as a Kantorovich dual formulation of the optimal transport (OT) problem using Wasserstein-1 distance. However, it was not clear how CycleGAN-type generative models can be derived from the OT theory. Here we show that a novel CycleGAN architecture can be derived as a Kantorovich dual OT formulation if a penalized least squares (PLS) cost with deep learning--based inverse path penalty is used as a transportation cost. One of the most important advantages of this formulation is that depending on the knowledge of the forward problem, distinct variations of CycleGAN architecture can be derived: for example, one with two pairs of generators and discriminators, and the other with only a single pair of generator and discriminator. Even for the two generator cases, we show that the structural knowledge of the forward operator can lead to a simpler generator architecture which significantly simplifies the neural network training. The new CycleGAN formulation, which we call the OT-CycleGAN, has been applied for various biomedical imaging problems, such as accelerated magnetic resonance imaging (MRI), super-resolution microscopy, and low-dose X-ray computed tomography (CT). Experimental results confirm the efficacy and flexibility of the theory. Byeongsu Sim, Gyutaek Oh, Jeongsol Kim, Chanyong Jung, Jong Chul Ye |
SIAM J. Imaging Sci. | 5 |
| 2020 | Mumford-Shah Loss Functional for Image Segmentation With Deep LearningabstractRecent state-of-the-art image segmentation algorithms are mostly based on deep neural networks, thanks to their high performance and fast computation time. However, these methods are usually trained in a supervised manner, which requires large number of high quality ground-truth segmentation masks. On the other hand, classical image segmentation approaches such as level-set methods are formulated in a self-supervised manner by minimizing energy functions such as Mumford-Shah functional, so they are still useful to help generation of segmentation masks without labels. Unfortunately, these algorithms are usually computationally expensive and often have limitation in semantic segmentation. In this paper, we propose a novel loss function based on Mumford-Shah functional that can be used in deep-learning based image segmentation without or with small labeled data. This loss function is based on the observation that the softmax layer of deep neural networks has striking similarity to the characteristic function in the Mumford-Shah functional. We show that the new loss function enables semi-supervised and unsupervised segmentation. In addition, our loss function can be also used as a regularized function to enhance supervised semantic segmentation algorithms. Experimental results on multiple datasets demonstrate the effectiveness of the proposed method. Boah Kim, Jong Chul Ye |
IEEE Trans. Image Process. | 2 |
| 2020 | Differentiated Backprojection Domain Deep Learning for Conebeam Artifact RemovalabstractConebeam CT using a circular trajectory is quite often used for various applications due to its relative simple geometry. For conebeam geometry, Feldkamp, Davis and Kress algorithm is regarded as the standard reconstruction method, but this algorithm suffers from so-called conebeam artifacts as the cone angle increases. Various model-based iterative reconstruction methods have been developed to reduce the cone-beam artifacts, but these algorithms usually require multiple applications of computational expensive forward and backprojections. In this paper, we develop a novel deep learning approach for accurate conebeam artifact removal. In particular, our deep network, designed on the differentiated backprojection domain, performs a data-driven inversion of an ill-posed deconvolution problem associated with the Hilbert transform. The reconstruction results along the coronal and sagittal directions are then combined using a spectral blending technique to minimize the spectral leakage. Experimental results under various conditions confirmed that our method generalizes well and outperforms the existing iterative methods despite significantly reduced runtime complexity. Yoseob Han, Junyoung Kim 0002, Jong Chul Ye |
IEEE Trans. Medical Imaging | 3 |
| 2020 | k-Space Deep Learning for Accelerated MRIabstractThe annihilating filter-based low-rank Hankel matrix approach (ALOHA) is one of the state-of-the-art compressed sensing approaches that directly interpolates the missing k -space data using low-rank Hankel matrix completion. The success of ALOHA is due to the concise signal representation in the k -space domain, thanks to the duality between structured low-rankness in the k -space domain and the image domain sparsity. Inspired by the recent mathematical discovery that links convolutional neural networks to Hankel matrix decomposition using data-driven framelet basis, here we propose a fully data-driven deep learning algorithm for k -space interpolation. Our network can be also easily applied to non-Cartesian k -space trajectories by simply adding an additional regridding layer. Extensive numerical experiments show that the proposed deep learning method consistently outperforms the existing image-domain deep learning approaches. Yoseob Han, Leonard Sunwoo, Jong Chul Ye |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Deep Learning COVID-19 Features on CXR Using Limited Training Data SetsabstractUnder the global pandemic of COVID-19, the use of artificial intelligence to analyze chest X-ray (CXR) image for COVID-19 diagnosis and patient triage is becoming important. Unfortunately, due to the emergent nature of the COVID-19 pandemic, a systematic collection of CXR data set for deep neural network training is difficult. To address this problem, here we propose a patch-based convolutional neural network approach with a relatively small number of trainable parameters for COVID-19 diagnosis. The proposed method is inspired by our statistical analysis of the potential imaging biomarkers of the CXR radiographs. Experimental results show that our method achieves state-of-the-art performance and provides clinically interpretable saliency maps, which are useful for COVID-19 diagnosis and patient triage. Yujin Oh, Sangjoon Park, Jong Chul Ye |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Deep Learning Diffuse Optical TomographyabstractDiffuse optical tomography (DOT) has been investigated as an alternative imaging modality for breast cancer detection thanks to its excellent contrast to hemoglobin oxidization level. However, due to the complicated non-linear photon scattering physics and ill-posedness, the conventional reconstruction algorithms are sensitive to imaging parameters such as boundary conditions. To address this, here we propose a novel deep learning approach that learns non-linear photon scattering physics and obtains an accurate three dimensional (3D) distribution of optical anomalies. In contrast to the traditional black-box deep learning approaches, our deep network is designed to invert the Lippman-Schwinger integral equation using the recent mathematical theory of deep convolutional framelets. As an example of clinical relevance, we applied the method to our prototype DOT system. We show that our deep neural network, trained with only simulation data, can accurately recover the location of anomalies within biomimetic phantoms and live animals without the use of an exogenous contrast agent. Jaejun Yoo 0001, Sohail Sabir, Duchang Heo, Kee Hyun Kim, Abdul Wahab 0002, Yoonseok Choi, Seul-I Lee, Eun Young Chae, Hak Hee Kim, Young Min Bae, Young-Wook Choi, Seungryong Cho, Jong Chul Ye |
IEEE Trans. Medical Imaging | 13 |
| 2019 | CollaGAN: Collaborative GAN for Missing Image Data ImputationabstractIn many applications requiring multiple inputs to obtain a desired output, if any of the input data is missing, it often introduces large amounts of bias. Although many techniques have been developed for imputing missing data, the image imputation is still difficult due to complicated nature of natural images. To address this problem, here we proposed a novel framework for missing image data imputation, called Collaborative Generative Adversarial Network (CollaGAN). CollaGAN convert the image imputation problem to a multi-domain images-to-image translation task so that a single generator and discriminator network can successfully estimate the missing data using the remaining clean data set. We demonstrate that CollaGAN produces the images with a higher visual quality compared to the existing competing approaches in various image imputation tasks. Dongwook Lee 0005, Junyoung Kim 0002, Won-Jin Moon, Jong Chul Ye |
CVPR | 4 |
| 2019 | Understanding Geometry of Encoder-Decoder CNNsabstractEncoder-decoder networks using convolutional neural network (CNN) architecture have been extensively used in deep learning literatures thanks to its excellent performance for various inverse problems in computer vision, medical imaging, etc. However, it is still difficult to obtain coherent geometric view why such an architecture gives the desired performance. Inspired by recent theoretical understanding on generalizability, expressivity and optimization landscape of neural networks, as well as the theory of convolutional framelets, here we provide a unified theoretical framework that leads to a better understanding of geometry of encoder-decoder CNNs. Our unified mathematical framework shows that encoder-decoder CNN architecture is closely related to nonlinear basis representation using combinatorial convolution frames, whose expressibility increases exponentially with the network depth. We also demonstrate the importance of skipped connection in terms of expressibility, and optimization landscape. Jong Chul Ye, Woon Kyoung Sung |
ICML | 1 |
| 2019 | Deep Learning-Based Universal Beamformer for Ultrasound Imaging
Shujaat Khan, Jaeyoung Huh, Jong Chul Ye |
MICCAI (5) | 3 |
| 2019 | Unsupervised Deformable Image Registration Using Cycle-Consistent CNN
Boah Kim, June-Goo Lee, Dong Hwan Kim, Seong Ho Park, Jong Chul Ye |
MICCAI (6) | 6 |
| 2019 | Efficient B-Mode Ultrasound Image Reconstruction From Sub-Sampled RF Data Using Deep LearningabstractIn portable, 3-D, and ultra-fast ultrasound imaging systems, there is an increasing demand for the reconstruction of high-quality images from a limited number of radio-frequency (RF) measurements due to receiver (Rx) or transmit (Xmit) event sub-sampling. However, due to the presence of side lobe artifacts from RF sub-sampling, the standard beamformer often produces blurry images with less contrast, which are unsuitable for diagnostic purposes. Existing compressed sensing approaches often require either hardware changes or computationally expensive algorithms, but their quality improvements are limited. To address this problem, in this paper, we propose a novel deep learning approach that directly interpolates the missing RF data by utilizing redundancy in the Rx-Xmit plane. Our extensive experimental results using sub-sampled RF data from a multi-line acquisition B-mode system confirm that the proposed method can effectively reduce the data rate without sacrificing the image quality. Yeo Hun Yoon, Shujaat Khan, Jaeyoung Huh, Jong Chul Ye |
IEEE Trans. Medical Imaging | 4 |
| 2018 | Deep Learning for Accelerated Ultrasound ImagingabstractIn portable, 3-D, or ultra-fast ultrasound (US) imaging systems, there is an increasing demand to reconstruct high quality images from limited number of data. However, the existing solutions require either hardware changes or computationally expansive algorithms. To overcome these limitations, here we propose a novel deep learning approach that interpolates the missing RF data by utilizing the sparsity of the RF data in the Fourier domain. Extensive experimental results from sub-sampled RF data from a real US system confirmed that the proposed method can effectively reduce the data rate without sacrificing the image quality. Yeo Hun Yoon, Jong Chul Ye |
ICASSP | 2 |
| 2018 | Deep Convolutional Framelets: A General Deep Learning Framework for Inverse ProblemsabstractRecently, deep learning approaches with various network architectures have achieved significant performance improvement over existing iterative reconstruction methods in various imaging problems. However, it is still unclear why these deep learning architectures work for specific inverse problems. Moreover, in contrast to the usual evolution of signal processing theory around the classical theories, the link between deep learning and the classical signal processing approaches, such as wavelets, nonlocal processing, and compressed sensing, are not yet well understood. To address these issues, here we show that the long-sought missing link is the convolution framelets for representing a signal by convolving local and nonlocal bases. The convolution framelets were originally developed to generalize the theory of low-rank Hankel matrix approaches for inverse problems, and this paper further extends this idea so as to obtain a deep neural network using multilayer convolution framelets with perfect reconstruction (PR) under rectified linear unit (ReLU) nonlinearity. Our analysis also shows that the popular deep network components such as residual blocks, redundant filter channels, and concatenated ReLU (CReLU) do indeed help to achieve PR, while the pooling and unpooling layers should be augmented with high-pass branches to meet the PR condition. Moreover, by changing the number of filter channels and bias, we can control the shrinkage behaviors of the neural network. This discovery reveals the limitations of many existing deep learning architectures for inverse problems, and leads us to propose a novel theory for a deep convolutional framelet neural network. Using numerical experiments with various inverse problems, we demonstrate that our deep convolutional framelets network shows consistent improvement over existing deep architectures. This discovery suggests that the success of deep learning stems not from a magical black box, but rather from the power of a novel signal representation using a nonlocal basis combined with a data-driven local basis, which is indeed a natural extension of classical signal processing theory. Jong Chul Ye, Yoseob Han, Eun Ju Cha |
SIAM J. Imaging Sci. | 1 |
| 2018 | Sparse and Low-Rank Decomposition of a Hankel Structured Matrix for Impulse Noise RemovalabstractRecently, the annihilating filter-based low-rank Hankel matrix (ALOHA) approach was proposed as a powerful image inpainting method. Based on the observation that smoothness or textures within an image patch correspond to sparse spectral components in the frequency domain, ALOHA exploits the existence of annihilating filters and the associated rank-deficient Hankel matrices in an image domain to estimate any missing pixels. By extending this idea, we propose a novel impulse-noise removal algorithm that uses the sparse and low-rank decomposition of a Hankel structured matrix. This method, referred to as the robust ALOHA, is based on the observation that an image corrupted with the impulse noise has intact pixels; consequently, the impulse noise can be modeled as sparse components, whereas the underlying image can still be modeled using a low-rank Hankel structured matrix. To solve the sparse and low-rank matrix decomposition problem, we propose an alternating direction method of multiplier approach, with initial factorized matrices coming from a low-rank matrix-fitting algorithm. To adapt local image statistics that have distinct spectral distributions, the robust ALOHA is applied in a patch-by-patch manner. Experimental results from impulse noise for both single-channel and multichannel color images demonstrate that the robust ALOHA is superior to existing approaches, especially during the reconstruction of complex texture patterns. Kyong Hwan Jin, Jong Chul Ye |
IEEE Trans. Image Process. | 2 |
| 2018 | Grid-Free Localization Algorithm Using Low-Rank Hankel Matrix for Super-Resolution MicroscopyabstractLocalization microscopy, such as STORM / PALM, can reconstruct super-resolution images with a nanometer resolution through the iterative localization of fluorescence molecules. Recent studies in this area have focused mainly on the localization of densely activated molecules to improve temporal resolutions. However, higher density imaging requires an advanced algorithm that can resolve closely spaced molecules. Accordingly, sparsitydriven methods have been studied extensively. One of the major limitations of existing sparsity-driven approaches is the need for a fine sampling grid or for Taylor series approximation which may result in some degree of localization bias toward the grid. In addition, prior knowledge of the point-spread function (PSF) is required. To address these drawbacks, here we propose a true grid-free localization algorithm with adaptive PSF estimation. Specifically, based on the observation that sparsity in the spatial domain implies a low rank in the Fourier domain, the proposed method converts source localization problems into Fourier-domain signal processing problems so that a truly gridfree localization is possible. We verify the performance of the newly proposed method with several numerical simulations and a live-cell imaging experiment. Junhong Min, Kyong Hwan Jin, Michael Unser, Jong Chul Ye |
IEEE Trans. Image Process. | 4 |
| 2018 | Unified Theory for Recovery of Sparse Signals in a General Transform DomainabstractCompressed sensing is provided a data-acquisition paradigm for sparse signals. Remarkably, it has been shown that the practical algorithms provide robust recovery from noisy linear measurements acquired at a near optimal sampling rate. In many real-world applications, a signal of interest is typically sparse not in the canonical basis but in a certain transform domain, such as wavelets or the finite difference. The theory of compressed sensing was extended to the analysis sparsity model, but known extensions are limited to the specific choices of sensing matrix and sparsifying transform. In this paper, we propose a unified theory for robust recovery of sparse signals in a general transform domain by convex programming. In particular, our results apply to the general acquisition and sparsity models and show how the number of measurements for recovery depends on properties of measurement and sparsifying transforms. Moreover, we also provide extensions of our results to the scenarios where the atoms in the transform have varying incoherence parameters and the unknown signal exhibits a structured sparsity pattern. In particular, for the partial Fourier recovery of sparse signals over a circulant transform, our main results suggest a uniformly random sampling. Numerical results demonstrate that the variable density random sampling by our main results provides a superior recovery performance over the known sampling strategies. Kiryung Lee, Yanjun Li 0001, Kyong Hwan Jin, Jong Chul Ye |
IEEE Trans. Inf. Theory | 4 |
| 2018 | Framing U-Net via Deep Convolutional Framelets: Application to Sparse-View CTabstractX-ray computed tomography (CT) using sparse projection views is a recent approach to reduce the radiation dose. However, due to the insufficient projection views, an analytic reconstruction approach using the filtered back projection (FBP) produces severe streaking artifacts. Recently, deep learning approaches using large receptive field neural networks such as U-Net have demonstrated impressive performance for sparse-view CT reconstruction. However, theoretical justification is still lacking. Inspired by the recent theory of deep convolutional framelets, the main goal of this paper is, therefore, to reveal the limitation of U-Net and propose new multi-resolution deep learning schemes. In particular, we show that the alternative U-Net variants such as dual frame and tight frame U-Nets satisfy the so-called frame condition which makes them better for effective recovery of high frequency edges in sparse-view CT. Using extensive experiments with real patient data set, we demonstrate that the new network architectures provide better reconstruction performance. Yoseob Han, Jong Chul Ye |
IEEE Trans. Medical Imaging | 2 |
| 2018 | Deep Convolutional Framelet Denosing for Low-Dose CT via Wavelet Residual NetworkabstractModel-based iterative reconstruction algorithms for low-dose X-ray computed tomography (CT) are computationally expensive. To address this problem, we recently proposed a deep convolutional neural network (CNN) for low-dose X-ray CT and won the second place in 2016 AAPM Low-Dose CT Grand Challenge. However, some of the textures were not fully recovered. To address this problem, here we propose a novel framelet-based denoising algorithm using wavelet residual network which synergistically combines the expressive power of deep learning and the performance guarantee from the framelet-based denoising algorithms. The new algorithms were inspired by the recent interpretation of the deep CNN as a cascaded convolution framelet signal representation. Extensive experimental results confirm that the proposed networks have significantly improved performance and preserve the detail texture of the original images. Eunhee Kang, Won Chang, Jaejun Yoo 0001, Jong Chul Ye |
IEEE Trans. Medical Imaging | 4 |
| 2018 | Image Reconstruction is a New Frontier of Machine LearningabstractOver past several years, machine learning, or more generally artificial intelligence, has generated overwhelming research interest and attracted unprecedented public attention. As tomographic imaging researchers, we share the excitement from our imaging perspective [item 1) in the Appendix], and organized this special issue dedicated to the theme of "Machine learning for image reconstruction." This special issue is a sister issue of the special issue published in May 2016 of this journal with the theme "Deep learning in medical imaging" [item 2) in the Appendix]. While the previous special issue targeted medical image processing/analysis, this special issue focuses on data-driven tomographic reconstruction. These two special issues are highly complementary, since image reconstruction and image analysis are two of the main pillars for medical imaging. Together we cover the whole workflow of medical imaging: from tomographic raw data/features to reconstructed images and then extracted diagnostic features/readings. Ge Wang 0001, Jong Chul Ye, Klaus Mueller 0001, Jeffrey A. Fessler |
IEEE Trans. Medical Imaging | 2 |
| 2017 | A Joint Sparse Recovery Framework for Accurate Reconstruction of Inclusions in Elastic MediaabstractA robust algorithm is proposed to reconstruct the spatial support and the Lamé parameters of multiple inclusions in a homogeneous background elastic material using a few measurements of the displacement field over a finite collection of boundary points. The algorithm does not require any linearization or iterative update of Green's function but still allows very accurate reconstruction. The breakthrough comes from a novel interpretation of Lippmann--Schwinger type integral representation of the displacement field in terms of unknown densities having common sparse support on the location of inclusions. Accordingly, the proposed algorithm consists of a two-step approach. First, the localization problem is recast as a joint sparse recovery problem that renders the densities and the inclusion support simultaneously. Then, a noise robust constrained optimization problem is formulated for the reconstruction of elastic parameters. An efficient algorithm is designed for numerical implementation using the Multiple Sparse Bayesian Learning (M-SBL) for joint sparse recovery problem and the Constrained Split Augmented Lagrangian Shrinkage Algorithm (C-SALSA) for the constrained optimization problem. The efficacy of the proposed framework is manifested through extensive numerical simulations. To the best of our knowledge, this is the first algorithm tailored for parameter reconstruction problems in elastic media using highly under-sampled data in the sense of Nyquist rate. Jaejun Yoo 0001, Younghoon Jung, Mikyoung Lim, Jong Chul Ye, Abdul Wahab 0002 |
SIAM J. Imaging Sci. | 4 |
| 2017 | Compressive Sampling Using Annihilating Filter-Based Low-Rank InterpolationabstractWhile the recent theory of compressed sensing provides an opportunity to overcome the Nyquist limit in recovering sparse signals, a solution approach usually takes the form of an inverse problem of an unknown signal, which is crucially dependent on specific signal representation. In this paper, we propose a drastically different two-step Fourier compressive sampling framework in a continuous domain that can be implemented via measurement domain interpolation, after which signal reconstruction can be done using classical analytic reconstruction methods. The main idea originates from the fundamental duality between the sparsity in the primary space and the low-rankness of a structured matrix in the spectral domain, showing that a low-rank interpolator in the spectral domain can enjoy all of the benefits of sparse recovery with performance guarantees. Most notably, the proposed low-rank interpolation approach can be regarded as a generalization of recent spectral compressed sensing to recover large classes of finite rate of innovations (FRI) signals at a near-optimal sampling rate. Moreover, for the case of cardinal representation, we can show that the proposed low-rank interpolation scheme will benefit from inherent regularization and an optimal incoherence parameter. Using a powerful dual certificate and the golfing scheme, we show that the new framework still achieves a near-optimal sampling rate for a general class of FRI signal recovery, while the sampling rate can be further reduced for a class of cardinal splines. Numerical results using various types of FRI signals confirm that the proposed low-rank interpolation approach offers significantly better phase transitions than conventional compressive sampling approaches. Jong Chul Ye, Jong Min Kim 0002, Kyong Hwan Jin, Kiryung Lee |
IEEE Trans. Inf. Theory | 1 |
| 2016 | Recent progresses of accelerated MRI using annihilating filter-based low-rank interpolationabstractRecently, an annihilating filter based low-rank Hankel matrix approach (ALOHA) was proposed as a general framework for sparsity-driven k-space interpolation method for compressed sensing MRI (CS-MRI). The principle of ALOHA framework is based on the fundamental duality between the transform domain sparsity in the primary space and the low-rankness of weighted Hankel matrix in Fourier domain, which converts CS-MRI to a k-space interpolation problem using structured matrix completion. In this review, we explain the theory behind ALOHA. Experimental results with in vivo data for multi-coil dynamic imaging, parametric mapping as well as Nyquist ghost correction confirmed that the proposed method has potential to be a general solution of various MR imaging problems. Kyong Hwan Jin, Dongwook Lee 0005, Jong Chul Ye |
ICIP | 4 |
| 2016 | Random impulse noise removal using sparse and low rank decomposition of annihilating filter-based Hankel matrixabstractAnnihilating filer-based low rank Hankel matrix (ALOHA) approach was recently proposed as an intrinsic image model for image inpainting estimation. Based on the observation that smoothness or textures within an image patch are represented as sparse spectral components in the frequency domain, ALOHA exploits the existence of annihilating filters and the associated rank-deficient Hankel matrices in the image domain to estimate the missing pixels. As a extension, here we propose a novel impulse noise removal algorithm using sparse + low rank decomposition of an annihilating filter-based Hankel matrix. This novel approach, what we call robust ALOHA, is inspired by the observation that an image corrupted with impulse noises has intact pixels; so the impulse noises can be modeled as sparse outliers, whereas the underlying image can be still modeled using a low-rank Hankel structured matrix. Numerical results confirm that robust ALOHA has significant performance improvements compared to the state-of-the-art impulse removal algorithms. Kyong Hwan Jin, Jong Chul Ye |
ICIP | 2 |
| 2015 | Interior Tomography Using 1D Generalized Total Variation. Part II: Multiscale ImplementationabstractTo address the classic interior tomography problem where projections at each view extend only to the shadow of a circular region completely interior to the subject being scanned, previously we showed that the exact recovery of two- and three-dimensional piecewise smooth images is guaranteed using a one-dimensional generalized total variation seminorm penalty which allows a much faster reconstruction. To further accelerate the algorithm up to a level for clinical use, this paper proposes a novel multiscale reconstruction method by exploiting the Bedrosian identity of the Hilbert transform. More specifically, we show that the high frequency parts of the one-dimensional signals can be quickly recovered analytically with the Hilbert transform because of the Bedrosian identity. This implies that computationally expensive iterative reconstruction need only be applied to low resolution images in the downsampled domain, which significantly reduces the computational burden. Moreover, even for incomplete trajectories such as circular cone-beam geometry, we demonstrate that the proposed multiscale interior tomography approach can be combined with a novel spectral blending method in order to mitigate cone-beam artifacts from missing frequency regions. We show the efficacy of the proposed multiscale algorithm using circular fan-beam, helical cone-beam data, and circular cone-beam geometry. With a graphics processing unit implementation, we demonstrate that the speed of the algorithm can be significantly accelerated up to the level for clinical use for various acquisition geometries. Yoseob Han, John Paul Ward, Michael Unser, Jong Chul Ye |
SIAM J. Imaging Sci. | 5 |
| 2015 | Interior Tomography Using 1D Generalized Total Variation. Part I: Mathematical FoundationabstractMotivated by the interior tomography problem, we propose a method for exact reconstruction of a region of interest of a function from its local Radon transform in any number of dimensions. Our aim is to verify the feasibility of a one-dimensional reconstruction procedure that can provide the foundation for an efficient algorithm. For a broad class of functions, including piecewise polynomials and generalized splines, we prove that an exact reconstruction is possible by minimizing a generalized total variation seminorm along lines. The main difference with previous works is that our approach is inherently one-dimensional and that it imposes less constraints on the class of admissible signals. Within this formulation, we derive unique reconstruction results using properties of the Hilbert transform, and we present numerical examples of the reconstruction. John Paul Ward, Jong Chul Ye, Michael Unser |
SIAM J. Imaging Sci. | 3 |
| 2015 | Annihilating Filter-Based Low-Rank Hankel Matrix Approach for Image InpaintingabstractIn this paper, we propose a patch-based image inpainting method using a low-rank Hankel structured matrix completion approach. The proposed method exploits the annihilation property between a shift-invariant filter and image data observed in many existing inpainting algorithms. In particular, by exploiting the commutative property of the convolution, the annihilation property results in a low-rank block Hankel structure data matrix, and the image inpainting problem becomes a low-rank structured matrix completion problem. The block Hankel structured matrices are obtained patch-by-patch to adapt to the local changes in the image statistics. To solve the structured low-rank matrix completion problem, we employ an alternating direction method of multipliers with factorization matrix initialization using the low-rank matrix fitting algorithm. As a side product of the matrix factorization, locally adaptive dictionaries can be also easily constructed. Despite the simplicity of the algorithm, the experimental results using irregularly subsampled images as well as various images with globally missing patterns showed that the proposed method outperforms existing state-of-the-art image inpainting methods. Kyong Hwan Jin, Jong Chul Ye |
IEEE Trans. Image Process. | 2 |
| 2015 | Sparse-View Spectral CT Reconstruction Using Spectral Patch-Based Low-Rank PenaltyabstractSpectral computed tomography (CT) is a promising technique with the potential for improving lesion detection, tissue characterization, and material decomposition. In this paper, we are interested in kVp switching-based spectral CT that alternates distinct kVp X-ray transmissions during gantry rotation. This system can acquire multiple X-ray energy transmissions without additional radiation dose. However, only sparse views are generated for each spectral measurement; and the spectra themselves are limited in number. To address these limitations, we propose a penalized maximum likelihood method using spectral patch-based low-rank penalty, which exploits the self-similarity of patches that are collected at the same position in spectral images. The main advantage is that the relatively small number of materials within each patch allows us to employ the low-rank penalty that is less sensitive to intensity changes while preserving edge directions. In our optimization formulation, the cost function consists of the Poisson log-likelihood for X-ray transmission and the nonconvex patch-based low-rank penalty. Since the original cost function is difficult to minimize directly, we propose an optimization method using separable quadratic surrogate and concave convex procedure algorithms for the log-likelihood and penalty terms, which results in an alternating minimization that provides a computational advantage because each subproblem can be solved independently. We performed computer simulations and a real experiment using a kVp switching-based spectral CT with sparse-view measurements, and compared the proposed method with conventional algorithms. We confirmed that the proposed method improves spectral images both qualitatively and quantitatively. Furthermore, our GPU implementation significantly reduces the computational cost. Kyung Sang Kim, Jong Chul Ye, William Worstell, Jinsong Ouyang, Yothin Rakvongthai, Georges El Fakhri, Quanzheng Li |
IEEE Trans. Medical Imaging | 2 |
| 2015 | A Unified Sparse Recovery and Inference Framework for Functional Diffuse Optical Tomography Using Random Effect ModelabstractDiffuse optical tomography (DOT) is a non-invasive imaging technique to reconstruct optical properties of biological tissues using near-infrared light, and it has been successfully used to measure functional brain activities via changes in cerebral blood volume and cerebral blood oxygenation. However, DOT presents a severely ill-posed inverse problem, so various types of regularization should be incorporated to overcome low spatial resolution and lack of depth sensitivity. Another limitation of the conventional DOT reconstruction methods is that an inference step is separately performed after the reconstruction, so complicated interaction between reconstruction and regularization is difficult to analyze. To overcome these technical difficulties, we propose a unified sparse recovery framework using a random effect model whose termination criterion is determined by the statistical inference. Both numerical and experimental results confirm that the proposed method outperforms the conventional approaches. Ok Kyun Lee, Sungho Tak, Jong Chul Ye |
IEEE Trans. Medical Imaging | 3 |
| 2014 | Ultra-Fast Hybrid CPU-GPU Multiple Scatter Simulation for 3-D PETabstractScatter correction is very important in 3-D PET reconstruction due to a large scatter contribution in measurements. Currently, one of the most popular methods is the so-called single scatter simulation (SSS), which considers single Compton scattering contributions from many randomly distributed scatter points. The SSS enables a fast calculation of scattering with a relatively high accuracy; however, the accuracy of SSS is dependent on the accuracy of tail fitting to find a correct scaling factor, which is often difficult in low photon count measurements. To overcome this drawback as well as to improve accuracy of scatter estimation by incorporating multiple scattering contribution, we propose a multiple scatter simulation (MSS) based on a simplified Monte Carlo (MC) simulation that considers photon migration and interactions due to photoelectric absorption and Compton scattering. Unlike the SSS, the MSS calculates a scaling factor by comparing simulated prompt data with the measured data in the whole volume, which enables a more robust estimation of a scaling factor. Even though the proposed MSS is based on MC, a significant acceleration of the computational time is possible by using a virtual detector array with a larger pitch by exploiting that the scatter distribution varies slowly in spatial domain. Furthermore, our MSS implementation is nicely fit to a parallel implementation using graphic processor unit (GPU). In particular, we exploit a hybrid CPU-GPU technique using the open multiprocessing and the compute unified device architecture, which results in 128.3 times faster than using a single CPU. Overall, the computational time of MSS is 9.4 s for a high-resolution research tomograph (HRRT) system. The performance of the proposed MSS is validated through actual experiments using an HRRT. Kyung Sang Kim, Young-Don Son, Zang-Hee Cho, Jong Beom Ra, Jong Chul Ye |
IEEE J. Biomed. Health Informatics | 5 |
| 2014 | Motion Adaptive Patch-Based Low-Rank Approach for Compressed Sensing Cardiac Cine MRIabstractOne of the technical challenges in cine magnetic resonance imaging (MRI) is to reduce the acquisition time to enable the high spatio-temporal resolution imaging of a cardiac volume within a short scan time. Recently, compressed sensing approaches have been investigated extensively for highly accelerated cine MRI by exploiting transform domain sparsity using linear transforms such as wavelets, and Fourier. However, in cardiac cine imaging, the cardiac volume changes significantly between frames, and there often exist abrupt pixel value changes along time. In order to effectively sparsify such temporal variations, it is necessary to exploit temporal redundancy along motion trajectories. This paper introduces a novel patch-based reconstruction method to exploit geometric similarities in the spatio-temporal domain. In particular, we use a low rank constraint for similar patches along motion, based on the observation that rank structures are relatively less sensitive to global intensity changes, but make it easier to capture moving edges. A Nash equilibrium formulation with relaxation is employed to guarantee convergence. Experimental results show that the proposed algorithm clearly reconstructs important anatomical structures in cardiac cine image and provides improved image quality compared to existing state-of-the-art methods such as k-t FOCUSS, k-t SLR, and MASTeR. Huisu Yoon, Kyung Sang Kim, Daniel Kim 0002, Yoram Bresler, Jong Chul Ye |
IEEE Trans. Medical Imaging | 5 |
| 2013 | Subspace penalized sparse learning for joint sparse recoveryabstractThe multiple measurement vector problem (MMV) is a generalization of the compressed sensing problem that addresses the recovery of a set of jointly sparse signal vectors. One of the important contributions of this paper is to reveal that the seemingly least related state-of-art MMV joint sparse recovery algorithms - M-SBL (multiple sparse Bayesian learning) and subspace-based hybrid greedy algorithms - have a very important link. More specifically, we show that replacing the log det(·) term in M-SBL by a log det(·) rank proxy that exploits the spark reduction property discovered in subspace-based joint sparse recovery algorithms, provides significant improvements. Theoretical analysis demonstrates that even thoughM-SBL is often unable to remove all localminimizers, the proposed method can do so under fairly mild conditions, without affecting the global minimizer. Jong Chul Ye, Jong Min Kim 0002, Yoram Bresler |
ICASSP | 1 |
| 2013 | Corrections to "Compressive MUSIC: Revisiting the Link Between Compressive Sensing and Array Signal Processing"abstractThere are a few corrections for the above titled paper (IEEE Trans. Inf. Theory, vol. 58, no. 1, pp. 278-301, Jan. 2012). They are presented here. Jong Min Kim 0002, Jong Chul Ye |
IEEE Trans. Inf. Theory | 2 |
| 2012 | Dynamic sparse support tracking with multiple measurement vectors using compressive MUSICabstractDynamic tracking of sparse targets has been one of the important topics in array signal processing. Recently, compressed sensing (CS) approaches have been extensively investigated as a new tool for this problem using partial support information obtained by exploiting temporal redundancy. However, most of these approaches are formulated under single measurement vector compressed sensing (SMV-CS) framework, where the performance guarantees are only in a probabilistic manner. The main contribution of this paper is to allow deterministic tracking of time varying supports with multiple measurement vectors (MMV) by exploiting multi-sensor diversity. In particular, we show that a novel compressive MUSIC (CS-MUSIC) algorithm with optimized partial support selection not only allows removal of inaccurate portion of previous support estimation but also enables addition of newly emerged part of unknown support. Numerical results confirm the theory. Jong Min Kim 0002, Ok Kyun Lee, Jong Chul Ye |
ICASSP | 3 |
| 2012 | Resting-state fMRI analysis of Alzheimer's disease progress using sparse dictionary learningabstractA novel data-driven resting state fMRI analysis based on sparse dictionary learning is presented. Although ICA has been a popular data-driven method for resting state fMRI data, the assumption that sources are independent often leads to a paradox in analyzing closely interconnected brain networks. Rather than using independency, the proposed approach starts from an assumption that a temporal dynamics at each voxel position is a sparse combination of global brain dynamics and then proposes a novel sparse dictionary learning method for analyzing the resting state fMRI analysis. Moreover, using a mixed model, we provide a statistically rigorous group analysis. Using extensive data set obtained from normal, Mild Cognitive Impairment (MCI), Clinical Dementia Rating scale (CDR) 0.5, CDR 1.0, and CDR 2.0 patients groups, we demonstrated that the changes of default mode network extracted by the proposed method is more closely correlated with the progress of Alzheimer disease. Jeonghyeon Lee, Jong Chul Ye |
SMC | 2 |
| 2012 | Compressive MUSIC: Revisiting the Link Between Compressive Sensing and Array Signal ProcessingabstractThe multiple measurement vector (MMV) problem addresses the identification of unknown input vectors that share common sparse support. Even though MMV problems have been traditionally addressed within the context of sensor array signal processing, the recent trend is to apply compressive sensing (CS) due to its capability to estimate sparse support even with an insufficient number of snapshots, in which case classical array signal processing fails. However, CS guarantees the accurate recovery in a probabilistic manner, which often shows inferior performance in the regime where the traditional array signal processing approaches succeed. The apparent dichotomy between the probabilistic CS and deterministic sensor array signal processing has not been fully understood. The main contribution of the present article is a unified approach that revisits the link between CS and array signal processing first unveiled in the mid 1990s by Feng and Bresler. The new algorithm, which we call compressive MUSIC, identifies the parts of support using CS, after which the remaining supports are estimated using a novel generalized MUSIC criterion. Using a large system MMV model, we show that our compressive MUSIC requires a smaller number of sensor elements for accurate support recovery than the existing CS methods and that it can approach the optimal -bound with finite number of snapshots even in cases where the signals are linearly dependent. Jong Min Kim 0002, Ok Kyun Lee, Jong Chul Ye |
IEEE Trans. Inf. Theory | 3 |
| 2011 | Compressive MUSIC with optimized partial support for joint sparse recoveryabstractThe multiple measurement vector (MMV) problem addresses the identification of unknown input vectors that share common sparse support. The MMV problem has been traditionally addressed either by sensor array signal processing or compressive sensing. However, recent breakthroughs in this area such as compressive MUSIC (CS-MUSIC) or subspace-augumented MUSIC (SA-MUSIC) optimally combine the compressive sensing (CS) and array signal processing such that k - r supports are first found by CS and the remaining r supports are determined by a generalized MUSIC criterion, where k and r denote the sparsity and the number of independent snapshots, respectively. Even though such a hybrid approach significantly outperforms the conventional algorithms, its performance heavily depends on the correct identification of k-r partial support by the compressive sensing step, which often deteriorates the overall performance. The main contribution of this paper is, therefore, to show that as long as k - r + 1 correct supports are included in any k-sparse CS solution, the optimal k - r partial support can be found using a subspace fitting criterion, significantly improving the overall performance of CS-MUSIC. Furthermore, unlike the single measurement CS counterpart that requires infinite SNR for a perfect support recovery, we can derive an information theoretic sufficient condition for the perfect recovery using CS-MUSIC under a finite SNR scenario. Jong Min Kim 0002, Ok Kyun Lee, Jong Chul Ye |
ISIT | 3 |
| 2011 | Compressive Diffuse Optical Tomography: Noniterative Exact Reconstruction Using Joint SparsityabstractDiffuse optical tomography (DOT) is a sensitive and relatively low cost imaging modality that reconstructs optical properties of a highly scattering medium. However, due to the diffusive nature of light propagation, the problem is severely ill-conditioned and highly nonlinear. Even though nonlinear iterative methods have been commonly used, they are computationally expensive especially for three dimensional imaging geometry. Recently, compressed sensing theory has provided a systematic understanding of high resolution reconstruction of sparse objects in many imaging problems; hence, the goal of this paper is to extend the theory to the diffuse optical tomography problem. The main contributions of this paper are to formulate the imaging problem as a joint sparse recovery problem in a compressive sensing framework and to propose a novel noniterative and exact inversion algorithm that achieves the l(0) optimality as the rank of measurement increases to the unknown sparsity level. The algorithm is based on the recently discovered generalized MUSIC criterion, which exploits the advantages of both compressive sensing and array signal processing. A theoretical criterion for optimizing the imaging geometry is provided, and simulation results confirm that the new algorithm outperforms the existing algorithms and reliably reconstructs the optical inhomogeneities when we assume that the optical background is known to a reasonable accuracy. Ok Kyun Lee, Jong Min Kim 0002, Yoram Bresler, Jong Chul Ye |
IEEE Trans. Medical Imaging | 4 |
| 2011 | A Data-Driven Sparse GLM for fMRI Analysis Using Sparse Dictionary Learning With MDL CriterionabstractWe propose a novel statistical analysis method for functional magnetic resonance imaging (fMRI) to overcome the drawbacks of conventional data-driven methods such as the independent component analysis (ICA). Although ICA has been broadly applied to fMRI due to its capacity to separate spatially or temporally independent components, the assumption of independence has been challenged by recent studies showing that ICA does not guarantee independence of simultaneously occurring distinct activity patterns in the brain. Instead, sparsity of the signal has been shown to be more promising. This coincides with biological findings such as sparse coding in V1 simple cells, electrophysiological experiment results in the human medial temporal lobe, etc. The main contribution of this paper is, therefore, a new data driven fMRI analysis that is derived solely based upon the sparsity of the signals. A compressed sensing based data-driven sparse generalized linear model is proposed that enables estimation of spatially adaptive design matrix as well as sparse signal components that represent synchronous, functionally organized and integrated neural hemodynamics. Furthermore, a minimum description length (MDL)-based model order selection rule is shown to be essential in selecting unknown sparsity level for sparse dictionary learning. Using simulation and real fMRI experiments, we show that the proposed method can adapt individual variation better compared to the conventional ICA methods. Kangjoo Lee, Sungho Tak, Jong Chul Ye |
IEEE Trans. Medical Imaging | 3 |
| 2009 | Single channel 2-D and 3-D blind image deconvolution for circularly symmetric fir blursabstractThe circular symmetry of point spread function (PSF) is common for many optical imaging systems such as optical microscopes, cameras, and astigmatism corrected electron microscopes. In our previous work, we showed that an accurate PSF estimation from a single measured image on a flat background is possible by exploiting the circular symmetry. This paper extends the results to 2-D single blind deconvolution from partial data as well as 3-D single channel blind deconvolution. This work allows simple computation section of biological phantom from z-stack images of brightfield and fluorescence microscope. Experimental results using a fluorescence microscope confirm our theory. Kwang Eun Jang, Hee Won Yang, Jong Chul Ye |
ICIP | 3 |
| 2008 | Non-iterative exact inverse scattering using simultaneous orthogonal matching pursuit (S-OMP)abstractEven though recently proposed time-reversal MUSIC approach for inverse scattering problem is non-iterative and exact, the approach breaks down when there are more targets than sensors. The main contribution of this paper is a novel non-iterative exact inverse scattering algorithm that still guarantees the exact recovery of the extended targets under a very relaxed constraint on the number of source and receivers, where the conventional time-reversal MUSIC fails. Such breakthrough was possible from the observation that the induced currents on the unknown targets assume the same sparse support, which can be recovered accurately using the simultaneous orthogonal matching pursuit developed for multiple measurement vector problems. Simulation results demonstrate that perfect reconstruction can be quickly obtained from a very limited number of samples. Jong Chul Ye, Su Yeon Lee |
ICASSP | 1 |
| 2007 | Compressed Sensing Shape Estimation of Star-Shaped Objects in Fourier ImagingabstractRecent theory of compressed sensing informs us that near-exact recovery of an unknown sparse signal is possible from a very limited number of Fourier samples by solving a convex L1optimization problem. The main contribution of the present letter is a compressed sensing-based novel nonparametric shape estimation framework and a computational algorithm for binary star shape objects, whose radius functions belong to the space of bounded-variation functions. Specifically, in contrast with standard compressed sensing, the present approach involves directly reconstructing the shape boundary under sparsity constraint. This is done by converting the standard pixel-based reconstruction approach into estimation of a nonparametric shape boundary on a wavelet basis. This results in an L1minimization under a nonlinear constraint, which makes the optimization problem nonconvex. We solve the problem by successive linearization and application of one-dimensional L1minimization, which significantly reduces the number of sampling requirements as well as the computational burden. Fourier imaging simulation results demonstrate that high quality reconstruction can be quickly obtained from a very limited number of samples. Furthermore, the algorithm outperforms the standard compressed sensing reconstruction approach using the total variation norm. Jong Chul Ye |
IEEE Signal Process. Lett. | 1 |
| 2006 | Asymptotic Global Confidence Regions for 3-D Parametric Shape Estimation in Inverse ProblemsabstractThis paper derives fundamental performance bounds for statistical estimation of parametric surfaces embedded in R3. Unlike conventional pixel-based image reconstruction approaches, our problem is reconstruction of the shape of binary or homogeneous objects. The fundamental uncertainty of such estimation problems can be represented by global confidenceregions, which facilitate geometric inference and optimization ofthe imaging system. Compared to our previous work on global confidence region analysis for curves [two-dimensional (2-D) shapes], computation of the probability that the entire surface estimate lies within the confidence region is more challenging because a surface estimate is an inhomogeneous random field continuously indexed by a 2-D variable. We derive an asymptotic lower bound to this probability by relating it to the exceedence probability of a higher dimensional Gaussian random field, which can, in turn, be evaluated using the tube formula due to Sun. Simulation results demonstrate the tightness of the resulting bound and the usefulness of the three-dimensional global confidence region approach. Jong Chul Ye, Pierre Moulin, Yoram Bresler |
IEEE Trans. Image Process. | 1 |
| 2005 | In vivo optical molecular imaging: principles and signal processing issuesabstractIn vivo optical molecular imaging involves the use of light emitting tracers combined with sophisticated sensing modalities to perform in vivo imaging of genetic and molecular information. In contrast to the classical diagnostic imaging tools which image the end effects of the diseases, optical molecular imaging could enhance our knowledge of biological phenomena, monitor genetic expression and the alteration of cells, and lead to earlier detection of diseases. With the development of exotic molecular probes with easily detectable bioluminescence and fluorescence labels, optical molecular imaging has emerged as an important new field within biomedical imaging. This paper reviews this state-of-the-art imaging technology and signal processing issues to monitor molecular and cellular events in living organisms. Jong Chul Ye, Kevin J. Webb, Rick P. Millane, Charles A. Bouman |
ICASSP (5) | 1 |
| 2003 | Rate-distortion optimized data partitioning for video using backward adaptationabstractWhile data partitioning, in conjunction with unequal error protection, provides superb error resiliency, insufficient video quality when only the base partition is available prevents its wide deployment in high-quality video applications. We develop a new scheme for data partitioning of motion-compensated DCT coded video in an operational rate-distortion context. Unlike the conventional data partitioning scheme, which adapts the DCT break points at slice or video packet level, our new partitioning algorithm adapts the partitioning points at as low as the DCT block level with virtually no overhead using backward adaptation; hence it produces superior video quality over the conventional data partitioning scheme. Simulation results show that significant PSNR gain can be achieved using the new algorithm. Jong Chul Ye, Yingwei Chen |
ICASSP (3) | 1 |
| 2003 | Video streaming over wireless LAN with efficient scalable coding and prioritized adaptive transmissionabstractWe propose a video streaming scheme over wireless LAN using novel rate-distortion optimized data partitioning scheme and prioritized adaptive packet transmission. The new algorithm enables DCT data partitioning up to the DCT block level without rate overhead using backward adaptation. For channel adaptation, we selectively drop enhancement layer packets when the channel throughput is not sufficient. Real video streaming experiments with 802.11b testbed have demonstrated excellent performance of the algorithm under severe interference such as microwave oven. Yingwei Chen, Jong Chul Ye, Carles Ruiz Floriach, Kiran S. Challapali |
ICIP (3) | 2 |
| 2003 | Channel adaptive prioritized transmission of layered video over wireless LANabstractRobust video transmission over wireless LAN remains a key technical obstacle against the mass deployment of wireless multimedia home networking, due to the lack of guaranteed Quality of Service over the wireless medium and the sensitivity of single-layer coded video (MPEG-2, MPEG-4, etc) to packet losses. This paper describes a robust video transmission system over wireless LAN utilizing layered and prioritized transmission mechanisms. Specifically, a single layer coded video is first organized into multiple layers of different importance. The different layers are then packetized and transmitted with different priorities to ensure that vital video information is sent with higher priority while the rest (enhancement layer(s)) are sent with only best effort. Realtime streaming tests show that our streaming system can effectively minimize the video quality degradation under channel throughput variation. Yingwei Chen, Carles Ruiz Floriach, Jong Chul Ye, Kiran S. Challapali |
PIMRC | 3 |
| 2003 | Cramer-Rao bounds for parametric shape estimation in inverse problemsabstractWe address the problem of computing fundamental performance bounds for estimation of object boundaries from noisy measurements in inverse problems, when the boundaries are parameterized by a finite number of unknown variables. Our model applies to multiple unknown objects, each with its own unknown gray level, or color, and boundary parameterization, on an arbitrary known background. While such fundamental bounds on the performance of shape estimation algorithms can in principle be derived from the Cramér-Rao lower bounds, very few results have been reported due to the difficulty of computing the derivatives of a functional with respect to shape deformation. We provide a general formula for computing Cramér-Rao lower bounds in inverse problems where the observations are related to the object by a general linear transform, followed by a possibly nonlinear and noisy measurement system. As an illustration, we derive explicit formulas for computed tomography, Fourier imaging, and deconvolution problems. The bounds reveal that highly accurate parametric reconstructions are possible in these examples, using severely limited and noisy data. Jong Chul Ye, Yoram Bresler, Pierre Moulin |
IEEE Trans. Image Process. | 1 |
| 2002 | Cramer-Rao bounds for parametric shape estimationabstractWe address the problem of computing fundamental performance bounds for estimation of object boundaries from noisy measurements in inverse problems, when the boundaries are parameterized by a finite number of unknown variables. Our model applies to multiple unknown objects, each with its own unknown gray level, or color, and boundary parameterization, on an arbitrary known background. While such fundamental bounds on the performance of shape estimation algorithms can in principle be derived from the Cramer-Rao lower bounds, very few results have been reported due to the difficulty of computing the derivatives of a functional with respect to shape deformation. We provide a general formula for computing Cramer-Rao lower bounds in inverse problems where the observations are related to the object by a general linear transform, followed by a possibly nonlinear and noisy measurement system. Jong Chul Ye, Yoram Bresler, Pierre Moulin |
ICIP (2) | 1 |
| 2002 | A Self-Referencing Level-Set Method for Image Reconstruction from Sparse Fourier Samples
Jong Chul Ye, Yoram Bresler, Pierre Moulin |
Int. J. Comput. Vis. | 1 |
| 2001 | A self-referencing level-set method for image reconstruction from sparse Fourier samplesabstractWe address image estimation from sparse Fourier samples. The problem is formulated as joint estimation of the supports of unknown sparse objects in the image, and pixel values on these supports. The domain and the pixel values are alternately estimated using the level-set method and the conjugate gradient method, respectively. Our level-set evolution shows a unique switching behavior, which stabilizes the level-set evolution and removes the re-initialization steps in conventional level set approaches. Jong Chul Ye, Yoram Bresler, Pierre Moulin |
ICIP (2) | 1 |
| 2001 | Nonlinear multigrid algorithms for Bayesian optical diffusion tomographyabstractOptical diffusion tomography is a technique for imaging a highly scattering medium using measurements of transmitted modulated light. Reconstruction of the spatial distribution of the optical properties of the medium from such data is a difficult nonlinear inverse problem. Bayesian approaches are effective, but are computationally expensive, especially for three-dimensional (3-D) imaging. This paper presents a general nonlinear multigrid optimization technique suitable for reducing the computational burden in a range of nonquadratic optimization problems. This multigrid method is applied to compute the maximum a posteriori (MAP) estimate of the reconstructed image in the optical diffusion tomography problem. The proposed multigrid approach both dramatically reduces the required computation and improves the reconstructed image quality. Jong Chul Ye, Charles A. Bouman, Kevin J. Webb, Rick P. Millane |
IEEE Trans. Image Process. | 1 |
| 2000 | Global confidence regions in parametric shape estimationabstractWe introduce confidence region techniques for analyzing and visualizing the performance of two-dimensional parametric shape estimators. Assuming an asymptotically normal and efficient estimator for a finite parameterization of the object boundary, Cramer-Rao bounds are used to define a confidence region, centered around the true boundary. Computation of the probability that an entire boundary estimate lies within the confidence region is a challenging problem, because the estimate is a two-dimensional nonstationary random process. We derive lower bounds on this probability using level crossing statistics. The results make it possible to generate confidence regions for arbitrary prescribed probabilities. These global confidence regions conveniently display the uncertainty in various geometric parameters such as shape, size, orientation, and position of the estimated object, and facilitate geometric inferences. Numerical simulations suggest that the new bounds are quite tight. Jong Chul Ye, Yoram Bresler, Pierre Moulin |
ICASSP | 1 |
| 2000 | Asymptotic global confidence regions in parametric shape estimation problemsabstractWe introduce confidence region techniques for analyzing and visualizing the performance of two-dimensional parametric shape estimators. Assuming an asymptotically normal and efficient estimator for a finite parameterization of the object boundary, Cramer-Rao bounds are used to define an asymptotic confidence region, centered around the true boundary. Computation of the probability that an entire boundary estimate lies within the confidence region is a challenging problem, because the estimate is a two-dimensional nonstationary random process. We derive lower bounds on this probability using level crossing statistics. The same bounds also apply to asymptotic confidence regions formed around the estimated boundaries, lower-bounding the probability that the entire true boundary lies within the confidence region. The results make it possible to generate asymptotic confidence regions for arbitrary prescribed probabilities. These asymptotic global confidence regions conveniently display the uncertainty in various geometric parameters such as shape, size, orientation, and position of the estimated object, and facilitate geometric inferences. Numerical simulations suggest that the new bounds are quite tight. Jong Chul Ye, Yoram Bresler, Pierre Moulin |
IEEE Trans. Inf. Theory | 1 |
| 1999 | Nonlinear Multigrid Optimization for Bayesian Diffusion TomographyabstractOptical diffusion tomography attempts to reconstruct an object cross section (a highly scattering media such as tissue) from measurements of scattered and attenuated light. While Bayesian approaches are well suited to this difficult nonlinear inverse problem, the resulting optimization problem is very computationally expensive. In this paper, we propose a nonlinear multigrid technique for computing the maximum a posteriori (MAP) reconstruction in the optical diffusion tomography problem. The multigrid approach improves reconstruction quality by avoiding a local minimum. In addition, it dramatically reduces computation. Each iteration of the algorithm alternates a Born approximation step with a single cycle of a nonlinear multigrid algorithm. Jong Chul Ye, Charles A. Bouman, Rick P. Millane, Kevin J. Webb |
ICIP (2) | 1 |
| 1998 | Optimal Parameter Updating for Optical Diffusion Imaging
Jong Chul Ye, Kevin J. Webb, Rick P. Millane, Thomas J. Downar |
ICIP (3) | 1 |