EDBT 2026 Demo / reviewers in the wild / expert
Juhyung Ha
dblp:356/6937
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-6596-5116ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Video understanding and tracking · 87% Language models and text generation · 13% | |
| Computer graphics and multimedia
1 paper |
Geometric modeling and processing · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
action segmentation |
0.9 | 1 | 2025 | What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning · ICCV 2025 |
Geometric modeling and processing › point cloud processing
point cloud upsampling |
0.9 | 1 | 2025 | HVPUNet: Hybrid-Voxel Point-Cloud Upsampling Network · ICCV 2025 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation Learning · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
state-change counterfactuals · 0.9LLM-generated descriptions · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker DetectionabstractActive Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails to capture fine-grained cross-modal interactions, which can be critical for robust performance in unconstrained scenarios. In this paper, we introduce GateFusion, a novel architecture that combines strong pretrained unimodal encoders with a Hierarchical Gated Fusion Decoder (HiGate). HiGate enables progressive, multi-depth fusion by adaptively injecting contextual features from one modality into the other at multiple layers of the Transformer backbone, guided by learnable, bimodally-conditioned gates. To further strengthen multimodal learning, we propose two auxiliary objectives: Masked Alignment Loss (MAL) to align unimodal outputs with multimodal predictions, and Over-Positive Penalty (OPP) to suppress spurious video-only activations. GateFusion establishes new state-of-the-art results on several challenging ASD benchmarks, achieving 77.8% mAP (+9.4%), 86.1% mAP (+2.9%), and 96.1% mAP (+0.5%) on Ego4D-ASD, UniTalk, and WASD benchmarks, respectively, and delivering competitive performance on AVA-ActiveSpeaker. Out-of-domain experiments demonstrate the generalization of our model, while comprehensive ablations show the complementary benefits of each component. Yu Wang 0164, Juhyung Ha, Frangil Ramirez, David Crandall |
WACV | 2 |
| 2025 | HVPUNet: Hybrid-Voxel Point-Cloud Upsampling Network
Juhyung Ha, Vibhas Vats, Soon-Heung Jung, Md. Alimoor Reza, David Crandall |
ICCV | 1 |
| 2025 | What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation LearningabstractUnderstanding a procedural activity requires modeling both how action steps transform the scene, and how evolving scene transformations can influence the sequence of action steps, even those that are accidental or erroneous. Existing work has studied procedure-aware video representations by modeling the temporal order of actions, but has not explicitly learned the state changes (scene transformations). In this work, we study procedure-aware video representation learning by incorporating state-change descriptions generated by Large Language Models (LLMs) as supervision signals for video encoders. Moreover, we generate state-change counterfactuals that simulate hypothesized failure outcomes, allowing models to learn by imagining unseen "What if" scenarios. This counterfactual reasoning facilitates the model's ability to understand the cause and effect of each step in an activity. We conduct extensive experiments on procedure-aware tasks, including temporal action segmentation, error detection, action phase classification, frame retrieval, multi-instance retrieval, and action recognition. Our results demonstrate the effectiveness of the proposed state-change descriptions and their counterfactuals, and achieve significant improvements on multiple tasks. Chi-Hsi Kung, Frangil Ramirez, Juhyung Ha, Yi-Ting Chen 0001, David Crandall, Yi-Hsuan Tsai |
ICCV | 3 |
| 2025 | Multi-Resolution Guided 3D GANs for Medical Image TranslationabstractMedical image translation is the process of converting from one imaging modality to another, in order to reduce the need for multiple image acquisitions from the same patient. This can enhance the efficiency of treatment by reducing the time, equipment, and labor needed. In this paper, we introduce a multi-resolution guided Generative Adversarial Network (GAN)-based framework for 3D medical image translation. Our framework uses a 3D multi-resolution DenseAttention UNet (3D-mDAUNet) as the generator and a 3D multi-resolution UNet as the discriminator, optimized with a unique combination of loss functions including voxel-wise GAN loss and 2.5D perception loss. Our approach yields promising results in volumetric image quality assessment (IQA) across a variety of imaging modalities, body regions, and age groups, demonstrating its robustness. Furthermore, we propose a synthetic-to-real applicability assessment as an additional evaluation to assess the effectiveness of synthetic data in downstream applications such as segmentation. This comprehensive evaluation shows that our method produces synthetic medical images not only of high-quality but also potentially useful in clinical applications. Our code is available at github.com/juhha/3D-mADUNet. Juhyung Ha, Jong Sung Park, David Crandall, Eleftherios Garyfallidis, Xuhong Zhang 0001 |
WACV | 1 |
| 2024 | Generating synthetic computed tomography for radiotherapy: SynthRAD2023 challenge reportabstractRadiation therapy plays a crucial role in cancer treatment, necessitating precise delivery of radiation to tumors while sparing healthy tissues over multiple days. Computed tomography (CT) is integral for treatment planning, offering electron density data crucial for accurate dose calculations. However, accurately representing patient anatomy is challenging, especially in adaptive radiotherapy, where CT is not acquired daily. Magnetic resonance imaging (MRI) provides superior soft-tissue contrast. Still, it lacks electron density information, while cone beam CT (CBCT) lacks direct electron density calibration and is mainly used for patient positioning. Adopting MRI-only or CBCT-based adaptive radiotherapy eliminates the need for CT planning but presents challenges. Synthetic CT (sCT) generation techniques aim to address these challenges by using image synthesis to bridge the gap between MRI, CBCT, and CT. The SynthRAD2023 challenge was organized to compare synthetic CT generation methods using multi-center ground truth data from 1080 patients, divided into two tasks: (1) MRI-to-CT and (2) CBCT-to-CT. The evaluation included image similarity and dose-based metrics from proton and photon plans. The challenge attracted significant participation, with 617 registrations and 22/17 valid submissions for tasks 1/2. Top-performing teams achieved high structural similarity indices (≥0.87/0.90) and gamma pass rates for photon (≥98.1%/99.0%) and proton (≥97.3%/97.0%) plans. However, no significant correlation was found between image similarity metrics and dose accuracy, emphasizing the need for dose evaluation when assessing the clinical applicability of sCT. SynthRAD2023 facilitated the investigation and benchmarking of sCT generation techniques, providing insights for developing MRI-only and CBCT-based adaptive radiotherapy. It showcased the growing capacity of deep learning to produce high-quality sCT, reducing reliance on conventional CT for treatment planning. Evi M. C. Huijben, Maarten L. Terpstra, Arthur Jr Galapon, Suraj Pai, Adrian Thummerer, Peter J. Koopmans, Manya Afonso, Maureen van Eijnatten, Oliver J. Gurney-Champion, Zeli Chen, Kaiyi Zheng, Chuanpu Li, Haowen Pang, Chuyang Ye, Runqi Wang, Fuxin Fan, Jingna Qiu, Yixing Huang, Juhyung Ha, Jong Sung Park, Alexandra Alain-Beaudoin, Silvain Bériault, Pengxin Yu, Zhanyao Huang, Gengwan Li, Xueru Zhang, Yubo Fan, Bowen Xin, Aaron Nicolson, Lujia Zhong, Zhiwei Deng, Gustav Mueller-Franzes, Firas Khader, Xia Li 0005, Ye Zhang 0039, Cédric Hémon, Valentin Boussot, Shaobin Wang, Derk Mus, Bram Kooiman, Chelsea A. H. Sargeant, Edward G. A. Henderson, Satoshi Kondo, Satoshi Kasai, Reza Karimzadeh, Bulat Ibragimov, Thomas Helfer, Jessica Dafflon, Enpei Wang, Zoltán Perkó, Matteo Maspero |
Medical Image Anal. | 21 |