Junghyun Cho

dblp:18/4874 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Systems, architecture and hardware · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 Multi-View Pedestrian Occupancy Prediction with a Novel Synthetic Dataset
abstract
We address an advanced challenge of predicting pedestrian occupancy as an extension of multi-view pedestrian detection in urban traffic. To support this, we have created a new synthetic dataset called MVP-Occ, designed for dense pedestrian scenarios in large-scale scenes. Our dataset provides detailed representations of pedestrians using voxel structures, accompanied by rich semantic scene understanding labels, facilitating visual navigation and insights into pedestrian spatial information. Furthermore, we present a robust baseline model, termed OmniOcc, capable of predicting both the voxel occupancy state and panoptic labels for the entire scene from multi-view images. Through in-depth analysis, we identify and evaluate the key elements of our proposed model, highlighting their specific contributions and importance.
Sithu Aung, Min-Cheol Sagong, Junghyun Cho
AAAI3
2025 Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor Scenes
abstract
We propose a diffusion-based inverse rendering framework that decomposes a single RGB image into geometry, material, and lighting. Inverse rendering is inherently ill-posed, making it difficult to predict a single accurate solution. To address this challenge, recent generative model-based methods aim to present a range of possible solutions. However, finding a single accurate solution and generating diverse solutions can be conflicting. In this paper, we propose a channel-wise noise scheduling approach that allows a single diffusion model architecture to achieve two conflicting objectives. The resulting two diffusion models, trained with different channel-wise noise schedules, can predict a single highly accurate solution and present multiple possible solutions. The experimental results demonstrate the superiority of our two models in terms of both diversity and accuracy, which translates to enhanced performance in downstream applications such as object insertion and material editing.
Junyong Choi, Min-Cheol Sagong, SeokYeong Lee, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho
CVPR6
2025 VIGFace: Virtual Identity Generation for Privacy-Free Face Recognition Dataset
Min-Cheol Sagong, Gi Pyo Nam, Junghyun Cho, Ig-Jae Kim
ICCV4
2025 MAIR++: Improving Multi-View Attention Inverse Rendering With Implicit Lighting Representation
abstract
In this paper, we propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, SVBRDF, and 3D spatially-varying lighting. While multi-view images have been widely used for object-level inverse rendering, scene-level inverse rendering has primarily been studied using single-view images due to the lack of a dataset containing high dynamic range multi-view images with ground-truth geometry, material, and spatially-varying lighting. To improve the quality of scene-level inverse rendering, a novel framework called Multi-view Attention Inverse Rendering (MAIR) was recently introduced. MAIR performs scene-level multi-view inverse rendering by expanding the OpenRooms dataset, designing efficient pipelines to handle multi-view images, and splitting spatially-varying lighting. Although MAIR showed impressive results, its lighting representation is fixed to spherical Gaussians, which limits its ability to render images realistically. Consequently, MAIR cannot be directly used in applications such as material editing. Moreover, its multi-view aggregation networks have difficulties extracting rich features because they only focus on the mean and variance between multi-view features. In this paper, we propose its extended version, called MAIR++. MAIR++ addresses the aforementioned limitations by introducing an implicit lighting representation that accurately captures the lighting conditions of an image while facilitating realistic rendering. Furthermore, we design a directional attention-based multi-view aggregation network to infer more intricate relationships between views. Experimental results show that MAIR++ not only outperforms MAIR and single-view-based methods but also demonstrates robust performance on unseen real-world scenes.
Junyong Choi, SeokYeong Lee, Haesol Park, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Few-Shot Neural Radiance Fields under Unconstrained Illumination
abstract
In this paper, we introduce a new challenge for synthesizing novel view images in practical environments with limited input multi-view images and varying lighting conditions. Neural radiance fields (NeRF), one of the pioneering works for this task, demand an extensive set of multi-view images taken under constrained illumination, which is often unattainable in real-world settings. While some previous works have managed to synthesize novel views given images with different illumination, their performance still relies on a substantial number of input multi-view images. To address this problem, we suggest ExtremeNeRF, which utilizes multi-view albedo consistency, supported by geometric alignment. Specifically, we extract intrinsic image components that should be illumination-invariant across different views, enabling direct appearance comparison between the input and novel view under unconstrained illumination. We offer thorough experimental results for task evaluation, employing the newly created NeRF Extreme benchmark—the first in-the-wild benchmark for novel view synthesis under multiple viewing directions and varying illuminations.
SeokYeong Lee, Junyong Choi, Seungryong Kim, Ig-Jae Kim, Junghyun Cho
AAAI5
2024 Enhancing Multi-view Pedestrian Detection Through Generalized 3D Feature Pulling
abstract
The main challenge in multi-view pedestrian detection is integrating view-specific features into a unified space for comprehensive end-to-end perception. Prior multi-view detection methods have focused on projecting perspective-view features onto the ground plane, creating a "bird’s eye view" (BEV) representation of the scene. This paper proposes a simple but effective architecture that utilizes a nonparametric 3D feature-pulling strategy. This strategy directly extracts the corresponding 2D features for each valid voxel within the 3D feature volume, addressing the feature loss that may arise in previous methods. The proposed framework introduces three novel modules, each crafted to bolster the generalization capabilities of multi-view detection systems. Through extensive experiments, the efficacy of the proposed model is demonstrated. The results show a new state-of-the-art accuracy, both in conventional scenarios and particularly in the context of scene generalization benchmarks.
Sithu Aung, Haesol Park, Hyungjoo Jung, Junghyun Cho
WACV4
2023 MAIR: Multi-View Attention Inverse Rendering with 3D Spatially-Varying Lighting Estimation
abstract
We propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, a SVBRDF, and 3D spatially-varying lighting. Because multi-view images provide a variety of information about the scene, multi-view images in object-level inverse rendering have been taken for granted. However, owing to the absence of multi-view HDR synthetic dataset, scene-level inverse rendering has mainly been studied using single-view image. We were able to successfully perform scene-level inverse rendering using multi-view images by expanding OpenRooms dataset and designing efficient pipelines to handle multi-view images, and splitting spatially-varying lighting. Our experiments show that the proposed method not only achieves better performance than single-view-based methods, but also achieves robust performance on unseen real-world scene. Also, our sophisticated 3D spatially-varying lighting volume allows for photorealistic object insertion in any 3D location.
Junyong Choi, SeokYeong Lee, Haesol Park, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho
CVPR6
2022 Quality of Satellite Communication Signals Related to Problems in RF Components as Hardware
abstract
As the generation of mobile communication develops as is slated to be hooked up to satellite communication, the importance of digital or software is highlighted more and more as a fad that the notion of SDR will work things out. Communication is a product of impartially integrating the fields of software and hardware. In order to stop the waste of time in finding causes of failure in communication by scrutinizing the software of a system, this paper as a problem diagnosis report shares an idea that problems in the hardware of the system end up with malfunctions in satellite communication. Since the portals of the satellite wireless payload are taken up by the antenna and the filter, which account for forming the wireless link, the consequences of problematic RF components are investigated. Physical errors in the components are brought up to demonstrate a functional degradation like weakened signals.
Sungtek Kahng, Junghyun Cho, Yejune Seo, Yejin Lee 0007, Jiyeon Jang, Hosub Lee 0003
APCC2
2021 A 3d Model-Based Approach For Fitting Masks To Faces In The Wild
abstract
Face recognition now requires a large number of labelled masked face images in the era of this unprecedented COVID19 pandemic. Unfortunately, the rapid spread of the virus has left us little time to prepare for such dataset in the wild. To circumvent this issue, we present a 3D model-based approach called WearMask3D for augmenting face images of various poses to the masked face counterparts. Our method proceeds by first fitting a 3D morphable model on the input image, second overlaying the mask surface onto the face model and warping the respective mask texture, and last projecting the 3D mask back to 2D. The mask texture is adapted based on the brightness and resolution of the input image. By working in 3D, our method can produce more natural masked faces of diverse poses from a single mask texture. To compare precisely between different augmentation approaches, we have constructed a dataset comprising masked and unmasked faces with labels called MFW-mini. Experimental results demonstrate WearMask3D1produces more realistic masked faces, and utilizing these images for training leads to state-of-the-art recognition accuracy for masked faces.
Je Hyeong Hong, Hanjo Kim, Gi Pyo Nam, Junghyun Cho, Hyeong-Seok Ko, Ig-Jae Kim
ICIP5
2020 EdNet: A Large-Scale Hierarchical Dataset in Education
Youngduck Choi, Youngnam Lee, Dongmin Shin, Junghyun Cho, Seoyon Park, Seewoo Lee, Jineon Baek, Chan Bae, Byungsoo Kim 0002, Jaewe Heo
AIED (2)4
2020 Deep Attentive Study Session Dropout Prediction in Mobile Learning Environment
abstract
Student dropout prediction provides an opportunity to improve student engagement, which maximizes the overall effectiveness of learning experiences. However, researches on student dropout were mainly conducted on school dropout or course dropout, and study session dropout in a mobile learning environment has not been considered thoroughly. In this paper, we investigate the study session dropout prediction problem in a mobile learning environment. First, we define the concept of the study session, study session dropout and study session dropout prediction task in a mobile learning environment. Based on the definitions, we propose a novel Transformer based model for predicting study session dropout, DAS: Deep Attentive Study Session Dropout Prediction in Mobile Learning Environment. DAS has an encoder-decoder structure which is composed of stacked multi-head attention and point-wise feed-forward networks. The deep attentive computations in DAS are capable of capturing complex relations among dynamic student interactions. To the best of our knowledge, this is the first attempt to investigate study session dropout in a mobile learning environment. Empirical evaluations on a large-scale dataset show that DAS achieves the best performance with a significant improvement in area under the receiver operating characteristic curve compared to baseline models.
Youngnam Lee, Dongmin Shin, Hyunbin Loh, Piljae Chae, Junghyun Cho, Seoyon Park, Jinhwan Lee, Jineon Baek, Byungsoo Kim 0002, Youngduck Choi
CSEDU (1)6
2020 Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint Estimation
abstract
Existing techniques to encode spatial invariance within deep convolutional neural networks only model 2D transformation fields. This does not account for the fact that objects in a 2D space are a projection of 3D ones, and thus they have limited ability to severe object viewpoint changes. To overcome this limitation, we introduce a learnable module, cylindrical convolutional networks (CCNs), that exploit cylindrical representation of a convolutional kernel defined in the 3D space. CCNs extract a view-specific feature through a view-specific convolutional kernel to predict object category scores at each viewpoint. With the view-specific feature, we simultaneously determine objective category and viewpoints using the proposed sinusoidal soft-argmax module. Our experiments demonstrate the effectiveness of the cylindrical convolutional networks on joint object detection and viewpoint estimation.
Sunghun Joung, Seungryong Kim, Hanjae Kim, Ig-Jae Kim, Junghyun Cho, Kwanghoon Sohn
CVPR6
2020 Towards an Appropriate Query, Key, and Value Computation for Knowledge Tracing
abstract
In this paper, we propose a novel Transformer-based model for knowledge tracing, SAINT: Separated Self-AttentIve Neural Knowledge Tracing. SAINT has an encoder-decoder structure where the exercise and response embedding sequences separately enter, respectively, the encoder and the decoder. The encoder applies self-attention layers to the sequence of exercise embeddings, and the decoder alternately applies self-attention layers and encoder-decoder attention layers to the sequence of response embeddings. This separation of input allows us to stack attention layers multiple times, resulting in an improvement in area under receiver operating characteristic curve (AUC). To the best of our knowledge, this is the first work to suggest an encoder-decoder model for knowledge tracing that applies deep self-attentive layers to exercises and responses separately. We empirically evaluate SAINT on a large-scale knowledge tracing dataset, EdNet, collected by an active mobile education application, Santa, which has 627,347 users, 72,907,005 response data points as well as a set of 16,175 exercises gathered since 2016. The results show that SAINT achieves state-of-the-art performance in knowledge tracing with an improvement of 1.8% in AUC compared to the current state-of-the-art model.
Youngduck Choi, Youngnam Lee, Junghyun Cho, Jineon Baek, Byungsoo Kim 0002, Yeongmin Cha, Dongmin Shin, Chan Bae, Jaewe Heo
L@S3
2016 3D Modeling from Photos Given Topological Information
abstract
Reconstructing 3D models given a single-view 2D information is inherently an ill-posed problem and requires additional information such as shape prior or user input.We introduce a method to generate multiple 3D models of a particular category given corresponding photographs when the topological information is known. While there is a wide range of shapes for an object of a particular category, the basic topology usually remains constant.In consequence, the topological prior needs to be provided only once for each category and can be easily acquired by consulting an existing database of 3D models or by user input. The input of topological description is only connectivity information between parts; this is in contrast to previous approaches that have required users to interactively mark individual parts. Given the silhouette of an object and the topology, our system automatically finds a skeleton and generates a textured 3D model by jointly fitting multiple parts. The proposed method, therefore, opens the possibility of generating a large number of 3D models by consulting a massive number of photographs. We demonstrate examples of the topological prior and reconstructed 3D models using photos.
Young Min Kim 0001, Junghyun Cho, Sang Chul Ahn
IEEE Trans. Vis. Comput. Graph.2
2013 Geometry-Aware Volume-of-Fluid Method
abstract
Abstract We present a new framework to simulate moving interfaces in viscous incompressible two phase flows. The goal is to achieve both conservation of the fluid volume and a detailed reconstruction of the fluid surface. To these ends, we incorporate sub‐grid refinement of the level set with the volume‐of‐fluid method. In the context of this refined level set grid we propose the algorithms needed for the coupling of the level set and the volume‐of‐fluid, which include techniques for computing volume, redistancing the level set, and handling surface tension. We report the experimental results produced with the proposed method via simulations of the two phase fluid phenomena such as air‐cushioning and deforming large bubbles.
Junghyun Cho, Hyeong-Seok Ko
Comput. Graph. Forum1
2007 A Low Power Phase-Change Random Access Memory using a Data-Comparison Write Scheme
abstract
A low power PRAM using a data-comparison write (DCW) scheme is proposed. The PRAM consumes large write power because large write currents are required during long time. At first, the DCW scheme reads a stored data during write operation. And then, it writes an input data only when the input and stored data are different. Therefore, it can reduce the write power consumption to a half. The 1K-bit PRAM test chip with 128×8bits is implemented with a 0.8μm CMOS technology with a 0.5μm GST cell.
Byung-Do Yang, Jae-Eun Lee, Jang-Su Kim, Junghyun Cho, Seung-Yun Lee, Byoung-Gon Yu
ISCAS4
2006 e-AIRS: An e-Science Collaboration Portal for Aerospace Applications
Yoonhee Kim, Jeuyoung Kim, Junghyun Cho, Chongam Kim, Kum Won Cho
HPCC4
2005 An analog front-end IP for 13.56MHz RFID interrogators
abstract
An analog front-end circuit for 13.56MHz RFID interrogators compatible with ISO14443, ISO15693 and ISO18000-3 Mode 1 RFID interrogators was designed and fabricated by using 0.35μm double poly CMOS process. The fabricated chip was operated at 3.3 volt single supply. The results of this work can be provided as reusable IPs in a form of hard or firm IPs for designing single chip 13.56MHz RFID interrogators.
Junghyun Cho, Suk-Byung Chai, Chung-Gi Song, Kyung-Won Min, Shiho Kim
ASP-DAC1