EDBT 2026 Demo / reviewers in the wild / expert
Sujin Jang
dblp:146/6241
· DBLP profile ↗
13ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-2723-5606ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
3D vision · 37% Transfer learning and domain adaptation · 24% Efficient and distributed learning · 8% | |
| Human-computer interaction and pervasive computing
4 papers |
Interaction techniques and input · 65% Usability and user experience research · 18% Human-robot interaction · 17% |
Topics — the 27 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d object detection |
2.2 | 3 | 2024 | Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection · NeurIPS 2024 CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection · AAAI 2024 STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object Detection · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation
test-time adaptation |
1.7 | 2 | 2025 | Active Test-time Vision-Language Navigation · NeurIPS 2025 Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement Learning · ICML 2025 |
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
multi-view 3d object detection |
1.4 | 2 | 2024 | Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection · NeurIPS 2024 STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object Detection · NeurIPS 2023 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
1.3 | 2 | 2024 | CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection · AAAI 2024 DaDA: Distortion-aware Domain Adaptation for Unsupervised Semantic Segmentation · NeurIPS 2022 |
Computer vision › Vision and language
vision-and-language navigation |
1.1 | 2 | 2025 | Active Test-time Vision-Language Navigation · NeurIPS 2025 Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement Learning · ICML 2025 |
Interaction techniques and input › gesture input
mid-air interaction |
0.9 | 2 | 2023 | Advanced modeling method for quantifying cumulative subjective fatigue in mid-air interaction · Int. J. Hum. Comput. Stud. 2023 Modeling Cumulative Arm Fatigue in Mid-Air Interaction based on Perceived Exertion and Kinetics of Arm Motion · CHI 2017 |
Computer vision › 3D vision
3d scene understanding |
0.9 | 1 | 2025 | 3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation · CVPR 2025 |
Computer vision › 3D vision › 3d scene understanding
semantic scene completion |
0.9 | 1 | 2025 | 3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation · CVPR 2025 |
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty calibration |
0.9 | 1 | 2025 | Active Test-time Vision-Language Navigation · NeurIPS 2025 |
Computer vision › 3D vision
view transformation |
0.9 | 1 | 2025 | 3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation · CVPR 2025 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
cross-modal domain adaptation |
0.8 | 1 | 2024 | CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection · AAAI 2024 |
Machine learning › Transfer learning and domain adaptation
domain adaptation and generalization |
0.8 | 1 | 2024 | Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection · NeurIPS 2024 |
Robotics › Autonomous driving
HD map construction |
0.8 | 1 | 2024 | Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and Propagation · NeurIPS 2024 |
Computer vision › 3D vision › 3d object detection › point cloud object detection
LiDAR-based 3D object detection |
0.8 | 1 | 2024 | CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection · AAAI 2024 |
Computer vision › Video understanding and tracking › temporal modeling
temporal fusion |
0.8 | 1 | 2024 | Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and Propagation · NeurIPS 2024 |
Robotics › Autonomous driving › HD map construction
vectorized map construction |
0.8 | 1 | 2024 | Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and Propagation · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
cross-modal distillation |
0.7 | 1 | 2023 | STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object Detection · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.7 | 1 | 2023 | STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object Detection · NeurIPS 2023 |
Computer vision › Segmentation and scene understanding › semantic segmentation › transfer learning for semantic segmentation
domain adaptive semantic segmentation |
0.6 | 1 | 2022 | DaDA: Distortion-aware Domain Adaptation for Unsupervised Semantic Segmentation · NeurIPS 2022 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.6 | 1 | 2022 | DaDA: Distortion-aware Domain Adaptation for Unsupervised Semantic Segmentation · NeurIPS 2022 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2025 | 3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation · CVPR 2025 |
Visualization and visual analytics › data visualization › animated visualization
motion visualization |
0.2 | 1 | 2016 | MotionFlow: Visual Abstraction and Aggregation of Sequential Patterns in Human Motion Tracking Data · IEEE Trans. Vis. Comput. Graph. 2016 |
Visualization and visual analytics
visual analytics |
0.2 | 1 | 2016 | MotionFlow: Visual Abstraction and Aggregation of Sequential Patterns in Human Motion Tracking Data · IEEE Trans. Vis. Comput. Graph. 2016 |
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation |
0.2 | 1 | 2015 | A Collaborative Filtering Approach to Real-Time Hand Pose Estimation · ICCV 2015 |
Usability and user experience research
user experience evaluation |
0.1 | 1 | 2017 | Modeling Cumulative Arm Fatigue in Mid-Air Interaction based on Perceived Exertion and Kinetics of Arm Motion · CHI 2017 |
Interaction techniques and input
gesture input |
0.1 | 1 | 2016 | MotionFlow: Visual Abstraction and Aggregation of Sequential Patterns in Human Motion Tracking Data · IEEE Trans. Vis. Comput. Graph. 2016 |
Recommender systems
collaborative filtering |
0.1 | 1 | 2015 | A Collaborative Filtering Approach to Real-Time Hand Pose Estimation · ICCV 2015 |
Methods — techniques the papers use, named apart from their topics
self-active learning · 1.7mixture entropy optimization · 1.7entropy minimization · 1.7prototype learning · 0.9gradient regularization · 0.9cross-attention · 0.9binary episodic feedback · 0.9self-training · 0.8label-efficient adaptation · 0.8adversarial training · 0.8cumulative fatigue modeling · 0.7partition-based clustering · 0.5flow visualization · 0.5cross-validation · 0.3biomechanical modeling · 0.3nearest neighbor search · 0.2matrix factorization · 0.2local shape descriptors · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View TransformationabstractThe resolution of voxel queries significantly influences the quality of view transformation in camera-based 3D occupancy prediction. However, computational constraints and the practical necessity for real-time deployment require smaller query resolutions, which inevitably leads to an information loss. Therefore, it is essential to encode and preserve rich visual details within limited query sizes while ensuring a comprehensive representation of 3D occupancy. To this end, we introduce ProtoOcc, a novel occupancy network that leverages prototypes of clustered image segments in view transformation to enhance low-resolution context. In particular, the mapping of 2D prototypes onto 3D voxel queries encodes high-level visual geometries and complements the loss of spatial information from reduced query resolutions. Additionally, we design a multi-perspective decoding strategy to efficiently disentangle the densely compressed visual cues into a high-dimensional 3D occupancy scene. Experimental results on both Occ3D and SemanticKITTI benchmarks demonstrate the effectiveness of the proposed method, showing clear improvements over the baselines. More importantly, ProtoOcc achieves competitive performance against the baselines even with 75% reduced voxel resolution. Project page: https://kuai-lab.github.io/cvpr2025protoocc. Gyeongrok Oh, Sungjune Kim, Heeju Ko, Hyung-gun Chi, Jinkyu Kim 0001, Daehyun Ji, Sujin Jang, Sangpil Kim |
CVPR | 9 |
| 2025 | Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement LearningabstractNavigating in an unfamiliar environment during deployment poses a critical challenge for a vision-language navigation (VLN) agent. Yet, test-time adaptation (TTA) remains relatively underexplored in robotic navigation, leading us to the fundamental question: what are the key properties of TTA for online VLN? In our view, effective adaptation requires three qualities: 1) flexibility in handling different navigation outcomes, 2) interactivity with external environment, and 3) maintaining a harmony between plasticity and stability. To address this, we introduce FeedTTA, a novel TTA framework for online VLN utilizing feedback-based reinforcement learning. Specifically, FeedTTA learns by maximizing binary episodic feedback, a practical setup in which the agent receives a binary scalar after each episode that indicates the success or failure of the navigation. Additionally, we propose a gradient regularization technique that leverages the binary structure of FeedTTA to achieve a balance between plasticity and stability during adaptation. Our extensive experiments on challenging VLN benchmarks demonstrate the superior adaptability of FeedTTA, even outperforming the state-of-the-art offline training methods in REVERIE benchmark with a single stream of learning. Sungjune Kim, Gyeongrok Oh, Heeju Ko, Daehyun Ji, Sujin Jang, Sangpil Kim |
ICML | 7 |
| 2025 | Active Test-time Vision-Language NavigationabstractVision-Language Navigation (VLN) policies trained on offline datasets often exhibit degraded task performance when deployed in unfamiliar navigation environments at test time, where agents are typically evaluated without access to external interaction or feedback. Entropy minimization has emerged as a practical solution for reducing prediction uncertainty at test time; however, it can suffer from accumulated errors, as agents may become overconfident in incorrect actions without sufficient contextual grounding. To tackle these challenges, we introduce ATENA (Active TEst-time Navigation Agent), a test-time active learning framework that enables a practical human-robot interaction via episodic feedback on uncertain navigation outcomes. In particular, ATENA learns to increase certainty in successful episodes and decrease it in failed ones, improving uncertainty calibration. Here, we propose mixture entropy optimization, where entropy is obtained from a combination of the action and pseudo-expert distributions—a hypothetical action distribution assuming the agent's selected action to be optimal—controlling both prediction confidence and action preference. In addition, we propose a self-active learning strategy that enables an agent to evaluate its navigation outcomes based on confident predictions. As a result, the agent stays actively engaged throughout all iterations, leading to well-grounded and adaptive decision-making. Extensive evaluations on challenging VLN benchmarks—REVERIE, R2R, and R2R-CE—demonstrate that ATENA successfully overcomes distributional shifts at test time, outperforming the compared baseline methods across various settings. Heeju Ko, Sung June Kim, Gyeongrok Oh, Jeongyoon Yoon, Honglak Lee, Sujin Jang, Seungryong Kim, Sangpil Kim |
NeurIPS | 6 |
| 2024 | CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object DetectionabstractRecent LiDAR-based 3D Object Detection (3DOD) methods show promising results, but they often do not generalize well to target domains outside the source (or training) data distribution. To reduce such domain gaps and thus to make 3DOD models more generalizable, we introduce a novel unsupervised domain adaptation (UDA) method, called CMDA, which (i) leverages visual semantic cues from an image modality (i.e., camera images) as an effective semantic bridge to close the domain gap in the cross-modal Bird's Eye View (BEV) representations. Further, (ii) we also introduce a self-training-based learning strategy, wherein a model is adversarially trained to generate domain-invariant features, which disrupt the discrimination of whether a feature instance comes from a source or an unseen target domain. Overall, our CMDA framework guides the 3DOD model to generate highly informative and domain-adaptive features for novel data distributions. In our extensive experiments with large-scale benchmarks, such as nuScenes, Waymo, and KITTI, those mentioned above provide significant performance gains for UDA tasks, achieving state-of-the-art performance. Gyusam Chang, Wonseok Roh, Sujin Jang, Daehyun Ji, Gyeongrok Oh, Jinsun Park, Jinkyu Kim 0001, Sangpil Kim |
AAAI | 3 |
| 2024 | Unified Domain Generalization and Adaptation for Multi-View 3D Object DetectionabstractRecent advances in 3D object detection leveraging multi-view cameras have demonstrated their practical and economical value in various challenging vision tasks.
However, typical supervised learning approaches face challenges in achieving satisfactory adaptation toward unseen and unlabeled target datasets (i.e., direct transfer) due to the inevitable geometric misalignment between the source and target domains.
In practice, we also encounter constraints on resources for training models and collecting annotations for the successful deployment of 3D object detectors.
In this paper, we propose Unified Domain Generalization and Adaptation (UDGA), a practical solution to mitigate those drawbacks.
We first propose Multi-view Overlap Depth Constraint that leverages the strong association between multi-view, significantly alleviating geometric gaps due to perspective view changes.
Then, we present a Label-Efficient Domain Adaptation approach to handle unfamiliar targets with significantly fewer amounts of labels (i.e., 1$\%$ and 5$\%)$, while preserving well-defined source knowledge for training efficiency.
Overall, UDGA framework enables stable detection performance in both source and target domains, effectively bridging inevitable domain gaps, while demanding fewer annotations.
We demonstrate the robustness of UDGA with large-scale benchmarks: nuScenes, Lyft, and Waymo, where our framework outperforms the current state-of-the-art methods. Gyusam Chang, Donghyun Kim 0006, Jinkyu Kim 0001, Daehyun Ji, Sujin Jang, Sangpil Kim |
NeurIPS | 7 |
| 2024 | Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and PropagationabstractPredicting and constructing road geometric information (e.g., lane lines, road markers) is a crucial task for safe autonomous driving, while such static map elements can be repeatedly occluded by various dynamic objects on the road. Recent studies have shown significantly improved vectorized high-definition (HD) map construction performance, but there has been insufficient investigation of temporal information across adjacent input frames (i.e., clips), which may lead to inconsistent and suboptimal prediction results. To tackle this, we introduce a novel paradigm of clip-level vectorized HD map construction, MapUnveiler, which explicitly unveils the occluded map elements within a clip input by relating dense image representations with efficient clip tokens. Additionally, MapUnveiler associates inter-clip information through clip token propagation, effectively utilizing long- term temporal map information. MapUnveiler runs efficiently with the proposed clip-level pipeline by avoiding redundant computation with temporal stride while building a global map relationship. Our extensive experiments demonstrate that MapUnveiler achieves state-of-the-art performance on both the nuScenes and Argoverse2 benchmark datasets. We also showcase that MapUnveiler significantly outperforms state-of-the-art approaches in a challenging setting, achieving +10.7% mAP improvement in heavily occluded driving road scenes. The project page can be found at https://mapunveiler.github.io. Nayeon Kim 0006, Hongje Seong, Daehyun Ji, Sujin Jang |
NeurIPS | 4 |
| 2023 | STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object Detectionabstract3D object detection (3DOD) from multi-view images is an economically appealing alternative to expensive LiDAR-based detectors, but also an extremely challenging task due to the absence of precise spatial cues. Recent studies have leveraged the teacher-student paradigm for cross-modal distillation, where a strong LiDAR-modality teacher transfers useful knowledge to a multi-view-based image-modality student. However, prior approaches have only focused on minimizing global distances between cross-modal features, which may lead to suboptimal knowledge distillation results. Based on these insights, we propose a novel structural and temporal cross-modal knowledge distillation (STXD) framework for multi-view 3DOD. First, STXD reduces redundancy of the feature components of the student by regularizing the cross-correlation of cross-modal features, while maximizing their similarities. Second, to effectively transfer temporal knowledge, STXD encodes temporal relations of features across a sequence of frames via similarity maps. Lastly, STXD also adopts a response distillation method to further enhance the quality of knowledge distillation at the output-level. Our extensive experiments demonstrate that STXD significantly improves the NDS and mAP of the based student detectors by 2.8%~4.5% on the nuScenes testing dataset. Sujin Jang, Sung Ju Hwang, Daehyun Ji |
NeurIPS | 1 |
| 2023 | Advanced modeling method for quantifying cumulative subjective fatigue in mid-air interaction
Ana M. Villanueva, Sujin Jang, Wolfgang Stuerzlinger, Satyajit Ambike, Karthik Ramani |
Int. J. Hum. Comput. Stud. | 2 |
| 2022 | DaDA: Distortion-aware Domain Adaptation for Unsupervised Semantic SegmentationabstractDistributional shifts in photometry and texture have been extensively studied for unsupervised domain adaptation, but their counterparts in optical distortion have been largely neglected. In this work, we tackle the task of unsupervised domain adaptation for semantic image segmentation where unknown optical distortion exists between source and target images. To this end, we propose a distortion-aware domain adaptation (DaDA) framework that boosts the unsupervised segmentation performance. We first present a relative distortion learning (RDL) approach that is capable of modeling domain shifts in fine-grained geometric deformation based on diffeomorphic transformation. Then, we demonstrate that applying additional global affine transformations to the diffeomorphically transformed source images can further improve the segmentation adaptation. Besides, we find that our distortion-aware adaptation method helps to enhance self-supervised learning by providing higher-quality initial models and pseudo labels. To evaluate, we propose new distortion adaptation benchmarks, where rectilinear source images and fisheye target images are used for unsupervised domain adaptation. Extensive experimental results highlight the effectiveness of our approach over state-of-the-art methods under unknown relative distortion across domains. Datasets and more information are available at https://sait-fdd.github.io/. Sujin Jang, Joohan Na, Dokwan Oh |
NeurIPS | 1 |
| 2017 | Modeling Cumulative Arm Fatigue in Mid-Air Interaction based on Perceived Exertion and Kinetics of Arm MotionabstractQuantifying cumulative arm muscle fatigue is a critical factor in understanding, evaluating, and optimizing user experience during prolonged mid-air interaction. A reasonably accurate estimation of fatigue requires an estimate of an individual's strength. However, there is no easy-to-access method to measure individual strength to accommodate inter-individual differences. Furthermore, fatigue is influenced by both psychological and physiological factors, but no current HCI model provides good estimates of cumulative subjective fatigue. We present a new, simple method to estimate the maximum shoulder torque through a mid-air pointing task, which agrees with direct strength measurements. We then introduce a cumulative fatigue model informed by subjective and biomechanical measures. We evaluate the performance of the model in estimating cumulative subjective fatigue in mid-air interaction by performing multiple cross-validations and a comparison with an existing fatigue metric. Finally, we discuss the potential of our approach for real-time evaluation of subjective fatigue as well as future challenges. Sujin Jang, Wolfgang Stuerzlinger, Satyajit Ambike, Karthik Ramani |
CHI | 1 |
| 2016 | MotionFlow: Visual Abstraction and Aggregation of Sequential Patterns in Human Motion Tracking DataabstractPattern analysis of human motions, which is useful in many research areas, requires understanding and comparison of different styles of motion patterns. However, working with human motion tracking data to support such analysis poses great challenges. In this paper, we propose MotionFlow, a visual analytics system that provides an effective overview of various motion patterns based on an interactive flow visualization. This visualization formulates a motion sequence as transitions between static poses, and aggregates these sequences into a tree diagram to construct a set of motion patterns. The system also allows the users to directly reflect the context of data and their perception of pose similarities in generating representative pose states. We provide local and global controls over the partition-based clustering process. To support the users in organizing unstructured motion data into pattern groups, we designed a set of interactions that enables searching for similar motion sequences from the data, detailed exploration of data subsets, and creating and modifying the group of motion patterns. To evaluate the usability of MotionFlow, we conducted a user study with six researchers with expertise in gesture-based interaction design. They used MotionFlow to explore and organize unstructured motion tracking data. Results show that the researchers were able to easily learn how to use MotionFlow, and the system effectively supported their pattern analysis activities, including leveraging their perception and domain knowledge. Sujin Jang, Niklas Elmqvist, Karthik Ramani |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2015 | A Collaborative Filtering Approach to Real-Time Hand Pose EstimationabstractCollaborative filtering aims to predict unknown user ratings in a recommender system by collectively assessing known user preferences. In this paper, we first draw analogies between collaborative filtering and the pose estimation problem. Specifically, we recast the hand pose estimation problem as the cold-start problem for a new user with unknown item ratings in a recommender system. Inspired by fast and accurate matrix factorization techniques for collaborative filtering, we develop a real-time algorithm for estimating the hand pose from RGB-D data of a commercial depth camera. First, we efficiently identify nearest neighbors using local shape descriptors in the RGB-D domain from a library of hand poses with known pose parameter values. We then use this information to evaluate the unknown pose parameters using a joint matrix factorization and completion (JMFC) approach. Our quantitative and qualitative results suggest that our approach is robust to variation in hand configurations while achieving real time performance (≈ 29 FPS) on a standard computer. Chiho Choi, Ayan Sinha, Joon Hee Choi, Sujin Jang, Karthik Ramani |
ICCV | 4 |
| 2014 | PuppetX: a framework for gestural interactions with user constructed playthingsabstractWe present PuppetX, a framework for both constructing playthings and playing with them using spatial body and hand gestures. This framework allows users to construct various playthings similar to puppets with modular components representing basic geometric shapes. It is topologically-aware, i.e. depending on its configuration; PuppetX automatically determines its own topological construct. Once the plaything is made the users can interact with them naturally via body and hand gestures as detected by depth-sensing cameras. This gives users the freedom to create playthings using our components and the ability to control them using full body interactions. Our framework creates affordances for a new variety of gestural interactions with physically constructed objects. As its by-product, a virtual 3D model is created, which can be animated as a proxy to the physical construct. Our algorithms can recognize hand and body gestures in various configurations of the playthings. Through our work, we push the boundaries of interaction with user-constructed objects using large gestures involving the whole body or fine gestures involving the fingers. We discuss the results of a study to understand how users interact with the playthings and conclude with a demonstration of the abilities of gestural interactions with PuppetX by exploring a variety of interaction scenarios. Saikat Gupta, Sujin Jang, Karthik Ramani |
AVI | 2 |