VLDB 2026 Research / reviewers in the wild / expert
Minsu Jang
dblp:64/4831
· DBLP profile ↗
22ranked-venue papers
7as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 12 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorSystems, architecture and hardware · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Designing Artificial Identity: The Identity Design Framework and Research AgendaabstractThe identity design of artificial agents carries growing ethical, psychological, and cultural weight, as ubiquitous language models and diverse robotic forms are blended into everyday use. However, structured approaches to designing coherent and interpretable artificial identities remain limited. To address urgent challenges in artificial identity design, including harmful stereotypes and deceptive practices, we introduce the Identity Design (ID) Framework and an accompanying research agenda. Drawing on emerging work on artificial identity in human-robot interaction and taking an interdisciplinary perspective, we propose twelve design principles across three levels: individual (recognisability, behavioural consistency, identity continuity, memory, persistent goals), group (membership signalling, social alignment, role clarity), and societal (benevolence, artificiality, social justice, transparency). The research agenda outlines open questions around the operationalisation and measurement of identity, social dynamics, and ethical considerations for identity design. Together, they lay the groundwork for future research and responsible practice in robotic, virtual, and multi-embodied agents. Karla Bransky, Penny Kyburz, Patrick Holthaus, Guy Laban, Katie Winkle, Neziha Akalin, Ashita Ashok, Rucha Khot, Alexandra Bejarano, Jorrit Thijn, Roger K. Moore, Minsu Jang, Joel E. Fischer, Minha Lee |
DIS | 13 |
| 2025 | Adaptive Robot-Mediated Assessment using LLM for Enhanced Survey Quality in Older Adults Care ProgramsabstractThis study presents an adaptive human-robot interaction (HRI) system that evaluates older adult participants' satisfaction with personalized health care programs. By integrating the CLOi robot with a large language model (LLM), the system conducts satisfaction surveys that adapt in real-time to participant responses. The system was applied to evaluate healthcare programs that include physical health measurements, exercise assessments, and virtual reality (VR) experiences. The system utilizes the CLOi robot and Claude API to analyze response clarity in real-time, automatically generating contextually appropriate follow-up questions when responses are deemed ambiguous. This adaptive questioning strategy ensures comprehensive response quality before proceeding to subsequent survey items. We conducted a preliminary feasibility study with five older adult participants to evaluate our approach. The system leverages LLM prompts to analyze gaps between question intent and participant responses, generating targeted follow-up questions as needed. Results demonstrate that our LLM-enhanced robotic interview system effectively reduced response ambiguity through dynamic follow-up questioning, achieving an 85% response resolution rate. This adaptive approach improved the clarity and specificity of healthcare satisfaction assessments for older adults. Cheonshu Park, Miyoung Cho, Minjung Shin, Jeh-Kwang Ryu, Minsu Jang |
HRI | 5 |
| 2024 | LoTa-Bench: Benchmarking Language-oriented Task Planners for Embodied AgentsabstractLarge language models (LLMs) have recently received considerable attention as alternative solutions for task planning. However, comparing the performance of language-oriented task planners becomes difficult, and there exists a dearth of detailed exploration regarding the effects of various factors such as pre-trained model selection and prompt construction. To address this, we propose a benchmark system for automatically quantifying performance of task planning for home-service embodied agents. Task planners are tested on two pairs of datasets and simulators: 1) ALFRED and AI2-THOR, 2) an extension of Watch-And-Help and VirtualHome. Using the proposed benchmark system, we perform extensive experiments with LLMs and prompts, and explore several enhancements of the baseline planner. We expect that the proposed benchmark tool would accelerate the development of language-oriented task planners. Jaewoo Choi 0001, Youngwoo Yoon, Hyobin Ong, Jaehong Kim 0001, Minsu Jang |
ICLR | 5 |
| 2024 | Rotational Angle Estimation Using an Acceleration Sensor Array and Real-Time Detection AlgorithmabstractThe instantaneous center of rotation (ICOR) is essential engineering information, providing wheel slip data for fully autonomous skid-steering mobile robots driven by deep learning-based dead-reckoning algorithms. Conventionally, cost-effective and compact inertial measurement unit (IMU) sensors based on microelectromechanical systems (MEMSs) have been employed for ICOR measurement. However, these sensors inherently rely on the differential value of gyroscope data, leading to significant estimation errors due to the amplified noise at low sampling rates. To address this challenge, we developed an acceleration sensor array based on a fiber Bragg grating, inspired by existing MEMS acceleration sensor arrays, which allows ICOR estimation solely from the acceleration data provided by two sensors at the beginning of rotation. The estimation accuracies in the clockwise and counter-clockwise directions were measured to be 3.1 ± 5.1 mm and 3.7 ± 9.9 mm, respectively. Furthermore, a peak detection algorithm was proposed for extracting the peak acceleration under different ICOR conditions using the variance and Bollinger bands. The precision level was comparable to those of optical cameras and MEMS IMU sensors that use computationally intensive algorithms, such as Kalman filters, even without differentiation. This breakthrough in measurement capability opens up the possibility for real-time ICOR estimation, allowing the complete automation of skid-steering-based robots and benefiting fields constrained by the structural limitations of conventional MEMS IMU sensors. Byung Kook Kim, Hyowon Moon, Minsu Jang, Jinseok Kim 0002 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | A Structured Prompting based on Belief-Desire-Intention Model for Proactive and Explainable Task PlanningabstractWe investigate the potential of the belief-desire-intention (BDI) model for enhancing proactive action planning and transparency in large language models (LLMs). Our proposed method, BDIPrompting, integrates the knowledge representation framework of the BDI model into prompt design. This allows agents to generate motivational and goal-directed service plans proactively while offering human users insights into the rationale behind the decision-making process. Through preliminary experiments with OpenAI’s GPT-4, we highlight the effectiveness of our approach in planning motivational actions and providing improved explanations during human-agent interactions. Minsu Jang, Youngwoo Yoon, Jaewoo Choi 0001, Hyobin Ong, Jaehong Kim 0001 |
HAI | 1 |
| 2022 | Real-world Validation Study of Daily Activity Detection for the ElderlyabstractThe development of artificial intelligence has led to significant progress in activity detection; however, a proper evaluation of the activity detection performance is not guaranteed and is still lacking in real-life environments. In this study, to verify the stability and usefulness of activity detection in a real-world setting, we analyzed the activity detection performance for 40 elderly people using a human care robot. Miyoung Cho, Jinhyeok Jang, Jaeyeon Lee 0001, Minsu Jang, Do-Hyung Kim 0004, Jaehong Kim 0001 |
RO-MAN | 4 |
| 2021 | SGToolkit: An Interactive Gesture Authoring Toolkit for Embodied Conversational AgentsabstractNon-verbal behavior is essential for embodied agents like social robots, virtual avatars, and digital humans. Existing behavior authoring approaches including keyframe animation and motion capture are too expensive to use when there are numerous utterances requiring gestures. Automatic generation methods show promising results, but their output quality is not satisfactory yet, and it is hard to modify outputs as a gesture designer wants. We introduce a new gesture generation toolkit, named SGToolkit, which gives a higher quality output than automatic methods and is more efficient than manual authoring. For the toolkit, we propose a neural generative model that synthesizes gestures from speech and accommodates fine-level pose controls and coarse-level style controls from users. The user study with 24 participants showed that the toolkit is favorable over manual authoring, and the generated gestures were also human-like and appropriate to input speech. The SGToolkit is platform agnostic, and the code is available at https://github.com/ai4r/SGToolkit. Youngwoo Yoon, Keun-Woo Park, Minsu Jang, Jaehong Kim 0001, Geehyuk Lee |
UIST | 3 |
| 2020 | ETRI-Activity3D: A Large-Scale RGB-D Dataset for Robots to Recognize Daily Activities of the ElderlyabstractDeep learning, based on which many modern algorithms operate, is well known to be data-hungry. In particular, the datasets appropriate for the intended application are difficult to obtain. To cope with this situation, we introduce a new dataset called ETRI-Activity3D, focusing on the daily activities of the elderly in robot-view. The major characteristics of the new dataset are as follows: 1) practical action categories that are selected from the close observation of the daily lives of the elderly; 2) realistic data collection, which reflects the robot's working environment and service situations; and 3) a large-scale dataset that overcomes the limitations of the current 3D activity analysis benchmark datasets. The proposed dataset contains 112,620 samples including RGB videos, depth maps, and skeleton sequences. During the data acquisition, 100 subjects were asked to perform 55 daily activities. Additionally, we propose a novel network called four-stream adaptive CNN (FSA-CNN). The proposed FSA-CNN has three main properties: robustness to spatio-temporal variations, input-adaptive activation function, and extension of the conventional two-stream approach. In the experiment section, we confirmed the superiority of the proposed FSA-CNN using NTU RGB+D and ETRI-Activity3D. Further, the domain difference between both groups of age was verified experimentally. Finally, the extension of FSA-CNN to deal with the multimodal data was investigated. Jinhyeok Jang, Do-Hyung Kim 0004, Cheonshu Park, Minsu Jang, Jaeyeon Lee 0001, Jaehong Kim 0001 |
IROS | 4 |
| 2020 | End-to-End Learning of Social Behaviors for Humanoid RobotsabstractSocial robots should understand the user's nonverbal behavior and respond appropriately. Machine learning is one way of implementing the social intelligence. It provides the ability to automatically learn and improve from experience instead of explicitly telling the robot what to do. This paper proposes an end-to-end machine learning method to learn social behaviors for humanoid robots. We adapt sequence-to-sequence architecture consisting of two long short-term memory (LSTM) units. One is an LSTM encoder for encoding the previous sequence of human poses, and the other is an LSTM decoder for generating the next sequence of robot poses. The weights of the LSTMs are trained using human-human interaction data such as greeting and handshaking. The trained model is implemented in a humanoid robot, Pepper, to show its feasibility. Experimental results show that the robot can generate gestures appropriate to the situation and recognize subtle differences in user behavior. In addition, when a user's behavior changes, the transition to another behavior occurs naturally. Woo-Ri Ko, Jaeyeon Lee 0001, Minsu Jang, Jaehong Kim 0001 |
SMC | 3 |
| 2020 | Speech gesture generation from the trimodal context of text, audio, and speaker identityabstractFor human-like agents, including virtual avatars and social robots, making proper gestures while speaking is crucial in human-agent interaction. Co-speech gestures enhance interaction experiences and make the agents look alive. However, it is difficult to generate human-like gestures due to the lack of understanding of how people gesture. Data-driven approaches attempt to learn gesticulation skills from human demonstrations, but the ambiguous and individual nature of gestures hinders learning. In this paper, we present an automatic gesture generation model that uses the multimodal context of speech text, audio, and speaker identity to reliably generate gestures. By incorporating a multimodal context and an adversarial training scheme, the proposed model outputs gestures that are human-like and that match with speech content and rhythm. We also introduce a new quantitative evaluation metric for gesture generation models. Experiments with the introduced metric and subjective human evaluation showed that the proposed gesture generation model is better than existing end-to-end generation models. We further confirm that our model is able to work with synthesized audio in a scenario where contexts are constrained, and show that different gesture styles can be generated for the same speech by specifying different speaker identities in the style embedding space that is learned from videos of various speakers. All the code and data is available at https://github.com/ai4r/Gesture-Generation-from-Trimodal-Context. Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee 0001, Jaehong Kim 0001, Geehyuk Lee |
ACM Trans. Graph. | 4 |
| 2019 | Social Human-Robot Interaction of Human-Care Service RobotsabstractService robots with social intelligence are starting to be integrated into our everyday lives. The robots are intended to help improve aspects of quality of life as well as improve efficiency. We are organizing an exciting workshop at HRI 2019 that is oriented towards sharing the ideas amongst participants with diverse backgrounds ranging from Human-Robot Interaction design, social intelligence, decision making, social psychology and aspects and robotic social skills. The purpose of this workshop is to explore how social robots can interact with humans socially and facilitate the integration of social robots into our daily lives. This workshop focuses on three social aspects of human-robot interaction: (1) technical implementation of social robots and products, (2) form, function and behavior, and (3) human behavior and expectations as a means to understand the social aspects of interacting with these robots and products. Ho Seok Ahn, Jongsuk Choi, Hyungpil Moon, Minsu Jang, Sonya S. Kwak, Yoonseob Lim |
HRI | 4 |
| 2019 | Robots Learn Social Skills: End-to-End Learning of Co-Speech Gesture Generation for Humanoid RobotsabstractCo-speech gestures enhance interaction experiences between humans as well as between humans and robots. Most existing robots use rule-based speech-gesture association, but this requires human labor and prior knowledge of experts to be implemented. We present a learning-based co-speech gesture generation that is learned from 52 h of TED talks. The proposed end-to-end neural network model consists of an encoder for speech text understanding and a decoder to generate a sequence of gestures. The model successfully produces various gestures including iconic, metaphoric, deictic, and beat gestures. In a subjective evaluation, participants reported that the gestures were human-like and matched the speech content. We also demonstrate a co-speech gesture with a NAO robot working in real time. Youngwoo Yoon, Woo-Ri Ko, Minsu Jang, Jaeyeon Lee 0001, Jaehong Kim 0001, Geehyuk Lee |
ICRA | 3 |
| 2019 | Development of Wearable Motion Capture System Using Fiber Bragg Grating Sensors for Measuring Arm MotionabstractMotion capture systems are gaining much attention in various fields, including entertainment, medical and sports fields. Although many types of motion capture sensor have been emerging, they have limitations and disadvantages such as occlusion, drift and interference by electromagnetic fields. Here, we introduce the novel wearable motion capture system using fiber Bragg gratings (FBGs) sensors. Since the human joints have different degrees of freedom (DOF), we developed three types of sensors to reconstruct the human body motion from the strains induced on the FBGs. First, a shape sensor using three fibers provides the position and orientation of joints in three dimensional space. Second, we introduce the angle sensor which is capable of measuring bending angle with high curvature using single fiber. Lastly, to detect the twisting of joints, a sensor with fiber attached on a soft material spirally is used. With the optical fiber based motion capture sensors, we reconstruct the motion of arm in realtime. In detail, the joints of the arm include the sternoclavicular, acromioclavicular, shoulder and elbow. By arranging the three types of sensors on the joints in accordance with the DOF, the accuracy of the reconstructed motion is evaluated, resulting in an average error below 2.42°. Finally, to prove the feasibility of applying in virtual reality, we successfully manipulate the virtual avatar in real-time. Minsu Jang, Jun Sik Kim, Kyumin Kang, Soong Ho Um, Sungwook Yang, Jinseok Kim 0002 |
VR | 1 |
| 2014 | Building an automated engagement recognizer based on video analysisabstractThis paper presents a process to build a classifier in a data-driven way for recognizing engagement of children in a robot-based math quiz game. The process consists of collecting video recordings from HRI experiments; annotating the social signals and engagement states via video analysis; extracting feature vectors from the annotations and training classifiers. We conducted an experiment with 7 participants of 10 -- 11 years of age using an android robot EveR-4. With three coders annotating the video recordings and extracting features by snapshot model with 1-second time window, we achieved 84.83% recall performance. Minsu Jang, Cheonshu Park, Hyun-Seung Yang, Jaehong Kim 0001, Young-Jo Cho, Hye-Kyung Cho, Young-Ae Kim, Kyoungwha Chae, Byeong-Kyu Ahn |
HRI | 1 |
| 2013 | Identifying principal social signals in private student-teacher interactions for robot-enhanced educationabstractProviding robots with social intelligence is critical for making entertaining and sustainable human-robot interactions. The first step to get good social intelligence is to appropriately understand the meaning of social signals emitted by interactors. In this paper, we introduce a preliminary study on identifying principal social signals in interpreting participant's engagement and confirmation intention in 1:1 interactions. We annotated 6 video recordings of private teacher-student interactions with 20 social signals and their interpretations, and built pattern data sets with different subsets of social signals. C4.5 based decision trees were generated using the pattern data sets and the recall rates were compared. Also attribute selection was performed to find principal social signals. The results showed that verbal signal was the most principal for determining engagement, and the combination of gaze and verbal signal for confirmation intention. Minsu Jang, Daeha Lee, Jaehong Kim 0001, Young-Jo Cho |
RO-MAN | 1 |
| 2009 | Design and Implementation of a Ubiquitous Robotic SpaceabstractThis paper describes a concerted effort to design and implement a robotic service framework. The proposed framework is comprised of three conceptual spaces: physical, semantic, and virtual spaces, collectively referred to as a ubiquitous robotic space. We implemented a prototype robotic security application in an office environment, which confirmed that the proposed framework is an efficient tool for developing a robotic service employing IT infrastructure, particularly for integrating heterogeneous technologies and robotic platforms. Wonpil Yu, Jae-Yeong Lee, Young-Guk Ha, Minsu Jang, Joo-Chan Sohn, Yong-Moo Kwon, Hyo-Sung Ahn |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2008 | Robot task control utilizing human-in-the-loop perceptionabstractWe propose a network robot application utilizing a human-in-the-loop structure. The proposed structure hybridizes robot perception with perception capability of a human, rendering a robot immune to the uncontrolled environment. Network infrastructure enables the user to intervene in the operation of a robot whenever necessary, enabling seamless transition of task control between the robot and the remote user. We have developed a prototype security robot system by implementing the proposed human-in-the-loop structure. Introduction of network infrastructure, however, also invokes system instability due to network fluctuation, time delay, or jittering. We could resolve the instability by imparting task intelligence to the robot and employing wireless distribution system (WDS) configuration for seamless data communication beyond the coverage of a single access point. Although our approach may solve only part of problems inherent in networked robots, we found that the developed robot system bears sufficient applicability to real world robotic systems incorporating large-area robot navigation. Wonpil Yu, Jae-Yeong Lee, Heesung Chae, Kyuseo Han, Yu-Cheol Lee, Minsu Jang |
RO-MAN | 6 |
| 2007 | Automated Question Answering Using Semantic Web ServicesabstractQuestion answering(QA) can vastly improve the quality and effectiveness of knowledge acquisition on the Web. Web search provides a list of documents targeted based on the keywords coined by the users; users need to investigate further the documents to obtain the very knowledge they wanted. QA can come up with the knowledge itself ready to be conceived directly by the users. We present a framework for building QA systems on the open Web, based on semantic Web services(SWS). In our framework, natural language (NL) questions are translated into a frame-based structure; semantic service discovery is performed to extract the knowledge for answering the questions; and knowledge mediation is performed to formulate answers. Based on SWS framework and Web ontology, knowledge sources can be dynamically expanded, discovered, and mediated for QA. We show the plausibility of our approach by describing an implementation and a step-wise answering scenario with a sample query. Minsu Jang, Joo-Chan Sohn, Hyunkyu Cho |
APSCC | 1 |
| 2007 | Building Semantic Robot Space based on the Semantic WebabstractThe robot space is a ubiquitous environment in which the networked robot plays the role of main mediator that accepts the sensor data from the environment, interprets the data to understand the situation, and decides the actions to perform. Semantic robot space provides the high-level data model of the robot space that enables context-aware service execution. We designed the semantic robot space based on the semantic Web technology, which provided crucial advantages in architecting dynamic reconfiguration mechanism. The standard nature of the technology leverages the interaction among stake-holders in building robot spaces, which makes it possible to establish public registry-based construction of the semantic robot space. We conclude that semantic robot space based on the semantic Web not only embellishes the intelligence of robots but also reduces the cost of installing and maintaining robot spaces. Minsu Jang, Joo-Chan Sohn, Young-Jo Cho |
RO-MAN | 1 |
| 2005 | Automated Teleoperation of Web-Based Devices Using Semantic Web Services
Young-Guk Ha, Jaehong Kim 0001, Minsu Jang, Joo-Chan Sohn, Hyunsoo Yoon |
IEA/AIE | 3 |
| 2005 | MoA: OWL Ontology Merging and Alignment Tool for the Semantic Web
Jaehong Kim 0001, Minsu Jang, Young-Guk Ha, Joo-Chan Sohn, Sang-Jo Lee |
IEA/AIE | 2 |
| 2005 | Ubiquitous robot simulation framework and its applicationsabstractWe describe in this paper a framework, called URSF, for simulating ubiquitous computing environment and ubiquitous robots. URSF provides in/out channel with which ubiquitous robot platforms can be plugged. Once connected to the framework, ubiquitous robot platform can percept and affect the world simulated in the framework. The simulated world is built by composing a simulation space and placing in it operational components such as appliances, sensors, persons, robots etc. Each operational component continuously generates context data, which are fed to the plugged platform. The platform builds world model by interpreting the context data stream. Operational components in the simulation space expose services through semantic Web service scheme, through which plugged platform can affect the simulated world. Semantic Web service scheme enables automated discovery of services dynamically coming and going in the simulation space. The framework can be used to observe how robots interact with environment via ubiquitous network. We discuss two advanced usages of URSF: intelligent robot manipulation and mixed robot-reality manifestation. Minsu Jang, Jaehong Kim 0001, Meekyoung Lee, Joo-Chan Sohn |
IROS | 1 |