EDBT 2026 Demo / reviewers in the wild / expert
Sang-Yun Shin
dblp:237/8407 · also Sangyun Shin
· DBLP profile ↗
12ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0003-0766-4101ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiffRefine: Diffusion-Based Proposal Specific Point Cloud Densification for Cross-Domain Object Detection
Sang-Yun Shin, Xinyu Hou, Samuel Hodgson, Andrew Markham, Agathoniki Trigoni |
ICCV | 1 |
| 2025 | SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic CameraabstractAccurately localizing 3D sound sources and estimating their semantic labels - where the sources may not be visible, but are assumed to lie on the physical surface of objects in the scene - have many real applications, including detecting gas leak and machinery malfunction. The audio-visual weak-correlation in such setting poses new challenges in deriving innovative methods to answer if or how we can use cross-modal information to solve the task. Towards this end, we propose to use an acoustic-camera rig consisting of a pinhole RGB-D camera and a coplanar four-channel microphone array (Mic-Array). By using this rig to record audio-visual signals from multiviews, we can use the cross-modal cues to estimate the sound sources 3D locations. Specifically, our framework SoundLoc3D treats the task as a set prediction problem, each element in the set corresponds to a potential sound source. Given the audio-visual weak-correlation, the set representation is initially learned from a single view microphone array signal, and then refined by actively incorporating physical surface cues revealed from multiview RGB-D images. We demonstrate the efficiency and superiority of SoundLoc3D on large-scale simulated dataset, and further show its robustness to RGB-D measurement inaccuracy and ambient noise interference. Sang-Yun Shin, Anoop Cherian, Agathoniki Trigoni, Andrew Markham |
WACV | 2 |
| 2024 | Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical RepresentationabstractCoarse-to-fine 3D instance segmentation methods show weak performances compared to recent Grouping-based, Kernel-based and Transformer-based methods. We argue that this is due to two limitations: 1) Instance size over-estimation by axis-aligned bounding box(AABB) 2) False negative error accumulation from inaccurate box to the re-finement phase. In this work, we introduce Spherical Mask, a novel coarse-to-fine approach based on spherical repre-sentation, overcoming those two limitations with several benefits. Specifically, our coarse detection estimates each in-stance with a 3D polygon using a center and radial distance predictions, which avoids excessive size estimation of AABB. To cut the error propagation in the existing coarse-to-fine approaches, we virtually migrate points based on the polygon, allowing all foreground points, including false negatives, to be refined. During inference, the proposal and point mi-gration modules run in parallel and are assembled to form binary masks of instances. We also introduce two margin-based losses for the point migration to enforce corrections for the false positives/negatives and cohesion of foreground points, significantly improving the performance. Experimen-tal results from three datasets, such as ScanNetV2, S3DIS, and STPLS3D, show that our proposed method outperforms existing works, demonstrating the effectiveness of the new in-stance representation with spherical coordinates. The code is available at: https://github.com/yunshin/SphericalMask Sang-Yun Shin, Kaichen Zhou, Madhu Vankadari, Andrew Markham, Agathoniki Trigoni |
CVPR | 1 |
| 2024 | Dusk Till Dawn: Self-supervised Nighttime Stereo Depth Estimation using Visual Foundation ModelsabstractSelf-supervised depth estimation algorithms rely heavily on frame-warping relationships, exhibiting substantial performance degradation when applied in challenging circumstances, such as low-visibility and nighttime scenarios with varying illumination conditions. Addressing this challenge, we introduce an algorithm designed to achieve accurate selfsupervised stereo depth estimation focusing on nighttime conditions. Specifically, we use pretrained visual foundation models to extract generalised features across challenging scenes and present an efficient method for matching and integrating these features from stereo frames. Moreover, to prevent pixels violating photometric consistency assumption from negatively affecting the depth predictions, we propose a novel masking approach designed to filter out such pixels. Lastly, addressing weaknesses in the evaluation of current depth estimation algorithms, we present novel evaluation metrics. Our experiments, conducted on challenging datasets including Oxford RobotCar and MultiSpectral Stereo, demonstrate the robust improvements realized by our approach. Madhu Vankadari, Samuel Hodgson, Sang-Yun Shin, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni |
ICRA | 3 |
| 2024 | Towards Learning Group-Equivariant Features for Domain Adaptive 3D DetectionabstractThe performance of 3D object detection in large outdoor point clouds deteriorates significantly in an unseen environment due to the inter-domain gap. To address these challenges, most existing methods for domain adaptation harness self-training schemes and attempt to bridge the gap by focusing on a single factor that causes the inter-domain gap, such as objects' sizes, shapes, and foreground density variation. However, the resulting adaptations suggest that there is still a substantial inter-domain gap left to be minimized. We argue that this is due to two limitations: 1) Biased pseudo-label collection from self-training. 2) Multiple factors jointly contributing to how the object is perceived in the unseen target domain. In this work, we propose a grouping-exploration strategy framework, Group Explorer Domain Adaptation ($\textbf{GroupEXP-DA}$), to addresses those two issues. Specifically, our grouping divides the available label sets into multiple clusters and ensures all of them have equal learning attention with the group-equivariant spatial feature, avoiding dominant types of objects causing imbalance problems. Moreover, grouping learns to divide objects by considering inherent factors in a data-driven manner, without considering each factor separately as existing works. On top of the group-equivariant spatial feature that selectively detects objects similar to the input group, we additionally introduce an explorative group update strategy that reduces the false negative detection in the target domain, further reducing the inter-domain gap. During inference, only the learned group features are necessary for making the group-equivariant spatial feature, placing our method as a simple add-on that can be applicable to most existing detectors. We show how each module contributes to substantially bridging the inter-domain gaps compared to existing works across large urban outdoor datasets such as NuScenes, Waymo, and KITTI. Sang-Yun Shin, Madhu Vankadari, Ta Ying Cheng, Qian Xie 0001, Andrew Markham, Agathoniki Trigoni |
NeurIPS | 1 |
| 2024 | Sound3DVDet: 3D Sound Source Detection using Multiview Microphone Array and RGB ImagesabstractSpatial localization of 3D sound sources is an important problem in many real world scenarios, especially when the sources may not have any visually distinguishable characteristic; e.g., finding a gas leak, a malfunctioning motor, etc. In this paper, we cast this task in a novel audio-visual setting, by introducing an acoustic-camera rig consisting of a centered pinhole RGB camera and a uniform circular array of four coplanar microphones. Using this setup, we propose Sound3DVDet – a 3D sound source localization Transformer model that treats this task as a set prediction problem. It first learns a set of initial sound source locations (dubbed queries) from a single view of the microphone array signal, then feeds the query set to a sequence of Transformerlike layers for refinement. Each query arising from each layer repeatedly aggregates sound source cues from other views. We deeply supervise the initial sound source queries, intermediate layer queries, and the final output by measuring their respective discrepancy against ground truth queries via bipartite matching. To evaluate our method, we introduce a new dataset: Sound3DVDet Dataset, consisting of nearly 6k scenes produced using the SoundSpaces simulator. We conduct extensive experiments on our dataset and show the efficacy of our approach against closely related methods, demonstrating significant improvements in the localization accuracy. Code is available at https://github.com/yuhanghe01/Sound3DVDet. Sang-Yun Shin, Anoop Cherian, Agathoniki Trigoni, Andrew Markham |
WACV | 2 |
| 2023 | Sample, Crop, Track: Self-Supervised Mobile 3D Object Detection for Urban Driving LiDARabstractDeep learning has led to great progress in the detection of mobile (i.e. movement-capable) objects in urban driving scenes in recent years. Supervised approaches typically require the annotation of large training sets; there has thus been great interest in leveraging weakly, semi- or self- supervised methods to avoid this, with much success. Whilst weakly and semi-supervised methods require some annotation, self-supervised methods have used cues such as motion to relieve the need for annotation altogether. However, a complete absence of annotation typically degrades their performance, and ambiguities that arise during motion grouping can inhibit their ability to find accurate object boundaries. In this paper, we propose a new self-supervised mobile object detection approach called SCT. This uses both motion cues and expected object sizes to improve detection performance, and predicts a dense grid of 3$D$oriented bounding boxes to improve object discovery. We significantly outperform the state-of-the-art self-supervised mobile object detection method TCR on the KITTI tracking benchmark, and achieve performance that is within 30 % of the fully supervised PV-RCNN++ method for IoUs$\leq$0.5. Our source code will be made available online. Sang-Yun Shin, Stuart Golodetz, Madhu Vankadari, Kaichen Zhou, Andrew Markham, Agathoniki Trigoni |
ICRA | 1 |
| 2023 | DynPoint: Dynamic Neural Point For View SynthesisabstractThe introduction of neural radiance fields has greatly improved the effectiveness of view synthesis for monocular videos. However, existing algorithms face difficulties when dealing with uncontrolled or lengthy scenarios, and require extensive training time specific to each new scenario.
To tackle these limitations, we propose DynPoint, an algorithm designed to facilitate the rapid synthesis of novel views for unconstrained monocular videos.
Rather than encoding the entirety of the scenario information into a latent representation, DynPoint concentrates on predicting the explicit 3D correspondence between neighboring frames to realize information aggregation.
Specifically, this correspondence prediction is achieved through the estimation of consistent depth and scene flow information across frames.
Subsequently, the acquired correspondence is utilized to aggregate information from multiple reference frames to a target frame, by constructing hierarchical neural point clouds.
The resulting framework enables swift and accurate view synthesis for desired views of target frames.
The experimental results obtained demonstrate the considerable acceleration of training time achieved - typically an order of magnitude - by our proposed method while yielding comparable outcomes compared to prior approaches. Furthermore, our method exhibits strong robustness in handling long-duration videos without learning a canonical representation of video content. Kaichen Zhou, Jia-Xing Zhong, Sang-Yun Shin, Kai Lu 0003, Yiyuan Yang, Andrew Markham, Agathoniki Trigoni |
NeurIPS | 3 |
| 2022 | Real-Time Hybrid Mapping of Populated Indoor Scenes using a Low-Cost Monocular UAVabstractUnmanned aerial vehicles (UAVs) have been used for many applications in recent years, from urban search and rescue, to agricultural surveying, to autonomous underground mine exploration. However, deploying UAVs in tight, indoor spaces, especially close to humans, remains a challenge. One solution, when limited payload is required, is to use micro-UAVs, which pose less risk to humans and typically cost less to replace after a crash. However, micro-UAVs can only carry a limited sensor suite, e.g. a monocular camera instead of a stereo pair or LiDAR, complicating tasks like dense mapping and markerless multi-person 3D human pose estimation, which are needed to operate in tight environments around people. Monocular approaches to such tasks exist, and dense monocular mapping approaches have been successfully deployed for UAV applications. However, despite many recent works on both marker-based and markerless multi-UAV single-person motion capture, markerless single-camera multi-person 3D human pose estimation remains a much earlier-stage technology, and we are not aware of existing attempts to deploy it in an aerial context. In this paper, we present what is thus, to our knowledge, the first system to perform simultaneous mapping and multi-person 3D human pose estimation from a monocular camera mounted on a single UAV. In particular, we show how to loosely couple state-of-the-art monocular depth estimation and monocular 3D human pose estimation approaches to reconstruct a hybrid map of a populated indoor scene in real time. We validate our component-level design choices via extensive experiments on the large-scale ScanNet and GTA-IM datasets. To evaluate our system-level performance, we also construct a new Oxford Hybrid Mapping dataset of populated indoor scenes. Stuart Golodetz, Madhu Vankadari, Aluna Everitt, Sang-Yun Shin, Andrew Markham, Agathoniki Trigoni |
IROS | 4 |
| 2020 | Android-GAN: Defending against android pattern attacks using multi-modal generative network as anomaly detector
Sang-Yun Shin, Yong-Won Kang, Yong-Guk Kim |
Expert Syst. Appl. | 1 |
| 2020 | Reward-driven U-Net training for obstacle avoidance drone
Sang-Yun Shin, Yong-Won Kang, Yong-Guk Kim |
Expert Syst. Appl. | 1 |
| 2019 | Automatic Drone Navigation in Realistic 3D Landscapes using Deep Reinforcement LearningabstractWe present a study where a drone navigates through diverse 3D obstacles by finding a 3D path and reaches the goal using deep reinforcement learning (RL) in a 3D realistic landscape. The drone has two inputs: first RGB provides a first person view of the landscape and secondly depth map gives it 3D information of the environment. For training the drone for automatic navigation, deep reinforcement learning is extensively used. For the same task, human pilot navigates through the obstacles with a radio controller (RC) using a hardware-in-the-loop setup. Racing performance between human and several deep RL algorithms such as Deep Q-Network (DQN), Double DQN, Dueling DQN and Double Dueling DQN (DD-DQN) are evaluated. Results suggest that DD-DQN outperforms other algorithms and, for the racing between humans and algorithms, DD-DQN performs better than a novice and yet an expert or intermediate-level pilot outperforms any other algorithms. The present study demonstrates that time and resource for training a drone can be saved using a realistic and yet controllable platform. Sang-Yun Shin, Yong-Won Kang, Yong-Guk Kim |
CoDIT | 1 |