VLDB 2026 Research / reviewers in the wild / expert
Seonwook Park
dblp:203/4739
· DBLP profile ↗
15ranked-venue papers
6as first author
4since 2021 · last 2023
0000-0001-7992-3876ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Face, body and person analysis · 35% Image recognition and object detection · 15% Language models and text generation · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Medical and health informatics · 100% | |
| Human-computer interaction and pervasive computing
3 papers |
Human-AI interaction · 37% Wearable and physiological sensing · 28% Interaction techniques and input · 22% |
Topics — the 27 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
gaze estimation |
2.1 | 5 | 2021 | Weakly-Supervised Physically Unconstrained Gaze Estimation · CVPR 2021 Self-Learning Transformations for Improving Gaze and Head Redirection · NeurIPS 2020 ETH-XGaze: A Large Scale Dataset for Gaze Estimation Under Extreme Head Pose and Gaze Variation · ECCV (5) 2020 |
Medical and health informatics
computational pathology |
1.3 | 2 | 2023 | OCELOT: Overlapped Cell on Tissue Dataset for Histopathology · CVPR 2023 Benchmarking Self-Supervised Learning on Diverse Pathology Datasets · CVPR 2023 |
Natural language and speech › Language models and text generation › large language model training
domain-adaptive pre-training |
0.7 | 1 | 2023 | Benchmarking Self-Supervised Learning on Diverse Pathology Datasets · CVPR 2023 |
Computer vision › Segmentation and scene understanding › medical image segmentation
tissue segmentation |
0.7 | 1 | 2023 | OCELOT: Overlapped Cell on Tissue Dataset for Histopathology · CVPR 2023 |
Medical and health informatics › computational pathology
histopathology image analysis |
0.7 | 1 | 2023 | Benchmarking Self-Supervised Learning on Diverse Pathology Datasets · CVPR 2023 |
Computer vision › Image recognition and object detection
object detection |
0.6 | 1 | 2022 | Interactive Multi-Class Tiny-Object Detection · CVPR 2022 |
Computer vision › Image recognition and object detection › object detection
small object detection |
0.6 | 1 | 2022 | Interactive Multi-Class Tiny-Object Detection · CVPR 2022 |
Human-AI interaction › human-in-the-loop
human-in-the-loop annotation |
0.6 | 1 | 2022 | Interactive Multi-Class Tiny-Object Detection · CVPR 2022 |
Computer vision › Face, body and person analysis › gaze estimation
3d gaze estimation |
0.5 | 1 | 2021 | Weakly-Supervised Physically Unconstrained Gaze Estimation · CVPR 2021 |
Machine learning › Generative modeling › face synthesis
controllable face generation |
0.4 | 1 | 2020 | Self-Learning Transformations for Improving Gaze and Head Redirection · NeurIPS 2020 |
Wearable and physiological sensing
eye tracking |
0.4 | 1 | 2020 | Towards End-to-End Video-Based Eye-Tracking · ECCV (12) 2020 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
personalization |
0.4 | 1 | 2019 | Few-Shot Adaptive Gaze Estimation · ICCV 2019 |
Machine learning › Transfer learning and domain adaptation › meta-learning
personalized meta-learning |
0.4 | 1 | 2019 | Few-Shot Adaptive Gaze Estimation · ICCV 2019 |
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation |
0.3 | 1 | 2018 | Cross-Modal Deep Variational Hand Pose Estimation · CVPR 2018 |
Machine learning › Generative modeling
variational autoencoder |
0.3 | 1 | 2018 | Cross-Modal Deep Variational Hand Pose Estimation · CVPR 2018 |
Interaction techniques and input › cross-device interaction
cross-device UI distribution |
0.3 | 1 | 2018 | AdaM: Adapting Multi-User Interfaces for Collaborative Environments in Real-Time · CHI 2018 |
Robotics › Robot navigation and mapping › SLAM › visual SLAM
direct visual SLAM |
0.3 | 1 | 2017 | Illumination change robustness in direct visual SLAM · ICRA 2017 |
Robotics › Autonomous driving › perception › perception robustness
illumination change robustness |
0.3 | 1 | 2017 | Illumination change robustness in direct visual SLAM · ICRA 2017 |
Robotics › Robot navigation and mapping › SLAM
visual SLAM |
0.3 | 1 | 2017 | Illumination change robustness in direct visual SLAM · ICRA 2017 |
Computer vision › Image recognition and object detection › object detection
multi-class object detection |
0.2 | 1 | 2022 | Interactive Multi-Class Tiny-Object Detection · CVPR 2022 |
Computer vision › Video understanding and tracking › video analytics › behavior analysis › human behavior analysis
human interaction analysis |
0.1 | 1 | 2021 | Weakly-Supervised Physically Unconstrained Gaze Estimation · CVPR 2021 |
Computer vision › Face, body and person analysis
head pose estimation |
0.1 | 1 | 2020 | ETH-XGaze: A Large Scale Dataset for Gaze Estimation Under Extreme Head Pose and Gaze Variation · ECCV (5) 2020 |
Machine learning › Representation and self-supervised learning › multimodal representation learning
cross-modal representation learning |
0.1 | 1 | 2018 | Cross-Modal Deep Variational Hand Pose Estimation · CVPR 2018 |
Collaborative and social computing › collaborative systems
collaborative environments |
0.1 | 1 | 2018 | AdaM: Adapting Multi-User Interfaces for Collaborative Environments in Real-Time · CHI 2018 |
Collaborative and social computing › collaborative systems
multi-user interfaces |
0.1 | 1 | 2018 | AdaM: Adapting Multi-User Interfaces for Collaborative Environments in Real-Time · CHI 2018 |
Computer vision › 3D vision
direct image alignment |
0.1 | 1 | 2017 | Illumination change robustness in direct visual SLAM · ICRA 2017 |
Robotics › Robot navigation and mapping
visual odometry |
0.1 | 1 | 2017 | Illumination change robustness in direct visual SLAM · ICRA 2017 |
Methods — techniques the papers use, named apart from their topics
self-supervised pretraining · 1.3nuclei instance segmentation · 1.3multi-task learning · 1.3linear evaluation · 1.3fine-tuning · 1.3point-based user input · 1.1late fusion · 1.1feature correlation · 1.1deep learning · 0.8loss function design · 0.5end-to-end learning · 0.4mixed integer programming · 0.3lab study · 0.3combinatorial optimization · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Benchmarking Self-Supervised Learning on Diverse Pathology DatasetsabstractComputational pathology can lead to saving human lives, but models are annotation hungry and pathology images are notoriously expensive to annotate. Self-supervised learning (SSL) has shown to be an effective method for utilizing unlabeled data, and its application to pathology could greatly benefit its downstream tasks. Yet, there are no principled studies that compare SSL methods and discuss how to adapt them for pathology. To address this need, we execute the largest-scale study of SSL pre-training on pathology image data, to date. Our study is conducted using 4 representative SSL methods on diverse downstream tasks. We establish that large-scale domain-aligned pre-training in pathology consistently out-performs ImageNet pre-training in standard SSL settings such as linear and fine-tuning evaluations, as well as in low-label regimes. Moreover, we propose a set of domain-specific techniques that we experimentally show leads to a performance boost. Lastly, for the first time, we apply SSL to the challenging task of nuclei instance segmentation and show large and consistent performance improvements. We release the pre-trained model weights11https://lunit-io.github.io/research/publications/pathology_ssl. Mingu Kang, Heon Song, Seonwook Park, Donggeun Yoo, Sérgio Pereira |
CVPR | 3 |
| 2023 | OCELOT: Overlapped Cell on Tissue Dataset for HistopathologyabstractCell detection is a fundamental task in computational pathology that can be used for extracting high-level medical information from whole-slide images. For accurate cell detection, pathologists often zoom out to understand the tissue-level structures and zoom in to classify cells based on their morphology and the surrounding context. However, there is a lack of efforts to reflect such behaviors by pathologists in the cell detection models, mainly due to the lack of datasets containing both cell and tissue annotations with overlapping regions. To overcome this limitation, we propose and publicly release OCELOT, a dataset purposely dedicated to the study of cell-tissue relationships for cell detection in histopathology. OCELOT provides overlapping cell and tissue annotations on images acquired from multiple organs. Within this setting, we also propose multi-task learning approaches that benefit from learning both cell and tissue tasks simultaneously. When compared against a model trained only for the cell detection task, our proposed approaches improve cell detection performance on 3 datasets: proposed OCELOT, public TIGER, and internal CARP datasets. On the OCELOT test set in particular, we show up to 6.79 improvement in F1-score. We believe the contributions of this paper, including the release of the OCELOT dataset at https://lunit-io.github.io/research/publications/OCELOT are a crucial starting point toward the important research direction of incorporating cell-tissue relationships in computation pathology. Jeongun Ryu, Aaron Valero Puche, Jaewoong Shin, Seonwook Park, Biagio Brattoli, Wonkyung Jung, Soo Ick Cho, Kyunghyun Paeng, Chan-Young Ock, Donggeun Yoo, Sérgio Pereira |
CVPR | 4 |
| 2022 | Interactive Multi-Class Tiny-Object DetectionabstractAnnotating tens or hundreds of tiny objects in a given image is laborious yet crucial for a multitude of Computer Vision tasks. Such imagery typically contains objects from various categories, yet the multi-class interactive annotation setting for the detection task has thus far been unex-plored. To address these needs, we propose a novel interactive annotation method for multiple instances of tiny objects from multiple classes, based on a few point-based user in-puts. Our approach, C3Det, relates the full image context with annotator inputs in a local and global manner via late-fusion andfeature-correlation, respectively. We perform ex-periments on the Tiny-DOTA. and LCell datasets using both two-stage and one-stage object detection architectures to verify the efficacy of our approach. Our approach outper-forms existing approaches in interactive annotation, achieving higher mAP with fewer clicks. Furthermore, we validate the annotation efficiency of our approach in a user study where it is shown to be 2.85x faster and yield only 0.36x task load (NASA-TLX, lower is better) compared to manual annotation. The code is available at https://github.com/ChungYi347/Interactive-Multi-Class-Tiny-Object-Detection. Chunggi Lee, Seonwook Park, Heon Song, Jeongun Ryu, Haejoon Kim, Sérgio Pereira, Donggeun Yoo |
CVPR | 2 |
| 2021 | Weakly-Supervised Physically Unconstrained Gaze EstimationabstractA major challenge for physically unconstrained gaze estimation is acquiring training data with 3D gaze annotations for in-the-wild and outdoor scenarios. In contrast, videos of human interactions in unconstrained environments are abundantly available and can be much more easily annotated with frame-level activity labels. In this work, we tackle the previously unexplored problem of weakly-supervised gaze estimation from videos of human interactions. We leverage the insight that strong gaze-related geometric constraints exist when people perform the activity of "looking at each other" (LAEO). To acquire viable 3D gaze supervision from LAEO labels, we propose a training algorithm along with several novel loss functions especially designed for the task. With weak supervision from two large scale CMU-Panoptic and AVA-LAEO activity datasets, we show significant improvements in (a) the accuracy of semisupervised gaze estimation and (b) cross-domain generalization on the state-of-the-art physically unconstrained in-the-wild Gaze360 gaze estimation benchmark. We open source our code at https://github.com/NVlabs/weaklysupervised-gaze. Rakshit Sunil Kothari, Shalini De Mello, Umar Iqbal 0001, Wonmin Byeon, Seonwook Park, Jan Kautz |
CVPR | 5 |
| 2020 | Towards End-to-End Video-Based Eye-Tracking
Seonwook Park, Emre Aksan, Xucong Zhang, Otmar Hilliges |
ECCV (12) | 1 |
| 2020 | ETH-XGaze: A Large Scale Dataset for Gaze Estimation Under Extreme Head Pose and Gaze Variation
Xucong Zhang, Seonwook Park, Thabo Beeler, Derek Bradley, Siyu Tang 0001, Otmar Hilliges |
ECCV (5) | 2 |
| 2020 | Detecting Relevance during Decision-Making from Eye Movements for UI AdaptationabstractThis paper proposes an approach to detect information relevance during decision-making from eye movements in order to enable user interface adaptation. This is a challenging task because gaze behavior varies greatly across individual users and tasks and ground-truth data is difficult to obtain. Thus, prior work has mostly focused on simpler target-search tasks or on establishing general interest, where gaze behavior is less complex. From the literature, we identify six metrics that capture different aspects of the gaze behavior during decision-making and combine them in a voting scheme. We empirically show, that this accounts for the large variations in gaze behavior and out-performs standalone metrics. Importantly, it offers an intuitive way to control the amount of detected information, which is crucial for different UI adaptation schemes to succeed. We show the applicability of our approach by developing a room-search application that changes the visual saliency of content detected as relevant. In an empirical study, we show that it detects up to 97% of relevant elements with respect to user self-reporting, which allows us to meaningfully adapt the interface, as confirmed by participants. Our approach is fast, does not need any explicit user input and can be applied independent of task and user. Anna Maria Feit, Lukas Vordemann, Seonwook Park, Caterina Bérubé, Otmar Hilliges |
ETRA | 3 |
| 2020 | Self-Learning Transformations for Improving Gaze and Head RedirectionabstractMany computer vision tasks rely on labeled data. Rapid progress in generative modeling has led to the ability to synthesize photorealistic images. However, controlling specific aspects of the generation process such that the data can be used for supervision of downstream tasks remains challenging. In this paper we propose a novel generative model for images of faces, that is capable of producing high-quality images under fine-grained control over eye gaze and head orientation angles. This requires the disentangling of many appearance related factors including gaze and head orientation but also lighting, hue etc. We propose a novel architecture which learns to discover, disentangle and encode these extraneous variations in a self-learned manner. We further show that explicitly disentangling task-irrelevant factors results in more accurate modelling of gaze and head orientation. A novel evaluation scheme shows that our method improves upon the state-of-the-art in redirection accuracy and disentanglement between gaze direction and head orientation changes. Furthermore, we show that in the presence of limited amounts of real-world training data, our method allows for improvements in the downstream task of semi-supervised cross-dataset gaze estimation. Please check our project page at: https://ait.ethz.ch/projects/2020/STED-gaze/ Seonwook Park, Xucong Zhang, Shalini De Mello, Otmar Hilliges |
NeurIPS | 2 |
| 2020 | Accurate Real-time 3D Gaze Tracking Using a Lightweight Eyeball CalibrationabstractAbstract 3D gaze tracking from a single RGB camera is very challenging due to the lack of information in determining the accurate gaze target from a monocular RGB sequence. The eyes tend to occupy only a small portion of the video, and even small errors in estimated eye orientations can lead to very large errors in the triangulated gaze target. We overcome these difficulties with a novel lightweight eyeball calibration scheme that determines the user‐specific visual axis, eyeball size and position in the head. Unlike the previous calibration techniques, we do not need the ground truth positions of the gaze points. In the online stage, gaze is tracked by a new gaze fitting algorithm, and refined by a 3D gaze regression method to correct for bias errors. Our regression is pre‐trained on several individuals and works well for novel users. After the lightweight one‐time user calibration, our method operates in real time. Experiments show that our technique achieves state‐of‐the‐art accuracy in gaze angle estimation, and we demonstrate applications of 3D gaze target tracking and gaze retargeting to an animated 3D character. Q. Wen, Derek Bradley, Thabo Beeler, Seonwook Park, Otmar Hilliges, J. Yong |
Comput. Graph. Forum | 4 |
| 2019 | Few-Shot Adaptive Gaze EstimationabstractInter-personal anatomical differences limit the accuracy of person-independent gaze estimation networks. Yet there is a need to lower gaze errors further to enable applications requiring higher quality. Further gains can be achieved by personalizing gaze networks, ideally with few calibration samples. However, over-parameterized neural networks are not amenable to learning from few examples as they can quickly over-fit. We embrace these challenges and propose a novel framework for Few-shot Adaptive GaZE Estimation (Faze) for learning person-specific gaze networks with very few (≤ 9) calibration samples. Faze learns a rotation-aware latent representation of gaze via a disentangling encoder-decoder architecture along with a highly adaptable gaze estimator trained using meta-learning. It is capable of adapting to any new person to yield significant performance gains with as few as 3 samples, yielding state-of-the-art performance of 3.18-deg on GazeCapture, a 19% improvement over prior art. We open-source our code at https://github.com/NVlabs/few_shot_gaze. Seonwook Park, Shalini De Mello, Pavlo Molchanov 0001, Umar Iqbal 0001, Otmar Hilliges, Jan Kautz |
ICCV | 1 |
| 2018 | AdaM: Adapting Multi-User Interfaces for Collaborative Environments in Real-TimeabstractDeveloping cross-device multi-user interfaces (UIs) is a challenging problem. There are numerous ways in which content and interactivity can be distributed. However, good solutions must consider multiple users, their roles, their preferences and access rights, as well as device capabilities. Manual and rule-based solutions are tedious to create and do not scale to larger problems nor do they adapt to dynamic changes, such as users leaving or joining an activity. In this paper, we cast the problem of UI distribution as an assignment problem and propose to solve it using combinatorial optimization. We present a mixed integer programming formulation which allows real-time applications in dynamically changing collaborative settings. It optimizes the allocation of UI elements based on device capabilities, user roles, preferences, and access rights. We present a proof-of-concept designer-in-the-loop tool, allowing for quick solution exploration. Finally, we compare our approach to traditional paper prototyping in a lab study. Seonwook Park, Christoph Gebhardt, Roman Rädle, Anna Maria Feit, Hana Vrzakova, Niraj Ramesh Dayama, Hui-Shyong Yeo, Clemens Nylandsted Klokmose, Aaron J. Quigley, Antti Oulasvirta, Otmar Hilliges |
CHI | 1 |
| 2018 | Cross-Modal Deep Variational Hand Pose EstimationabstractThe human hand moves in complex and high-dimensional ways, making estimation of 3D hand pose configurations from images alone a challenging task. In this work we propose a method to learn a statistical hand model represented by a cross-modal trained latent space via a generative deep neural network. We derive an objective function from the variational lower bound of the VAE framework and jointly optimize the resulting cross-modal KL-divergence and the posterior reconstruction objective, naturally admitting a training regime that leads to a coherent latent space across multiple modalities such as RGB images, 2D keypoint detections or 3D hand configurations. Additionally, it grants a straightforward way of using semi-supervision. This latent space can be directly used to estimate 3D hand poses from RGB images, outperforming the state-of-the art in different settings. Furthermore, we show that our proposed method can be used without changes on depth images and performs comparably to specialized methods. Finally, the model is fully generative and can synthesize consistent pairs of hand configurations across modalities. We evaluate our method on both RGB and depth datasets and analyze the latent space qualitatively. Adrian Spurr, Jie Song 0006, Seonwook Park, Otmar Hilliges |
CVPR | 3 |
| 2018 | Deep Pictorial Gaze Estimation
Seonwook Park, Adrian Spurr, Otmar Hilliges |
ECCV (13) | 1 |
| 2018 | Learning to find eye region landmarks for remote gaze estimation in unconstrained settingsabstractConventional feature-based and model-based gaze estimation methods have proven to perform well in settings with controlled illumination and specialized cameras. In unconstrained real-world settings, however, such methods are surpassed by recent appearance-based methods due to difficulties in modeling factors such as illumination changes and other visual artifacts. We present a novel learning-based method for eye region landmark localization that enables conventional methods to be competitive to latest appearance-based methods. Despite having been trained exclusively on synthetic data, our method exceeds the state of the art for iris localization and eye shape registration on real-world imagery. We then use the detected landmarks as input to iterative model-fitting and lightweight learning-based gaze estimation methods. Our approach outperforms existing model-fitting and appearance-based methods in the context of person-independent and personalized gaze estimation. Seonwook Park, Xucong Zhang, Andreas Bulling, Otmar Hilliges |
ETRA | 1 |
| 2017 | Illumination change robustness in direct visual SLAMabstractDirect visual odometry and Simultaneous Localization and Mapping (SLAM) methods determine camera poses by means of direct image alignment. This optimizes a photometric cost term based on the Lucas-Kanade method. Many recent works use the brightness constancy assumption in the alignment cost formulation and therefore cannot cope with significant illumination changes. Such changes are especially likely to occur for loop closures in SLAM. Alternatives exist which attempt to match images more robustly. In our paper, we perform a systematic evaluation of real-time capable methods. We determine their accuracy and robustness in the context of odometry and of loop closures, both on real images as well as synthetic datasets with simulated lighting changes. We find that for real images, a Census-based method outperforms the others. We make our new datasets available online. Seonwook Park, Thomas Schöps, Marc Pollefeys |
ICRA | 1 |