Jussi Keppo

dblp:05/1442 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0002-4571-5566ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficacy of High-Fidelity VR Threat-and-Error Simulation for Competency-based Pilot Training
abstract
The aviation industry faces increasing pilot training demands, and reliance on conventional Full Flight Simulators (FFS) limits training capacity. Virtual Reality (VR) offers scalable, remote training opportunities, but its role as a complement to FFS requires empirical validation. In collaboration with Singapore Airlines (SIA) instructor pilots, we developed a VR training prototype for a visual approach into Gimhae International Airport, emphasizing Competency-Based Training and Assessment (CBTA)-based Threat and Error Management (TEM). An empirical study with 39 SIA Boeing 737-MAX 8 type-rated first officers evaluated VR against FFS. An equivalence analysis showed that VR achieved performance outcomes comparable to FFS in 13 of the 16 Observable Behaviors (OBs) and across 4 Competencies. In addition, a comparative analysis indicated measurable performance improvements when VR was used to supplement FFS training. Our results suggest VR can meaningfully complement FFS in targeted competency areas, with future work required to assess its broader integration across additional scenarios and performance metrics.
Teong Leong Chuah, Brian Soon Wei Chiam, Ahmad Iqbal bin Othman, Lindy Li Wen Lim, Jia Wang Tay, Vinh-Thuyen Nguyen-Truong, Catherine Wan Ting Leo, Leslie Siew Mun Chong, Jussi Keppo, Eng Tat Khoo
IEEE Trans. Vis. Comput. Graph.9
2025 Credit Rating Design Under Adverse Selection
abstract
Credit ratings are essential in capital markets, serving as a bridge to reduce information asymmetry between firms and investors. However, the prevalent issuer-pay model creates a potential conflict of interest: since CRAs are paid by the issuers, they may have incentives to inflate ratings to retain clients, compromising the credibility of ratings and threatening market stability. This concern is often linked to the observed discrepancies between solicited ratings (for paying firms) and unsolicited ratings (for non-paying firms), as documented in empirical studies.
Jussi Keppo, Qinzhen Li
EC2
2025 CroCoDai: A Stablecoin for Cross-Chain Commerce
abstract
Decentralized Finance (DeFi), in which digital assets are exchanged without trusted intermediaries, has grown rapidly in value in recent years. Stablecoins , which are pegged to a non-volatile asset such as the US dollar, are a prominent feature of DeFi as they mitigate the risk associated with price fluctuations. However, existing stablecoin systems are tied to individual blockchain platforms, and trusted parties or complex protocols are needed to exchange stablecoin tokens between blockchains. Our goal is to design a practical stablecoin system for cross-chain commerce, and to do so we must overcome two main challenges. The first is to support a large and growing number of blockchains efficiently. The second is for the stablecoin to be resilient to blockchain platform failures and to price fluctuations that affect its collateral. We present CroCoDai to address these challenges. We demonstrate CroCoDai ’s efficiency by comparing the performance of a prototype implementation to related baselines, and its resilience through an empirical analysis of historical token price data.
Daniël Reijsbergen, Bretislav Hajek, Tien Tuan Anh Dinh, Jussi Keppo, Henry F. Korth, Anwitaman Datta
Distributed Ledger Technol. Res. Pract.4
2025 Correction: Instant3D: Instant Text-to-3D Generation
Ming Li 0073, Pan Zhou 0002, Jia-Wei Liu, Jussi Keppo, Shuicheng Yan, Xiangyu Xu 0002
Int. J. Comput. Vis.4
2024 DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing
abstract
Despite recent progress in diffusion-based video editing, existing methods are limited to short-length videos due to the contradiction between long-range consistency and frame-wise editing. Prior attempts to address this challenge by introducing video-2D representations encounter significant difficulties with large motion- and view-change videos, especially in human-centric scenarios. To overcome this, we propose to introduce the dynamic Neural Radiance Fields (NeRF) as the innovative video representation, where the editing can be performed in the 3D spaces and propagated to the entire video via the deformation field. To provide consistent and controllable editing, we propose the image-based video-NeRF editing pipeline with a set of innovative designs, including multi-view multi-pose Score Distillation Sampling (SDS) from both the 2D personalized diffusion prior and 3D diffusion prior, reconstruction losses, text-guided local parts super-resolution, and style transfer. Extensive experiments demonstrate that our method dubbed as DynVideo-E, significantly outperforms SOTA approaches on two challenging datasets by a large margin of 50% ~ 95% for human preference. Code will be released at https://showlab.github.io/DynVideo-E/.
Jia-Wei Liu, Yan-Pei Cao 0001, Jay Zhangjie Wu, Weijia Mao, Yuchao Gu, Rui Zhao 0001, Jussi Keppo, Ying Shan, Zheng Shou 0001
CVPR7
2024 X- Adapter: Universal Compatibility of Plugins for Upgraded Diffusion Model
abstract
We introduce X-Adapter, a universal upgrader to enable the pretrained plug-and-play modules (e.g., ControlNet, LoRA) to work directly with the upgraded text-to-image diffusion model (e.g., SDXL) without further retraining. We achieve this goal by training an additional network to control the frozen upgraded model with the new text-image data pairs. In detail, X-Adapter keeps a frozen copy of the old model to preserve the connectors of different plugins. Additionally, X-Adapter adds trainable mapping layers that bridge the decoders from models of different versions for feature remapping. The remapped features will be used as guidance for the upgraded model. To enhance the guidance ability of X-Adapter, we employ a null-text training strategy for the upgraded model. After training, we also introduce a two-stage denoising strategy to align the initial latents of X Adapter and the upgraded model. Thanks to our strategies, X-Adapter demonstrates universal compatibility with various plugins and also enables plugins of different versions to work together, thereby expanding the functionalities of diffusion community. To verify the effectiveness of the proposed method, we conduct extensive experiments and the results show that X-Adapter may facilitate wider application in the upgraded foundational diffusion model. Project page at: https://showlab.github.io/X-Adapter/.
Lingmin Ran, Xiaodong Cun, Jia-Wei Liu, Rui Zhao 0001, Song Zijie, Xintao Wang 0002, Jussi Keppo, Zheng Shou 0001
CVPR7
2024 MotionDirector: Motion Customization of Text-to-Video Diffusion Models
Rui Zhao 0001, Yuchao Gu, Jay Zhangjie Wu, Junhao Zhang 0001, Jia-Wei Liu, Weijia Wu 0001, Jussi Keppo, Zheng Shou 0001
ECCV (56)7
2024 Exocentric-to-Egocentric Video Generation
abstract
We introduce Exo2Ego-V, a novel exocentric-to-egocentric diffusion-based video generation method for daily-life skilled human activities where sparse 4-view exocentric viewpoints are configured 360° around the scene. This task is particularly challenging due to the significant variations between exocentric and egocentric viewpoints and high complexity of dynamic motions and real-world daily-life environments. To address these challenges, we first propose a new diffusion-based multi-view exocentric encoder to extract the dense multi-scale features from multi-view exocentric videos as the appearance conditions for egocentric video generation. Then, we design an exocentric-to-egocentric view translation prior to provide spatially aligned egocentric features as a concatenation guidance for the input of egocentric video diffusion model. Finally, we introduce the temporal attention layers into our egocentric video diffusion pipeline to improve the temporal consistency cross egocentric frames. Extensive experiments demonstrate that Exo2Ego-V significantly outperforms SOTA approaches on 5 categories from the Ego-Exo4D dataset with an average of 35% in terms of LPIPS. Our code and model will be made available on https://github.com/showlab/Exo2Ego-V.
Jia-Wei Liu, Weijia Mao, Zhongcong Xu, Jussi Keppo, Zheng Shou 0001
NeurIPS4
2024 Instant3D: Instant Text-to-3D Generation
Ming Li 0073, Pan Zhou 0002, Jia-Wei Liu, Jussi Keppo, Shuicheng Yan, Xiangyu Xu 0002
Int. J. Comput. Vis.4
2024 DR-FER: Discriminative and Robust Representation Learning for Facial Expression Recognition
abstract
Learning discriminative and robust representations is important for facial expression recognition (FER) due to subtly different emotional faces and their subjective annotations. Previous works usually address one representation solely because these two goals seem to be contradictory for optimization. Their performances inevitably suffer from challenges from the other representation. In this article, by considering this problem from two novel perspectives, we demonstrate that discriminative and robust representations can be learned in a unified approach, i.e., DR-FER, and mutually benefit each other. Moreover, we make it with the supervision from only original annotations. Specifically, to learn discriminative representations, we propose performing masked image modeling (MIM) as an auxiliary task to force our network to discover expression-related facial areas. This is the first attempt to employ MIM to explore discriminative patterns in a self-supervised manner. To extract robust representations, we present a category-aware self-paced learning schedule to mine high-quality annotated (easy) expressions and incorrectly annotated (hard) counterparts. We further introduce a retrieval similarity-based relabeling strategy to correct hard expression annotations, exploiting them more effectively. By enhancing the discrimination ability of the FER classifier as a bridge, these two learning goals significantly strengthen each other. Extensive experiments on several popular benchmarks demonstrate the superior performance of our DR-FER. Moreover, thorough visualizations and extra experiments on manually annotation-corrupted datasets show that our approach successfully accomplishes learning both discriminative and robust representations simultaneously.
Ming Li 0073, Huazhu Fu, Shengfeng He, Hehe Fan, Jun Liu 0036, Jussi Keppo, Zheng Shou 0001
IEEE Trans. Multim.6
2023 STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition
abstract
Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks. First, they may compromise temporal dynamics in input videos, which are critical for accurate action recognition. Second, they are vulnerable to practical attacking scenarios where attackers probe for privacy from an entire video rather than individual frames. To address these issues, we propose a novel framework STPrivacy to perform video-level PPAR. For the first time, we introduce vision Transformers into PPAR by treating a video as a tubelet sequence, and accordingly design two complementary mechanisms, i.e., sparsification and anonymization, to remove privacy from a spatio-temporal perspective. In specific, our privacy sparsification mechanism applies adaptive token selection to abandon action-irrelevant tubelets. Then, our anonymization mechanism implicitly manipulates the remaining action-tubelets to erase privacy in the embedding space through adversarial learning. These mechanisms provide significant advantages in terms of privacy preservation for human eyes and action-privacy trade-off adjustment during deployment. We additionally contribute the first two large-scale PPAR benchmarks, VP-HMDB51 and VP-UCF101, to the community. Extensive evaluations on them, as well as two other tasks, validate the effectiveness and generalization capability of our framework.
Ming Li 0073, Xiangyu Xu 0002, Hehe Fan, Pan Zhou 0002, Jun Liu 0036, Jia-Wei Liu, Jiahe Li 0009, Jussi Keppo, Zheng Shou 0001, Shuicheng Yan
ICCV8
2023 HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video
abstract
We introduce HOSNeRF, a novel 360° free-viewpoint rendering method that reconstructs neural radiance fields for dynamic human-object-scene from a single monocular in-the-wild video. Our method enables pausing the video at any frame and rendering all scene details (dynamic humans, objects, and backgrounds) from arbitrary viewpoints. The first challenge in this task is the complex object motions in human-object interactions, which we tackle by introducing the new object bones into the conventional human skeleton hierarchy to effectively estimate large object deformations in our dynamic human-object model. The second challenge is that humans interact with different objects at different times, for which we introduce two new learnable object state embeddings that can be used as conditions for learning our human-object representation and scene representation, respectively. Extensive experiments show that HOSNeRF significantly outperforms SOTA approaches on two challenging datasets by a large margin of 40%~50% in terms of LPIPS. The code, data, and compelling examples of 360° free-viewpoint renderings from single videos: https://showlab.github.io/HOSNeRF.
Jia-Wei Liu, Yan-Pei Cao 0001, Tianyuan Yang, Zhongcong Xu, Jussi Keppo, Ying Shan, Xiaohu Qie, Zheng Shou 0001
ICCV5
2022 DeVRF: Fast Deformable Voxel Radiance Fields for Dynamic Scenes
abstract
Modeling dynamic scenes is important for many applications such as virtual reality and telepresence. Despite achieving unprecedented fidelity for novel view synthesis in dynamic scenes, existing methods based on Neural Radiance Fields (NeRF) suffer from slow convergence (i.e., model training time measured in days). In this paper, we present DeVRF, a novel representation to accelerate learning dynamic radiance fields. The core of DeVRF is to model both the 3D canonical space and 4D deformation field of a dynamic, non-rigid scene with explicit and discrete voxel-based representations. However, it is quite challenging to train such a representation which has a large number of model parameters, often resulting in overfitting issues. To overcome this challenge, we devise a novel static-to-dynamic learning paradigm together with a new data capture setup that is convenient to deploy in practice. This paradigm unlocks efficient learning of deformable radiance fields via utilizing the 3D volumetric canonical space learnt from multi-view static images to ease the learning of 4D voxel deformation field with only few-view dynamic sequences. To further improve the efficiency of our DeVRF and its synthesized novel view's quality, we conduct thorough explorations and identify a set of strategies. We evaluate DeVRF on both synthetic and real-world dynamic scenes with different types of deformation. Experiments demonstrate that DeVRF achieves two orders of magnitude speedup (100× faster) with on-par high-fidelity results compared to the previous state-of-the-art approaches. The code and dataset are released in https://github.com/showlab/DeVRF.
Yan-Pei Cao 0001, Weijia Mao, Wenqiao Zhang, Junhao Zhang 0001, Jussi Keppo, Ying Shan, Xiaohu Qie, Zheng Shou 0001
NeurIPS6