Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Simin Kou

dblp:307/5946 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-7222-2214ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Visual content generation and editing · 48% Virtual and augmented reality · 26% Rendering · 26%
Artificial intelligence
1 paper
Representation and self-supervised learning · 100%
Human-computer interaction and pervasive computing
1 paper
Immersive interaction · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Virtual and augmented reality › immersive video
360-degree video
0.912025
OmniPlane: A Recolorable Representation for Dynamic Scenes in Omnidirectional Videos · IEEE Trans. Vis. Comput. Graph. 2025
Rendering
dynamic scene representation
0.912025
OmniPlane: A Recolorable Representation for Dynamic Scenes in Omnidirectional Videos · IEEE Trans. Vis. Comput. Graph. 2025
Visual content generation and editing › video editing
video recoloring
0.912025
OmniPlane: A Recolorable Representation for Dynamic Scenes in Omnidirectional Videos · IEEE Trans. Vis. Comput. Graph. 2025
Visual content generation and editing
video editing
0.812024
Neural Panoramic Representation for Spatially and Temporally Consistent 360° Video Editing · ISMAR 2024
Immersive interaction › head-mounted display
head-mounted display interaction
0.812024
Neural Panoramic Representation for Spatially and Temporally Consistent 360° Video Editing · ISMAR 2024
Machine learning › Representation and self-supervised learning
subspace clustering
0.712023
Structure-Aware Subspace Clustering · IEEE Trans. Knowl. Data Eng. 2023

Methods — techniques the papers use, named apart from their topics

spherical implicit content layers · 1.5neural representation · 1.5MLP-based networks · 1.5weighted sampling · 0.9spherical spatiotemporal feature grids · 0.9palette-based color decomposition · 0.9subspace clustering · 0.7structure-aware learning · 0.7
YearPublicationVenuePosition
2026 OmniPrior: A Multi-Prior-Guided Omnidirectional Representation of Dynamic Scenes in Overlapping Ultra-Wide Multi-Fisheye Videos
abstract
Omnidirectional capture of dynamic scenes facilitates the creation of immersive virtual reality assets and holistic scene understanding. Outward-facing multi-fisheye camera rigs offer an efficient solution for full-scene coverage, using fewer lenses than conventional pinhole arrays while enabling all-directional observation of complex, time-varying environments. By continuously recording scene evolution from every angle, these systems naturally enable a richer characterization of dynamic interactions. Despite these advantages, dynamic scene modeling in this setting remains underexplored. Existing methods, typically designed for fixed pinhole configurations or monocular setups, rely heavily on photometric cues and often neglect the strong geometric and semantic priors inherent in multi-fisheye omnidirectional data. To address this gap, we present OmniPrior, a Gaussian Splatting-based framework for outward-facing, multi-fisheye omnidirectional capture. Our approach incorporates metric-geometry-aware initialization with multi-prior guidance, introducing a dynamicness-aware Gaussian representation that encodes both object motion and subtle temporal variations. The resulting representations are physically consistent and temporally stable. Extensive experiments validate the effectiveness of our method in novel view synthesis across new viewpoints and timestamps. We demonstrate its utility in two representative applications derived from our learned representations: 6DoF rendering with flexible FoV and motion-freeze rendering.
Simin Kou, Jakob Nazarenus, Reinhard Koch, Can Wang 0006, Neil A. Dodgson
IEEE Trans. Vis. Comput. Graph.1
2025 OmniPlane: A Recolorable Representation for Dynamic Scenes in Omnidirectional Videos
abstract
Consumer-level omnidirectional video offers an economically viable means to create virtual reality (VR) assets, enabling users to explore and interact within a fully immersive visual environment. However, editing such videos, particularly those with 360${}^{\circ }$∘ views and dynamic objects, poses significant challenges. Existing approaches to representing and manipulating omnidirectional content-whether designed for typical 2D perspective imagery or panoramas-often fail to adequately capture the complex spatiotemporal relationships crucial for producing high-quality, editable outputs in dynamic, panoramic settings. To overcome these challenges, we introduce OmniPlane, a novel method that leverages spherical spatiotemporal feature grids to empower the representation and editability of real-world dynamic omnidirectional environments casually captured by commodity omnidirectional cameras. OmniPlane computes spatiotemporal features by fusing vectors or matrices from each learnable spatial and spatiotemporal feature plane within a spherical coordinate system, complemented by a specifically designed weighted sampling strategy respecting the inherent spherical distribution of omnidirectional content. These learned feature planes can be flexibly decomposed into palette-based color bases. This innovative method not only enhances the representation capability of omnidirectional content and dynamics but also enables the recoloring of omnidirectional videos. Extensive experiments and a dedicated user study validate the superior performance of our proposed method in facilitating recolorable representations of dynamic omnidirectional environments.
Simin Kou, Jakob Nazarenus, Reinhard Koch, Neil A. Dodgson
IEEE Trans. Vis. Comput. Graph.1
2024 Neural Panoramic Representation for Spatially and Temporally Consistent 360° Video Editing
abstract
Content-based 360° video editing allows users to manipulate panoramic content for interaction in a dynamic visual world. However, the current related methods (2D neural representation and optical flow) show limitations in producing high-quality panoramic content from 360° videos due to their lack of capacity to model the inherent spatiotemporal relationships among pixels in the true panoramic space. To address this issue, we propose a Neural Panoramic Representation (NPR) method to model the global inter-pixel relationships, facilitating immersive video editing. Specifically, our method utilizes MLP-based networks to learn spherical implicit content layers, by encoding the spherical spatiotemporal positions and appearance details within the panoramic video, and bi-directional mapping between the original video frames and the learned content layers, to capture the interpretable and global omnidirectional visual characteristics of individual dynamic scenes. Additionally, we introduce innovative loss functions (spherical neighborhood consistency and unit spherical regularization) to ensure the creation of appropriate implicit spherical content layers. We further provide an interactive layer neural panoramic editing approach based on the proposed NPR, in the head-mounted display device. We evaluate this framework on diverse real-world 360° videos, showing superior performance on both reconstruction and consistent editing compared to existing state-of-the-art (SOTA) neural representation techniques.
Simin Kou, Yukun Lai, Neil A. Dodgson
ISMAR1
2024 WebLFR: An interactive light field renderer in web browsers
Xiaofei Ai, Yigang Wang, Simin Kou
Multim. Tools Appl.4
2023 Structure-Aware Subspace Clustering
abstract
Subspace clustering has attracted much attention because of its ability to group unlabeled high-dimensional data into multiple subspaces. Existing graph-based subspace clustering methods focus on either the sparsity of data affinity or the low rank of data affinity. Thus, the quality of data affinity plays an essential role in the performance of subspace clustering. However, the real-world data are generally high-dimensional, complex, and heterogeneous multi-source data, so that the data affinity learned by these methods cannot be completely dependent. Moreover, since these approaches always ignore the intrinsic structure of data, their grouping effect is relatively low. In this paper, we propose a novel unsupervised algorithm, called Structure-Aware Subspace Clustering (SASC), to address the above issues. SASC considers local and global correlation structures simultaneously to capture the intrinsic structure. Further, it integrates the captured structure into representation learning to gain a relatively precise data affinity. It is powerful to promote an all-around grouping effect and enhances the robustness and applicability of subspace clustering. Experiments on various benchmark datasets, including bioinformatics, handwritten digit, object image, and speech signal, demonstrate the effectiveness of the proposed algorithm.
Simin Kou, Xuesong Yin, Yigang Wang, Songcan Chen, Tieming Chen, Zizhao Wu
IEEE Trans. Knowl. Data Eng.1