EDBT 2026 Demo / reviewers in the wild / expert
Sipeng Yang
dblp:314/5934
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | F.A.C.U.L.: Language-Based Interaction with AI Companions in GamingabstractIn cooperative video games, traditional AI companions are deployed to assist players, who control them using hotkeys or command wheels to issue predefined commands such as ''attack'', ''defend'', or ''retreat''. Despite their simplicity, these methods, which lack target specificity, limit players' ability to give complex tactical instructions and hinder immersive gameplay experiences. To address this, we propose the FPS AI Companion who Understands Language (F.A.C.U.L.), the first real-time AI system that enables players to communicate and collaborate with AI companions using natural language. By integrating natural language processing with a confidence-based framework, F.A.C.U.L. efficiently decomposes complex commands and interprets player intent. It also employs a dynamic entity retrieval method for environmental awareness, aligning human intentions with decision-making. Unlike traditional rule-based systems, our method supports real-time language interactions, enabling players to issue complex commands such as ''clear the second floor,'' ''take cover behind that tree,'' or ''retreat to the river''. The system provides real-time behavioral responses and vocal feedback, ensuring seamless tactical collaboration. Using the popular FPS game Arena Breakout: Infinite as a case study, we present comparisons demonstrating the efficacy of our approach and discuss the advantages and limitations of AI companions based on real-world user feedback. Wenya Wei, Sipeng Yang, Qixian Zhou, Xuelei Zhang, Yifu Yuan, Yongle Luo, Tianzhou Wang, Peipei Jin, Wangtong Liu, Xiaogang Jin 0001, Elvis S. Liu |
AAAI | 2 |
| 2026 | Lightmap Compression with Color-Coherent UV Clustering and Cascade Texture OptimizationabstractAbstract To address the storage overhead of lightmaps and the limitations of existing compression techniques, we propose a novel UV‐space compression framework based on per‐triangle processing. By mapping triangles to a standardized domain, we cluster and repack color‐coherent regions into a compact atlas, generating a cascade texture refined via differentiable rendering. Experimental results show an average storage reduction of 83% with approximately 10 dB higher PSNR than existing methods. Our approach is the first dedicated lightmap compression framework compatible with standard block‐based formats, offering an effective solution for memory‐efficient 3D asset delivery. Dehan Chen, Hongyu Huang 0001, Yuzhe Luo, Hao Xu 0049, Yuqing Zhang 0005, Sipeng Yang, Xifeng Gao, Heng Cai, Xiaogang Jin 0001 |
Comput. Graph. Forum | 6 |
| 2025 | Towards Realistic Example-based Modeling via 3D Gaussian StitchingabstractUsing parts of existing models to rebuild new models, commonly termed as example-based modeling, is a classical methodology in the realm of computer graphics. Previous works mostly focus on shape composition, making them very hard to use for realistic composition of 3D objects captured from real-world scenes. This leads to combining multiple NeRFs into a single 3D scene to achieve seamless appearance blending. However, the current SeamlessNeRF method struggles to achieve interactive editing and harmonious stitching for real-world scenes due to its gradient-based strategy and grid-based representation. To this end, we present an example-based modeling method that combines multiple Gaussian fields in a point-based representation using sample-guided synthesis. Specifically, as for composition, we create a GUI to segment and transform multiple fields in real time, easily obtaining a semantically meaningful composition of models represented by 3D Gaussian Splatting (3DGS). For texture blending, due to the discrete and irregular nature of 3DGS, straightforwardly applying gradient propagation as SeamlssNeRF is not supported. Thus, a novel sampling-based cloning method is proposed to harmonize the blending while preserving the original rich texture and content. Our workflow consists of three steps: 1) real-time segmentation and transformation of 3DGS using a well-tailored GUI, 2) KNN analysis to identify boundary points in the intersecting area between the source and target models, and 3) two-phase optimization of the target model using sampling-based cloning and gradient constraints. Extensive experimental results validate that our approach significantly outperforms previous works in realistic synthesis, demonstrating its practicality. Ziyi Yang 0008, Bingchen Gong, Xiaoguang Han 0001, Sipeng Yang, Xiaogang Jin 0001 |
CVPR | 5 |
| 2025 | Lightweight, Edge-Aware, and Temporally Consistent Supersampling for Mobile Real-Time RenderingabstractSupersampling has proven highly effective in enhancing visual fidelity by reducing aliasing, increasing resolution, and generating interpolated frames. It has become a standard component of modern real-time rendering pipelines. However, on mobile platforms, deep learning-based supersampling methods remain impractical due to stringent hardware constraints, while non-neural supersampling techniques often fall short in delivering perceptually high-quality results. In particular, producing visually pleasing reconstructions and temporally coherent interpolations is still a significant challenge in mobile settings. In this work, we present a novel, lightweight supersampling framework tailored for mobile devices. Our approach substantially improves both image reconstruction quality and temporal consistency while maintaining real-time performance. For super-resolution, we propose an intra-pixel object coverage estimation method for reconstructing high-quality anti-aliased pixels in edge regions, a gradient-guided strategy for non-edge areas, and a temporal sample accumulation approach to improve overall image quality. For frame interpolation, we develop an efficient motion estimation module coupled with a lightweight fusion scheme that integrates both estimated optical flow and rendered motion vectors, enabling temporally coherent interpolation of object dynamics and lighting variations. Extensive experiments demonstrate that our method consistently outperforms existing baselines in both perceptual image quality and temporal smoothness, while maintaining real-time performance on mobile GPUs. A demo application and supplementary materials are available on the project page. Sipeng Yang, Jiayu Ji, Junhao Zhuge, Jinzhe Zhao, Chen Li 0062, Yuzhong Yan, Kerong Wang, Lingqi Yan 0001, Xiaogang Jin 0001 |
ACM Trans. Graph. | 1 |
| 2025 | Accelerating Stereo Rendering via Image Reprojection and Spatio-Temporal SupersamplingabstractAchieving immersive virtual reality (VR) experiences typically requires extensive computational resources to ensure high-definition visuals, high frame rates, and low latency in stereoscopic rendering. This challenge is particularly pronounced for lower-tier and standalone VR devices with limited processing power. To accelerate rendering, existing supersampling and image reprojection techniques have shown significant potential, yet to date, no previous work has explored their combination to minimize stereo rendering overhead. In this paper, we introduce a lightweight supersampling framework that integrates image projection with spatio-temporal supersampling to accelerate stereo rendering. Our approach effectively leverages the temporal and spatial redundancies inherent in stereo videos, enabling rapid image generation for unshaded viewpoints and providing resolution-enhanced and anti-aliased images for binocular viewpoints. We first blend a rendered low-resolution (LR) frame with accumulated temporal samples to construct an high-resolution (HR) frame. This HR frame is then reprojected to the other viewpoint to directly synthesize a new image. To address disocclusions in reprojected images, we utilize accumulated history data and low-pass filtering for filling, ensuring high-quality results with minimal delay. Extensive evaluations on both the PC and the standalone device confirm that our framework requires short runtime to generate high-fidelity images, making it an effective solution for stereo rendering across various VR platforms. Sipeng Yang, Junhao Zhuge, Jiayu Ji, Qingchuan Zhu, Xiaogang Jin 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Fast best viewpoint selection with geometry-enhanced multiple views and cross-modal distillation
Zidi Cao, Jiayi Han, Sipeng Yang, Xiaogang Jin 0001 |
Vis. Comput. | 3 |
| 2024 | SR-VFA: Accurate Self-Refined Face Alignment in VideosabstractFace alignment is a critical and difficult task for many facial analysis applications. Existing VFA methods frequently ignore the consistency of facial geometries and textures across video sequences, limiting their ability to handle accurate and stable face alignment. This paper describes a robust and highly accurate 3D Morphable Model (3DMM)-based VFA approach that employs a novel texture generation method and a self-refined face alignment procedure. Our method iteratively fine-tunes facial geometries, textures, and poses by using a differentiable rendering technique and a self-refined optimization method. Experiment results show that our method outperforms existing state-of-the-art methods in terms of both accuracy and temporal stability. Visual results and source code are available at: https://pawindergit.github.io/SR-VFA/ Sipeng Yang, Hongyu Huang 0001, Qingchuan Zhu, Xiaogang Jin 0001 |
ICASSP | 1 |
| 2024 | Generated realistic noise and rotation-equivariant models for data-driven mesh denoising
Sipeng Yang, Wenhui Ren, Xiwen Zeng, Qingchuan Zhu, Hongbo Fu 0001, Kaijun Fan, Lei Yang 0048, Jingping Yu, Qilong Kou, Xiaogang Jin 0001 |
Comput. Aided Geom. Des. | 1 |
| 2024 | Facial action units detection using temporal context and feature reassignmentabstractAbstract Facial action units (AUs) encode the activations of facial muscle groups, playing a crucial role in expression analysis and facial animation. However, current deep learning AU detection methods primarily focus on single‐image analysis, which limits the exploitation of rich temporal context for robust outcomes. Moreover, the scale of available datasets remains limited, leading models trained on these datasets to tend to suffer from overfitting issues. This paper proposes a novel AU detection method integrating spatial and temporal data with inter‐subject feature reassignment for accurate and robust AU predictions. Our method first extracts regional features from facial images. Then, to effectively capture both the temporal context and identity‐independent features, we introduce a temporal feature combination and feature reassignment (TC&FR) module, which transforms single‐image features into a cohesive temporal sequence and fuses features across multiple subjects. This transformation encourages the model to utilize identity‐independent features and temporal context, thus ensuring robust prediction outcomes. Experimental results demonstrate the enhancements brought by the proposed modules and the state‐of‐the‐art (SOTA) results achieved by our method. Sipeng Yang, Hongyu Huang 0001, Ying Sophie Huang, Xiaogang Jin 0001 |
Comput. Animat. Virtual Worlds | 1 |
| 2024 | MNSS: Neural Supersampling Framework for Real-Time Rendering on Mobile DevicesabstractAlthough neural supersampling has achieved great success in various applications for improving image quality, it is still difficult to apply it to a wide range of real-time rendering applications due to the high computational power demand. Most existing methods are computationally expensive and require high-performance hardware, preventing their use on platforms with limited hardware, such as smartphones. To this end, we propose a new supersampling framework for real-time rendering applications to reconstruct a high-quality image out of a low-resolution one, which is sufficiently lightweight to run on smartphones within a real-time budget. Our model takes as input the renderer-generated low resolution content and produces high resolution and anti-aliased results. To maximize sampling efficiency, we propose using an alternate sub-pixel sample pattern during the rasterization process. This allows us to create a relatively small reconstruction model while maintaining high image quality. By accumulating new samples into a high-resolution history buffer, an efficient history check and re-usage scheme is introduced to improve temporal stability. To our knowledge, this is the first research in pushing real-time neural supersampling on mobile devices. Due to the absence of training data, we present a new dataset containing 57 training and test sequences from three game scenes. Furthermore, based on the rendered motion vectors and a visual perception study, we introduce a new metric called inter-frame structural similarity (IF-SSIM) to quantitatively measure the temporal stability of rendered videos. Extensive evaluations demonstrate that our supersampling model outperforms existing or alternative solutions in both performance and temporal stability. Sipeng Yang, Yunlu Zhao, Yuzhe Luo, He Wang 0002, Hongyu Sun 0001, Chen Li 0062, Binghuang Cai, Xiaogang Jin 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | 3D Shape Segmentation Using Soft Density Peak Clustering and Semi-Supervised Learning
Zhenyu Shu, Sipeng Yang, Shi-Qing Xin, Chaoyi Pang, Ladislav Kavan, Ligang Liu 0001 |
Comput. Aided Des. | 2 |
| 2022 | Detecting 3D Points of Interest Using Projective Neural NetworksabstractDetecting points of interest on 3D shapes is a fundamental research problem in geometry processing. Due to the complicated relationship between points of interest and their geometric features, detecting points of interest on any given 3D shape remains challenging. Due to the lack of training data, previous data-driven methods for detecting 3D points of interest mainly focus on utilizing hand-crafted geometric features to predict the probabilities of each point being a POI, which greatly limits detection performance. In this paper, we propose a novel algorithm for detecting 3D points of interest by using projective neural networks. Our method first projects the labeled training 3D shapes into multiple 2D views and then learns the required features from the 2D views in an end-to-end fashion. The points of interest on test 3D shapes are then automatically detected by applying the learned neural network and our improved density peak clustering. Our method relies neither on hand-crafted feature descriptors nor a large quantity of expensive 3D training data to obtain satisfactory results. Experimental results show significantly superior detection performance of our method over the state-of-the-art methods. Zhenyu Shu, Sipeng Yang, Shi-Qing Xin, Chaoyi Pang, Xiaogang Jin 0001, Ladislav Kavan, Ligang Liu 0001 |
IEEE Trans. Multim. | 2 |