EDBT 2026 Demo / reviewers in the wild / expert
Yuehao Wang
dblp:272/0852
· DBLP profile ↗
17ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Oscillation Inversion: Training-Free Image and Video Enhancement Through Oscillated Latents in Large Flow ModelsabstractWe explore the oscillatory behavior observed in inversion methods applied to large-scale flow models, including text-to-image and text-to-video. By employing an augmented fixed-point-inspired iterative approach to invert real-world images, we observe that the solution does not achieve convergence, instead oscillating between distinct clusters. Through both experiments on synthetic data, text-to-image and text-to-video, we demonstrate that these oscillating clusters exhibit notable semantic coherence. We offer theoretical insights, showing that this behavior arises from oscillatory dynamics in flow models. Building on this understanding, we introduce a simple and fast distribution transfer technique that facilitates training-free image and video editing/enhancement. Furthermore, we provide quantitative results demonstrating the effectiveness of our method on tasks such as image enhancement, editing, and reconstruction. Notably, our approach enables the transformation of image-only enhancers and editors into lightweight, video-capable tools—without additional training—highlighting its practical versatility and impact. Zhenxiao Liang, Xiaoyan Cong, Yi Yang 0001, Lanqing Guo, Yuehao Wang, Peihao Wang, Zhangyang Wang |
AAAI | 6 |
| 2026 | Understanding Emotional Closeness in Distanced Intergenerational Relationships Between Young Children and Older Relatives: A Scoping Review for HCIabstractEmotional closeness (EC) is central to family relationships, however in Human-Computer Interaction (HCI) it is often regarded as self-evident, invoked through adjacent constructs such as connection or co-presence. This ambiguity is particularly limiting for remote relationships between young children (aged 4-8 years) and their older relatives, where developmental asymmetries and generational roles shape how EC unfolds. To clarify how EC is understood in this specific intergenerational context, we conducted a scoping review of 30 papers (2010 - 2025) examining how EC is defined, evaluated, and technologically mediated. Our analysis reveals three key patterns: reliance on self-report evaluations, a persistent interaction-closeness assumption, and under-exploration of embodied and cultural framings. We synthesise a multidimensional definition of EC comprising Affective Expression, Relational Practices, Embodied Presence, and Cultural Belonging. We conclude with implications for HCI, including the need for multimodal and longitudinal methods and technologies that support multi-dimensional, culturally grounded, and meaningful intergenerational connection. Yuehao Wang, Alethea Blackler, Li Jiang 0016, Beheshteh Atrian, Shital Desai, Nicole Vickery, Bernd Ploderer, Jane Turner, Linda Knight |
CHI | 1 |
| 2026 | DE-UNet: an enhanced UNet with dual-branch attention convolution and efficient multi-scale attention aggregation for UAV lane line segmentation
Yuehao Wang, Haiqing Liu |
Vis. Comput. | 1 |
| 2025 | FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian Splattingabstract3D Gaussian splatting (3DGS) has enabled various applications in 3D scene representation and novel view synthesis due to its efficient rendering capabilities. However, 3DGS demands relatively significant GPU memory, limiting its use on devices with restricted computational resources. Previous approaches have focused on pruning less important Gaussians, effectively compressing 3DGS but often requiring a fine-tuning stage and lacking adaptability for the specific memory needs of different devices. In this work, we present an elastic inference method for 3DGS. Given an input for the desired model size, our method selects and transforms a subset of Gaussians, achieving substantial rendering performance without additional fine-tuning. We introduce a tiny learnable module that controls Gaussian selection based on the input percentage, along with a transformation module that adjusts the selected Gaussians to complement the performance of the reduced model. Comprehensive experiments on ZipNeRF, MipNeRF and Tanks&Temples scenes demonstrate the effectiveness of our approach. Code is available at https://flexgs.github.io/. Hengyu Liu 0007, Yuehao Wang, Chenxin Li, Ruisi Cai, Wuyang Li, Pavlo Molchanov 0001, Peihao Wang, Zhangyang Wang |
CVPR | 2 |
| 2025 | Steepest Descent Density Control for Compact 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has emerged as a powerful technique for real-time, high-resolution novel view synthesis. By representing scenes as a mixture of Gaussian primitives, 3DGS leverages GPU rasterization pipelines for efficient rendering and reconstruction. To optimize scene coverage and capture fine details, 3DGS employs a densification algorithm to generate additional points. However, this process often leads to redundant point clouds, resulting in excessive memory usage, slower performance, and substantial storage demands–posing significant challenges for deployment on resource-constrained devices. To address this limitation, we propose a theoretical framework that demystifies and improves density control in 3DGS. Our analysis reveals that splitting is crucial for escaping saddle points. Through an optimization-theoretic approach, we establish the necessary conditions for densification, determine the minimal number of offspring Gaussians, identify the optimal parameter update direction, and provide an analytical solution for normalizing off-spring opacity. Building on these insights, we introduce SteepGS, incorporating steepest density control, a principled strategy that minimizes loss while maintaining a compact point cloud. SteepGS achieves a ~ 50% reduction in Gaussian points without compromising rendering quality, significantly enhancing both efficiency and scalability. Peihao Wang, Yuehao Wang, Dilin Wang, Sreyas Mohan, Zhiwen Fan, Lemeng Wu, Ruisi Cai, Yu-Ying Yeh, Zhangyang Wang, Qiang Liu 0001 |
CVPR | 2 |
| 2025 | Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothingabstractStructured State Space Models (SSMs) have emerged as alternatives to transformers.
While SSMs are often regarded as effective in capturing long-sequence dependencies, we rigorously demonstrate that they are inherently limited by strong recency bias.
Our empirical studies also reveal that this bias impairs the models' ability to recall distant information and introduces robustness issues. Our scaling experiments then discovered that deeper structures in SSMs can facilitate the learning of long contexts.
However, subsequent theoretical analysis reveals that as SSMs increase in depth, they exhibit another inevitable tendency toward over-smoothing, e.g., token representations becoming increasingly indistinguishable.
This *fundamental dilemma* between recency and over-smoothing hinders the scalability of existing SSMs.
Inspired by our theoretical findings, we propose to *polarize* two channels of the state transition matrices in SSMs, setting them to zero and one, respectively, simultaneously addressing recency bias and over-smoothing.
Experiments demonstrate that our polarization technique consistently enhances the associative recall accuracy of long-range tokens and unlocks SSMs to benefit further from deeper architectures.
All source codes are released at https://github.com/VITA-Group/SSM-Bottleneck. Peihao Wang, Ruisi Cai, Yuehao Wang, Pragya Srivastava, Zhangyang Wang, Pan Li 0005 |
ICLR | 3 |
| 2025 | SAS: Simulated Attention ScoreabstractThe attention mechanism is a core component of the Transformer architecture.
Various methods have been developed to compute attention scores, including multi-head attention (MHA), multi-query attention, group-query attention and so on. We further analyze the MHA and observe that its performance improves as the number of attention heads increases, provided the hidden size per head remains sufficiently large. Therefore, increasing both the head count and hidden size per head with minimal parameter overhead can lead to significant performance gains at a low cost.
Motivated by this insight, we introduce Simulated Attention Score (SAS), which **maintains a compact model size while simulating a larger number of attention heads and hidden feature dimension per head.** This is achieved by projecting a low-dimensional head representation into a higher-dimensional space, effectively increasing attention capacity without increasing parameter count. Beyond the head representations, we further extend the simulation approach to feature dimension of the key and query embeddings, enhancing expressiveness by mimicking the behavior of a larger model while preserving the original model size.
**To control the parameter cost, we also propose Parameter-Efficient Attention Aggregation (PEAA).**
Comprehensive experiments on a variety of datasets and tasks demonstrate the effectiveness of the proposed SAS method, achieving significant improvements over different attention variants. Chuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang, Peihao Wang, Liliang Ren, Hao Cheng 0002, Janardhan Kulkarni, Yelong Shen, Zhangyang Wang, Mac Schwager, Anderson Schneider, Jianfeng Gao 0001 |
NeurIPS | 4 |
| 2025 | ICDNSGA: Identification of Potential circRNA-Disease Associations Based on Improved Non-Dominated Sorting Genetic AlgorithmabstractIncreasing biological research indicates that the expression levels of circRNAs fluctuate during the onset of various diseases, making them potential biomarkers for multiple conditions. Although numerous artificial intelligence-based computational methods are currently employed for circRNA-disease associations prediction, these methods often rely on a single objective function, which can lead to suboptimal prediction accuracy. To date, no method has designed a set of multi-objective functions specifically for the circRNA-disease prediction problem and optimized it using a non-dominated sorting genetic algorithm. This paper introduces a novel approach by utilizing multi-objective functions and an improved non-dominated sorting genetic algorithm (ICDNSGA) to identify potential associations of circRNA-disease. The method constructs a solution space through matrix factorization and network community characteristics, designing four distinct objective functions optimized via the enhanced multi-objective non-dominated sorting genetic algorithm. ICDNSGA incorporates a population-based adaptive normalization strategy, improving algorithm convergence and solution diversity. Experimental results show that ICDNSGA outperforms pure matrix factorization methods, non-dominated sorting genetic algorithms and other machine learning techniques in predictive performance. Additionally, the prediction results can be validated through existing research and biological analyses, underscoring ICDNSGA's potential as a valuable tool for biomedical experimentation. Yuehao Wang, Pengli Lu |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | EndoGSLAM: Real-Time Dense Reconstruction and Tracking in Endoscopic Surgeries Using Gaussian Splatting
Kailing Wang, Chen Yang 0023, Yuehao Wang, Sikuang Li, Yan Wang 0033, Qi Dou 0001, Xiaokang Yang 0001, Wei Shen 0002 |
MICCAI (6) | 3 |
| 2024 | RDGAN: Prediction of circRNA-Disease Associations via Resistance Distance and Graph Attention NetworkabstractAs a series of single-stranded RNAs, circRNAs have been implicated in numerous diseases and can serve as valuable biomarkers for disease therapy and prevention. However, traditional biological experiments demand significant time and effort. Therefore, various computational methods have been proposed to address this limitation, but how to extract features more comprehensively remains a challenge that needs further attention in the future. In this study, we propose a unique approach to predict circRNA-disease associations based on resistance distance and graph attention network (RDGAN). First, the associations of circRNA and disease are obtained by fusing multiple databases, and resistance distance as a similarity matrix is used to further deal with the sparse of the similarity matrices. Then the circRNA-disease heterogeneous network is constructed based on the similiarity of circRNA-circRNA, disease-disease and the known circRNA-disease adjacency matric. Second, leveraging the three neural network modules-ResGatedGraphConv, GAT and MFConv-we gather node feature embeddings collected from the heterogeneous network. Subsequently, all the characteristics are supplied to the self-attention mechanism to predict new potential connections. Finally, our model obtains a remarkable AUC value of 0.9630 through five-fold cross-validation, surpassing the predictive performance of the other eight state-of-the-art models. Pengli Lu, Yuehao Wang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | Efficient Deformable Tissue Reconstruction via Orthogonal Neural PlaneabstractIntraoperative imaging techniques for reconstructing deformable tissues in vivo are pivotal for advanced surgical systems. Existing methods either compromise on rendering quality or are excessively computationally intensive, often demanding dozens of hours to perform, which significantly hinders their practical application. In this paper, we introduce Fast Orthogonal Plane (Forplane), a novel, efficient framework based on neural radiance fields (NeRF) for the reconstruction of deformable tissues. We conceptualize surgical procedures as 4D volumes, and break them down into static and dynamic fields comprised of orthogonal neural planes. This factorization discretizes the four-dimensional space, leading to a decreased memory usage and faster optimization. A spatiotemporal importance sampling scheme is introduced to improve performance in regions with tool occlusion as well as large motions and accelerate training. An efficient ray marching method is applied to skip sampling among empty regions, significantly improving inference speed. Forplane accommodates both binocular and monocular endoscopy videos, demonstrating its extensive applicability and flexibility. Our experiments, carried out on two in vivo datasets, the EndoNeRF and Hamlyn datasets, demonstrate the effectiveness of our framework. In all cases, Forplane substantially accelerates both the optimization process (by over 100 times) and the inference process (by over 15 times) while maintaining or even improving the quality across a variety of non-rigid deformations. This significant performance improvement promises to be a valuable asset for future intraoperative surgical applications. The code of our project is now available at https://github.com/Loping151/ForPlane. Chen Yang 0023, Kailing Wang, Yuehao Wang, Qi Dou 0001, Xiaokang Yang 0001, Wei Shen 0002 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Bilateral Guided Radiance Field ProcessingabstractNeural Radiance Fields (NeRF) achieves unprecedented performance in synthesizing novel view synthesis, utilizing multi-view consistency. When capturing multiple inputs, image signal processing (ISP) in modern cameras will independently enhance them, including exposure adjustment, color correction, local tone mapping, etc. While these processings greatly improve image quality, they often break the multi-view consistency assumption, leading to "floaters" in the reconstructed radiance fields. To address this concern without compromising visual aesthetics, we aim to first disentangle the enhancement by ISP at the NeRF training stage and re-apply user-desired enhancements to the reconstructed radiance fields at the finishing stage. Furthermore, to make the re-applied enhancements consistent between novel views, we need to perform imaging signal processing in 3D space (i.e. "3D ISP"). For this goal, we adopt the bilateral grid, a locally-affine model, as a generalized representation of ISP processing. Specifically, we optimize per-view 3D bilateral grids with radiance fields to approximate the effects of camera pipelines for each input view. To achieve user-adjustable 3D finishing, we propose to learn a low-rank 4D bilateral grid from a given single view edit, lifting photo enhancements to the whole 3D scene. We demonstrate our approach can boost the visual quality of novel view synthesis by effectively removing floaters and performing enhancements from user retouching. The source code and our data are available at: https://bilarfpro.github.io. Yuehao Wang, Bingchen Gong, Tianfan Xue |
ACM Trans. Graph. | 1 |
| 2023 | Neural LerPlane Representations for Fast 4D Reconstruction of Deformable Tissues
Chen Yang 0023, Kailing Wang, Yuehao Wang, Xiaokang Yang 0001, Wei Shen 0002 |
MICCAI (9) | 3 |
| 2023 | RecolorNeRF: Layer Decomposed Radiance Fields for Efficient Color Editing of 3D ScenesabstractRadiance fields have gradually become a main representation of media. Although its appearance editing has been studied, how to achieve view-consistent recoloring in an efficient manner is still under explored. We present RecolorNeRF, a novel user-friendly color editing approach for the neural radiance fields. Our key idea is to decompose the scene into a set of pure-colored layers, forming a palette. By this means, color manipulation can be conducted by altering the color components of the palette directly. To support efficient palette-based editing, the color of each layer needs to be as representative as possible. In the end, the problem is formulated as an optimization problem, where the layers and their blending weights are jointly optimized with the NeRF itself. Extensive experiments show that our jointly-optimized layer decomposition can be used against multiple backbones and produce photo-realistic recolored novel-view renderings. We demonstrate that RecolorNeRF outperforms baseline methods both quantitatively and qualitatively for color editing even in complex real-world scenes. Bingchen Gong, Yuehao Wang, Xiaoguang Han 0001, Qi Dou 0001 |
ACM Multimedia | 2 |
| 2023 | SeamlessNeRF: Stitching Part NeRFs with Gradient PropagationabstractNeural Radiance Fields (NeRFs) have emerged as promising digital mediums of 3D objects and scenes, sparking a surge in research to extend the editing capabilities in this domain. The task of seamless editing and merging of multiple NeRFs, resembling the “Poisson blending” in 2D image editing, remains a critical operation that is under-explored by existing work. To fill this gap, we propose SeamlessNeRF, a novel approach for seamless appearance blending of multiple NeRFs. In specific, we aim to optimize the appearance of a target radiance field in order to harmonize its merge with a source field. We propose a well-tailored optimization procedure for blending, which is constrained by 1) pinning the radiance color in the intersecting boundary area between the source and target fields and 2) maintaining the original gradient of the target. Extensive experiments validate that our approach can effectively propagate the source appearance from the boundary area to the entire target field through the gradients. To the best of our knowledge, SeamlessNeRF is the first work that introduces gradient-guided appearance editing to radiance fields, offering solutions for seamless stitching of 3D objects represented in NeRFs. Our code and more results are available at https://sites.google.com/view/seamlessnerf. Bingchen Gong, Yuehao Wang, Xiaoguang Han 0001, Qi Dou 0001 |
SIGGRAPH Asia | 2 |
| 2022 | Neural Rendering for Stereo 3D Reconstruction of Deformable Tissues in Robotic Surgery
Yuehao Wang, Yonghao Long 0001, Siu Hin Fan, Qi Dou 0001 |
MICCAI (8) | 1 |
| 2020 | Multi-View Neural Human RenderingabstractWe present an end-to-end Neural Human Renderer (NHR) for dynamic human captures under the multi-view setting. NHR adopts PointNet++ for feature extraction (FE) to enable robust 3D correspondence matching on low quality, dynamic 3D reconstructions. To render new views, we map 3D features onto the target camera as a 2D feature map and employ an anti-aliased CNN to handle holes and noises. Newly synthesized views from NHR can be further used to construct visual hulls to handle textureless and/or dark regions such as black clothing. Comprehensive experiments show NHR significantly outperforms the state-of-the-art neural and image-based rendering techniques, especially on hands, hair, nose, foot, etc. Minye Wu, Yuehao Wang, Qiang Hu 0003, Jingyi Yu 0001 |
CVPR | 2 |