VLDB 2026 Research / reviewers in the wild / expert
Yu Guo 0006
dblp:53/382-6
· DBLP profile ↗
34ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0002-5489-8288ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 5 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improve noise tolerance of robust feature selection via block-sparse projection learning
Jie Wang 0164, Zheng Wang 0037, Yu Guo 0006, Rong Wang 0001, Fei Wang 0008, Feiping Nie 0001 |
Pattern Recognit. | 3 |
| 2025 | Diffusion-based Realistic Listening Head Generation via Hybrid Motion ModelingabstractListening head generation aims to synthesize non-verbal responsive listening head videos that naturally react to a certain speaker, for which, both realistic head movements, expressive facial expressions, and high visual qualities are expected. Previous approaches typically follow a two-stage pipeline that first generates intermediate 3D motion signals such as 3DMM coefficients, and then synthesizes the videos by deterministic rendering, suffering from limited motion expressiveness and low visual quality (e.g. 256×256). In this work, we propose a novel listening head generation method that harnesses the generative capabilities of the diffusion model for both motion generation and high-quality rendering. Crucially, we propose an effective hybrid motion modeling module that addresses training difficulties caused by the scarcity of listening head data while preserving the intricate details that may be lost in explicit motion representations. We further develop a tailored control guidance for head pose and facial expression, by integrating their intrinsic motion characteristics. Our method enables high-fidelity video generation with 512 × 512 resolution and delivers vivid listener motion feedback. We conduct comprehensive experiments and obtain superior performance in terms of both visual quality and motion expressiveness compared with existing methods. Yanbo Fan, Xuan Wang 0009, Yu Guo 0006, Fei Wang 0008 |
CVPR | 4 |
| 2025 | Fine-Grained 3D Gaussian Head Avatars Modeling from Static Captures Via Joint Reconstruction and Registration
Yuan Sun 0003, Xuan Wang 0009, WeiLi Zhang, Yanbo Fan, Yu Guo 0006, Fei Wang 0008 |
ICCV | 6 |
| 2025 | UltraVSR: Achieving Ultra-Realistic Video Super-Resolution with Efficient One-Step Diffusion Space
Yong Liu 0031, Jinshan Pan, Yinchuan Li, Qingji Dong, Chao Zhu 0007, Yu Guo 0006, Fei Wang 0008 |
ACM Multimedia | 6 |
| 2025 | Revitalizing Image Dehazing in the Real World: A High-Quality Dataset and a Customized MethodabstractExisting dehazing methods face challenges in generalization due to the lack of paired real-world training data and tailored models. Recently, some semi-supervised/unsupervised schemes have been explored, achieving impressive performance. However, their performance still depends heavily on synthetic training data and the introduced prior-based strong constraints do not always hold. In this paper, we first introduce RealHQ-HAZE, a new dataset with 200 collected real-world hazy images, 200 corresponding carefully rendered haze-free images, and an additional 1000 varicolored hazy images transferred from the collected images. We also propose a prior-compensated multi-stage dehazing network, PMDN, which can learn different levels of real-world haze distribution through multi-stage progressive learning. To utilize prior knowledge effectively, we introduce a prior-based feature compensation module, guiding intermediate results with an adaptive weight. Additionally, we propose a MixCut consistent dehazing strategy to mix paired and derived images using a cross-cutting scheme, reinforcing dehazing through consistency principles. Extensive experiments demonstrate the effectiveness of our dataset and the superiority of PMDN compared to existing state-of-the-art dehazing methods. Yong Liu 0031, Qingji Dong, Chao Zhu 0007, Yu Guo 0006, Fei Wang 0008 |
Comput. Vis. Media | 4 |
| 2024 | Multi-View Subspace Clustering With Consensus Graph Contrastive LearningabstractA significant challenge in multi-view clustering lies in the comprehensive extraction of consistency and complementary information from heterogeneous multi-view data. Numerous methods employ contrastive learning techniques to explore the information between views. However, the basic contrastive learning strategy does not consider cluster information when constructing sample pairs, potentially leading to the emergence of false negative pairs (FNPs). To tackle this concern, we propose a Multi-view Subspace Clustering with Consensus Graph Contrastive Learning (CGCL) model. Specifically, a self-representation layer is designed to acquire a consensus graph that elucidates the overall data distribution. Furthermore, a contrastive learning layer utilizes the cluster information embedded in the consensus graph to yield reliable sample pairs, resulting in a reduction of the detrimental FNPs and the extraction of complementary information from the various views. Extensive experiments on public datasets demonstrate the effectiveness of CGCL. Jie Zhang 0090, Yuan Sun 0003, Yu Guo 0006, Zheng Wang 0037, Feiping Nie 0001, Fei Wang 0008 |
ICASSP | 3 |
| 2024 | Spectral Aggregation Cross-Square Transformer for Hyperspectral Image Denoising
Yang Liu 0385, Yantao Ji, Jiahua Xiao, Yu Guo 0006, Peilin Jiang, Haiwei Yang, Fei Wang 0008 |
ICPR (15) | 4 |
| 2024 | Match Normalization: Learning-Based Point Cloud Registration for 6D Object Pose Estimation in the Real WorldabstractIn this work, we tackle the task of estimating the 6D pose of an object from point cloud data. While recent learning-based approaches have shown remarkable success on synthetic datasets, we have observed them to fail in the presence of real-world data. We investigate the root causes of these failures and identify two main challenges: The sensitivity of the widely-used SVD-based loss function to the range of rotation between the two point clouds, and the difference in feature distributions between the source and target point clouds. We address the first challenge by introducing a directly supervised loss function that does not utilize the SVD operation. To tackle the second, we introduce a new normalization strategy, Match Normalization. Our two contributions are general and can be applied to many existing learning-based 3D object registration frameworks, which we illustrate by implementing them in two of them, DCP and IDAM. Our experiments on the real-scene TUD-L Hodan et al. 2018, LINEMOD Hinterstoisser et al. 2012 and Occluded-LINEMOD Brachmann et al. 2014 datasets evidence the benefits of our strategies. They allow for the first-time learning-based 3D object registration methods to achieve meaningful results on real-world data. We therefore expect them to be key to the future developments of point cloud registration methods. Zheng Dang, Lizhou Wang, Yu Guo 0006, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Dual-path dehazing network with spatial-frequency feature fusion
Li Wang 0072, Hang Dong 0001, Chao Zhu 0007, Huibin Tao, Yu Guo 0006, Fei Wang 0008 |
Pattern Recognit. | 6 |
| 2024 | Joint learning of latent subspace and structured graph for multi-view clustering
Yu Guo 0006, Zheng Wang 0037, Fei Wang 0008 |
Pattern Recognit. | 2 |
| 2024 | Double-Structured Sparsity Guided Flexible Embedding Learning for Unsupervised Feature SelectionabstractIn this article, we propose a novel unsupervised feature selection model combined with clustering, named double-structured sparsity guided flexible embedding learning (DSFEL) for unsupervised feature selection. DSFEL includes a module for learning a block-diagonal structural sparse graph that represents the clustering structure and another module for learning a completely row-sparse projection matrix using the$\ell_{2,0}$-norm constraint to select distinctive features. Compared with the commonly used$\ell_{2,1}$-norm regularization term, the$\ell_{2,0}$-norm constraint can avoid the drawbacks of sparsity limitation and parameter tuning. The optimization of the$\ell_{2,0}$-norm constraint problem, which is a nonconvex and nonsmooth problem, is a formidable challenge, and previous optimization algorithms have only been able to provide approximate solutions. In order to address this issue, this article proposes an efficient optimization strategy that yields a closed-form solution. Eventually, through comprehensive experimentation on nine real-world datasets, it is demonstrated that the proposed method outperforms existing state-of-the-art unsupervised feature selection methods. Yu Guo 0006, Yuan Sun 0003, Zheng Wang 0037, Feiping Nie 0001, Fei Wang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Local-to-Global Registration for Bundle-Adjusting Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have achieved photorealistic novel views synthesis; however, the requirement of accurate camera poses limits its application. Despite analysis-by-synthesis extensions for jointly learning neural3D representations and registering camera frames exist, they are susceptible to suboptimal solutions if poorly initialized. We propose L2G-NeRF, a Local-to-Global registration method for bundle-adjusting Neural Radiance Fields: first, a pixel-wise flexible alignment, followed by a framewise constrained parametric alignment. Pixel-wise local alignment is learned in an unsupervised way via a deep network which optimizes photometric reconstruction errors. framewise global alignment is performed using differentiable parameter estimation solvers on the pixel-wise correspondences to find a global transformation. Experiments on synthetic and real-world data show that our method outperforms the current state-of-the-art in terms of high-fidelity reconstruction and resolving large camera pose misalignment. Our module is an easy-to-use plugin that can be applied to NeRF variants and other neural field applications. The Code and supplementary materials are available at https://rover-xingyu.github.io/L2G-NeRF/. Xuan Wang 0009, Qi Zhang 0029, Yu Guo 0006, Ying Shan, Fei Wang 0008 |
CVPR | 5 |
| 2023 | UV Volumes for Real-time Rendering of Editable Free-view Human PerformanceabstractNeural volume rendering enables photo-realistic renderings of a human performer in free-view, a critical task in immersive VR/AR applications. But the practice is severely limited by high computational costs in the rendering process. To solve this problem, we propose the UV Volumes, a new approach that can render an editable free-view video of a human performer in real-time. It separates the high-frequency (i.e., non-smooth) human appearance from the 3D volume, and encodes them into 2D neural texture stacks (NTS). The smooth UV volumes allow much smaller and shallower neural networks to obtain densities and texture coordinates in 3D while capturing detailed appearance in 2D NTS. For editability, the mapping between the parameterized human model and the smooth texture coordinates allows us a better generalization on novel poses and shapes. Furthermore, the use of NTS enables interesting applications, e.g., retexturing. Extensive experiments on CMU Panoptic, ZJU Mocap, and H36M datasets show that our model can render$960\times 540$images in 30FPS on average with comparable photo-realism to state-of-the-art methods. The project and supplementary materials are available at https://fanegg.github.io/UV-Volumes. Xuan Wang 0009, Qi Zhang 0029, Xiaoyu Li 0002, Yu Guo 0006, Jue Wang 0001, Fei Wang 0008 |
CVPR | 6 |
| 2023 | SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationabstractGenerating talking head videos through a face image and a piece of speech audio still contains many challenges. i.e., unnatural head movement, distorted expression, and identity modification. We argue that these issues are mainly caused by learning from the coupled 2D motion fields. On the other hand, explicitly using 3D information also suffers problems of stiff expression and incoherent video. We present SadTalker, which generates 3D motion coefficients (head pose, expression) of the 3DMM from audio and implicitly modulates a novel 3D-aware face render for talking head generation. To learn the realistic motion coefficients, we explicitly model the connections between audio and different types of motion coefficients individually. Precisely, we present ExpNet to learn the accurate facial expression from audio by distilling both coefficients and 3D-rendered faces. As for the head pose, we design PoseVAE via a conditional VAE to synthesize head motion in different styles. Finally, the generated 3D motion coefficients are mapped to the unsupervised 3D keypoints space of the proposed face render to synthesize the final video. We conducted extensive experiments to demonstrate the superiority of our method in terms of motion and video quality.11The code and demo videos are available at https://sadtalker.github.io. Xiaodong Cun, Xuan Wang 0009, Yong Zhang 0034, Xi Shen 0001, Yu Guo 0006, Ying Shan, Fei Wang 0008 |
CVPR | 6 |
| 2023 | Boosting Video Super Resolution with Patch-Based Temporal Redundancy Optimization
Hang Dong 0001, Jinshan Pan, Chao Zhu 0007, Boyang Liang, Yu Guo 0006, Ding Liu 0001, Lean Fu, Fei Wang 0008 |
ICANN (7) | 6 |
| 2022 | Learning-Based Point Cloud Registration for 6D Object Pose Estimation in the Real World
Zheng Dang, Lizhou Wang, Yu Guo 0006, Mathieu Salzmann |
ECCV (1) | 3 |
| 2022 | Multi-View Stereo and Depth Priors Guided NeRF for View SynthesisabstractIn this paper, we present a new framework for view synthesis of novel view based on Neural Radiance Fields(NeRF). We aim to address two main limitations of NeRF. Firstly, we propose to combine multi-view stereo into NeRF to help construct general neural radiance fields across different scenes. Specifically, We build a MVS-Encoding Feature Volume with average groupwise correlation to aggregate the multi-view appearance and geometry feature for every source view. And then we use an MLP to encode neural radiance fields by using the scene-dependent features interpolated from the MVS-Encoding Feature Volumes. This makes our model can be applied to other unseen scenes without any per-scene fine-tuning, and render realistic images with few images. If more training images are provided, our method can be fine-tuned quickly to render more realistic images. In fine-tuning phase, we propose a depth priors guided sampling method, which can make the model represent more accurate geometry for corresponding scenes and so render high-quality images of novel view. We evaluate our method on three common datasets. The experiment results show that our method performs better than other baselines, neither without or with fine-tuning. And the depth priors guided sampling method can be easily applied on other methods based on Neural Radiance Fields to further improve the quality of rendered images. Wang Deng, Xuetao Zhang 0001, Yu Guo 0006 |
ICPR | 3 |
| 2022 | Frequency-aware Deep Dual-path Feature Enhancement Network for Image DehazingabstractSingle image dehazing is a challenging task due to the severe degradations caused by the particles in the air. Recently, various CNN-based methods have been proposed and they have achieved promising results on some dehazing tasks. However, the existing end-to-end dehazing networks process high-frequency information and low-frequency information at the same time. Therefore, most dehazing methods cannot restore dehazed image with satisfying high-frequency details. In this paper, we propose a Frequency-aware deep Dual-path Feature enhancement Network (FDF-Net) to better restore the high-frequency information while removing the haze. To achieve this, we introduce a Dual-path Feature Enhancement (DFE) block, which contains two branches: one path is to remedy the missing spatial information from high-resolution features, and the other one is to obtain new features to increase the variety of features. We believe the dual-path architecture can help the first path to focus on the recovering the high-frequency information. Furthermore, to reserve more detailed image information from the features with larger resolution, we adopt a wavelet transform module during the downsampling process of the encoder module to directly pass the high frequency information to the next level. The extensive experiments show the superiority of the proposed model over previous methods on the benchmark datasets as well as real-world hazy images. Hang Dong 0001, Li Wang 0072, Boyang Liang, Yu Guo 0006, Fei Wang 0008 |
ICPR | 5 |
| 2022 | Adaptive weighted robust iterative closest point
Yu Guo 0006, Luting Zhao, Xuetao Zhang 0001, Shaoyi Du, Fei Wang 0008 |
Neurocomputing | 1 |
| 2021 | Photometric Stereo Based on Multiple Kernel Learning
Yu Guo 0006, Xiaoxiao Yang, Xuetao Zhang 0001, Fei Wang 0008 |
ICIG (3) | 2 |
| 2021 | Dynamic Hypergraph Regularized Broad Learning System for Image Classification
Xiaoxiao Yang, Yu Guo 0006, Peilin Jiang, Fei Wang 0008 |
ICIG (1) | 2 |
| 2021 | Monocular 3D multi-person pose estimation via predicting factorized correction factors
Yu Guo 0006, Lichen Ma, Zhi Li 0055, Xuan Wang 0009, Fei Wang 0008 |
Comput. Vis. Image Underst. | 1 |
| 2021 | Online robust echo state broad learning system
Yu Guo 0006, Xiaoxiao Yang, Fei Wang 0008, Badong Chen |
Neurocomputing | 1 |
| 2020 | Deep Multi-Scale Gabor Wavelet Network for Image RestorationabstractDue to the limitations of the imaging processors and complex weather conditions, image degradation is often inevitable. Existing deep learning-based image restoration methods often rely on the powerful feature representation capacity of deep networks and pay less attention to the inherent properties of the degradation signal, e.g. variations in spatial scale and orientations across the image, which makes them ineffective for the image restoration tasks. In this paper, we propose a Multiscale Gabor Wavelet Network (MsGWN) for image restoration. We apply the multi-scale architecture to extract the contaminated feature from input at different spatial scales, and thus the contaminated feature can be effectively restored in a corse- to-fine manner. However, using multi-scale architecture alone cannot remove the degradations with different orientations. To overcome this problem, we introduce a Gabor Wavelet Module (GWM) to further extract the contaminated features from four orientations. By decomposing the features into four multi-orientation components, the restoration process can be facilitated by avoiding learning the mixed degradations all-in- one. We evaluate the proposed method on image demoirding, image deraining, and image dehazing. Experiments on these applications demonstrate that the proposed method can achieve favorable results against the state-of-the-art approaches. Hang Dong 0001, Xinyi Zhang 0005, Yu Guo 0006, Fei Wang 0008 |
ICASSP | 3 |
| 2020 | Recursive Maximum Correntropy Criterion Based Randomized Recurrent Broad Learning System
Yu Guo 0006, Fei Wang 0008 |
ICONIP (5) | 2 |
| 2019 | Coarse-to-Fine 3D Human Pose Estimation
Yu Guo 0006, Lin Zhao 0003, Shanshan Zhang 0001, Jian Yang 0003 |
ICIG (3) | 1 |
| 2019 | Gated Contiguous Memory U-Net for Single Image Dehazing
Hang Dong 0001, Fei Wang 0008, Yu Guo 0006, Kaisheng Ma |
ICONIP (2) | 4 |
| 2018 | A Deep Encoder-Decoder Networks for Joint Deblurring and Super-ResolutionabstractIn this paper, we propose an end-to-end convolution neural network (CNN) to restore a clear high-resolution image from a severely blurry image. It's a highly ill-posed problem and brings tremendous challenges to state-of-art deblurring or super-resolution (SR) methods. A straightforward way to solve this problem is to concatenate two types of networks directly. However, experiments show that the concatenation of independent networks increases computation complexity instead of generating satisfying high-resolution images. Consequently, we focus on designing a single deep network to solve the deblurring and SR problems in parallel. Our method, called ED-DSRN, extends the traditional Super-Resolution network by adding a deblurring branch that shares the same feature maps extracted from an encoder-decoder module with the original SR branch. Extensive experiments show that our method produces remarkable deblurred and super-resolved images simultaneously with high efficiency. Xinyi Zhang 0005, Fei Wang 0008, Hang Dong 0001, Yu Guo 0006 |
ICASSP | 4 |
| 2018 | Generalized Maximum Correntropy-Based Echo State Network for Robust Nonlinear System IdentificationabstractIn this paper, we propose a robust method for non-linear system identification that incorporates robustness to echo state networks (ESNs). In particular, the ESNs utilize generalized correntropy as a loss function to get optimal solutions. Generalized correntropy is a more flexible extension of correntropy in information theoretic learning (ITL). Generalized correntropy induced metric (GCIM) is robust to outliers with a proper shape parameter. The ESNs with GCIM can provide the anti-noise capacity and are insensitive outliers which are prevalent in real-world tasks. They also inherit the basic architecture of echo state network but replaces the commonly used mean square error (MSE) criterion with GCIM. The stochastic gradient descent method is adopted to optimize the generalized correntropy-based cost function. Numerical simulations are given to show that the proposed algorithm is robust to the non-Gaussian noise and outliers. Changhao Zhang, Yu Guo 0006, Fei Wang 0008, Badong Chen |
IJCNN | 2 |
| 2018 | Point-wise saliency detection on 3D point clouds via covariance descriptors
Yu Guo 0006, Fei Wang 0008, Jingmin Xin |
Vis. Comput. | 1 |
| 2017 | A neural filter-based scheme for synchronizing chaotic systemsabstractSynchronization of chaotic systems and/or maps is a key step to implement secure communication schemes with chaos. If the process to synchronize chaotic systems is modeled stochastic, schemes based on extended Kalman filter (EKF) and unscented Kalman filter (UKF) have been studied in the past. However, such nonlinear filters are employed with assumptions of Gaussian noise processes and the Markov property. Further, EKF and UKF are suboptimal filtering methods, incurring unacceptable errors for high nonlinear systems. In this paper, neural filter (NF) is proposed for chaotic synchronization. This new approach requires no mentioned assumptions and achieves optimal filter. Numerical comparisons between the proposed approach and existing schemes are presented in this paper, showing the superiority of the proposed approach. Yu Guo 0006, Fei Wang 0008, James Ting-Ho Lo |
ICASSP | 1 |
| 2017 | Saliency-Guided Smoothing for 3D Point Clouds
Fei Wang 0008, Yu Guo 0006, Peilin Jiang |
ICIC (1) | 3 |
| 2017 | Robust echo state networks based on correntropy induced loss function
Yu Guo 0006, Fei Wang 0008, Badong Chen, Jingmin Xin |
Neurocomputing | 1 |
| 2016 | Accommodative neural filtersabstractBy the fundamental neural filtering theorem, a properly trained recursive neural filter with fixed weights that processes only the measurement process generates recursively the conditional expectation of the signal process with respect to the joint probability distributions of the signal and measurement processes and any uncertain environmental process involved. This means that a recursive neural filter with fixed weights has the ability to adapt to the uncertain environmental parameter. The neural filter with this ability is called an accommodative neural filter. In this paper, we show that if the uncertain environmental process is observable from the measurement process, the accommodative neural filter outputs virtually the estimate of the signal process that would be generated by a non-adaptive minimal-variance filter as if the precise value of the uncertain environmental process were given. Numerical results comparing the accommodative neural filter and the existing non-adaptive filters each designed for a precise value of the environmental process confirm our theorem and show the advantages of the accommodative neural filter in both accuracy and efficiency. James Ting-Ho Lo, Yu Guo 0006 |
IJCNN | 2 |