VLDB 2026 Research / reviewers in the wild / expert
Yaoxing Wang
dblp:199/1182
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0004-8227-6509ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ThermalGate-GS: Frequency-Gated Graph Splatting for Thermal Novel View SynthesisabstractThermal infrared imaging is pivotal for all-weather 3D perception, yet analyzing thermal information remains a formidable challenge due to the complexity of heat conduction. Unlike visible light, heat conduction acts as a natural low-pass filter that suppresses high-frequency textural details, causing severe geometric ambiguities and "ghosting" artifacts in standard 3D reconstruction pipelines. To accurately model the inherently diffusive thermal field for high-fidelity reconstruction, we propose Frequency-Gated Graph Splatting (ThermalGate-GS), a framework that explicitly decouples the scene into diffusive thermal distributions (low-frequency) and sharp structural boundaries (high-frequency). Within this framework, we introduce a novel Frequency-Gated Anisotropic Diffusion mechanism. Specifically, the frequency-gating module utilizes extracted high-frequency structural cues to determine spatially-adaptive gating weights. Subsequently, these weights drive an anisotropic diffusion process that dynamically regulates thermal feature propagation, promoting smoothness on object surfaces while suppressing cross-boundary bleeding. Finally, these spectrally refined features are employed to regress 3D Gaussian attributes, substantially alleviating the ambiguity in thermal reconstruction. Extensive experiments demonstrate that ThermalGate-GS achieves state-of-the-art performance, with a notable 7.94 dB PSNR improvement on the ThermoScenes benchmark over prior physics-inspired baselines. Yaoxing Wang, Wenkang Chen, Guangqian Guo, Chaowei Wang, Yan Di, Shan Gao 0003 |
IEEE Trans. Image Process. | 1 |
| 2026 | A2Net: Affiliation Alignment Networks for Whole-Body Pose Estimation With Vision-Language ModelsabstractThe whole-body pose estimation task aims to predict the location of keypoints of the face, body, hands, and feet given an image. However, scale variation in different parts of the human body and semantic ambiguity in small-scale parts cause performance degradation in keypoint localization. The traditional paradigm for solving multiscale issues is to construct multiscale feature representations. Nevertheless, multiscale features extracted from visual images do not eliminate the semantic ambiguity issue in the small-scale part. In this article, we propose affiliation alignment network (A2Net), which solves the aforementioned problem by alignment of vision-language hierarchical affiliations. Specifically, text modality has the advantage of not being affected by the scaling problem and the small-scale semantic ambiguity problem, which is due to image scale variations. We construct a multisemantic hierarchical language latent space with clear semantic and affiliation relations by designing Text Affiliation Injection operations. Subsequently, we adopt the optimal transport (OT) method to align image features of different scales with text features of the corresponding hierarchical levels to build an image scale-independent visual-language latent space, which overcomes the image scale problem and the small-scale semantic ambiguity problem. Extensive experimental results on two whole-body pose estimation datasets show that our model achieves convincing performance compared to the current state-of-the-art methods. The code is openly available at https://github.com/LingLin-ll/A2Net. Ling Lin 0002, Yaoxing Wang, Congcong Zhu, Jingrun Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Segment Any-Quality Images with Generative Latent Space EnhancementabstractDespite their success, Segment Anything Models (SAMs) experience significant performance drops on severely degraded, low-quality images, limiting their effectiveness in real-world scenarios. To address this, we propose Gle-SAM, which utilizes Generative Latent space Enhancement to boost robustness on low-quality images, thus enabling generalization across various image qualities. Specifically, we adapt the concept of latent diffusion to SAM-based segmentation frameworks and perform the generative diffusion process in the latent space of SAM to reconstruct high-quality representation, thereby improving segmentation. Additionally, we introduce two techniques to improve compatibility between the pre-trained diffusion model and the segmentation framework. Our method can be applied to pre-trained SAM and SAM2 with only minimal additional learnable parameters, allowing for efficient optimization. We also construct the LQSeg dataset with a greater diversity of degradation types and levels for training and evaluating the model. Extensive experiments demonstrate that GleSAM significantly improves segmentation robustness on complex degradations while maintaining generalization to clear images. Furthermore, GleSAM also performs well on unseen degradations, underscoring the versatility of our approach and dataset. Guangqian Guo, Xuehui Yu, Yaoxing Wang, Shan Gao 0003 |
CVPR | 5 |
| 2024 | Language-Driven Ordinal Learning for Imbalanced Head Pose EstimationabstractHead pose estimation aims to predict three degrees of freedom pose angles in an unconstrained environment. Conventional ordinal learning methods project the input in a one-dimensional label distribution, with preserving ordinal relationship among labels. However, this assumption frequently fails to hold in multi-dimensional head pose data, resulting in performance in imbalanced data despite extensive training. To address this issue, our model calibrates to sparse multi-dimensional distributions by forging a connection between the order concept in human language and the ordinal characteristics of pose labels. Our approach endeavors to utilize linguistic ordering properties to compensate for potential data scarcity in certain continuous labels. Specifically, we incorporate an ordinal pose prompt to leverage the inherent ranking relationship, and in turn, expand the pose regression boundary. We further endorse the use of real-valued encoding for label representation to subtly model prediction distributions, thereby achieving a balanced prediction of the training sample distribution. Experimental results across multiple datasets confirm that our method achieves compelling performance existing state-of-the-art techniques without auxiliary data. Yaoxing Wang |
ICASSP | 1 |
| 2024 | HeadDiff: Exploring Rotation Uncertainty With Diffusion Models for Head Pose EstimationabstractIn this paper, we propose a probabilistic regression diffusion model for head pose estimation, dubbed HeadDiff, which typically addresses the rotation uncertainty, especially when faces are captured in wild conditions. Unlike conventional image-to-pose methods which cannot explicitly establish the rotational manifold of head poses, our HeadDiff aims to ensure the pose rotation via the diffusion process and in parallel, refine the mapping process iteratively. Specifically, we initially formulate the head pose estimation problem as a reverse diffusion process, defining a paradigm for progressive denoising on the manifold, which explores the uncertainty by decomposing the large gap into intermediate steps. Moreover, our HeadDiff is equipped with an isotropic Gaussian distribution by encoding the incoherence information in our rotation representation. Finally, we learn the facial relationship of nearest neighbors with a cycle-consistent constraint for robust pose estimation versus diverse shape variations. Experimental results on multiple datasets demonstrate that our proposed method outperforms existing state-of-the-art techniques without auxiliary data. Yaoxing Wang, Hao Liu 0019, Yaowei Feng, Xiangjuan Wu, Congcong Zhu |
IEEE Trans. Image Process. | 1 |
| 2023 | Structural Equivariance Self-Supervised Learning for Facial Pose EstimationabstractIn this paper, we propose a self-supervised learning method for robust facial pose estimation. Conventional methods usually split the coherent head motion into discrete and finite outputs, likely leading to bias prediction because the performance of head pose estimation highly relies on structural facial appearance. To address this issue, our model achieves structural equivariance to poses through a self-supervised learning strategy from extrinsic attributes of face neighbors and underlying local associations. Specifically, we construct a complete neighbor graph to capture the extrinsic properties of face neighbors, where different latent semantic attributes are assigned to each subgraph. Accordingly, we design a set of proxy tasks based on different attribute subgraphs, where the model is encouraged to learn the underlying relation of local features under pose variation. Extensive experimental results on the challenging, widely evaluated datasets indicate the effectiveness of our model compared with the state of the arts. Yaoxing Wang, Xian Mo, Hao Liu 0019 |
ICME | 1 |
| 2023 | Edge-Prior Contrastive Transformer for Optic Cup and Optic Disc Segmentation
Yaowei Feng, Yaoxing Wang |
PRCV (5) | 3 |
| 2021 | PML: Progressive Margin Loss for Long-Tailed Age ClassificationabstractIn this paper, we propose a progressive margin loss (PML) approach for unconstrained facial age classification. Conventional methods make strong assumption on that each class owns adequate instances to outline its data distribution, likely leading to bias prediction where the training samples are sparse across age classes. Instead, our PML aims to adaptively refine the age label pattern by enforcing a couple of margins, which fully takes in the in-between discrepancy of the intra-class variance, inter-class variance and class center. Our PML typically incorporates with the ordinal margin and the variational margin, simultaneously plugging in the globally-tuned deep neural network paradigm. More specifically, the ordinal margin learns to exploit the correlated relationship of the real-world age labels. Accordingly, the variational margin is leveraged to minimize the influence of head classes that misleads the prediction of tailed samples. Moreover, our optimization carefully seeks a series of indicator curricula to achieve robust and efficient model training. Extensive experimental results on three face aging datasets demonstrate that our PML achieves compelling performance compared to state of the art. Code will be made publicly. Zongyong Deng, Hao Liu 0019, Yaoxing Wang, Chenyang Wang 0004, Zekuan Yu, Xuehong Sun |
CVPR | 3 |