EDBT 2026 Demo / reviewers in the wild / expert
Haoxin Yang
dblp:315/9434
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StarPose: 3D Human Pose Estimation via Spatial-Temporal Autoregressive DiffusionabstractMonocular 3D human pose estimation remains a challenging task due to inherent depth ambiguities and occlusions. Compared to traditional methods based on Transformers or Convolutional Neural Networks (CNNs), recent diffusion-based approaches have shown superior performance, leveraging their probabilistic nature and high-fidelity generation capabilities. However, these methods often fail to account for the spatial and temporal correlations across predicted frames, resulting in limited temporal consistency and inferior accuracy in predicted 3D pose sequences. To address these shortcomings, this paper proposesStarPose, an autoregressive diffusion framework that effectively incorporates historical 3D pose predictions and spatial-temporal physical guidance to significantly enhance both the accuracy and temporal coherence of pose predictions. Unlike existing approaches,StarPosemodels the 2D-to-3D pose mapping as an autoregressive diffusion process. By synergically integrating previously predicted 3D poses with 2D pose inputs via a Historical Pose Integration Module (HPIM), the framework generates rich and informative historical pose embeddings that guide subsequent denoising steps, ensuring temporally consistent predictions. In addition, a fully plug-and-play Spatial-Temporal Physical Guidance (STPG) mechanism is tailored to refine the denoising process in an iterative manner, which further enforces spatial anatomical plausibility and temporal motion dynamics, rendering robust and realistic pose estimates. Extensive experiments on benchmark datasets demonstrate thatStarPoseoutperforms state-of-the-art methods, achieving superior accuracy and temporal consistency in 3D human pose estimation. Code is available at https://github.com/wileychan/StarPose. Haoxin Yang, Xuemiao Xu, Cuifeng Sun, Shaoyu Huang, Shengfeng He |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | MPC-FH: Privately Estimating Frequency Histogram for Advertising MeasurementabstractAn advertiser typically conducts a campaign across several publishers. Different publishers often reach different sets of persons, and some people may be exposed to the campaign via multiple publishers. Due to privacy restrictions, it is prohibited to collect all audience activities from multiple publishers in a centralized manner, therefore, we desire secure protocols for mining audience activities. In this paper, we propose a secure protocol for efficiently estimating the frequency histogram, i.e. the fraction of persons (or users) appearing a given number of times across all publishers. Our protocolMPC-FHis designed based on a novel sketch methodMPSthat uniformly samples a fraction of users across multiple parties and can be efficiently implemented on SPDZ, which is a popular secure multiparty computation (MPC) scheme. Our protocol provides more strict security guarantees than the state-of-the-art. In addition, our experimental results on a variety of large datasets demonstrate that our protocol can also accelerate the computation speed of existing protocols by orders of magnitudes. Pinghui Wang, Haoxin Yang, Dongxu Zeng, Hong Xie 0004, Junzhou Zhao, Xiaohong Guan |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | ViewSRD: 3D Visual Grounding Via Structured Multi-View Decomposition
Ronggang Huang, Haoxin Yang, Yan Cai 0021, Xuemiao Xu, Huaidong Zhang, Shengfeng He |
ICCV | 2 |
| 2025 | SCJD: Sparse Correlation and Joint Distillation for Efficient 3D Human Pose EstimationabstractExisting 3D Human Pose Estimation (HPE) methods achieve high accuracy but suffer from computational overhead and slow inference, while knowledge distillation methods fail to address spatial relationships between joints and temporal correlations in multi-frame inputs. In this paper, we propose Sparse Correlation and Joint Distillation (SCJD), a novel framework that balances efficiency and accuracy for 3D HPE. SCJD introduces Sparse Correlation Input Sequence Downsampling to reduce redundancy in student network inputs while preserving inter-frame correlations. For effective knowledge transfer, we propose Dynamic Joint Spatial Attention Distillation, which includes Dynamic Joint Embedding Distillation to enhance the student’s feature representation using the teacher’s multi-frame context feature, and Adjacent Joint Attention Distillation to improve the student network’s focus on adjacent joint relationships for better spatial understanding. Additionally, Temporal Consistency Distillation aligns the temporal correlations between teacher and student networks through upsampling and global supervision. Extensive experiments demonstrate that SCJD achieves state-of-the-art performance. Code is available at https://github.com/wileychan/SCJD. Xuemiao Xu, Haoxin Yang, Huaidong Zhang, Pheng-Ann Heng |
ICME | 3 |
| 2025 | StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion ModelsabstractThe advancement of diffusion models has enhanced the realism of AI-generated content but also raised concerns about misuse, necessitating robust copyright protection and tampering localization. Although recent methods have made progress toward unified solutions, their reliance on post hoc processing introduces considerable application inconvenience and compromises forensic reliability. We propose StableGuard, a novel framework that seamlessly integrates a binary watermark into the diffusion generation process, ensuring copyright protection and tampering localization in Latent Diffusion Models through an end-to-end design. We develop a Multiplexing Watermark VAE (MPW-VAE) by equipping a pretrained Variational Autoencoder (VAE) with a lightweight latent residual-based adapter, enabling the generation of paired watermarked and watermark-free images. These pairs, fused via random masks, create a diverse dataset for training a tampering-agnostic forensic network. To further enhance forensic synergy, we introduce a Mixture-of-Experts Guided Forensic Network (MoE-GFN) that dynamically integrates holistic watermark patterns, local tampering traces, and frequency-domain cues for precise watermark verification and tampered region detection. The MPW-VAE and MoE-GFN are jointly optimized in a self-supervised, end-to-end manner, fostering a reciprocal training between watermark embedding and forensic accuracy. Extensive experiments demonstrate that StableGuard consistently outperforms state-of-the-art methods in image fidelity, watermark verification, and tampering localization. Haoxin Yang, Bangzhen Liu, Xuemiao Xu, Yuyang Yu, Zikai Huang, Shengfeng He |
NeurIPS | 1 |
| 2025 | Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detectionabstract3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based methods often suffer from inconsistent feature transformations and limited discriminative capacity, particularly in capturing local geometric details and achieving rotation invariance. These limitations become more pronounced when registration fails, leading to unreliable detection results. We argue that point-cloud registration plays an essential role not only in aligning geometric structures but also in guiding feature extraction toward rotation-invariant and locally discriminative representations. To this end, we propose a registration-induced, rotation-invariant feature extraction framework that integrates the objectives of point-cloud registration and memory-based anomaly detection. Our key insight is that both tasks rely on modeling local geometric structures and leveraging feature similarity across samples. By embedding feature extraction into the registration learning process, our framework jointly optimizes alignment and representation learning. This integration enables the network to acquire features that are both robust to rotations and highly effective for anomaly detection. Extensive experiments on the Anomaly-ShapeNet and Real3D-AD datasets demonstrate that our method consistently outperforms existing approaches in effectiveness and generalizability. Yuyang Yu, Zhengwei Chen, Xuemiao Xu, Lei Zhang 0006, Haoxin Yang, Yongwei Nie, Shengfeng He |
NeurIPS | 5 |
| 2025 | SITA: Structurally Imperceptible and Transferable Adversarial Attacks for Stylized Image Generation
Jingdan Kang, Haoxin Yang, Yan Cai 0021, Huaidong Zhang, Xuemiao Xu, Yong Du 0003, Shengfeng He |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | G²Face: High-Fidelity Reversible Face Anonymization via Generative and Geometric PriorsabstractReversible face anonymization, unlike traditional face pixelization, seeks to replace sensitive identity information in facial images with synthesized alternatives, preserving privacy without sacrificing image clarity. Traditional methods, such as encoder-decoder networks, often result in significant loss of facial details due to their limited learning capacity. Additionally, relying on latent manipulation in pre-trained GANs can lead to changes in ID-irrelevant attributes, adversely affecting data utility due to GAN inversion inaccuracies. This paper introduces G2Face, which leverages both generative and geometric priors to enhance identity manipulation, achieving high-quality reversible face anonymization without compromising data utility. We utilize a 3D face model to extract geometric information from the input face, integrating it with a pre-trained GAN-based decoder. This synergy of generative and geometric priors allows the decoder to produce realistic anonymized faces with consistent geometry. Moreover, multi-scale facial features are extracted from the original face and combined with the decoder using our novel identity-aware feature fusion blocks (IFF). This integration enables precise blending of the generated facial patterns with the original ID-irrelevant features, resulting in accurate identity manipulation. Extensive experiments demonstrate that our method outperforms existing state-of-the-art techniques in face anonymization and recovery, while preserving high data utility. Code is available athttps://github.com/Harxis/G2Face. Haoxin Yang, Xuemiao Xu, Huaidong Zhang, Harry Qin, Yi Wang 0017, Pheng-Ann Heng, Shengfeng He |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Individual Property Inference Over Collaborative Learning in Deep Feature SpaceabstractCollaborative learning is used in multi-media applications to distribute computing tasks and data storage over multiple sites. Recent studies found that private data information can be derived from model updates between the server and clients. Yet, previous methods are limited by their capabilities of privacy inference in more general and practical situations. In this paper, we propose a novel property inference method in the deep feature space to overcome those limitations. In particular, our method can make inference decisions on the level of individual examples instead of a batch of examples. We can simultaneously perform multiple property inference attacks without the need of image reconstruction. The proposed method is evaluated on several image benchmark datasets, which demonstrates significant improvement of inference accuracy even in the presence of privacy protection schemes. Haoxin Yang, Yi Wang 0017, Bin Li 0011 |
ICME | 1 |