VLDB 2026 Research / reviewers in the wild / expert
Hao Ai
dblp:286/5383
· DBLP profile ↗
14ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RRAM based CIM, PUF and True Random Number Generator for High-Security AES Encryption System
Kefan Tao, Shiyue Song, Yading Yi, Ao Shi, Haokai Guan, Lianliang Wu, Hao Ai, Lifeng Liu, Yulin Feng, Peng Huang 0004 |
ISCAS | 7 |
| 2026 | RRAM-based CAM for Energy-Efficient In-Memory Text Compression System
Lianliang Wu, Hao Ai, Ao Shi, Haokai Guan, Kexun Li, Kefan Tao, Yulin Feng, Zongwei Wang 0001, Yimao Cai, Peng Huang 0004 |
ISCAS | 2 |
| 2026 | Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic SegmentationabstractBenefiting from the inductive biases learned from large-scale datasets, open-vocabulary semantic segmentation (OVSS) leverages the power of vision-language models, such as CLIP, to achieve remarkable progress without requiring task-specific training. However, due to CLIP’s pre-training nature on image-text pairs, it tends to focus on global semantic alignment, resulting in suboptimal performance when associating fine-grained visual regions with text. This leads to noisy and inconsistent predictions, particularly in local areas. We attribute this to a dispersed bias stemming from its contrastive training paradigm, which is difficult to alleviate using CLIP features alone. To address this, we propose a structure-aware feature rectification approach that incorporates instance-specific priors derived directly from the image. Specifically, we construct a region adjacency graph (RAG) based on low-level features (e.g. colour and texture) to capture local structural relationships and use it to refine CLIP features by enhancing local discrimination. Extensive experiments show that our method effectively suppresses segmentation noise, improves region-level consistency, and achieves strong performance on multiple open-vocabulary segmentation benchmarks. Project page: https://qiming-huang.github.io/RAG-OVS/. Qiming Huang, Hao Ai, Jianbo Jiao |
WACV | 2 |
| 2026 | HeatSim: A Highly Efficient Analytical Transient Thermal Simulator With Explicit Error Bound
Hao Ai, Liang Chen 0025, Wenxing Zhu |
IEEE Trans. Computers | 1 |
| 2026 | Fast Steady-State Thermal Analysis With Separation of Variables and Discrete Cosine Transform
Hao Ai, Liang Chen 0025, Bei Yu 0001, Wenxing Zhu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | PanDA: Towards Panoramic Depth Anything with Unlabeled Panoramas and Mobius Spatial AugmentationabstractRecently, Depth Anything Models (DAMs) [47], [48] - a type of depth foundation models – have demonstrated impressive zero-shot capabilities across diverse perspective images. Despite its success, it remains an open question regarding DAMs’ performance on panorama images that enjoy a large field-of-view (180° × 360°) but suffer from spherical distortions. To address this gap, we conduct an empirical analysis to evaluate the performance of DAMs on panoramic images and identify their limitations. For this, we undertake comprehensive experiments to assess the performance of DAMs from three key factors: panoramic representations, 360° camera positions for capturing scenarios, and spherical spatial transformations. This way, we reveal some key findings, e.g., DAMs are sensitive to spatial transformations. We then propose a semi-supervised learning (SSL) framework to learn a panoramic DAM, dubbed PanDA. Under the umbrella of SSL, PanDA first learns a teacher model by fine-tuning DAM through joint training on synthetic indoor and outdoor panoramic datasets. Then, a student model is trained using large-scale unlabeled data, leveraging pseudo-labels generated by the teacher model. To enhance PanDA’s generalization capability, Möbius transformation-based spatial augmentation (MTSA) is proposed to impose consistency regularization between the predicted depth maps from the original and spatially transformed ones. This subtly improves the student model’s robustness to various spatial transformations, even under severe distortions. Extensive experiments demonstrate that PanDA exhibits remarkable zero-shot capability across diverse scenes, and outperforms the data-specific panoramic depth estimation methods on two popular real-world benchmarks. Project page: https://caozidong.github.io/PanDA_Depth/. Zidong Cao, Jinjing Zhu, Weiming Zhang 0001, Hao Ai, Haotian Bai, Hengshuang Zhao, Lin Wang 0025 |
CVPR | 4 |
| 2025 | ST$^2$360D: Spatial-to-Temporal Consistency for Training-free 360 Monocular Depth Estimationabstract360-degree monocular depth estimation plays a crucial role in scene understanding owing to its 180-degree by 360-degree field-of-view (FoV). To mitigate the distortions brought by equirectangular projection, existing methods typically divide 360-degree images into distortion-less perspective patches. However, since these patches are processed independently, depth inconsistencies are often introduced due to scale drift among patches. Recently, video depth estimation (VDE) models have leveraged temporal consistency for stable depth predictions across frames. Inspired by this, we propose to represent a 360-degree image as a sequence of perspective frames, mimicking the viewpoint adjustments users make when exploring a 360-degree scenario in virtual reality. Thus, the spatial consistency among perspective depth patches can be enhanced by exploiting the temporal consistency inherent in VDE models. To this end, we introduce a training-free pipeline for 360-degree monocular depth estimation, called ST²360D. Specifically, ST²360D transforms a 360-degree image into perspective video frames, predicts video depth maps using VDE models, and seamlessly merges these predictions into a complete 360-degree depth map. To generate sequenced perspective frames that align with VDE models, we propose two tailored strategies. First, a spherical-uniform sampling (SUS) strategy is proposed to facilitate uniform sampling of perspective views across the sphere, avoiding oversampling in polar regions typically with limited structural details. Second, a latitude-guided scanning (LGS) strategy is introduced to organize the frames into a coherent sequence, starting from the equator, prioritizing low-latitude slices, and progressively moving toward higher latitudes. Extensive experiments demonstrate that ST²360D achieves strong zero-shot capability on several datasets, supporting resolutions up to 4K. Zidong Cao, Jinjing Zhu, Hao Ai, Lutao Jiang, Yuanhuiyi Lyu, Hui Xiong 0001 |
NeurIPS | 3 |
| 2025 | CSGO: Content-Style Composition in Text-to-Image GenerationabstractThe advancement of image style transfer has been fundamentally constrained by the absence of large-scale, high-quality datasets with explicit content-style-stylized supervision. Existing methods predominantly adopt training-free paradigms (e.g., image inversion), which limit controllability and generalization due to the lack of structured triplet data. To bridge this gap, we design a scalable and automated pipeline that constructs and purifies high-fidelity content-style-stylized image triplets. Leveraging this pipeline, we introduce IMAGStyle—the first large-scale dataset of its kind, containing 210K diverse and precisely aligned triplets for style transfer research. Empowered by IMAGStyle, we propose CSGO, a unified, end-to-end trainable framework that decouples content and style representations via independent feature injection. CSGO jointly supports image-driven style transfer, text-driven stylized generation, and text-editing-driven stylized synthesis within a single architecture. Extensive experiments show that CSGO achieves state-of-the-art controllability and fidelity, demonstrating the critical role of structured synthetic data in unlocking robust and generalizable style transfer. Source code: \url{https://github.com/instantX-research/CSGO} Peng Xing, Yanpeng Sun, Qixun Wang 0001, Hao Ai, Jen-Yuan Huang, Zechao Li |
NeurIPS | 6 |
| 2025 | A Survey of Representation Learning, Optimization Strategies, and Applications for Omnidirectional Vision
Hao Ai, Zidong Cao |
Int. J. Comput. Vis. | 1 |
| 2024 | Elite360D: Towards Efficient 360 Depth Estimation via Semantic- and Distance-Aware Bi-Projection Fusionabstract360 depth estimation has recently received great attention for 3D reconstruction owing to its omnidirectional field of view (FoV). Recent approaches are predominantly focused on cross-projection fusion with geometry-based reprojection: they fuse 360 images with equirectangular projection (ERP) and another projection type, e.g., cubemap projection to estimate depth with the ERP format. However, these methods suffer from 1) limited local receptive fields, making it hardly possible to capture large FoV scenes, and 2) prohibitive computational cost, caused by the complex cross-projection fusion module design. In this paper, we propose Elite360D, a novel framework that inputs the ERP image and icosahedron projection (ICOSAP) point set, which is undistorted and spatially continuous. Elite360D is superior in its capacity in learning a representation from a local-with-global perspective. With a flexible ERP image encoder, it includes an ICOSAP point encoder, and a Biprojection Bi-attention Fusion (B2F) module (totally ~1M parameters). Specifically, the ERP image encoder can take various perspective image-trained backbones (e.g., ResNet, Transformer) to extract local features. The point encoder extracts the global features from the ICOSAP. Then, the B2F module captures the semantic- and distance-aware dependencies between each pixel of the ERP feature and the entire ICOSAP feature set. Without specific backbone design and obvious computational cost increase, Elite360D outperforms the prior arts on several benchmark datasets. Hao Ai, Lin Wang 0025 |
CVPR | 1 |
| 2024 | Dream360: Diverse and Immersive Outdoor Virtual Scene Creation via Transformer-Based 360° Image Outpaintingabstract360° images, with a field-of-view (FoV) of $180^{\circ}\times 360^{\circ}$, provide immersive and realistic environments for emerging virtual reality (VR) applications, such as virtual tourism, where users desire to create diverse panoramic scenes from a narrow FoV photo they take from a viewpoint via portable devices. It thus brings us to a technical challenge: 'How to allow the users to freely create diverse and immersive virtual scenes from a narrow FoV image with a specified viewport?' To this end, we propose a transformer-based 360° image outpainting framework called Dream360, which can generate diverse, high-fidelity, and high-resolution panoramas from user-selected viewports, considering the spherical properties of 360° images. Compared with existing methods, e.g., [3], which primarily focus on inputs with rectangular masks and central locations while overlooking the spherical property of 360° images, our Dream360 offers higher outpainting flexibility and fidelity based on the spherical representation. Dream360 comprises two key learning stages: (I) codebook-based panorama outpainting via Spherical-VQGAN (S-VQGAN), and (II) frequency-aware refinement with a novel frequency-aware consistency loss. Specifically, S-VQGAN learns a sphere-specific codebook from spherical harmonic (SH) values, providing a better representation of spherical data distribution for scene modeling. The frequency-aware refinement matches the resolution and further improves the semantic consistency and visual fidelity of the generated results. Our Dream360 achieves significantly lower Frechet Inception Distance (FID) scores and better visual fidelity than existing methods. We also conducted a user study involving 15 participants to interactively evaluate the quality of the generated results in VR, demonstrating the flexibility and superiority of our Dream360 framework. Hao Ai, Zidong Cao, Haonan Lu, Chen Chen 0015, Jian Ma 0010, Peng Yuan Zhou, Tae-Kyun Kim 0001, Pan Hui 0001, Lin Wang 0025 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | HRDFuse: Monocular 360° Depth Estimation by Collaboratively Learning Holistic-with-Regional Depth DistributionsabstractDepth estimation from a monocular 360° image is a burgeoning problem owing to its holistic sensing of a scene. Recently, some methods, e.g., OmniFusion, have applied the tangent projection (TP) to represent a 360° image and predicted depth values via patch-wise regressions, which are merged to get a depth map with equirectangular projection (ERP) format. However, these methods suffer from 1) non-trivial process of merging plenty of patches; 2) capturing less holistic-with-regional contextual information by directly regressing the depth value of each pixel. In this paper, we propose a novel framework, HRDFuse, that subtly combines the potential of convolutional neural networks (CNNs) and transformers by collaboratively learning the holistic contextual information from the ERP and the regional structural information from the TP. Firstly, we propose a spatial feature alignment (SFA) module that learns feature similarities between the TP and ERP to aggregate the TP features into a complete ERP feature map in a pixelwise manner. Secondly, we propose a collaborative depth distribution classification (CDDC) module that learns the holistic-with-regional histograms capturing the ERP and TP depth distributions. As such, the final depth values can be predicted as a linear combination of histogram bin centers. Lastly, we adaptively combine the depth predictions from ERP and TP to obtain the final depth map. Extensive experiments show that our method predicts more smooth and accurate depth results while achieving favorably better results than the SOTA methods. Hao Ai, Zidong Cao, Yan-Pei Cao 0001, Ying Shan, Lin Wang 0025 |
CVPR | 1 |
| 2023 | OmniZoomer: Learning to Move and Zoom in on Sphere at High-ResolutionabstractOmnidirectional images (ODIs) have become increasingly popular, as their large field-of-view (FoV) can offer viewers the chance to freely choose the view directions in immersive environments such as virtual reality. The Möbius transformation is typically employed to further provide the opportunity for movement and zoom on ODIs, but applying it to the image level often results in blurry effect and aliasing problem. In this paper, we propose a novel deep learning-based approach, called OmniZoomer, to incorporate the Möbius transformation into the network for movement and zoom on ODIs. By learning various transformed feature maps under different conditions, the network is enhanced to handle the increasing edge curvatures, which alleviates the blurry effect. Moreover, to address the aliasing problem, we propose two key components. Firstly, to compensate for the lack of pixels for describing curves, we enhance the feature maps in the high-resolution (HR) space and calculate the transformed index map with a spatial index generation module. Secondly, considering that ODIs are inherently represented in the spherical space, we propose a spherical resampling module that combines the index map and HR feature maps to transform the feature maps for better spherical correlation. The transformed feature maps are decoded to output a zoomed ODI. Experiments show that our method can produce HR and high-quality ODIs with the flexibility to move and zoom in to the object of interest. Project page is available at http: //vlislab22.github.io/OmniZoomer/. Zidong Cao, Hao Ai, Yan-Pei Cao 0001, Ying Shan, Xiaohu Qie, Lin Wang 0025 |
ICCV | 2 |
| 2021 | Gaussian Mixture Distribution Makes Data Uncertainty Learning BetterabstractAs a mainstream method in face recognition, extracting separable facial features by deep CNNs in the latent space has achieved remarkable success. In most existing works, people often view the embedding features as points. Dealing with entirely unconstrained face images, DUL and PFE demonstrated that point estimation shows weak robusticity on the inherent noise in the input images (data uncertainty) and introduced the distribution estimation by modeling each latent feature using a Gaussian distribution. However, these two methods only apply a unimodal Gaussian prior distribution, which is insufficient to represent wild faces with complex variations. In this paper, we propose a novel face recognition framework based on the multivariate Gaussian mixture distribution (DUL-GM). Through numerous experiments, we show that compared with the prior works, the features modeled by multivariate Gaussian mixture distribution have a better interference suppression ability and achieve state-of-the-art performance on extensive challenging benchmarks. Hao Ai, Qingmin Liao |
FG | 1 |