EDBT 2026 Demo / reviewers in the wild / expert
Haifeng Zhao 0001
dblp:17/5245-1
· DBLP profile ↗
39ranked-venue papers
16as first author
30since 2021 · last 2026
0000-0002-5300-0683ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 14 · 7 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | APMVS: Learning Multi-View Stereo Based on Adjacent Stage and Pair-Wise Stage Uncertainty EstimationabstractMany multi-view stereo (MVS) networks with a cascaded structure can effectively estimate depth while saving memory. However, the accuracy of the depth map in the fine stage depends on the depth map estimated in the coarse stage. Additionally, the multi-stage depth maps generated by the cascaded structure are used to compute losses but are not reused, resulting in a loss of inter-stage differentiation information. To address these issues, we propose a dual-uncertainty estimation MVS method that learns an MVS network based on adjacent stage and pair-wise stage uncertainty estimation, named APMVS. The core of the proposed APMVS is to employ dual-uncertainty estimation to mitigate the adverse effects of the cascaded structure. Specifically, it involves two estimation modules: adjacent stage uncertainty (ASU) and pair-wise stage uncertainty (PSU). The ASU estimation module dynamically adjusts the depth-hypothesis range by leveraging uncertainty from the previous stage, thereby improving the accuracy of depth-map prediction in the current stage. The PSU estimation module estimates the uncertainty between each pair of stages. Thus, regions with high uncertainty have minimal impact. We evaluate the proposed APMVS on the DTU, Tanks and Temples, and BlendedMVS datasets. Experimental results show that our method achieves superior reconstruction quality compared with other state-of-the-art methods. Mingwei Cao, Siqi Nian, Haifeng Zhao 0001, Feng Xue 0002, Zhihan Lyu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | DOMVS: Unsupervised Multi-view Stereo for Dealing With Occlusion Scenes
Mingwei Cao, Qiuju Wang, Haifeng Zhao 0001 |
CGI (2) | 4 |
| 2025 | Texture and Geometry Optimization for 3D Reconstruction
Yanping Fu, Hongjing Zhang, Shaojie Zhang 0002, Dengdi Sun, Haifeng Zhao 0001 |
CGI (1) | 5 |
| 2025 | Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt TuningabstractModeling label correlations has always played a pivotal role in multi-label image classification (MLC), attracting significant attention from researchers. However, recent studies have overemphasized co-occurrence relationships among labels, which can lead to overfitting risk on this overemphasis, resulting in suboptimal models. To tackle this problem, we advocate for balancing correlative and discriminative relationships among labels to mitigate the risk of overfitting and enhance model performance. To this end, we propose the Multi-Label Visual Prompt Tuning framework, a novel and parameter-efficient method that groups classes into multiple class subsets according to label co-occurrence and mutual exclusivity relationships, and then models them respectively to balance the two relationships. In this work, since each group contains multiple classes, multiple prompt tokens are adopted within Vision Transformer (ViT) to capture the correlation or discriminative label relationship within each group, and effectively learn correlation or discriminative representations for class subsets. On the other hand, each group contains multiple group-aware visual representations that may correspond to multiple classes, and the mixture of experts (MoE) model can cleverly assign them from the group-aware to the label-aware, adaptively obtaining label-aware representation, which is more conducive to classification. Experiments on multiple benchmark datasets show that our proposed approach achieves competitive results and outperforms SOTA methods on multiple pre-trained models. Leilei Ma 0002, Ming-Kun Xie, Lei Wang 0095, Dengdi Sun, Haifeng Zhao 0001 |
CVPR | 6 |
| 2025 | StarGS: Towards Real-Time Dynamic Scene Rendering via Gaussian Splatting with Spatiotemporal-Aware Density ControlabstractThe 3D Gaussian Splatting (3DGS) technique has yielded significant achievements in scene rendering. However, it encounters difficulties when representing dynamic scenes marked by complexity, high degrees of freedom, and limited viewpoints. To overcome these challenges, we introduce StarGS, an innovative method for real-time dynamic scene rendering that employs spatiotemporal-aware density control in Gaussian splatting. StarGS features an adaptive density mechanism that dynamically modifies the Gaussian distribution based on spatiotemporal characteristics, thereby precisely capturing structural variations and intricate texture details in areas with complex motion. Firstly, we devise a motion-aware keypoints mechanism, guided by spatio-temporal features, which isolates notable moving objects within the scene and promotes local density augmentation around them, enhancing their structural representations. Secondly, we incorporate a pixel-wise error-guided pruning module to eliminate low-contribution Gaussians and noise, effectively minimizing computational redundancy and enhancing rendering efficiency. Thirdly, we suggest a structural feature-anchor rigidity constraint that bolsters the local consistency of Gaussians, reducing motion artifacts in dynamic scenes. We assess StarGS using the HyperNeRF and Neu3D datasets and compare it with state-of-the-art methods. Experimental results reveal that StarGS achieves 86 FPS at a$1352 \times 1014$resolution, while outperforming state-of-the-art methods in image quality. Source code available at: https://github.com/caomw/stargs Mingwei Cao, Haifeng Zhao 0001 |
CW | 3 |
| 2025 | Sparse-View X-ray 3D Reconstruction using Hybrid Representation Neural Attenuation FieldsabstractX-ray 3D reconstruction has achieved superior performance in medical imaging with traditional and deep learning methods. However, when sparse-view X-ray projections are used to minimize patient exposure to radiation, these methods tend to overfit and produce blurring. To overcome this problem, we propose a novel hybrid feature representation neural attenuation field framework for sparse-view X-ray 3D reconstruction. First, we integrate tri-plane features with hash coding features as the network input, enabling the capture of intricate local details and high-frequency information. Second, we enhance the modeling of radiation attenuation across different organs by designing a specialized attenuation weight estimation network. This network enables the attenuation field estimation network to more accurately focus on the varying attenuation rates of different tissues. Third, we introduce a new multi-skip strategy, which uses the skip connection strategy for each layer of MLPs to the attenuation value and weight prediction network, markedly enhancing the performance of our method. Experiments on public datasets demonstrate the superiority of our proposed method over state-of-the-art methods. Yanping Fu, Hao Geng, Zhuangzhuang Zhao, Shaojie Zhang 0002, Haifeng Zhao 0001 |
ICASSP | 5 |
| 2025 | Single-View Clothed Human Reconstruction using Symmetric FeatureabstractSingle-view clothed human reconstruction has emerged as a prominent research focus, yet reconstructing the occluded backside of the human body remains a significant challenge due to viewpoint limitations and occlusion. In this paper, we propose a novel single-view clothed human reconstruction method that leverages non-rigid symmetric features guidance to effectively mitigate these challenges. Firstly, we utilize the natural symmetry of the human body to construct the non-rigid symmetry through the parametric human body model SMPL-X. Secondly, these symmetric-driven features are then used to enhance the pixel-aligned features, enabling a more precise and complete reconstruction. Finally, we introduce an innovative occupancy query selector that intelligently identifies and prioritizes the most relevant features during the query stage to further improve the quality of reconstruction. Experimental results demonstrate that the proposed method can achieve better reconstruction results, especially in occluded areas. Yanping Fu, Zhuangzhuang Zhao, Hao Geng, Haifeng Zhao 0001 |
ICASSP | 4 |
| 2025 | PFGM-IQA: CT Image Quality Assessment using Poisson Flow Generative ModelsabstractImage Quality Assessment (IQA) is essential for optimizing radiation dose in Computed Tomography (CT) while ensuring radiologists achieve the highest diagnostic accuracy. To advance research in this field, we propose a novel algorithm for CT image quality assessment without the need for reference images. To overcome the challenge of no-reference IQA in CT scans, we employ a novel approach using Poisson Flow Generative models (PFGM) to generate pseudo-reference images for low-dose CT scans. These pseudo-reference images, paired with corresponding low-dose inputs, are fed into a robust regression network specifically designed for this task. To enhance feature extraction, we design a nested convolutional architecture using multi-scale features to improve the feature extraction capability of the regression Network. Furthermore, to strengthen the monotonic correlation between subjective and objective scores, we incorporate the relative distance information within each batch and enforce the relative ranking among the images. To propel research in this domain, we also build and release a new abdominal CT dataset with labeled IQA scores, tailored for benchmarking IQA methods. Extensive qualitative and quantitative experiments, conducted on both public datasets and our newly released dataset, demonstrate that our proposed PFGM-IQA method significantly outperforms state-of-the-art techniques in CT Image Quality Assessment. Haifeng Zhao 0001, Tianxia Yang, Shaojie Zhang 0002, Yanping Fu |
IJCNN | 1 |
| 2025 | Towards Space and Semantics: Object-Purified Representation Learning for Multi-Label Image ClassificationabstractMulti-label image classification requires simultaneously recognizing multiple objects with complex interdependencies. While existing attention-based methods are prominent, their performance is hampered by two forms of representation entanglement: 1) Spatial entanglement, where contextual interference from backgrounds and co-occurring objects confuses specific object representations; 2) Semantic entanglement, where models overfit label co-occurrence priors, thereby impairing a genuine semantic understanding of the image. To address these challenges, we propose an Object-Purified Representation Learning framework. Concretely, for spatial entanglement, we propose the Spatial-wise Representation Purification Module that employs Spatial-Purified Attention to eliminate object-irrelevant feature activations for contextual interference reduction, combined with Spatial-Aware Supervision to enhance object perception capability. For semantic entanglement, we develop the Semantic-wise Association Purification Module that synergistically integrates our proposed average message with the original co-occurrence-based message. This design effectively models co-occurrence relationships while preventing their overemphasis. Furthermore, we design the Bidirectional Representation Refinement Module to efficiently enhance representations, further boosting classification performance. Extensive experiments on multiple benchmark datasets with different configurations demonstrate that our proposed method achieves state-of-the-art performance. Haifeng Zhao 0001, Leilei Ma 0002, Lei Wang 0095, Dengdi Sun |
ACM Multimedia | 1 |
| 2025 | Segment Anything Model Meets Semi-supervised Medical Image Segmentation: A Novel PerspectiveabstractThe scarcity of annotated medical imaging data has driven significant progress in semi-supervised learning to alleviate reliance on expensive expert labeling. While foundational vision models such as the Segment Anything Model (SAM) exhibit robust generalization in generic segmentation tasks, their direct application to medical images often results in suboptimal performance. To address this challenge, in this work, we propose a novel fully SAM-based semi-supervised medical image segmentation framework and develop the corresponding knowledge distillation-based learning strategy. Specifically, we first employ an efficient SAM variant as the backbone network of the semi‑supervised framework and update the default prompt embedding of SAM to unleash its full potential. Then, we utilize an original SAM, which is rich in prior knowledge, as the teacher to optimize our efficient student SAM backbone through hierarchical knowledge distillation and a dynamic loss weighting strategy. Extensive experiments on various medical datasets demonstrate that our method outperforms state-of-the-art semi-supervised segmentation approaches. Especially, our model requires less than 10% of the parameter size of the original SAM, enabling substantially lower deployment and storage overhead in real-world clinical settings. Haifeng Zhao 0001, Leilei Ma 0002, Dengdi Sun |
NeurIPS | 1 |
| 2025 | Fully Automated SAM for Single-source Domain Generalization in Medical Image Segmentation
Huanli Zhuo, Leilei Ma 0002, Haifeng Zhao 0001, Dengdi Sun, Yanping Fu |
SMC | 3 |
| 2025 | Semantic knowledge transfer for semi-supervised medical image segmentation
Haifeng Zhao 0001, Leilei Ma 0002, Dengdi Sun |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Dual-level semantic alignment for video moment retrieval and highlight detection
Haifeng Zhao 0001, Wenhai Qin, Leilei Ma 0002, Dengdi Sun |
Multim. Syst. | 1 |
| 2025 | Deformation Field Fusion for Medical Image RegistrationabstractDeformable medical image registration is to find a series of non-linear spatial transformations to align a pair of fixed and moving voxel images. Deep learning based registration models are effective in learning differences between such image pair to obtain the deformation field which is specialized in describing non-rigid deformations in the 3D voxel context. However, existing models tend to learn either one single deformation field only or multi-stage (multi-level) deformation fields progressively arriving at a final optimal field. Actually, deformation fields resulting from different architectures or losses are capable of capturing diverse types of deformations, complementing to each other. In this article, we propose a novel framework of fusing different deformation fields to acquire an overall field to describe all-round deformations, in which multiple complementary cues regarding deformable 3D voxels can be strategically leveraged to improve the alignment of the given image pair. The key to the effect of deformation field fusion for registration lies in two aspects: the fusion network architecture and the loss function. Thus, we develop a well-designed fusion block using ingenious operations based on different types of pooling, convolution, and concatenation. Moreover, since calculating the deformation field using a conventional similarity loss cannot describe the contextual variations which are inter-dependent in each pair of fixed and moving images, we propose a novel Contrast-Structural loss to enhance the motion displacement between the image pair by calculating the similarity of pixels in density values, while being ranged in their spatial proximity. Extensive experimental results demonstrate that our proposed method achieves state-of-the-art performance on currently mainstream benchmark datasets. Haifeng Zhao 0001, Chi Zhang 0082, Deyin Liu, Lin Wu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Single image shadow removal using 2D signed distance field
Yanping Fu, Dengdi Sun, Shaojie Zhang 0002, Haifeng Zhao 0001 |
Vis. Comput. | 5 |
| 2024 | Towards Efficient Sparse Transformer based Medical Image RegistrationabstractDeformable medical image registration is a crucial task that involves extracting and aligning features from two images to establish precise correspondence, essentially for accurate registration. While visual transformers have propelled recent advancements in medical image analysis, training and inference with Transformers can become excessively computationally expensive, particularly due to the quadratic complexity of self-attention when handling long sequences of representations. This challenge becomes more pronounced in 3D medical image registration tasks. To tackle this issue, we propose an efficient Hierarchical Pyramid Converter for medical image registration. The proposed approach firstly capitalizes on the observation that early self-attention layers in Transformers mainly emphasize local patterns, though with limited benefits. Specifically, we employ the plain multi-layer perceptrons (MLP), i.e., Spatial shift MLP (S-MLP), in the early stages of feature extraction. This module employs a spatial offset operation to facilitate communication between patches, encoding rich local patterns and effectively reducing computational expenses. We further propose a sparse Transformer block that adaptively selects and preserves the most valuable self-attention values for feature extraction. We introduce a learnable top-k selection operator, allowing the model to selectively retain attention scores that contribute the most to each query keyword. This innovation significantly enhances feature extraction in later stages. We conducted extensive evaluations using publicly available datasets, and the experimental results confirm that our proposed method achieves state-of-the-art performance in deformable medical image registration tasks. Haifeng Zhao 0001, Quanshuang He, Deyin Liu |
CSCWD | 1 |
| 2024 | Single Image Reflection removal Using Feature Difference EnhancementabstractMost existing reflection removal methods pay too much attention to the transmission layer and ignore the mutual complementary mechanisms between the transmission layer and the reflection layer. To make full use of the complementarity and distinction between the reflection and transmission layers in reflection-contaminated images, we propose a novel single image reflection removal framework using the feature difference and adaptive information exchange between the transmission and reflection layers of reflection-contaminated images. First, we design a feature difference enhancement module to distinguish and enhance the feature difference of the reflective and transmissive layers. Second, we propose an adaptive information exchange module between transmission and reflection layers in the decoder, which can capture more complementary information. Finally, we introduce a 1/4 selective Instance Normalization strategy to improve our reflection removal tasks further. The experimental results demonstrate the efficiency of the proposed method and superior performance against state-of-the-art methods. Haifeng Zhao 0001, Shaojie Zhang 0002, Yanping Fu |
ICASSP | 1 |
| 2024 | Text-Region Matching for Multi-Label Image Recognition with Missing LabelsabstractRecently, large-scale visual language pre-trained (VLP) models have demonstrated impressive performance across various downstream tasks. Motivated by these advancements, pioneering efforts have emerged in multi-label image recognition with missing labels, leveraging VLP prompt-tuning technology. However, they usually cannot match text and vision features well, due to complicated semantics gaps and missing labels in a multi-label image. To tackle this challenge, we propose Text-Region Matching for optimizing Multi-Label prompt tuning, namely TRM-ML, a novel method for enhancing meaningful cross-modal matching. Compared to existing methods, we advocate exploring the information of category-aware regions rather than the entire image or pixels, which contributes to bridging the semantic gap between textual and visual representations in a one-to-one matching manner. Concurrently, we further introduce multimodal contrastive learning to narrow the semantic gap between textual and visual modalities and establish intra-class and inter-class relationships. Additionally, to deal with missing labels, we propose a multimodal category prototype that leverages intra- and inter-category semantic relationships to estimate unknown labels, facilitating pseudo-label generation. Extensive experiments on the MS-COCO, PASCAL VOC, Visual Genome, NUS-WIDE, and CUB-200-211 benchmark datasets demonstrate that our proposed framework outperforms the state-of-the-art methods by a significant margin. Our code is available here. Leilei Ma 0002, Hongxing Xie, Lei Wang 0095, Yanping Fu, Dengdi Sun, Haifeng Zhao 0001 |
ACM Multimedia | 6 |
| 2024 | BCS-NeRF: Bundle Cross-Sensing Neural Radiance Fields
Mingwei Cao, Fengna Wang, Dengdi Sun, Haifeng Zhao 0001 |
MMAsia | 4 |
| 2024 | Semantics Guided Disentangled GAN for Chest X-Ray Image Rib Segmentation
Lili Huang 0006, Dexin Ma, Chenglong Li 0002, Haifeng Zhao 0001, Jin Tang 0001, Chuanfu Li |
PRCV (14) | 5 |
| 2024 | Domain Adaptive Lung Nodule Detection in X-Ray ImageabstractMedical images from different healthcare centers exhibit varied data distributions, posing significant challenges for adapting lung nodule detection due to the domain shift between training and application phases. Traditional unsupervised domain adaptive detection methods often struggle with this shift, leading to suboptimal outcomes. To overcome these challenges, we introduce a novel domain adaptive approach for lung nodule detection that leverages mean teacher self-training and contrastive learning. First, we propose a hierarchical contrastive learning strategy to refine nodule representations and enhance the distinction between nodules and background. Second, we introduce a nodule-level domain-invariant feature learning (NDL) module to capture domain-invariant features through adversarial learning across different domains. Additionally, we propose a new annotated dataset of X-ray images to aid in advancing lung nodule detection research. Extensive experiments conducted on multiple X-ray datasets demonstrate the efficacy of our approach in mitigating domain shift impacts. Haifeng Zhao 0001, Lixiang Jiang, Leilei Ma 0002, Dengdi Sun, Yanping Fu |
SMC | 1 |
| 2024 | Channel-Wise Interactive Learning for Remote Heart Rate Estimation From Facial VideoabstractRemote photoplethysmography measurement (also called rPPG prediction) is a vision-based technique that allows for the non-contact monitoring of human physiological activity using facial video. However, precisely detecting subtle color changes on facial skin, especially in less-constrained real-life scenarios, remains a formidable challenge for rPPG prediction. In this work, we address a rPPG-based heart rate estimation task by proposing an end-to-end Channel-wise Interaction Network (CIN-rPPG), in which the core idea contains two specialized units: channel-temporal interactive learning (CIT) and channel-spatial interactive learning (CIS). The CITunit gets the periodicity of the rPPG signal by using temporal-wise shifting and channel-wise scaling to measure the interaction between channels and temporal dimensions. The CISunit does both spatial-wise scaling and channel-wise scaling at the same time to perform channel-spatial interaction. This is intended to reveal how rPPG-related visual responses are detected on the human face. We exploit the rPPG recovery through the alternation of CITand CISimplementations. The CIN-rPPG is completely conducted by convolutional operations on the sequential 2D feature maps of facial video in an end-to-end manner. Extensive experiments on three heart rate estimation datasets (UBFC-rPPG, PURE, and MMSE-HR) demonstrate that CIN-rPPG achieves state-of-the-art performance on both intra-dataset and cross-dataset testing. Qi Li 0045, Dan Guo 0001, Xilan Tian, Xiao Sun 0003, Haifeng Zhao 0001, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Semisupervised Subspace Learning With Adaptive Pairwise Graph EmbeddingabstractGraph-based semisupervised learning can explore the graph topology information behind the samples, becoming one of the most attractive research areas in machine learning in recent years. Nevertheless, existing graph-based methods also suffer from two shortcomings. On the one hand, the existing methods generate graphs in the original high-dimensional space, which are easily disturbed by noisy and redundancy features, resulting in low-quality constructed graphs that cannot accurately portray the relationships between data. On the other hand, most of the existing models are based on the Gaussian assumption, which cannot capture the local submanifold structure information of the data, thus reducing the discriminativeness of the learned low-dimensional representations. This article proposes a semisupervised subspace learning with adaptive pairwise graph embedding (APGE), which first builds a -nearest neighbor graph on the labeled data to learn local discriminant embeddings for exploring the intrinsic structure of the non-Gaussian labeled data, i.e., the submanifold structure. Then, a -nearest neighbor graph is constructed on all samples and mapped to GE learning to adaptively explore the global structure of all samples. Clustering unlabeled data and its corresponding labeled neighbors into the same submanifold, sharing the same label information, improves embedded data's discriminative ability. And the adaptive neighborhood learning method is used to learn the graph structure in the continuously optimized subspace to ensure that the optimal graph matrix and projection matrix are finally learned, which has strong robustness. Meanwhile, the rank constraint is added to the Laplacian matrix of the similarity matrix of all samples so that the connected components in the obtained similarity matrix are precisely equal to the number of classes in the sample, which makes the structure of the graph clearer and the relationship between the near-neighbor sample points more explicit. Finally, multiple experiments on several synthetic and real-world datasets show that the method performs well in exploring local structure and classification tasks. Hebing Nie, Qi Li 0045, Zheng Wang 0037, Haifeng Zhao 0001, Feiping Nie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Semantic-Aware Dual Contrastive Learning for Multi-Label Image ClassificationabstractExtracting image semantics effectively and assigning corresponding labels to multiple objects or attributes for natural images is challenging due to the complex scene contents and confusing label dependencies. Recent works have focused on modeling label relationships with graph and understanding object regions using class activation maps (CAM). However, these methods ignore the complex intra- and inter-category relationships among specific semantic features, and CAM is prone to generate noisy information. To this end, we propose a novel semantic-aware dual contrastive learning framework that incorporates sample-to-sample contrastive learning (SSCL) as well as prototype-to-sample contrastive learning (PSCL). Specifically, we leverage semantic-aware representation learning to extract category-related local discriminative features and construct category prototypes. Then based on SSCL, label-level visual representations of the same category are aggregated together, and features belonging to distinct categories are separated. Meanwhile, we construct a novel PSCL module to narrow the distance between positive samples and category prototypes and push negative samples away from the corresponding category prototypes. Finally, the discriminative label-level features related to the image content are accurately captured by the joint training of the above three parts. Experiments on five challenging large-scale public datasets demonstrate that our proposed method is effective and outperforms the state-of-the-art methods. Code and supplementary materials are released on https://github.com/yu-gi-oh-leilei/SADCL. Leilei Ma 0002, Dengdi Sun, Lei Wang 0095, Haifeng Zhao 0001, Bin Luo 0001 |
ECAI | 4 |
| 2023 | G2PL: Lexicon Enhanced Chinese Polyphone Disambiguation Using Bert Adapter with a New DatasetabstractPolyphone disambiguation is the core of grapheme-to-phoneme(G2P) module for the Chinese speech synthesis system. However, there is a lack of datasets and only one public for polyphone disambiguation. Moreover, due to the double long-tail distribution of polyphones, the ratio of pronunciation data for most polyphones is extremely unbalanced after sampling. To solve these problems, we propose a new dataset with 57,000 sentences from various domains by a new strategy for sampling. In addition, we propose the G2PL, which integrates word features into the bottom of BERT to assist in predicting the correct pronunciation of polyphone. In the experiment, we train the G2PL model to outperform other methods on our and public datasets. Our dataset, codes and user-friendly package are freely available. Haifeng Zhao 0001, Hongzhi Wan, Lili Huang 0006, Mingwei Cao |
ICASSP | 1 |
| 2023 | Simultaneous local clustering and unsupervised feature selection via strong space constraint
Zheng Wang 0037, Qi Li 0045, Haifeng Zhao 0001, Feiping Nie 0001 |
Pattern Recognit. | 3 |
| 2022 | Depth-Aware Shadow RemovalabstractAbstract Shadow removal from a single image is an ill‐posed problem because shadow generation is affected by the complex interactions of geometry, albedo, and illumination. Most recent deep learning‐based methods try to directly estimate the mapping between the non‐shadow and shadow image pairs to predict the shadow‐free image. However, they are not very effective for shadow images with complex shadows or messy backgrounds. In this paper, we propose a novel end‐to‐end depth‐aware shadow removal method without using depth images, which estimates depth information from RGB images and leverages the depth feature as guidance to enhance shadow removal and refinement. The proposed framework consists of three components, including depth prediction, shadow removal, and boundary refinement. First, the depth prediction module is used to predict the corresponding depth map of the input shadow image. Then, we propose a new generative adversarial network (GAN) method integrated with depth information to remove shadows in the RGB image. Finally, we propose an effective boundary refinement framework to alleviate the artifact around boundaries after shadow removal by depth cues. We conduct experiments on several public datasets and real‐world shadow images. The experimental results demonstrate the efficiency of the proposed method and superior performance against state‐of‐the‐art methods. Yanping Fu, Zhenyu Gai, Haifeng Zhao 0001, Shaojie Zhang 0002, Ying Shan, Yang Wu 0001, Jin Tang 0001 |
Comput. Graph. Forum | 3 |
| 2022 | ORSI Salient Object Detection via Multiscale Joint Region and Boundary ModelabstractSalient object detection (SOD) in optical remote sense images (ORSIs) is a valuable and challenging task. The factors in ORSI, such as background clutter, lighting shadows, imaging blur, and low resolution, significantly degrade the completeness and accuracy of salient objects. To handle this problem, we propose a novel model to learn robust multiscale region features of salient objects by simultaneously optimizing their boundaries. First, we extract multiscale region features of salient objects through a hierarchical attention module. Second, we generate the boundary features by combining the local cues and the global information generated by pyramid pooling. Finally, we embed the boundary features into region features at multiple scales. In particular, we design a joint learning scheme based on a bidirectional feature transformation to optimize boundary and region features simultaneously for accurate ORSI SOD. To provide a comprehensive evaluation platform, we construct a new dataset called ORSI-4199 for ORSI SOD. It contains 4199 finely annotated image pairs with diverse scenes, in which nine attributes (i.e., challenge types) are annotated to facilitate analyzing the strengths and weaknesses of SOD models from different perspectives. Extensive experiments on the public dataset ORSSD, EORRSD, and the newly created dataset ORSI-4199 show that the proposed approach achieves promising results against state-of-the-art methods.https://github.com/wchao1213/ORSI-SOD. Zhengzheng Tu, Chenglong Li 0002, Minghao Fan, Haifeng Zhao 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | CGFNet: cross-guided fusion network for RGB-thermal semantic segmentation
Yanping Fu, Qiaoqiao Chen, Haifeng Zhao 0001 |
Vis. Comput. | 3 |
| 2021 | Intracranial Hematoma Classification Based on the Pyramid Hierarchical Bilinear Pooling
Haifeng Zhao 0001, Dejun Bao, Shaojie Zhang 0002 |
PRCV (3) | 1 |
| 2020 | Multiclass discriminant analysis via adaptive weighted scheme
Haifeng Zhao 0001, Zheng Wang 0037, Feiping Nie 0001 |
Neurocomputing | 1 |
| 2019 | A New Formulation of Linear Discriminant Analysis for Robust Dimensionality ReductionabstractDimensionality reduction is a critical technology in the domain of pattern recognition, and linear discriminant analysis (LDA) is one of the most popular supervised dimensionality reduction methods. However, whenever its distance criterion of objective function uses$L_2$-norm, it is sensitive to outliers. In this paper, we propose a new formulation of linear discriminant analysis via joint$L_{2,1}$-norm minimization on objective function to induce robustness, so as to efficiently alleviate the influence of outliers and improve the robustness of proposed method. An efficient iterative algorithm is proposed to solve the optimization problem and proved to be convergent. Extensive experiments are performed on an artificial data set, on UCI data sets, and on four face data sets, which sufficiently demonstrates the efficiency of comparing to other methods and robustness to outliers of our approach. Haifeng Zhao 0001, Zheng Wang 0037, Feiping Nie 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | Adaptive Neighborhood MinMax Projections
Haifeng Zhao 0001, Zheng Wang 0037, Feiping Nie 0001 |
Neurocomputing | 1 |
| 2018 | Multiclass Classification and Feature Selection Based on Least Squares Regression with Large MarginabstractLeast squares regression (LSR) is a fundamental statistical analysis technique that has been widely applied to feature learning. However, limited by its simplicity, the local structure of data is easy to neglect, and many methods have considered using orthogonal constraint for preserving more local information. Another major drawback of LSR is that the loss function between soft regression results and hard target values cannot precisely reflect the classification ability; thus, the idea of the large margin constraint is put forward. As a consequence, we pay attention to the concepts of large margin and orthogonal constraint to propose a novel algorithm, orthogonal least squares regression with large margin (OLSLM), for multiclass classification in this letter. The core task of this algorithm is to learn regression targets from data and an orthogonal transformation matrix simultaneously such that the proposed model not only ensures every data point can be correctly classified with a large margin than conventional least squares regression, but also can preserve more local data structure information in the subspace. Our efficient optimization method for solving the large margin constraint and orthogonal constraint iteratively proved to be convergent in both theory and practice. We also apply the large margin constraint in the process of generating a sparse learning model for feature selection via joint [Formula: see text]-norm minimization on both loss function and regularization terms. Experimental results validate that our method performs better than state-of-the-art methods on various real-world data sets. Haifeng Zhao 0001, Zheng Wang 0037 |
Neural Comput. | 1 |
| 2016 | Orthogonal least squares regression for feature extraction
Haifeng Zhao 0001, Zheng Wang 0037, Feiping Nie 0001 |
Neurocomputing | 1 |
| 2015 | Image matching using a local distribution based outlier detection technique
Haifeng Zhao 0001, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
Neurocomputing | 1 |
| 2014 | A sparse nonnegative matrix factorization technique for graph matching problems
Bo Jiang 0002, Haifeng Zhao 0001, Jin Tang 0001, Bin Luo 0001 |
Pattern Recognit. | 2 |
| 2009 | Registration of blurred images for image mosaicabstractExisting methods for the registration of blurred images are efficient for the artificially blurred images or a planar registration. They are not suitable for image mosaic of the source images from a real camera with an almost fixed optical center. We propose a registration method so that a distortion-free registration on naturally captured images can be obtained. It adopts a multi-resolution and robust feature based inter-layer mosaic together. In each layer, Harris corner detector is chosen to effectively detect features and RANSAC is used to find reliable matches for further calibration as well as an initial homography as the initial motion of next layer. Simplex and subspace trust region methods are used consequently to estimate the stable focal length and rotation matrix through the transformation property of feature matches. Experimental results demonstrate the performance of our proposed method. Xianyong Fang, Bin Luo 0001, Jin Tang 0001, Haifeng Zhao 0001 |
CAD/Graphics | 4 |
| 2007 | 2D-LPP: A two-dimensional extension of locality preserving projections
Sibao Chen 0001, Haifeng Zhao 0001, Min Kong, Bin Luo 0001 |
Neurocomputing | 2 |