EDBT 2026 Demo / reviewers in the wild / expert
Haibo Ye
dblp:299/4980
· DBLP profile ↗
27ranked-venue papers
13as first author
14since 2021 · last 2025
0000-0001-9034-7013ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Computer networks · 6 · 4 first-authorHuman-computer interaction and ubiquitous computing · 5 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Embarrassingly Simple but Effective Knowledge-enhanced RecommenderabstractKnowledge graphs (KG) have demonstrated significant potential in recommender systems by providing complementary semantic information that is typically absent in user-item interaction graphs (IG). While contrastive learning has emerged as a powerful paradigm for integrating these dual information sources, we identify a critical limitation in existing approaches: current methods fail to effectively balance the contrastive views derived from IG and KG, often resulting in performance degradation compared to using IG alone. To address this fundamental challenge, we propose SimKGCL, a novel contrastive learning framework that introduces a simple yet principled solution -- cross-view, layer-wise fusion between IG and KG representations prior to contrastive learning. This design ensures effective knowledge transfer while maintaining the discriminative power of contrastive objectives. Comprehensive experiments across three real-world benchmarks demonstrate that our approach not only consistently outperforms existing methods but also achieves remarkable efficiency gains. Our code is available through this link: https://figshare.com/articles/conference_contribution/SimKGCL/22783382. Haibo Ye, Yuan Yao 0001, Xinjie Li 0007 |
CIKM | 1 |
| 2025 | Semantic Integrity Preserving Recommendation Framework with Large Language ModelsabstractRecommendation systems based on deep learning and graph neural networks have effectively captured complex user-item relationships but struggle to fully leverage textual data for personalized recommendations. This is primarily due to three key challenges: (1) shallow semantic integration in language model representations, (2) semantic degradation during the fusion of large language model (LLM) embeddings with collaborative filtering signals, and (3) scalability bottlenecks in sparse data and large-scale interaction scenarios. To address these challenges, we propose SIP-Rec, a Semantic Integrity Preserving Recommendation Framework based on LLM and contrastive learning. First, SIP-Rec utilizes LLM to generate semantically rich profiles, accurately capturing fine-grained differences in user preferences and item characteristics. Second, it introduces an innovative cross-modal fusion mechanism that employs a dynamic alignment strategy to bridge the representation space gap between high-dimensional LLM embeddings and collaborative filtering signals, effectively mitigating contextual information loss. Finally, by integrating knowledge distillation with contrastive learning objectives, SIP-Rec optimizes the consistency between semantically rich representations and interaction patterns. Rongbing Niu, Yunfeng Fan, Haibo Ye |
CW | 3 |
| 2025 | Dual-Phase Graph Refinement: Pruning Noisy Edges and Augmenting Tail Links for RecommendationabstractImplicit feedback recommender systems face significant performance challenges due to noisy interactions and sparse data from long-tail users. Noisy edges propagate errors through graph convolutions, while sparse interactions of long-tail users bias recommendations toward popular items, compromising fairness and accuracy. This paper proposes a novel Dual-Phase Graph Refinement (DPGR) framework, built upon GNN, to optimize the interaction graph structure through synergistic denoising and long-tail link augmentation. In the first phase, DPGR employs node neighborhood similarity with a dynamic thresholding mechanism to accurately identify and remove unreliable edges, constructing a high-quality subgraph to enhance embedding robustness. In the second phase, targeting longtail users-defined as the bottom 80% by interaction count-DPGR selectively adds high-confidence links via a probabilistic sampling model, enriching their neighborhood representations and mitigating sparsity. Comprehensive empirical studies on three diverse datasets demonstrate that DPGR outperforms other methods in both overall and long-tail performance, achieving average improvements of 7.98% and 6.53% over the baseline, respectively, across Recall and NDCG metrics. Ablation studies confirm the necessity of the dual-phase design, and noise injection experiments further highlight the method's robustness. By jointly optimizing noise reduction and long-tail augmentation, DPGR offers an efficient and equitable solution for graph-based collaborative filtering. Haibo Ye |
CW | 4 |
| 2025 | Divide, Weight, and Route: Difficulty-Aware Optimization with Dynamic Expert Fusion for Long-Tailed Recognition
Xiaolei Wei, Haibo Ye |
PRCV (6) | 3 |
| 2025 | Feature calibration and feature separation for long-tailed visual recognition
Fangyu Zhou, Xiangge Zhao, Yangtao Lin, Haibo Ye |
Neurocomputing | 5 |
| 2025 | Diffusion-Noise-Based Augmentation for Long-Tailed Remote Sensing Image ClassificationabstractRemote sensing image classification refers to the task that using algorithms to categorize satellite or aerial imagery into different land cover types. In the real world, long-tailed data distribution is commonly present in remote sensing image classification tasks, causing models to excessively favor sufficient head classes during training and degrading prediction accuracy for scarce tail classes. Although existing methods such as resampling and re-weighting can alleviate the issue of data imbalance to a certain extent, they struggle to sufficiently enhance the diversity of tail class samples. In recent years, some works have begun to use diffusion models to generate diverse samples to balance the class distribution. However, these approaches often overlook the distribution inconsistency between generated and real images, which will hinder the improvement of model performance. To tackle these challenges, this paper proposes a novel diffusion-noise-based augmentation method (DONA) with a two-stage training process. Before training, our specially designed conditional prompts are used together with the original training set to guide the diffusion model in image generation. Furthermore, we propose two strategies to effectively leverage the generated images, which are applied respectively at the end of the first training stage and during the second training stage. First, we design DiffCam-Mix to fuse the background of the generated data with the foreground of the original data, preserving the essential information of the original real images while incorporating the diversity of the generated ones. Second, we use cosine similarity to minimize the differences between the mixed data and their corresponding original data, further calibrating the distribution of different samples. Extensive experiments on three public datasets—SIRI-WHU-LT, PatternNet-LT, and RSI-CB256-LT—demonstrate the effectiveness of the proposed method. Qianqian Wang 0014, Haibo Ye, Dong Liang 0008, Sheng-Jun Huang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Bidirectional Uncertainty-Based Active Learning for Open-Set Annotation
Chen-Chen Zong, Ye-Wen Wang, Kun-Peng Ning, Haibo Ye, Sheng-Jun Huang |
ECCV (28) | 4 |
| 2024 | Class-Aware Feature Perturbation for Long-Tailed Visual RecognitionabstractThe distribution of data in the real world is often imbalanced, with a small number of classes having a large number of instances, while most other classes have relatively few sample instances, resulting in a long-tailed distribution. Supervised contrastive learning still tends to favor classes with more samples, making it challenging to learn discriminative and robust features in imbalanced settings. In our work, we theoretically analyze the advantages of hard samples in contrastive learning and propose an end-to-end training framework based on class-aware perturbation. Specifically, we introduce different levels of perturbation based on class priors to construct hard samples, improve the disadvantaged position of tail classes, balance the feature space distribution, and dynamically adjust the classifier output with a combination of local batch priors and global priors. Experimental results show that our method achieves competitive performance across multiple long-tailed benchmark datasets. Xicheng Chen, Haibo Ye, Fangyu Zhou |
ICME | 2 |
| 2024 | Denoised Graph Collaborative Filtering via Neighborhood Similarity and Dynamic ThresholdingabstractGraph collaborative filtering (GCF) has achieved great success in recommender systems due to its ability in mining high-order collaborative signals from historical user-item interactions. However, GCF's performance could be severely affected by the intrinsic noise within the user-item interactions. To this end, several denoised GCF frameworks have been proposed, whose heart is to estimate and handle the reliability of existing interactions. However, most of them suffer from two limitations: 1) the reliability computation itself is noisy, and 2) the reliability threshold is difficult to determine. To address the two limitations, in this paper, we propose a newNeighborhood-informedDenoising framework NiDen for GCF. Specifically, for an existing user-item interaction, NiDen first estimates its reliability by employing the neighborhood information of the user and the item, and then determines whether the interaction is noisy or not via a dynamic thresholding strategy. After that, NiDen mitigates the negative impact of noise by both structure denoising and sample re-weighting. We instantiate NiDen on two representative GCF models and conduct extensive experiments on four widely-used datasets. The results show that NiDen achieves the best performance compared to the existing denoising methods, especially on datasets with heavy noise. Haibo Ye, Yuan Yao 0001, Sheng-Jun Huang |
IEEE Trans. Big Data | 1 |
| 2023 | Balanced Mixup Loss for Long-Tailed Visual RecognitionabstractIn the real world, the data collected naturally are often long-tailed, which inevitably leads to class-imbalanced prediction and performance degradation. As a simple and effective data augmentation method, Mixup has been proven to be beneficial for the tail class in recent long-tail learning studies. However, samples selected by Mixup are still imbalanced, which could exacerbate the imbalance problem. Existing work always considers adjusting the input distribution to alleviate this problem, while this may lead to over-fitting in the tail class. In this paper, we detail the theoretical analysis of the data imbalance caused by Mixup, and propose a novel Balanced Mixup (BaMix) loss function from the output perspective. By adding a balance term, we theoretically prove the BaMix loss can overcome the imbalance caused by Mixup. In the experiments, our solution achieves the state-of-the-art performance on CIFAR-LT, ImageNet-LT, and iNaturalist 2018. Haibo Ye, Fangyu Zhou, Xinjie Li 0007, Qingheng Zhang |
ICASSP | 1 |
| 2023 | Mendam: Multi-Expert Network with Distribution-Aware Momentum for Long-Tailed RecognitionabstractLong-tailed data distribution (i.e., minority classes occupy most of the data, while most classes have very few samples) is a common problem in image classification. The existing deferred re-balancing methods suffer from low accuracy for tail classes due to their bias towards head classes. In this paper, we find that properly adjusting the momentum can significantly improve the performance for class-imbalanced tasks, which provides a novel perspective to solve this thorny problem. Based on this finding, we propose a new way to calculate a distribution-aware momentum (DAM) based on the data distribution to avoid bias towards head classes. Furthermore, we design a multi-expert network (MEN) to designate each expert to focus on different tail classes. We achieve new state-of-the-art on CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT and iNaturalist2018 for image classification. Qingheng Zhang, Haibo Ye, Kaicheng Yu |
ICASSP | 2 |
| 2023 | Space-Transform Margin Loss with Mixup for Long-Tailed Visual Recognition
Fangyu Zhou, Xicheng Chen, Haibo Ye |
PRCV (8) | 3 |
| 2023 | Towards Robust Neural Graph Collaborative Filtering via Structure Denoising and Embedding PerturbationabstractNeural graph collaborative filtering has received great recent attention due to its power of encoding the high-order neighborhood via the backbone graph neural networks. However, their robustness against noisy user-item interactions remains largely unexplored. Existing work on robust collaborative filtering mainly improves the robustness by denoising the graph structure, while recent progress in other fields has shown that directly adding adversarial perturbations in the embedding space can significantly improve the model robustness. In this work, we propose to improve the robustness of neural graph collaborative filtering via both denoising in the structure space and perturbing in the embedding space. Specifically, in the structure space, we measure the reliability of interactions and further use it to affect the message propagation process of the backbone graph neural networks; in the embedding space, we add in-distribution perturbations by mimicking the behavior of adversarial attacks and further combine it with contrastive learning to improve the performance. Extensive experiments have been conducted on four benchmark datasets to evaluate the effectiveness and efficiency of the proposed approach. The results demonstrate that the proposed approach outperforms the recent neural graph collaborative filtering methods especially when there are injected noisy interactions in the training data. Haibo Ye, Xinjie Li 0007, Yuan Yao 0001, Hanghang Tong |
ACM Trans. Inf. Syst. | 1 |
| 2021 | Asynchronous Active Learning with Distributed Label QueryingabstractActive learning tries to learn an effective model with lowest labeling cost. Most existing active learning methods work in a synchronous way, which implies that the label querying can be performed only after the model updating in each iteration. While training models is usually time-consuming, it may lead to serious latency between two queries, especially in the crowdsourcing environments where there are many online annotators working simultaneously. This will significantly decrease the labeling efficiency and strongly limit the application of active learning in real tasks. To overcome this challenge, we propose a multi-server multi-worker framework for asynchronous active learning in the distributed environment. By maintaining two shared pools of candidate queries and labeled data respectively, the servers, the workers and the annotators efficiently corporate with each other without synchronization. Moreover, diverse sampling strategies from distributed workers are incorporated to select the most useful instances for model improving. Both theoretical analysis and experimental study validate the effectiveness of the proposed approach. Sheng-Jun Huang, Chen-Chen Zong, Kun-Peng Ning, Haibo Ye |
IJCAI | 4 |
| 2019 | CBSC: A Crowdsensing System for Automatic Calibrating of Barometers
Haibo Ye, Xuansong Li, Kai Dong 0001 |
J. Comput. Sci. Technol. | 1 |
| 2019 | SMinder: Detect a Left-behind Phone using Sensor-based Context Awareness
Haibo Ye, Kai Dong 0001, Tao Gu 0001 |
Mob. Networks Appl. | 1 |
| 2018 | On the limitations of existing notions of location privacy
Kai Dong 0001, Taolin Guo, Haibo Ye, Xuansong Li, Zhen Ling 0001 |
Future Gener. Comput. Syst. | 3 |
| 2017 | Dealing with Insufficient Location Fingerprints in Wi-Fi Based Indoor Location FingerprintingabstractThe development of the Internet of Things has accelerated research in the indoor location fingerprinting technique, which provides value-added localization services for existing WLAN infrastructures without the need for any specialized hardware. The deployment of a fingerprinting based localization system requires an extremely large amount of measurements on received signal strength information to generate a location fingerprint database. Nonetheless, this requirement can rarely be satisfied in most indoor environments. In this paper, we target one but common situation when the collected measurements on received signal strength information are insufficient, and show limitations of existing location fingerprinting methods in dealing with inadequate location fingerprints. We also introduce a novel method to reduce noise in measuring the received signal strength based on the maximum likelihood estimation, and compute locations from inadequate location fingerprints by using the stochastic gradient descent algorithm. Our experiment results show that our proposed method can achieve better localization performance even when only a small quantity of RSS measurements is available. Especially when the number of observations at each location is small, our proposed method has evident superiority in localization accuracy. Kai Dong 0001, Zhen Ling 0001, Xiangyu Xia, Haibo Ye, Wenjia Wu, Ming Yang 0001 |
Wirel. Commun. Mob. Comput. | 4 |
| 2016 | A Reliability-Augmented Particle Filter for Magnetic Fingerprinting Based Indoor Localization on SmartphoneabstractUsing magnetic field data as fingerprints for smartphone indoor positioning has become popular in recent years. Particle filter is often used to improve accuracy. However, most of existing particle filter based approaches either are heavily affected by motion estimation errors, which result in unreliable systems, or impose strong restrictions on smartphone such as fixed phone orientation, which are not practical for real-life use. In this paper, we present a novel indoor positioning system for smartphones, which is built on our proposed reliability-augmented particle filter. We create several innovations on the motion model, the measurement model, and the resampling model to enhance the basic particle filter. To minimize errors in motion estimation and improve the robustness of the basic particle filter, we propose a dynamic step length estimation algorithm and a heuristic particle resampling algorithm. We use a hybrid measurement model, combining a new magnetic fingerprinting model and the existing magnitude fingerprinting model, to improve system performance, and importantly avoid calibrating magnetometers for different smartphones. In addition, we propose an adaptive sampling algorithm to reduce computation overhead, which in turn improves overall usability tremendously. Finally, we also analyze the “Kidnapped Robot Problem” and present a practical solution. We conduct comprehensive experimental studies, and the results show that our system achieves an accuracy of 1~2 m on average in a large building. Hongwei Xie, Tao Gu 0001, XianPing Tao, Haibo Ye, Jian Lu 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2016 | Scalable floor localization using barometer on smartphoneabstractAbstract Traditional fingerprint‐based localization techniques mainly rely on infrastructure support such as GSM and Wi‐Fi. They require war‐driving, which is both time‐consuming and labor‐intensive. With recent advances of smartphone sensors, sensor‐assisted localization techniques are emerging. However, they often need user‐specific training and more power intensive sensing, resulting in infeasible solutions for real deployment. In this paper, we present Barometer‐based floor Localization system (B‐Loc), a novel floor localization system to identify the floor level in a multi‐floor building on which a mobile user is located. It makes use of the barometer on smartphone. B‐Loc does not rely on any Wi‐Fi infrastructure and requires neither war‐driving nor prior knowledge of the buildings. Leveraging on crowdsourcing, B‐Loc builds the barometer fingerprint map, which contains the barometric pressure value for each floor level to locate users' floor levels. We conduct both simulation and field studies to demonstrate the accuracy, scalability, and robustness of B‐Loc. Our simulation shows that B‐Loc can locate the user fast and the field study in a 10‐floor building shows that B‐Loc achieves an accuracy of over 98%. Copyright © 2016 John Wiley & Sons, Ltd. Haibo Ye, Tao Gu 0001, XianPing Tao, Jian Lu 0001 |
Wirel. Commun. Mob. Comput. | 1 |
| 2015 | Infrastructure-Free Floor Localization Through Crowdsourcing
Haibo Ye, Tao Gu 0001, XianPing Tao, Jian Lu 0001 |
J. Comput. Sci. Technol. | 1 |
| 2014 | MaLoc: a practical magnetic fingerprinting approach to indoor localization using smartphonesabstractUsing magnetic field data as fingerprints for localization in indoor environment has become popular in recent years. Particle filter is often used to improve accuracy. However, most of existing particle filter based approaches either are heavily affected by motion estimation errors, which makes the system unreliable, or impose strong restrictions on smartphone such as fixed phone orientation, which is not practical for real-life use. In this paper, we present an indoor localization system named MaLoc, built on our proposed augmented particle filter. We create several innovations on the motion model, the measurement model and the resampling model to enhance the traditional particle filter. To minimize errors in motion estimation and improve the robustness of particle filter, we augment the particle filter with a dynamic step length estimation algorithm and a heuristic particle resampling algorithm. We use a hybrid measurement model which combines a new magnetic fingerprinting model and the existing magnitude fingerprinting model to improve the system performance and avoid calibrating different smartphone magnetometers. In addition, we present a novel localization quality estimation method and a localization failure detection method to address the "Kidnapped Robot Problem" and improve the overall usability. Our experimental studies show that MaLoc achieves a localization accuracy of 1~2.8m on average in a large building. Hongwei Xie, Tao Gu 0001, XianPing Tao, Haibo Ye, Jian Lu 0001 |
UbiComp | 4 |
| 2014 | F-Loc: Floor localization via crowdsourcingabstractTraditional fingerprint based localization techniques mainly rely on infrastructure support such as GSM, Wi-Fi or GPS. They work by war-driving the entire indoor spaces which is both time-consuming and labor-intensive. With recent advances of smartphone and sensing technologies, sensor-assisted localization techniques leveraging on mobile phone sensing are emerging. However, sensors are inherently noisy, making this technique challenging for real deployment. In this paper, we present F-Loc, a novel floor localization system to identify the floor level in a multi-floor building on which a mobile user is located. It does not need to war-drive the entire building. Leveraging on crowdsourcing and mobile phone sensing, we collect users' Wi-Fi traces and accelerometer readings. Through advanced clustering and cluster manipulating techniques, we are able to build the Wi-Fi map of the entire building, which can then be used for floor localization. We conduct both simulation and field studies to demonstrate the accuracy, scalability, and robustness of F-Loc. Our field study in a 10-floor building shows that F-Loc achieves an accuracy of over 98%. Haibo Ye, Tao Gu 0001, XianPing Tao, Jian Lu 0001 |
ICPADS | 1 |
| 2014 | B-Loc: Scalable Floor Localization Using Barometer on SmartphoneabstractTraditional fingerprint based localization techniques mainly rely on infrastructure support such as GSM and Wi-Fi. They require war-driving which is both time-consuming and labor-intensive. With recent advances of smartphone sensors, sensor-assisted localization techniques are emerging. However, they often need user-specific training and more power intensive sensing, resulting in infeasible solutions for real deployment. In this paper, we present B-Loc, a novel floor localization system to identify the floor level in a multi-floor building on which a mobile user is located. It makes use of the barometer on smartphone only. B-Loc does not rely on any Wi-Fi infrastructure and requires neither war-driving nor prior knowledge of the buildings. Leveraging on crowd sourcing, B-Loc builds the barometer fingerprint map which contains the barometric pressure value for each floor level to locate users' floor levels. We conduct both simulation and field studies to demonstrate the accuracy, scalability, and robustness of B-Loc. Our field study in a 10-floor building shows that B-Loc achieves an accuracy of over 98%. Haibo Ye, Tao Gu 0001, XianPing Tao, Jian Lu 0001 |
MASS | 1 |
| 2014 | SBC: scalable smartphone barometer calibration through crowdsourcingabstractWe have seen increasingly popularity in embedding barometer into smartphone today. A barometer measures the barometric pressure, and it can be used for a variety of applications. For example, in localization techniques, it is used to detect the altitude or altitude change of a user. Unfortunately, t Haibo Ye, Tao Gu 0001, XianPing Tao, Jian Lu 0001 |
MobiQuitous | 1 |
| 2014 | Crowdsourced smartphone sensing for localization in metro trainsabstractTraditional fingerprint based localization techniques mainly rely on infrastructure support such as RFID, Wi-Fi or GPS. They operate by war-driving the entire space which is both time-consuming and labor-intensive. In this paper, we present M-Loc, a novel infrastructure-free localization system to locate mobile users in a metro line. It does not rely on any Wi-Fi infrastructure, and does not need to war-drive the metro line. Leveraging crowdsourcing, we collect accelerometer, magnetometer and barometer readings on smartphones, and analyze these sensor data to extract patterns. Through advanced data manipulating techniques, we build the pattern map for the entire metro line, which can then be used for localization. We conduct field studies to demonstrate the accuracy, scalability, and robustness of M-Loc. The results of our field studies in 3 metro lines with 55 stations show that M-Loc achieves an accuracy of 93% when travelling 3 stations, 98% when travelling 5 stations. Haibo Ye, Tao Gu 0001, XianPing Tao, Jian Lu 0001 |
WoWMoM | 1 |
| 2012 | FTrack: Infrastructure-free floor localization via mobile phone sensingabstractMobile phone localization plays a key role in the fast-growing Location Based Applications domain. Most of the existing localization schemes rely on infrastructure support such as GSM, WiFi or GPS. In this paper, we present FTrack, a novel floor localization system to identify the floor level in a multi-floor building on which a mobile user is located. FTrack uses the mobile phone's accelerometer only without any infrastructure support. It does not require any prior knowledge of the building such as floor height. By capturing user encounters and analyzing user trails, FTrack finds the mapping from the traveling time (when taking the elevator) or the step counts (when walking on the stairs) between any two floors to the number of floor levels. The mapping can then be used for mobile users to pinpoint their current floor levels. We conduct both simulation and field studies to demonstrate the effectiveness of FTrack. Our field trial in a 10-floor building shows that FTrack achieves an accuracy of over 90% after two hours in our experiment. Haibo Ye, Tao Gu 0001, Jinwei Xu, XianPing Tao, Jian Lu 0001 |
PerCom | 1 |