VLDB 2026 Research / reviewers in the wild / expert
Hai Fang
dblp:55/4892
· DBLP profile ↗
17ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorComputer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Point cloud instance segmentation for building indoor scenes using deep learning and BIM-Generated synthetic point clouds
Hongzhe Yue, Qian Wang 0007, Xiang Nie, Hai Fang, Jack C. P. Cheng, Shuju Jing |
Adv. Eng. Informatics | 4 |
| 2026 | An Integrated OTFS-NOMA Framework for Multi-Beam LEO Systems: Reliability and Capacity AnalysisabstractMulti-beam low earth orbit (LEO) satellite communications, as an essential component for 6G systems, may encounter challenges from severe Doppler shifts and co-channel interference. This paper addresses a realistic problem in 6G-LEO systems, that is, how to meet the high-reliability demands of massive high-mobility terminals. We propose an integrated framework to exploit the synergy of non-orthogonal multiple access (NOMA) and orthogonal time frequency space (OTFS). OTFS modulation is employed to achieve full time-frequency diversity to combat Doppler shifts, while NOMA is used to accommodate more access requests. Specifically, within each beam, power domain superposition is applied to the delay-Doppler domain, enabling multiple terminals to share delay-Doppler grid resources. We analyze the performance of reliability, outage probability and ergodic capacity. Notably, we derive a novel closed-form expression to characterize the distribution of multi-beam interference with varying beam gains. Theoretical analysis and simulation results confirm that the proposed framework achieves a substantially lower outage probability compared to conventional OFDM schemes, with a system capacity improvement exceeding 11.9%. Xiaohui Zhao 0007, Lei Lei 0001, Zhiqiang Wei 0001, Hai Fang, Wenjie Wang 0001, Symeon Chatzinotas |
IEEE Trans. Wirel. Commun. | 4 |
| 2025 | Lightweight Adaptive PPO-AHP Enhanced Algorithm for Task Offloading in Vehicular Edge ComputingabstractModern vehicular networks leverage Road Side Units (RSUs) as distributed edge computing hubs to support delay-critical IoV applications. However, existing approaches face three critical limitations: slow convergence of tasking offloading algorithms in high-dimensional state spaces, ineffective temporal dependency modeling due to simplistic state representations, and reactive task offloading strategies lacking predictive caching mechanisms. To address these challenges, we propose a Lightweight Adaptive PPO-Transformer Enhanced System. The framework pioneers two synergistic innovations: First, a proactive caching engine that predicts content popularity through temporal demand patterns. Second, an Adaptive Hybrid Performer (AHP) architecture that unites Proximal Policy Optimization (PPO) with lightweight Transformer networks. The AHP achieves dynamic feature extraction via dual fixed-trainable attention streams while ensuring linear O(L) complexity through kernel-based linear attention. Experimental results demonstrate our method’s superior performance: 42.06% faster convergence, 22.02% lower latency, and 8.56% higher cache hits than state-of-the-art baselines. These improvements maintain consistency across diverse RSU deployments and mobility scenarios, demonstrating robust performance in dynamic IoV environments. Hai Fang, Jiadong Tang |
IJCNN | 2 |
| 2025 | Enhancing semantic segmentation of MEP scenes with deep learning and BIM-generated synthetic point clouds
Hongzhe Yue, Qian Wang 0007, Hai Fang, Jack C. P. Cheng |
Adv. Eng. Informatics | 5 |
| 2025 | Inferring pathway activity from single-cell and spatial transcriptomics data with PaaScabstractRecent advances in single-cell and spatial transcriptomics have revolutionized our understanding of cellular heterogeneity. However, translating high-dimensional data into functional pathway insights remains challenging. To address this obstacle, we developed PaaSc (Pathway activity analysis of Single-cell), a computational method for inferring pathway activity at single-cell resolution. PaaSc employs multiple correspondence analysis to simultaneously project cells and genes into a common latent space and selects pathway-associated dimensions through linear regression to infer pathway activity scores. We validated PaaSc across diverse benchmarking datasets, including those that jointly profiled protein and RNA levels, as well as large-scale cancer scRNA-seq cohorts. Compared with state-of-the-art methods, PaaSc demonstrated superior performance in multiple applications: scoring cell type-specific gene sets, identifying cell senescence-associated pathways, and exploring GWAS trait-associated cell types. Importantly, PaaSc maintained accuracy despite batch effects and demonstrated robust performance across different data modalities, including scATAC-seq and spatial transcriptomics data. Our results demonstrate that PaaSc accurately captures dynamic cellular states and spatial patterns, thereby advancing our understanding of cellular dynamics, aging, and disease mechanisms. Xiqi Liao, Yuyang Hong, Henghui Li, Hai Fang |
PLoS Comput. Biol. | 5 |
| 2024 | Heter-Train: A Distributed Training Framework Based on Semi-Asynchronous Parallel Mechanism for Heterogeneous Intelligent Transportation SystemsabstractTransportation big data (TBD) are increasingly combined with artificial intelligence to mine novel patterns and information due to the powerful representational capabilities of deep neural networks (DNNs), especially for anti-COVID19 applications. The distributed cloud-edge-vehicle training architecture has been applied to accelerate DNNs training while ensuring low latency and high privacy for TBD processing. However, multiple intelligent devices (e.g., intelligent vehicles, edge computing chips at base stations) and different networks in intelligent transportation systems lead to computing power and communication heterogeneity among distributed nodes. Existing parallel training mechanisms perform poorly on heterogeneous cloud-edge-vehicle clusters. The synchronous parallel mechanism may force fast workers to wait for the slowest worker for synchronization, thus wasting their computing power. The asynchronous mechanism has communication bottlenecks and can exacerbate the straggler problem, causing increased training iterations and even incorrect convergence. In this paper, we introduce a distributed training framework, Heter-Train. First, a communication-efficient semi-asynchronous parallel mechanism (SAP-SGD) is proposed, which can take full advantage of acceleration effect of asynchronous strategy on heterogeneous training and constrain the straggler problem by using global interval synchronization. Second, Considering the difference in node bandwidth, we design a solution for heterogeneous communication. Moreover, a novel weighted aggregation strategy is proposed to aggregate the model parameters with different versions. Finally, experimental results show that our proposed strategy can achieve up to$6.74 \times $speedups on training time, with almost no accuracy decrease. Jiawei Geng, Haipeng Jia, Zongwei Zhu, Hai Fang, Chengxi Gao, Cheng Ji 0002, Gangyong Jia, Guangjie Han, Xuehai Zhou |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Ensemble learning of multi-kernel Kriging surrogate models using regional discrepancy and space-filling criteria-based hybrid sampling method
Xiaobing Shang, Hai Fang, Yunhui Li |
Adv. Eng. Informatics | 3 |
| 2023 | Cost-Efficient Scheduling of Streaming Applications in Apache Flink on CloudabstractStream processing has been gaining extensive attention in the past few years. Apache Flink is a new generation of distributed stream processing engines that can process a great deal of data in real-time with low latency. But the default scheduler of Flink adopts a random task scheduling strategy, which does not consider the cost and load balancing in the cloud environment. In this article, a cost-efficient task scheduling algorithm (CETSA) and a cost-efficient load balancing algorithm (LBA-CE) for Flink are proposed to reduce the job execution cost while optimizing load balancing. First, a cost-efficient model and a load balancing model based on Flink are constructed. Then, the core mechanism of Flink task scheduling is improved based on the cost-efficient model and the improved task scheduler is implemented. In addition, the concept of node adaptation is introduced into cost-efficient scheduling according to the load balancing model, ensuring that the cluster load is balanced as much as possible while reducing the cost in a heterogeneous cluster. Extensive experiments have been performed with Hibench's Wordcount and Fixwindow workloads in the cloud environment. The experimental results indicate that compared to the baseline scheduling algorithm, the proposed algorithms reduce the cost by about 37.9% and 20.2% on average, and the load deviation of the cluster is reduced by about 23.1% and 24.6% on average, respectively. In summary, the proposed algorithms in this paper can significantly reduce the cost of executing jobs and optimize the load balancing of the cluster in Flink. Jianglin Xia, Hai Fang |
IEEE Trans. Big Data | 4 |
| 2022 | Matching Game based Task Offloading and Resource Allocation Algorithm for Satellite Edge Computing NetworksabstractAiming at the problem of computing offloading with resource constraints in edge computing of dual-layer satellite networks, each low earth orbit (LEO) satellite can offload computing workloads to Geostationary Orbit (GEO) satellites. In order to obtain the optimal strategy for radio resource allocation and offloading decision-making, the system overhead minimization problem is studied, and the offloading decision problem is modeled as a two-sided matching game problem with LEO satellites and GEO satellites as participants. The solution of the offloading decision is given by an improved two-sided many-to-one matching game algorithm with a preference for delay and energy cost. Simulation results show that by effectively offloading decision-making and power, bandwidth, and computing resource allocation, the proposed algorithm can achieve remarkable performance in satellite network edge computing. Meanwhile, compared with the existing methods, the proposed offloading algorithm can effectively reduce the running time of the offloading decision algorithm by 55.53% while the system overhead is reduced by 8.8%. Hai Fang, Yangyang Jia, Yuanle Wang |
ISNCC | 1 |
| 2021 | Oblivious Data Structure for Secure Multiple-Set Membership Testing
Yanjun An, Hai Fang |
WISA | 4 |
| 2019 | Expensive Inequality Constraints Handling Methods Suitable for Dynamic Surrogate-based OptimizationabstractIn modern engineering design optimization problems, high-fidelity analysis are always used for evaluating objectives and constraints, which might be quite expensive. Thus, efficient global optimization method should be developed to relieve the computational burden. This paper proposed a dynamic surrogate-based optimization (DSBO) using Kriging model, of which two criteria for selecting infill samples in refinement procedure are employed: maximizing expected improvement (EI) function and minimizing surrogate prediction. The DSBO are validated to be robust and efficient by six standard analytical tests. The inequality constraints are handled by three different means here: constraining EI function, penalizing surrogate prediction, and penalizing objective function. Analytical tests and an engineering optimization problem with inequality constraints are carried out. The results indicate that simultaneous constraining EI function and penalizing surrogate prediction is most efficient for DSBO, and there is no need of adjusting penalty factor. Hai Fang, Chunlin Gong |
CEC | 2 |
| 2019 | DeltaDou: Expert-level Doudizhu AI through Self-playabstractArtificial Intelligence has seen several breakthroughs in two-player perfect information game. Nevertheless, Doudizhu, a three-player imperfect information game, is still quite challenging. In this paper, we present a Doudizhu AI by applying deep reinforcement learning from games of self-play. The algorithm combines an asymmetric MCTS on nodes of information set of each player, a policy-value network that approximates the policy and value on each decision node, and inference on unobserved hands of other players by given policy. Our results show that self-play can significantly improve the performance of our agent in this multi-agent imperfect information game. Even starting with a weak AI, our agent can achieve human expert level after days of self-play and training. Qiqi Jiang, Kuangzheng Li, Boyao Du, Hai Fang |
IJCAI | 5 |
| 2014 | dcGOR: An R Package for Analysing Ontologies and Protein Domain AnnotationsabstractI introduce an open-source R package 'dcGOR' to provide the bioinformatics community with the ease to analyse ontologies and protein domain annotations, particularly those in the dcGO database. The dcGO is a comprehensive resource for protein domain annotations using a panel of ontologies including Gene Ontology. Although increasing in popularity, this database needs statistical and graphical support to meet its full potential. Moreover, there are no bioinformatics tools specifically designed for domain ontology analysis. As an add-on package built in the R software environment, dcGOR offers a basic infrastructure with great flexibility and functionality. It implements new data structure to represent domains, ontologies, annotations, and all analytical outputs as well. For each ontology, it provides various mining facilities, including: (i) domain-based enrichment analysis and visualisation; (ii) construction of a domain (semantic similarity) network according to ontology annotations; and (iii) significance analysis for estimating a contact (statistical significance) network. To reduce runtime, most analyses support high-performance parallel computing. Taking as inputs a list of protein domains of interest, the package is able to easily carry out in-depth analyses in terms of functional, phenotypic and diseased relevance, and network-level understanding. More importantly, dcGOR is designed to allow users to import and analyse their own ontologies and annotations on domains (taken from SCOP, Pfam and InterPro) and RNAs (from Rfam) as well. The package is freely available at CRAN for easy installation, and also at GitHub for version control. The dedicated website with reproducible demos can be found at http://supfam.org/dcGOR. Hai Fang |
PLoS Comput. Biol. | 1 |
| 2013 | Reversible Data Hiding for Multispectral Image with High Radiometric ResolutionabstractThis paper presents a reversible data hiding algorithm based on integer transform for high radiometric resolution multispectral images which have gradually been the main data source of spatial geographic information. Focusing on the characteristic of high radiometric resolution multispectral images, Reversible Karhunen-Loêve transform (RKLT) followed by integer wavelet transform (IWT) is applied to host images. Secret messages are embedded by a multilevel histogram modification technique in the transform domain. The Triangular Elementary Reversible Matrixes (TERM) are also embedded as secret messages into the input image which guarantees the host image can be perfectly restored after the embedded data extracted. Experimental results on multispectral images obtained from Quick Bird show that the proposed method has better visual quality and higher hiding capacity than the representative methods. Hai Fang |
ICIG | 1 |
| 2013 | A domain-centric solution to functional genomics via dcGO PredictorabstractBACKGROUND: Computational/manual annotations of protein functions are one of the first routes to making sense of a newly sequenced genome. Protein domain predictions form an essential part of this annotation process. This is due to the natural modularity of proteins with domains as structural, evolutionary and functional units. Sometimes two, three, or more adjacent domains (called supra-domains) are the operational unit responsible for a function, e.g. via a binding site at the interface. These supra-domains have contributed to functional diversification in higher organisms. Traditionally functional ontologies have been applied to individual proteins, rather than families of related domains and supra-domains. We expect, however, to some extent functional signals can be carried by protein domains and supra-domains, and consequently used in function prediction and functional genomics. RESULTS: Here we present a domain-centric Gene Ontology (dcGO) perspective. We generalize a framework for automatically inferring ontological terms associated with domains and supra-domains from full-length sequence annotations. This general framework has been applied specifically to primary protein-level annotations from UniProtKB-GOA, generating GO term associations with SCOP domains and supra-domains. The resulting 'dcGO Predictor', can be used to provide functional annotation to protein sequences. The functional annotation of sequences in the Critical Assessment of Function Annotation (CAFA) has been used as a valuable opportunity to validate our method and to be assessed by the community. The functional annotation of all completely sequenced genomes has demonstrated the potential for domain-centric GO enrichment analysis to yield functional insights into newly sequenced or yet-to-be-annotated genomes. This generalized framework we have presented has also been applied to other domain classifications such as InterPro and Pfam, and other ontologies such as mammalian phenotype and disease ontology. The dcGO and its predictor are available at http://supfam.org/SUPERFAMILY/dcGO including an enrichment analysis tool. CONCLUSIONS: As functional units, domains offer a unique perspective on function prediction regardless of whether proteins are multi-domain or single-domain. The 'dcGO Predictor' holds great promise for contributing to a domain-centric functional understanding of genomes in the next generation sequencing era. Hai Fang, Julian Gough |
BMC Bioinform. | 1 |
| 2004 | Complete Local Search for Propositional Satisfiability
Hai Fang, Wheeler Ruml |
AAAI | 1 |
| 2003 | A provably sound TAL for back-end optimizationabstractTyped assembly languages provide a way to generate machine-checkable safety proofs for machine-language programs. But the soundness proofs of most existing typed assembly languages are hand-written and cannot be machine-checked, which is worrisome for such large calculi. We have designed and implemented a low-level typed assembly language (LTAL) with a semantic model and established its soundness from the model. Compared to existing typed assembly languages, LTAL is more scalable and more secure; it has no macro instructions that hinder low-level optimizations such as instruction scheduling; its type constructors are expressive enough to capture dataflow information, support the compiler's choice of data representations and permit typed position-independent code; and its type-checking algorithm is completely syntax-directed.We have built a prototype system, based on Standard ML of New Jersey, that compiles most of core ML to Sparc code. We explain how we were able to make the untyped back end in SML/NJ preserve types during instruction selection and register allocation, without restricting low-level optimizations and without knowledge of any type system pervading the instruction selector and register allocator. Dinghao Wu, Andrew W. Appel, Hai Fang |
PLDI | 4 |