EDBT 2026 Demo / reviewers in the wild / expert
Haofeng Li
dblp:204/1032
· DBLP profile ↗
46ranked-venue papers
11as first author
39since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 14 since 2021Software engineering, systems software and programming languages · 13 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Context-Free Language Reachability via Efficient Relation ChainingabstractContext-free language (CFL) reachability is a fundamental framework widely used to model a variety of program analysis tasks, though it often suffers from inherent inefficiency due to its (sub)cubic time complexity. In this paper, we propose a novel perspective, relation chaining, which interprets CFL-reachability solving as the process of chaining labeled edges representing binary relations. This formulation exposes substantial derivation redundancy (in terms of frequent and repetitive chaining operations) arising from inefficient chaining strategies employed by existing approaches. To address this, we introduce Squid , a new algorithm that incorporates two simple yet effective chaining techniques—adaptive chaining and differential chaining—built upon an enhanced graph representation. We have implemented Squid as a standalone tool and evaluated it against two state-of-the-art CFL-reachability solvers and a leading Datalog solver across three key program analyses: field-sensitive alias analysis and context-sensitive value-flow analysis for C/C++, and field-sensitive points-to analysis for Java. Experimental results show that Squid substantially improves the scalability of CFL-reachability solving by effectively reducing a large portion of redundant chaining operations. Chenghang Shi, Haofeng Li, Jie Lu 0009, Lian Li 0002 |
Proc. ACM Program. Lang. | 2 |
| 2026 | Universal Scale Transformer for Histology Image SegmentationabstractHistology image segmentation is a critical prerequisite to pathological diagnosis. Accurate segmentation of these images can significantly aid physicians by facilitating quicker and more precise diagnostic decisions. A notable challenge in this area arises from the fact that different objects within histology images require segmentation at varying magnifications. However, most existing models are typically limited to performing segmentation at one pre-determined magnification. In this paper, we propose a novel universal scale transformer model (UniScaleFormer), which employs a scale-aware approach and integrates textual input to uniformly segment objects in histology images across various magnifications. Our method adopts an end-to-end architecture and utilizes a candidate mask query mechanism, specifically designed to identify and distinguish objects at different image scales. Moreover, we develop a Scale-Aware Module that enhances our network's ability to recognize the magnification level of input histology images, by using a scale query with extracted visual features. Then the scale query is integrated with mask queries, facilitating the incorporation of scale information. Experimental results demonstrate that the proposed method effectively achieves competitive results on various segmentation benchmarks at different magnifications. The code will be released at https://github.com/lhaof/UniScale. Junjia Huang, Haofeng Li, Yuanhuan Xiong, Guanbin Li |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Domain Generalized Medical Landmark Detection via Robust Boundary-Aware Pre-TrainingabstractIn recent years, deep learning has revenue in automated medical landmark detection. Nonetheless, prevailing research in this field predominantly addresses single-center scenarios or domain adaptation settings. In practical environments, the acquisition of multi-center data faces privacy concerns, coupled with the time-intensive and costly nature of data collection and annotation. These challenges substantially impede the broader application of deep learning-based medical landmark detection. To mitigate these issues, we propose a novel domain-generalized medical landmark detection framework that relies solely on single-center data for training. Considering the availability of numerous public medical segmentation datasets, we design a simple yet effective method that utilizes single-center segmentation to enhance the domain generalization capabilities of the landmark detection task. Specifically, we introduce a novel boundary-aware pre-training approach to focus the model on regions pertinent to landmarks. To further enhance the robustness and generalization capabilities during pre-training, we have derived a mixing loss term and proved its effectiveness in theory and practice. Extensive experiments conducted on our new domain generalization benchmark for medical landmark detection demonstrate the superiority of our approach. Haifan Gong, Haofeng Li |
AAAI | 4 |
| 2025 | Module-Aware Context Sensitive Pointer AnalysisabstractThe Java Platform Module System (JPMS) has found widespread applications since introduced in Java 9. However, existing pointer analyses fail to leverage the semantics of JPMS. This paper presents a novel module-aware approach to improving the performance of pointer analysis. We model the semantics of keywords provides and uses in JPMS to recover missing points-to relations. We design a module-aware context-sensitive analysis, which can propagate and apply critical contexts (by exploiting modularity) to balance precision and efficiency better. We have implemented our module-aware pointer analysis named MPA in TAI - E and conducted extensive experiments to compare it with standard object-sensitivity. The evaluation results demonstrate that MPA finds more reachable methods and enhances existing context-sensitive approaches, striking a good balance between efficiency and precision. MPA can increase the number of reachable methods up to 90.9× (lombok) under the same analysis. Performance-wise, MPA is nearly as fast as context-insensitivity for most benchmarks, while its precision is superior to that of 1-object-sensitivity on average. Haofeng Li, Chenghang Shi, Jie Lu 0009, Lian Li 0002 |
ICSE | 1 |
| 2025 | Intermediate Domain Alignment and Morphology Analogy for Patent-Product Image RetrievalabstractRecent advances in artificial intelligence have significantly impacted image retrieval tasks, yet Patent-Product Image Retrieval (PPIR) has received limited attention. PPIR, which retrieves patent images based on product images to identify potential infringements, presents unique challenges: (1) both product and patent images often contain numerous categories of artificial objects, but models pre-trained on standard datasets exhibit limited discriminative power to recognize some of those unseen objects; and (2) the significant domain gap between binary patent line drawings and colorful RGB product images further complicates similarity comparisons for product-patent pairs. To address these challenges, we formulate it as an open-set image retrieval task and introduce a comprehensive Patent-Product Image Retrieval Dataset (PPIRD) including a test set with 439 product-patent pairs, a retrieval pool of 727,921 patents, and an unlabeled pre-training set of 3,799,695 images. We further propose a novel Intermediate Domain Alignment and Morphology Analogy (IDAMA) strategy. IDAMA maps both image types to an intermediate sketch domain using edge detection to minimize the domain discrepancy, and employs a Morphology Analogy Filter to select discriminative patent images based on visual features via analogical reasoning. Extensive experiments on PPIRD demonstrate that IDAMA significantly outperforms baseline methods (+7.58 mAR) and offers valuable insights into domain mapping and representation learning for PPIR. (The PPIRD dataset is available at: \href{https://loslorien.github.io/idama-project/}{https://loslorien.github.io/idama-project/}) Haifan Gong, Xuanye Zhang, Ruifei Zhang, Anningzhe Gao, Haofeng Li |
NeurIPS | 9 |
| 2025 | Sparse diffusion models for multi-annotator medical image segmentation
Haofeng Li, Guanbin Li, Si-Qi Liu 0003 |
Knowl. Based Syst. | 2 |
| 2025 | Fast Client-Driven CFL-Reachability via Regularization-Based Graph SimplificationabstractContext-free language (CFL) reachability is a critical framework for various program analyses, widely adopted despite its computational challenges due to cubic or near-cubic time complexity. This often leads to significant performance degradation in client applications. Notably, in real-world scenarios, clients typically require reachability information only for specific source-to-sink pairs, offering opportunities for targeted optimization. We introduce MoYe, an effective regularization-based graph simplification technique designed to enhance the performance of client-driven CFL-reachability analyses by pruning non-contributing edges—those that do not participate in any specified CFL-reachable paths. MoYe employs a regular approximation to ensure exact reachability results for all designated node pairs and operates linearly with respect to the number of edges in the graph. This lightweight efficiency makes MoYe a valuable pre-processing step that substantially reduces both computational time and memory requirements for CFL-reachability analysis, outperforming a recent leading graph simplification approach. Our evaluations with two prominent CFL-reachability client applications demonstrate that MoYe can substantially improve performance and reduce resource consumption. Chenghang Shi, Dongjie He, Haofeng Li, Jie Lu 0009, Lian Li 0002, Jingling Xue |
Proc. ACM Program. Lang. | 3 |
| 2025 | Fetal Cerebellum Landmark Detection Based on 3D MRI: Method and BenchmarkabstractFetal cerebellum landmark detection is crucial for assessing fetal brain development. Although deep learning has become the standard for automatic landmark detection, most previous methods have focused on using 2D ultrasound or thick Magnetic Resonance Imaging (MRI). To improve accuracy, landmarks should be located on thin 3D MRIs. However, abnormal development, high noise, and fuzzy boundaries in 3D fetal brain images make traditional methods less effective for cerebellum landmark detection. To address this, we introduce the Anatomical Pseudo-label Guided Attention (APGA) network alongside a 3D MRI-based benchmark for fetal cerebellum landmark detection. During training, we use a shared encoder to extract image features and two decoders for landmark regression and anatomical pseudo-label segmentation. We design a Feature Decoupling Transformer (FDT) and embed it into the encoder to better calibrate the features for the two tasks. We only need the encoder, the FDT, and the landmark decoder during the inference phase. Extensive experiments on our proposed benchmark and out-of-domain test set have shown the effectiveness of our method. Our simulations also demonstrated that 3D biometrics are better than 2D biometrics. Haifan Gong, Huixian Liu, Yitao Wang 0001, Qiao Shi, Haofeng Li |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Boundary as the Bridge: Toward Heterogeneous Partially-Labeled Medical Image Segmentation and Landmark DetectionabstractMedical landmark detection and segmentation are crucial elements for computer-aided diagnosis and treatment. However, a common challenge arises because many datasets are exclusively annotated with either landmarks or segmentation masks: a situation we term the 'heterogeneous partially-labeled' problem. To address this, we propose a novel yet effective 'Boundary-as-Bridge' Loss (BaBLoss) that models the interplay between landmark detection and segmentation tasks. Specifically, our loss function is designed to maximize the correlation between the boundary distance map of the segmentation area and the heatmap deployed for landmark detection. Moreover, we introduce a prompt pipeline to use a segment anything model and landmarks to generate pseudo-segmentation labels for data with landmark annotation. To evaluate the effectiveness of our method, we collect and build two heterogeneous partially-labeled datasets on the brain and knee. Extensive experiments on these datasets using various backbone structures have shown the effectiveness of our method. Code is available at https://github.com/lhaof/HPL. Haifan Gong, Boyao Wan, Luoyao Kang, Haofeng Li |
IEEE Trans. Medical Imaging | 6 |
| 2025 | BCNet: Bronchus Classification via Structure Guided Representation LearningabstractCT-based bronchial tree analysis is a key step for the diagnosis of lung and airway diseases. However, the topology of bronchial trees varies across individuals, which presents a challenge to the automatic bronchus classification. To solve this issue, we propose the Bronchus Classification Network (BCNet), a structure-guided framework that exploits the segment-level topological information using point clouds to learn the voxel-level features. BCNet has two branches, a Point-Voxel Graph Neural Network (PV-GNN) for segment classification, and a Convolutional Neural Network (CNN) for voxel labeling. The two branches are simultaneously trained to learn topology-aware features for their shared backbone while it is feasible to run only the CNN branch for the inference. Therefore, BCNet maintains the same inference efficiency as its CNN baseline. Experimental results show that BCNet significantly exceeds the state-of-the-art methods by over 8.0% both on F1-score for classifying bronchus. Furthermore, we contribute BronAtlas: an open-access benchmark of bronchus imaging analysis with high-quality voxel-wise annotations of both anatomical and abnormal bronchial segments. The benchmark is available at https://osf.io/pskr9/?viewonly=94fa3d87274b4095ac9a4b88cc9a1341. Haifan Gong, Guanbin Li, Haofeng Li |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Instance-Aware Multi-Task Learning for Nuclei SegmentationabstractNuclei segmentation is critical for computational pathology analysis. Most previous methods employ pixel-wise classification or regression for automatic nuclei segmentation, without describing nucleus instances as individual entities at the feature level. To address the above limitation, we propose an instance-aware multi-task learning framework that strengthens a pixel-wise prediction branch with an instance-wise prediction branch. The instance-wise prediction branch leverages learnable cell-level queries, enabling the model to capture positional information and visual representations for individual nuclei. Concretely, we introduce an instance-disentangling feature learning module that effectively aligns the embeddings of the object-level queries with pixel-wise decoder features from the first branch. Further, we design a dual-branch unified post-processing algorithm that aggregates the complementary outputs of both branches for computing the instance segmentation results. Experimental results demonstrate that our framework achieves competitive performance on a wide range of nuclei segmentation benchmarks. The code and model weights are released in https://github.com/lhaof/IML. Wei Lou, Haofeng Li, Guanbin Li, Xiaoying Lou, Yuanhuan Xiong, Xusheng Wu |
IEEE Trans. Medical Imaging | 2 |
| 2024 | UniCell: Universal Cell Nucleus Classification via Prompt LearningabstractThe recognition of multi-class cell nuclei can significantly facilitate the process of histopathological diagnosis. Numerous pathological datasets are currently available, but their annotations are inconsistent. Most existing methods require individual training on each dataset to deduce the relevant labels and lack the use of common knowledge across datasets, consequently restricting the quality of recognition. In this paper, we propose a universal cell nucleus classification framework (UniCell), which employs a novel prompt learning mechanism to uniformly predict the corresponding categories of pathological images from different dataset domains. In particular, our framework adopts an end-to-end architecture for nuclei detection and classification, and utilizes flexible prediction heads for adapting various datasets. Moreover, we develop a Dynamic Prompt Module (DPM) that exploits the properties of multiple datasets to enhance features. The DPM first integrates the embeddings of datasets and semantic categories, and then employs the integrated prompts to refine image representations, efficiently harvesting the shared knowledge among the related cell types and data sources. Experimental results demonstrate that the proposed method effectively achieves the state-of-the-art results on four nucleus detection and classification benchmarks. Code and models are available at https://github.com/lhaof/UniCell Junjia Huang, Haofeng Li, Guanbin Li |
AAAI | 2 |
| 2024 | Cell Graph Transformer for Nuclei ClassificationabstractNuclei classification is a critical step in computer-aided diagnosis with histopathology images. In the past, various methods have employed graph neural networks (GNN) to analyze cell graphs that model inter-cell relationships by considering nuclei as vertices. However, they are limited by the GNN mechanism that only passes messages among local nodes via fixed edges. To address the issue, we develop a cell graph transformer (CGT) that treats nodes and edges as input tokens to enable learnable adjacency and information exchange among all nodes. Nevertheless, training the transformer with a cell graph presents another challenge. Poorly initialized features can lead to noisy self-attention scores and inferior convergence, particularly when processing the cell graphs with numerous connections. Thus, we further propose a novel topology-aware pretraining method that leverages a graph convolutional network (GCN) to learn a feature extractor. The pre-trained features may suppress unreasonable correlations and hence ease the finetuning of CGT. Experimental results suggest that the proposed cell graph transformer with topology-aware pretraining significantly improves the nuclei classification results, and achieves the state-of-the-art performance. Code and models are available at https://github.com/lhaof/CGT Wei Lou, Guanbin Li, Haofeng Li |
AAAI | 4 |
| 2024 | Detecting Broken Object-Level Authorization Vulnerabilities in Database-Backed ApplicationsabstractBroken object-level authorization (BOLA) vulnerabilities are among the most critical security risks facing database-backed applications. However, there is still a significant gap in our systematic understanding of these vulnerabilities. To bridge this gap, we conducted an in-depth study of 101 real-world BOLA vulnerabilities from opensource applications. Our study revealed the four most common object-level authorization models in database-backed application. Yongheng Huang, Chenghang Shi, Jie Lu 0009, Haofeng Li, Haining Meng, Lian Li 0002 |
CCS | 4 |
| 2024 | Boosting the Performance of Multi-Solver IFDS Algorithms with Flow-Sensitivity OptimizationsabstractThe IFDS (Inter-procedural, Finite, Distributive, Subset) algorithms are popularly used to solve a wide range of analysis problems. In particular, many interesting problems are formulated as multi-solver IFDS problems which expect multiple interleaved IFDS solvers to work together. For instance, taint analysis requires two IFDS solvers, one forward solver to propagate tainted data-flow facts, and one backward solver to solve alias relations at the same time. For such problems, large amount of additional data-flow facts need to be introduced for flow-sensitivity. This often leads to poor performance and scalability, as evident in our experiments and previous work. In this paper, we propose a novel approach to reduce the number of introduced additional data-flow facts while preserving flow-sensitivity and soundness. We have developed a new taint analysis tool, SADROID, and evaluated it on 1,228 open-source Android APPs. Evaluation results show that SADROID significantly outperforms FLowDROID (the state-of-the-art multi-solver IFDS taint analysis tool) without affecting precision and soundness: the run time performance is sped up by up to 17.89X and memory usage is optimized by up to 9X. Haofeng Li, Jie Lu 0009, Haining Meng, Liqing Cao, Lian Li 0002, Lin Gao 0002 |
CGO | 1 |
| 2024 | AutoWeb: Automatically Inferring Web Framework Semantics via Configuration Mutation
Haining Meng, Haofeng Li, Jie Lu 0009, Chenghang Shi, Liqing Cao, Lian Li 0002, Lin Gao 0002 |
ICECCS | 2 |
| 2024 | Intensity Confusion Matters: An Intensity-Distance Guided Loss For Bronchus SegmentationabstractAutomatic segmentation of the bronchial tree from CT imaging is important, as it provides structural information for disease diagnosis. Despite the merits of previous automatic bronchus segmentation methods, they have paied less attention to the issue we term as Intensity Confusion, wherein the intensity values of certain background voxels approach those of the foreground voxels within bronchi. Conversely, the intensity values of some foreground voxels are nearly identical to those of background voxels. This proximity in intensity values introduces significant challenges to neural network methodologies. To address the issue, we introduce a novel Intensity-Distance Guided loss function, which assigns adaptive weights to different image voxels for mining hard samples that cause the intensity confusion. The proposed loss estimates the voxel-level hardness of samples, on the basis of the following intensity and distance priors. We regard a voxel as a hard sample if it is in: (1) the background and has an intensity value close to the bronchus region; (2) the bronchus region and is of higher intensity than most voxels inside the bronchus; (3) the background region and at a short distance from the bronchus. Extensive experiments not only show the superiority of our method compared with the state-of-the-art methods, but also verify that tackling the intensity confusion issue helps to significantly improve bronchus segmentation. Project page: https://github.com/lhaof/ICM. Haifan Gong, Guanbin Li, Haofeng Li |
ICME | 8 |
| 2024 | Better Not Together: Staged Solving for Context-Free Language ReachabilityabstractContext-free language reachability (CFL-reachability) is a fundamental formulation for program analysis with many applications. CFL-reachability analysis is computationally expensive, with a slightly subcubic time complexity concerning the number of nodes in the input graph. This paper proposes staged solving: a new perspective on solving CFL-reachability. Our key observation is that the context-free grammar (CFG) of a CFL-based program analysis can be decomposed into (1) a smaller CFG, L, for matching parentheses, such as procedure calls/returns, field stores/loads, and (2) a regular grammar, R, capturing control/data flows. Instead of solving these two parts monolithically (as in standard algorithms), staged solving solves L-reachability and R-reachability in two distinct stages. In practice, L-reachability, though still context-free, involves only a small subset of edges, while R-reachability can be computed efficiently with close to quadratic complexity relative to the node size of the input graph. We implement our staged CFL-reachability solver, STG, and evaluate it using two clients: context-sensitive value-flow analysis and field-sensitive alias analysis. The empirical results demonstrate that STG achieves speedups of 861.59x and 4.1x for value-flow analysis and alias analysis on average, respectively, over the standard subcubic algorithm. Moreover, we also showcase that staged solving can help to significantly improve the performance of two state-of-the-art solvers, POCR and PEARL, by 74.82x (1.78x) and 37.66x (1.7x) for value-flow (alias) analysis, respectively. Chenghang Shi, Haofeng Li, Jie Lu 0009, Lian Li 0002 |
ISSTA | 2 |
| 2024 | Multi-modal Denoising Diffusion Pre-training for Whole-Slide Image ClassificationabstractWhole-slide image (WSI) classification methods play a crucial role in tumor diagnosis. Most of them use hematoxylin and eosin (H&E) stained images, while Immunohistochemistry (IHC) staining provides molecular markers and protein expression information that highlights cancer regions. However, obtaining IHC-stained images requires higher costs in practice. In this work, we propose a multi-modal denoising diffusion pre-training framework that harnesses the advantages of IHC staining to learn visual representations. The framework is trained with the H&E-to-IHC re-staining task and IHC-stained image reconstruction task, which helps capture the structural similarity and staining difference between two image modalities. The trained model can then provide IHC-guided features, by taking only H&E-stained images as inputs. Besides, we build a new class-constraint constrastive loss to achieve the semantic consistency between dual-modal features from our pre-training framework. To integrate with WSI classifiers based on multi-instance learning, we further propose a bag feature augmentation strategy to extend bags with the features extracted by our pre-trained model. Experimental results on three datasets show that our pre-training framework effectively improves WSI classification and surpasses the state-of-the-art pre-training approaches. Code and model are released via https://github.com/lhaof/MDDP Wei Lou, Guanbin Li, Haofeng Li |
ACM Multimedia | 4 |
| 2024 | Boosting the Performance of Alias-Aware IFDS Analysis with CFL-Based Environment TransformersabstractThe IFDS algorithm is pivotal in solving field-sensitive data-flow problems. However, its conventional use of access paths for field sensitivity leads to the generation of a large number of data-flow facts. This causes scalability challenges in larger programs, limiting its practical application in extensive codebases. In response, we propose a new field-sensitive technique that reinterprets the generation of access paths as a Context-Free Language (CFL) for field-sensitivity and formulates it as an IDE problem. This approach significantly reduces the number of data-flow facts generated and handled during the analysis, which is a major factor in performance degradation. To demonstrate the effectiveness of this approach, we developed a taint analysis tool, IDEDroid, in the IFDS/IDE framework. IDEDroid outperforms FlowDroid, an established IFDS-based taint analysis tool, in the analysis of 24 major Android apps while improving its precision (guaranteed theoretically). The speed improvement ranges from 2.1 × to 2,368.4 × , averaging at 222.0 × , with precision gains reaching up to 20.0 % (in terms of false positives reduced). This performance indicates that IDEDroid is substantially more effective in detecting information-flow leaks, making it a potentially superior tool for mobile app vetting in the market. Haofeng Li, Chenghang Shi, Jie Lu 0009, Lian Li 0002, Jingling Xue |
Proc. ACM Program. Lang. | 1 |
| 2024 | SENSE: Self-Evolving Learning for Self-Supervised Monocular Depth EstimationabstractSelf-supervised depth estimation methods can achieve competitive performance using only unlabeled monocular videos, but they suffer from the uncertainty of jointly learning depth and pose without any ground truths of both tasks. Supervised framework provides robust and superior performance but is limited by the scope of the labeled data. In this paper, we introduce SENSE, a novel learning paradigm for self-supervised monocular depth estimation that progressively evolves the prediction result using supervised learning, but without requiring labeled data. The key contribution of our approach stems from the novel use of the pseudo labels - the noisy depth estimation from the self-supervised methods. We surprisingly find that a fully supervised depth estimation network trained using the pseudo labels can produce even better results than its "ground truth". To push the envelope further, we then evolve the self-supervised backbone by replacing its depth estimation branch with that fully supervised network. Based on this idea, we devise a comprehensive training pipeline that alternatively enhances the two key branches (depth and pose estimation) of the self-supervised backbone network. Our proposed approach can effectively ease the difficulty of multi-task training in self-supervised depth estimation. Experimental results have shown that our proposed approach achieves state-of-the-art results on the KITTI dataset. Guanbin Li, Ricong Huang, Haofeng Li, Zunzhi You, Weikai Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Structure Embedded Nucleus Classification for Histopathology ImagesabstractNuclei classification provides valuable information for histopathology image analysis. However, the large variations in the appearance of different nuclei types cause difficulties in identifying nuclei. Most neural network based methods are affected by the local receptive field of convolutions, and pay less attention to the spatial distribution of nuclei or the irregular contour shape of a nucleus. In this paper, we first propose a novel polygon-structure feature learning mechanism that transforms a nucleus contour into a sequence of points sampled in order, and employ a recurrent neural network that aggregates the sequential change in distance between key points to obtain learnable shape features. Next, we convert a histopathology image into a graph structure with nuclei as nodes, and build a graph neural network to embed the spatial distribution of nuclei into their representations. To capture the correlations between the categories of nuclei and their surrounding tissue patterns, we further introduce edge features that are defined as the background textures between adjacent nuclei. Lastly, we integrate both polygon and graph structure learning mechanisms into a whole framework that can extract intra and inter-nucleus structural characteristics for nuclei classification. Experimental results show that the proposed framework achieves significant improvements compared to the previous methods. Code and data are made available via https://github.com/lhaof/SENC. Wei Lou, Guanbin Li, Xiaoying Lou, Chenghang Li, Feng Gao 0023, Haofeng Li |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Generic Sensitivity: Generics-Guided Context Sensitivity for Pointer AnalysisabstractGeneric programming has found widespread application in object-oriented languages like Java. However, existing context-sensitive pointer analyses fail to leverage the benefits of generic programming. This paper introducesgeneric sensitivity, a new context customization scheme targeting generics. We design our context customization scheme in such a way that generic instantiation sites, i.e., locations instantiating generic classes/methods with concrete types, are always preserved as key context elements. This is realized by augmenting contexts with a type variable lookup map, which is efficiently generated in a context-sensitive manner throughout the analysis process. We have implemented various variants of generic-sensitive analysis in WALA and conducted extensive experiments to compare it with state-of-the-art approaches, including both traditional and selective context-sensitivity methods. The evaluation results demonstrate that generic sensitivity effectively enhances existing context-sensitivity approaches, striking a new balance between efficiency and precision. For instance, it enables a 1-object-sensitive analysis to achieve overall better precision compared to a 2-object-sensitive analysis, with an average speedup of 12.6 times (up to 62 times). Haofeng Li, Tian Tan 0001, Yue Li 0006, Jie Lu 0009, Haining Meng, Liqing Cao, Yongheng Huang, Lian Li 0002, Lin Gao 0002, Peng Di, ChenXi Cui |
IEEE Trans. Software Eng. | 1 |
| 2024 | Pearl: A Multi-Derivation Approach to Efficient CFL-Reachability SolvingabstractContext-free language (CFL) reachability is a fundamental framework for formulating program analyses. CFL-reachability analysis works on top of an edge-labeled graph by deriving reachability relations and adding them as labeled edges to the graph. Existing CFL-reachability algorithms typically adopt a single-reachability relation derivation (SRD) strategy, i.e., one reachability relation is derived at a time. Unfortunately, this strategy can lead to redundancy, hindering the efficiency of the analysis. To address this problem, this paper proposesPearl, amulti-derivationapproach that reduces derivation redundancy for CFL-reachability solving, which significantly improves the efficiency of CFL-reachability analysis. Our key insight is that multiple edges can be simultaneously derived via batch propagation of reachability relations. We also tailor our multi-derivation approach to tackle transitive relations that frequently arise when solving CFL-reachability. Specifically, we present a highly efficient transitive-aware variant,PearlPG, which enhancesPearlwithpropagation graphs, a lightweight but effective graph representation, to further diminish redundant derivations. We evaluate the performance of our approach on two clients, i.e., context-sensitive value-flow analysis and field-sensitive alias analysis for C/C++. By eliminating a large amount of redundancy, our approach outperforms two baselines including the standard CFL-reachability algorithm and a state-of-the-art solverPocrspecialized for fast transitivity solving. In particular, the empirical results demonstrate that, for value-flow analysis and alias analysis respectively,PearlPGruns 3.09$\times$faster on average (up to 4.44$\times$) and 2.25$\times$faster on average (up to 3.31$\times$) thanPocr, while also consuming less memory. Chenghang Shi, Haofeng Li, Yulei Sui, Jie Lu 0009, Lian Li 0002, Jingling Xue |
IEEE Trans. Software Eng. | 2 |
| 2023 | Affine-Consistent Transformer for Multi-Class Cell Nuclei DetectionabstractMulti-class cell nuclei detection is a fundamental prerequisite in the diagnosis of histopathology. It is critical to efficiently locate and identify cells with diverse morphology and distributions in digital pathological images. Most existing methods take complex intermediate representations as learning targets and rely on inflexible post-refinements while paying less attention to various cell density and fields of view. In this paper, we propose a novel Affine-Consistent Transformer (AC-Former), which directly yields a sequence of nucleus positions and is trained collaboratively through two sub-networks, a global and a local network. The local branch learns to infer distorted input images of smaller scales while the global network outputs the large-scale predictions as extra supervision signals. We further introduce an Adaptive Affine Transformer (AAT) module, which can automatically learn the key spatial transformations to warp original images for local network training. The AAT module works by learning to capture the transformed image regions that are more valuable for training the model. Experimental results demonstrate that the proposed method significantly outperforms existing state-of-the-art algorithms on various benchmarks. Junjia Huang, Haofeng Li, Guanbin Li |
ICCV | 2 |
| 2023 | Two Birds with One Stone: Multi-Derivation for Fast Context-Free Language Reachability AnalysisabstractContext-free language (CFL) reachability is a fundamental framework for formulating program analyses. CFL-reachability analysis works on top of an edge-labeled graph by deriving reachability relations and adding them as labeled edges to the graph. Existing CFL-reachability algorithms typically adopt a single-reachability relation derivation (SRD) strategy, i.e., one reachability relation is derived at a time. Unfortunately, this strategy can lead to redundancy, hindering the efficiency of the analysis. To address this problem, this paper proposes Pearl, a multi-derivation approach that reduces derivation redundancy for transitive relations that frequently arise when solving reachability relations, significantly improving the efficiency of CFL-reachability analysis. Our key insight is that multiple edges involving transitivity can be simultaneously derived via batch propagation of reachability relations on the transitivity-aware subgraphs that are induced from the original edge-labeled graph. We evaluate the performance of Pearl on two clients, i.e., context-sensitive value-flow analysis and field-sensitive alias analysis for C/C++. By eliminating a large amount of redundancy, Pearl achieves average speedups of 82.73x for value-flow analysis and 155.26x for alias analysis over the standard CFL-reachability algorithm. The comparison with Pocr, a state-of-the-art CFL-reachability solver, shows that Pearl runs 10.1x (up to 29.2x) and 2.37x (up to 4.22x) faster on average respectively for value-flow analysis and alias analysis with less consumed memory. Chenghang Shi, Haofeng Li, Yulei Sui, Jie Lu 0009, Lian Li 0002, Jingling Xue |
ASE | 2 |
| 2023 | Prompt-Based Grouping Transformer for Nucleus Detection and Classification
Junjia Huang, Haofeng Li, Weijun Sun, Guanbin Li |
MICCAI (8) | 2 |
| 2023 | Visual-Attribute Prompt Learning for Progressive Mild Cognitive Impairment Prediction
Luoyao Kang, Haifan Gong, Haofeng Li |
MICCAI (5) | 4 |
| 2023 | ASC: Appearance and Structure Consistency for Unsupervised Domain Adaptation in Fetal Brain MRI Segmentation
Zihang Xu, Haifan Gong, Haofeng Li |
MICCAI (7) | 4 |
| 2023 | Diffusion-Based Data Augmentation for Nuclei Image Segmentation
Guanbin Li, Wei Lou, Si-Qi Liu 0003, Haofeng Li |
MICCAI (8) | 7 |
| 2023 | Unbiased curriculum learning enhanced global-local graph neural network for protein thermodynamic stability predictionabstractMOTIVATION: Proteins play crucial roles in biological processes, with their functions being closely tied to thermodynamic stability. However, measuring stability changes upon point mutations of amino acid residues using physical methods can be time-consuming. In recent years, several computational methods for protein thermodynamic stability prediction (PTSP) based on deep learning have emerged. Nevertheless, these approaches either overlook the natural topology of protein structures or neglect the inherent noisy samples resulting from theoretical calculation or experimental errors. RESULTS: We propose a novel Global-Local Graph Neural Network powered by Unbiased Curriculum Learning for the PTSP task. Our method first builds a Siamese graph neural network to extract protein features before and after mutation. Since the graph's topological changes stem from local node mutations, we design a local feature transformation module to make the model focus on the mutated site. To address model bias caused by noisy samples, which represent unavoidable errors from physical experiments, we introduce an unbiased curriculum learning method. This approach effectively identifies and re-weights noisy samples during the training process. Extensive experiments demonstrate that our proposed method outperforms advanced protein stability prediction methods, and surpasses state-of-the-art learning methods for regression prediction tasks. AVAILABILITY AND IMPLEMENTATION: All code and data is available at https://github.com/haifangong/UCL-GLGNN. Haifan Gong, Chenhe Dong, Yue Wang 0101, Guanqi Chen, Bilin Liang, Haofeng Li, Lanxuan Liu, Jie Xu 0068, Guanbin Li |
Bioinform. | 7 |
| 2023 | Which Pixel to Annotate: A Label-Efficient Nuclei Segmentation FrameworkabstractRecently deep neural networks, which require a large amount of annotated samples, have been widely applied in nuclei instance segmentation of H&E stained pathology images. However, it is inefficient and unnecessary to label all pixels for a dataset of nuclei images which usually contain similar and redundant patterns. Although unsupervised and semi-supervised learning methods have been studied for nuclei segmentation, very few works have delved into the selective labeling of samples to reduce the workload of annotation. Thus, in this paper, we propose a novel full nuclei segmentation framework that chooses only a few image patches to be annotated, augments the training set from the selected samples, and achieves nuclei segmentation in a semi-supervised manner. In the proposed framework, we first develop a novel consistency-based patch selection method to determine which image patches are the most beneficial to the training. Then we introduce a conditional single-image GAN with a component-wise discriminator, to synthesize more training samples. Lastly, our proposed framework trains an existing segmentation model with the above augmented samples. The experimental results show that our proposed method could obtain the same-level performance as a fully-supervised baseline by annotating less than 5% pixels on some benchmarks. Wei Lou, Haofeng Li, Guanbin Li, Xiaoguang Han 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Detecting Missing-Permission-Check Vulnerabilities in Distributed Cloud SystemsabstractMissing- Permission-Check (MPC) vulnerability is a type of bug where permission checks are not enforced for privileged operations. MPC vulnerability is prevalent and can cause severe security impacts. This paper proposes the first tool to detect MPC vulnerabilities in distributed cloud systems. We conduct an in-depth study of 95 real-world MPC vulnerabilities and our findings motivate a new tool named MPChecker. The tool introduces a combined log-static analysis to automatically identify privileged operations by inferring variables representing user owned data and critical system states, whose accesses need to be protected. We have evaluated MPChecker with 6 popular distributed systems. The tool reports 44 new vulnerabilities, and 43 of them have been confirmed and labeled as critical bugs. Moreover, 1 bug is particular dangerous and the developers requested to keep it undisclosed. Jie Lu 0009, Haofeng Li, Lian Li 0002 |
CCS | 2 |
| 2022 | Attentive Symmetric Autoencoder for Brain MRI Segmentation
Junjia Huang, Haofeng Li, Guanbin Li |
MICCAI (5) | 2 |
| 2022 | Generic sensitivity: customizing context-sensitive pointer analysis for genericsabstractGeneric programming has been extensively used in object-oriented programs such as Java. However, existing context-sensitive pointer analyses perform poorly in analyzing generics. This paper introduces generic sensitivity, a new context customization scheme targeting generics. We design our context customization scheme in such a way that generic instantiation sites, i.e., locations instantiating generic classes/methods with concrete types, are always preserved as key context elements. This is realized by augmenting contexts with a type variable lookup map, which is efficiently updated during the analysis in a context-sensitive manner. Haofeng Li, Jie Lu 0009, Haining Meng, Liqing Cao, Yongheng Huang, Lian Li 0002, Lin Gao 0002 |
ESEC/SIGSOFT FSE | 1 |
| 2021 | Scaling Up the IFDS Algorithm with Efficient Disk-Assisted ComputingabstractThe IFDS algorithm can be memory-intensive, requiring a memory budget of more than 100 GB of RAM for some applications. The large memory requirements significantly restrict the deployment of IFDS-based tools in practise. To improve this, we propose a disk-assisted solution that drastically reduces the memory requirements of traditional IFDS solvers. Our solution saves memory by 1) recomputing instead of memorizing intermediate analysis data, and 2) swapping in-memory data to disk when memory usages reach a threshold. We implement sophisticated scheduling schemes to swap data between memory and disks efficiently. We have developed a new taint analysis tool, DiskDroid, based on our disk-assisted IFDS solver. Compared to FlowDroid, a state-of-the-art IFDS-based taint analysis tool, for a set of 19 apps which take from 10 to 128 GB of RAM by FlowDroid, DiskDroid can analyze them with less than 10GB of RAM at a slight performance improvement of 8.6%. In addition, for 21 apps requiring more than 128GB of RAM by FlowDroid, DiskDroid can analyze each app in 3 hours, under the same memory budget of 10GB. This makes the tool deployable to normal desktop environments. We make the tool publicly available at https://github.com/HaofLi/DiskDroid. Haofeng Li, Haining Meng, Hengjie Zheng, Liqing Cao, Jie Lu 0009, Lian Li 0002, Lin Gao 0002 |
CGO | 1 |
| 2021 | Robust Real-World Image Super-Resolution against Adversarial AttacksabstractRecently deep neural networks (DNNs) have achieved significant success in real-world image super-resolution (SR). However, adversarial image samples with quasi-imperceptible noises could threaten deep learning SR models. In this paper, we propose a robust deep learning framework for real-world SR that randomly erases potential adversarial noises in the frequency domain of input images or features. The rationale is that on the SR task clean images or features have a different pattern from the attacked ones in the frequency domain. Observing that existing adversarial attacks usually add high-frequency noises to input images, we introduce a novel random frequency mask module that blocks out high-frequency components possibly containing the harmful perturbations in a stochastic manner. Since the frequency masking may not only destroys the adversarial perturbations but also affects the sharp details in a clean image, we further develop an adversarial sample classifier based on the frequency domain of images to determine if applying the proposed mask module. Based on the above ideas, we devise a novel real-world image SR framework that combines the proposed frequency mask modules and the proposed adversarial classifier with an existing super-resolution backbone network. Experiments show that our proposed method is more insensitive to adversarial attacks and presents more stable SR results than existing models and defenses. Jiutao Yue, Haofeng Li, Pengxu Wei, Guanbin Li, Liang Lin 0004 |
ACM Multimedia | 2 |
| 2021 | SSMD: Semi-Supervised medical image detection with adaptive consistency and heterogeneous perturbation
Chengdi Wang, Haofeng Li, Shu Zhang 0001, Weimin Li 0003, Yizhou Yu |
Medical Image Anal. | 3 |
| 2021 | Depthwise Nonlocal Module for Fast Salient Object Detection Using a Single ThreadabstractRecently, deep convolutional neural networks have achieved significant success in salient object detection. However, existing state-of-the-art methods require high-end GPUs to achieve real-time performance, which makes it hard to adapt to low cost or portable devices. Although generic network architectures have been proposed to speed up inference on mobile devices, they are tailored to the task of image classification or semantic segmentation, and struggle to capture intrachannel and interchannel correlations that are essential for contrast modeling in salient object detection. Motivated by the above observations, we design a new deep-learning algorithm for fast salient object detection. The proposed algorithm for the first time achieves competitive accuracy and high inference efficiency simultaneously with a single CPU thread. Specifically, we propose a novel depthwise nonlocal module (DNL), which implicitly models contrast via harvesting intrachannel and interchannel correlations in a self-attention manner. In addition, we introduce a depthwise nonlocal network architecture that incorporates both DNLs module and inverted residual blocks. The experimental results show that our proposed network attains very competitive accuracy on a wide range of salient object detection datasets while achieving state-of-the-art efficiency among all existing deep-learning-based algorithms. Haofeng Li, Guanbin Li, Guanqi Chen, Liang Lin 0004, Yizhou Yu |
IEEE Trans. Cybern. | 1 |
| 2020 | ROSA: Robust Salient Object Detection Against Adversarial AttacksabstractRecently, salient object detection has witnessed remarkable improvement owing to the deep convolutional neural networks which can harvest powerful features for images. In particular, the state-of-the-art salient object detection methods enjoy high accuracy and efficiency from fully convolutional network (FCN)-based frameworks which are trained from end to end and predict pixel-wise labels. However, such framework suffers from adversarial attacks which confuse neural networks via adding quasi-imperceptible noises to input images without changing the ground truth annotated by human subjects. To our knowledge, this paper is the first one that mounts successful adversarial attacks on salient object detection models and verifies that adversarial samples are effective on a wide range of existing methods. Furthermore, this paper proposes a novel end-to-end trainable framework to enhance the robustness for arbitrary FCN-based salient object detection models against adversarial attacks. The proposed framework adopts a novel idea that first introduces some new generic noise to destroy adversarial perturbations, and then learns to predict saliency maps for input images with the introduced noise. Specifically, our proposed method consists of a segment-wise shielding component, which preserves boundaries and destroys delicate adversarial noise patterns and a context-aware restoration component, which refines saliency maps through global contrast modeling. The experimental results suggest that our proposed framework improves the performance significantly for state-of-the-art models on a series of datasets. Haofeng Li, Guanbin Li, Yizhou Yu |
IEEE Trans. Cybern. | 1 |
| 2020 | Online Alternate Generator Against Adversarial AttacksabstractThe field of computer vision has witnessed phenomenal progress in recent years partially due to the development of deep convolutional neural networks. However, deep learning models are notoriously sensitive to adversarial examples which are synthesized by adding quasi-perceptible noises on real images. Some existing defense methods require to re-train attacked target networks and augment the train set via known adversarial attacks, which is inefficient and might be unpromising with unknown attack types. To overcome the above issues, we propose a portable defense method, online alternate generator, which does not need to access or modify the parameters of the target networks. The proposed method works by online synthesizing another image from scratch for an input image, instead of removing or destroying adversarial noises. To avoid pretrained parameters exploited by attackers, we alternately update the generator and the synthesized image at the inference stage. Experimental results demonstrate that the proposed defensive scheme and method outperforms a series of state-of-the-art defending models against gray-box adversarial attacks. Haofeng Li, Yirui Zeng, Guanbin Li, Liang Lin 0004, Yizhou Yu |
IEEE Trans. Image Process. | 1 |
| 2019 | Non-Local Context Encoder: Robust Biomedical Image Segmentation against Adversarial AttacksabstractRecent progress in biomedical image segmentation based on deep convolutional neural networks (CNNs) has drawn much attention. However, its vulnerability towards adversarial samples cannot be overlooked. This paper is the first one that discovers that all the CNN-based state-of-the-art biomedical image segmentation models are sensitive to adversarial perturbations. This limits the deployment of these methods in safety-critical biomedical fields. In this paper, we discover that global spatial dependencies and global contextual information in a biomedical image can be exploited to defend against adversarial attacks. To this end, non-local context encoder (NLCE) is proposed to model short- and long-range spatial dependencies and encode global contexts for strengthening feature activations by channel-wise attention. The NLCE modules enhance the robustness and accuracy of the non-local context encoding network (NLCEN), which learns robust enhanced pyramid feature representations with NLCE modules, and then integrates the information across different levels. Experiments on both lung and skin lesion segmentation datasets have demonstrated that NLCEN outperforms any other state-of-the-art biomedical image segmentation methods against adversarial attacks. In addition, NLCE modules can be applied to improve the robustness of other CNN-based biomedical image segmentation methods. Sibei Yang, Guanbin Li, Haofeng Li, HuiYou Chang, Yizhou Yu |
AAAI | 4 |
| 2019 | Motion Guided Attention for Video Salient Object DetectionabstractVideo salient object detection aims at discovering the most visually distinctive objects in a video. How to effectively take object motion into consideration during video salient object detection is a critical issue. Existing state-of-the-art methods either do not explicitly model and harvest motion cues or ignore spatial contexts within optical flow images. In this paper, we develop a multi-task motion guided video salient object detection network, which learns to accomplish two sub-tasks using two sub-networks, one sub-network for salient object detection in still images and the other for motion saliency detection in optical flow images. We further introduce a series of novel motion guided attention modules, which utilize the motion saliency sub-network to attend and enhance the sub-network for still images. These two sub-networks learn to adapt to each other by end-to-end training. Experimental results demonstrate that the proposed method significantly outperforms existing state-of-the-art algorithms on a wide range of benchmarks. We hope our simple and effective approach will serve as a solid baseline and help ease future research in video salient object detection. Code and models will be made available. Haofeng Li, Guanqi Chen, Guanbin Li, Yizhou Yu |
ICCV | 1 |
| 2019 | Performance-Boosting Sparsification of the IFDS Algorithm with Applications to Taint AnalysisabstractThe IFDS algorithm can be compute-and memoryintensive for some large programs, often running for a long time (more than expected) or terminating prematurely after some time and/or memory budgets have been exhausted. In the latter case, the corresponding IFDS data-flow analyses may suffer from false negatives and/or false positives. To improve this, we introduce a sparse alternative to the traditional IFDS algorithm. Instead of propagating the data-flow facts across all the program points along the program’s (interprocedural) control flow graph, we propagate every data-flow fact directly to its next possible use points along its own sparse control flow graph constructed on the fly, thus reducing significantly both the time and memory requirements incurred by the traditional IFDS algorithm. In our evaluation, we compare FLOWDROID, a taint analysis performed by using the traditional IFDS algorithm, with our sparse incarnation, SPARSEDROID, on a set of 40 Android apps selected. For the time budget (5 hours) and memory budget (220GB) allocated per app, SPARSEDROID can run every app to completion but FLOWDROID terminates prematurely for 9 apps, resulting in an average speedup of 22.0x. This implies that when used as a market-level vetting tool, SPARSEDROID can finish analyzing these 40 apps in 2.13 hours (by issuing 228 leak warnings) while FLOWDROID manages to analyze only 30 apps in the same time period (by issuing only 147 leak warnings). Dongjie He, Haofeng Li, Lei Wang 0004, Haining Meng, Hengjie Zheng, Jie Liu 0020, Shuangwei Hu, Lian Li 0002, Jingling Xue |
ASE | 2 |
| 2019 | Context-Aware Semantic InpaintingabstractIn recent times, image inpainting has witnessed rapid progress due to the generative adversarial networks (GANs) that are able to synthesize realistic contents. However, most existing GAN-based methods for semantic inpainting apply an auto-encoder architecture with a fully connected layer, which cannot accurately maintain spatial information. In addition, the discriminator in existing GANs struggles to comprehend high-level semantics within the image context and yields semantically consistent content. Existing evaluation criteria are biased toward blurry results and cannot well characterize edge preservation and visual authenticity in the inpainting results. In this paper, we propose an improved GAN to overcome the aforementioned limitations. Our proposed GAN-based framework consists of a fully convolutional design for the generator which helps to better preserve spatial structures and a joint loss function with a revised perceptual loss to capture high-level semantics in the context. Furthermore, we also introduce two novel measures to better assess the quality of image inpainting results. The experimental results demonstrate that our method outperforms the state-of-the-art under a wide range of criteria. Haofeng Li, Guanbin Li, Liang Lin 0004, Hongchuan Yu, Yizhou Yu |
IEEE Trans. Cybern. | 1 |
| 2019 | Harvesting Visual Objects from Internet Images via Deep-Learning-Based Objectness AssessmentabstractThe collection of internet images has been growing in an astonishing speed. It is undoubted that these images contain rich visual information that can be useful in many applications, such as visual media creation and data-driven image synthesis. In this article, we focus on the methodologies for building a visual object database from a collection of internet images. Such database is built to contain a large number of high-quality visual objects that can help with various data-driven image applications. Our method is based on dense proposal generation and objectness-based re-ranking. A novel deep convolutional neural network is designed for the inference of proposal objectness , the probability of a proposal containing optimally located foreground object. In our work, the objectness is quantitatively measured in regard of completeness and fullness , reflecting two complementary features of an optimal proposal: a complete foreground and relatively small background. Our experiments indicate that object proposals re-ranked according to the output of our network generally achieve higher performance than those produced by other state-of-the-art methods. As a concrete example, a database of over 1.2 million visual objects has been built using the proposed method, and has been successfully used in various data-driven image applications. Kan Wu 0005, Guanbin Li, Haofeng Li, Jian J. Zhang 0001, Yizhou Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |