VLDB 2026 Research / reviewers in the wild / expert
Haozhe Lin
dblp:226/0573
· DBLP profile ↗
18ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-3707-3575ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GigaMoE: Sparsity-Guided Mixture of Experts for Efficient Gigapixel Object DetectionabstractObject detection in High-Resolution Wide (HRW) shots, or gigapixel images, presents unique challenges due to extreme object sparsity and vast scale variations. State-of-the-art methods like SparseFormer have pioneered sparse processing by selectively focusing on important regions, yet they apply a uniform computational model to all selected regions, overlooking their intrinsic complexity differences. This leads to a suboptimal trade-off between performance and efficiency. In this paper, we introduce GigaMoE, a novel backbone architecture that pioneers adaptive computation for this domain by replacing the standard Feed-Forward Networks (FFNs) with a Mixture-of-Experts (MoE) module. Our architecture first employs a shared expert to provide a robust feature baseline for all selected regions. Upon this foundation, our core innovation---a novel Sparsity-Guided Routing mechanism---insightfully repurposes importance scores from the sparse backbone to provide a "computational bonus,'' dynamically engaging a variable number of specialized experts based on content complexity. The entire system is trained efficiently via a loss-free load-balancing technique, eliminating the need for cumbersome auxiliary losses. Extensive experiments show that GigaMoE sets a new state-of-the-art on the PANDA benchmark, improving detection accuracy by 1.1% over SparseFormer while simultaneously reducing the computational cost (FLOPs) by a remarkable 32.3%. Wenxi Li, Yuetong Wang, Chenyang Lyu, Haozhe Lin, Guiguang Ding |
AAAI | 5 |
| 2025 | Conditional Diffusion Anomaly Modeling on GraphsabstractGraph anomaly detection (GAD) has become a critical research area, with successful applications in financial fraud and telecommunications. Traditional Graph Neural Networks (GNNs) face significant challenges: at the topology level, they suffer from over-smoothing that averages out anomalous signals; at the feature level, discriminative models struggle when fraudulent nodes obfuscate their features to evade detection. In this paper, we propose a Conditional Graph Anomaly Diffusion Model (CGADM) that addresses these issues through the iterative refinement and denoising reconstruction properties of diffusion models. Our approach incorporates a prior-guided diffusion process that injects a pre-trained conditional anomaly estimator into both forward and reverse diffusion chains, enabling more accurate anomaly detection. For computational efficiency on large-scale graphs, we introduce a prior confidence-aware mechanism that adaptively determines the number of reverse denoising steps based on prior confidence. Experimental results on benchmark datasets demonstrate that CGADM achieves state-of-the-art performance while maintaining significant computational advantages for large-scale graph applications. Chunyu Wei, Haozhe Lin, Yueguo Chen, Yunhai Wang |
NeurIPS | 2 |
| 2024 | GigaTraj: Predicting Long-term Trajectories of Hundreds of Pedestrians in Gigapixel Complex ScenesabstractPedestrian trajectory prediction is a well-established task with significant recent advancements. However, existing datasets are unable to fulfill the demand for studying minute-level long-term trajectory prediction, mainly due to the lack of high-resolution trajectory observation in the wide field of view (FoV). To bridge this gap, we introduce a novel dataset named GigaTraj, featuring videos capturing a wide FoV with ~4 ×104m2and high-resolution imagery at the gigapixel level. Furthermore, GigaTraj in-cludes comprehensive annotations such as bounding boxes, identity associations, world coordinates, group/interaction relationships, and scene semantics. Leveraging these multimodal annotations, we evaluate and validate the state-of-the-art approaches for minute-level long-term trajectory prediction in large-scale scenes. Extensive experiments and analyses have revealed that long-term prediction for pedestrian trajectories presents numerous challenges, indicating a vital new direction for trajectory research. The dataset is available at WWW.gigavision ai. Haozhe Lin, Chunyu Wei, Yunqi Zhao, Shanglong Li, Lu Fang 0001 |
CVPR | 1 |
| 2024 | When Visual Grounding Meets Gigapixel-Level Large-Scale Scenes: Benchmark and ApproachabstractVisual grounding refers to the process of associating natural language expressions with corresponding regions within an image. Existing benchmarks for visual grounding primarily operate within small-scale scenes with a few objects. Nevertheless, recent advances in imaging technology have enabled the acquisition of gigapixel-level images, providing high-resolution details in large-scale scenes containing numerous objects. To bridge this gap between imaging and computer vision benchmarks and make grounding more practically valuable, we introduce a novel dataset, named GigaGrounding, designed to challenge visual grounding models in gigapixel-level large-scale scenes. We extensively analyze and compare the dataset with existing benchmarks, demonstrating that GigaGrounding presents unique challenges such as large-scale scene understanding, gigapixel-level resolution, significant variations in object scales, and the “multi-hop expressions”. Furthermore, we introduced a simple yet effective grounding approach, which employs a “glance-to-zoom-in” paradigm and exhibits enhanced capabilities for addressing the GigaGrounding task. The dataset is available at www.gigavision.ai. M. Tao, Haozhe Lin, Heyuan Wang 0001, Lu Fang 0001 |
CVPR | 3 |
| 2024 | SaccadeMOT: Enhancing Object Detection and Tracking in Gigapixel Images via Scale-Aware Density EstimationabstractThe proliferation of gigapixel imaging has ushered in unprecedented challenges in object detection and tracking due to the intense computational demands. Previous deep learning approaches, often tailored for megapixel images, fall short in addressing the unique complexities presented by the gigapixel level. To bridge this gap, we introduce SaccadeMOT, a novel architecture designed for efficient gigapixel-level multi-object tracking. Based on our observations of density map regression in crowd counting and small object detection in object detection tasks, we propose a novel gigapixel detection paradigm that combines the strengths of both approaches. Firstly, the “saccade” stage swiftly identifies regions likely containing objects, followed by the “gaze” stage that refines the detection within these areas. This strategic region selection is complemented by a robust tracking mechanism that combines head and body tracking, enhancing accuracy in environments with potential occlusions. Validated on the PANDA dataset, SaccadeMOT not only demonstrates an 13× speed improvement over existing state-of-the-art tracker BotSORT but also exhibits promising applications in gigapixel-level pathology analysis, particularly in Whole Slide Imaging (WSI). This approach sets a new benchmark for handling super high-resolution images, offering significant advancements in both the speed and precision of object tracking technologies. Wenxi Li, Ruxin Zhang, Haozhe Lin, Chao Ma 0004, Xiaokang Yang 0001 |
ECAI | 3 |
| 2024 | SparseFormer: Detecting Objects in HRW Shots via Sparse Vision TransformerabstractRecent years have seen an increase in the use of gigapixel-level image and video capture systems and benchmarks with high-resolution wide (HRW) shots. However, unlike close-up shots in the MS COCO dataset, the higher resolution and wider field of view raise unique challenges, such as extreme sparsity and huge scale changes, causing existing close-up detectors inaccuracy and inefficiency. In this paper, we present a novel model-agnostic sparse vision transformer, dubbed SparseFormer, to bridge the gap of object detection between close-up and HRW shots. The proposed SparseFormer selectively uses attentive tokens to scrutinize the sparsely distributed windows that may contain objects. In this way, it can jointly explore global and local attention by fusing coarse- and fine-grained features to handle huge scale changes. SparseFormer also benefits from a novel Cross-slice non-maximum suppression (C-NMS) algorithm to precisely localize objects from noisy windows and a simple yet effective multi-scale strategy to improve accuracy. Extensive experiments on two HRW benchmarks, PANDA and DOTA-v1.0, demonstrate that the proposed SparseFormer significantly improves detection accuracy (up to 5.8%) and speed (up to 3x) over the state-of-the-art approaches. Wenxi Li, Jilai Zheng, Haozhe Lin, Chao Ma 0004, Lu Fang 0001, Xiaokang Yang 0001 |
ACM Multimedia | 4 |
| 2024 | SaccadeDet: A Novel Dual-Stage Architecture for Rapid and Accurate Detection in Gigapixel Images
Wenxi Li, Ruxin Zhang, Haozhe Lin, Chao Ma 0004, Xiaokang Yang 0001 |
ECML/PKDD (2) | 3 |
| 2024 | Bridging the gap between object detection in close-up and high-resolution wide shots
Wenxi Li, Jilai Zheng, Haozhe Lin, Chao Ma 0004, Lu Fang 0001, Xiaokang Yang 0001 |
Comput. Vis. Image Underst. | 4 |
| 2023 | DartBlur: Privacy Preservation with Detection Artifact SuppressionabstractNowadays, privacy issue has become a top priority when training AI algorithms. Machine learning algorithms are expected to benefit our daily life, while personal information must also be carefully protected from exposure. Facial information is particularly sensitive in this regard. Multiple datasets containing facial information have been taken offline, and the community is actively seeking solutions to remedy the privacy issues. Existing methods for privacy preservation can be divided into blur-based and face replacement-based methods. Owing to the advantages of review convenience and good accessibility, blur-based based methods have become a dominant choice in practice. However, blur-based methods would inevitably introduce training artifacts harmful to the performance of downstream tasks. In this paper, we propose a novel De-artifact Blurring (DartBlur) privacy-preserving method, which capitalizes on a DNN architecture to generate blurred faces. DartBlur can effectively hide facial privacy information while detection artifacts are simultaneously suppressed. We have designed four training objectives that particularly aim to improve review convenience and maximize detection artifact suppression. We associate the algorithm with an adversarial training strategy with a second-order optimization pipeline. Experimental results demonstrate that DartBlur outperforms the existing face-replacement method from both perspectives of review convenience and accessibility, and also shows an exclusive advantage in suppressing the training artifact compared to traditional blur-based methods. Our implementation is available at https://github.com/JaNg2333/DartBlur. Baowei Jiang, Haozhe Lin, Lu Fang 0001 |
CVPR | 3 |
| 2023 | Crowd3D: Towards Hundreds of People Reconstruction from a Single ImageabstractImage-based multi-person reconstruction in wide-field large scenes is critical for crowd analysis and security alert. However, existing methods cannot deal with large scenes containing hundreds of people, which encounter the challenges of large number of people, large variations in human scale, and complex spatial distribution. In this paper, we propose Crowd3D, the first framework to reconstruct the 3D poses, shapes and locations of hundreds of people with global consistency from a single large-scene image. The core of our approach is to convert the problem of complex crowd localization into pixel localization with the help of our newly defined concept, Human-scene Virtual Interaction Point (HVIP). To reconstruct the crowd with global consistency, we propose a progressive reconstruction network based on HVIP by pre-estimating a scene-level camera and a ground plane. To deal with a large number of persons and various human sizes, we also design an adaptive human-centric cropping scheme. Besides, we contribute a benchmark dataset, LargeCrowd, for crowd reconstruction in a large scene. Experimental results demonstrate the effectiveness of the proposed method. The code and the dataset are available at http://cic.tju.edu.cn/faculty/likun/projects/Crowd3D. Huili Cui, Haozhe Lin, Yukun Lai, Lu Fang 0001, Kun Li 0001 |
CVPR | 4 |
| 2023 | RealGraph: A Multiview Dataset for 4D Real-world Context Graph GenerationabstractUnderstanding 4D scene context in real world has become urgently critical for deploying sophisticated AI systems. In this paper, we propose a brand new scene understanding paradigm called "Context Graph Generation (CGG)", aiming at abstracting holistic semantic information in the complicated 4D world. The CGG task capitalizes on the calibrated multiview videos of a dynamic scene, and targets at recovering semantic information (coordination, trajectories and relationships) of the presented objects in the form of spatio-temporal context graph in 4D space. We also present a benchmark 4D video dataset "RealGraph", the first dataset tailored for the proposed CGG task. The raw data of RealGraph is composed of calibrated and synchronized multiview videos. We exclusively provide manual annotations including object 2D&3D bounding boxes, category labels and semantic relationships. We also make sure the annotated ID for every single object is temporally and spatially consistent. We propose the first CGG baseline algorithm, Multiview-based Context Graph Generation Network (MCGNet), to empirically investigate the legitimacy of CGG task on RealGraph dataset. We nevertheless reveal the great challenges behind this task and encourage the community to explore beyond our solution. Our project page is at https://github.com/THU-luvision/RealGraph. Haozhe Lin, Zequn Chen, Jinzhi Zhang, Ruqi Huang, Lu Fang 0001 |
ICCV | 1 |
| 2023 | Toward Knowledge as a Service (KaaS): Predicting Popularity of Knowledge Services Leveraging Graph Neural NetworksabstractKnowledge services are becoming a rising star in the family of XaaS (Everything as a Service). In recent years, people are more willing to search for answers and share their knowledge directly over the Internet, which makes the knowledge service ecosystem prosperous. In this paper, we aim to predict the popularity of knowledge services, which will benefit the downstream industries. Toward such a task, the spatial interactions (e.g., hyperlinks in Wikipedia) and temporal observations (e.g., page views) provide crucial information. However, it is difficult to utilize this information due to: (i) complicated and different usage observations, (ii) intricate and evolutionary spatial interactions, and (iii) small world trait of the network. To tackle such issues, we propose evolutionary graph convolutional recurrent neural networks (E-GCRNNs) to simultaneously model both temporal and spatial dependencies of knowledge services from their evolving networks. Additionally, a localized mini-batch training scheme is developed, which allows the E-GCRNNs to work on large-scale knowledge services network and reduce the prediction bias caused by the small world trait. Extensive experiments on real-world datasets have demonstrated that the proposed E-GCRNNs outperform baselines in terms of prediction accuracy, especially with the prediction range being longer, while remaining computationally efficient. Haozhe Lin, Yushun Fan, Jia Zhang 0001, Zhenghua Xu 0001, Thomas Lukasiewicz |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Multi-Modal Reciprocal Spatiotemporal Framework for Predicting Usage Trend of Knowledge ServicesabstractAs an emerging concept, Knowledge as a Service (KaaS) aims to provide on-demand content-based (data, information, knowledge) delivery to meet the needs of users. With the prosperity of knowledge services, the prediction of the usage tendency of knowledge services has become an important and timely research topic. This study focuses on speculating the possible popularity of knowledge services in the next period of time, which can assist other downstream service tasks such as service recommendations. The interactions among knowledge services and their rich information (such as historical usage observation and text information) provide grounding for predicting the usage trend of services. However, recent spatial-temporal prediction based on graph neural networks usually depends heavily on the quality of manually created graphs, which may be expensive for knowledge services. To tackle such a limitation, this article proposes a novel Multi-modal Reciprocal SpatioTemporal (MRST) framework, which can jointly mine spatial dependencies and model time patterns for spatiotemporal coupling prediction. Two types of Edge Inference Networks (called EIN-o and EIN-t) are designed to sufficiently discover the spatial dependencies among knowledge services based on the data of usage observation sequences and service descriptions, respectively, and generate multi-modal directed weighted knowledge service graphs. Based on these graphs, MRST integrates GCN-based spatiotemporal prediction models as backbones to make predictions. Particularly, MRST features a unique reciprocal framework. On the one hand, EINs infer and generate multi-modal graphs to serve GCNs; on the other hand, GCNs utilize such spatial dependencies to make predictions and then introduce feedback to optimize EINs. In the meantime, to facilitate reproducible research, we collect a new knowledge service dataset fromWikipediacalled Wiki-EN dataset. Experiments on this real data set show that the proposed MRST framework significantly surpasses the baselines and can learn meaningful spatial dependencies outside the predefined graphic structure. Ruyu Yan, Haozhe Lin, Yushun Fan, Jia Zhang 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | MSP-RNN: Multi-Step Piecewise Recurrent Neural Network for Predicting the Tendency of Services InvocationabstractDriven by the widespread application of Service-Oriented Architecture (SOA), an increasing number of services and mashups have been developed and published onto the Internet in the past decades. With the number keeping on burgeoning, predicting the tendency of services invocation will provide various roles in service ecosystems with promising opportunities. However, services invocation bear three unique characteristics, which give rise to difficulties in predicting them. First, enormous services show different and complicated traits, like periodicity, nonlinearity and nonstationarity. Second, services providing similar or compensatory functions make up intricate relationship. Third, the combination dependencies between mashups and their comprising component services further amplify the difficulty. Given these factors, we have developed a tailored model Multi-Step Piecewise Recurrent Neural Network (MSP-RNN) to predict the tendency of services invocation. In MSP-RNN, Long Short Term Memory (LSTM) units are used to extract universal features. Based on these features, we have developed a piecewise regressive mechanism to make prediction discriminatingly. Besides, we have developed a multi-step prediction strategy to further enhance prediction accuracy and robustness. Extensive experiments in real-world data set with interpretable analysis show that MSP-RNN predicts the tendency of services invocation more accurately, i.e., by 3.7 percent in terms of symmetric mean absolute percentage error (SMAPE), than state-of-the-art baseline methods. Haozhe Lin, Yushun Fan, Jia Zhang 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2021 | Service Recommendation for Composition Creation based on Collaborative Attention Convolutional NetworkabstractService recommendation for composition creation is a widely applied technique, which expedites mashup development by reusing existing services. The core of service recommendations is to simultaneously understand user needs as well as the functions of available services. However, the descriptions provided by users and service providers may not always be accurate or up to date, which poses significant challenges to composition creating. To tackle this problem, in this paper we propose a deep learning-based service recommendation framework named coACN, short for Collaborative Attention Convolutional Network, which can effectively learn the bilateral information toward service recommendation. On the one hand, a domain-level attention module is constructed to refine user needs embeddings by drawing messages from related service domains. On the other hand, a graph convolutional network is established to excavate the service-composition graph and fuse structured information into service embeddings. For a service node in the graph, the information of its compositions as its first-order neighbor nodes is used to supplement the latest functions and features of the service; and the information of the services as its second-order neighbor nodes may bring collaborative relationships into the service. Extensive experiments on the real-world ProgrammableWeb dataset show the significant improvement of our proposed coACN framework over state-of-the-art methods. Ruyu Yan, Yushun Fan, Jia Zhang 0001, Haozhe Lin |
ICWS | 5 |
| 2021 | REST: Reciprocal Framework for Spatiotemporal-coupled PredictionsabstractIn recent years, Graph Convolutional Networks (GCNs) have been applied to benefit spatiotemporal predictions. The current shell for spatiotemporal predictions often relies heavily on the quality of handcraft, fixed graphical structures, however, we argue that such a paradigm could be expensive and sub-optimal in many applications. To raise the bar, this paper proposes to jointly mine the spatial dependencies and model temporal patterns in a coupled framework, i.e., to make spatiotemporal-coupled predictions. We come up with a novel Reciprocal SpatioTemporal (REST) framework, which introduces Edge Inference Networks (EINs) to couple with GCNs. From the temporal side to the spatial side, EINs infer spatial dependencies among time series vertices and generate multi-modal directed weighted graphs to serve GCNs. And from the temporal side to the spatial side, GCNs utilize these spatial dependencies to make predictions and then introduce feedback to optimize EINs. The REST framework is incrementally trained for higher performance of spatiotemporal prediction, powered by the reciprocity between its comprised two components from such an iterative joint learning process. Additionally, to maximize the power of the REST framework, we design a phased heuristic approach, which effectively stabilizes training procedure and prevents early-stop. Extensive experiments on two real-world datasets have demonstrated that the proposed REST framework significantly outperforms baselines, and can learn meaningful spatial dependencies beyond predefined graphical structures. Haozhe Lin, Yushun Fan, Jia Zhang 0001 |
WWW | 1 |
| 2020 | A-HSG: Neural Attentive Service Recommendation based on High-order Social GraphabstractWith the widespread application of Service-Oriented Architecture, the quantity of web services keeps increasing rapidly over the Internet. Providing personalized service recommendation to users remains to be an important research topic. Recent studies have proved social connections helpful for modeling users' potential preference thus improving the performance of service recommendation. To date, however, one special type of social relation, called high-order social relation, has not been thoroughly studied. In reality, a user's preference may not only be affected by the user's direct neighbors, but also indirect ones. Furthermore, such influences may not remain static in the context of various attentions. To tackle such issues, we have developed a novel neural Attentive network based on High-order Social Graph (A-HSG) toward offering social-aware service recommendation. First, a graph convolution-based, multi-hop propagation module is devised to extract the high-order similarity signals from users' local social networks, and inject them into the users' general representations. Second, a neighbor-level attention module is constructed to adaptively select informative neighbors to model the users' specific preference. Extensive experiments over a real-life service dataset show that A-HSG outperforms baseline methods in terms of prediction accuracy. Chunyu Wei, Yushun Fan, Jia Zhang 0001, Haozhe Lin |
ICWS | 4 |
| 2018 | PRNN: Piecewise Recurrent Neural Networks for Predicting the Tendency of Services InvocationabstractDriven by the widespread application of Service-Oriented Architecture (SOA), the quantity of web services and their users keeps increasing in the service ecosystem. Since services are hosted by service providers, it will be very helpful to predict the tendency of services invocation for service providers, so that proper actions may be taken to ensure the quality of services. Two major challenges exist in predicting the tendency of services invocation, however. First, different service invocation sequences may bear different and complicated characteristics, which is hard to be modeled generally. Second, the intricate relations between service invocation sequences are valuable but hard to be discriminated and utilized. To address these issues, a deep neural network, named Piecewise Recurrent Neural Network (PRNN), is developed by taking both generality and pertinence into consideration. For generality, PRNN extracts complicated characteristics of all service invocation sequences through Long Short-Term Memory (LSTM) units. For pertinence, PRNN develops a piecewise mechanism, through which service invocation sequences can be clustered automatically and predicted discriminatingly. Extensive experiments in real-world dataset show that PRNN outperforms baseline methods in predicting the tendency of services invocation. Haozhe Lin, Yushun Fan, Jia Zhang 0001 |
ICWS | 1 |