EDBT 2026 Demo / reviewers in the wild / expert
Heyuan Huang
dblp:32/6989
· DBLP profile ↗
13ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon TasksabstractThe rapid advancement of multimodal large language models has enabled agents to operate mobile devices by directly interacting with graphical user interfaces, opening new possibilities for mobile automation. However, real-world mobile tasks are often complex and allow for multiple valid solutions. This contradicts current mobile agent evaluation standards: offline static benchmarks can only validate a single predefined ''golden path'', while online dynamic testing is constrained by the complexity and non-reproducibility of real devices, making both approaches inadequate for comprehensively assessing agent capabilities. To bridge the gap between offline and online evaluation and enhance testing stability, this paper introduces a novel graph-structured benchmarking framework. By modeling the finite states observed during real-device interactions, it achieves static simulation of dynamic behaviors. Building on this, we develop ColorBench, a benchmark focused on complex long-horizon tasks. It supports evaluation of multiple valid solutions, subtask completion rate statistics, and atomic-level capability analysis. ColorBench contains 175 tasks (74 single-app, 101 cross-app) with an average length of over 13 steps. Each task includes at least two correct paths and several typical error paths, enabling quasi-dynamic interaction. Yuanyi Song, Heyuan Huang, Qiqiang Lin, Yin Zhao, Xiangmou Qu, Jun Wang 0152, Xingyu Lou, Weiwen Liu, Zhuosheng Zhang 0001, Jun Wang 0020, Zhaoxiang Wang, Yong Yu 0001, Weinan Zhang 0001 |
WWW | 2 |
| 2025 | RAG+: Enhancing Retrieval-Augmented Generation with Application-Aware ReasoningabstractYu Wang, Shiwan Zhao, Zhihu Wang, Ming Fan, Xicheng Zhang, Yubo Zhang, Zhengfan Wang, Heyuan Huang, Ting Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yu Wang 0093, Shiwan Zhao, Zhihu Wang, Ming Fan 0002, Yubo Zhang 0006, Zhengfan Wang, Heyuan Huang, Ting Liu 0002 |
EMNLP | 8 |
| 2025 | Training-free Periodic Interest Augmentation in Incremental RecommendationabstractIndustrial recommender systems usually train models incrementally to grasp recent interests of users. However, a fundamental issue of these incremental updated models is their tendency to overfit current data while neglecting past information. Specifically, we have observed that the data distribution of real systems exhibits periodic drifts, leading to periodic fluctuations of prediction bias. To alleviate the above bias fluctuations while minimizing the loss of recent interests, we propose TPIA, a Training-free approach for Periodic Interest Augmentation in incremental recommendation. Specifically, after the latest model is trained, we first calculate the importance score of each model in the previous period. Then, we merge these models based on the importance scores. To minimize information loss due to interference of parameters during model merging, we further develop a method for trimming redundant and abnormal parameters. Offline experiments on both public and private datasets demonstrate the effectiveness of TPIA. It has also been deployed on a large-scale industrial recommender system, and has shown a notable 1.61% increase in CVR and a 1.97% increase in CPM, along with enhanced stability in prediction bias. Heyuan Huang, Xingyu Lou, Changwang Zhang, Chaochao Chen 0001, Kuiyao Dong, Han Lei, Yihao Wang 0007, Wangchunshu Zhou, Jun Wang 0020 |
SIGIR | 1 |
| 2025 | Public data-enhanced multi-stage differentially private graph neural networks
Heyuan Huang, Lingbo Wei, Chi Zhang 0001 |
J. Inf. Secur. Appl. | 2 |
| 2024 | DIIT: A Domain-Invariant Information Transfer Method for Industrial Cross-Domain RecommendationabstractCross-Domain Recommendation (CDR) have received widespread attention due to their ability to utilize rich information across domains. However, most existing CDR methods assume an ideal static condition that is not practical in industrial recommendation systems (RS). Therefore, simply applying existing CDR methods in the industrial RS environment may lead to low effectiveness and efficiency. To fill this gap, we propose DIIT, an end-to-end Domain-Invariant Information Transfer method for industrial cross-domain recommendation. Specifically, We first simulate the industrial RS environment that maintains respective models in multiple domains, each of them is trained in the incremental mode. Then, for improving the effectiveness, we design two extractors to fully extract domain-invariant information from the latest source domain models at the domain level and the representation level respectively. Finally, for improving the efficiency, we design a migrator to transfer the extracted information to the latest target domain model, which only need the target domain model for inference. Experiments conducted on one production dataset and two public datasets verify the effectiveness and efficiency of DIIT. Heyuan Huang, Xingyu Lou, Chaochao Chen 0001, Pengxiang Cheng 0003, Chengwei He, Jun Wang 0020 |
CIKM | 1 |
| 2024 | Balancing Centralized and Local Differential Privacy in Pneumonia DiagnosisabstractAs a collaborative learning method, federated learning trains models on large datasets without sharing raw data. However, sharing gradients generated with the aggregation server could reveal sensitive information related to the raw data in federated learning. Privacy-preserving methods such as differential privacy (DP) are needed for federated learning clinical applications. Existing centralized DP relies on a trusted server, while local DP suffers from the accuracy limitations of adding noise from each data owner. In this paper, we propose a privacy-preserving federated learning method by centralized differential privacy without a trusted server. By encrypting each local gradient with multikey homomorphic encryption, we can prevent the aggregation server from inferring local data privacy during the training process. By adding noise to the aggregated gradient with centralized DP, we prevent medical institutions from inferring private information about other institutions from the aggregated global update. The utilization of multikey homo-morphic encryption eliminates the need for a trusted server. Our experimental evaluation on chest X-ray dataset demonstrates that our proposed approach has a 9% loss in accuracy compared to non-private federated learning. Liwei Luo, Heyuan Huang |
ICC | 4 |
| 2024 | Vessel-targeted compensation of deformable motion in interventional cone-beam CT
Alexander Lu, Heyuan Huang, Wojciech Zbijewski, Mathias Unberath, Jeffrey H. Siewerdsen, Clifford R. Weiss, Alejandro Sisniega |
Medical Image Anal. | 2 |
| 2023 | A PATE-based Approach for Training Graph Neural Networks under Label Differential PrivacyabstractAs a standard solution to the problem of private deep learning, differential privacy (DP) is widely used in graph neural networks (GNNs) to protect sensitive information about the input graph data. However, most existing DP algorithms for GNNs protect the privacy of every attribute for each node. This results in the need for injecting a large amount of noise, making these methods significantly underperform their non-private counterparts. We argue that in some practical scenarios, node labels serve as the only or the most sensitive attribute, where label differential privacy, a more fine-grained notion of differential privacy that only protects the labels is more appropriate. To better capture these scenarios and improve the trade-off between data privacy and model accuracy, we propose a novel method of training GNNs under label differential privacy. Instead of naively adding noise to the node labels before training the GNN, our method follows the strategy of Private Aggregation of Teacher Ensembles (PATE) to generate differentially private node labels with both high accuracy and strong privacy guarantee. We also propose a label denoising module that takes advantage of the graph structure to further improve the accuracy of the trained model. Additionally, our method is model-agnostic, making it applicable to any GNN architecture. We evaluate its performance on two commonly used benchmark datasets and demonstrate its capability to learn high-performance models while ensuring privacy. Heyuan Huang, Liwei Luo, Yankai Xie, Chi Zhang 0001, Jianqing Liu |
GLOBECOM | 1 |
| 2023 | Synthesizing High-Utility Tabular Data with Enhanced Privacy Via Split-and-Discard Pre-TrainingabstractData sharing has led to the emergence of the deep generative model (DGM) with differential privacy for synthesizing tabular data. However, existing methods struggle to synthesize high-utility tabular data with enhanced privacy. One challenge is degraded data utility due to the limited number of training iterations available under strong privacy guarantees. The other challenge is that widely-used encoding schemes may leak the sensitive distribution of continuous features. To this end, we propose a novel pipeline incorporating split-and-discard pre-training and an embedding module to synthesize data. To reduce the impact of limited iterations, we employ the split-and-discard pre-training method. This method leverages the intrinsic structure of DGM, which can be split into discriminative and generative sub-models. By conducting pre-training and discarding specific sub-models of DGM on private data, we address these challenges while training models with differential privacy. To preserve the privacy of continuous features, we propose a piecewise linear one-hot encoding scheme followed by an embedding layer. We instantiate this pipeline using variational autoencoders and generative adversarial networks respectively and compare them against popular models and variants. Results show that our pipeline on private data effectively balances privacy and utility. Liwei Luo, Heyuan Huang, Yankai Xie, Chi Zhang 0001, Lingbo Wei |
GLOBECOM | 2 |
| 2023 | A Graph Convolutional Neural Network for Recommendation Based on Community Detection and Combination of Multiple Heterogeneous GraphsabstractGraph Convolutional Neural Networks (GCNs) have performed well in many recommendation scenarios. In spite of this, recommendation models based on GCNs still face problems such as insufficient information mining and high complexity for some existing models. To address the above problems, we propose a Graph Convolutional Neural Network for Recommendation Based on Community Detection and the Combination of Multiple Heterogeneous Graphs (GCN-CMHG). This model uses the community detection algorithm to detect the communities in the user-item interaction heterogeneous graph (UIIHG), Finds the regional central nodes of communities, and then creates edges between the regional central node of each community and all other nodes in the UIIHG to construct the heterogeneous partial adjacent graph. Then, a Heterogeneous Partial Adjacent Auxiliary (HPAA) layer is designed to aggregate information on the heterogeneous partial adjacent graph. HPAA layer expands the influence of distant nodes on target nodes, enables target nodes to receive global information, and enhances the ability of GCN-CMHG to mine information. Specially, due to the low complexity of HPAA layer and the abandonment of redundant information, GCN-CMHG is easier to implement and train. Under the exact same experimental setting, GCN-CMHG’s time consumption is only about 1/10 of another model based on GCN called Graph Convolutional Neural Network for Recommendation Based on the Combination of Multiple Heterogeneous Graphs (GCN-MHG). Experiments on multiple real-world datasets show that GCN-CMHG achieves better results compared with several advanced models. The implementation of our work can be found at https://github.com/GCNRSs/GCN-CMHG. Caihong Mu, Heyuan Huang, Yunfei Fang 0002, Yi Liu 0051 |
ICDM | 2 |
| 2023 | Automatic Detection of Tooth-Gingiva Trim Lines on Dental SurfacesabstractDetecting the tooth-gingiva trim line from a dental surface plays a critical role in dental treatment planning and aligner 3D printing. Existing methods treat this task as a segmentation problem, which is resolved with geometric deep learning based mesh segmentation techniques. However, these methods can only provide indirect results (i.e., segmented teeth) and suffer from unsatisfactory accuracy due to the incapability of making full use of high-resolution dental surfaces. To this end, we propose a two-stage geometric deep learning framework for automatically detecting tooth-gingiva trim lines from dental surfaces. Our framework consists of a trim line proposal network (TLP-Net) for predicting an initial trim line from the low-resolution dental surface as well as a trim line refinement network (TLR-Net) for refining the initial trim line with the information from the high-resolution dental surface. Specifically, our TLP-Net predicts the initial trim line by fusing the multi-scale features from a U-Net with a proposed residual multi-scale attention fusion module. Moreover, we propose feature bridge modules and a trim line loss to further improve the accuracy. The resulting trim line is then fed to our TLR-Net, which is a deep-based LDDMM model with the high-resolution dental surface as input. In addition, dense connections are incorporated into TLR-Net for improved performance. Our framework provides an automatic solution to trim line detection by making full use of raw high-resolution dental surfaces. Extensive experiments on a clinical dental surface dataset demonstrate that our TLP-Net and TLR-Net are superior trim line detection methods and outperform cutting-edge methods in both qualitative and quantitative evaluations. Geng Chen 0001, Jie Qin 0004, Boulbaba Ben Amor, Weiming Zhou, Hang Dai, Tao Zhou 0002, Heyuan Huang, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2005 | A practical pattern recovery approach based on both structural and behavioral analysis
Heyuan Huang, Shensheng Zhang, Jian Cao 0001, Yonghong Duan |
J. Syst. Softw. | 1 |
| 2003 | Hierarchical process patterns: construct software processes in a stepwise wayabstractPatterns are widely used to capture design decisions and rationale of software, but they could also be used to document and guide the development of software process. This paper proposes a framework called hierarchical process patterns (HPP), which includes three types of pattern: lifecycle pattern, activity pattern, and workflow pattern. This division makes it easier to tailor and refine software processes in a stepwise way. To describe the workflow of process pattern and relationships between roles and artifacts of realized activity and those of sub activities, this paper presents a set of representation mechanism and defines role inheritance diagram and artifact decomposition diagram. To support variable number of same category of activities, this paper introduce parameterized compound activity in workflow pattern. Finally, this paper gives an example to illustrate how to apply the framework in SPDM, which is a process-centered software engineering environment. Heyuan Huang, Shensheng Zhang |
SMC | 1 |