Jin Huang 0007

dblp:49/2488-7 · DBLP profile ↗
← Back
10ranked-venue papers in the field
2as first author
6since 2021 · last 2026
0000-0003-2285-5248ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 7 (1 first)Database Systems & Data Management · 3 (1 first)
YearPublicationVenuePosition
2026 Mitigating Generic Token Dominance in Cross-Domain Foundation Model for Text-Attributed Graphs
Haochen You, Lubin Gan, Jin Huang 0007
DASFAA (2)7
2026 Gbf 2rammar: Bilingual Grammar Modeling for Enhanced Text-Attributed Graph Learning
Heng Zheng 0008, Haochen You, Lubin Gan, Jin Huang 0007
DASFAA (2)8
2026 EaSFE: Scalable and Efficient Feature Engineering for Boosting Machine Learning Performance
abstract
Feature engineering plays a critical role in machine learning (ML), but existing methods often struggle with high computational cost and limited scalability when applied to large-scale and sparse datasets. In this article, we propose EaSFE, an efficient and scalable feature engineering framework that unifies feature generation, filtering, and evaluation in an end-to-end manner. EaSFE is designed to efficiently construct and select informative features while explicitly considering computational and memory constraints. To achieve scalability, EaSFE incorporates parallel and distributed execution mechanisms, as well as a chunk-based data processing strategy that enables memory-efficient feature engineering on large datasets. In addition, EaSFE adopts tailored storage and execution strategies to handle high-dimensional sparse data effectively. Extensive experiments on multiple real-world datasets demonstrate that EaSFE consistently improves predictive performance (e.g., 5% accuracy improvement in poker ) while substantially enhancing efficiency (i.e., over 10x speedup) compared to existing feature engineering methods. In addition, EaSFE is demonstrated to scale to large and sparse datasets, successfully handling datasets with over 119 million training instances and 54 million features.
Jian Chen 0011, Yile Chen 0004, Zhenya Zheng, Zeyi Wen, Jin Huang 0007
ACM Trans. Knowl. Discov. Data6
2026 DyGHydra: A Hierarchical State-Space Model with Time Dynamics and Interactive-Relational Selectivity for Link Prediction
abstract
The task of dynamic graph link prediction is to forecast the evolution of complex systems. Empirical observations reveal that interactions within these systems exhibit an Entangled Spatio-Temporal Pattern, which manifests through three interrelated phenomena, namely Latent High-Order Bridges, Multi-Frequency Temporal Dynamics, and Spatio-Temporal Entanglement, with stronger structural ties facilitating tolerance for longer temporal gaps. However, limited by computationally prohibitive multi-hop sampling or inefficient long-sequence modeling, existing methods struggle to capture this complex pattern. Inspired by State-Space Models (SSMs) like Mamba for efficient long-range modeling yet aiming to address their native agnosticism to structural and multi-frequency dynamics, we propose a framework named DyGHydra, which couples a tailored Continuous-Time Hierarchical Mamba (CT-HMamba) backbone with a multi-hop structural encoder. The framework first employs the multi-hop structural encoder to reveal latent high-order interactions, extracting interaction-level cross-hop features. Subsequently, the CT-HMamba backbone utilizes these features to address multi-frequency dynamics through a hierarchical architecture, decomposing interaction history to simultaneously model high-frequency bursts and long-term trends. To capture the spatio-temporal entanglement, CT-HMamba further tailors its core state-space mechanism to be co-driven by physical time and structural context. Specifically, physical time governs the state transition decay to reflect temporal forgetting, while structural context modulates the input-output projections to prioritize topologically significant events. Extensive experiments on eleven real-world datasets show that DyGHydra achieves state-of-the-art performance across most settings for both transductive and inductive link prediction, validating its effectiveness in modeling complex temporal dynamics with superior efficiency.
Yueqi Guo, Weihao Yu 0002, Jin Huang 0007
ACM Trans. Knowl. Discov. Data4
2025 Towards Recommendation on Good Quality Data Science Solutions
abstract
Data science aims to solve real-world problems with the knowledge derived from data. Successfully tackling a data science problem requires practitioners to choose an appropriate solution, which potentially comprises various components such as pre-processing techniques, learning algorithms, hyper-parameters, and so on. Therefore, a problem-driven recommendation for the promising solution is invaluable, as it facilitates efficient and convenient problem-solving. However, existing solution recommendation approaches confront notable challenges when dealing with limited and sparse prior experience in practical applications. Learning from such prior easily leads to overfitting and poor generalization in solution recommendations. To address this issue, we propose a novel solution recommendation method that can predict a good-quality data science solution, including the pre-processing, the learning algorithm, and hyper-parameters, for a given problem. The foundation of our method is a carefully designed ranking model that exploits a weight-sharing structure and a newly proposed loss. The ranking model focuses on incorporating relative ranking information into the predicted performance score of each solution. With these techniques, our method can recommend the solution with the highest score and effectively mitigate the limitations of using sparse prior experience. Our experiments demonstrate the superiority of our method in predicting solutions with higher accuracy and rank, even trained on highly sparse historical performance records. It also reduces recommendation time significantly compared to the baselines, offering remarkable efficiency and convenience for practitioners.
Jian Chen 0011, Yile Chen 0004, Zeyi Wen, Jin Huang 0007
ACM Trans. Knowl. Discov. Data5
2023 Two-Stage Denoising Diffusion Model for Source Localization in Graph Inverse Problems
Bosong Huang, Weihao Yu 0002, Ruzhong Xie, Jing Xiao 0005, Jin Huang 0007
ECML/PKDD (3)5
2015 A privacy-enhancing model for location-based personalized recommendations
Jin Huang 0007, Jianzhong Qi 0001, Yabo Xu, Jian Chen 0011
Distributed Parallel Databases1
2013 Recommendations for two-way selections using skyline view queries
Jian Chen 0011, Jin Huang 0007, Bin Jiang 0009, Jian Pei 0001, Jian Yin 0001
Knowl. Inf. Syst.2
2013 Skyline distance: a measure of multidimensional competence
Jin Huang 0007, Bin Jiang 0009, Jian Pei 0001, Jian Chen 0011, Yong Tang 0001
Knowl. Inf. Syst.1
2005 Mining Correlated Rules for Associative Classification
Jian Chen 0011, Jian Yin 0001, Jin Huang 0007
ADMA3