EDBT 2026 Demo / reviewers in the wild / expert
Yanyong Huang
dblp:171/3049
· DBLP profile ↗
15ranked-venue papers in the field
4as first author
12since 2021 · last 2026
0000-0001-9322-2777ORCID · corroborated
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 6 (2 first)Database Systems & Data Management · 4 (2 first)Data Mining & Knowledge Discovery · 3Information Retrieval & Web Search · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Data-Centric AI: A Comprehensive Survey of Traditional, Reinforcement, and Generative Approaches for Tabular Data TransformationabstractTabular data is one of the most widely used formats across industries, driving critical applications in areas such as finance, healthcare, and marketing. In the era of data-centric AI, improving data quality and representation has become essential for enhancing model performance, particularly in applications centered around tabular data. This survey examines the key aspects of tabular data-centric AI, emphasizing feature selection and feature generation as essential techniques for data space refinement. We provide a systematic review of feature selection methods, which identify and retain the most relevant data attributes, and feature generation approaches, which create new features to simplify the capture of complex data patterns. This survey offers a comprehensive overview of current methodologies through an analysis of recent advancements, practical applications, and the strengths and limitations of these techniques. Finally, we outline open challenges and suggest future perspectives to inspire continued innovation in this field. Dongjie Wang 0001, Yanyong Huang, Wangyang Ying, Haoyue Bai 0002, Nanxu Gong, Xinyuan Wang 0011, Sixun Dong, Tao Zhe, Kunpeng Liu 0001, Meng Xiao 0001, Pengfei Wang 0008, Pengyang Wang, Hui Xiong 0001, Yanjie Fu |
ACM Trans. Knowl. Discov. Data | 2 |
| 2026 | CONDEN-FI: Consistency and Diversity Learning-Based Multi-View Unsupervised Feature and Instance Co-Selection
Yanyong Huang, Yuxin Cai 0001, Dongjie Wang 0001, Xiuwen Yi, Tianrui Li 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | Adaptive Context-Infused Performance Evaluator for Iterative Feature Space OptimizationabstractIterative feature space optimization includes continuously evaluating and refining the feature space to improve downstream task performance. However, existing methods commonly suffer from three major limitations: 1) ignoring differences between samples leads to evaluation bias; 2) the feature space is overly tailored to specific models, resulting in overfitting and poor generalization; and 3) retraining the evaluator from scratch in each iteration significantly reduces overall efficiency. To bridge these gaps, we introduce EASE (gEneralized Adaptive feature Space Evaluator), a generalized framework for efficient and objective evaluation of iteratively generated feature spaces. This framework includes two key components: Feature-Sample Subspace Generator and Contextual Attention Evaluator. The first component aims to mitigate evaluation bias by decoupling the information distribution within the feature space. To achieve this, based on feedback from the subsequent evaluator, we identify the samples most challenging for evaluation and the features most relevant to prediction tasks. The second component intends to incrementally capture evolving patterns of the feature space for efficient evaluation. Specifically, we propose a weighted-sharing multi-head attention mechanism to encode the feature space into an embedding vector for evaluation, and update the evaluator incrementally to retain prior knowledge while incorporating new information. Extensive experiments on fifteen public datasets demonstrate the effectiveness of EASE. We have released our code and data to the public. Yanyong Huang, Zijun Yao 0001, Yanjie Fu, Kunpeng Liu 0001, Xiao Luo 0001, Dongjie Wang 0001 |
CIKM | 2 |
| 2024 | Progressive Multimodal Pivot Learning: Towards Semantic Discordance Understanding as HumansabstractMultimodal recognition can achieve enhanced performance by leveraging the complementary information from different modali- ties. However, in real-world scenarios, multimodal samples often express discordant semantic meanings across modalities, lacking evident complementary information. Unlike humans who can easily understand the intrinsic semantic information of these semantically discordant samples, existing multimodal recognition models show poor performance on them. With the motivation of improving the robustness of multimodal recognition models in practical scenar- ios, this work poses a new challenge in multimodal recognition, which is coined as Semantic Discordance Understanding. Unlike ex- isting works only focusing on detecting semantically discordant samples as noisy data, this new challenge requires deep models to follow humans’ ability in understanding the inherent seman- tic meanings of semantically discordant samples. To address this challenge, we further propose the Progressive Multimodal Pivot Learning (PMPL) approach by introducing a learnable pivot mem- ory to explore the inherent semantics meaning hidden under dis- cordant modalities. To this end, our approach inserts Pivot Memory Learning (PML) modules into multiple layers of unimodal foun- dation models to progressively trade-off the conflict information across modalities. By introducing the multimodal pivot learning paradigm for multimodal recognition, the proposed PMPL approach can alleviate the negative effect of semantic discordance caused by the cross-modal information exchange mechanism of existingmultimodal recognition models. Experiments on different bench- marks validate the superiority of our approach. Code is available at https://github.com/tiggers23/PMPL. Junlin Fang, Wenya Wang 0001, Tianze Luo, Yanyong Huang, Fengmao Lv |
CIKM | 4 |
| 2024 | Spatio-Temporal Consistency Enhanced Differential Network for Interpretable Indoor Temperature PredictionabstractIndoor temperature prediction is crucial for decision-making in central heating systems. Beyond accuracy, predictions shall be interpretable, i.e. conform to the laws of physics; otherwise, it may lead to system failures or unsafe conditions. However, deep learning models often face criticism regarding interpretability, which limits their application in such settings. To this end, we propose a Spatio-Temporal Consistency enhanced Differential Network (CONST) for interpretable indoor temperature prediction. Our approach mainly consists of a differential predictive module and a spatio-temporal consistency module. Modeling the influential factors, the first module solves the issue of multicollinearity through the differential operation. Considering the heterogeneity of global and local data distributions, the second module characterizes the temporal and spatial consistency to mine the universal pattern by multi-task learning, thereby improving the prediction interpretability. Besides, we propose a set of interpretability metrics to overcome the drawbacks of partial dependence plot metric, which are more practical, zero-centered, flexible, and numerical. We conclude experiments on a real-world dataset with four heating stations. The results demonstrate the advantages of our approach over various baselines, where the interpretability can be improved by more than 8 times on cRPD while maintaining high accuracy. We developed CONST on the SmartHeat system, providing hourly indoor temperature forecasts for 13 heating stations in northern China. Dekang Qi, Xiuwen Yi, Chengjie Guo, Yanyong Huang, Junbo Zhang 0004, Tianrui Li 0001, Yu Zheng 0004 |
KDD | 4 |
| 2023 | Automated urban planning aware spatial hierarchies and human instructions
Dongjie Wang 0001, Kunpeng Liu 0001, Yanyong Huang, Leilei Sun, Bowen Du 0001, Yanjie Fu |
Knowl. Inf. Syst. | 3 |
| 2023 | C2IMUFS: Complementary and Consensus Learning-Based Incomplete Multi-View Unsupervised Feature SelectionabstractMulti-view unsupervised feature selection (MUFS) has been demonstrated as an effective technique to reduce the dimensionality of multi-view unlabeled data. The existing methods assume that all of views are complete. However, multi-view data are usually incomplete, i.e., a part of instances are presented on some views but not all views. Besides, learning the complete similarity graph, as an important promising technology in existing MUFS methods, cannot achieve due to the missing views. In this paper, we propose a complementary and consensus learning-based incomplete multi-view unsupervised feature selection method (C$^{2}$IMUFS) to address the aforementioned issues. Concretely, C$^{2}$IMUFS integrates feature selection into an extended weighted non-negative matrix factorization model equipped with adaptive learning of view-weights and a sparse$\ell _{2,p}$-norm, which can offer better adaptability and flexibility. By the sparse linear combinations of multiple similarity matrices derived from different views, a complementary learning-guided similarity matrix reconstruction model is presented to obtain the complete similarity graph in each view. Furthermore, C$^{2}$IMUFS learns a consensus clustering indicator matrix across different views and embeds it into a spectral graph term to preserve the local geometric structure. Comprehensive experimental results on real-world datasets demonstrate the effectiveness of C$^{2}$IMUFS compared with state-of-the-art methods. Yanyong Huang, Zongxin Shen, Yuxin Cai 0001, Xiuwen Yi, Dongjie Wang 0001, Fengmao Lv, Tianrui Li 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Multi-memory Enhanced Separation Network for Indoor Temperature Prediction
Zhewen Duan, Xiuwen Yi, Dekang Qi, Yexin Li, Haoran Xu 0003, Yanyong Huang, Junbo Zhang 0004, Yu Zheng 0004 |
DASFAA (2) | 7 |
| 2022 | Matrix representation of the conditional entropy for incremental feature selection on multi-source data
Yanyong Huang, Kejun Guo, Xiuwen Yi, Zhong Li 0001, Tianrui Li 0001 |
Inf. Sci. | 1 |
| 2022 | Dynamic three-way neighborhood decision model for multi-dimensional variation of incomplete hybrid data
Yanyong Huang, Tianrui Li 0001, Xin Yang 0012 |
Inf. Sci. | 2 |
| 2022 | Orthogonally constrained matrix factorization for robust unsupervised feature selection with local preserving
Chuan Luo 0001, Tianrui Li 0001, Hongmei Chen 0001, Yanyong Huang, Xi Peng 0001 |
Inf. Sci. | 5 |
| 2022 | Gas-Theft Suspect Detection Among Boiler Room Users: A Data-Driven ApproachabstractThe natural gas tightly correlates with our everyday life. However, driven by gray incomes, some users are prone to stealing gas by refitting the equipment without permission. Especially for the boiler room users in winter, this phenomenon appears more rampant. Traditional gas-theft detection methods highly rely on the on-site inspection, where exists ineffective and randomness. With the rapidly deployed IoT sensors, we can collect real-time gas consumption data to analyze users’ behavior patterns, where the gas-theft suspects could be discovered early and accurately. In this paper, we propose a data-driven approach, named SVOC, to detect gas-theft suspects among boiler room users. Our approach consists of a scenario-based data quality detection algorithm, a deformation-based normality detection algorithm, and an One-Class Support Vector Machine (OCSVM) based anomaly detection algorithm. Specifically, considering the temporal proximity between the gas consumption and the outdoor temperature, the normality detection algorithm adopts a similarity-based deformation correlation to detect normal boiler room users out of abnormal ones. Then, we employ OCSVM as the anomaly detection algorithm to capture various features across multiple data sources, aiming to distinguish gas-theft suspects from the remaining irregular users. Here, the detected normal and abnormal users are fed into the OCSVM for training and prediction, respectively, which can overcome the label scarcity problem. We conduct extensive experiments on a real-world dataset during one heating season. The results demonstrate distinct advantages of our approach over various baselines. We have developed a real-time system on the cloud, providing daily gas-theft suspects for gas companies. Xiuwen Yi, Yanyong Huang, Songyu Ke, Junbo Zhang 0004, Tianrui Li 0001, Yu Zheng 0004 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Incremental three-way neighborhood approach for dynamic incomplete hybrid data
Tianrui Li 0001, Yanyong Huang, Xin Yang 0012 |
Inf. Sci. | 3 |
| 2020 | Dynamic maintenance of rough approximations in multi-source hybrid information systems
Yanyong Huang, Tianrui Li 0001, Chuan Luo 0001, Hamido Fujita, Shi-Jinn Horng, Bin Wang 0045 |
Inf. Sci. | 1 |
| 2019 | Updating three-way decisions in incomplete multi-scale information systems
Chuan Luo 0001, Tianrui Li 0001, Yanyong Huang, Hamido Fujita |
Inf. Sci. | 3 |