EDBT 2026 Demo / reviewers in the wild / expert
Honggang Wang 0001
dblp:70/5417-1
· DBLP profile ↗
5ranked-venue papers in the field
0as first author
1since 2021 · last 2024
0000-0001-9475-2630ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 4Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Feature Interaction Detection in Big Data Through a New Choquet Integral based Deep Neural NetworkabstractLearning from massive amounts of domain-specific information requires new algorithms and models for parsing the ever-expanding field of big data. Such algorithms for exploring and identifying key features in vast databases require analysis of complex interactions to uncover critical features under a variety of circumstances. We study a comprehensive collection of health-related data, showing that our novel Choquet Integral activation function for deep neural networks transforms high-dimensional data into simpler sub-feature sets that better model complex interactions. While standard methods account for unitary feature tracking, they do not extend to multiple feature subsets, an impactful and necessary knowledge base. To this end, our novel activation function creates a sub-additive tool that better considers the weighted compilation of features within a robust set of standard benchmarks, advancing the synergistic and antagonistic relationships among features, capturing non-linear dependencies. We present the theoretical underpinnings, highlighting balanced fuzzy measures and sub-additivity for an optimized model based on real-world health data targeting weight loss. We further test different model settings, akin to hyper-parameter optimization. Despite computational time consumption, which could be improved via nowadays more powerful computing units, this novel method can be implemented as a pre-trained model using big data to identify heretofore unknown sub-additive feature interactions in a variety of fields such as biomedicine, fraud detection, cyber-security, and finance. Matthew Fried, Honggang Wang 0001, Hua Fang 0001 |
IEEE Big Data | 2 |
| 2018 | K-nearest Neighbor Search by Random Projection ForestsabstractK-nearest neighbor (kNN) search has wide applications in many areas, including data mining, machine learning, statistics and many applied domains. Inspired by the success of ensemble methods and the flexibility of tree-based methodology, we propose random projection forests, rpForests, for kNN search. rpForests finds kNNs by aggregating results from an ensemble of random projection trees with each constructed recursively through a series of carefully chosen random projections. rpForests achieves a remarkable accuracy in terms of fast decay in the missing rate of kNNs and that of discrepancy in the kNN distances. rpForests has a very low computational complexity. The ensemble nature of rpForests makes it easily run in parallel on multicore or clustered computers; the running time is expected to be nearly inversely proportional to the number of cores or machines. We give theoretical insights by showing the exponential decay of the probability that neighboring points would be separated by ensemble random projection trees when the ensemble size increases. Our theory can be used to refine the choice of random projections in the growth of trees, and experiments show that the effect is remarkable. Donghui Yan, Honggang Wang 0001 |
IEEE BigData | 4 |
| 2015 | Using probabilistic approach to joint clustering and statistical inference: Analytics for big investment dataabstractThis paper proposes a Contrarian Probabilistic Model (CPM) to evaluate the effectiveness of contrarians' investment in preferred stocks using big data from Tradeline. CPM accommodates the unique features of investment data which are often correlated, nested, heterogeneous, non-normal with missing values. The clustering and statistical inference are integrated in CPM, which enables joint investment behavior trajectory pattern recognition and risk analyses based on the entire variance-covariance structure between and within clusters. The empirical study using CPM provides a finer and comprehensive evaluation of contrarian investment in preferred stocks. Two distinctive investment behavior trajectory clusters were identified, showing a few high-risk-seeking contrarians achieved high returns over five year long-term investment, while the majority of contrarians did not outperform glamour stockholders in preferred stock investment. Although CPM was developed using historical data, it could be developed into an analytical tool for online near real time big investment data analyses. Hua Fang 0001, Honggang Wang 0001, Chonggang Wang, Mahmoud Daneshmand |
IEEE BigData | 2 |
| 2015 | A novel initialization method for particle swarm optimization-based FCM in big biomedical dataabstractBased on empirical studies, the feature of random initialization in Particle Swarm Optimization (PSO) based Fuzzy c-means (FCM) methods affects the computational performance especially in big data. As the data points in high-density areas are more likely near the cluster centroids, we design a new algorithm to guide the initialization according to the data density patterns. Our algorithm is initialized by fusing the data characteristics near the cluster centers. Our evaluation results from real data show that our approach can significantly improve the computational performance of PSO-based Fuzzy clustering methods, while preserving comparable clustering performance. Chanpaul Jin Wang, Hua Fang 0001, Chonggang Wang, Mahmoud Daneshmand, Honggang Wang 0001 |
IEEE BigData | 5 |
| 2012 | User preferences based software defect detection algorithms selection using MCDM
Yi Peng 0001, Guoxun Wang, Honggang Wang 0001 |
Inf. Sci. | 3 |