EDBT 2026 Demo / reviewers in the wild / expert
Hyun Ah Song
dblp:38/10461
· DBLP profile ↗
21ranked-venue papers
4as first author
1since 2021 · last 2024
0000-0001-6337-9036ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 15 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 10 · 4 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Data mining · 60% Data integration and cleaning · 28% Spatial and temporal data management · 8% | |
| Theoretical computer science
3 papers |
Graph algorithms and graph theory · 47% Information theory · 31% Mathematical optimization · 22% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › dimensionality reduction
canonical correlation analysis |
0.6 | 2 | 2019 | Efficient and Distributed Generalized Canonical Correlation Analysis for Big Multiview Data · IEEE Trans. Knowl. Data Eng. 2019 Efficient and Distributed Algorithms for Large-Scale Generalized Canonical Correlations Analysis · ICDM 2016 |
Data mining
dimensionality reduction |
0.6 | 2 | 2019 | Efficient and Distributed Generalized Canonical Correlation Analysis for Big Multiview Data · IEEE Trans. Knowl. Data Eng. 2019 Efficient and Distributed Algorithms for Large-Scale Generalized Canonical Correlations Analysis · ICDM 2016 |
Data integration and cleaning › data preprocessing › data cleaning
data repair |
0.4 | 1 | 2020 | TurboLift: fast accuracy lifting for historical data recovery · VLDB J. 2020 |
Data integration and cleaning
data reconstruction |
0.3 | 1 | 2018 | Ares: Automatic Disaggregation of Historical Data · ICDE 2018 |
Data mining
time series analysis |
0.3 | 1 | 2018 | Ares: Automatic Disaggregation of Historical Data · ICDE 2018 |
Information theory › signal processing › compressed sensing › sparse recovery
basis pursuit |
0.3 | 1 | 2018 | HomeRun: Scalable Sparse-Spectrum Reconstruction of Aggregated Historical Data · Proc. VLDB Endow. 2018 |
Data mining
anomaly detection |
0.2 | 1 | 2016 | FRAUDAR: Bounding Graph Fraud in the Face of Camouflage · KDD 2016 |
Data mining › big data analytics › large-scale data mining
distributed data mining |
0.2 | 1 | 2016 | Efficient and Distributed Algorithms for Large-Scale Generalized Canonical Correlations Analysis · ICDM 2016 |
Data mining › anomaly detection
fraud detection |
0.2 | 1 | 2016 | FRAUDAR: Bounding Graph Fraud in the Face of Camouflage · KDD 2016 |
Graph algorithms and graph theory
dense subgraph discovery |
0.2 | 1 | 2016 | FRAUDAR: Bounding Graph Fraud in the Face of Camouflage · KDD 2016 |
Graph algorithms and graph theory
graph mining |
0.2 | 1 | 2016 | FRAUDAR: Bounding Graph Fraud in the Face of Camouflage · KDD 2016 |
Mathematical optimization
distributed optimization |
0.1 | 1 | 2019 | Efficient and Distributed Generalized Canonical Correlation Analysis for Big Multiview Data · IEEE Trans. Knowl. Data Eng. 2019 |
Mathematical optimization
large-scale optimization |
0.1 | 1 | 2019 | Efficient and Distributed Generalized Canonical Correlation Analysis for Big Multiview Data · IEEE Trans. Knowl. Data Eng. 2019 |
Query processing and optimization
aggregation |
0.1 | 1 | 2018 | Ares: Automatic Disaggregation of Historical Data · ICDE 2018 |
Methods — techniques the papers use, named apart from their topics
karush-kuhn-tucker convergence analysis · 0.8alternating optimization · 0.8discrete cosine transform · 0.7alternating direction method of multipliers · 0.7sparse matrix computation · 0.5distributed optimization · 0.5constraint-based repair · 0.4sparsity constraints · 0.3sparsity constraint · 0.3least-squares approximation · 0.3annihilating filter · 0.3bounding · 0.2bipartite graph analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Correction to: TurboLift: fast accuracy lifting for historical data recovery
Faisal M. Almutairi, Hyun Ah Song, Christos Faloutsos, Nicholas D. Sidiropoulos, Vladimir Zadorozhny |
VLDB J. | 3 |
| 2020 | TurboLift: fast accuracy lifting for historical data recovery
Faisal M. Almutairi, Hyun Ah Song, Christos Faloutsos, Nicholas D. Sidiropoulos, Vladimir Zadorozhny |
VLDB J. | 3 |
| 2019 | TensorCast: forecasting and mining with coupled tensors
Miguel Araujo, Pedro Ribeiro 0004, Hyun Ah Song, Christos Faloutsos |
Knowl. Inf. Syst. | 3 |
| 2019 | Efficient and Distributed Generalized Canonical Correlation Analysis for Big Multiview DataabstractGeneralized canonical correlation analysis (GCCA) integrates information from data samples that are acquired at multiple feature spaces (or `views') to produce low-dimensional representations-which is an extension of classical two-view CCA. Since the 1960s, (G)CCA has attracted much attention in statistics, machine learning, and data mining because of its importance in data analytics. Despite these efforts, the existing GCCA algorithms have serious complexity issues. The memory and computational complexities of the existing algorithms usually grow as a quadratic and cubic function of the problem dimension (the number of samples / features), respectively-e.g., handling views with ≈1,000 features using such algorithms already occupies ≈106memory and the periteration complexity is ≈109flops-which makes it hard to push these methods much further. To circumvent such difficulties, we first propose a GCCA algorithm whose memory and computational costs scale linearly in the problem dimension and the number of nonzero data elements, respectively. Consequently, the proposed algorithm can easily handle very large sparse views whose sample and feature dimensions both exceed 100,000. Our second contribution lies in proposing two distributed algorithms for GCCA, which compute the canonical components of different views in parallel and thus can further reduce the runtime significantly if multiple computing agents are available. We provide detailed convergence analyses of the proposed algorithms and show that all the largescale GCCA algorithms converge to a Karush-Kuhn-Tucker (KKT) point at least sublinearly. Judiciously designed synthetic and realdata experiments are employed to showcase the effectiveness of the proposed algorithms. Xiao Fu 0001, Kejun Huang, Evangelos E. Papalexakis, Hyun Ah Song, Partha P. Talukdar, Nicholas D. Sidiropoulos, Christos Faloutsos, Tom M. Mitchell |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2018 | Ares: Automatic Disaggregation of Historical DataabstractWe address the challenge of reconstructing historical counts from aggregated, possibly overlapping historical reports. For example, given the monthly and weekly sums, how can we find the daily counts of people infected with flu? We propose an approach, called ARES (Automatic REStoration), that performs automatic data reconstruction in two phases: (1) first, it estimates the sequence of historical counts utilizing domain knowledge, such as smoothness and periodicity of historical events; (2) then, it uses the estimated sequence to learn notable patterns in the target sequence to refine the reconstructed time series. In order to derive such patterns, ARES uses an annihilating filter technique. The idea is to learn a linear shift-invariant operator whose response to the desired sequence is (approximately) zero-yielding a set of null-space equations that the desired signal should satisfy, without the need for the accompanying data. The reconstruction accuracy can be further improved by applying the second phase iteratively. We evaluate ARES on the real epidemiological data from the Tycho project and demonstrate that ARES recovers historical data from aggregated reports with high accuracy. In particular, it considerably outperforms top competitors, including least squares approximation and the more advanced H-FUSE method (42% and 34% improvement based on average RMSE, respectively). Hyun Ah Song, Zongge Liu, Christos Faloutsos, Vladimir Zadorozhny, Nicholas D. Sidiropoulos |
ICDE | 2 |
| 2018 | GridWatch: Sensor Placement and Anomaly Detection in the Electrical Grid
Bryan Hooi, Dhivya Eswaran, Hyun Ah Song, Amritanshu Pandey, Marko Jereminov, Lawrence T. Pileggi, Christos Faloutsos |
ECML/PKDD (1) | 3 |
| 2018 | StreamCast: Fast and Online Mining of Power Grid Time SequencesabstractHow can we efficiently forecast the power consumption of a location for the next few days? More challengingly, how can we forecast the power consumption if the temperature increases by 10° C, the number of appliances in the grid increase by 20%, and voltage levels increase by 5%? Such ‘what-if scenarios' are crucial for future planning, to ensure that the grid remains reliable even under extreme conditions. Our contributions are as follows: 1) Domain knowledge infusion: we propose a novel Temporal BIG model that extends the physics-based BIG model, allowing it to capture changes over time, trends, and seasonality, and temperature effects. 2) Forecasting: our StreamCast algorithm forecasts multiple steps ahead and outperforms baselines in accuracy. Our algorithm is online, requiring constant update time per new data point and bounded memory. 3) What-if scenarios and anomaly detection: our approach can handle scenarios in which the voltage levels, temperature, or number of appliances change. It also spots anomalies in real data, and provides confidence intervals for its forecasts, to assist in planning for various scenarios. Experimental results show that StreamCast has 27% lower forecasting error than baselines on real data, scales linearly, and runs in 4 minutes on a time sequence of 40 million points. Bryan Hooi, Hyun Ah Song, Amritanshu Pandey, Marko Jereminov, Lawrence T. Pileggi, Christos Faloutsos |
SDM | 2 |
| 2018 | HomeRun: Scalable Sparse-Spectrum Reconstruction of Aggregated Historical DataabstractRecovering a time sequence of events from multiple aggregated and possibly overlapping reports is a major challenge in historical data fusion. The goal is to reconstruct a higher resolution event sequence from a mixture of lower resolution samples as accurately as possible. For example, we may aim to disaggregate overlapping monthly counts of people infected with measles into weekly counts. In this paper, we propose a novel data disaggregation method, called H ome R un , that exploits an alternative representation of the sequence and finds the spectrum of the target sequence. More specifically, we formulate the problem as so-called basis pursuit using the Discrete Cosine Transform (DCT) as a sparsifying dictionary and impose non-negativity and smoothness constraints. H ome R un utilizes the energy compaction feature of the DCT by finding the sparsest spectral representation of the target sequence that contains the largest (most important) coefficients. We leverage the Alternating Direction Method of Multipliers to solve the resulting optimization problem with scalable and memory efficient steps. Experiments using real epidemiological data show that our method considerably outperforms the state-of-the-art techniques, especially when the DCT of the sequence has a high degree of energy compaction. Faisal M. Almutairi, Hyun Ah Song, Christos Faloutsos, Nicholas D. Sidiropoulos, Vladimir Zadorozhny |
Proc. VLDB Endow. | 3 |
| 2017 | PowerCast: Mining and Forecasting Power Grid Sequences
Hyun Ah Song, Bryan Hooi, Marko Jereminov, Amritanshu Pandey, Lawrence T. Pileggi, Christos Faloutsos |
ECML/PKDD (2) | 1 |
| 2017 | BrainZoom: High Resolution Reconstruction from Multi-modal Brain SignalsabstractHow close can we zoom in to observe brain activity? Our understanding is limited by the resolution of imaging modalities that exhibit good spatial but poor temporal resolution, or vice-versa. In this paper, we propose BrainZoom, an efficient imaging algorithm that cross-leverages multi-modal brain signals. BrainZoom (a) constructs high resolution brain images from multi-modal signals, (b) is scalable, and (c) is flexible in that it can easily incorporate various priors on the brain activities, such as sparsity, low rank, or smoothness. We carefully formulate the problem to tackle nonlinearity in the measurements (via variable splitting) and auto-scale between different modal signals, and judiciously design an inexact alternating optimization-based algorithmic framework to handle the problem with provable convergence guarantees. Our experiments using a popular realistic brain signal simulator to generate fMRI and MEG demonstrate that high spatio-temporal resolution brain imaging is possible from these two modalities. The experiments also suggest that smoothness seems to be the best prior, among several we tried. Xiao Fu 0001, Kejun Huang, Otilia Stretcu, Hyun Ah Song, Evangelos E. Papalexakis, Partha P. Talukdar, Tom M. Mitchell, Nicholas D. Sidiropoulos, Christos Faloutsos, Barnabás Póczos |
SDM | 4 |
| 2017 | H-Fuse: Efficient Fusion of Aggregated Historical DataabstractIn this paper, we address the challenge of recovering a time sequence of counts from aggregated historical data. For example, given a mixture of the monthly and weekly sums, how can we find the daily counts of people infected with flu? In general, what is the best way to recover historical counts from aggregated, possibly overlapping historical reports, in the presence of missing values? Equally importantly, how much should we trust this reconstruction? We propose H-Fuse, a novel method that solves above problems by allowing injection of domain knowledge in a principled way, and turning the task into a well-defined optimization problem. H-Fuse has the following desirable properties: (a) Effectiveness, recovering historical data from aggregated reports with high accuracy; (b) Self-awareness, providing an assessment of when the recovery is not reliable; (c) Scalability, computationally linear on the size of the input data. Experiments on the real data (epidemiology counts from the Tycho project [13]) demonstrates that H-FUSE reconstructs the original data 30 – 81% better than the least squares method. Zongge Liu, Hyun Ah Song, Vladimir Zadorozhny, Christos Faloutsos, Nicholas D. Sidiropoulos |
SDM | 2 |
| 2017 | Graph-Based Fraud Detection in the Face of CamouflageabstractGiven a bipartite graph of users and the products that they review, or followers and followees, how can we detect fake reviews or follows? Existing fraud detection methods (spectral, etc.) try to identify dense subgraphs of nodes that are sparsely connected to the remaining graph. Fraudsters can evade these methods using camouflage , by adding reviews or follows with honest targets so that they look “normal.” Even worse, some fraudsters use hijacked accounts from honest users, and then the camouflage is indeed organic. Our focus is to spot fraudsters in the presence of camouflage or hijacked accounts. We propose FRAUDAR, an algorithm that (a) is camouflage resistant, (b) provides upper bounds on the effectiveness of fraudsters, and (c) is effective in real-world data. Experimental results under various attacks show that FRAUDAR outperforms the top competitor in accuracy of detecting both camouflaged and non-camouflaged fraud. Additionally, in real-world experiments with a Twitter follower--followee graph of 1.47 billion edges, FRAUDAR successfully detected a subgraph of more than 4, 000 detected accounts, of which a majority had tweets showing that they used follower-buying services. Bryan Hooi, Kijung Shin, Hyun Ah Song, Alex Beutel, Neil Shah, Christos Faloutsos |
ACM Trans. Knowl. Discov. Data | 3 |
| 2016 | Efficient and Distributed Algorithms for Large-Scale Generalized Canonical Correlations AnalysisabstractGeneralized canonical correlation analysis (GCCA) aims at extracting common structure from multiple 'views', i.e., high-dimensional matrices representing the same objects in different feature domains – an extension of classical two-view CCA. Existing (G)CCA algorithms have serious scalability issues, since they involve square root factorization of the correlation matrices of the views. The memory and computational complexity associated with this step grow as a quadratic and cubic function of the problem dimension (the number of samples / features), respectively. To circumvent such difficulties, we propose a GCCA algorithm whose memory and computational costs scale linearly in the problem dimension and the number of nonzero data elements, respectively. Consequently, the proposed algorithm can easily handle very large sparse views whose sample and feature dimensions both exceed 100,000 – while the current approaches can only handle thousands of features / samples. Our second contribution is a distributed algorithm for GCCA, which computes the canonical components of different views in parallel and thus can further reduce the runtime significantly (by ≥ 30% in experiments) if multiple cores are available. Judiciously designed synthetic and real-data experiments using a multilingual dataset are employed to showcase the effectiveness of the proposed algorithms. Xiao Fu 0001, Kejun Huang, Evangelos E. Papalexakis, Hyun Ah Song, Partha P. Talukdar, Nicholas D. Sidiropoulos, Christos Faloutsos, Tom M. Mitchell |
ICDM | 4 |
| 2016 | FRAUDAR: Bounding Graph Fraud in the Face of CamouflageabstractGiven a bipartite graph of users and the products that they review, or followers and followees, how can we detect fake reviews or follows? Existing fraud detection methods (spectral, etc.) try to identify dense subgraphs of nodes that are sparsely connected to the remaining graph. Fraudsters can evade these methods using camouflage, by adding reviews or follows with honest targets so that they look "normal". Even worse, some fraudsters use hijacked accounts from honest users, and then the camouflage is indeed organic. Our focus is to spot fraudsters in the presence of camouflage or hijacked accounts. We propose FRAUDAR, an algorithm that (a) is camouflage-resistant, (b) provides upper bounds on the effectiveness of fraudsters, and (c) is effective in real-world data. Experimental results under various attacks show that FRAUDAR outperforms the top competitor in accuracy of detecting both camouflaged and non-camouflaged fraud. Additionally, in real-world experiments with a Twitter follower-followee graph of 1.47 billion edges, FRAUDAR successfully detected a subgraph of more than 4000 detected accounts, of which a majority had tweets showing that they used follower-buying services. Bryan Hooi, Hyun Ah Song, Alex Beutel, Neil Shah, Kijung Shin, Christos Faloutsos |
KDD | 2 |
| 2016 | Matrices, Compression, Learning Curves: Formulation, and the GroupNteach Algorithms
Bryan Hooi, Hyun Ah Song, Evangelos E. Papalexakis, Rakesh Agrawal 0001, Christos Faloutsos |
PAKDD (2) | 2 |
| 2015 | Hierarchical feature extraction by multi-layer non-negative matrix factorization network for classification task
Hyun Ah Song, Bo-Kyeong Kim, Thanh Xuan Luong, Soo-Young Lee |
Neurocomputing | 1 |
| 2015 | Noise-Robust Detection of Symmetric Axes by Self-Correcting Artificial Neural Network
Wonil Chang, Hyun Ah Song, Sang-Hoon Oh, Soo-Young Lee |
Neural Process. Lett. | 2 |
| 2013 | Flexible Reasoning of Boolean Constraints in Recurrent Neural Networks with Dual Representation
Wonil Chang, Hyun Ah Song, Soo-Young Lee |
ICONIP (1) | 2 |
| 2013 | Hierarchical Representation Using NMF
Hyun Ah Song, Soo-Young Lee |
ICONIP (1) | 1 |
| 2012 | Self-correcting Symmetry Detection Network
Wonil Chang, Hyun Ah Song, Sang-Hoon Oh, Soo-Young Lee |
ICONIP (2) | 2 |
| 2011 | Enhanced Discrimination of Face Orientation Based on Gabor Filters
Hyun Ah Song, Sung-Do Choi, Soo-Young Lee |
ICONIP (1) | 1 |