Dantong Yu

dblp:45/3429 · DBLP profile ↗
← Back
47ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0001-5921-5231ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 1 since 2021Databases, data management, data science and information retrieval · 17 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 ESPNet: Edge-Aware Graph Representation Learning Over Analyst-Firm Bipartite Networks for Earnings Surprise Prediction
Siqi Jiang, Xinyuan Tao, Ajim Uddin, Zhi Wei 0001, Dantong Yu
IEEE Big Data5
2025 TrackGNN: A Highly Parallelized and Self-Adaptive GNN Accelerator for Track Reconstruction on FPGAs
abstract
Real-time track reconstruction in high energy physics imposes stringent latency constraints, hindering the deployment of graph neural networks (GNNs) on general-purpose platforms. We present TrackGNN11https//github.com/silvenachen/TrackGNN, an open-sourced GNN accelerator for track reconstruction. Using a dataflow architecture with multiple parallelism and a self-adaptive renaming mechanism, TrackGNN shows 27.6× speedup over CPUs, up to 101.1× over GPUs, and 5.7× over an FPGA overlay. Compared with FlowGNN, the renaming mechanism also reduces end-to-end latency by 1.12-1.16× with negligible resource overhead.
Ruiqi Chen 0001, Bruno da Silva 0001, Giorgian Borca-Tasciuc, Dantong Yu, Cong Hao
FCCM6
2023 End-to-End Pipeline for Trigger Detection on Hit and Track Graphs
abstract
There has been a surge of interest in applying deep learning in particle and nuclear physics to replace labor-intensive offline data analysis with automated online machine learning tasks. This paper details a novel AI-enabled triggering solution for physics experiments in Relativistic Heavy Ion Collider and future Electron-Ion Collider. The triggering system consists of a comprehensive end-to-end pipeline based on Graph Neural Networks that classifies trigger events versus background events, makes online decisions to retain signal data, and enables efficient data acquisition. The triggering system first starts with the coordinates of pixel hits lit up by passing particles in the detector, applies three stages of event processing (hits clustering, track reconstruction, and trigger detection), and labels all processed events with the binary tag of trigger versus background events. By switching among different objective functions, we train the Graph Neural Networks in the pipeline to solve multiple tasks: the edge-level track reconstruction problem, the edge-level track adjacency matrix prediction, and the graph-level trigger detection problem. We propose a novel method to treat the events as track-graphs instead of hit-graphs. This method focuses on intertrack relations and is driven by underlying physics processing. As a result, it attains a solid performance (around 72% accuracy) for trigger detection and outperforms the baseline method using hit-graphs by 2% higher accuracy.
Tingting Xuan, Giorgian Borca-Tasciuc, Ming Xiong Liu, Yu Sun 0075, Cameron Dean, Yasser Corrales Morales, Zhaozhong Shi, Dantong Yu
AAAI9
2022 Trigger Detection for the sPHENIX Experiment via Bipartite Graph Networks with Set Transformer
Tingting Xuan, Giorgian Borca-Tasciuc, Yu Sun 0075, Cameron Dean, Zhaozhong Shi, Dantong Yu
ECML/PKDD (3)7
2021 Perturbing Eigenvalues with Residual Learning in Graph Convolutional Neural Networks
abstract
Network structured data is ubiquitous in natural and social science applications. Graph Convolutional Neural Network (GCN) has attracted significant attention recently due to its success in representing, modeling, and predicting large-scale network data. Various types of graph convolutional filters were proposed to process graph signals to boost the performance of graph-based semi-supervised learning. This paper introduces a novel spectral learning technique called EigLearn, which uses residual learning to perturb the eigenvalues of the graph filter matrix to optimize its capability. EigLearn is relatively easy to implement, and yet thorough experimental studies reveal that it is more effective and efficient than the prior works on the specific issue, such as LanczosNet and FisherGCN. EigLearn only perturbs a small number of eigenvalues and does not require a complete eigendecomposition. Our investigation shows that EigLearn reaches the maximal performance improvement by perturbing about 30 to 40 eigenvalues, and the EigLearn-based GCN has comparable efficiency as the standard GCN. Furthermore, EigLearn bears a clear explanation in the spectral domain of the graph filter and shows aggregation effects in performance improvement when coupled with different graph filters. Hence, we anticipate that EigLearn may serve as a useful neural unit in various graph-involved neural net architectures.
Shibo Yao, Dantong Yu, Xiangmin Jiao
ACML2
2021 Attention Based Dynamic Graph Learning Framework for Asset Pricing
abstract
Recent studies suggest that financial networks play an essential role in asset valuation and investment decisions. Unlike road networks, financial networks are neither given nor static, posing significant challenges in learning meaningful networks and promoting their applications in price prediction. In this paper, we first apply the attention mechanism to connect the "dots" (firms) and learn dynamic network structures among stocks over time. Next, the end-to-end graph neural networks pipeline diffuses and propagates the firms' accounting fundamentals into the learned networks and ultimately predicts stock future returns. The proposed model reduces the prediction errors by 6% compared to the state-of-the-art models. Our results are robust with different assessment measures. We also show that portfolios based on our model outperform the S&P-500 index by 34% in terms of Sharpe Ratio, suggesting that our model is better at capturing the dynamic inter-connection among firms and identifying stocks with fast recovery from major events. Further investigation on the learned networks reveals that the network structure aligns closely with the market conditions. Finally, with an ablation study, we investigate different alternative versions of our model and the contribution of each component.
Ajim Uddin, Xinyuan Tao, Dantong Yu
CIKM3
2020 Multi-Label Visual Feature Learning with Attentional Aggregation
abstract
Today convolutional neural networks (CNNs) have reached out to specialized applications in science communities that otherwise would not be adequately tackled. In this paper, we systematically study a multi-label annotation problem of x-ray scattering images in material science. For this application, we tackle an open challenge with training CNNs - identifying weak scattered patterns with diffuse background interference, which is common in scientific imaging. We articulate an Attentional Aggregation Module (AAM) to enhance feature representations. First, we reweight and highlight important features in the images using data-driven attention maps. We decompose the attention maps into channel and spatial attention components. In the spatial attention component, we design a mechanism to generate multiple spatial attention maps tailored for diversified multi-label learning. Then, we condense the enhanced local features into non-local representations by performing feature aggregation. Both attention and aggregation are designed as network layers with learnable parameters so that CNN training remains fluidly end-to-end, and we apply it in-network a few times so that the feature enhancement is multi-scale. We conduct extensive experiments on CNN training and testing, as well as transfer learning, and empirical studies confirm that our method enhances the discriminative power of visual features of scientific imaging.
Ziqiao Guan, Kevin G. Yager, Dantong Yu, Hong Qin 0001
WACV3
2019 Enhancing Domain Word Embedding via Latent Semantic Imputation
abstract
We present a novel method named Latent Semantic Imputation (LSI) to transfer external knowledge into semantic space for enhancing word embedding. The method integrates graph theory to extract the latent manifold structure of the entities in the affinity space and leverages non-negative least squares with standard simplex constraints and power iteration method to derive spectral embeddings. It provides an effective and efficient approach to combining entity representations defined in different Euclidean spaces. Specifically, our approach generates and imputes reliable embedding vectors for low-frequency words in the semantic space and benefits downstream language tasks that depend on word embedding. We conduct comprehensive experiments on a carefully designed classification problem and language modeling and demonstrate the superiority of the enhanced embedding via LSI over several well-known benchmark embeddings. We also confirm the consistency of the results under different parameter settings of our method.
Shibo Yao, Dantong Yu, Keli Xiao
KDD2
2019 A Lightweight Multi-Section CNN for Lung Nodule Classification and Malignancy Estimation
abstract
The size and shape of a nodule are the essential indicators of malignancy in lung cancer diagnosis. However, effectively capturing the nodule's structural information from CT scans in a computer-aided system is a challenging task. Unlike previous models that proposed computationally intensive deep ensemble models or three-dimensional CNN models, we propose a lightweight, multiple view sampling based multi-section CNN architecture. The model obtains a nodule's cross sections from multiple view angles and encodes the nodule's volumetric information into a compact representation by aggregating information from its different cross sections via a view pooling layer. The compact feature is subsequently used for the task of nodule classification. The method does not require the nodule's spatial annotation and works directly on the cross sections generated from volume enclosing the nodule. We evaluated the proposed method on lung image database consortium (LIDC) and image database resource initiative (IDRI) dataset. It achieved the state-of-the-art performance with a mean 93.18% classification accuracy. The architecture could also be used to select the representative cross sections determining the nodule's malignancy that facilitates in the interpretation of results. Because of being lightweight, the model could be ported to mobile devices, which brings the power of artificial intelligence (AI) driven application directly into the practitioner's hand.
Pranjal Sahu, Dantong Yu, Mallesham Dasari, Fei Hou 0001, Hong Qin 0001
IEEE J. Biomed. Health Informatics2
2018 Automatic X-ray Scattering Image Annotation via Double-View Fourier-Bessel Convolutional Networks
Ziqiao Guan, Hong Qin 0001, Kevin G. Yager, Youngwoo Choo, Dantong Yu
BMVC5
2017 PhiOpenSSL: Using the Xeon Phi Coprocessor for Efficient Cryptographic Calculations
abstract
The Secure Sockets Layer (SSL) is the main protocol used to secure Internet traffic and cloud computing. It relies on the computation-intensive RSA cryptography, which primarily limits the throughput of the handshake process. In this paper, we design and implement an OpenSSL library, termed PhiOpenSSL, which targets the Intel Xeon Phi (KNC) coprocessor, and utilizes Intel Phi's SIMD and multi-threading capability to reduce the SSL computation latency. In particular, PhiOpenSSL vectorizes all big integer multiplications and Montgomery operations involved in RSA calculations and employs theChinese Remainder Theorem and fixed-window exponentiation in its customized library. In an experiment involving the computation of Montgomery exponentiation, PhiOpenSSL was as much as 15.3 times faster than the two other reference libcrypto libraries, one from the Intel Many-core Platform Software Stack (MPSS) and the other from the default OpenSSL. Our RSA private key cryptography routines in PhiOpenSSL are 1.6-5.7 times faster than those in these two reference systems.
Dantong Yu
IPDPS2
2017 X-Ray Scattering Image Classification Using Deep Learning
abstract
Visual inspection of x-ray scattering images is a powerful technique for probing the physical structure of materials at the molecular scale. In this paper, we explore the use of deep learning to develop methods for automatically analyzing x-ray scattering images. In particular, we apply Convolutional Neural Networks and Convolutional Autoencoders for x-ray scattering image classification. To acquire enough training data for deep learning, we use simulation software to generate synthetic x-ray scattering images. Experiments show that deep learning methods outperform previously published methods by 10% on synthetic and real datasets.
Boyu Wang 0001, Kevin G. Yager, Dantong Yu, Minh Hoai
WACV3
2017 Analysis of NUMA effects in modern multicore systems for the design of high-performance data transfer applications
Yufei Ren, Dantong Yu, Shudong Jin
Future Gener. Comput. Syst.3
2017 RAMSYS: Resource-Aware Asynchronous Data Transfer with Multicore SYStems
abstract
High-speed data transfer is vital to data-intensive computing that often requires moving large data volumes efficiently within a local data center and among geographically dispersed facilities. Effective utilization of the abundant resources in modern multicore environments for data transfer remains a persistent challenge, particularly, for Non-Uniform Memory Access (NUMA) systems wherein the locality of data accessing is an important factor. This requires rethinking how to exploit parallel access to data and to optimize the storage and network I/Os. We address this challenge and present a novel design of asynchronous processing and resource-aware task scheduling in the context of high-throughput data replication. Our software allocates multiple sets of threads to different stages of the processing pipeline, including storage I/O and network communication, based on their capacities. Threads belonging to each stage follow an asynchronous model, and attain high performance via multiple locality-aware and peer-aware mechanisms, such as task grouping, buffer sharing, affinity control and communication protocols. Our design also integrates high performance features to enhance the scalability of data transfer in several scenarios, e.g., file-level sorting, block-level asynchrony, and thread-level pipelining. Our experiments confirm the advantages of our software under different types of workloads and dynamic environments with contention for shared resources, including a 28-160 percent increase in bandwidth for transferring large files, 1.7-66 times speed-up for small files, and up to 108 percent larger throughput for mixed workloads compared with three state of the art alternatives, GridFTP , BBCP and Aspera.
Yufei Ren, Dantong Yu, Shudong Jin
IEEE Trans. Parallel Distributed Syst.3
2016 Diverse Power Iteration Embeddings: Theory and Practice
abstract
Manifold learning, especially spectral embedding, is known as one of the most effective learning approaches on high dimensional data, but for real-world applications it raises a serious computational burden in constructing spectral embeddings for large datasets. To overcome this computational complexity, we propose a novel efficient embedding construction, Diverse Power Iteration Embedding (DPIE). DPIE shows almost the same effectiveness of spectral embeddings and yet is three order of magnitude faster than spectral embeddings computed from eigen-decomposition. Our DPIE is unique in that (1) it finds linearly independent embeddings and thus shows diverse aspects of dataset; (2) the proposed regularized DPIE is effective if we need many embeddings; (3) we show how to efficiently orthogonalize DPIE if one needs; and (4) Diverse Power Iteration Value (DPIV) provides the importance of each DPIE like an eigen value. Such various aspects of DPIE and DPIV ensure that our algorithm is easy to apply to various applications, and we also show the effectiveness and efficiency of DPIE on clustering, anomaly detection, and feature selection as our case studies.
Hao Huang 0007, Shinjae Yoo, Dantong Yu, Hong Qin 0001
IEEE Trans. Knowl. Data Eng.3
2015 Vectorized Big Integer Operations for Cryptosystems on the Intel MIC Architecture
abstract
Cryptosystems play a vital role in cyber security. To accelerate their big integer operations without jeopardizing the security level has become one of the main objectives of the cryptography research. However, popular big integer libraries highly optimized for CPU and GPGPU perform poorly on the emerging Intel Xeon Phi coprocessor mainly because they cannot take advantage of the 512-bit Single Instruction Multiple Data (SIMD) vector parallelism on Intel Many Integrated Core (MIC) architecture. In this paper, we design and implement big integer algorithms of addition, multiplication, square and modular exponentiation for Intel MIC architecture. Our algorithms offer two key improvements over other big integer algorithms: 1) They explicitly and efficiently vectorize calculations via the 512-bit SIMD intrinsics, achieving a remarkable computing efficiency (vectorization), 2) They minimize memory footprints by restricting all operations within available registers and L1 cache once operands are loaded from the memory (memory-efficient). Furthermore, we apply our algorithms to implement a cryptography application, RSA private key decryption, to demonstrate their high performance. The benchmark results on Intel Phi confirm that our big integer algorithms consistently and unambiguously outperform the best available libraries (GNU MP, OpenSSL and MAPM) on Intel Phi in terms of both latency and throughput. Our exemplary RSA implementation achieves 4.3 to 5.9 times shorter latency and 5.4 to 9.1 times more throughput compared to that in OpenSSL.
Dantong Yu
HiPC3
2015 sRSA: High Speed RSA on the Intel MIC Architecture
abstract
RSA cryptography provides key functions for signing/verifying digital signatures and encrypting/decrypting shared secrets, and is broadly deployed to ensure secure endto-end communications nowadays. However, the adoption of RSA-enabled applications is fairly limited mainly due to the large computation overheads, in particular, of cryptographic operations with the RSA private key. In this paper, we design and implement sRSA, a high speed RSA on the new Intel®Many Integrated Core (MIC) architecture. We introduce several optimization strategies to sRSA for the MIC architecture without jeopardizing the security level. For example, 1) sRSA explicitly and efficiently vectorizes the underlying cryptographic primitives with the 512-bit vector registers; 2) It also integrates the advanced algorithmic features of fast RSA variants and many other fine-grained implementations; 3) It is thoroughly designed to resist applicable RSA system attacks, i.e., the factoring attacks and the CRT exponent attack. In the end, we evaluate the performance of sRSA and compare it with the industry-standard OpenSSL. The benchmark result shows that sRSA retains a comparable latency with, but demonstrates a much higher throughput than OpenSSL on both CPU and MIC based Phi coprocessor.
Dantong Yu
ICPADS3
2015 Resources-Conscious Asynchronous High-Speed Data Transfer in Multicore Systems: Design, Optimizations, and Evaluation
abstract
One constant challenge in multicourse systems is to utilize fully the abundant resources, while assuring superior performance for individual tasks, particularly, in Non-uniform Memory Access (NUMA) systems where the locality of access is an important factor. To achieve this goal requires rethinking how to exploit parallel data access and I/O related optimizations. In the context of developing software for high-speed data transfer, we offer a novel design using asynchronous processing, and detail the advantages of resources-conscious task scheduling. In our design, multiple sets of threads are allocated to the different stages of the processing pipeline based on the capacity of resources, including storage I/O, and network communication operations. The threads in these stages are executed in an asynchronous mode, and they communicate efficiently via localized mechanisms in NUMA systems, e.g., task grouping, buffer memory, and locks. With this design, multiple effective optimizations are seamlessly integrated particularly for improving the performance and scalability of end-to-end data transfer. To validate the benefits of the design and optimizations therein, we conducted extensive experiments on the state-of-the-art multicourse systems. Our results highlighted the performance advantages of our software across different typical workloads, compared to the widely adopted data transfer tools, Graft and BBCP.
Yufei Ren, Dantong Yu, Shudong Jin
IPDPS3
2015 A Stochastic Framework for Solar Irradiance Forecasting Using Condition Random Field
Shinjae Yoo, Dantong Yu, Hao Huang 0007, John Heiser, Paul Kalb
PAKDD (1)3
2015 Prior knowledge driven Granger causality analysis on gene regulatory network discovery
abstract
BACKGROUND: Our study focuses on discovering gene regulatory networks from time series gene expression data using the Granger causality (GC) model. However, the number of available time points (T) usually is much smaller than the number of target genes (n) in biological datasets. The widely applied pairwise GC model (PGC) and other regularization strategies can lead to a significant number of false identifications when n>>T. RESULTS: In this study, we proposed a new method, viz., CGC-2SPR (CGC using two-step prior Ridge regularization) to resolve the problem by incorporating prior biological knowledge about a target gene data set. In our simulation experiments, the propose new methodology CGC-2SPR showed significant performance improvement in terms of accuracy over other widely used GC modeling (PGC, Ridge and Lasso) and MI-based (MRNET and ARACNE) methods. In addition, we applied CGC-2SPR to a real biological dataset, i.e., the yeast metabolic cycle, and discovered more true positive edges with CGC-2SPR than with the other existing methods. CONCLUSIONS: In our research, we noticed a " 1+1>2" effect when we combined prior knowledge and gene expression data to discover regulatory networks. Based on causality networks, we made a functional prediction that the Abm1 gene (its functions previously were unknown) might be related to the yeast's responses to different levels of glucose. Our research improves causality modeling by combining heterogeneous knowledge, which is well aligned with the future direction in system biology. Furthermore, we proposed a method of Monte Carlo significance estimation (MCSE) to calculate the edge significances which provide statistical meanings to the discovered causality networks. All of our data and source codes will be available under the link https://bitbucket.org/dtyu/granger-causality/wiki/Home.
Shinjae Yoo, Dantong Yu
BMC Bioinform.3
2015 End-to-End Delay Minimization for Scientific Workflows in Clouds under Budget Constraint
abstract
Next-generation e-Science features large-scale, compute-intensive workflows of many computing modules that are typically executed in a distributed manner. With the recent emergence of cloud computing and the rapid deployment of cloud infrastructures, an increasing number of scientific workflows have been shifted or are in active transition to cloud environments. As cloud computing makes computing a utility, scientists across different application domains are facing the same challenge of reducing financial cost in addition to meeting the traditional goal of performance optimization. We develop a prototype generic workflow system by leveraging existing technologies for a quick evaluation of scientific workflow optimization strategies. We construct analytical models to quantify the network performance of scientific workflows using cloud-based computing resources, and formulate a task scheduling problem to minimize the workflow end-to-end delay under a user-specified financial constraint. We rigorously prove that the proposed problem is not only NP-complete but also non-approximable. We design a heuristic solution to this problem, and illustrate its performance superiority over existing methods through extensive simulations and real-life workflow experiments based on proof-of-concept implementation and deployment in a local cloud testbed.
Chase Qishi Wu, Xiangyu Lin, Dantong Yu
IEEE Trans. Cloud Comput.3
2015 Density-Aware Clustering Based on Aggregated Heat Kernel and Its Transformation
abstract
Current spectral clustering algorithms suffer from the sensitivity to existing noise and parameter scaling and may not be aware of different density distributions across clusters. If these problems are left untreated, the consequent clustering results cannot accurately represent true data patterns, in particular, for complex real-world datasets with heterogeneous densities. This article aims to solve these problems by proposing a diffusion-based Aggregated Heat Kernel (AHK) to improve the clustering stability, and a Local Density Affinity Transformation (LDAT) to correct the bias originating from different cluster densities. AHK statistically models the heat diffusion traces along the entire time scale, so it ensures robustness during the clustering process, while LDAT probabilistically reveals the local density of each instance and suppresses the local density bias in the affinity matrix. Our proposed framework integrates these two techniques systematically. As a result, it not only provides an advanced noise-resisting and density-aware spectral mapping to the original dataset but also demonstrates the stability during the processing of tuning the scaling parameter (which usually controls the range of neighborhood). Furthermore, our framework works well with the majority of similarity kernels, which ensures its applicability to many types of data and problem domains. The systematic experiments on different applications show that our proposed algorithm outperforms state-of-the-art clustering algorithms for the data with heterogeneous density distributions and achieves robust clustering performance with respect to tuning the scaling parameter and handling various levels and types of noise.
Hao Huang 0007, Shinjae Yoo, Dantong Yu, Hong Qin 0001
ACM Trans. Knowl. Discov. Data3
2015 Design, Implementation, and Evaluation of a NUMA-Aware Cache for iSCSI Storage Servers
abstract
In an iSCSI based storage area network, target hosts serve concurrent I/O requests from initiators to achieve both high throughput and low latency. Existing iSCSI leverages the OS page cache to ensure data sharing and reuse. However, the non-uniform memory access (NUMA) architecture introduces another dimension of complexity, i.e., asymmetric memory access in multi-core and many-core platforms. Within a NUMA platform, an iSCSI target often dispatches an access request with a cache hit to an I/O thread remote to cached data, and thus cannot fully utilize multi-core systems. We encounter this problem in the context of ultra high-speed data transfer between two iSCSI storage systems, during which inferior NUMA remote memory access lags behind available high network bandwidth, and thereby becomes a bottleneck of the entire end-to-end data transfer path. We design a NUMA-aware cache mechanism to align cache memory with local NUMA nodes and threads, and then schedule I/O requests to those threads that are local to the data being accessed. This NUMA-aware solution results in lower access latency and higher system throughput. We implement a cache system within the Linux SCSI target framework, and evaluated it on our NUMA-based iSCSI testbed. Experimental results show the NUMA-aware cache can significantly improve the performance of iSCSI as measured by several benchmark tools and confirm its viability in data intensive applications and real-life workloads.
Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi
IEEE Trans. Parallel Distributed Syst.3
2014 An Improved Ratio-Based (IRB) Batch Effects Removal Algorithm for Cancer Data in a Co-Analysis Framework
abstract
Ratio-based algorithms are proven to be effective methods for removing batch effects that exist among micro array expression data from different data sources. They are outperforming than other methods in the enhancement of cross-batch prediction, especially for cancer data sets. However, their overall power is limited by: (1) Not every batch has control samples. The original method uses all negative samples to calculate the subtrahend. (2) Micro array experimental data may not have clear labels, especially in the prediction application, the labels of test data set are unknown. In this paper, we propose an Improved Ratio-Based (IRB) method to relieve these two constraints for cross-batch prediction applications. For each batch in a single study, we select one reference sample based on the idea of aligning probability density functions (pdfs) of each gene in different batches. Moreover, for data sets without label information, we transfer the problem of finding reference sample to the dense sub graph problem in graph theory. Our newly-proposed IRB method is straightforward and efficient, and can be extended for integrating large volume micro array data sets. The experiments show that our method is stable and has high performance in tumor/non-tumor prediction.
Shuchu Han, Hong Qin 0001, Dantong Yu
BIBE3
2014 Diverse Power Iteration Embeddings and Its Applications
abstract
Spectral Embedding is one of the most effective dimension reduction algorithms in data mining. However, its computation complexity has to be mitigated in order to apply it for real-world large scale data analysis. Many researches have been focusing on developing approximate spectral embeddings which are more efficient, but meanwhile far less effective. This paper proposes Diverse Power Iteration Embeddings (DPIE), which not only retains the similar efficiency of power iteration methods but also produces a series of diverse and more effective embedding vectors. We test this novel method by applying it to various data mining applications (e.g. Clustering, anomaly detection and feature selection) and evaluating their performance improvements. The experimental results show our proposed DPIE is more effective than popular spectral approximation methods, and obtains the similar quality of classic spectral embedding derived from eigen-decompositions. Moreover it is extremely fast on big data applications. For example in terms of clustering result, DPIE achieves as good as 95% of classic spectral clustering on the complex datasets but 4000+ times faster in limited memory environment.
Hao Huang 0007, Shinjae Yoo, Dantong Yu, Hong Qin 0001
ICDM3
2014 Noise-Resistant Unsupervised Feature Selection via Multi-perspective Correlations
abstract
Unsupervised feature selection is an important issue for high dimensional dataset analysis. However popular methods are susceptible to noisy instances (observations) or noisy features. We propose a noise-resistant feature selection algorithm by capturing multi-perspective correlations. Our proposed approach, called Noise-Resistant Unsupervised Feature Selection (NRFS), is based on multi-perspective correlation that reflects the importance of feature with respect to noise-resistant representative instances and various global trends from spectral decomposition. In this way, the model concisely captures a wide variety of local patterns. Experimental results demonstrate the effectiveness of our algorithm.
Hao Huang 0007, Shinjae Yoo, Dantong Yu, Hong Qin 0001
ICDM3
2014 Virtual data center allocation with dynamic clustering in clouds
abstract
Clouds are being widely used for leasing resources to users in the form of on-demand virtual data centers, which comprise sets of virtual machines interconnected by sets of virtual links. Given a user request for a virtual data center with specific resource requirements, a critical problem is to select a set of servers and links in the physical data center of a cloud to satisfy the request in a manner that minimizes the amount of reserved resources. In this paper, we study the main aspects of this Virtual Data Center Allocation (VDCA) problem, and decompose it into three subproblems: virtual data center clustering, virtual machine allocation, and virtual link allocation. We prove the NP-hardness of VDCA and propose an algorithm that solves the problem by dynamically clustering the requested virtual data center and jointly optimizing virtual machine and virtual link allocation. We further compare the performance and scalability of the proposed algorithm with two existing algorithms, called LoCo and SecondNet, through simulations. We demonstrate that our algorithm generates 30%-200% more revenue than LoCo and 55%-300% than SecondNet, while being up to 12 times faster.
Dimitrios Katramatos, Dantong Yu
IPCCC3
2014 Physics-Based Anomaly Detection Defined on Manifold Space
abstract
Current popular anomaly detection algorithms are capable of detecting global anomalies but often fail to distinguish local anomalies from normal instances. Inspired by contemporary physics theory (i.e., heat diffusion and quantum mechanics), we propose two unsupervised anomaly detection algorithms. Building on the embedding manifold derived from heat diffusion, we devise Local Anomaly Descriptor (LAD), which faithfully reveals the intrinsic neighborhood density. It uses a scale-dependent umbrella operator to bridge global and local properties, which makes LAD more informative within an adaptive scope of neighborhood. To offer more stability of local density measurement on scaling parameter tuning, we formulate Fermi Density Descriptor (FDD), which measures the probability of a fermion particle being at a specific location. By choosing the stable energy distribution function, FDD steadily distinguishes anomalies from normal instances with any scaling parameter setting. To further enhance the efficacy of our proposed algorithms, we explore the utility of anisotropic Gaussian kernel (AGK), which offers better manifold-aware affinity information. We also quantify and examine the effect of different Laplacian normalizations for anomaly detection. Comprehensive experiments on both synthetic and benchmark datasets verify that our proposed algorithms outperform the existing anomaly detection algorithms.
Hao Huang 0007, Hong Qin 0001, Shinjae Yoo, Dantong Yu
ACM Trans. Knowl. Discov. Data4
2013 Characterization of Input/Output Bandwidth Performance Models in NUMA Architecture for Data Intensive Applications
abstract
Data-intensive applications frequently rely on multicore computer systems, in which Non-Uniform Memory Access (NUMA) is a dominant architecture. To transfer data into and out from these high-performance computers becomes a bottleneck, and thus it is crucial to understand their I/O performance characteristics. However, the complexity in NUMA architecture presents a new challenge in modeling its I/O access cost, and thus lead to difficulties in configuring proper processor and memory affinity. In this paper, we show that existing NUMA experimental methods and metrics are inappropriate on contemporary high-end systems. We characterize a state-of-the-art NUMA host, and propose, to the best of our knowledge, the first methodology to simulate I/O operations using memory semantics, and model the I/O bandwidth performance. Our methodology is thoroughly tested and validated by mapping multiple parallel I/O streams to different sets of hardware components (CPU, memory, network cards, and SSDs) and by measuring the performance of each mapping. The experimental results and analysis reveal that our methodology can dramatically reduce characterization workload, accurately estimate the overall I/O performance, and effectively mitigate resource contention among I/O tasks.
Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi
ICPP3
2013 Design and performance evaluation of NUMA-aware RDMA-based end-to-end data transfer systems
abstract
Data-intensive applications place stringent requirements on the performance of both back-end storage systems and front-end network interfaces. However, for ultra high-speed data transfer, for example, at 100 Gbps and higher, the effects of multiple bottlenecks along a full end-to-end path, have not been resolved efficiently. In this paper, we describe our implementation of an end-to-end data transfer software at such high-speeds. At the back-end, we construct a storage area network with the iSCSI protocols, and utilize efficient RDMA technology. At the front-end, we design network communication software to transfer data in parallel, and utilize NUMA techniques to maximize the performance of multiple network interfaces. We demonstrate that our system can deliver the full 100 Gbps end-to-end data transfer throughput. The software product is tested rigorously and demonstrated applicable to supporting various data-intensive applications that constantly move bulk data within and across data centers.
Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi
SC3
2013 Distributed Throughput Optimization for Large-Scale Scientific Workflows Under Fault-Tolerance Constraint
Chase Qishi Wu, Xin Liu 0056, Dantong Yu
J. Grid Comput.4
2013 Design and testbed evaluation of RDMA-based middleware for high-performance data transfer applications
Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi
J. Syst. Softw.3
2012 Local anomaly descriptor: a robust unsupervised algorithm for anomaly detection based on diffusion space
abstract
Current popular anomaly detection algorithms are capable of detecting global anomalies but oftentimes fail to distinguish local anomalies from normal instances. This paper aims to improve unsupervised anomaly detection via the exploration of physics-based diffusion space. Building upon the embedding manifold derived from diffusion maps, we devise Local Anomaly Descriptor (LAD) whose originality results from faithfully preserving intrinsic and informative density-relevant neighborhood information. This robust and effective algorithm is designed with a weighted umbrella Laplacian operator to bridge global and local properties. To further enhance the efficacy of our proposed algorithm, we explore the utility of anisotropic Gaussian kernel (AGK) which can offer better manifold-aware affinity information. Comprehensive experiments on both synthetic and UCI real datasets verify that our LAD outperforms existing anomaly detection algorithms.
Hao Huang 0007, Hong Qin 0001, Shinjae Yoo, Dantong Yu
CIKM4
2012 A New Anomaly Detection Algorithm Based on Quantum Mechanics
abstract
The primary originality of this paper lies at the fact that we have made the first attempt to apply quantum mechanics theory to anomaly (outlier) detection in high-dimensional datasets for data mining. We propose Fermi Density Descriptor (FDD) which represents the probability of measuring a fermion at a specific location for anomaly detection. We also quantify and examine different Laplacian normalization effects and choose the best one for anomaly detection. Both theoretical proof and quantitative experiments demonstrate that our proposed FDD is substantially more discriminative and robust than the commonly-used algorithms.
Hao Huang 0007, Hong Qin 0001, Shinjae Yoo, Dantong Yu
ICDM4
2012 Protocols for wide-area data-intensive applications: design and performance issues
abstract
Providing high-speed data transfer is vital to various data-intensive applications.While there have been remarkable technology advances to provide ultra-high-speed network bandwidth, existing protocols and applications may not be able to fully utilize the bare-metal bandwidth due to their inefficient design.We identify the same problem remains in the field of Remote Direct Memory Access (RDMA) networks.RDMA offloads TCP/IP protocols to hardware devices.However, its benefits have not been fully exploited due to the lack of efficient software and application protocols, in particular in wide-area networks.In this paper, we address the design choices to develop such protocols.We describe a protocol implemented as part of a communication middleware.The protocol has its flow control, connection management, and task synchronization.It maximizes the parallelism of RDMA operations.We demonstrate its performance benefit on various local and wide-area testbeds, including the DOE ANI testbed with RoCE links and InfiniBand links.
Yufei Ren, Dantong Yu, Shudong Jin, Thomas G. Robertazzi, Brian Tierney, Eric Pouyoul
SC3
2012 Design and implementation of an intelligent end-to-end network QoS system
abstract
End-to-End guaranteed network QoS is a requirement for predictable data transfers between geographically distant end-hosts. Existing QoS systems, however, do not have the capability/intelligence to decide what resources to reserve and which paths to choose when there are multiple and flexible resource reservation requests. In this paper, we design and implement an intelligent system that can guarantee end-to-end network QoS for multiple flexible reservation requests. At the heart of this system is a polynomial time algorithm called resource reservation and path construction (RRPC). The RRPC algorithm schedules multiple flexible end-to-end data transfer requests by jointly optimizing the path construction and bandwidth reservation along these paths. We show that constructing such schedules is NP-hard. We implement our intelligent QoS system, and present the results of deployment on real world production networks (ESnet and Internet2). Our implementation does not require modifications or new software to be deployed on the routers within network.
Sushant Sharma, Dimitrios Katramatos, Dantong Yu
SC3
2011 Improving Throughput and Reliability of Distributed Scientific Workflows for Streaming Data Processing
abstract
With the advent of next-generation scientific applications, the workflow-based computing technology has become an indispensable research method for managing and streamlining large-scale distributed data processing. This paper investigates a problem of mapping distributed workflows for streaming data processing in faulty networks where nodes and links are subject to probabilistic failures. We formulate this problem as a bi-objective optimization problem in terms of both throughput and reliability, and propose a decentralized layer-oriented method to achieve high throughput for smooth data flow while satisfying a prespecified overall failure rate bound for a guaranteed level of reliability. The superiority of the proposed mapping solution is illustrated by both extensive simulation-based performance comparisons with existing algorithms and experimental results from a real-life scientific workflow deployed in wide-area networks.
Chase Qishi Wu, Xin Liu 0056, Dantong Yu
HPCC4
2011 A Robust Clustering Algorithm Based on Aggregated Heat Kernel Mapping
abstract
Current spectral clustering algorithms suffer from both sensitivity to scaling parameter selection in similarity matrix construction, and data perturbation. This paper aims to improve robustness in clustering algorithms and combat these two limitations based on heat kernel theory. Heat kernel can statistically depict traces of random walk, so it has an intrinsic connection with diffusion distance, with which we can ensure robustness during any clustering process. By integrating heat distributed along time scale, we propose a novel method called Aggregated Heat Kernel (AHK) to measure the distance between each point pair in their eigen space. Using AHK and Laplace-Beltrami Normalization (LBN) we are able to apply an advanced noise-resisting robust spectral mapping to original dataset. Moreover it offers stability on scaling parameter tuning. Experimental results show that, compared to other popular spectral clustering methods, our algorithm can achieve robust clustering results on both synthetic and UCI real datasets.
Hao Huang 0007, Shinjae Yoo, Hong Qin 0001, Dantong Yu
ICDM4
2011 End-to-end network QoS via scheduling of flexible resource reservation requests
abstract
Modern data-intensive applications move vast amounts of data between multiple locations around the world. To enable predictable and reliable data transfers, next generation networks allow such applications to reserve network resources for exclusive use. In this paper, we solve an important problem (called SMR3) to accommodate multiple and concurrent network reservation requests between a pair of end sites. Given the varying availability of bandwidth within the network, our goal is to accommodate as many reservation requests as possible while minimizing the total time needed to complete the data transfers. First, we prove that SMR3 is an NP-hard problem. Then, we solve it by developing a polynomial-time heuristic called RRA. The RRA algorithm hinges on an efficient mechanism to accommodate large number of requests in an iterative manner. Finally, we show via numerical results that RRA constructs schedules that accommodate significantly larger number of requests compared to other, seemingly efficient, heuristics.
Sushant Sharma, Dimitrios Katramatos, Dantong Yu
SC3
2006 TeraPaths: End-to-End Network Path QoS Configuration Using Cross-Domain Reservation Negotiation
abstract
TeraPaths is a DOE MICS/SciDAC-fundedproject conceived to address the needs of the high energy and nuclear physics scientific community for effectively protecting data flows of various levels of priority through modern high-speed networks. TeraPaths is rapidly evolving from a last-mile, LAN QoS provider to a distributed end-to-end network path QoS negotiator through multiple administrative domains. Developed as a web service-based software system, TeraPaths automates the establishment of network paths with QoS guarantees between end sites by configuring their corresponding LANs and requesting MPLS paths through WANs on behalf of end users. The primary mechanism for the creation of such paths is the negotiation and placement of advance reservations across all involved domains. This paper describes the status of the project, our experiences so far, as well as the directions of our continued work.
Bruce Gibbard, Dimitrios Katramatos, Dantong Yu, Shawn McKee
BROADNETS3
2004 The Grid2003 Production Grid: Principles and Practice
Ian T. Foster, Jerry Gieraltowski, Scott Gose, Natalia Maltsev, Edward N. May, Alexis A. Rodriguez, Dinanath Sulakhe, A. Vaniachine, Jim Shank, Saul Youssef, David Adams, Richard Baker 0003, Wensheng Deng, Dantong Yu, Iosif Legrand, Conrad Steenberg, M. Anzar Afaq, Eileen Berman, James Annis, L. A. T. Bauerdick, Michael Ernst, Ian Fisk, Lisa Giacchetti, Gregory E. Graham, Anne Heavey, Joseph Kaiser, Nickolai Kuropatkin, Ruth Pordes, Vijay Sekhri, John Weigand, Yujun Wu, Keith Baker, Lawrence Sorrillo, John Huth, Matthew Allen, Leigh Grundhoefer, John Hicks, Fred Luehring, Steve Peck, Robert Quick, Stephen C. Simms, George Fekete, Jan vandenBerg, Kihyeon Cho, Kihwan Kwon, Dongchul Son, Hyoungwoo Park, Shane Canon, Keith R. Jackson, David E. Konerding, Jason Lee 0001, Doug Olson, Iwona Sakrejda, Brian Tierney, Mark Green 0001, Russ Miller, James Letts, Terrence Martin, David Bury, Catalin Dumitrescu, Daniel Engh, Robert W. Gardner, Marco Mambelli, Yuri Smirnov, Jens-S. Vöckler, Michael Wilde, Yong Zhao 0009, Paul Avery, Richard Cavanaugh, Bockjoo Kim, Craig Prescott, Jorge Rodríguez 0002, Andrew Zahn, Shawn McKee, Christopher T. Jordan, James E. Prewett, Timothy L. Thomas, Horst Severini, Ben Clifford, Ewa Deelman, Larry Flon, Carl Kesselman, Gaurang Mehta, Nosa Olomu, Karan Vahi, Kaushik De, Patrick McGuigan, Mark Sosebee, Dan Bradley, Peter Couvares, Alan DeSmet, Carey Kireyev, Erik Paulson 0001, Alain J. Roy, Scott Koranda, Brian Moe, Bobby Brown, Paul Sheldon
HPDC15
2003 ClusterTree: Integration of Cluster Representation and Nearest-Neighbor Search for Large Data Sets with High Dimensions
abstract
We introduce the ClusterTree, a new indexing approach for representing clusters generated by any existing clustering approach. A cluster is decomposed into several subclusters and represented as the union of the subclusters. The subclusters can be further decomposed, which isolates the most related groups within the clusters. A ClusterTree is a hierarchy of clusters and subclusters which incorporates the cluster representation into the index structure to achieve effective and efficient retrieval. Our cluster representation is highly adaptive to any kind of cluster. It is well accepted that most existing indexing techniques degrade rapidly as the dimensions increase. The ClusterTree provides a practical solution to index clustered data sets and supports the retrieval of the nearest-neighbors effectively without having to linearly scan the high-dimensional data set. We also discuss an approach to dynamically reconstruct the ClusterTree when new data is added. We present the detailed analysis of this approach and justify it extensively with experiments.
Dantong Yu, Aidong Zhang 0001
IEEE Trans. Knowl. Data Eng.1
2002 FindOut: Finding Outliers in Very Large Datasets
Dantong Yu, Gholamhosein Sheikholeslami, Aidong Zhang 0001
Knowl. Inf. Syst.1
2000 ACQ: An Automatic Clustering and Querying Approach for Large Image Databases
Dantong Yu, Aidong Zhang 0001
ICDE1
1999 Efficiently Detecting Arbitrary Shaped Clusters in Image Databases
abstract
Image databases contain data with high dimensions. Finding interesting patterns in these databases poses a very challenging problem because of the scalability, lack of domain knowledge and complex structures of the embedded clusters. High dimensionality adds severely to the scalability problem. In this paper, we introduce WaveCluster/sup +/, a novel approach to apply wavelet-based techniques for clustering high-dimensional data. Using a hash-based data structure to represent the data set, we offer a detailed technique to apply a wavelet transform on the hashed feature space. We demonstrate that the cost of clustering can be reduced dramatically yet maintaining all the advantages of wavelet-based clustering. This hash-based data representation can be applied for any grid-based clustering approaches. The experimental results show the effectiveness and efficiency of our method on high-dimensional data sets.
Dantong Yu, Surojit Chatterjee, Aidong Zhang 0001
ICTAI1
1999 ACQ: an automatic clustering and querying approach for large image databases
abstract
Article Free Access Share on ACQ: an automatic clustering and querying approach for large image databases Authors: Dantong Yu Department of Computer Science and Engineering, State University of New York at Buffalo, Buffalo, NY Department of Computer Science and Engineering, State University of New York at Buffalo, Buffalo, NYView Profile , Aidong Zhang Department of Computer Science and Engineering, State University of New York at Buffalo, Buffalo, NY Department of Computer Science and Engineering, State University of New York at Buffalo, Buffalo, NYView Profile Authors Info & Claims MULTIMEDIA '99: Proceedings of the seventh ACM international conference on Multimedia (Part 2)October 1999 Pages 95–98https://doi.org/10.1145/319878.319904Published:01 October 1999Publication History 2citation298DownloadsMetricsTotal Citations2Total Downloads298Last 12 Months5Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Dantong Yu, Aidong Zhang 0001
ACM Multimedia (2)1
1999 GiView: A Multi-Resolution Geographical Data Retrieval System
abstract
The State University of New York at Buffalo is developing a geographic image retrieval system termed GiView. From a large-scale geographic image database, the system retrieves relevant images that contain parts similar to a given query image. The system extracts texture and color features of geographic images. One key design of the system is a practical image decomposition scheme called nona-tree that uses a hierarchical over-lapping window structure. This treatment offers a compromise to alleviate the difficulty of object recognition which is still an open problem in computer vision. The system also takes the advantage of multi-resolution properties of wavelet transforms so that the system can extract texture features in different scales. This design takes into account the multi-scale nature of geographic images and improves retrieval efficiency. A third key design of the system is to cluster database images based on their similarity to a set of geographic image templates. The templates are selected according to needs of particular projects, and multiple cluster approaches are implemented to accommodate different needs.
Dantong Yu, Lei Zhu 0013, Aidong Zhang 0001, Ling Bian
SSDBM1