Baijian Yang 0001

dblp:184/3556 · also Baijian Justin Yang · DBLP profile ↗
← Back
41ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0003-4440-3701ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 6 since 2021Computer networks · 9 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Security and privacy · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
YearPublicationVenuePosition
2025 Cognitive Load-Based Affective Workload Allocation for Multihuman Multirobot Teams
abstract
The interaction and collaboration between humans and multiple robots represent a novel field of research known as human multirobot systems. Adequately designed systems within this field allow teams composed of both humans and robots to work together effectively on tasks, such as monitoring, exploration, and search and rescue operations. This article presents a deep reinforcement learning-based affective workload allocation controller specifically for multihuman multirobot teams. The proposed controller can dynamically reallocate workloads based on the performance of the operators during collaborative missions with multirobot systems. The operators' performances are evaluated through the scores of a self-reported questionnaire (i.e., subjective measurement) and the results of a deep learning-based cognitive workload prediction algorithm that uses physiological and behavioral data (i.e., objective measurement). To evaluate the effectiveness of the proposed controller, we conduct an exploratory user experiment with various allocation strategies. The user experiment uses a multihuman multirobot CCTV monitoring task as an example and carry out comprehensive real-world experiments with 32 human subjects for both quantitative measurement and qualitative analysis. Our results demonstrate the performance and effectiveness of the proposed controller and highlight the importance of incorporating both subjective and objective measurements of the operators' cognitive workload as well as seeking consent for workload transitions, to enhance the performance of multihuman multirobot teams.
Wonse Jo, Baijian Yang 0001, Daniel Foti, Mohammad Rastgaar, Byung-Cheol Min
IEEE Trans. Hum. Mach. Syst.3
2024 Unsupervised Machine Learning for Detecting and Locating Human-Made Objects in 3D Point Cloud
abstract
3D point clouds are unstructured, sparse, and irregular data collected by airborne LiDAR systems over a geological region. Laser pulses emitted from the systems reflect off objects both on and above the ground, resulting in data with the longitude, latitude, and elevation of the points, and the corresponding laser pulse strengths. Ground filtering is important. The aim is to partition the points into ground and non-ground subsets. In addition, this research introduces a novel task: detecting and identifying human-made objects amidst natural tree structures. The task is performed on the non-ground subset derived given by the ground filtering stage. Marked Point Fields (MPFs) are used to these tasks. The proposed methodology consists of three stages: ground filtering, local information extraction (LIE), and clustering. In the ground filtering stage, a statistical method called One-Sided Regression (OSR) is devised to overcome the limitations of prior ground filtering methods on uneven terrains. In the LIE stage, a kernel-based method for the Hessian matrix of the MPF is developed. In the clustering stage, the Gaussian Mixture Model (GMM) is applied to the results of the LIE for partitioning the non-ground points into trees and human-made objects. The underlying assumption is that LiDAR points from trees exhibit a three-dimensional distribution, while those from human-made objects follow a two-dimensional distribution. The Hessian matrix of the MPF effectively captures the difference. Experimental results demonstrate that the proposed ground filtering method outperforms previous techniques, and the LIE method successfully distinguishes between points representing trees and human-made objects.
Huyunting Huang, Tonglin Zhang, Baijian Yang 0001, Jin Wei-Kocsis, Songlin Fei
IEEE Big Data4
2024 Label-Efficient Video Object Segmentation With Motion Clues
abstract
Video object segmentation (VOS) plays an important role in video analysis and understanding, which in turn facilitates a number of diverse applications, including video editing, video rendering, and augmented reality / virtual reality. However, existing deep learning-based approaches rely heavily on a large number of pixel-wise annotated video frames to achieve promising results, which is notoriously laborious and costly. To address this, in this paper, we formulate unsupervised video object detection by exploring simulated dense labels and explicit motion clues. Specifically, we first propose an effective video label generator network based on the sparsely annotated frames and the flow motion between them. It can largely alleviate our dependence and limitation on the sparse labels. Furthermore, we propose a transformer-based architecture to model the appearance and motion clues simultaneously with the cross-attention module, in order to maximally overcome non-linear motion with potential occlusions. Extensive experiments show that the proposed method outperforms recent VOS methods on four popular benchmarks (i.e., DAVIS-16, FBMS, Youtube-VOS and SegTrack-v2). Moreover, the proposed method can be further applied to a wide range of wild scenes such as wild forests and animals. Because of its effectiveness and generalization, we believe that our method could serve as a useful basis for alleviating the dependence on dense annotation in video data.
Yawen Lu, Jie Zhang 0066, Su Sun, Zhiwen Cao, Songlin Fei, Baijian Yang 0001, Victor Y. Chen
IEEE Trans. Circuits Syst. Video Technol.7
2023 Improved Clustering Using Nice Initialization
abstract
Popular clustering methods, such as the k-means and the Gaussian mixture model (GMM), are composed of an initialization stage and an iteration stage. Although this is well known, the role of the two stages has not been well studied for existing clustering methods yet. To understand this issue, the current research investigates well-known existing k-means and the GMM methods. The finding is that the initialization stage is more critical than the iteration stage. Based on this issue, the research develops improved k-means and GMM methods by using a nice initialization, which significantly enhances the performance of the corresponding methods. Experiments based on Monte Carlo simulations show that the proposed improved k-means and GMM methods can completely identify the true clusters if they are well separated from each, but this cannot be achieved by the existing k-means and GMM methods. Ex-periments based on a real-world dataset show that the proposed methods can be efficiently combined with dimension-reduction techniques for clustering high-dimensional massive data.
Huyunting Huang, Ziyang Tang, Tonglin Zhang, Baijian Yang 0001
GLOBECOM4
2023 SpaRx: elucidate single-cell spatial heterogeneity of drug responses for personalized treatment
abstract
Spatial cellular authors heterogeneity contributes to differential drug responses in a tumor lesion and potential therapeutic resistance. Recent emerging spatial technologies such as CosMx, MERSCOPE and Xenium delineate the spatial gene expression patterns at the single cell resolution. This provides unprecedented opportunities to identify spatially localized cellular resistance and to optimize the treatment for individual patients. In this work, we present a graph-based domain adaptation model, SpaRx, to reveal the heterogeneity of spatial cellular response to drugs. SpaRx transfers the knowledge from pharmacogenomics profiles to single-cell spatial transcriptomics data, through hybrid learning with dynamic adversarial adaption. Comprehensive benchmarking demonstrates the superior and robust performance of SpaRx at different dropout rates, noise levels and transcriptomics coverage. Further application of SpaRx to the state-of-the-art single-cell spatial transcriptomics data reveals that tumor cells in different locations of a tumor lesion present heterogenous sensitivity or resistance to drugs. Moreover, resistant tumor cells interact with themselves or the surrounding constituents to form an ecosystem for drug resistance. Collectively, SpaRx characterizes the spatial therapeutic variability, unveils the molecular mechanisms underpinning drug resistance and identifies personalized drug targets and effective drug combinations.
Ziyang Tang, Xiang Liu 0016, Zuotian Li, Tonglin Zhang, Baijian Yang 0001, Jing Su 0003, Qianqian Song 0002
Briefings Bioinform.5
2023 spaCI: deciphering spatial cellular communications through adaptive graph model
abstract
Cell-cell communications are vital for biological signalling and play important roles in complex diseases. Recent advances in single-cell spatial transcriptomics (SCST) technologies allow examining the spatial cell communication landscapes and hold the promise for disentangling the complex ligand-receptor (L-R) interactions across cells. However, due to frequent dropout events and noisy signals in SCST data, it is challenging and lack of effective and tailored methods to accurately infer cellular communications. Herein, to decipher the cell-to-cell communications from SCST profiles, we propose a novel adaptive graph model with attention mechanisms named spaCI. spaCI incorporates both spatial locations and gene expression profiles of cells to identify the active L-R signalling axis across neighbouring cells. Through benchmarking with currently available methods, spaCI shows superior performance on both simulation data and real SCST datasets. Furthermore, spaCI is able to identify the upstream transcriptional factors mediating the active L-R interactions. For biological insights, we have applied spaCI to the seqFISH+ data of mouse cortex and the NanoString CosMx Spatial Molecular Imager (SMI) data of non-small cell lung cancer samples. spaCI reveals the hidden L-R interactions from the sparse seqFISH+ data, meanwhile identifies the inconspicuous L-R interactions including THBS1-ITGB1 between fibroblast and tumours in NanoString CosMx SMI data. spaCI further reveals that SMAD3 plays an important role in regulating the crosstalk between fibroblasts and tumours, which contributes to the prognosis of lung cancer patients. Collectively, spaCI addresses the challenges in interrogating SCST data for gaining insights into the underlying cellular communications, thus facilitates the discoveries of disease mechanisms, effective biomarkers and therapeutic targets.
Ziyang Tang, Tonglin Zhang, Baijian Yang 0001, Jing Su 0003, Qianqian Song 0002
Briefings Bioinform.3
2022 The Alcoholic Hepatitis Network Research Data Commons (ARDaC): Design and Development
Jing Su 0003, Nanxin Jin, Zuotan Li, Carla Kettler, Bruce Barton, Greg Puetz, Chi Mai Nguyen, Donna McGrath, Victor Y. Chen, Baijian Yang 0001, Vijay Shah, Svetlana Radaeva, Samer Gawrieh, Wanzhu Tu
AMIA10
2022 Cybersecurity Education in the Age of Artificial Intelligence: A Novel Proactive and Collaborative Learning Paradigm
abstract
This Innovative Practice Work-in-Progress paper presents a virtual, proactive, and collaborative learning paradigm that can engage learners with different backgrounds and enable effective retention and transfer of the multidisciplinary AI-cybersecurity knowledge. While progress has been made to better understand the trustworthiness and security of artificial intelligence (AI) techniques, little has been done to translate this knowledge to education and training. There is a critical need to foster a qualified cybersecurity workforce that understands the usefulness, limitations, and best practices of AI technologies in the cybersecurity domain. To address this import issue, in our proposed learning paradigm, we leverage multidisciplinary expertise in cybersecurity, AI, and statistics to systematically investigate two cohesive research and education goals. First, we develop an immersive learning environment that motivates the students to explore AI/machine learning (ML) development in the context of real-world cybersecurity scenarios by constructing learning models with tangible objects. Second, we design a proactive education paradigm with the use of hackathon activities based on game-based learning, lifelong learning, and social constructivism. The proposed paradigm will benefit a wide range of learners, especially underrepresented students. It will also help the general public understand the security implications of AI. In this paper, we describe our proposed learning paradigm and present our current progress of this ongoing research work. In the current stage, we focus on the first research and education goal and have been leveraging cost-effective Minecraft platform to develop an immersive learning environment where the learners are able to investigate the insights of the emerging AI/ML concepts by constructing related learning modules via interacting with tangible AI/ML building blocks.
Jin Wei-Kocsis, Moein Sabounchi, Baijian Yang 0001, Tonglin Zhang
FIE3
2022 Uncovering Brain Differences in Preschoolers and Young Adolescents with Autism Spectrum Disorder Using Deep Learning
abstract
Identifying brain abnormalities in autism spectrum disorder (ASD) is critical for early diagnosis and intervention. To explore brain differences in ASD and typical development (TD) individuals by detecting structural features using T1-weighted magnetic resonance imaging (MRI), we developed a deep learning-based approach, three-dimensional (3D)-ResNet with inception (I-ResNet), to identify participants with ASD and TD and propose a gradient-based backtracking method to pinpoint image areas that I-ResNet uses more heavily for classification. The proposed method was implemented in a preschool dataset with 110 participants and a public autism brain imaging data exchange (ABIDE) dataset with 1099 participants. An extra epilepsy dataset with 200 participants with clear degeneration in the parahippocampal area was applied as a verification and an extension. Among the datasets, we detected nine brain areas that differed significantly between ASD and TD. From the ROC in PASD and ABIDE, the sensitivity was 0.88 and 0.86, specificity was 0.75 and 0.62, and area under the curve was 0.787 and 0.856. In a word, I-ResNet with gradient-based backtracking could identify brain differences between ASD and TD. This study provides an alternative computer-aided technique for helping physicians to diagnose and screen children with an potential risk of ASD with deep learning model.
Ziyang Tang, Nanxin Jin, Qiansu Yang, Tiefang Liu, Jianxing Hu, Sijun Liu, Jingru Hao, Baijian Yang 0001
Int. J. Neural Syst.17
2021 DenserNet: Weakly Supervised Visual Localization Using Multi-Scale Feature Aggregation
abstract
In this work, we introduce a Denser Feature Network(DenserNet) for visual localization. Our work provides three principal contributions. First, we develop a convolutional neural network (CNN) architecture which aggregates feature maps at different semantic levels for image representations. Using denser feature maps, our method can produce more key point features and increase image retrieval accuracy. Second, our model is trained end-to-end without pixel-level an-notation other than positive and negative GPS-tagged image pairs. We use a weakly supervised triplet ranking loss to learn discriminative features and encourage keypoint feature repeatability for image representation. Finally, our method is computationally efficient as our architecture has shared features and parameters during forwarding propagation. Our method is flexible and can be crafted on a light-weighted backbone architecture to achieve appealing efficiency with a small penalty on accuracy. Extensive experiment results indicate that our method sets a new state-of-the-art on four challenging large-scale localization benchmarks and three image retrieval benchmarks with the same level of supervision. The code is available at https://github.com/goodproj13/DenserNet
Dongfang Liu, Yiming Cui 0002, Liqi Yan, Christos Mousas, Baijian Yang 0001, Victor Y. Chen
AAAI5
2021 PASC CKD: revealing the progression trajectories of sustained COVID-19-related renal injury using real-world evidence
Jing Su 0003, Pengyue Zhang, Zuoyi Zhang, Michael Eadon, Xiaochun Li 0003, Stanley Taylor, Travis Johnson, Zhaorui Liu, Ziyang Tang, Baijian Yang 0001, Qianqian Song 0002, Kun Huang 0001
AMIA11
2021 ADNet: Identify biomarkers of Alzheimer Disease with MRI and EMR data using Deep Neural Networks
Ziyang Tang, Qianqian Song 0002, Jing Su 0003, Baijian Yang 0001
AMIA4
2021 Anomaly detection of core failures in die casting X-ray inspection images using a convolutional autoencoder
Weitao Tang, Corey M. Vian, Ziyang Tang, Baijian Yang 0001
Mach. Vis. Appl.4
2020 CHEESE: Cyber Human Ecosystem of Engaged Security Education
abstract
This Innovative Practice Full Paper presents CHEESE, a platform for cybersecurity education that complements formal classroom instruction with hands-on experience. With the ubiquitous use of computing devices and applications today, the protection of personal and privileged information is a persistent challenge. Modern software applications are typically complex pieces of code that borrow from various preexisting software libraries. Consequently, a flaw in one piece of software can have far-reaching and often unintended security implications that malicious actors can exploit. Thus, cybersecurity education needs to be transformed from a purely academic enterprise for cybersecurity researchers into a necessary skill that is imparted to the current and future IT workforce at large. CHEESE aims to impart such skills. CHEESE is composed of CHEESEHub, a public web-platform hosting demonstrations of cybersecurity concepts, a set of lessons complementing the demonstrations, and a community-driven approach to the contribution of new demonstrations and lessons. CHEESE is intended to supplement and enhance traditional cybersecurity education with hands-on training that has been shown to improve concept retention and understanding. Instructors can incorporate CHEESE into their teaching in several ways: by utilizing one or more of the demonstrations hosted on the publicly-accessible CHEESEHub in conjunction with the web-accessible lessons; by deploying their own version of CHEESEHub with a custom set of demonstrations and lessons; or by developing their own lesson plan which borrows from and combines one or more demonstrations on CHEESEHub. The use of CHEESEHub only requires a web-browser and can hence be employed in a wide variety of educational and training settings from K-12 schools through university.
Rajesh Kalyanam, Baijian Yang 0001, Craig Willis, Mike Lambert, Christine R. Kirkpatrick
FIE2
2020 Regression PCA for Moving Objects Separation
abstract
This work proposed a new approach called regression PCA (RegPCA) for statistical machine learning and big data analyses. One of the potential use cases investigated in this work is to separate the moving objects (foreground) from the background images. This is achieved by performing regression before conducting Robust PCA (RPCA). RegPCA works well in the moving object detection task because the background information can be conceived as the regression portion of the images, while the residual portion of the regression can then be fed into RPCA to fine tune the foreground detection. The experiments show that in moving object detection problems RegPCA provides much better results than applying only RPCA, especially in color videos and when the moving objects are relatively big. Further studies are needed to leverage the interesting features of RegPCA approach and apply it to solve more real world problems.
Huyunting Huang, Xiang Liu 0016, Tonglin Zhang, Baijian Yang 0001
GLOBECOM4
2020 Low-Rank Sparse Tensor Approximations for Large High-Resolution Videos
abstract
Tensor decomposition techniques are becoming increasingly important in processing videos with large sizes and dimensions. Under the framework of CANDECOMP/PARAFAC decomposition (CPD), this work studies low-rank sparse tensor approximations (LRSTAs) to higher-order tensors. Both theoretical and practical properties are evaluated for LRSTAs to represent large high-resolution videos. The evaluation brings three major contributions of this work. Firstly, the theoretical connection between CPD for high-order tensors and traditional singular value decomposition (SVD) for matrices are established, and the tensor rank for traditional SVD is defined. This provides a theoretical basis to compare tensor-based approach against matrix-based approach under the framework of tensor decompositions. Secondly, the non-orthogonality of CPD and its implications are revealed. The solution set of an LRSTA can only be used as a whole. Thirdly, a computationally efficient algorithm is developed. Its practical properties are also investigated in object detection and recognition in high-resolution videos. The results of the experiments showed that the proposed algorithm can handle large high-resolution videos very efficiently in terms of memory allocation. Results also revealed that commonly used total variations may not be a good evaluation metric for real world applications in computer vision. LRSTAs should be evaluated using the end goal of the applications, such as the accuracy of object detection and recognition.
Xiang Liu 0016, Huyunting Huang, Weitao Tang, Tonglin Zhang, Baijian Yang 0001
ICMLA5
2020 PENet: Object Detection Using Points Estimation in High Definition Aerial Images
abstract
Aerial imagery has been increasingly adopted in mission-critical tasks, such as traffic surveillance, smart cities, and disaster assistance. However, identifying objects from aerial images faces the following challenges: 1) objects of interests are often too small and too dense relative to the images; 2) objects of interests are often in different relative sizes; and 3) the number of objects in each category is imbalanced. A novel network structure, Points Estimated Network (PENet), is proposed in this work to answer these challenges. PENet uses a Mask Resampling Module (MRM) to augment the imbalanced datasets, a coarse anchor-free detector (CPEN) to effectively predict the center points of the small object clusters, and a fine anchor-free detector FPEN to locate the precise positions of the small objects. An adaptive merge algorithm Non-maximum Merge (NMM) is implemented in CPEN to address the issue of detecting dense small objects, and a hierarchical loss is defined in FPEN to further improve the classification accuracy. Our extensive experiments on aerial datasets visDrone [1] and UAVDT [2] showed that PENet achieved higher precision results than existing state-of-the-art approaches. Our best model achieved 8.7% improvement on visDrone and 20.3% on UAVDT.
Ziyang Tang, Xiang Liu 0016, Baijian Yang 0001
ICMLA3
2020 Visual Localization for Autonomous Driving: Mapping the Accurate Location in the City Maze
abstract
Accurate localization is a foundational capacity, required for autonomous vehicles to accomplish other tasks such as navigation or path planning. It is a common practice for vehicles to use GPS to acquire location information. However, the application of GPS can result in severe challenges when vehicles run within the inner city where different kinds of structures may shadow the GPS signal and lead to inaccurate location results. To address the localization challenges of urban settings, we propose a novel feature voting technique for visual localization. Different from the conventional front-view-based method, our approach employs views from three directions (front, left, and right) and thus significantly improves the robustness of location prediction. In our work, we craft the proposed feature voting method into three state-of-the-art visual localization networks and modify their architectures properly so that they can be applied for vehicular operation. Extensive field test results indicate that our approach can predict location robustly even in challenging inner-city settings. Our research sheds light on using the visual localization approach to help autonomous vehicles to find accurate location information in a city maze, within a desirable time constraint. The source code is available at github.com/HappyDonkey13/Visual- Localization-for- Autonomous-Driving.
Dongfang Liu, Yiming Cui 0002, Baijian Yang 0001, Victor Y. Chen
ICPR5
2019 Sparse Block Regression (SBR) for Big Data with Categorical Variables
abstract
Categorical variables are nominal variables that classify observations by groups. The treatment of categorical variables in regression is a well-studied yet vital problem, with the most popular solution to perform a one hot encoding. However, challenges arise if a categorical variable has millions of levels. It will cause the memory needed for the computation far exceeds the total available memory in a given computer system or even a computer cluster. Thus, it is fair to state that one hot encoding approach has its limitations when a categorical variable has a large number of levels. The common workaround is the sparse matrix approach because it requires much fewer resources to cache the dummy variables. However, existing sparse matrix approaches are still not sufficient to handle extreme cases when a categorical variable has millions of levels. For instance, the number of subnets in network traffic analyses can easily exceeds tens of millions. In this paper, we proposed an innovative approach called sparse block regression (SBR) to address this challenge. SBR constructs a sparse block matrix using sufficient statistics. The benefits include but not limited to: 1) overcome the memory barrier issue caused by one hot encoding, 2) obtain multiple models with a single scan of data stored in the secondary storage; and 3) update the models with simple matrix operations. The study compared proposed SBR against conventional sparse matrix approaches. The experiments proved that SBR can efficiently and accurately solve the regression problem with large category number. Compared to the sparse matrix approach, SBR saved 90% memory in size during the computation.
Xiang Liu 0016, Huyunting Huang, Ziyang Tang, Tonglin Zhang, Baijian Yang 0001
IEEE BigData5
2019 Multiple Learning for Regression in Big Data
abstract
Regression problems that have closed-form solutions are well understood and can be easily implemented when the dataset is small enough to be all loaded into the RAM. Challenges arise when data are too big to be stored in RAM to compute the closed form solutions. Many techniques were proposed to overcome or alleviate the memory barrier problem but the solutions are often local optima. In addition, most approaches require loading the raw data to the memory again when updating the models. Parallel computing clusters are often expected in practice if multiple models need to be computed and compared. We propose multiple learning approaches that utilize an array of sufficient statistics (SS) to address the aforementioned big data challenges. The memory oblivious approaches break the memory barrier when computing regressions with closed-form solutions, including but not limited to linear regression, weighted linear regression, linear regression with Box-Cox transformation (Box-Cox regression) and ridge regression models. The computation and update of the SS arrays can be handled at per row level or per mini-batch level. And updating a model is as easy as matrix addition and subtraction. Furthermore, the proposed approaches also enable the computational parallelizability of multiple models because multiple SS arrays for different models can be computed simultaneously with a single pass of slow disk I/O access to the dataset. We implemented our approaches on Spark and evaluated over the simulated datasets. Results showed our approaches can achieve exact solutions of multiple models. The training time saved compared to the traditional methods is proportional to the number of models need to be investigated.
Xiang Liu 0016, Ziyang Tang, Huyunting Huang, Tonglin Zhang, Baijian Yang 0001
ICMLA5
2019 Computer Vision-based Algae Removal Planner for Multi-robot Teams
abstract
Water pollution has caused increased incidence of algal growth around the globe. Harmful algae blooms result in massive economic losses. In this paper, a multi-robot based task planner is designed to remove excessive algae from water bodies and to identify algae build-up so that prompt action can be taken against its accumulation. Computer vision is incorporated to enable algae detection and area estimation based on training, comparing, and evaluating various advanced deep learning models using our custom algae dataset. We further propose a novel algorithm for robot resource allocation between bounding boxes of detected algae based on multivariable optimization. This systematic solution is evaluated in a simulated environment, demonstrating how the robots are optimally assigned to the detected algae patches for algae removal.
Manoj Penmetcha, Shaocheng Luo, Arabinda Samantaray, J. Eric Dietz, Baijian Yang 0001, Byung-Cheol Min
SMC5
2018 A Study of Exact Ridge Regression for Big Data
abstract
Ridge regression is a regularization technique that can be used together with other regression algorithms to model highly correlated data. Like many other traditional techniques, ridge regression of big data version requires a large number of iterations over the dataset to converge. As the dataset cannot all be stored in memory, the dataset is split into RAM-accommodable subsets for training, however, this strategy is time-consuming for reading all subsets from hard drive to memory over and over. To overcome the memory barrier, we proposed to use working sufficient statistics to solve the problem [1]. The parameters of the working sufficient matrix is small enough to be stored in RAM all the time. They can be updated at per row level to allow online computation. This strategy only requires one iteration over the dataset. While our previous work proved its theoretical correctness, it was not clear how our innovative algorithm would work in practice. In this study, we aims to validate and evaluate the performance improvement of the algorithm we proposed in earlier work-Three sets of experiments were conducted using large data-sets published by FAA and BTS to examine the computation time, memory requirement, and the accuracy of the output. Results showed that our exact ridge aggression algorithm enjoyed many benefits, such as faster computing time, minimal memory requirements and more accurate estimates.
Wanchih Chiang, Xiang Liu 0016, Tonglin Zhang, Baijian Yang 0001
IEEE BigData4
2018 File Toolkit for Selective Analysis & Reconstruction (FileTSAR) for Large-Scale Networks
abstract
There are many challenges in digital forensic investigations involving large-scale computer networks; these include large volume of data, the limited scope of tools, the financial burdens of purchasing and licensing those tools, and identifying salient evidence from the vast amounts of network data. We have implemented a collection of open-source tools and code wrappers to provide a tool for network forensic investigators to capture, selectively analyze, and reconstruct files from network traffic. The main functions of this tool (FileTSAR) are capturing data flows and providing a mechanism to selectively reconstruct documents, images, email, and VoIP conversations. To validate the large-scale capabilities of the toolkit, we conducted a "stress test" of the system using approximately 123,500,000 packets from a collection of packet capture files totaling nearly 100GB. Additionally, sixteen (16) digital forensic examiners participated in a 3-day law enforcement training workshop for FileTSAR from across the United States; the examiners expressed substantial support for FileTSAR with large-scale investigations as well as an interest in a scaled-down version for smaller agencies with storage, budget, and back-end support limitations.
Raymond A. Hansen, Kathryn C. Seigfried-Spellar, Siddarth S. Chowdhury, Niveah Abraham, John A. Springer, Baijian Yang 0001, Marcus K. Rogers
IEEE BigData7
2018 Streaming Algorithm for Big Data Logistic Regression
abstract
This research proposed a novel fitting algorithm for big data logistic regression by combining Fisher Scoring and IRWLS. The algorithm enables streamed operation to fit and update the model at per-row level, without the need to store the entire dataset in RAM. The algorithm is fully parallelizable and was implemented on Spark for this study. A set of experiments were conducted on Spark and Scikit-Learn to compare its performance against existing approaches. The results showed that the proposed method can provide exact results, rather than approximate results, with significantly fewer iterations. It also converged in merely 2 iterations in terms of prediction accuracy. More importantly, the proposed algorithm has a very small memory footprint and completely broke the memory barrier problem that debilitated conventional logistic regression approaches. Compared to the SGD sampling approach in scikit-learn, the proposed approach also demonstrated a clear advantage in accuracy.
Baijian Yang 0001, Zhenzhi Xu, Tonglin Zhang
IEEE BigData1
2017 Finding the best box-cox transformation from massive datasets on spark
abstract
In order to find the best linear regression model or polynomial regression model that fits the data, traditional methods have to read the whole datasets repetitively and incur many unnecessary slow I/O operations. Apache Spark can train regression models significantly more efficiently with distributed clusters due to its well-crafted in-memory computing architecture. However, if the dataset itself or the temporary data during computation is even bigger for the total physical memory space of a spark system, in-memory data has to be spilled to the secondary storage (such as hard drives or solid state disks) and read it back later if it is needed. These frequent I/O operations will negatively affect the efficiency of Spark computation. Built on top of the per-row update-able data modeling concept we proposed before, this work investigated the cases of finding the best Box-Cox transformation model on a Spark system. The major contribution of this work is that the information needed to compute a linear regression model, or a polynomial regression model can be summarized in an Information Array. The size of this information array does not grow with the datasets. Rather, it is only related to the number of features and the number of models need to be considered. Because the information array is usually very small, it can be stored in memory all the time. With the propose information array approach, the best linear or polynomial regression model could be obtained after one scan of the raw data. The experiment results proved that this approach is fast and efficient on Spark. When training 41 models, the proposed Box-Cox Information Array method is about 8 times faster than the existing Spark APIs and it has better performance of prediction than using linear regression models.
Huayi Fang, Baijian Yang 0001, Tonglin Zhang
IEEE BigData2
2017 Vulnerabilities in hub architecture IoT devices
abstract
This paper introduces new methods for gaining sensitive information from Internet of Things (IoT) devices. Generally, manufacturers tend to forget about security when putting a new product on the market, this notion does not exclude IoT. This has been proven by current research using simple techniques and devices to obtain unencrypted information. Using similar attacks, the IoT hubs themselves was chosen as a target instead of the devices connected to them. Two different attacks will be used. First, using various methods for sniffing Hub traffic, we will attempt to gain credentials on the uplink side. Secondly, port scans will be used to gain any exploited services in order to obtain root access. After our attempts, it was found that IoT hubs can be very susceptible to tech savvy attackers using simple Man in the Middle attacks. If network access was achieved, the security of ones home is greatly compromised allowing intruders an 8-10 minute window to enter a house undetected. Using modest methodologies, consumers and companies can prevent future exploits.
Bogdan Alexandra Visan, Baijian Yang 0001, Anthony H. Smith, Eric T. Matson
CCNC3
2017 Finding the Best Box-Cox Transformation in Big Data with Meta-Model Learning: A Case Study on QCT Developer Cloud
abstract
Finding the best model to reveal potential relationships of a given set of data is not an easy job and often requires many iterations of trial and errors for model sections, feature selections and parameters tuning. This problem is greatly complicated in the big data era where the I/O bottlenecks significantly slowed down the time needed to finding the best model. In this article, we examine the case of Box-Cox transformation when assumptions of a regression model are violated. Specifically, we construct and compute a set of summary statistics and transformed the maximum likelihood computation into a per-role operational fashion. The innovative algorithms reduced the big data machine learning problem into a stream based small data learning problem. Once the Box-Cox information array is obtained, the optimal power transformation as well as the corresponding estimates of model parameters can be quickly computed. To evaluate the performance, we implemented the proposed Box-Cox algorithms on QCT developer cloud. Our results showed that by leveraging both the algorithms and the QCT cloud technology, find the fittest model from 101 potential parameters is much faster than the conventional approach.
Yuxiang Gao, Tonglin Zhang, Baijian Yang 0001
CSCloud3
2016 A Black-Box Self-Learning Scheduler for Cloud Block Storage Systems
abstract
A major requirement of cloud block storage services is guaranteed performance and high availability. However, offering guaranteed Service Level Agreements (SLAs) in cloud block storage services is often not straightforward. Cloud block storage performance may be affected by physical disk background operations, like garbage collection, storage cluster features, workload interference and the chraracteristcs of the workload itself. On the other hand, the underlying physical storage drives do not expose the internal states to higher level block storage service offerings. Therefore, SLAs can only be satisfied by over-provisioning the storage resources. To address this issue, we propose a self-learning scheduler that can dynamically adapt based on the workload, and efficiently provide a scheduling decision with zero knowledge of the underlying hardware. We study two candidate algorithms based on Feedback learning and Two-Phase learning. We used workloads that were deducted from real-world block-level traces of an enterprise data center, and conducted extensive simulations. Our results indicate that the self-learning scheduling approach can reduce the SLA violations by mitigating the unexpected resource fluctuation, and the scheduler can also adapt dynamically to various workloads.
Babak Ravandi, Ioannis Papapanagiotou, Baijian Yang 0001
CLOUD3
2015 Forensically Sound Retrieval and Recovery of Images from GPU Memory
Baijian Yang 0001, Marcus K. Rogers, Raymond A. Hansen
ICDF2C2
2015 A visual analytics approach to detecting server redirections and data exfiltration
abstract
How to better find potential cyberattacks is a challenging question for security researchers and practitioners. In recent years, visualization has been applied in the field of analyzing cybersecurity issues, but most work has not been able to provide better than non-visualization based techniques. In this paper, we innovatively designed a visual analytics system to allow analysts to overview network traffic and identify such suspicious such activities as server redirection attack and data exfiltration. Because of the nature of the problem, the overview design must be scalable, accurate, and fast. Through aggregating traffic data along the two dimensions of duration and payload, the system reveals key network traffic characteristics for the analyst to identify security events. The system is evaluated with the test data sets from VAST 2013 mini-challenge 3. The results are very encouraging and shed a more positive light on applying visual analytics in information security.
Baijian Yang 0001, Victor Y. Chen
ISI2
2014 Trust-based incentive mechanism to motivate cooperation in hybrid P2P networks
Chunqi Tian, Baijian Yang 0001, Jidong Zhong
Comput. Networks2
2014 A D-S evidence theory based fuzzy trust model in file-sharing P2P networks
Chunqi Tian, Baijian Yang 0001
Peer-to-Peer Netw. Appl.2
2011 R2 Trust, a reputation and risk based trust management framework for large-scale, fully decentralized overlay networks
Chunqi Tian, Baijian Yang 0001
Future Gener. Comput. Syst.2
2009 tk-coverage: Time-Based K-Coverage for Energy Efficient Monitoring
abstract
K-coverage is a classic issue in wireless sensor network (WSN) deployment. Existing works typically assumes that every position in the monitoring field is covered by at least k sensor nodes at any given time. This may not always be necessary because certain events to be monitored will last only for a short period of time. Based upon such observation, we include the time dimension in the original k-coverage problem, and denote it as tk-coverage In the context of tk-coverage, sensor nodes can apply periodical sleeping strategies to save energy use. A corresponding tk-coverage (TKC) model is proposed to analyze the energy consumption and detection delay., The proposed tk-coverage can balance energy consumption and detection delay by adjusting the sleeping strategies of sensor nodes. Comprehensive simulations have been conducted to validate the effectiveness and demonstrate the efficiency of this solution.
Zheng Yang 0002, Bin Xu 0001, Saier Ye, Baijian Yang 0001
ICPADS4
2009 LORP: a load-balancing based optimal routing protocol for sensor networks with bottlenecks
abstract
The performance of wireless sensor networks (WSNs) is tightly coupled with the geometric environment in which sensors are deployed. In a practical environment, bottleneck regions, for example bridges, may exist due to the existence of physical obstacles or energy depletion. In this paper, we propose a load-balancing based optimal routing protocol (LORP). By finding the boundaries of holes in a sensor network with bottlenecks, LORP first identifies the bridges in the sensor field using our MACB algorithm. A centralized routing algorithm, "balance-first" routing is then employed to prolong the lifetime of a WSN with bottlenecks. Theoretical analysis prove that LORP can improve the load distribution among different bridges, increase the lifetime of a WSN, and enhance the quality of network services. This conclusion is reinforced in our simulation results.
Lijie Xu, Guihai Chen, Xinchun Yin, Panlong Yang, Baijian Yang 0001
WCNC5
2009 On the reliability of large-scale distributed systems - A topological view
Yuan He 0004, Hao Ren 0013, Yunhao Liu 0001, Baijian Yang 0001
Comput. Networks4
2008 On the Reliability of Large-Scale Distributed Systems A Topological View
abstract
In large-scale, self-organized and distributed systems, such as peer-to-peer (P2P) overlays and wireless sensor networks (WSN), a small proportion of nodes are likely to be more critical to the system's reliability than the others. This paper focuses on detecting cut vertices so that we can either neutralize or protect these critical nodes. Detection of cut vertices is trivial if the global knowledge of the whole system is known but it is very challenging when the global knowledge is missing. In this paper, we propose a completely distributed scheme where every single node can determine whether it is a cut vertex or not. In addition, our design can also confine the detection overhead to a constant instead of being proportional to the size of a network. The correctness of this algorithm is theoretically proved and a number of performance measures are verified through trace driven simulations.
Yuan He 0004, Hao Ren 0013, Yunhao Liu 0001, Baijian Yang 0001
ICPP4
2006 A Framework to Provide Trust and Incentive in CROWN Grid for Dynamic Resource Management
abstract
In order to maximize resource utilization as well as providing trust management in grid, we propose a novel framework - Trust-Incentive Resource Management (TIM). Having child model, club model, bid model and trust model, TIM dynamically manages grid resource by integrating values of prices, trust, and incentive. In this mechanism, providers set the price according to demand and supply, and consumers maximize the surplus upon budget and deadline. A weighted voting scheme is also proposed to secure the grid system by declining the join request from malicious nodes. A TIM prototype has been successfully implemented in a real grid system, CROWN grid. We evaluate the proposed approach through comprehensive experiments and achieve improved results in resource allocation efficiency, system completion time, and aggregated resource utilization.
Jinpeng Huai, Yunhao Liu 0001, Li Lin 0008, Baijian Yang 0001
ICCCN5
2004 Efficient Gnutella-like P2P Overlay Construction
abstract
Without assuming any knowledge of the underlying physical topology, the conventional P2P mechanisms are designed to randomly choose logical neighbors, causing a serious topology mismatch problem between the P2P overlay network and the underlying physical network. This mismatch problem incurs a great stress in the Internet infrastructure and adversely restraints the performance gains from the various search or routing techniques. In order to alleviate the mismatch problem, reduce the unnecessary traffic and response time, we propose two schemes, namely, location-aware topology matching (LTM) and scalable bipartite overlay (SBO) techniques. Both LTM and SBO achieve the above goals without bringing any noticeable extra overheads. More-over, both techniques are scalable because the P2P over-lay networks are constructed in a fully distributed manner where global knowledge of the network is not necessary. This paper demonstrates the effectiveness of LTM and SBO, and compares the performance of these two approaches through simulation studies. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Yunhao Liu 0001, Li Xiao 0001, Lionel M. Ni, Baijian Yang 0001
NPC4
2004 Multicasting in MPLS domains
Baijian Yang 0001, Prasant Mohapatra
Comput. Commun.1
2002 Multicasting in differentiated service domains
abstract
Advances in the areas of QoS and IP multicasting have necessitated the need of integration of these two important features of Internet. Differentiated services (DiffServ) has been proposed as a scalable solution for supporting QoS in the Internet. Coexistence of multicasting and DiffServ is promising since the DiffServ model can provide a scalable framework and may reduce the computational complexity to locate a QoS-satisfied multicast tree. We first identify the problems of provisioning multicasting in DiffServ domains. Next, we propose an efficient DiffServ-Aware Multicasting (DAM) scheme which has three novel features: weighted traffic conditioning (WTC), receiver-initiated marking (RIM) scheme, and Heterogeneous DSCP Headers encapsulation (HDE). The proposed technique solves many problems with the integration of DiffServ and multicasting while accommodating heterogeneous QoS requirements. The framework is scalable, flexible, and feasible. Performance evaluation through analyses and simulations demonstrate conformance of the QoS requirements and the potential benefits of DAM.
Baijian Yang 0001, Prasant Mohapatra
GLOBECOM1