VLDB 2026 Research / reviewers in the wild / expert
Biao Yin
dblp:141/2035
· DBLP profile ↗
17ranked-venue papers
11as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 9 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ensembles of Statistically Independent Models: A Method for Semi-Supervised Domain AdaptationabstractOver recent years, many sophisticated methods for the critical problem of domain adaptation (DA) have been proposed in the literature. Unfortunately, there is no single DA model that always outperforms others regardless of datasets and setups. Motivated by this, we investigate the effect of ensembles of (imperfect) classic DA models – each pre-trained on distinct folds of data – on the ability to learn a simple yet strong DA model for solving a new target problem. We show that a simple fusion of approximately statistically independent models, modeled with a simple linear ensemble model under semi-supervision (1-/3-shot examples), boosts the performance over each individual DA model. We demonstrate that, when taken together, the rich variety of ensembled models succeeds to better cover the feature space for domain adaptation, achieving improvements over state-of-the-art semi-supervised single-source performance on benchmark data sets: DomainNet (+8.8% 1-shot, +3.1% 3-shot) and Office-31 (+15.8% 1-shot, +8.4% 3-shot). We release our code for reproducibility. Nicholas Josselyn, Walter Gerych, Biao Yin, Elke A. Rundensteiner |
ICMLA | 3 |
| 2024 | A multi-channel spatial information feature based human pose estimation algorithmabstractAbstract Human pose estimation is an important task in computer vision, which can provide key point detection of human body and obtain bone information. At present, human pose estimation is mainly utilized for detection of large targets, and there is no solution for detection of small targets. This paper proposes a multi-channel spatial information feature based human pose (MCSF-Pose) estimation algorithm to address the issue of medium and small targets inaccurate detection of human key points in scenarios involving occlusion and multiple poses. The MCSF-Pose network is a bottom-up regression network. Firstly, an UP-Focus module is designed to expand the feature information while reducing parameter computation during the up-sampling process. Then, the channel segmentation strategy is adopted to cut the features, and the feature information of multiple dimensions is retained through different convolutional groups, which reduces the parameter lightweight network model and makes up for the loss of the feature information associated with the depth of the network. Finally, the three-layer PANet structure is designed to reduce the complexity of the model. With the aid of the structure, it also to improve the detection accuracy and anti-interference ability of human key points. The experimental results indicate that the proposed algorithm outperforms YOLO-Pose and other human pose estimation algorithms on COCO2017 and MPII human pose datasets. Yinghong Xie, Yan Hao, Biao Yin |
Cybersecur. | 5 |
| 2023 | MOSS: AI Platform for Discovery of Corrosion-Resistant MaterialsabstractAmid corrosion degradation of metallic structures causing expenses nearing 3 trillion or 4% of the GDP annually along with major safety risks, the adoption of AI technologies for accelerating the materials science life-cycle for developing materials with better corrosive properties is paramount. While initial machine learning models for corrosion assessment are being proposed in the literature, their incorporation into end-to-end tools for field experimentation by corrosion scientists remains largely unexplored. To fill this void, our university data science team in collaboration with the materials science unit at the Army Research Lab have jointly developed MOSS, an innovative AI-based digital platform to support material science corrosion research. MOSS features user-friendly iPadOS app for in-field corrosion progression data collection, deep-learning corrosion assessor, robust data repository system for long-term experimental data modeling, and visual analytics web portal for material science research. In this demonstration, we showcase the key innovations of the MOSS platform via use cases supporting the corrosion exploration processes, with the promise of accelerating the discovery of new materials. We open a MOSS video demo at: https://www.youtube.com/watch?v=CzcxMMRsxkE Biao Yin, Nicholas Josselyn, Elke A. Rundensteiner, Thomas A. Considine, John V. Kelley, Berend Christopher Rinderspacher, Robert E. Jensen, James F. Snyder |
CIKM | 1 |
| 2023 | AlloyGAN: Domain-Promptable Generative Adversarial Network for Generating Aluminum Alloy MicrostructuresabstractThe global metal market, expected to exceed $18.5 trillion by 2030, faces costly inefficiencies from defects in alloy manufacturing. Although microstructure analysis has improved alloy performance, current numerical models struggle to accurately simulate solidification. In this research, we thus introduce AlloyGAN - the first domain-driven Conditional Generative Adversarial Network (cGAN) involving domain prior for generating alloy microstructures of previously not considered chemical and manufactural compositions. AlloyGAN improves cGAN process by involving prior factors from solidification reaction to generate scientifically valid images of alloy microstructure given basic alloy manufacturing compositions. It achieves a faster and equally accurate alternative to traditional material science methods for assessing alloy microstructures. We contribute (1) a novel Alloy-GAN design for rapid alloy optimization; (2) unique methods that inject prior knowledge of the chemical reaction into cGAN-based models; and (3) metrics from machine learning and chemistry for generation evaluation. Our approach highlights the promise of GAN-based models in the scientific discovery of materials. AlloyGAN has successfully transitioned into an AIGC startup with a core focus on model-generated metallography. We open its interactive demo at: https://deepalloy.com/ Biao Yin, Yangyang Fan, Nicholas Josselyn, Elke A. Rundensteiner |
ICMLA | 1 |
| 2023 | DeepSC-Edge: Scientific Corrosion Segmentation with Edge-Guided and Class-Balanced LossesabstractCorrosion is a prevalent issue in numerous industrial fields, causing expenses nearing $3 trillion or 4% of the GDP annually with safety threats and environmental pollution. To timely qualify and validate new corrosion-inhibiting materials on a large scale, accurate and efficient corrosion assessment is crucial. Yet it is hindered by a lack of automatic tools for expert-level corrosion segmentation of material science experimental images. Developing such tools is challenging due to limited domain-valid data, image artifacts visually similar to corrosion, various corrosion morphology, strong class imbalance, and millimeter-precision corrosion boundaries. To help the community address these challenges, we curate the first expert-level segmentation annotations for a real-world image dataset [1] for scientific corrosion segmentation. In addition, we design a deep learning based model, called DeepSC-Edge that achieves guidance of ground-truth edge learning by adopting a novel loss that avoids over-fitting to edges. It also is enriched by integrating a class-balanced loss that improves segmentation with small area but crucial edges of interest for scientific corrosion assessment. Our dataset and methods pave the way to advanced deep-learning models for corrosion assessment and generation – promoting new research to connect computer vision and material science discovery. Once the appropriate approvals have been cleared, we expect to release the code and data at: https://arl.wpi.edu/ Biao Yin, Nicholas Josselyn, Thomas A. Considine, John V. Kelley, Berend Christopher Rinderspacher, Robert E. Jensen, James F. Snyder, Elke A. Rundensteiner |
ICMLA | 1 |
| 2022 | An Empirical Study of Domain Adaptation: Are We Really Learning Transferable Representations?abstractDeep learning often relies on the availability of a large amount of high-quality labeled data, which can be very limited in novel domains. To address such data scarcity, domain adaptation is one promising approach that allows for deep networks to leverage large amounts of available data from a source domain to enhance the model’s efficacy on the target domain of interest. However, while there is a plethora of alternate models for domain adaptation proposed over many years in the literature, there is a dearth of studies that objectively compare the relative effectiveness of these models in a rigorous, empirical study. To fill this gap, we provide a thorough, unbiased, empirical study of five state-of-the-art (SOTA) deep domain adaptation models proposed over the past 6 years whose codes are publicly available. Models are evaluated on the complex and diverse domain adaptation tasks featured in the DomainNet benchmark dataset as well as the popular Office-31 dataset. Our results suggest that (1) all 5 models perform similarly, on average, and do not even significantly beat the oldest model, and (2) counter to their intended purpose, the transfer loss functions in the literature do not contribute significantly to learning transferable representations. Our observations suggest that domain adaptation research needs to more thoroughly compare newly proposed models against existing works, along with assessing their loss functions’ utility thoroughly. Our code and data splits are made public for reproducibility of results by the community. Nicholas Josselyn, Biao Yin, Elke A. Rundensteiner |
IEEE Big Data | 2 |
| 2022 | Transferring Indoor Corrosion Image Assessment Models to Outdoor Images via Domain AdaptationabstractCorrosion of materials impacts critical economic sectors from infrastructure, transportation, defense, health, to the environment. The development of safe anti-corrosive materials is thus an important area of study in materials science. Corrosion science of preparing materials and then monitoring their corrosion under adverse conditions is labor intensive, time consuming, and extremely costly. While deep learning has become popular in automating various engineering tasks, the development of deep models for corrosion assessment is lacking. We are the first to study deep domain adaptation (DA) models for the automated assessment of the corrosion status of anti-corrosive materials. Corrosion data, i.e., photographic images of treated corroding materials, is abundant when produced in artificially controlled laboratory settings, while corrosion image data sets from rich natural outdoor environments are more challenging to produce and thus much smaller. We leverage the more readily available indoor corrosion data to train a classifier and then transfer it via deep domain adaptation to also perform well on the small yet more realistic outdoor corrosion image data set – without requiring target labels. We empirically compare 5 popular domain adaptation models on real-world corrosion image data sets. Our study finds that DA achieves 27% improvement in test accuracy compared to the performance of the no-DA baseline for classifying real-world outdoor corrosion data. Nicholas Josselyn, Biao Yin, Thomas A. Considine, John V. Kelley, Berend Christopher Rinderspacher, Robert E. Jensen, James F. Snyder, Elke A. Rundensteiner |
ICMLA | 2 |
| 2021 | Corrosion Image Data Set for Automating Scientific Assessment of Materials
Biao Yin, Nicholas Josselyn, Thomas A. Considine, John V. Kelley, Berend Christopher Rinderspacher, Robert E. Jensen, James F. Snyder, Elke A. Rundensteiner |
BMVC | 1 |
| 2021 | NeurJudge: A Circumstance-aware Neural Framework for Legal Judgment PredictionabstractLegal Judgment Prediction is a fundamental task in legal intelligence of the civil law system, which aims to automatically predict the judgment results of multiple subtasks, such as charge, law article, and term of penalty prediction. Existing studies mainly focus on the impact of the entire fact description on all subtasks. They ignore the practical judicial scenario, where judges adopt circumstances of crime (i.e., various parts of the fact) to decide judgment results. To this end, in this paper, we propose a circumstance-aware legal judgment prediction framework (i.e., NeurJudge) by exploring circumstances of crime. Specifically, NeurJudge utilizes the results of intermediate subtasks to separate the fact description into different circumstances and exploits them to make the predictions of other subtasks. In addition, considering the popularity of confusing verdicts (i.e., charges and law articles), we further extend NeurJudge to a more comprehensive framework which is denoted by NeurJudge+. Particularly, NeurJudge+ utilizes a label embedding method to incorporate the semantics of labels (i.e., charges and law articles) into facts to generate more expressive fact representations for confusing verdicts problems. Extensive experimental results on two real-world datasets clearly validate the effectiveness of our proposed frameworks. Linan Yue, Qi Liu 0003, Binbin Jin, Han Wu 0002, Kai Zhang 0038, Yanqing An, Mingyue Cheng 0004, Biao Yin, Dayong Wu |
SIGIR | 8 |
| 2020 | Human Attention Maps for Text Classification: Do Humans and Neural Networks Focus on the Same Words?abstractMotivated by human attention, computational attention mechanisms have been designed to help neural networks adjust their focus on specific parts of the input data.While attention mechanisms are claimed to achieve interpretability, little is known about the actual relationships between machine and human attention.In this work, we conduct the first quantitative assessment of human versus computational attention mechanisms for the text classification task.To achieve this, we design and conduct a large-scale crowd-sourcing study to collect human attention maps that encode the parts of a text that humans focus on when conducting text classification.Based on this new resource of human attention dataset for text classification, YELP-HAT, collected on the publicly available YELP dataset, we perform a quantitative comparative analysis of machine attention maps created by deep learning models and human attention maps.Our analysis offers insights into the relationships between human versus machine attention maps along three dimensions: overlap in word selections, distribution over lexical categories, and context-dependency of sentiment polarity.Our findings open promising future research opportunities ranging from supervised attention to the design of human-centric attentionbased explanations. Cansu Sen, Thomas Hartvigsen, Biao Yin, Xiangnan Kong, Elke A. Rundensteiner |
ACL | 3 |
| 2020 | Corrosion Assessment: Data Mining for Quantifying Associations between Indoor Accelerated and Outdoor Natural TestsabstractMaterial scientists study corrosion degradation of metallic structures due to its heavy economic and maintenance burdens. Assessing corrosion is both time consuming and labor intensive when utilizing outdoor tests under natural exposure conditions. Accelerated indoor corrosion tests are conducted in laboratory settings by material scientists to gage performance in a shorter period of time than outdoor tests. However, these indoor tests do not always correlate well with the actual performance in outdoor environments. Thus, there is a need to apply data-science methodologies to analyze and establish quantitative associations between indoor accelerated and outdoor exposure assessments to optimize artificially accelerated methods. We work with material experimental records including images, notes, meta-data and human-rated assessments collected over years. We apply data mining methods, such as Canonical Correlation Analysis (CCA) and its variants, to extract latent associations with corresponding feature mappings between the indoor and outdoor assessments in a projected data subspace. We find that CCA provides not only reliable quantitative associations but also interpretive mappings between the indoor and outdoor assessments. Further, three methods applied to control bias due to distinctive coating system stack up yield Pearson's correlation coefficients ranging from 0.70 to 0.83 in the optimized CCA subspaces, respectively. Moreover, predicting outdoor assessments from the CCA-projected indoor test data compared to using the original indoor data is shown to result in an increase in accuracy from 88% to 90% - confirming the effectiveness of our approach. Lastly, our results facilitate interesting domain-relevant discovery that could potentially lead to experimentalists better understanding corrosion resistance. Biao Yin, Thomas A. Considine, Fatemeh Emdad, John V. Kelley, Robert E. Jensen, Elke A. Rundensteiner |
IEEE BigData | 1 |
| 2019 | Recursive least-squares temporal difference learning for adaptive traffic signal control at intersection
Biao Yin, Mahjoub Dridi, Abdellah El Moudni |
Neural Comput. Appl. | 1 |
| 2017 | Causal Forest vs. Naive Causal Forest in Detecting Personalization: An Empirical Study in ASSISTments
Biao Yin, Anthony Botelho, Thanaporn Patikorn, Neil T. Heffernan |
EDM | 1 |
| 2017 | Observing Personalizations in Learning: Identifying Heterogeneous Treatment Effects Using Causal TreesabstractThe incorporation of computer-based platforms in the classroom has introduced the ability to conduct numerous randomized control trials at scale with student-level randomization. Such systems are able to collect vast amounts of data on each student while completing work in the classroom and at home. It is often the case, however, that the effects of these trials are reported across all students, ignoring the potential for personalized learning. Personalized learning, or the observation of heterogeneous treatment effects, considers that the effects of a studied learning intervention may differ for individual students; while an intervention may work well for low-performing students, for example, it may have no effect for higher performing students. Personalized learning can lead to better instructional practices that maximizes the learning benefits for each individual student, and with the use of computer-based platforms, such individualized instruction is made feasible at scale. In this work we use a causal decision tree to observe treatment effects in 9 experiments run in the ASSISTments online learning platform. Biao Yin, Thanaporn Patikorn, Anthony Botelho, Neil T. Heffernan |
L@S | 1 |
| 2015 | Adaptive Traffic Signal Control for Multi-intersection Based on Microscopic ModelabstractIn this paper, we mainly propose an online learning method for adaptive traffic signal control in a multi-intersection system. The method uses approximate dynamic programming (ADP) to achieve a near-optimal solution of the signal optimization in a distributed network, which is modeled in a microscopic way. The traffic network loading model and traffic signal control model are presented to serve as the basis of discrete-time control environment. The learning process of linear function approximation in ADP approach adopts the tunable parameters of the traffic states, including the vehicle queue length and the signal indication. ADP overcomes the computational complexity, which usually appears in large scale problems solved by exact algorithms, such as dynamic programming. Moreover, the proposed adaptive phase sequence (APS) mode improves the performance by comparing with other control methods. The results in simulation show that our method performs quite well for adaptive traffic signal control problem. Biao Yin, Mahjoub Dridi, Abdellah El Moudni |
ICTAI | 1 |
| 2014 | Traffic control model and algorithm based on decomposition of MDPabstractIn this paper, a new method based on decomposition of Markov Decision Process (MDP) for traffic control at isolated intersection is proposed. The conflicting traffic flows should be grouped into different combinations which can occupy the conflict zone concurrently. Thus, for purpose of traffic delay reduction, the optimal policy of signal sequence and duration among different combinations is studied by minimizing the number of vehicles waiting in the queue. In order to reduce the computation of probabilities in large state transition matrix, the decomposition method proposed classifies states into several parts as rule of traffic signal transition. Each part contains the vehicle states in all traffic flows. This method firstly achieves the full-states calculation in stochastic traffic control system. Moreover, the simulation results indicate that MDP approach is more efficient to improve the performance of traffic control than other comparing methods, such as fixed-time control and actuated control. Biao Yin, Mahjoub Dridi, Abdellah El Moudni |
CoDIT | 1 |
| 2013 | Markov Decision Process for Traffic Control at an Isolated IntersectionabstractIn this paper, a new control method based on Markov Decision Process for a simple traffic intersection is proposed. At the intersection, different directions have the arrival traffic flows and each has a single queue. The non-conflicting flows constitute a combination which possesses the optimal sequence to occupy the intersection. The iterative algorithm is used to minimize the number of vehicles waiting in the queue and the vehicles average waiting time. However, in this dynamic control, the more important thing is that the fairness among the combinations for maintaining the system stability is taken into account. As shown by the comparison of simulation results, it is necessary to use fairness control when global optimization is acquired. Statistical analysis of simulation results at different arrival rates indicates that the approach has a good performance. Biao Yin, Mahjoub Dridi, Abdellah El Moudni |
ICTAI | 1 |