Yafan Huang

dblp:268/9947 · DBLP profile ↗
← Back
27ranked-venue papers
11as first author
27since 2021 · last 2026
0000-0001-7370-6766ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 9 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference
abstract
Large language models (LLMs) are increasingly integrated into high-performance computing (HPC) workflows, accelerating scientific discovery through diverse perspectives such as code generation and domain-specific decision-making. Yet, how soft errors propagate and affect LLM inference remains largely unexplored. To bridge this gap, we present a comprehensive study on error propagation in LLM inference, enabled by our proposed LLMFI, a configurable and deterministic fault-injection framework. Using LLMFI, we systematically inject faults across three open-weighted LLMs and thirteen representative tasks, covering reasoning, multilingual, mathematical, and coding domains. In addition, we conduct fine-grained case studies that reveal critical vulnerability patterns. Overall, our study yields 17 takeaways that advance the understanding of error propagation in LLM inference and introduces four low-overhead directions to improve reliability through software-only modification, offering practical guidance for future error detection and mitigation.
Yafan Huang, Sheng Di, Guanpeng Li
ICS1
2026 GPZ: GPU-Accelerated Lossy Compressor for Particle Data
abstract
Particle-based simulations and point-cloud applications generate massive, irregular datasets that challenge storage, I/O, and real-time analytics. Traditional compression techniques struggle with irregular particle distributions and GPU architectural constraints, often resulting in limited throughput and suboptimal compression ratios. In this paper, we present GPZ, a high-performance, error-bounded lossy compressor designed specifically for large-scale particle data on modern GPUs. GPZ employs a novel four-stage parallel pipeline that synergistically balances high compression efficiency with the architectural demands of massively parallel hardware. We introduce a suite of targeted optimizations for computation, memory access, and GPU occupancy that enable GPZ to achieve near-hardware-limit throughput. We conduct an extensive evaluation on three distinct GPU architectures (workstation, data center, and edge) using six large-scale, real-world scientific datasets from four distinct domains. The results demonstrate that GPZ consistently and significantly outperforms four state-of-the-art GPU compressors, delivering up to 8x higher end-to-end throughput while achieving superior compression ratios and data quality.
Yafan Huang, Zhuoxun Yang, Sheng Di, Boyuan Zhang 0002, Jiajun Huang 0001, Jinyang Liu 0003, Jiannan Tian, Guanpeng Li, Fengguang Song, Hanqi Guo 0001, Franck Cappello, Kai Zhao 0008
ICS2
2026 Near-Zero Cost KV Cache Compression for Large Language Model Inference
Boyuan Zhang 0002, Yafan Huang, Shihui Song, Jinda Jia, Chengming Zhang 0006, Zhi Zhang 0005
IPDPS3
2025 Understanding Error Sensitivity in Checkpointing for Linear System Solvers
abstract
Fault tolerance in large-scale iterative solvers is critical, yet traditional checkpointing methods often impose significant storage overhead. In this study, we evaluate the compression error introduced from lossy compression in the Conjugate Gradient method. We systematically investigate how various compressor configurations such as numerical error bounds, error modes, and prediction algorithms influence compressed checkpoint size, compression error, and extra iterations after recovery. Our analysis reveals key trade-offs between storage efficiency and recovery overhead.
Yafan Huang, Guanpeng Li
HPDC2
2025 ghZCCL: Advancing GPU-aware Collective Communications with Homomorphic Compression
abstract
In the exascale computing era, collective communication has emerged as a significant bottleneck for GPU-based applications, as network bandwidth lags behind rapid GPU advancements.While traditional GPU-aware approaches employ error-bounded lossy compression to mitigate this issue, they incur substantial decompression-operation-compression (DOC) overhead.To overcome these limitations, we introduce ghZCCL, a first-ever GPU-aware homomorphic compression-accelerated collective communications library that enables direct computation and communication on compressed data, eliminating the DOC workflow.We design the first GPU homomorphic compressor, surpassing the fastest existing GPU lossy compressor, cuSZp2, by 3.47-3.89×for DOC workloads.We also propose co-design strategies to further optimize GPU-aware collective communications with homomorphic compression.Experiments on up to 512 NVIDIA A100 GPUs show that ghZCCL outperforms three state-of-the-art communication libraries-gZCCL, NCCL, and Cray MPI-by achieving speedups of up to 2.29×, 5.81×, and 188×, respectively, while maintaining high data accuracy.
Jiajun Huang 0001, Sheng Di, Yafan Huang, Zizhong Chen, Franck Cappello, Yanfei Guo, Rajeev Thakur
ICS3
2025 Pushing the Limits of GPU Lossy Compression: A Hierarchical Delta Approach
Boyuan Zhang 0002, Yafan Huang, Sheng Di, Fengguang Song, Guanpeng Li, Franck Cappello
ICS2
2025 A Memory-Efficient and Computation-Balanced Lossy Compressor on Wafer-Scale Engine
abstract
Cerebras system has demonstrated immense potential across various scientific domains. However, modern scientific simulations frequently generate vast volumes of data in a short time, leading to bottlenecks in runtime performance and memory footprint. While an ultra-fast error-bounded lossy compressor can mitigate such limitations with high compression ratios and guaranteed data quality, deploying it into Cerebras dataflow architecture poses significant difficulties. Specifically, Cerebras faces memory challenges, such as the absence of shared memory and limited local memory, alongside computational challenges, including specialized parallelism and sensitivity to imbalanced workloads. In this work, we propose CERESZII, an error-bounded lossy compressor that computes within Cerebras system. CereSZ-II addresses these challenges with a carefully optimized four-stage compression workflow, consisting of Pre-quantization, Lightweight Prediction, Fixed-size Huffman Encoding, and Spatial-aware Offset Computation, ensuring both memory efficiency and computational balance. Evaluation of several real-world scientific datasets shows that CERESZ-II achieves over 800 GB/s throughput, delivering high compression ratios and reliable reconstructed data quality.
Shihui Song, Robert Underwood, Sheng Di, Yafan Huang, Peng Jiang 0004, Franck Cappello
IPDPS4
2025 What to Support When You're Compressing: The State of Practice Gaps and Opportunities for Scientific Data Compression
abstract
Over the last nearly 20 years, lossy compression has become an essential aspect of HPC applications’ data pipelines, allowing them to overcome limitations in storage capacity and bandwidth and, in some cases, increase computational throughput and capacity. However, with the adoption of lossy compression comes the requirement to assess and control the impact lossy compression has on scientific outcomes. In this work, we take a major step forward in describing the state of practice and by characterizing workloads. We examine applications’ needs and compressors’ capabilities across 9 different supercomputing application domains. We present 24 takeaways that provide best practices for applications, operational impacts for facilities achieving compressed data, and gaps in application needs not addressed by production compressors that point towards opportunities for future compression research.
Franck Cappello, Robert Underwood, Yuri Alexeev, Allison H. Baker, Ebru Bozdag, Martin Burtscher, Kyle Chard, Sheng Di, Kyle Gerard Felker, Paul Christopher O'Grady, Hanqi Guo 0001, Yafan Huang, Peng Jiang 0004, Sian Jin, Petter Johansson, Shaomeng Li, Xin Liang 0001, Erik Lindahl, Peter Lindstrom 0001, Zarija Lukic, Magnus Lundborg, Danylo Lykov, Masaru Nagaso, Kento Sato, Amarjit Singh, Seung Woo Son 0001, Shihui Song, William Tang 0002, Dingwen Tao, Jiannan Tian, Kazutomo Yoshii, Kai Zhao 0008
SC12
2025 GPU Lossy Compression for HPC Can Be Versatile and Ultra-Fast
abstract
This work proposes VGC, a versatile and ultra-fast GPU lossy compression framework designed to address the growing data challenges in high-performance computing (HPC). VGC captures dimension information in scientific data and supports three compression algorithms, achieving high compression ratios across diverse HPC domains. Built with a highly optimized GPU kernel, VGC delivers state-of-the-art throughput with error control. In addition to compression ratio and speed, VGC supports two distinctive modes that enhance its versatility. Memory-efficient Compression uses a kernel fission design to compute compressed size, allocate only the required GPU memory, and compress data without waste, effectively reducing memory footprint. Selective Decompression introduces an early stopping mechanism that enables direct access to regions of interest without decompressing the entire dataset.
Yafan Huang, Sheng Di, Guanpeng Li, Franck Cappello
SC1
2025 lsCOMP: Efficient Light Source Compression
abstract
Light source facilities, which generate X-rays for probing microstructures and dynamic processes, produce intense data streams, reaching up to 250 GB/s and projected to exceed 1 TB/s by the end of this decade. Managing such massive data poses critical challenges due to limited local processing capacity and bandwidth constraints when offloading data to HPC systems. To address these challenges, we propose lsCOMP, a GPU compressor that operates within a single kernel. lsCOMP supports both lossless and configurable lossy compression, ensuring high compression ratios and preserved data quality across diverse light source applications. On a single NVIDIA A100 GPU, lsCOMP achieves compression throughputs of 380.89 to 509.21 GB/s in lossless mode, delivering up to 20 times higher performance than industry-leading GPU compressors while achieving superior compression ratios. In lossy modes, lsCOMP further improves throughput and ratios significantly. Additionally, lsCOMP demonstrates versatile performance across various integer datasets and supports TB/s-level random access throughput.
Yafan Huang, Sheng Di, Robert Underwood, Peco Myint, Miaoqi Chu, Guanpeng Li, Nicholas Schwarz, Franck Cappello
SC1
2024 CereSZ: Enabling and Scaling Error-bounded Lossy Compression on Cerebras CS-2
abstract
Today's scientific applications running on supercomputers produce large volumes of data, leading to critical data storage and communication challenges. To tackle the challenges, error-bounded lossy compression is commonly adopted since it can reduce data size drastically within a user-defined error threshold. Previous work has shown that compression techniques can significantly reduce the storage and I/O overhead while retaining good data quality. However, the existing compressors are mainly designed for CPU and GPU. As new AI chips are being incorporated into supercomputers and increasingly used for accelerating scientific computing, there is a growing demand for efficient data compression on the new architecture. In this paper, we propose an efficient lossy compressor, CereSZ, based on the Cerebras CS-2 system. The compression algorithm is mapped onto Cerebras using both data parallelism and pipeline parallelism. In order to achieve a balanced workload on each processing unit, we propose an algorithm to evenly distribute the pipeline stages. Our experiments with six scientific datasets demonstrate that CereSZ can achieve a throughput from 227.93 GB/s to 773.8 GB/s, 2.43x to 10.98x faster than existing GPU compressors.
Shihui Song, Yafan Huang, Peng Jiang 0004, Xiaodong Yu 0001, Weijian Zheng, Sheng Di, Qinglei Cao, Yunhe Feng, Franck Cappello
HPDC2
2024 gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
abstract
GPU-aware collective communication has become a major bottleneck for modern computing platforms as GPU computing power rapidly rises. A traditional approach is to directly integrate lossy compression into GPU-aware collectives, which can lead to serious performance issues such as underutilized GPU devices and uncontrolled data distortion. In order to address these issues, in this paper, we propose gZCCL, a first-ever general framework that designs and optimizes GPU-aware, compression-enabled collectives with an accuracy-aware design to control error propagation. To validate our framework, we evaluate the performance on up to 512 NVIDIA A100 GPUs with real-world applications and datasets. Experimental results demonstrate that our gZCCL-accelerated collectives, including both collective computation (Allreduce) and collective data movement (Scatter), can outperform NCCL as well as Cray MPI by up to 4.5 × and 28.7 ×, respectively. Furthermore, our accuracy evaluation with an image-stacking application confirms the high reconstructed data quality of our accuracy-aware framework.
Jiajun Huang 0001, Sheng Di, Xiaodong Yu 0001, Jinyang Liu 0003, Yafan Huang, Kenneth Raffenetti, Hui Zhou 0012, Kai Zhao 0008, Xiaoyi Lu 0001, Zizhong Chen, Franck Cappello, Yanfei Guo, Rajeev Thakur
ICS6
2024 POSTER: Optimizing Collective Communications with Error-bounded Lossy Compression for GPU Clusters
abstract
GPU-aware collective communication has become a major bottleneck for modern computing platforms as GPU computing power rapidly rises. To address this issue, traditional approaches integrate lossy compression directly into GPU-aware collectives, which still suffer from serious issues such as underutilized GPU devices and uncontrolled data distortion. In this paper, we propose GPU-LCC, a general framework that designs and optimizes GPU-aware, compression-enabled collectives with well-controlled error propagation. To validate our framework, we evaluate the performance on up to 64 NVIDIA A100 GPUs with real-world applications and datasets. Experimental results demonstrate that our GPU-LCC-accelerated collective computation (Allreduce), can outperform NCCL as well as Cray MPI by up to 3.4× and 18.7×, respectively. Furthermore, our accuracy evaluation with an image-stacking application confirms the high reconstructed data quality of our accuracy-aware framework.
Jiajun Huang 0001, Sheng Di, Xiaodong Yu 0001, Jinyang Liu 0003, Yafan Huang, Kenneth Raffenetti, Hui Zhou 0012, Kai Zhao 0008, Zizhong Chen, Franck Cappello, Yanfei Guo, Rajeev Thakur
PPoPP6
2024 cuSZ-i: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
abstract
Error-bounded lossy compression is a critical technique for significantly reducing scientific data volumes. Compared to CPU-based compressors, GPU-based compressors exhibit substantially higher throughputs, fitting better for today’s HPC applications. However, the critical limitations of existing GPU-based compressors are their low compression ratios and qualities, severely restricting their applicability. To overcome these, we introduce a new GPU-based error-bounded scientific lossy compressor named CUSZ-i, with the following contributions: (1) A novel GPU-optimized interpolation-based prediction method significantly improves the compression ratio and decompression data quality. (2) The Huffman encoding module in CUSZ-i is optimized for better efficiency. (3) CUSZ-i is the first to integrate the NVIDIA Bitcomp-lossless as an additional compression-ratio-enhancing module. Evaluations show that CUSZ-i significantly outperforms other latest GPU-based lossy compressors in compression ratio under the same error bound (hence, the desired quality), showcasing a 476% advantage over the second-best. This leads to CUSZ-i’s optimized performance in several real-world use cases.
Jinyang Liu 0003, Jiannan Tian, Shixun Wu, Sheng Di, Boyuan Zhang 0002, Robert Underwood, Yafan Huang, Jiajun Huang 0001, Kai Zhao 0008, Guanpeng Li, Dingwen Tao, Zizhong Chen, Franck Cappello
SC7
2024 cuSZp2: A GPU Lossy Compressor with Extreme Throughput and Optimized Compression Ratio
abstract
Existing GPU lossy compressors suffer from expensive data movement overheads, inefficient memory access patterns, and high synchronization latency, resulting in limited throughput. This work proposes cuSZP2, a generic single-kernel error-bounded lossy compressor purely on GPUs designed for applications that require high speed, such as large-scale GPU simulation and large language model training. In particular, CUSZP2 proposes a novel lossless encoding method, optimizes memory access patterns, and hides synchronization latency, achieving extreme end-to-end throughput and optimized compression ratio. Experiments on NVIDIA A100 GPU with 9 real-world HPC datasets demonstrate that, even with higher compression ratios and data quality, CUSZP2 can deliver on average 332.42 and $513.04 \mathrm{~GB} / \mathrm{s}$ end-to-end throughput for compression and decompression, respectively, which is around $2 \times$ of existing pure-GPU compressors and $200 \times$ of CPU-GPU hybrid compressors.
Yafan Huang, Sheng Di, Guanpeng Li, Franck Cappello
SC1
2024 Versatile Datapath Soft Error Detection on the Cheap for HPC Applications
abstract
With the ongoing reduction in technology sizes and voltage levels, modern microprocessors are increasingly susceptible to soft errors, corrupting datapath units during program execution. While these error types have received considerable attention recently, existing solutions either confine themselves to limited scopes or incur massive overheads in performance and power consumption, hindering practical usage. In this work, we propose CONDA, a novel error detection technique based on code transformation and static program analysis, achieving versatile datapath protection at low cost. At compile time, ConDa analyzes program characteristics and transforms the original program code without complicating its control-flow and memory access patterns. At runtime, ConDa detects datapath errors with low overhead and latency. The evaluation of 38 benchmarks and a parallel HPC simulation reveals that CONDA only incurs 57.79% runtime overhead, which is 41.84% faster than existing state-of-the-art, with the same level of error detection effectiveness and low detection latency.
Yafan Huang, Sheng Di, Xiaoyi Lu 0001, Guanpeng Li
SC1
2023 Towards Improving Reverse Time Migration Performance by High-speed Lossy Compression
abstract
Seismic imaging is an exploration method for estimating the seismic characteristics of the earth's sub-surface for geologists and geophysicists. Reverse time migration (RTM) is a critical method in seismic imaging analysis. It can produce huge volumes of data that need to be stored for later use during its execution. The traditional solution transfers the vast amount of data to peripheral devices and loads them back to memory whenever needed, which may cause a substantial burden to I/O and storage space. As such, an efficient data compressor turns out to be a very critical solution. In order to get the best overall RTM analysis performance, we develop a novel hybrid lossy compression method (called HyZ), which is not only fairly fast in both compression and decompression but also has a good compression ratio with satisfactory reconstructed data quality for post hoc analysis. We evaluate several state-of-the-art error-controlled lossy compression algorithms (including HyZ, BR, SZx, SZ, SZ-Interp, ZFP, etc.) in a supercomputer. Experiments show that HyZ not only significantly improves the overall performance for RTM by 6.29∼6.60× but also obtains fairly good qualities for both RTM single snapshots and the final stacking image.
Yafan Huang, Kai Zhao 0008, Sheng Di, Guanpeng Li, Maxim Dmitriev, Thierry-Laurent D. Tonellot, Franck Cappello
CCGrid1
2023 Characterizing Runtime Performance Variation in Error Detection by Duplicating Instructions
abstract
Soft error rate has been increasing due to the shrinking size of transistors, leading to an elevated risk of catastrophic failures in modern computer systems. Error detection by duplicating instructions (EDDI) is a software-based technique to mitigate soft errors with a low runtime performance overhead and has been widely adopted in many safety- and mission-critical real-time systems such as space applications. However, these systems are commonly sensitive to runtime performance overheads the protection techniques incur. Few studies have investigated the performance of EDDI across various system designs and operational parameters, hence lacking a complete understanding in the literature. In this paper, we conduct comprehensive experiments to study the variation of EDDI runtime performance overhead and characterize the root causes. We find that there exist significant variations in performance overheads of EDDI, due to a few architectural and program-level factors. Based on the findings, we propose two practical techniques FuzzyB and Celer: FuzzyB uses an input searching technique to bound EDDI runtime performance overhead across different inputs for a given program; while Celer reduces EDDI run-time performance overheads using compiler transformation (by 25.08% reduction).
Yafan Huang, Zhengyang He, Lingda Li, Guanpeng Li
ISSRE1
2023 Demystifying and Mitigating Cross-Layer Deficiencies of Soft Error Protection in Instruction Duplication
abstract
Soft errors are prevalent in modern High-Performance Computing (HPC) systems, resulting in silent data corruptions (SDCs), compromising system reliability. Instruction duplication is a widely used software-based protection technique against SDCs. Existing instruction duplication techniques are mostly implemented at LLVM level and may suffer from low SDC coverage at assembly level. In this paper, we evaluate instruction duplication at both LLVM and assembly levels. Our study shows that existing instruction duplication techniques have protection deficiency at assembly level and are usually over-optimistic in the protection. We investigate the root-causes of the protection deficiency and propose a mitigation technique, Flowery, to solve the problem. Our evaluation shows that Flowery can effectively protect programs from SDCs evaluated at assembly level.
Zhengyang He, Yafan Huang, Hui Xu 0009, Dingwen Tao, Guanpeng Li
SC2
2023 cuSZp: An Ultra-fast GPU Error-bounded Lossy Compression Framework with Optimized End-to-End Performance
abstract
Modern scientific applications and supercomputing systems are generating large amounts of data in various fields, leading to critical challenges in data storage footprints and communication times. To address this issue, error-bounded GPU lossy compression has been widely adopted, since it can reduce the volume of data within a customized threshold on data distortion. In this work, we propose an ultra-fast error-bounded GPU lossy compressor cuSZp. Specifically, cuSZp computes the linear recurrences with hierarchical parallelism to fuse the massive computation into one kernel, drastically improving the end-to-end throughput. In addition, cuSZp adopts a block-wise design along with a lightweight fixed-length encoding and bit-shuffle inside each block such that it achieves high compression ratios and data quality. Our experiments on NVIDIA A100 GPU with 6 representative scientific datasets demonstrate that cuSZp can achieve an ultra-fast end-to-end throughput (95.53x compared with cuSZ) along with a high compression ratio and high reconstructed data quality.
Yafan Huang, Sheng Di, Xiaodong Yu 0001, Guanpeng Li, Franck Cappello
SC1
2022 Hardening selective protection across multiple program inputs for HPC applications
abstract
With the ever-shrinking size of transistors and increasing scale of applications, silent data corruptions (SDCs) have become a common yet serious issue in HPC applications. Selective instruction duplication (SID) is a popular fault-tolerance technique that can obtain a high SDC coverage with low-performance overhead, as it selects the most vulnerable parts of a program for protection with priority. However, existing studies of SID are confined to single program input in the evaluation, assuming that the error resilience of the program remains similar across inputs, leading to a drastic loss of SDC coverage from SID when the protected program runs different inputs. Hence, we proposed Sentinel, an automated compiler-based framework to mitigate the loss of SDC coverage. Evaluation results show that Sentinel can effectively mitigate the loss of SDC coverage (up to 97.00%) across multiple inputs, which significantly hardens existing SID techniques.
Yafan Huang, Shengjian Guo, Sheng Di, Guanpeng Li, Franck Cappello
PPoPP1
2022 Salus: A Novel Data-Driven Monitor that Enables Real-Time Safety in Autonomous Driving Systems
abstract
This paper proposes Salus, a data-driven real-time safety monitor, that detects and mitigates safety violations of an autonomous vehicle (AV). The key insight is that traffic situations that lead to AV safety violations fall into patterns and can be identified by learning from the safety violations of the AV. Our approach is to use machine learning (ML) techniques to model the traffic behaviors that result in safety violations in the AV, characterize their early symptoms for training a preemptive model, hence deploy and detect real-time safety violations before the actual crashes happen to the AV. In order to train our ML model, we leverage a pipeline of fuzzing techniques to tailor AV-specific safety violation symptoms and generate the training data via data argumentation techniques. Our evaluation demonstrates our proposed technique is effective in reducing over 97.2% of safety violations in industry-level autonomous driving systems, such as Baidu Apollo, with no more than 0.018 false positive values.
Yafan Huang, Guanpeng Li
QRS2
2022 Mitigating Silent Data Corruptions in HPC Applications across Multiple Program Inputs
abstract
With the ever-shrinking size of transistors, silent data corruptions (SDCs) are becoming a common yet serious issue in HPC. Selective instruction duplication (SID) is a widely used fault-tolerance technique that can obtain high SDC coverage with low performance overhead. However, existing SID methods are confined to single program input in its assessment, assuming that error resilience of a program remains similar across inputs. Nevertheless, we observe that the assumption cannot always hold, leading to a drastic loss in SDC coverage across different inputs, compromising HPC reliability. We notice that the SDC coverage loss correlates with a small set of instructions - we call them incubative instructions, which reveal elusive error propagation characteristics across multiple inputs. We propose Minpsid, an automated SID framework that automatically identifies and re-prioritizes incubative instructions in a given program to enhance SDC coverage. Evaluation shows Minpsid can effectively mitigate the loss of SDC coverage across multiple inputs.
Yafan Huang, Shengjian Guo, Sheng Di, Guanpeng Li, Franck Cappello
SC1
2022 MATT. A Multiple-instance Attention Mechanism for Long-tail Music Genre Classification
abstract
Imbalanced music genre classification is a crucial task in the Music Information Retrieval (MIR) field for identifying the long-tail, data-poor genre based on the related music audio segments, which is very prevalent in real-world scenarios. Most of the existing models are designed for class-balanced music datasets, resulting in poor performance in accuracy and generalization when identifying the music genres at the tail of the distribution. Inspired by the success of introducing Multi-instance Learning (MIL) in various classification tasks, we propose a novel mechanism named Multi-instance Attention (MATT)1to boost the performance for identifying tail classes. Specifically, we first construct the bag-level datasets by generating the album-artist pair bags. Second, we leverage neural networks to encode the music audio segments. Finally, under the guidance of a multi-instance attention mechanism, the neural network-based models could select the most informative genre to match the given music segment. Comprehensive experimental results on a large-scale music genre benchmark dataset with long-tail distribution demonstrate MATT significantly outperforms other state-of-the-art baselines.1Github: https://github.com/JohannesLiu/Music-Genre-Classification
Shihui Song, Menghua Zhang, Yafan Huang
SMC4
2022 Dynamic Entity-Based Named Entity Recognition Under Unconstrained Tagging Schemes
abstract
As increasingly more textual information becomes available, named entity recognition (NER) systems are thriving, benefiting from powerful models and expressive tagging schemes that promote the full use of diverse features at different levels. To improve performance, traditional approaches have focused mainly on changing the structures of NER models but have always ignored the hard constraints and left the NER tagging schemes unchanged. To solve this problem, this article proposes a dynamic entity-based NER approach under unconstrained tagging schemes. To eliminate the constraints, we reorganize widely used tagging schemes and propose two novel unconstrained schemes: one in which tags are assigned to words and entities separately, and one where words and entities are labeled indiscriminately by uniformly taking them as chunks. Associated with the unconstrained tagging schemes, two entity-based neural architectures are also presented that recognize entities at the same time that the sentence is dynamically segmented. Unlike other static NER models that process all the tags after labeling each word, our models address the inputs dynamically by the interactions between the input text and the output labels. The dynamic mechanism can ensure that the entity-level features are included in the NER system, which is helpful for correctly recognizing entities. Except for word embeddings pretrained from unlabeled corpora, no external language-specific knowledge or other resources such as gazetteers are used. The experiments with English, German, Dutch, and Spanish datasets show that our methods can perform very well with different languages. Particularly, the results of the recall rate against the entity’s length reveal that the proposed entity-based models are suitable for recognizing entities with long lengths.
Feng Zhao 0003, Xiangyu Gui, Yafan Huang, Hai Jin 0001, Laurence T. Yang
IEEE Trans. Big Data3
2021 Rumor Detection on Social Media with Out-In-Degree Graph Convolutional Networks
abstract
With the tremendous development in hardware computing and the widespread use of mobile terminal devices, there are increasingly more people who prefer to share their lives and opinions on social media. Though social media plat-forms allow everyone to express their opinions freely, they create convenience for rumor propagation in the meantime, which brings huge negative influence on the public and makes rumor detection extremely necessary. Currently, the most effective methods regard rumor propagation network as a graph and adopt graph convolutional networks (GCN) to detect rumor automatically. Such methods achieve promising performance in rumor detection, however, we argue that they have two critical defects: 1) they neglect the position contributions of rumor nodes in a graph, reducing the accuracy of rumor detection results; 2) they are inadequate in dealing with imbalanced data, which also indicates the inflexibility and the poor generalization ability of the model. To overcome these issues, we incorporate Katz centrality into spectral-domain graph convolution and propose a novel model named Out-In-Degree Graph Convolutional Networks (OID-GCN). Specifically, besides enhancing accuracy, Katz centrality can efficiently capture the position information of nodes, while the rest structure of OID-GCN shows a superb ability in dealing with imbalanced data. Comprehensive experimental results on two real-world datasets Twitter-15 and Twitter-16 demonstrate our OID-GCN outperforms existing methods.
Shihui Song, Yafan Huang, Hongwei Lu
SMC2
2021 Path-enhanced explainable recommendation with knowledge graphs
Yafan Huang, Feng Zhao 0003, Xiangyu Gui, Hai Jin 0001
World Wide Web1