Yonghua Lin

dblp:95/9301 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
12since 2021 · last 2026
0009-0009-6500-9640ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 7 since 2021Systems, architecture and hardware · 9 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models
abstract
Jian Gao, Richeng Xuan, Zhaolu Kang, Dingshi Liao, Wenxin Huang, Zongmou Huang, Yangdi Xu, Bowen Qin, Zheqi He, Xi Yang, Changjinli, Yonghua Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Richeng Xuan, Zhaolu Kang, Dingshi Liao, Wenxin Huang, Zongmou Huang, Yangdi Xu, Bowen Qin, Zheqi He, Changjin Li, Yonghua Lin
ACL (1)12
2026 PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
abstract
Large Language Models (LLMs) demonstrate exceptional capabilities across various tasks, but their deployment is constrained by high computational and memory costs.Model pruning provides an effective means to alleviate these demands.However, existing methods often ignore the characteristics of prefill-decode (PD) disaggregation in practice.In this paper, we propose a pruning method that is highly integrated with PD disaggregation, enabling more precise pruning of blocks.Our approach constructs pruning and distillation sets to perform iterative block removal, obtaining better pruning solutions.Moreover, we analyze the pruning sensitivity of the prefill and decode stages and identify removable blocks specific to each stage, making it well suited for PD disaggregation deployment.Extensive experiments demonstrate our approach consistently achieves strong performance in both PD disaggregation and PD unified (non-PD disaggregation) settings, and can also be extended to other non-block pruning methods.Under the same settings, our method achieves improved performance and faster inference.
Mengsi Lyu, Zhuo Chen 0040, Yulong Ao, Yonghua Lin
ACL (1)5
2026 D-S evidence theory-driven attribute reduction for partial label heterogeneous data using label confidence and weighted neighborhood rough set model
Zhaowen Li, Yumei Nong, Guangming Xue, Yonghua Lin, Ning Lin
Eng. Appl. Artif. Intell.4
2026 Feature selection for partially labeled hybrid data via self information, prediction label using k -nearest neighbor and Student-t kernel rough set
Sujuan Pan, Yonghua Lin
Eng. Appl. Artif. Intell.2
2026 Local attribute reduction for large scale hybrid data with limited missing labels via self information and overlap degree
Run Guo, Yonghua Lin, Zhaowen Li
Expert Syst. Appl.2
2025 ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
abstract
Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children’s speech remains challenging due to differences in pronunciation, tone, and pace compared to adult speech. In this paper, we introduce a new Mandarin speech dataset focused on children aged 3 to 5, addressing the scarcity of resources in this area. The dataset comprises 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. We provide a comprehensive analysis of speaker demographics, speech duration distribution and geographic coverage. Additionally, we evaluate ASR performance on models trained from scratch, such as Conformer, as well as fine-tuned pre-trained models like HuBERT and Whisper, where fine-tuning demonstrates significant performance improvements. Furthermore, we assess speaker verification (SV) on our dataset, showing that, despite the challenges posed by the unique vocal characteristics of young children, the dataset effectively supports both ASR and SV tasks. This dataset is a valuable contribution to Mandarin child speech research and holds potential for applications in educational technology and child-computer interaction. It will be open-source and freely available for all academic purposes.
Jiaming Zhou 0001, Shiwan Zhao, Jiabei He 0001, Haoqin Sun, Hui Wang 0075, Aobo Kong, Xi Yang 0023, Yequan Wang, Yonghua Lin
ACL (1)12
2025 SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors
abstract
While voice technologies increasingly serve aging populations, current systems exhibit significant performance gaps due to inadequate training data capturing elderly-specific vocal characteristics like presbyphonia and dialectal variations. The limited data available on super-aged individuals in existing elderly speech datasets, coupled with overly simple recording styles and annotation dimensions, exacerbates this issue. To address the critical scarcity of speech data from individuals aged 75 and above, we introduce SeniorTalk, a carefully annotated Chinese spoken dialogue dataset. This dataset contains 55.53 hours of speech from 101 natural conversations involving 202 participants, ensuring a strategic balance across gender, region, and age. Through detailed annotation across multiple dimensions, it can support a wide range of speech tasks. We perform extensive experiments on speaker verification, speaker diarization, speech recognition, and speech editing tasks, offering crucial insights for the development of speech technologies targeting this age group. Code is available at https://github.com/flageval-baai/SeniorTalk and data at https://huggingface.co/datasets/evan0617/seniortalk.
Yang Chen 0034, Hui Wang 0075, Jiabei He 0001, Jiaming Zhou 0001, Xi Yang 0023, Yequan Wang, Yonghua Lin
NeurIPS9
2025 Conditional information entropy-based feature selection for partially labeled heterogeneous data via matrix operation and prediction label using k-nearest neighbor
Yumei Nong, Liwei Tian, Yonghua Lin, Zhaowen Li
J. Supercomput.3
2025 The ensemble of self-information-based feature selection for heterogeneous data via k-nearest neighborhood rough set model
Yong Cai Zhang, Yonghua Lin, Mohamed Rizon, Meng Choung Chiong, Yih Bing Chu
J. Supercomput.2
2024 A novel cooperative swarm intelligence feature selection method for hybrid data based on fuzzy β covering and fuzzy self-information
Zhaowen Li, Xiaopeng Cai, Yonghua Lin
Inf. Sci.4
2024 Attribute reduction for hybrid data based on statistical distribution of data and fuzzy evidence theory
Zhaowen Li, Haixin Huang, Yonghua Lin
Inf. Sci.4
2021 Exploring HW/SW Co-Optimizations for Accelerating Large-scale Texture Identification on Distributed GPUs
abstract
Texture identification has been developed recently to support one-to-one verification and one-to-many search, which provides much broader support than texture classification in real-life applications. It has demonstrated great potentials to enable product traceability by identifying the unique texture information on the surface of the targeted objects. However, existing hardware acceleration schemes are not enough to support a large-scale texture identification, especially for the search task, where the number of texture images being searched can reach millions, creating enormous compute and memory demands and making real-time texture identification infeasible. To address these problems, we propose a comprehensive toolset with jointly optimization strategies from both hardware and software to deliver optimized GPU acceleration and leverage large-scale texture identification with real-time responses. Novel technologies include: 1) a highly-optimized cuBLAS implementation for efficiently running 2-nearest neighbors algorithm; 2) a hybrid cache design to incorporate host memory for streaming data toward GPUs, which delivers a 5 × larger memory capacity while running the targeted workloads; 3) a batch process to fully exploit the data reuse opportunities by considering available compute resources and memory bandwidth constraints. 4) an asymmetric local feature extraction to reduce the memory footprint for keeping feature matrices of reference texture images. To the best of our knowledge, this work is the first implementation to provide real-time large-scale texture identification on GPUs. By exploring the co-optimizations from both hardware and software, we can deliver 31 × faster search and 20 × larger feature cache capacity compared to a conventional CUDA implementation. We also demonstrate our proposed designs by proposing a distributed texture identification system with 14 Nvidia Tesla P100 GPUs which can complete 872,984 texture similarity comparisons in just one second.
Xiaofan Zhang 0001, Yonghua Lin
ICPP4
2020 DNNExplorer: A Framework for Modeling and Exploring a Novel Paradigm of FPGA-based DNN Accelerator
abstract
Existing FPGA-based DNN accelerators typically fall into two design paradigms. Either they adopt a generic reusable architecture to support different DNN networks but leave some performance and efficiency on the table because of the sacrifice of design specificity. Or they apply a layer-wise tailor-made architecture to optimize layer-specific demands for computation and resources but loose the scalability of adaptation to a wide range of DNN networks. To overcome these drawbacks, this paper proposes a novel FPGA-based DNN accelerator design paradigm and its automation tool, called DNNExplorer, to enable fast exploration of various accelerator designs under the proposed paradigm and deliver optimized accelerator architectures for existing and emerging DNN networks. Three key techniques are essential for DNNExplorer's improved performance, better specificity, and scalability, including (1) a unique accelerator design paradigm with both high-dimensional design space support and fine-grained adjustability, (2) a dynamic design space to accommodate different combinations of DNN workloads and targeted FPGAs, and (3) a design space exploration (DSE) engine to generate optimized accelerator architectures following the proposed paradigm by simultaneously considering both FPGAs' computation and memory resources and DNN networks' layer-wise characteristics and overall complexity. Experimental results show that, for the same FPGAs, accelerators generated by DNNExplorer can deliver up to 4.2x higher performances (GOP/s) than the state-of-the-art layer-wise pipelined solutions generated by DNNBuilder [1] for VGG-like DNN with 38 CONV layers. Compared to accelerators with generic reusable computation units, DNNExplorer achieves up to 2.0x and 4.4x DSP efficiency improvement than a recently published accelerator design from academia (HybridDNN [2]) and a commercial DNN accelerator IP (Xilinx DPU [3]), respectively.
Xiaofan Zhang 0001, Hanchen Ye, Yonghua Lin, Jinjun Xiong, Wen-Mei W. Hwu, Deming Chen
ICCAD4
2019 Transferable AutoML by Model Sharing Over Grouped Datasets
abstract
Automated Machine Learning (AutoML) is an active area on the design of deep neural networks for specific tasks and datasets. Given the complexity of discovering new network designs, methods for speeding up the search procedure are becoming important. This paper presents a so-called transferable AutoML approach that Automated Machine Learning (AutoML) is an active area on the design of deep neural networks for specific tasks and datasets. Given the complexity of discovering new network designs, methods for speeding up the search procedure are becoming important. This paper presents a so-called transferable AutoML approach that leverages previously trained models to speed up the search process for new tasks and datasets. Our approach involves a novel meta-feature extraction technique based on the performance of benchmark models, and a dynamic dataset clustering algorithm based on Markov process and statistical hypothesis test. As such multiple models can share a common structure while with different learned parameters. The transferable AutoML can either be applied to search from scratch, search from predesigned models, or transfer from basic cells according to the difficulties of the given datasets. The experimental results on image classification show notable speedup in overall search time for multiple datasets with negligible loss in accuracy.
Chao Xue 0003, Junchi Yan, Stephen M. Chu, Yonggang Hu, Yonghua Lin
CVPR6
2019 Few-Shot Audio Classification with Attentional Graph Neural Networks
Shilei Zhang, Yong Qin 0001, Kewei Sun, Yonghua Lin
INTERSPEECH4
2018 AccDNN: An IP-Based DNN Generator for FPGAs
abstract
Using FPGA to accelerate Deep Neural Networks (DNNs) requires RTL programming, hardware verification, and precise resource allocation, which is both time-consuming and challenging. To address this issue, we present AccDNN, an end-to-end automation tool that can generate high-performance DNN designs on FPGAs automatically. Highlights of this tool include high-quality RTL network layer IPs, a fine-grained layer-based pipeline architecture, and a column-based cache scheme for high throughput, low latency, and reduced on-chip memory utilization. AccDNN also includes an automatic design space exploration tool, called A-REALM, used to generate optimized parallelism schemes by considering external memory access bandwidth, data reuse behaviors, resource availability, and network complexity. We demonstrate AccDNN on four DNNs (Alexnet, ZF, VGG16, and YOLO) on two Xilinx FPGAs (ZC706 and KU115) for edge- and cloud-computing, respectively. AccDNN generates designs that deliver 263 GOPS and 36.4 GOPS/W on ZC706 without any batching and 2109 GOPS and 94.5 GOPS/W on KU115.
Xiaofan Zhang 0001, Yonghua Lin, Jinjun Xiong, Wen-Mei W. Hwu, Deming Chen
FCCM4
2018 Design Flow of Accelerating Hybrid Extremely Low Bit-Width Neural Network in Embedded FPGA
abstract
Neural network accelerators with low latency and low energy consumption are desirable for edge computing. To create such accelerators, we propose a design flow for accelerating the extremely low bit-width neural network (ELB-NN) in embedded FPGAs with hybrid quantization schemes. This flow covers both network training and FPGA-based network deployment, which facilitates the design space exploration and simplifies the tradeoff between network accuracy and computation efficiency. Using this flow helps hardware designers to deliver a network accelerator in edge devices under strict resource and power constraints. We present the proposed flow by supporting hybrid ELB settings within a neural network. Results show that our design can deliver very high performance peaking at 10.3 TOPS and classify up to 325.3 image/s/watt while running large-scale neural networks for less than 5W using embedded FPGA. To the best of our knowledge, it is the most energy efficient solution in comparison to GPU or other FPGA implementations reported so far in the literature.
Qiuwen Lou, Xiaofan Zhang 0001, Yonghua Lin, Deming Chen
FPL5
2018 DNNBuilder: an automated tool for building high-performance DNN hardware accelerators for FPGAs
abstract
Building a high-performance FPGA accelerator for Deep Neural Networks (DNNs) often requires RTL programming, hardware verification, and precise resource allocation, all of which can be time-consuming and challenging to perform even for seasoned FPGA developers. To bridge the gap between fast DNN construction in software (e.g., Caffe, TensorFlow) and slow hardware implementation, we propose DNNBuilder for building high-performance DNN hardware accelerators on FPGAs automatically. Novel techniques are developed to meet the throughput and latency requirements for both cloud- and edge-devices. A number of novel techniques including high-quality RTL neural network components, a fine-grained layer-based pipeline architecture, and a column-based cache scheme are developed to boost throughput, reduce latency, and save FPGA on-chip memory. To address the limited resource challenge, we design an automatic design space exploration tool to generate optimized parallelism guidelines by considering external memory access bandwidth, data reuse behaviors, FPGA resource availability, and DNN complexity. DNNBuilder is demonstrated on four DNNs (Alexnet, ZF, VGG16, and YOLO) on two FPGAs (XC7Z045 and KU115) corresponding to the edge- and cloud-computing, respectively. The fine-grained layer-based pipeline architecture and the column-based cache scheme contribute to 7.7x and 43x reduction of the latency and BRAM utilization compared to conventional designs. We achieve the best performance (up to 5.15x faster) and efficiency (up to 5.88x more efficient) compared to published FPGA-based classification-oriented DNN accelerators for both edge and cloud computing cases. We reach 4218 GOPS for running object detection DNN which is the highest throughput reported to the best of our knowledge. DNNBuilder can provide millisecond-scale real-time performance for processing HD video input and deliver higher efficiency (up to 4.35x) than the GPU-based solutions.
Xiaofan Zhang 0001, Yonghua Lin, Jinjun Xiong, Wen-Mei W. Hwu, Deming Chen
ICCAD4
2016 Auto-tuning Spark Big Data Workloads on POWER8: Prediction-Based Dynamic SMT Threading
abstract
Much research work devotes to tuning big data analytics in modern data centers, since %the truth that even a small percentage of performance improvement immediately translates to huge cost savings because of the large scale. Simultaneous multithreading (SMT) receives great interest from data center communities, as it has the potential to boost performance of big data analytics by increasing the processor resources utilization. For example, the emerging processor architectures like POWER8 support up to 8-way multithreading. However, as different big data workloads have disparate architectural characteristics, how to identify the most efficient SMT configuration to achieve the best performance is challenging in terms of both complex application behaviors and processor architectures. In this paper, we specifically focus on auto-tuning SMT configuration for Spark-based big data workloads on POWE-R8. However, our methodology could be generalized and extended to other programming software stacks and other architectures.
Zhen Jia 0001, Guancheng Chen, Jianfeng Zhan, Lixin Zhang 0002, Yonghua Lin, H. Peter Hofstee
PACT6
2016 FPGA as service in public Cloud: Why and how
abstract
Summary form only given. IBM is the leader of Accelerator cloud technology innovation in industry. IBM Supervessel Cloud is the first cloud providing FPGA accelerator service and FPGA DevOps service to developers. In this talk, Yonghua Lin, Supervessel Cloud leader will share her view on why FPGA service in cloud is important, and how FPGA service could accelerate the Cognitive Computing in Cloud. Meanwhile, Yonghua will introduce the key technologies supporting FPGA service in cloud. Supervessel Cloud has launched FPGA as service for more than 18 months. The FPGA service has been used by developers from different countries. In this talk, Yonghua will also share the gap learned from all these users, and her vision for future.
Yonghua Lin
FPT1
2010 Uplink carrier frequency offset estimation for WiMAX OFDMA-based ranging
abstract
Abstract Ranging is one of the most important processes in the mobile worldwide interoperability for microwave access (WiMAX) standard, for resolving the uplink synchronization and near/far problems. In this paper, we focus on the multi‐user carrier frequency offset (CFO) estimation in both initial ranging and periodic ranging. After the analysis of some existing ranging methods, we propose two algorithms based on the correlation properties of pseudo noise (PN) sequences in time domain and frequency domain respectively. The root mean square error (RMSE) performance is evaluated in both additive white Gaussian noise (AWGN) channel and multi‐path fading channel. Simulation results show that the proposed frequency‐domain cross‐correlation method performs better than the proposed time‐domain cross‐correlation method, and is more robust to multi‐user interference and residual timing offset. Copyright © 2009 John Wiley & Sons, Ltd.
Yonghua Lin, Da Fan, Qing Wang 0045
Wirel. Commun. Mob. Comput.1
2009 An Efficient Software Radio Framework for WiMAX Physical Layer on Cell Multicore Platform
abstract
Wireless baseband processing, which is characterized by high computation complexity and high data throughput, is regarded as the most challenging issue for software radio (SR) systems. To relieve this implementation difficulty in SR systems, the multicore architecture is proposed. However, due to the lack of the universal parallel programming framework for SR systems, it is difficult to take full advantage of the multicore architecture. To fill this gap, in this paper, an efficient parallel SR framework is proposed under the Cell multicore architecture for WiMAX base station (BS). With the proposed framework, one Cell blade server (including two Cell processors) can support up to three sectors, and each sector can support 20Mbps data rate for both uplink and downlink. The system performance results verify the effectiveness of the proposed SR multicore framework.
Qing Wang 0045, Zhenbo Zhu, Yonghua Lin
ICC4
2008 Design of BS transceiver for IEEE 802.16E OFDMA mode
abstract
Wordwide interoperability for microwave access (WiMAX), or the IEEE 802.16d/e standard, is a technology for broadband wireless access (BWA) with significant market potential. In this paper, we propose a base station (BS) physical layer (PHY) transceiver solution for WiMAX orthogonal- frequency division-multiplexing access (OFDMA) mode. Our aim is to provide a single chip solution for baseband processing of WiMAX PHY. The data throughput should achieve 20 Mbps both for uplink and downlink. The solution is implemented on cell broadband engine, which is a multicore processor jointly developed by IBM, SONY and Toshiba for high performance computing(HPC). Algorithms for symbol timing offset, carrier frequency offset, channel estimation and space-frequency block code (SFBC) are embedded in this transceiver. Simulation and real test on Cell show that the proposed solution can fulfill the bit-error-rate (BER) requirements under most situations.
Qing Wang 0045, Da Fan, Yonghua Lin, Zhenbo Zhu
ICASSP3
2002 Adaptive diversity combination of space-time block coding in OFDM-CDMA systems with pilot-based channel estimation
abstract
We investigate the transmit diversity for orthogonal frequency division multiplexing code division multiple access (OFDM-CDMA). Diversity transmission at the base station is an effective technique to combat channel fading. Space-time (ST) coding has been proposed as coding technique to improve the system performance with additional gains. By combing the spatial transmit diversity with the space-time codes for OFDM systems it shows high performance improvement. But it is important to remark that ST decoding requires multichannel state information at the receiver. Thus the achievable diversity gain comes at the price of proportional increase in the amount of training or computation complexity, which incurs efficiency loss especially in a rapidly varying environment. In an OFDM system, there exist pilots used for frame detection, carrier frequency offset estimation and channel estimation. We proposed a receive scheme, which use the pilot structure to estimate the multichannel state information and use the CMA-based channel tracking in the space-time coded transmit diversity OFDM-CDMA systems. Numerical simulations illustrate the effectiveness of this algorithm.
Aigang Feng, Qin-Ye Yin 0001, Yonghua Lin
VTC Spring4