Yu Pu

dblp:13/2109 · DBLP profile ↗
← Back
28ranked-venue papers
11as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Security and privacy · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 WN-Sleep: Modeling Whole-Night Data for Improved Sleep Staging Classification
abstract
Sleep staging, crucial for diagnosing sleep disorders, requires precise recognition of physiological signals within 30-second epochs, a task fundamentally different from managing long-term semantic dependencies in natural language processing (NLP). Our model aims to refine the integration of local and global features for more accurate sleep stage classification. Following the American Academy of Sleep Medicine (AASM) guidelines, it focuses on rigorous intra-epoch feature extraction to ensure reliable identification of sleep stages. Moreover, our approach incorporates a global perspective by analyzing whole-night data, which is essential for handling transitional periods and ambiguities. Existing sequential modeling techniques often overlook the unique requirements of sleep staging, leading to performance declines when epochs extend beyond approximately 200. Our model addresses this by structurally processing local and global information and carefully balancing detailed intra-epoch analysis with an overarching view of sleep cycles through a gating mechanism. This gate mechanism selectively integrates long-term dependencies, optimizing the balance between local accuracy and global context. This approach represents a significant advancement over existing models, offering more accurate, reliable, and clinically relevant sleep staging. Extensive experiments on the SHHS, SleepEDF-20, and SleepEDF-78 datasets demonstrate that our method outperforms state-of-the-art approaches.
Gaohan Ye, Lingjie Shu, Yu Pu, Beilei Wang, Dong Zhang 0015, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
IEEE J. Biomed. Health Informatics6
2025 Integrating Pause Information with Word Embeddings in Language Models for Alzheimer's Disease Detection from Spontaneous Speech
abstract
Alzheimer’s disease (AD) is a progressive neurodegenerative disorder characterized by cognitive decline and memory loss. Early detection of AD is crucial for effective intervention and treatment. In this paper, we propose a novel approach to AD detection from spontaneous speech, which incorporates pause information into language models. Our method involves encoding pause information into embeddings and integrating them into the typical transformer-based language model, enabling it to capture both semantic and temporal features of speech data. We conduct experiments on the Alzheimer’s Dementia Recognition through Spontaneous Speech (ADReSS) dataset and its extension, the ADReSSo dataset, comparing our method with existing approaches. Our method achieves an accuracy of 83.1% in the ADReSSo test set. The results demonstrate the effectiveness of our approach in discriminating between AD patients and healthy individuals, highlighting the potential of pauses as a valuable indicator for AD detection. By leveraging speech analysis as a non-invasive and cost-effective tool for AD detection, our research contributes to early diagnosis and improved management of this debilitating disease.
Yu Pu
ICASSP1
2025 Development and Validation of a Usability Scale for B2B Cloud Products
Haoran Shao, Yue Mi, Yu Pu
INTERACT (3)4
2025 Empowering Large Language Models for End-to-End Speech Translation Leveraging Synthetic Data
Yu Pu, Weiqiang Zhang 0001, Xie Chen 0001
INTERSPEECH1
2025 PPGs-BERT: Leveraging Phoneme Sequence and BERT for Alzheimer's Disease Detection from Spontaneous Speech
Ziyue Qiu, Yu Pu, Xuchu Chen
INTERSPEECH3
2024 Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
abstract
Whisper and other large-scale automatic speech recognition models have made significant progress in performance. However, their performance on many low-resource languages, such as Kazakh, is not satisfactory. It is worth researching how to utilize low-cost data to improve the performance of Whisper on under-represented languages. In this study, we utilized easily accessible unpaired speech and text data and combined the language model GPT with Whisper on Kazakh. We implemented end of transcript (EOT) judgment modification and hallucination penalty to improve the performance of speech recognition. Further, we employed the decoding average token log probability as a criterion to select samples from unlabeled speech data and used pseudo-labeled data to fine-tune the model to further improve its performance. Ultimately, we achieved more than 10\% absolute WER reduction in multiple experiments, and the whole process has the potential to be generalized to other under-represented languages.
Yu Pu
INTERSPEECH2
2023 Cross-Lingual Alzheimer's Disease Detection Based on Paralinguistic and Pre-Trained Features
abstract
We present our submission to the ICASSP-SPGC-2023 ADReSS-M Challenge Task, which aims to investigate which acoustic features can be generalized and transferred across languages for Alzheimer’s Disease (AD) prediction. The challenge consists of two tasks: one is to classify the speech of AD patients and healthy individuals, and the other is to infer Mini Mental State Examination (MMSE) score based on speech only. The difficulty is mainly embodied in the mismatch of the dataset, in which the training set is in English while the test set is in Greek. We extract paralinguistic features using openSmile toolkit and acoustic features using XLSR-53. In addition, we extract linguistic features after transcribing the speech into text. These features are used as indicators for AD detection in our method. Our method achieves an accuracy of 69.6% on the classification task and a root mean squared error (RMSE) of 4.788 on the regression task. The results show that our proposed method is expected to achieve automatic multilingual Alzheimer’s Disease detection through spontaneous speech.
Xuchu Chen, Yu Pu
ICASSP2
2023 A Workload-Aware DVFS Robust to Concurrent Tasks for Mobile Devices
abstract
Power governing is a critical component of modern mobile devices, reducing heat generation and extending device battery life. A popular technology of power governing is dynamic voltage and frequency scaling (DVFS), which adjusts the operating frequency of a processor to balance its performance and energy consumption. With the emergence of diverse workloads on mobile devices, traditional DVFS methods that do not consider workload characteristics become suboptimal. Recent application-oriented methods propose dedicated and effective DVFS governors for individual application tasks. Since their approach is only tailored to the targeted task, performance drops significantly when other tasks run concurrently, which is however common on today's mobile devices. In this paper, our key insight is that hardware meta-data, widely used in existing DVFS designs, has great potential to enable capable workload awareness and task concurrency adaptability for DVFS, but they are underexplored. We find that workload characteristics can be described in a hyperspace composed of multiple dimensions derived from these metadata to form a novel workload contextual indicator to profile task dynamics and concurrency. On this basis, we propose a meta-state metric to capture this relationship and design a new solution, GearDVFS. We evaluate it for a rich set of application tasks, and it outperforms state-of-the-art methods.
Chengdong Lin, Kun Wang 0051, Zhenjiang Li 0001, Yu Pu
MobiCom4
2021 Hermes: an efficient federated learning framework for heterogeneous mobile clients
abstract
Federated learning (FL) has been a popular method to achieve distributed machine learning among numerous devices without sharing their data to a cloud server. FL aims to learn a shared global model with the participation of massive devices under the orchestration of a central server. However, mobile devices usually have limited communication bandwidth to transfer local updates to the central server. In addition, the data residing across devices is intrinsically statistically heterogeneous (i.e., non-IID data distribution). Learning a single global model may not work well for all devices participating in the FL under data heterogeneity. Such communication cost and data heterogeneity are two critical bottlenecks that hinder from applying FL in practice. Moreover, mobile devices usually have limited computational resources. Improving the inference efficiency of the learned model is critical to deploy deep learning applications on mobile devices. In this paper, we present Hermes - a communication and inference-efficient FL framework under data heterogeneity. To this end, each device finds a small subnetwork by applying the structured pruning; only the updates of these subnetworks will be communicated between the server and the devices. Instead of taking the average over all parameters of all devices as conventional FL frameworks, the server performs the average on only overlapped parameters across each subnetwork. By applying Hermes, each device can learn a personalized and structured sparse deep neural network, which can run efficiently on devices. Experiment results show the remarkable advantages of Hermes over the status quo approaches. Hermes achieves as high as 32.17% increase in inference accuracy, 3.48× reduction on the communication cost, 1.83× speedup in inference efficiency, and 1.8× savings on energy consumption.
Ang Li 0005, Jingwei Sun 0002, Pengcheng Li 0001, Yu Pu, Hai Li 0001, Yiran Chen 0001
MobiCom4
2021 Detection of the interictal epileptic discharges based on wavelet bispectrum interaction and recurrent neural network
Nabil Sabor, Yongfu Li 0002, Zhe Zhang 0008, Yu Pu, Guoxing Wang, Yong Lian 0001
Sci. China Inf. Sci.4
2021 A robust QRS detection and accurate R-peak identification algorithm for wearable ECG sensors
Yongfu Li 0002, Guoxing Wang, Yu Pu, Yong Lian 0001
Sci. China Inf. Sci.4
2021 A differential game view of antagonistic dynamics for cybersecurity
Shengling Wang 0001, Yu Pu, Yinhao Xiao
Comput. Networks2
2020 Xuantie-910: Innovating Cloud and Edge Computing by RISC-V
abstract
This article consists only of a collection of slides from the author's conference presentation.
Chen Chen 0058, Xiaoyan Xiang, Chang Liu 0021, Yunhai Shang, Ren Guo, Dongqi Liu 0003, Ziyi Hao, Chunqiang Li, Yu Pu, Jian-Yi Meng, Xiaolang Yan, Yuan Xie 0001, Xiaoning Qi
Hot Chips Symposium12
2020 Xuantie-910: A Commercial Multi-Core 12-Stage Pipeline Out-of-Order 64-bit High Performance RISC-V Processor with Vector Extension : Industrial Product
abstract
The open source RISC-V ISA has been quickly gaining momentum. This paper presents Xuantie-910, an industry leading 64-bit high performance embedded RISC-V processor from Alibaba T-Head division. It is fully based on the RV64GCV instruction set and it features custom extensions to arithmetic operation, bit manipulation, load and store, TLB and cache operations. It also implements the 0.7.1 stable release of RISCV vector extension specification for high efficiency vector processing. Xuantie-910 supports multi-core multi-cluster SMP with cache coherence. Each cluster contains 1 to 4 core(s) capable of booting the Linux operating system. Each single core utilizes the state-of-the-art 12-stage deep pipeline, out-of-order, multi-issue superscalar architecture, achieving a maximum clock frequency of 2.5 GHz in the typical process, voltage and temperature condition in a TSMC 12nm FinFET process technology. Each single core with the vector execution unit costs an area of 0.8 mm2, (excluding the L2 cache). The toolchain is enhanced significantly to support the vector extension and custom extensions. Through hardware and toolchain co-optimization, to date Xuantie-910 delivers the highest performance (in terms of IPC, speed, and power efficiency) for a number of industrial control flow and data computing benchmarks, when compared with its predecessors in the RISC-V family. Xuantie-910 FPGA implementation has been deployed in the data centers of Alibaba Cloud, for applicationspecific acceleration (e.g., blockchain transaction). The ASIC deployment at low-cost SoC applications, such as IoT endpoints and edge computing, is planned to facilitate Alibaba’s end-to-end and cloud-to-edge computing infrastructure.
Chen Chen 0058, Xiaoyan Xiang, Chang Liu 0021, Yunhai Shang, Ren Guo, Dongqi Liu 0003, Ziyi Hao, Chunqiang Li, Yu Pu, Jian-Yi Meng, Xiaolang Yan, Yuan Xie 0001, Xiaoning Qi
ISCA12
2020 Energy-Efficient Arbitrary Precision Multi-Bit Multiplication with Bi-Serial In/Near Memory Computing
abstract
Recent works show that multi-bit multiplications can be achieved in multi-cycles via serial in/near memory computing, aiming at reducing data transfers hence improving energy efficiency. However, the pure serial approach suffers from a long latency and leaves additional room to optimize energy efficiency. We propose three new techniques to develop a novel SRAM structure to realize multi-bit multiplication with bi-serial in/near memory computing. Firstly, we use 2-bit binary numbers (bi-serial) as the smallest unit of operation working in the mixed-signal mode to achieve arbitrary precision. This significantly reduces latency compared to the pure serial approaches. Secondly, we take the unique advantages of bi-serial (2-bit by 2-bit) multiplication and reduce the number of voltage bands that we need to differentiate from seven to four. This reduces the voltage swing on analog bit-lines (ABLs), in this way, reducing the dynamic power and improving accuracy. Thirdly, we optimize a low-cost voltage comparator based on the inverter chain to reduce static power further. When normalized to 8-bit by 8-bit multiply operation and implemented a 16KB SRAM array in SMIC 55-nm CMOS technology, the energy efficiency of our design is 1.47 TOPS/W, which is 2.7 times better than the state-of-the-art.
Yuqi Wang 0004, Yu Pu, Yajun Ha
ISCAS3
2017 VaultIME: Regaining User Control for Password Managers Through Auto-Correction
Le Guan, Sadegh Farhang, Yu Pu, Pinyao Guo, Jens Grossklags, Peng Liu 0005
SecureComm3
2017 Valuating Friends' Privacy: Does Anonymity of Sharing Personal Data Matter?
Yu Pu, Jens Grossklags
SOUPS1
2016 Sharing Is Caring, or Callous?
Yu Pu, Jens Grossklags
CANS1
2016 Towards a Model on the Factors Influencing Social App Users' Valuation of Interdependent Privacy
abstract
Abstract In the context of third-party social apps, the problem of interdependency of privacy refers to users making app adoption decisions which cause the collection and utilization of personal information of users’ friends. In contrast, users’ friends have typically little or no direct influence over these decision-making processes. We conduct a conjoint analysis study with two treatment conditions which vary the app data collection context (i.e., to which degree the functionality of the app makes it necessary for the app developer to collect friends’ information). Analyzing the data, we are able to quantify the monetary value which app users place on their friends’ and their own personal information in each context. Combining these valuations with the responses to a comprehensive survey, we apply structural equation modeling (SEM) analysis to investigate the roles of privacy concern, its antecedents, as well as app data collection context to work towards a model of interdependent privacy for the scenario of third-party social app adoption. We find that individuals’ past experiences regarding privacy invasions are negatively associated with their trust for third-party social apps’ proper handling of their personal information, which in turn influences their concerns for their own privacy associated with third-party social apps. In addition, positive effects of users’ privacy knowledge on concerns for their own privacy and concerns for friends’ privacy regarding app adoption are partially supported. These privacy concerns are further found to affect how users value their own and their friends’ personal information. However, we are unable to support an association between users’ online social capital and their concerns for friends’ privacy. Nor do we have enough evidence to show that treatment conditions moderate the association between the concern for friends’ personal information and the value of such information in app adoption contexts.
Yu Pu, Jens Grossklags
Proc. Priv. Enhancing Technol.1
2015 Ultra-low-energy adiabatic dynamic logic circuits using nanoelectromechanical switches
abstract
Nanoelectromechanical (NEM) switches have the unique property of virtually zero leakage current, making it an appealing candidate for implementing dynamic logic circuits. However, NEM switches have relatively slow switching times (tens of nanoseconds) due to their mechanical nature. Adiabatic Dynamic Logic (ADL) is an approach that can achieve ultra-low-energy dissipation when operating at electrically slow speeds. We show through simulation that NEM switches can be used to implement ADL circuits that exploit its zero leakage property while operating at mechanically fast switching speeds. At least an order of magnitude reduction in energy consumption compared to conventional circuit design is shown in simulation using analytical compact modeling.
Christopher Lawrence Ayala, Antonios Bazigos, Daniel Grogg, Yu Pu, Christoph Hagleitner
ISCAS4
2014 Logic synthesis of low-power ICs with ultra-wide voltage and frequency scaling
abstract
For low-power digital ICs with ultra-wide voltage and frequency scaling (e.g., from the nominal supply voltage to the sub/near-threshold regime), achieving design closure can be a big challenge, especially when speed limits are pushed at very different voltages. This paper shares a practical logic synthesis recipe that helps to fulfill tight timing constraints. Our method includes: i) synthesizing circuits at a high voltage; ii) over-constraining maximal transition time; iii) pruning standard cell library based on cell delay degradation factor across voltages. This approach shows effectiveness on an industrial 90nm low-power micro-controller.
Yu Pu, Juan Diego Echeverri, Maurice Meijer, José Pineda de Gyvez
DATE1
2013 Modelling NEM relays for digital circuit applications
abstract
A reduced-order model for NEM relays is presented that combines electro-mechanical beam actuation and landing of beam tip on the surface electrode. This model shows a deviation of less than 2%, for the DC as well as the transient response for beam actuation in a circuit simulation, when compared to a finite-element simulation. It also shows an excellent match for the energy. The model allows accurate circuit simulation to aid in NEM-relay based logic design, and facilitates the quantification of key gate-level metrics.
Sunil Rana, Dinesh Pamunuwa, Daniel Grogg, Michel Despont, Yu Pu, Christoph Hagleitner
ISCAS6
2011 From Xetal-II to Xetal-Pro: On the Road Toward an Ultralow-Energy and High-Throughput SIMD Processor
abstract
Looking forward to the next generation of mobile streaming computing, the demanded energy efficiency of end-user terminals will become ever stringent. The Xetal-Pro processor, which is the latest member of the Xetal low-power single-instruction multiple data (SIMD) processor family from Philips, is presented in this paper. The predecessor of Xetal-Pro, known as Xetal-II, already ranks as one of the most computational-efficient [in terms of giga operations per second (GOPS)/Watt] processors available today, however, it cannot yet achieve the demanded energy efficiency (less than 1 pJ per operation). Unlike Xetal-II, Xetal-Pro supports ultrawide supply voltage (Vdd) scaling from the nominal supply to the subthreshold region. Although aggressiveVddscaling causes severe throughput degradation, this can be partly compensated for by the massive parallelism in the Xetal family. Xetal-II includes a large on-chip frame memory (FM), which cannot be scaled well to an ultralowVddhence creating a big obstacle to increase energy efficiency. Therefore, we investigate both different FM realizations and memory organization alternatives. A hybrid memory system (HMS), which reduces the non-local memory traffic and enables furtherVddscaling, is proposed. For design space exploration of the right number of the scratchpad memory (SM) entries, the corresponding data locality analysis is provided, too. Moreover, some unique circuit implementation issues of Xetal-Pro such as the customized level-shifter are also discussed. Compared to Xetal-II operating at the nominal voltage, Xetal-Pro provides up to two times energy efficiency improvement even withoutVddscaling (essentially a consequence of data localization in the SM) when delivering the same amount of ultrahigh throughput. WithVddscaling into the sub/near threshold region, Xetal-Pro could gain more than ten times energy reduction while still delivering a high throughput of 0.69 GOPS (counting multiply and add operations only). The new insight of Xetal-Pro sheds light on the direction of future ultralow-energy SIMD processors.
Yu Pu, Yifan He 0002, Zhenyu Ye, Sebastian M. Londono, Anteneh A. Abbo, Richard P. Kleihorst, Henk Corporaal
IEEE Trans. Circuits Syst. Video Technol.1
2010 Xetal-Pro: an ultra-low energy and high throughput SIMD processor
abstract
This paper presents Xetal-Pro SIMD processor, which is based on Xetal-II, one of the most computational-efficient (in terms of GOPS/Watt) processors available today. Xetal-Pro supports ultra wide V DD scaling from nominal supply to the sub-threshold region. Although aggressive V DD scaling causes severe throughput degradation, this can be compensated by the nature of massive parallelism in the Xetal family. The predecessor of Xetal-Pro, Xetal-II, includes a large on-chip frame memory (FM), which cannot operate reliably at ultra low voltage. Therefore we investigate both different FM realizations and memory organization alternatives. We propose a hybrid memory architecture which reduces the non-local memory traffic and enables further V DD scaling. Compared to Xetal-II operating at nominal voltage, we could gain more than 10× energy reduction while still delivering a sufficiently high throughput of 0.69 GOPS (counting multiply and add operations only). This work gives a new insight to the design of ultra-low energy SIMD processors, which are suitable for portable streaming applications.
Yifan He 0002, Yu Pu, Richard P. Kleihorst, Zhenyu Ye, Anteneh A. Abbo, Sebastian M. Londono, Henk Corporaal
DAC2
2010 Misleading energy and performance claims in sub/near threshold digital systems
abstract
Many of us in the field of ultra-low-Vddprocessors experience difficulty in assessing the sub/near threshold circuit techniques proposed by earlier papers. This paper investigates five major pitfalls which are often not appreciated by researchers when claiming that their circuits outperform others by working at a lower Vddwith a higher energy-efficiency. These pitfalls include: i) overlook the impacts of different technologies and different Vthdefinitions, ii) only emphasize energy reduction but ignore severe throughput degradation, or expect impractical pipelining depth and parallelism degree to compensate this throughput degradation, iii) unrealistically assume that memory's Vddand energy could scale as well as standard cells, iv) use the highest temperature as the worst timing corner as in the super-threshold, but in fact negative temperature becomes much more detrimental in the sub/near threshold regime, v) pursue just-in-need Vddto compensate effects of PVT, but without considering the high energy loss on DC-DC converters. Therefore, the actual energy benefit from using a sub/near threshold Vddcan be greatly overestimated. This work provides some design guidelines and silicon evidence to ultra-low-Vddsystems. The outlined pitfalls also shed light on future directions in this field.
Yu Pu, Xin Zhang 0025, Jim Huang, Atsushi Muramatsu, Masahiro Nomura, Koji Hirairi, Hidehiro Takata, Taro Sakurabayashi, Shinji Miyano, Makoto Takamiya, Takayasu Sakurai
ICCAD1
2008 Statistical noise margin estimation for sub-threshold combinational circuits
abstract
The increasingly popular sub-threshold design is strongly calling for EDA support to estimate noise margins, minimum functional supply voltage, as well as the functional yield. In this paper, we propose a fast, accurate and statistical approach to accomplish these goals. First, we derive close-form functions based on a new equivalent resistance model which enables the fast estimation of noise margins of individual cells at the gate-level. Second, we propose to calculate and propagate the noise margin information with an affine arithmetic model that takes into account process variations and correspondent inter-cell correlations. Experiments with ISCAS benchmarks have shown that the new approach has an accuracy of 98.5% w.r.t. transistor-level Monte Carlo simulations. The running time per input vector of the new approach only needs a few seconds, in contrast to the many hours required by transistor-level DC Monte-Carlo simulations. To the best of our knowledge, we are the first to provide a fast, accurate and statistical methodology other than Monte-Carlo simulation for the noise margin estimation of sub-threshold combinational circuits.
Yu Pu, José Pineda de Gyvez, Henk Corporaal, Yajun Ha
ASP-DAC1
2007 Vt balancing and device sizing towards high yield of sub-threshold static logic gates
abstract
Operating digital circuits in the sub-threshold region is potentially a solution for ultra low-power applications. However, simply reducing supply voltage well below threshold voltage causes functional yield degradation. In this paper, we show that imbalanced VT of pMOS and nMOS transistors and VT mismatch of paired transistors are especially detrimental to sub-threshold functional yield. We propose a variability-driven digital gate design approach which includes balancing process-corner VT shifts of nMOS/pMOS transistors with a low-overhead bulk-bias circuitry and a gate-sizing approach that yields close to minimum size transistor dimensions. Results of Monte-Carlo simulations of a ring oscillator with 31 stages show that our solution can help to achieve a mean frequency speedup of 51.91% and energy/cycle saving of 19.67% on average.
Yu Pu, José Pineda de Gyvez, Henk Corporaal, Yajun Ha
ISLPED1
2006 An automated, efficient and static bit-width optimization methodology towards maximum bit-width-to-error tradeoff with affine arithmetic model
abstract
Ideally, bit-width analysis methods should be able to find the most appropriate bit-widths to achieve the optimum bit-width-to-error tradeoff for variables and constants in high level DSP algorithms when they are implemented into hardware. The tradeoff enables the fixed-point hardware implementation to be area efficient but still within the allowed error tolerance. Unfortunately, almost all the existing static bit-width analysis methods are Interval Arithmetic (IA) based that may overestimate bit-widths and enable fairly pessimistic bit-width-to-error tradeoff. We have developed an automated and efficient bit-width optimization methodology that is Affine Arithmetic (AA) based. Experiments have proven that, compared to the previous static analysis methods, our methodology not only dramatically reduces the fractional bit-width by more than 35% but also slightly reduces the integer bit-width. In addition, our probabilistic error analysis method further enlarges the bit-width-to-error tradeoff.
Yu Pu, Yajun Ha
ASP-DAC1