EDBT 2026 Demo / reviewers in the wild / expert
Xiaoya Fan
dblp:73/4858
· DBLP profile ↗
52ranked-venue papers
8as first author
46since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 24 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Symmetrical RHP-Zero-Free Hybrid Boost Converter With Eliminated Flying Capacitor Inrush Current and Enhanced Output Current Continuity
Zhitong Chen, Xingyuan Niu, Yudi Ren, Xiaoya Fan, Yanzhao Ma |
ISCAS | 5 |
| 2026 | A DLSK Global Phase-Shift Controlled Wireless Power Transfer System Achieving Above 90% RX Efficiency
Fantao Wang, Junhong Yan, Fang An, Xiaoya Fan, Yanzhao Ma |
ISCAS | 5 |
| 2026 | BinEncoder: Binary code similarity detection across compilation configurations via microcode
Xiaoya Fan |
Appl. Intell. | 6 |
| 2026 | Mamba6mA: a Mamba-based DNA N6-methyladenine site prediction modelabstractMOTIVATION: N6-methyladenine (6 mA) is an important epigenetic modification of DNA that regulates biological processes such as gene expression, transcription, replication, DNA repair, and cell cycle without altering the DNA sequence. It also plays a key role in many diseases including cancer and autoimmune diseases. Although experimental approaches such as SMRT sequencing and methylated DNA immunoprecipitation can identify 6 mA sites, they suffer from drawbacks including suboptimal sequencing quality, low signal-to-noise ratios, high costs, and time-consuming procedures. In recent years, deep learning approaches have demonstrated significant advantages in predicting 6 mA sites; however, their generalization ability still requires further improvement. RESULTS: Inspired by the state space model Mamba, we propose a novel model for 6 mA site prediction, named Mamba6mA. In the Mamba6mA model, we design position-specific linear layers to replace traditional convolutional layers to facilitate capture specific positional information. Meanwhile, we construct a multi-scale feature extraction module and integrate features captured by sliding windows of different scales, feeding them into the classifier for prediction. Experimental results show that Mamba6mA achieves the best MCC on 9 out of 11 species datasets, surpassing existing state-of-the-art models. Ablation studies confirm that the position-specific linear layers and the multi-scale fusion module contribute MCC performance gains of 2.36% and 2.31%, respectively. Feature visualization analysis further reveals that the model effectively captures sequence patterns upstream and downstream of 6 mA sites providing a new technical approach for studying epigenetic modification mechanisms. AVAILABILITY AND IMPLEMENTATION: The source code for Mamba6mA is available at: https://github.com/XploreAI-Lab/Mamba6mA. Qi Zhao 0008, Haoxuan Shi, Xiaoya Fan |
Bioinform. | 8 |
| 2026 | A near CXL memory processing architecture for distributed graph neural network inference and trainingabstractDistributed Graph Neural Networks (GNNs) require efficient handling of both fine-grained memory accesses and cross memory-device communication, particularly when scaling to large graphs. However, existing acceleration solutions fail to adequately address low bandwidth utilization and scalability across different models and graph sizes. In this paper, we present OptGNN, a scalable heterogeneous distributed architecture tailored for GNNs. OptGNN addresses the challenges of fine-grained memory access by incorporating a near-memory processing mechanism, which improves internal bandwidth utilization. To optimize external communication, we introduce data packing and scheduling strategies that enhance cross memory-device data transfer efficiency. OptGNN achieves 5.7x performance improvement over baseline distributed GNN acceleration methods and 1.29x performance improvement over SOTA distributed GNN acceleration architecture CLAY. Additionally, the system is designed to support various GNN models and large-scale graphs while ensuring load balancing and high hardware utilization. Shengbing Zhang, Xiaoya Fan |
Connect. Sci. | 3 |
| 2026 | CQR-UC: A color QR code-based underwater wireless communication method with GAN-based image enhancement
Yufan Feng, Shangxin Li, Qi Zhao 0008, Xiaoya Fan |
J. Vis. Commun. Image Represent. | 7 |
| 2025 | PRDSE: A Prior-Driven Design Space Exploration MethodabstractDesign space exploration is essential for optimizing deep neural network accelerators, which face increasing computational and energy demands as model complexity grows. Previous approaches rely heavily on local insights, often neglecting the need for extensive exploration to improve the global perspective. This leads to challenges such as blind exploration and a higher likelihood of getting trapped in local optima. Without dynamic adjustments or adaptive strategies, these methods struggle to navigate large, complex design spaces effectively. In this paper, we propose PRDSE, a design space exploration framework based on reinforcement learning that integrates both intrinsic and extrinsic metrics, guided by prior knowledge. The proposed method incorporates an adaptive adjustment mechanism that dynamically balances intrinsic and extrinsic rewards based on the progress of exploration, improving both search efficiency and optimization performance. Compared to state-of-the-art methods, PRDSE achieves substantial improvements, with latency speedups of up to$3.49 \mathrm{x}$in the cloud environment and$3.43 \mathrm{x}$in the edge environment, respectively. This work demonstrates that PRDSE effectively balances exploration and optimization objectives, providing a more efficient and scalable approach to design space exploration in accelerator design. Junda Zhu 0007, Xiaoya Fan, Jianfeng An, Kaijie Feng |
ASAP | 2 |
| 2025 | Multi-Stain Attention Multiple Instance Learning for Prognosis Prediction in Esophageal Squamous Cell CarcinomaabstractEsophageal squamous cell carcinoma (ESCC) is a highly aggressive cancer with poor prognosis. Accurate prognostic models are essential for guiding personalized treatment strategies. However, most existing models rely solely on hematoxylineosin (HE) stained images, without fully leveraging the complementary information provided by multiple immunohistochemical (IHC) stains. IHC markers are crucial for capturing immunerelated features within the tumor microenvironment, which are closely associated with patient outcomes. To address this limitation, we propose Multi-Stain Attention Multiple Instance Learning (MSAMIL), a novel framework that integrates information from multiple staining modalities-including key IHC markers such as CD4 and PD-L1 alongside HE-to better capture tumor-immune interactions and tissue heterogeneity. We evaluate MSAMIL on a private dataset, which includes whole slide images (WSIs) from nine staining modalities across 208 ESCC patients. The model predicts$\mathbf{4}$-year disease-free survival (DFS) with superior accuracy and F1 score compared to existing methods. We further observe that the choice of feature extractor significantly affects performance: domain-specific backbones such as ProvGigapath consistently outperform generic encoders, underscoring the importance of tailored pretraining for histopathological representation learning. Ablation confirms the fusion module is critical for modeling cross-modality interactions. Modality contribution analysis shows that CD4 and PD-L1 contribute most significantly, underscoring their importance in capturing tumor immune microenvironment (TIME) features relevant to prognosis. MSAMIL thus provides a powerful tool for ESCC prognostic prediction and may support personalized treatment decisions. The code is available at http://github.com/wzhy2000/MSAMIL. Xiaoya Fan, Zhengde Jia, Ruishan Geng |
BIBM | 1 |
| 2025 | EEG-Based Seizure Detection and Type Classification with Structured State Space Modeling and Graph Neural NetworksabstractAccurate seizure detection and type classification from Electroencephalography (EEG) are crucial for epilepsy diagnosis and clinical treatment. The intricate nature of seizure dynamics presents a significant challenge in effectively extracting distinguishing features from multivariate EEG signals, mainly due to long-range temporal dynamics and complex spatial dependencies between electrodes. To tackle these challenges, we introduce TS-S4GNN, a two-stream graph neural network (GNN) with structured state space modeling. Specifically, we combine channel-independent 1D-convolutional neural network with structured state space models to capture local and longrange temporal dependencies. The complex spatial dependencies between electrodes are learned with two GNN layers in parallel, one with static graph structure constructed according to the distance between electrodes, one with dynamically evolving graph structures learned from data. We validated TS-S4GNN on TUSZv1.5.2, the largest public EEG corpus with seizure type annotations. Experiments demonstrate that TS-S4GNN achieves an AUROC score of 0.913 in seizure detection and a weighted F1-score of 0.781 in seizure type classification, surpassing the current state-of-the-art models. Ablation studies validate the effectiveness of multi-scale temporal modeling and the twostream GNN approach. This study provides new insights into spatiotemporal modeling of multivariate signals and can be easily extended to other relevant learning tasks. The code is available at https://github.com/XploreAI-Lab/TS-S4GNN. Xiaoya Fan, Pengyuan Ganzhang, Gensheng Pei, Shi-chun Bao |
BIBM | 1 |
| 2025 | Evo2-Virus: Ultra-Short Viral Sequence Identification with EVO2abstractThe accurate identification of viral sequences in sequencing data is critical for pathogen surveillance, especially in metagenomic environments where viral reads are ultra-short, noisy, and deeply embedded in host material. Existing methods struggle with generalizability and robustness, particularly on short fragments that are common in real-world sequencing data. In this work, we present Ev02- Virus, a lightweight and scalable framework that combines autoregressive pretraining with bidirectional Transformer fine-tuning to classify viral fragments as short as 105 bp. We curate a large-scale pretraining corpus from the Virus-Host DB and a balanced classification dataset constructed from human RNA-seq samples. Evo2-Virus is trained end-to-end and requires no hand-crafted features or sequence alignment. Extensive experiments show that Evo2-Virus significantly outperforms state-of-the-art models including Seeker, DeepVirFinder, RNA-FM, and RNN-VirSeeker, achieving an Fl-score of 0.889 and AUC of 0.957. The model demonstrates remarkable robustness under synthetic sequence noise and exhibits clear separability between viral and non-viral embed dings in low-dimensional space. Our results highlight Ev02- Virus as a practical and effective solution for ultra-short viral sequence detection, with direct implications for real-world pathogen screening, biosurveillance, and genomic diagnostics. Zhixiang Xu, Hankai Yang, Xiaoya Fan |
BIBM | 3 |
| 2025 | MambaHM: High-Resolution Histone Modification Prediction from ATAC-seq and DNA SequenceabstractHistone modifications play a crucial role in transcriptional regulation and are essential targets for genome annotation and gene expression modeling. However, experimentally profiling histone modifications across species, tissues, and cellular states is costly and often impractical. While numerous deep learning models have been proposed to predict histone modifications, most operate at relatively low resolution (128 bp), limiting their utility in fine-scale genomic analysis. In this study, we present MambaHM (Mamba for predicting Histone Modifications), a novel deep learning framework built upon Mamba architecture. MambaHM integrates DNA sequence features with chromatin accessibility data (ATAC-seq) to predict ten histone modifications. Leveraging the linear computational complexity of Mamba's state-space modeling, our model avoids excessive compression of feature length during embedding. This enables the model to retain sufficient contextual information while simultaneously achieving 16-bp resolution, thereby enhancing the granularity and accuracy of prediction. Experiments demonstrate that MambaHM achieves a mean Pearson correlation of 0.836 (± 0.063) on the K562 cell line test set. Furthermore, the model generalizes well across cell types, tissues, and species, with a mean cross-context correlation of$0.557 (\pm 0.105)$, approaching the reliability of experimental assays. Compared to state-of-the-art models, MambaHM achieves a$\mathbf{7. 0 3 3 \%}(\mathbf{\pm 0. 0 4 1})$improvement in cross-cell-type performance and over a 58.978 % (877 bp) improvement in prediction deviation evaluation for peak calling. Overall, MambaHM provides a powerful, cost-effective, and fineresolution tool for histone modification prediction, offering precise epigenomic references for downstream analysis and potential applications in drug discovery. The source code is available at https://github.com/zhichunlizzx/MambaHM. Lijuan Jia, Zhixiang Xu, Zengyou He, Zhong Wang 0001, Xiaoya Fan |
BIBM | 6 |
| 2025 | Joint Trajectory and Task Offloading Optimization in UAV-assisted Edge Computing Networks via Deep Reinforcement LearningabstractThe rapid advancement of 5G and 6G technologies has introduced various innovative applications, such as autonomous driving and augmented reality, which significantly increase network traffic and computational demands, particularly in densely populated areas. This study focuses on the utilization of unmanned aerial vehicles (UAVs) for Mobile Edge Computing (MEC) to address these challenges. Specifically, we explore the joint optimization of base station (BS) selection, computing resource allocation and UAV trajectory to minimize system delays and energy consumption in scenarios where computational requirements are related to user movement. A deep reinforcement learning-based approach, UM-DDPG, is proposed to optimize task offloading strategies and UAV trajectories. With our simulation platform, the results show that our method has been effective in reducing the system cost, including both delay and energy consumption, compared to traditional methods. Zhekun Zhang, Xin Chen 0018, Libo Jiao, Mingyang Xu, Xiaoya Fan |
CSCWD | 5 |
| 2025 | Load-Aware Offloading and Resource Allocation in SAGIN via LSTM-Enhanced Deep Reinforcement Learning
Mingyang Xu, Libo Jiao, Zhekun Zhang, Xiaoya Fan |
ICA3PP (7) | 5 |
| 2025 | Visual Feature Learning from Randomized EEG Trials for Object RecognitionabstractObject recognition from electroencephalography (EEG) responses to visual stimuli has received growing attention but has remained challenging, with prior methods achieving only marginally above-chance accuracy for randomized EEG trials. This paper introduces GVFL-EEG (Guided Visual Feature Learning from EEG), a novel framework that leverages well-trained computer vision models to enhance EEG object decoding. Specifically, a pre-trained image encoder is used to guide EEG feature learning via contrastive learning, aligning EEG embeddings with image representations. We present EEGMambaformer, a custom EEG encoder incorporating a residual Mamba block to capture temporal dynamics, an inverted transformer encoder to extract spatial dependencies, a temporal-spatial convolution block for feature fusion, and a projection layer for dimension alignment with the image encoder. Evaluation on the EEG40000 dataset show that our framework achieves 60.94% accuracy in 40-way classification, outperforming existing state-of-the-art methods. This work advances neural decoding and provides insights into brain-computer interface development. The code is available at https://github.com/xuehaixiao/GVFL_EEG/tree/main/GVFL. Xiaoya Fan, Haixiao Xue, Yufan Feng, Qi Zhao 0008, Zhong Wang 0001 |
ICME | 1 |
| 2025 | MSFF-SNet: An End-to-end Object Sorting Model with Multi-head Self-attention and Multi-scale Feature FusionabstractRobotic arm sorting aims to determine the optimal grasping poses and categories of objects to be sorted based on perception information and then place objects into corresponding positions. Most methods combine an object detection algorithm and a grasp detection algorithm sequentially. Such a two-stage strategy results in unsatisfactory grasp accuracy and speed. This paper proposes MSFF-SNet, which directly work with grasp poses as learning targets, while simultaneously recognizing the objects to be sorted. Specifically, multi-scale features from different layers of a ResNet50 are first extracted and intra-scale features are fused using a self-attention mechanism, emphasizing the global contextual information. The features of multiple scales are then fused through a cross-scale feature fusion module. Finally, a multi-task detection head is used to directly make pixel-level predictions of grasping poses and object categories. MSFF-SNet achieved a grasp accuracy of 99.62% on the Cornell Grasping Dataset and a mean Average Precision with Grasp of 79.52% on the Visual Manipulation Relationship Dataset, achieving the state-of-the-art performance. Besides, it achieved the highest inference speed on the Visual Manipulation Relationship Dataset. The code is available at https://github.com/wzhy2000/MSFF-SNet. Xiaoya Fan, Jiaxiao Wang, Zhonghua Huang |
IJCNN | 1 |
| 2025 | XGBoost6mA: A Framework for 6mA Site Prediction Based on Deep Learning and XGBoostabstractN6-methyladenine (6mA) is a crucial epigenetic regulator involved in various biological processes. Existing high-throughput methods like 6mA-DIP-seq, ChIP-exo/6mACE-seq, SMRT-seq, and nanopore sequencing offer high accuracy but are typically time-consuming, labor-intensive, and expensive. To address these limitations and promote 6mA research, we introduce XGBoost6mA, a predictive framework integrating deep learning and XGBoost. XGBoost6mA utilizes XGBoost to classify deep learning-derived sequence features, effectively capturing nonlinear patterns and improving classification performance. We evaluated the framework using BERT6mA and CNN6mA feature extractors on benchmark datasets from 11 species. Results showed that XGBoost6mA consistently outperformed both BERT6mA and CNN6mA across all tested species. To validate its robustness, we tested XGBoost6mA by introducing random nucleotide mutations in test sequences. The model maintained stable accuracy, indicating strong resilience to small sequence variations. Additionally, XGBoost6mA offers interpretability by pinpointing key k-mers and positional factors, enhancing the biological significance of its predictions. These findings underscore the advantages of XGBoost6mA in 6mA prediction and its promise as a widely applicable bioinformatics tool. To support future research, XGBoost6mA’s source code is publicly available at https://github.com/xzx0554/xgboost6ma. Zhixiang Xu, Yufei Bao, Xiaoya Fan, Xiangtao Liu |
IJCNN | 5 |
| 2025 | MambaCpG: an accurate model for single-cell DNA methylation status imputation using mambaabstractDNA methylation is a key epigenetic modification involved in biological processes and disease development. The accurate analysis of DNA methylation site information is of significant biological importance. Despite advances in single-cell sequencing, data sparsity due to low cytosine-phosphate-guanine (CpG) coverage remains a challenge. To address this, we introduce MambaCpG, a single-cell DNA methylation state imputation model based on the Mamba block. MambaCpG integrates the methylation matrix and DNA sequence context and uses bidirectional Mamba blocks to obtain embeddings, effectively capturing the long-range dependencies between CpG sites within DNA methylation patterns. Experiments on seven datasets of various scales demonstrate that MambaCpG outperforms existing models on large, highly sparse datasets while showing competitive performance on smaller datasets. MambaCpG has lower parameter and memory requirements, making it suitable for practical applications. MambaCpG also reveals ultra-long-range dependencies and provides new insights into DNA methylation patterning. Qi Zhao 0008, Bingle Li, Xiaoya Fan |
Briefings Bioinform. | 8 |
| 2025 | CSDSE: An efficient design space exploration framework for deep neural network accelerator based on cooperative search
Kaijie Feng, Xiaoya Fan, Jianfeng An, Haoyang Wang 0014, Chuxi Li |
Neurocomputing | 2 |
| 2025 | An Inductor-First Tri-Path Hybrid Buck Converter With Reduced Inductor Current Suitable for USB Power Delivery AdapterabstractThis paper presents an inductor-first tri-path (IFTP) buck converter suitable for USB power delivery adapter to charge 1-2 cell battery. The proposed topology adopts the inductor-first strategy and the tri-path strategy of one inductor path and two capacitor paths to extend the output voltage conversion range, realize the continuous input current, eliminate the input EMI noise, and reduce the inductor current. In addition, a phase-interleaved symmetric inductor-first tri-path (PIS-IFTP) buck converter is proposed to alleviate the inrush current in the flying capacitor under extreme duty cycle of IFTP converter, while further reducing inductor current ripple. Two experimental prototypes for 9 V input to 3-8.4 V ouput have been developed, demonstrating excellent ability of IFTP and PIS-IFTP topologies to reduce inductor current and achieve continuous input current over the whole duty cycle and load range. The experimental results validate that the prototypes provide a wide voltage conversion range of 1/3-1 and a maximum output current of 1.8 A. The peak efficiency of IFTP is 93% at$V_{OUT} \,\, =6.6$V, while the peak efficiency of PIS-IFTP is 94.5% at$V_{OUT} \,\, =3.3$V. Zhitong Chen, Xiaoya Fan, Yanzhao Ma |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | Partition, Predict and Assemble: Targeting Long RNA Secondary Structure PredictionabstractAccurately predicting the secondary structure of Non-coding RNAs is crucial for understanding their biological roles. However, current methods face challenges in accuracy and computational efficiency, particularly for long RNA sequences. Here, we present a novel three-step framework for RNA secondary structure prediction, PPA (Partition, Predict, and Assemble). It first partitions the RNA sequence into independent fragments based on exterior loops, then independently predicts the secondary structure of each fragment, and finally assembles them to construct the complete RNA secondary structure. Following the PPA framework, we introduce ELPBert, a model for RNA exterior loop prediction. We partition the RNA sequence at the central nucleotide of the predicted exterior-loop bases. The performance of six state-of-the-art RNA secondary structure prediction methods with and without RNA partition using our ELPBert were compared. Results demonstrate a great improvement of these methods with RNA partition, especially for long sequences. We believe our PPA framework could serve as an universal framework for RNA secondary structure prediction, particularly for long RNA sequences. Data and code are provided at https://github.com/Kali-ym/PPAwithELPBert. Xiaoya Fan, Yuming Cui, Zengyou He, Qi Zhao 0008, Zhong Wang 0001 |
BIBM | 1 |
| 2024 | EEG-based Seizure Type Classification with Temporal-Spatial-Spectral AttentionabstractSeizure detection and type classification from electroencephalogram (EEG) has the potential to improve the diagnosis and treatment of epilepsy. Although the task of seizure detection has been well-investigated, seizure type classification remains largely unexplored. The intricate nature of seizure dynamics presents a significant challenge in effectively extracting distinguishing features from noisy and high-dimensional EEG signals. Previous studies mainly focus on extracting features from temporal, spectral or special domain. Extracting temporal-spatial-spectral features simultaneously from EEG remains challenging. In this paper, we introduce an attention-based neural network to effectively extract temporal-spatial-spectral EEG features, for seizure type classification. It utilizes an attention module with one-shot aggregation to extract multi-level temporal-spatial-spectral EEG features, aiming to differentiate the complex patterns of various seizure types. Specifically, we construct the 3-dimensional representation of EEG by stacking the time-frequency matrices obtained from short time Fourier transform. The attention module consists of paralleled temporal and spatial-spectral attention blocks, allowing the model to focus on the most distinctive time stamps, sensor locations, and frequency bands. The proposed approach is validated on the largest public seizure EEG database, TUSZ v1.5.2. Five-fold cross validation demonstrates that our framework achieved 0.951 weighted F1 score on seizure type classification, achieving state-of-the-art performance. Ablation study confirmed the effectiveness of the temporal and spatial-spectral attention blocks. Xiaoya Fan, Pengzhi Xu, Wenkui Sun, Qi Zhao 0008, Chenru Hao, Zengyou He, Zhong Wang 0001 |
BIBM | 1 |
| 2024 | KAS-former: a transformer-based model for predicting histone modifications using KAS-seqabstractHistone modifications (HMs) play a critical role in various biological processes, but annotating histone modifications across different cell types using experimental methods alone is extremely challenging. Although many deep learning methods have been developed to predict histone modifications, most rely solely on DNA sequences and do not incorporate novel cell-specific features. In this study, we propose KAS-former, a transformer-based model that integrates DNA sequences with cell-specific features derived from KAS-seq data, enabling effective prediction of histone modifications. Leveraging this transformer architecture coupled with dilated convolution, KAS-former achieves a broad receptive field, effectively capturing cell type-specific specificity from KAS-seq data. Our results demonstrate that KAS-former achieves high accuracy in predicting histone modifications across multiple cell types and shows strong potential for transcription factor prediction. By capturing cell-specific features, this approach not only improves the accuracy of histone modification predictions but also offers valuable insights into the interplay between histone modifications and transcription regulation. The code for KAS-former is available on GitHub at https://github.com/wzhy2000/KAS-former. Lijuan Jia, Wen Wen 0004, Xiaoya Fan, Zengyou He, Ruitu Lyu, Zhong Wang 0001 |
BIBM | 5 |
| 2024 | Electroencephalogram Helps Few-Shot LearningabstractLearning to categorize images with limited samples is a challenge for machines. However, humans can easily generalize from just a few examples. In this study, we propose that the remarkable ability of the human brain to generalize can be reflected in electroencephalogram (EEG) signals. These EEG-related features have the potential to enhance few-shot image classification. Our novel two-stage approach involves the following: first, we learn transferable knowledge from large labeled auxiliary sets by multimodal learning of images and EEG signals using contrastive learning. Then, we finetune the image encoder with novel classes that have only a few samples. We integrate this approach with the Triplet and ProxyNCA framework. Experimental results demonstrate an average improvement of 6.1% and 8.5% in terms of top-1 recall compared to the original Triplet and ProxyNCA methods, respectively. This work showcases the feasibility of leveraging brain signals to enhance few-shot learning. Xiaoya Fan, Zhong Wang 0001 |
ICASSP | 1 |
| 2024 | A 6.78MHz Wireless Power Transfer System With Efficient Global Hysteresis Control for Implantable Medical DevicesabstractThis paper presents an efficient global hysteresis control for output voltage regulation in order to improve the overall system efficiency. Compared with the conventional regulating-rectifying schemes in receiver, the proposed method achieves high light-load end-to-end (E2E) efficiency by shutting down the transmitter (TX) during freewheeling operation period of the rectifier. In addition, a carrier-less pulse-resonant-current modulation technique is proposed to efficiently wake up TX in time. Besides, the residual power of the TX coil in shutdown phase is recycled to further enhance the efficiency. This proposed wireless power transfer system is designed in a 0.18µm BCD process, and regulates an output voltage of 3V. Simulation results show that the maximum output power is 450mW with a high light-load E2E efficiency of 55.9%. The efficiency improvement reached 27.6% compared with the RX regulation only. For a load switching between 15mA and 150mA, the transient response is less than 100ns. Kai Cui 0006, Fantao Wang, Ba Peng, Xiaoya Fan, Yanzhao Ma |
ISCAS | 4 |
| 2024 | A Fully Integrated LDO Using Synchronous VTC and Asynchronous Step Detection Recovery for Under-1 V Supply Voltage ApplicationabstractIn this paper, a fully integrated low-dropout regulator (LDO) using voltage-to-time conversion (VTC) technique is presented for under-1 V supply voltage application. A synchronous VTC technique is proposed using constant-current (CC) charging and discharging to achieve high loop gain. A high-gain charge pump (CP) is proposed to improve power-supply-rejection (PSR). Furthermore, an asynchronous step detection recovery technique is proposed to achieve fast transient response. A frequency-adaptive oscillator is proposed to remove the noise of the clock signal. The proposed LDO is designed in 28-nm process to achieve a droop voltage of 104 mV at load current transient of 90 mA. The proposed LDO achieves PSR of -77 dB at ILOAD=100 mA and PSR of -65 dB at ILOAD=10 mA for 1-kHz supply ripple frequency. The quiescent current is 32 µA and the peak current efficiency is 99.98%. Wan Wang, Na Kang, Xiaoya Fan, Yanzhao Ma |
ISCAS | 5 |
| 2024 | Joint Location Deployment, Offloading and Resource Allocation in Multi-UAV Collaborative Edge Computing NetworksabstractUnmanned aerial vehicles (UAVs) play a crucial role in mobile edge computing (MEC). On one hand, UAVs can serve as relay nodes, being flexibly deployed in various complex areas to provide network coverage services for users in remote regions. On the other hand, UAVs can carry edge servers, enabling them to approach terminal devices more conveniently and enhance the efficiency of task computation. However, in the face of a large number of computing tasks, the load balancing, network performance and computational resource management of the network are facing serious challenges due to the differences in task density caused by the irregular movement of ground users (GUs), as well as the limitations of the UAV’s computational capability and coverage. In order to address the above challenges, in this paper, We propose a multi-UAV collaborative MEC system. Our goal is to offload tasks as efficiently as possible with full and reasonable utilisation of computational resources, and to evenly distribute computational resources to meet the needs of varying levels of task intensity in a region through rational deployment of UAV locations and collaboration among UAVs. Specifically, in our proposed system framework, GUs can offload tasks to their associated UAVs and further decide whether or not to offload tasks to collaborative UAV or remote Base station (BS) based on the load of computational resources. We study the problem of location deployment of multiple UAVs for better offloading and resource allocation to tasks. Then, we model the task offloading and resource allocation process as a Markov decision process (MDP) focusing on minimising the average user cost of the system and propose a deep reinforcement learning (DRL)-based multi-uav collaborative computation offloading and resource allocation (DMCOA) algorithm. It has been shown through extensive simulation experiments that the algorithm can effectively reduce the average user cost of the MEC system. Mingyang Xu, Xin Chen 0018, Libo Jiao, Xiaoya Fan |
ISPA | 5 |
| 2024 | A Domain Adaption Approach for EEG-Based Automated Seizure Classification with Temporal-Spatial-Spectral Attention
Xiaoya Fan, Pengzhi Xu, Qi Zhao 0008, Chenru Hao, Zhong Wang 0001 |
MICCAI (5) | 1 |
| 2024 | NDPGNN: A Near-Data Processing Architecture for GNN Training and Inference AccelerationabstractGraph neural networks (GNNs) require a large number of fine-grained memory accesses, which results in inefficient use of bandwidth resources. In this article, we introduce a near-data processing architecture tailored for GNN acceleration, named NDPGNN. NDPGNN provides different operating modes to meet the acceleration needs of various GNN frameworks while ensuring the configurability and scalability of the system. NDPGNN takes advantage of data locality characteristics to repeatedly distribute and utilize data, thereby reducing memory access requirements, and further improving memory access efficiency by combining a subgraph sparse node scheduling strategy with intermediate result reuse. We use data packaging to provide a higher effective data ratio for long-distance data transmission, thereby improving the utilization of the system’s limited bandwidth resources. Compared with the previous method, NDPGNN brings 5.68 times improvement in system performance while reducing energy consumption overhead by 8.49 times. Haoyang Wang 0014, Shengbing Zhang, Xiaoya Fan, Zhao Yang 0005, Meng Zhang 0047 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | RE-Specter: Examining the Architectural Features of Configurable CNN With Power Side-ChannelabstractAs domain-specific training data is recognized as valuable intellectual property, acquiring well-trained weights in Convolutional Neural Networks (CNN) has emerged as a new threat to the neural network design community. To design a CNN accelerator that is resilient to side-channel threats, it is crucial to have an accurate and efficient security-driven framework at the early design stage. However, there is no standard way to perform root-cause analysis on the power side channel that exists in FPGA-based CNN accelerators. Therefore, we build RE-Specter, a framework that facilitates security-driven design space exploration (DSE) across various building components, combination patterns, and parallelism configurations in CNNs. The goal is to fully understand the power side-channel effects resulting from architectural modifications or optimization decisions. We further compare the benchmarks considering precision, resource utilization, and power side-channel leakage. Finally, we experimentally explore the design space of various architectural features. The experimental results show that low-bit precision delivers more secure architectures (68.9× among DSPs, 2439× among LUTs) in Measurement-To-Disclosure (MTD), but mixed-precision strategies are necessary to maintain the model accuracy. For loop optimization, in 16-parallel scenario, accumulator-based architecture outperforms the architecture featuring an adder tree with the improvements of 8.28× in MTD and 1.38× in PST. Lu Zhang 0074, Jingyu Wang 0004, Ruoyang Liu, Yifan He 0003, Yaolei Li, Yu Tai, Shengbing Zhang, Xiaoya Fan, Huazhong Yang, Yongpan Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2024 | More Modalities Mean Better: Vessel Target Recognition and Localization Through Symbiotic Transformer and Multiview RegressionabstractVessel target recognition and localization are typically modeled using underwater acoustic signals, which contain a large amount of vessel operating characteristics and condition information. However, extracting operating characteristics from single signals faces heavy noise and non-stationarity challenges. Meanwhile, feature extraction using multimodal data faces the challenges of conflicting gradients between different modalities and ensuring the separability of vessel targets. To tackle these issues, we propose an audio-visual-textual features fusion method to recognize and localize vessel targets through Symbiotic Transformer (Symb-Trans) and Multi-View Regression (MVR) models. Specifically, the audio-visual samples are first preprocessed into paired time series and then projected into a unified optimization landscape via a Heterogeneous Batch Normalization (HetBN) layer to avoid gradient conflicts. Second, the Symb-Trans trains parallel encoders with cross-modal attention and embeds audio-visual representations for vessel target recognition. Finally, the MVR method learns neighboring target properties of a graph model from different perspectives, audio-visual-textual representations, to infer the collector-target distance. Since no off-the-shell multimodal dataset is available for vessel targets, we combine multiple public datasets, consisting of acoustic, and/or visual, and/or textural data, to obtain multimodal materials for model training and validation. Through experimental results and theoretical analysis, we show that Symb-Trans and MVR models outperform unimodal and generic multimodal state-of-the-art solutions for vessel target recognition and localization. Shipei Liu, Xiaoya Fan, Guowei Wu 0001, Lin Yao 0001, Shisong Geng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | CSDSE: Apply Cooperative Search to Solve the Exploration-Exploitation Dilemma of Design Space Exploration
Kaijie Feng, Xiaoya Fan, Jianfeng An, Haoyang Wang 0014, Chuxi Li |
ICA3PP (4) | 2 |
| 2023 | A Wide Conversion Ratio Three-Level DC-DC Converter With Loop-Free Self-Balancing Technique of Flying CapacitorabstractThis paper proposes a loop-free self-balancing technique of flying capacitor for three-level buck converter. The pro-posed converter utilizes two balancing switches in self-balancing circuit, which can adaptively change the series and parallel structure of flying capacitors at different working phases. In two operating modes of duty cycle less than 0.5 and greater than 0.5, the converter can stabilize the flying capacitor voltages at$V_{IN}\mathbf{/2}$without any balancing feedback calibration loops. This three-level converter prototype is designed in$\mathbf{0.18\mu \mathrm{m}}$BCD process. When input voltage is set to 6V, the converter can provide a 0.4-5.6V wide output voltage range and automatically calibrate flying capacitor voltage to 3V under different conversion ratios and load currents. Simulation results show that the maximum output voltage ripple and inductor current ripple are 2.5mV and 174mA respectively when conversion ratio is equal to 0.25 or 0.75. Zhitong Chen, Shiying Liu, Xiaoya Fan, Yanzhao Ma |
ISCAS | 6 |
| 2023 | The Power of Fragmentation: A Hierarchical Transformer Model for Structural Segmentation in Symbolic Music GenerationabstractSymbolic music generation relies on the contextual representation capabilities of the generative model, where the most prevalent approach is the Transformer-based model. Learning contextual representations are also related to the structural elements in music, i.e., intro, verse, and chorus, which have not received much attention of scientific publications. In this paper, we propose a hierarchical Transformer model to learn multiscale contexts in music. In the encoding phase, we first design a fragment scope localization module to separate the music parts into chords and sections. Then, we use a multiscale attention mechanism to learn note-, chord-, and section-level contexts. In the decoding phase, we propose a hierarchical Transformer model that uses fine decoders to generate sections in parallel and a coarse decoder to decode the combined music. We also designed a music style normalization layer to achieve a consistent music style between the generated sections. Our model is evaluated on two open MIDI datasets. Experiments show that our model outperforms other comparative models in 50% (6 out of 12 metrics) and 83.3% (10 out of 12 metrics) of the quantitative metrics for short- and long-term music generation, respectively. Preliminary visual analysis also suggests its potential in following compositional rules, such as reuse of rhythmic patterns and critical melodies, which are associated with improved music quality. Guowei Wu 0001, Shipei Liu, Xiaoya Fan |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | A Noise-Driven Heterogeneous Stochastic Computing Multiplier for Heuristic Precision Improvement in Energy-Efficient DNNsabstractStochastic computing (SC) has become a promising approximate computing solution by its negligible resource occupancy and ultralow energy consumption. As a potential replacement of accurate multiplication, SC can dramatically mitigate the problematic power consumption by DNNs. However, current SC-multipliers illustrate an extremely imbalanced accuracy across product space, i.e., neglectable noise with large products but significant noise for small ones, which is discordant to the distribution of products by the sparse matrix in neural computing. In this article, we present a heterogeneous SC-multiplier that heuristically performs three divergent approximating multiplication, including “set-to-0,” “look-up-table,” and “low-discrepancy-SC,” for appropriate precision-provision in the whole space of products. Due to those popular DNN models cannot achieve consensus on the boundaries of above operations, a training-involved method is proposed to determine the settings with limited overhead. In this way, those models successively learn the SC-operation characters and exhibit a definitely improvement on network precision. The experiment shows that, for single multiplication, the product noise can be restrained by 36.86% on average, and for multiplication in multiple network models, the accuracy improvement reaches to 5.5% on average. Furthermore, a group of proposed logic-reduction techniques can improve the energy efficiency by 65% in the system-level evaluation. Danghui Wang, Shengbing Zhang, Xiaoya Fan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | ACDSE: A Design Space Exploration Method for CNN Accelerator based on Adaptive Compression MechanismabstractCustomized accelerators for Convolutional Neural Network (CNN) can achieve better energy efficiency than general computing platforms. However, the design of a high-performance accelerator should take into account a variety of parameters and physical constraints. The increasing parameters and tighter constraints gradually complicate the design space, which poses new challenges to the capacity and efficiency of design space exploration methods. In this paper, we provide a novel design space exploration method named ACDSE for optimizing the design process of CNN accelerators. ACDSE implements the adaptive compression mechanism to dynamically adjust the search range and prune low-value design points according to the exploration states. As a result, it can focus on valuable subspace while also improving exploration capacity and efficiency. Additionally, we implement ACDSE to address the problem of CNN accelerator latency optimization. The experiment indicates that, compared to former DSE methods, ACDSE can reduce latency and increase efficiency by 1.39x-5.07x and 2.07x-43.87x, respectively, under the most stringent constraint conditions, demonstrating its superior adaptability to the complicated design space. Kaijie Feng, Xiaoya Fan, Jianfeng An, Chuxi Li, Kaiyue Di, Jiangfei Li |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | A 6.78MHz Regulating Rectifier With Constant On-Time Control for High Resolution and Ultra-Fast Transient ResponseabstractThis paper presents a novel constant on-time (COT) single-stage reconfigurable regulating rectifier for wireless power transfer system. Compared with the discrete multi-cycle pulsewidth modulation technique adopting in the prior-art regulating rectifiers, the proposed rectifier achieves high regulating resolution by adjusting the duty ratio continuously and smoothly. In addition, the load-transient response is much faster with the COT control method. In order to optimize the output voltage ripples, the regulating cycle is designed to be small via minimizing the on-time. This single-stage wireless power receiver is designed in a 0.18$\mu$m BCD process, and regulates an output voltage of 5V. Simulation results show that the maximum output power is 850mW with a peak efficiency of 93.2% and a minimum regulating cycle of 0.12$\mu$s. For a load switching between 50mA and 170mA, the response of the proposed rectifier is instant with the unnoticeable overshoot and undershoot. Kai Cui 0006, Xiaoya Fan, Yanzhao Ma |
ISCAS | 3 |
| 2022 | A Current-Injection-Based Flying Capacitor Balancing Circuit for Three-Level DC-DC ConverterabstractThis paper presents a current-injection-based balancing circuit for three-level DC-DC converter. Compared with the method of adjusting duty cycle to achieve flying capacitor balance, the proposed circuit avoids the unbalanced problem without affecting the frequency. In order to adjust the flying capacitor voltage adaptively, a digital controlled feedback loop with 9-bit counter is adopted. The three-level DC-DC converter is designed in a 0.18μ m BCD process, and the input voltage range is 3-6V. Simulation results show that the proposed converter has a good performance of flying capacitor balance with a peak efficiency of 96.8%. The output voltage ripple and inductor current ripple can reach 0.15% and 1.06% of the output voltage and load current, respectively. Zhitong Chen, Shiying Liu, Xiaoya Fan, Yanzhao Ma |
ISCAS | 4 |
| 2022 | A High-Voltage Inverting Converter Based on COT Controlled Buck Regulator with On-Chip Ripple Compensation TechniqueabstractThis paper presents a high-voltage inverting converter implemented by reconfiguring a buck DC-DC regulator. The negative output voltage is generated from a positive input voltage by exchanging the output and ground on a buck converter. In this way, the complexity of circuit design is greatly reduced. The converter adopts constant on-time control with on-chip ripple compensation and output DC offset elimination techniques to achieve fast transient response and highly accurate output voltage, which also allows the use of output capacitor with low equivalent series resistance (RESR). In addition, an internal bootstrap circuit for high-side driver is proposed in place of conventional LDO structure to improve the efficiency. The converter is fabricated in 0.18-$\mu$m BCD process. The input voltage is in the range from 5V to 48V. The measurement results show that the transient response is about 220$\mu$s and the overshoot or undershoot voltage is less than 190mV when the load transient between 100mA and 500mA for switching frequency of 300kHz, input voltage of 24V and output voltage of-12V. Yanzhao Ma, Zhitong Chen, Xiaoya Fan |
ISCAS | 6 |
| 2022 | DCNN search and accelerator co-design: Improve the adaptability between NAS frameworks and embedded platforms
Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047 |
Integr. | 2 |
| 2022 | An Automatic-Addressing Architecture With Fully Serialized Access in Racetrack Memory for Energy-Efficient CNNsabstractRacetrack memory, an emerging low-power magnetic memory, promises a competitive replacement for traditional memory in the accelerators. However, random access in racetrack memory is time and energy expenditure for CNN accelerators because of its large amount of invalid-shifts. In this article, we propose an automatic-addressing architecture that builds a novel data layout to guarantee that the next round of memory access can be always satisfied at the in-situ or rigorously adjacent cells of current round, producing a fully serialized access footprint that can drive instant port-alignment without any invalid-shifts in racetrack memory. By this way, original address-based access degrades to the selections repeated among the three candidates, i.e., onein-situcell and two neighbor cells. Based on this simplification, a lightweight access management can generate the sequence of one-out-three selections according to the deterministic access behaviors defined by CNN hyper-parameters. The evaluation shows that, when deploying the five popular CNN applications to our architecture, the physical shifts of racetrack is curtailed by 74.64 percent over legacy layout, which achieves 54.2 and 42.1 percent energy reduction on read and write, respectively. A case study of YOLOv2 indicates that our architecture performs 6.503 GOp/J that achieves$18.5 \times$improvement to server-level GPUs. Danghui Wang, Jianfeng An, Xiaoya Fan |
IEEE Trans. Computers | 5 |
| 2022 | MemUnison: A Racetrack-ReRAM-Combined Pipeline Architecture for Energy-Efficient in-Memory CNNsabstractThough ReRAM has been greatly successful in reducing energy consumption of various neural networks, it still suffers write amplification in energy, which impedes ReRAM to provide efficient storage for the ubiquitous streaming data in CNNs, such as feature-maps. Racetrack memory, an emerging magnetic memory technique, is a proper candidate to hold streaming data since it enjoys fast sequential-access with ultra-low operating energy in read and write. In this work, we propose a hybrid processing-in-memory architecture, called MemUnison, that coordinates ReRAM and racetrack to overcome the expenditure storage of streaming data in ReRAM. By placing feature-maps in racetrack and leaving weights in ReRAM, a datapath is constructed between the two sides to form a fetch-process-writeback pipeline. As the invalid-shifts of the racetrack memory incurs a large amount of pipeline bubble, we propose a row-based access that can read and write a feature-map without any invalid-shifts. For the row-based operation, a cohesive controlling method is proposed to coordinate racetrack and ReRAM. In runtime, convolution kernels are scheduled in ReRAM banks for cross-channel calculations of one row, by which computing complexity of a convolutional layer can be reduced by 4 orders of magnitude, excessing the 2 order of reduction by traditional ReRAM. Danghui Wang, Shengbing Zhang, Xiaoya Fan |
IEEE Trans. Computers | 5 |
| 2022 | Memory-Computing Decoupling: A DNN Multitasking Accelerator With Adaptive Data ArrangementabstractMultiple deep neural networks (DNNs) are increasingly used in real-world intelligent applications, such as intelligent robotics and autonomous vehicles to collectively complete complicated tasks running on edge devices. Because each layer of the subtasks prefers a distinct dataflow due to the heterogeneity in shape and scale of the network layers, a variable dataflow approach on the DNN accelerators is urgently required. On DNN accelerators that enable multiple dataflows, however, we detect a dimension mismatch between parallel processing under the dataflow approach and linear data memory arrangement. When multiple DNN tasks share partial features or weights, the issue is further exacerbated. During processing, this mismatch causes a sluggish data supply from both off-chip and on-chip memory. Consequently, the overall throughput, performance, and energy efficiency suffer since DNN models are sensitive to data density. In this work, we reveal the mechanism behind this data dimension mismatch and present a series of metrics that quantify the influence on system performance. On this foundation, we offer a framework that tracks the data tensor dimension conversion and employs a flexible data arrangement over multi-DNN computation to adapt to dataflow variability. An accelerator architecture named data arrangement multi-DNN accelerator (DARMA) that features a data arrangement and distribution circuit and hierarchical memory for data dimension conversion is also presented. Since the mismatch is mitigated, the suggested accelerator outperforms current accelerators in terms of bandwidth and processing unit utilization. Through tests on VR/AR, MLperf, and other multitask applications, the evaluation results show that the proposed architecture provides both energy-efficiency and throughput improvements. Chuxi Li, Xiaoya Fan, Xiaoti Wu, Zhao Yang 0005, Meng Zhang 0047, Shengbing Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Hardware-Aware NAS Framework with Layer Adaptive Scheduling on Embedded SystemabstractNeural Architecture Search (NAS) has been proven to be an effective solution for building Deep Convolutional Neural Network (DCNN) models automatically. Subsequently, several hardware-aware NAS frameworks incorporate hardware latency into the search objectives to avoid the potential risk that the searched network cannot be deployed on target platforms. However, the mismatch between NAS and hardware persists due to the absent of rethinking the applicability of the searched network layer characteristics and hardware mapping. A convolution neural network layer can be executed on various dataflows of hardware with different performance, with which the characteristics of on-chip data using varies to fit the parallel structure. This mismatch also results in significant performance degradation for some maladaptive layers obtained from NAS, which might achieved a much better latency when the adopted dataflow changes. To address the issue that the network latency is insufficient to evaluate the deployment efficiency, this paper proposes a novel hardware-aware NAS framework in consideration of the adaptability between layers and dataflow patterns. Beside, we develop an optimized layer adaptive data scheduling strategy as well as a coarse-grained reconfigurable computing architecture so as to deploy the searched networks with high power-efficiency by selecting the most appropriate dataflow pattern layer-by-layer under limited resources. Evaluation results show that the proposed NAS framework can search DCNNs with the similar accuracy to the state-of-the-art ones as well as the low inference latency, and the proposed architecture provides both power-efficiency improvement and energy consumption saving. Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047 |
ASP-DAC | 2 |
| 2021 | A Fully-Integrated Reference-Free Relaxation Oscillator with No ComparatorsabstractA fully-integrated relaxation oscillator with a typical frequency of 1.37 MHz is proposed. Neither comparators nor the reference voltage is required in the proposed oscillator. A constant with temperature (CWT) current source has been utilized to achieve good temperature stability. Moreover, a conventional comparator has been replaced with a voltage controlled delay element (VCDE) to avoid the comparator offset effect. Furthermore, frequency variation against supply voltage has been eliminated by matching bias voltages of the current starved delay element (CSDE) in the ramp generator and VCDE. The proposed relaxation oscillator is implemented with 0.25 μm BCD process. The measured results show that the frequency variation against supply voltage is within ±0.6%. Yanzhao Ma, Zhengjie Ye, Zhitong Chen, Kai Cui 0006, Xiaoya Fan |
ISCAS | 6 |
| 2021 | Balancing memory-accessing and computing over sparse DNN accelerator via efficient data packaging
Xiaoya Fan, Tengteng Yao, Danghui Wang |
J. Syst. Archit. | 2 |
| 2021 | Review of machine learning methods for RNA secondary structure predictionabstractSecondary structure plays an important role in determining the function of noncoding RNAs. Hence, identifying RNA secondary structures is of great value to research. Computational prediction is a mainstream approach for predicting RNA secondary structure. Unfortunately, even though new methods have been proposed over the past 40 years, the performance of computational prediction methods has stagnated in the last decade. Recently, with the increasing availability of RNA structure data, new methods based on machine learning (ML) technologies, especially deep learning, have alleviated the issue. In this review, we provide a comprehensive overview of RNA secondary structure prediction methods based on ML technologies and a tabularized summary of the most important methods in this field. The current pending challenges in the field of RNA secondary structure prediction and future trends are also discussed. Qi Zhao 0008, Xiaoya Fan, Zhengwei Yuan, Yu-Dong Yao |
PLoS Comput. Biol. | 3 |
| 2020 | ENAS oriented layer adaptive data scheduling strategy for resource limited hardware
Chuxi Li, Xiaoya Fan, Yuling Geng, Danghui Wang |
Neurocomputing | 2 |
| 2013 | Correctly rounded architectures for Floating-Point multi-operand addition and dot-product computationabstractThis study presents hardware architectures performing correctly rounded Floating-Point (FP) multioperand addition and dot-product computation, both of which are widely used in various fields, such as scientific computing, digital signal processing, and 3D graphic applications. A novel realignment method is proposed to solve the catastrophic cancellation and multi-sticky bits. Only one rounding operation is performed in both of the proposed FP multi-operand adder and dot-product computation unit. Implementation results show that our architectures not only can produce correctly rounded results, whose errors are less than 0.5 ULP (Unit in the Last Place), but also have reduced delay comparing with the traditional network architecture, which uses 2-operand FP adders and multipliers to perform multi-operand addition and dot-product computation. Deyuan Gao, Xiaoya Fan, Jari Nurmi |
ASAP | 3 |
| 2012 | Analog layout retargeting with geometric programming and constrains symbolization methodabstractTo satisfy the requirements of complex and special analog layout constraints, a constrains symbolization method based on geometric programming for analog layout retargeting is presented in this paper. The approach is to build symbolic template for layouts, then uses geometric programming (GP) to achieve new technology design rules, implement device symmetry and matching constraints, and manage parasitics optimization. The GP, a class of non-linear optimization problem, can be transferred or fitted into a convex optimization problem. Therefore, a global optimum solution can be achieved. The symbolization method ensures the layout retargeting automatically. The efficiency and effectiveness of the proposed algorithm, as compared with the other existing methods, are demonstrated by a basic case-study example and a two-stage Miller-compensated operational amplifier. Shaoxi Wang, Xiaoya Fan, Shengbing Zhang, Ming-e Jing |
ISCAS | 2 |
| 2008 | Improving Performance of Partial Reconfiguration Using Strategy of Virtual DeletionabstractIn a partially reconfigurable system with online placement algorithm, we try to avoid mapping some redundant tasks by caching modules on the reconfigurable area. This paper proposes an elaborate strategy named virtual deletion and a low cost board- level hardware named recycle cache to accomplish the goal. In our strategy, the record of corresponding module is deleted from placer and indexed in the recycle cache. If the module might be used by following tasks, it can be restored from reconfigurable area by recycle cache immediately, without mapping the module again. Recycle cache can shorten average configuring time of partial reconfiguration without increasing arithmetic complex and placing time of the placer. Compared with large size of local register file which cache context of modules, the recycle cycle is much smaller and cheaper. Simulation results on large random tasks sets have shown that the recycle cache can improve performance of partially reconfigurable system effectively. Hangpei Tian, Deyuan Gao, Xiaoya Fan |
FCCM | 4 |
| 2007 | Embedded System's Performance Analysis with RTC and QT
Fulong Chen 0002, Xiaoya Fan |
APPT | 2 |
| 2007 | DAC Circuit with Multi-threshold Voltage for TFT-LCD Driver ICabstractA new DAC circuit with multi-threshold voltage for large panel TFT-CLD source driver is proposed based on its binary-tree structural characters. Through setting different bulk voltages VBfor different CMOS analog switches, the threshold voltages and the on-resistance of CMOS analog switches are reduced and the signal transmission speed from resistance network to output buffer is increased greatly. Physical implementation of this structure is simple and no extra components are required. The proposed DAC circuit structure with 10-bit resolution is designed and simulated using 0.35 mum 13.5 V CMOS high-voltage process, the SPICE simulation results show that the step response delay time is reduced from 35 ns to 17.4 ns (by 50%), as compared to conventional DAC structure. Wei Wu 0014, Tingcun Wei, Xiaoya Fan, Fulong Chen 0002 |
CAD/Graphics | 3 |