VLDB 2026 Research / reviewers in the wild / expert
Yawen Wu
dblp:230/8649
· DBLP profile ↗
37ranked-venue papers
12as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpecAgent: A Speculative Retrieval and Forecasting Agent for Code CompletionabstractGeorge Ma, Anurag Koul, Qi Chen, Yawen Wu, Sachit Kuhar, Yu Yu, Aritra Sengupta, Varun Kumar, Murali Krishna Ramanathan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. George Ma, Anurag Koul, Yawen Wu, Sachit Kuhar, Aritra Sengupta, Murali Krishna Ramanathan |
ACL (1) | 4 |
| 2025 | DLF: Disentangled-Language-Focused Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis (MSA) leverages heterogeneous modalities, such as language, vision, and audio, to enhance the understanding of human sentiment. While existing models often focus on extracting shared information across modalities or directly fusing heterogeneous modalities, such approaches can introduce redundancy and conflicts due to equal treatment of all modalities and the mutual transfer of information between modality pairs. To address these issues, we propose a Disentangled-Language-Focused (DLF) multimodal representation learning framework, which incorporates a feature disentanglement module to separate modality-shared and modality-specific information. To further reduce redundancy and enhance language-targeted features, four geometric measures are introduced to refine the disentanglement process. A Language-Focused Attractor (LFA) is further developed to strengthen language representation by leveraging complementary modality-specific information through a language-guided cross-attention mechanism. The framework also employs hierarchical predictions to improve overall accuracy. Extensive experiments on two popular MSA datasets, CMU-MOSI and CMU-MOSEI, demonstrate the significant performance gains achieved by the proposed DLF framework. Comprehensive ablation studies further validate the effectiveness of the feature disentanglement module, language-focused attractor, and hierarchical predictions. Pan Wang 0013, Yawen Wu, Tianlong Chen 0001, Jingtong Hu |
AAAI | 3 |
| 2025 | Rethinking Medical Anomaly Detection in Brain MRI: An Image Quality Assessment PerspectiveabstractReconstruction-based methods, particularly those leveraging autoencoders, have been widely adopted for anomaly detection task in brain MRI. Unlike most existing works try to improve the task accuracy through architectural or algorithmic innovations, we tackle this task from image quality assessment (IQA) perspective, an under-explored direction in the field. Due to the limitations of conventional metrics such as £1 in capturing the nuanced differences in reconstructed images for medical anomaly detection, we propose fusion quality, a novel metric that wisely integrates the structure-level sensitivity of Structural Similarity Index Measure (SSIM) with the pixel-level precision of £1. The metric offers a more comprehensive assessment of reconstruction quality, considering intensity (subtractive property of l1and divisive property of SSIM), contrast, and structural similarity. Furthermore, the proposed metric makes subtle regional variations more impactful in the final assessment. Thus, considering the inherent divisive properties of SSIM, we design an average intensity ratio (AIR)-based data transformation that amplifies the divisive discrepancies between normal and abnormal regions, thereby enhancing anomaly detection. By fusing the aforementioned two components, we devise the IQA approach. Experimental results on two distinct brain MRI datasets show that our IQA approach significantly enhances medical anomaly detection performance when integrated with state-of-the-art baselines. Code is provided here. Zixuan Pan, Jun Xia 0003, Zheyu Yan, Guoyue Xu, Yawen Wu, Zhenge Jia, Jianxu Chen 0001, Yiyu Shi 0001 |
BIBM | 7 |
| 2025 | STHVC: Spatial-Temporal Hybrid Video Compression for UAV-Assisted IoV SystemsabstractRecent rapid advancements in intelligent vehicular systems and deep learning techniques have led to the emergence of diverse applications utilizing high-quality automotive videos in the Internet-of-Vehicles (IoV), often assisted by uncrewed aerial vehicles (UAVs). These applications aim to provide convenience and security for users. However, transmitting automotive videos with high-quality and low-bit-rate poses a challenge due to the inherent lossiness of traditional compression codecs in current UAV-assisted IoV systems, thereby affecting the performance of subsequent tasks. To address this, we propose a spatial-temporal hybrid video compression framework (STHVC), which integrates Space-Time Super-Resolution (STSR) with conventional codecs to enhance the compression efficiency on automotive videos. In our hybrid design, the encoder generates a low-frame-rate and low-resolution version of the source video, which is then compressed using a traditional codec. During the decoding stage, an effective STSR network is developed to increase both the resolution and the frame rate, and mitigate compression artifacts for automotive videos simultaneously. Additionally, we introduce a rectified intermediate flow estimation technique (RecIFE) within the proposed STSR network to address the challenge of noisy and inaccurate motions during the compression pipeline. Extensive experiments on various benchmark datasets demonstrate that our approach achieves bit-rate reductions of 29.97% compared to H.265 (slow) and 31.27% compared to H.266, while also exhibiting superior restoration performance compared to other state-of-the-art learning-based approaches. Lvcheng Chen, Jianing Deng, Xudong Zeng, Liangwei Liu, Yawen Wu, Jingtong Hu, Qi Sun 0002, Zhiguo Shi 0001, Cheng Zhuo |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Enhancing 3D Transformer Segmentation Model for Medical Image with Token-level Representation LearningabstractIn the field of medical images, although various works find Swin Transformer has promising effectiveness on pixelwise dense prediction, whether pre-training these models without using extra dataset can further boost the performance for the downstream semantic segmentation remains unexplored. Applications of previous representation learning methods are hindered by the limited number of 3D volumes and high computational cost. In addition, most of pretext tasks designed specifically for Transformer are not applicable to hierarchical structure of Swin Transformer. Thus, this work proposes a token-level representation learning loss that maximizes agreement between token embeddings from different augmented views individually instead of volume-level global features. Moreover, we identify a potential representation collapse exclusively caused by this new loss. To prevent collapse, we invent a simple "rotate-and-restore" mechanism, which rotates and flips one augmented view of input volume, and later restores the order of tokens in the feature maps. We also modify the contrastive loss to address the discrimination between tokens at the same position but from different volumes. Results on a public experiment more improvement of our methods than other state-of-the-art pre-trainig methods. Xinrong Hu, Dewen Zeng, Yawen Wu, Yiyu Shi 0001 |
BIBM | 3 |
| 2024 | Contrastive Learning with Synthetic Positives
Dewen Zeng, Yawen Wu, Xinrong Hu, Xiaowei Xu 0004, Yiyu Shi 0001 |
ECCV (37) | 2 |
| 2024 | Unlocking Memorization in Large Language Models with Dynamic Soft PromptingabstractZhepeng Wang, Runxue Bao, Yawen Wu, Jackson Taylor, Cao Xiao, Feng Zheng, Weiwen Jiang, Shangqian Gao, Yanfu Zhang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhepeng Wang 0001, Runxue Bao, Yawen Wu, Jackson Taylor, Cao Xiao, Feng Zheng 0001, Weiwen Jiang, Shangqian Gao, Yanfu Zhang |
EMNLP | 3 |
| 2024 | Self-guided Knowledge-Injected Graph Neural Network for Alzheimer's Diseases
Zhepeng Wang 0001, Runxue Bao, Yawen Wu, Lei Yang 0018, Liang Zhan, Feng Zheng 0001, Weiwen Jiang, Yanfu Zhang |
MICCAI (2) | 3 |
| 2024 | Achieve fairness without demographics for dermatological disease diagnosis
Ching-Hao Chiu, Yawen Wu, Yiyu Shi 0001, Tsung-Yi Ho |
Medical Image Anal. | 3 |
| 2023 | Synthetic Data Can Also Teach: Synthesizing Effective Data for Unsupervised Visual Representation LearningabstractContrastive learning (CL), a self-supervised learning approach, can effectively learn visual representations from unlabeled data. Given the CL training data, generative models can be trained to generate synthetic data to supplement the real data. Using both synthetic and real data for CL training has the potential to improve the quality of learned representations. However, synthetic data usually has lower quality than real data, and using synthetic data may not improve CL compared with using real data. To tackle this problem, we propose a data generation framework with two methods to improve CL training by joint sample generation and contrastive learning. The first approach generates hard samples for the main model. The generator is jointly learned with the main model to dynamically customize hard samples based on the training state of the main model. Besides, a pair of data generators are proposed to generate similar but distinct samples as positive pairs. In joint learning, the hardness of a positive pair is progressively increased by decreasing their similarity. Experimental results on multiple datasets show superior accuracy and data efficiency of the proposed data generation methods applied to CL. For example, about 4.0%, 3.5%, and 2.6% accuracy improvements for linear classification are observed on ImageNet-100, CIFAR-100, and CIFAR-10, respectively. Besides, up to 2× data efficiency for linear classification and up to 5× data efficiency for transfer learning are achieved. Yawen Wu, Zhepeng Wang 0001, Dewen Zeng, Yiyu Shi 0001, Jingtong Hu |
AAAI | 1 |
| 2023 | Transfer Learning-Assisted Survival Analysis of Breast Cancer Relying on the Spatial Interaction Between Tumor-Infiltrating Lymphocytes and Tumors
Yawen Wu, Yingli Zuo, Qi Zhu 0001, Jianpeng Sheng, Daoqiang Zhang, Wei Shao 0005 |
MICCAI (6) | 1 |
| 2023 | Additional Positive Enables Better Representation Learning for Medical Images
Dewen Zeng, Yawen Wu, Xinrong Hu, Xiaowei Xu 0004, Jingtong Hu, Yiyu Shi 0001 |
MICCAI (1) | 2 |
| 2023 | Self-Supervised On-Device Federated Learning From Unlabeled StreamsabstractThe ubiquity of edge devices has led to a growing amount of unlabeled data produced at the edge. Deep learning models deployed on edge devices are required to learn from these unlabeled data to continuously improve accuracy. Self-supervised representation learning has achieved promising performances using centralized unlabeled data. However, the increasing awareness of privacy protection limits centralizing the distributed unlabeled image data on edge devices. While federated learning has been widely adopted to enable distributed machine learning with privacy preservation, without a data selection method to efficiently select streaming data, the traditional federated learning framework fails to handle these huge amounts of decentralized unlabeled data with limited storage resources on edge. To address these challenges, we propose a self-supervised on-device federated learning framework with coreset selection, which we call SOFed, to automatically select a coreset that consists of the most representative samples into the replay buffer on each device. It preserves data privacy as each client does not share raw data while learning good visual representations. Experiments demonstrate the effectiveness and significance of the proposed method in visual representation learning. Jiahe Shi, Yawen Wu, Dewen Zeng, Jun Tao 0001, Jingtong Hu, Yiyu Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Characterizing the Survival-Associated Interactions Between Tumor-Infiltrating Lymphocytes and Tumors From Pathological Images and Multi-Omics DataabstractThe tumor-infiltrating lymphocytes (TILs) and its correlation with tumors have shown significant values in the development of cancers. Many observations indicated that the combination of the whole-slide pathological images (WSIs) and genomic data can better characterize the immunological mechanisms of TILs. However, the existing image-genomic studies evaluated the TILs by the combination of pathological image and single-type of omics data (e.g., mRNA), which is difficulty in assessing the underlying molecular processes of TILs holistically. Additionally, it is still very challenging to characterize the intersections between TILs and tumor regions in WSIs and the high dimensional genomic data also brings difficulty for the integrative analysis with WSIs. Based on the above considerations, we proposed an end-to-end deep learning framework i.e., IMO-TILs that can integrate pathological image with multi-omics data (i.e., mRNA and miRNA) to analyze TILs and explore the survival-associated interactions between TILs and tumors. Specifically, we firstly apply the graph attention network to describe the spatial interactions between TILs and tumor regions in WSIs. As to genomic data, the Concrete AutoEncoder (i.e., CAE) is adopted to select survival-associated Eigengenes from the high-dimensional multi-omics data. Finally, the deep generalized canonical correlation analysis (DGCCA) accompanied with the attention layer is implemented to fuse the image and multi-omics data for prognosis prediction of human cancers. The experimental results on three cancer cohorts derived from the Cancer Genome Atlas (TCGA) indicated that our method can both achieve higher prognosis results and identify consistent imaging and multi-omics bio-markers correlated strongly with the prognosis of human cancers. Wei Shao 0005, Yingli Zuo, Yangyang Shi, Yawen Wu, Jiao Tang, Junyong Zhao, Liang Sun 0009, Zixiao Lu, Jianpeng Sheng, Qi Zhu 0001, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 4 |
| 2022 | The larger the fairer?: small neural networks can achieve fairness for edge devicesabstractAlong with the progress of AI democratization, neural networks are being deployed more frequently in edge devices for a wide range of applications. Fairness concerns gradually emerge in many applications, such as face recognition and mobile medical. One fundamental question arises: what will be the fairest neural architecture for edge devices? By examining the existing neural networks, we observe that larger networks typically are fairer. But, edge devices call for smaller neural architectures to meet hardware specifications. To address this challenge, this work proposes a novel Fairness- and Hardware-aware Neural architecture search framework, namely FaHaNa. Coupled with a model freezing approach, FaHaNa can efficiently search for neural networks with balanced fairness and accuracy, while guaranteed to meet hardware specifications. Results show that FaHaNa can identify a series of neural networks with higher fairness and accuracy on a dermatology dataset. Target edge devices, FaHaNa finds a neural architecture with slightly higher accuracy, 5.28X smaller size, 15.14% higher fairness score, compared with MobileNetV2; meanwhile, on Raspberry PI and Odroid XU-4, it achieves 5.75X and 5.79X speedup. Yi Sheng 0001, Junhuan Yang, Yawen Wu, Kevin Mao, Yiyu Shi 0001, Jingtong Hu, Weiwen Jiang, Lei Yang 0018 |
DAC | 3 |
| 2022 | Opportunistic Communication with Latency Guarantees for Intermittently-Powered DevicesabstractEnergy-harvesting wireless sensor nodes have found widespread adoption due to their low cost and small form factor. However, uncertainty in the available power supply introduces significant challenges in engineering communications between intermittently-powered nodes. We propose a constraint-based model for energy harvests that together with a hardware model can be used to enable consistent, opportunistic communication with worst-case latency guarantees. We show that greedy approaches that attempt communication whenever energy is available lead to prolonged latencies in real-world environments. Our approach offers bounded worst-case latency while providing a performance improvement over a conservative, offline approach planned around the worst-case energy harvest. Kacper Wardega, Wenchao Li 0001, Hyoseung Kim 0001, Yawen Wu, Zhenge Jia, Jingtong Hu |
DATE | 4 |
| 2022 | Decentralized Unsupervised Learning of Visual RepresentationsabstractCollaborative learning enables distributed clients to learn a shared model for prediction while keeping the training data local on each client. However, existing collaborative learning methods require fully-labeled data for training, which is inconvenient or sometimes infeasible to obtain due to the high labeling cost and the requirement of expertise. The lack of labels makes collaborative learning impractical in many realistic settings. Self-supervised learning can address this challenge by learning from unlabeled data. Contrastive learning (CL), a self-supervised learning approach, can effectively learn visual representations from unlabeled image data. However, the distributed data collected on clients are usually not independent and identically distributed (non-IID) among clients, and each client may only have few classes of data, which degrades the performance of CL and learned representations. To tackle this problem, we propose a collaborative contrastive learning framework consisting of two approaches: feature fusion and neighborhood matching, by which a unified feature space among clients is learned for better data representations. Feature fusion provides remote features as accurate contrastive information to each client for better local learning. Neighborhood matching further aligns each client’s local features to the remote features such that well-clustered features among clients can be learned. Extensive experiments show the effectiveness of the proposed framework. It outperforms other methods by 11% on IID data and matches the performance of centralized learning. Yawen Wu, Zhepeng Wang 0001, Dewen Zeng, Meng Li 0004, Yiyu Shi 0001, Jingtong Hu |
IJCAI | 1 |
| 2022 | FairPrune: Achieving Fairness Through Pruning for Dermatological Disease Diagnosis
Yawen Wu, Dewen Zeng, Xiaowei Xu 0004, Yiyu Shi 0001, Jingtong Hu |
MICCAI (1) | 1 |
| 2022 | Identify Consistent Imaging Genomic Biomarkers for Characterizing the Survival-Associated Interactions Between Tumor-Infiltrating Lymphocytes and Tumors
Yingli Zuo, Yawen Wu, Zixiao Lu, Qi Zhu 0001, Kun Huang 0001, Daoqiang Zhang, Wei Shao 0005 |
MICCAI (2) | 2 |
| 2022 | Distributed contrastive learning for medical image segmentation
Yawen Wu, Dewen Zeng, Zhepeng Wang 0001, Yiyu Shi 0001, Jingtong Hu |
Medical Image Anal. | 1 |
| 2022 | Enabling Weakly Supervised Temporal Action Localization From On-Device Learning of the Video StreamabstractDetecting actions in videos have been widely applied in on-device applications, such as cars, robots, etc. Practical on-device videos are always untrimmed with both action and background. It is desirable for a model to both recognize the class of action and localize the temporal position where the action happens. Such a task is called temporal action location (TAL), which is always trained on the cloud where multiple untrimmed videos are collected and labeled. It is desirable for a TAL model to continuously and locally learn from new data, which can directly improve the action detection precision while protecting customers’ privacy. However, directly training a TAL model on the device is nontrivial. To train a TAL model which can precisely recognize and localize each action, tremendous video samples with temporal annotations are required. However, annotating videos frame by frame is exorbitantly time consuming and expensive. Although weakly supervised temporal action localization (W-TAL) has been proposed to learn from untrimmed videos with only video-level labels, such an approach is also not suitable for on-device learning scenarios. In practical on-device learning applications, data are collected in streaming. For example, the camera on the device keeps collecting video frames for hours or days, and the actions of nearly all classes are included in a single long video stream. Dividing such a long video stream into multiple video segments requires lots of human effort, which hinders the exploration of applying the TAL tasks to realistic on-device learning applications. To enable W-TAL models to learn from a long, untrimmed streaming video, we propose an efficient video learning approach that can directly adapt to new environments. We first propose a self-adaptive video dividing approach with a contrast score-based segment merging approach to convert the video stream into multiple segments. Then, we explore different sampling strategies on the TAL tasks to request as few labels as possible. To the best of our knowledge, we are the first attempt to directly learn from the on-device, long video stream. Experimental results on the THUMOS’14 dataset show that the performance of our approach is comparable to the current W-TAL state-of-the-art (SOTA) work without any laborious manual video splitting. Yue Tang 0002, Yawen Wu, Peipei Zhou 0001, Jingtong Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Cooperative Communication Between Two Transiently Powered Sensor Nodes by Reinforcement LearningabstractEnergy harvesting (EH)-powered sensor nodes can achieve theoretically unlimited lifetime by scavenging energy from ambient power sources, such as radio-frequency (RF) and kinetic energy. The nodes can collect and transmit data wirelessly with the harvested energy. However, the transmission between two sensor nodes is successful only when both nodes have enough energy at the same time. While the receiver can be actively listening, it may deplete the energy long before the sender has accumulated enough energy. Thus, given the scarce, unpredictable, and unevenly distributed energy among sensor nodes, it is challenging to ensure efficient data transmission between them. To address this challenge, we propose a sensor node architecture with multiple radios, each with different energy consumption on the sender and receiver. A node can be put into sleep when charged up and wakes up for communication when it infers that both nodes have enough energy based on its observations. What is more, two nodes can cooperatively and dynamically select different radios according to the stored energy and historical information to maximize the data throughput. To achieve cooperative communication adaptively, the communication procedure is modeled as a cooperative Markov game with partial observability on each node, and multiagent reinforcement learning (MARL) is employed to achieve the best results. Experimental results on hardware prototype and by simulation show that the proposed approaches achieve up to 89.1% of the optimal throughput and significantly outperform other online algorithms. Yawen Wu, Zhenge Jia, Fei Fang 0001, Jingtong Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Energy-Aware Adaptive Multi-Exit Neural Network Inference Implementation for a Millimeter-Scale Sensing SystemabstractImplementing a neural network (NN) inference in a millimeter-scale system is challenging due to limited energy and storage size. This article proposes an energy-aware adaptive NN inference implementation that utilizes one of two exits with different accuracies and computation options. The early-exit path provides a shorter processing time but less accuracy than the main-exit path. To compensate for the reduced accuracy, it additionally applies the main-exit path if the entropy of the early-exit inference is higher than a predetermined value. The NN is implemented with a custom low-power 180-nm CMOS processor chip and a 90-nm embedded flash memory chip and tested by the CIFAR-10 dataset. The measurement results show that the implemented convolutional NN (CNN) reduces processing time and thus energy consumption by 43.9% compared with a main-exit-only method while sacrificing its accuracy from 69.9% to 66.2%. Also, we explore the required minimum battery capacity at each optimal configuration for accuracy and/or energy consumption to achieve energy-autonomous operation under measured exemplary light profiles. It requires a minimum battery capacity of 855 mJ, acceptable for the target miniature system with two millimeter-scale batteries (684 mJ each). Compared with the state-of-the-art CNN technique (BranchyNet) allowing early stopping, the proposed design improves the accuracy by 0.7% and 3.3% to maintain energy-autonomous operation with two and one millimeter-scale batteries, respectively. Compared with the state-of-the-art lightweight CNN technique (MobileNet), this work provides flexibility with a tradeoff between accuracy and processing time for different application requirements. Yuyang Li 0001, Yawen Wu, Xincheng Zhang, Jingtong Hu, Inhee Lee 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | Lightweight Run-Time Working Memory Compression for Deployment of Deep Neural Networks on Resource-Constrained MCUsabstractThis work aims to achieve intelligence on embedded devices by deploying deep neural networks (DNNs) onto resource-constrained microcontroller units (MCUs). Apart from the low frequency (e.g., 1-16 MHz) and limited storage (e.g., 16KB to 256KB ROM), one of the largest challenges is the limited RAM (e.g., 2KB to 64KB), which is needed to save the intermediate feature maps of a DNN. Most existing neural network compression algorithms aim to reduce the model size of DNNs so that they can fit into limited storage. However, they do not reduce the size of intermediate feature maps significantly, which is referred to as working memory and might exceed the capacity of RAM. Therefore, it is possible that DNNs cannot run in MCUs even after compression. To address this problem, this work proposes a technique to dynamically prune the activation values of the intermediate output feature maps in the runtime to ensure that they can fit into limited RAM. The results of our experiments show that this method could significantly reduce the working memory of DNNs to satisfy the hard constraint of RAM size, while maintaining satisfactory accuracy with relatively low overhead on memory and run-time latency. Zhepeng Wang 0001, Yawen Wu, Zhenge Jia, Yiyu Shi 0001, Jingtong Hu |
ASP-DAC | 2 |
| 2021 | Enabling On-Device Self-Supervised Contrastive Learning with Selective Data ContrastabstractAfter a model is deployed on edge devices, it is desirable for these devices to learn from unlabeled data to continuously improve accuracy. Contrastive learning has demonstrated its great potential in learning from unlabeled data. However, the online input data are usually none independent and identically distributed (non-iid) and edge devices’ storages are usually too limited to store enough representative data from different data classes. We propose a framework to automatically select the most representative data from the unlabeled input stream, which only requires a small data buffer for dynamic learning. Experiments show that accuracy and learning speed are greatly improved. Yawen Wu, Zhepeng Wang 0001, Dewen Zeng, Yiyu Shi 0001, Jingtong Hu |
DAC | 1 |
| 2021 | Federated Contrastive Learning for Dermatological Disease Diagnosis via On-device Learning (Invited Paper)abstractDeep learning models have been deployed in an increasing number of edge and mobile devices to provide healthcare. These models rely on training with a tremendous amount of labeled data to achieve high accuracy. However, for medical applications such as dermatological disease diagnosis, the private data collected by mobile dermatology assistants exist on distributed mobile devices of patients, and each device only has a limited amount of data. Directly learning from limited data greatly deteriorates the performance of learned models. Federated learning (FL) can train models by using data distributed on devices while keeping the data local for privacy. Existing works on FL assume all the data have ground-truth labels. However, medical data often comes without any accompanying labels since labeling requires expertise and results in prohibitively high labor costs. The recently developed self-supervised learning approach, contrastive learning (CL), can leverage the unlabeled data to pre-train a model for learning data representations, after which the learned model can be fine-tuned on limited labeled data to perform dermatological disease diagnosis. However, simply combining CL with FL as federated contrastive learning (FCL) will result in ineffective learning since CL requires diverse data for accurate learning but each device in FL only has limited data diversity. In this work, we propose an on-device FCL framework for dermatological disease diagnosis with limited labels. Features are shared among devices in the FCL pre-training process to provide diverse and accurate contrastive information without sharing raw data for privacy. After that, the pre-trained model is fine-tuned with local labeled data independently on each device or collaboratively with supervised federated learning on all devices. Experiments on dermatological disease datasets show that the proposed framework effectively improves the recall and precision of dermatological disease diagnosis compared with state-of-the-art methods. Yawen Wu, Dewen Zeng, Zhepeng Wang 0001, Yi Sheng 0001, Lei Yang 0018, Alaina J. James, Yiyu Shi 0001, Jingtong Hu |
ICCAD | 1 |
| 2021 | Developing a Miniature Energy-Harvesting-Powered Edge Device with Multi-Exit Neural NetworkabstractThis paper describes a miniature edge device that performs neural network inference with different exit options depending on available energy. In addition to the main-exit path, it provides an alternative, early-exit path that requires less computation and thus increase the number of inference operations for given energy. To compensate its degraded accuracy, the proposed device provides entropy as a confidence level for the early exit. The network is implemented with a custom low-power 180 nm CMOS processor chip and a 90 nm embedded flash memory chip and tested by images from CIFAR-10 dataset. The measurement results show the proposed neural network reduces processing time and thus energy consumption by 41.3% compared with the main-exit only method while sacrificing its accuracy from 69.5% to 66.0%. Yuyang Li 0001, Yawen Wu, Xincheng Zhang, Ehab A. Hamed, Jingtong Hu, Inhee Lee 0001 |
ISCAS | 2 |
| 2021 | Federated Contrastive Learning for Volumetric Medical Image Segmentation
Yawen Wu, Dewen Zeng, Zhepeng Wang 0001, Yiyu Shi 0001, Jingtong Hu |
MICCAI (3) | 1 |
| 2021 | Positional Contrastive Learning for Volumetric Medical Image Segmentation
Dewen Zeng, Yawen Wu, Xinrong Hu, Xiaowei Xu 0004, Haiyun Yuan, Meiping Huang, Jian Zhuang, Jingtong Hu, Yiyu Shi 0001 |
MICCAI (2) | 2 |
| 2021 | Algorithm-hardware Co-design of Attention Mechanism on FPGA DevicesabstractMulti-head self-attention (attention mechanism) has been employed in a variety of fields such as machine translation, language modeling, and image processing due to its superiority in feature extraction and sequential data analysis. This is benefited from a large number of parameters and sophisticated model architecture behind the attention mechanism. To efficiently deploy attention mechanism on resource-constrained devices, existing works propose to reduce the model size by building a customized smaller model or compressing a big standard model. A customized smaller model is usually optimized for the specific task and needs effort in model parameters exploration. Model compression reduces model size without hurting the model architecture robustness, which can be efficiently applied to different tasks. The compressed weights in the model are usually regularly shaped (e.g. rectangle) but the dimension sizes vary (e.g. differs in rectangle height and width). Such compressed attention mechanism can be efficiently deployed on CPU/GPU platforms as their memory and computing resources can be flexibly assigned with demand. However, for Field Programmable Gate Arrays (FPGAs), the data buffer allocation and computing kernel are fixed at run time to achieve maximum energy efficiency. After compression, weights are much smaller and different in size, which leads to inefficient utilization of FPGA on-chip buffer. Moreover, the different weight heights and widths may lead to inefficient FPGA computing kernel execution. Due to the large number of weights in the attention mechanism, building a unique buffer and computing kernel for each compressed weight on FPGA is not feasible. In this work, we jointly consider the compression impact on buffer allocation and the required computing kernel during the attention mechanism compressing. A novel structural pruning method with memory footprint awareness is proposed and the associated accelerator on FPGA is designed. The experimental results show that our work can compress Transformer (an attention mechanism based model) by 95x. The developed accelerator can fully utilize the FPGA resource, processing the sparse attention mechanism with the run-time throughput performance of 1.87 Tops in ZCU102 FPGA. Xinyi Zhang 0001, Yawen Wu, Peipei Zhou 0001, Xulong Tang, Jingtong Hu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2020 | Intermittent Inference with Nonuniformly Compressed Multi-Exit Neural Network for Energy Harvesting Powered DevicesabstractThis work aims to enable persistent, event-driven sensing and decision capabilities for energy-harvesting (EH)-powered devices by deploying lightweight DNNs onto EH-powered devices. However, harvested energy is usually weak and unpredictable and even lightweight DNNs take multiple power cycles to finish one inference. To eliminate the indefinite long wait to accumulate energy for one inference and to optimize the accuracy, we developed a power trace-aware and exit-guided network compression algorithm to compress and deploy multi-exit neural networks to EH-powered microcontrollers (MCUs) and select exits during execution according to available energy. The experimental results show superior accuracy and latency compared with state-of-the-art techniques. Yawen Wu, Zhepeng Wang 0001, Zhenge Jia, Yiyu Shi 0001, Jingtong Hu |
DAC | 1 |
| 2020 | A dual-domain deep lattice network for rapid MRI reconstruction
Liyan Sun, Yawen Wu, Binglin Shu, Xinghao Ding, Congbo Cai, Yue Huang 0001, John W. Paisley |
Neurocomputing | 2 |
| 2020 | Enabling On-Device CNN Training by Self-Supervised Instance Filtering and Error Map PruningabstractThis work aims to enable on-device training of convolutional neural networks (CNNs) by reducing the computation cost at training time. CNN models are usually trained on high-performance computers and only the trained models are deployed to edge devices. But the statically trained model cannot adapt dynamically in a real environment and may result in low accuracy for new inputs. On-device training by learning from the real-world data after deployment can greatly improve accuracy. However, the high computation cost makes training prohibitive for resource-constrained devices. To tackle this problem, we explore the computational redundancies in training and reduce the computation cost by two complementary approaches: 1) self-supervised early instance filtering on data level and 2) error map pruning (EMP) on the algorithm level. The early instance filter selects important instances from the input stream to train the network and drops trivial ones. The EMP further prunes out insignificant computations when training with the selected instances. Extensive experiments show that the computation and energy cost is substantially reduced without any or with marginal accuracy loss. For example, when training ResNet-110 on CIFAR-10, we achieve 67.8% computation saving while preserving full accuracy and 75.1% computation saving with a marginal accuracy loss of 1.3%. When training LeNet on MNIST, we save 79% computation while boosting accuracy by 0.2%. Besides, practical energy saving is measured on edge platforms. We achieve 67.6% energy saving when training ResNet-110 on mobile GPU and 74.1% energy saving when training LeNet on MCU without any accuracy loss. Yawen Wu, Zhepeng Wang 0001, Yiyu Shi 0001, Jingtong Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Lung Nodule Detection with a 3D ConvNet via IoU Self-normalization and Maxout UnitabstractThe automatic pulmonary nodule detection in thoracic computed tomography (CT) scans plays a crucial role in the early diagnosis of lung cancer. In this paper, we propose a novel framework with a 3D convolutional network (ConvNet) for pulmonary nodule detection. To improve the efficiency and flexibility, we adopt one-stage process without the false positive reduction stage. Specially, the great challenge of the nodule detection is the recall rate of small nodules. We propose two methods to solve this issue. Firstly, we set the classification label by the intersection over union (IoU) self-normalization, which enables to eliminate the loss of regression information caused by misleading classification confidence. Secondly, pulmonary nodules differ in size, shape and density, leading to large intra-class variations. We introduce maxout unit to solve this problem. Overall, we achieve an average FROC score of 0.912 on LUNA16 dataset, outperforming all other one-stage models as far as we know. Fei Li 0021, Yawen Wu, Congbo Cai, Yue Huang 0001, Xinghao Ding |
ICASSP | 3 |
| 2018 | High Efficient Reconstruction of Single-Shot Magnetic Resonance T_2 Mapping Through Overlapping Echo Detachment and DenseNet
Yawen Wu, Xinghao Ding, Yue Huang 0001, Congbo Cai |
ICONIP (6) | 2 |
| 2018 | A Deep Ensemble Network for Compressed Sensing MRI
Huafeng Wu, Yawen Wu, Liyan Sun, Congbo Cai, Yue Huang 0001, Xinghao Ding |
ICONIP (1) | 2 |
| 2018 | Prototyping Energy Harvesting Powered Systems with Nonvolatile Processor (Invited Paper)abstractEnergy harvesting is a promising solution to power ubiquitous Internet-of-Things (IoT) devices. But the frequent and inevitable power failure incurs significant backup overhead, greatly degrading performance and energy efficiency. Nonvolatile processor (NVP), which can checkpoint processor states, is designed to tackle this problem. The conventional system-level design method involves repeated system modification and verification on hardware, in which measurement on hardware consumes the majority time. To expedite the NVP-based system design process, we propose a rapid system prototyping flow to eliminate repeated hardware measurement in the design flow. This method involves an NVP system-level simulator, which takes the harvester power trace, system characteristics extracted from hardware, and user design as the input, and analyzes system energy and time profile under this power trace. Iterative system optimization and verification are conducted on the simulator, with only the final verification on hardware. We demonstrate the advantages of this method by two design cases, in which time, energy efficiency and the impact of different capacitor size are optimized. Yawen Wu, Zhenge Jia, Lefan Zhang, Yongpan Liu, Jingtong Hu |
RSP | 1 |