EDBT 2026 Demo / reviewers in the wild / expert
Dewen Zeng
dblp:128/0618
· DBLP profile ↗
24ranked-venue papers
8as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ripple Shapley: Data Influence Attribution in One Federated Training RunabstractContribution evaluation is essential for incentivizing high-quality data sharing in federated learning (FL), yet existing Shapley-value-based methods are prohibitively expensive and overlook temporal influence propagation. In this paper, we propose Ripple Shapley, a novel attribution framework that enables accurate, real-time data valuation within a single federated training run. Our method decomposes each sample’s impact into an instantaneous drop term and a recursive ripple term, the latter capturing downstream influence via a Jacobian chain over global updates. To scale computation, we introduce a low-rank approximation of the Jacobian product and construct a shared subspace for efficient ripple accumulation. Extensive experiments on CIFAR-10 and MNIST show that Ripple Shapley achieves up to 62× speedup over existing Shapley-based FL methods while maintaining high attribution fidelity, significantly improving efficiency, robustness, and fairness in federated environments. We further demonstrate its effectiveness in dynamic federated learning scenarios and its potential for real-time data pricing. Dewen Zeng, Haozhao Wang, Jianfeng Lu 0002, Weijun Xiao, Zhiyong Xu 0003 |
AAAI | 1 |
| 2026 | Smart-to-Compress: A Predictive and Game-Theoretic Framework for Data Reduction DecisionsabstractWith the rapid growth of data, redundancy among different users in cloud environments has become increasingly prominent. Detecting and removing these redundant parts can effectively improve storage efficiency. But these processes may dramatically degrade the system performance, especially when dealing with similar data. Although deduplication and delta compression are common data reduction techniques, their high overhead can outweigh the benefits. As a result, users often cannot determine in advance whether compression is worthwhile for their datasets. Some approaches have attempted to solve this, but each has important limitations. Danny Harnik et al. proposed a sampling-based deduplication estimation method using linear programming, which efficiently estimates redundancy from exact duplicates. However, it fails to capture redundancy arising from similar data, thus underestimating the full compression potential. To address this limitation, we propose Smart-to-Compress, a predictive compression decision framework. We introduce the Super Feature Frequency Histogram (SFH) to capture redundancy among similar data. Combined with the Duplication Frequency Histogram (DFH), our method estimates the overall Data Reduction Ratio (DRR) without scanning the entire dataset. Furthermore, we design a game-theoretic decision model to weigh compression benefits against predicted costs, providing users with guidance on whether compression should be applied. Experiments on real-world datasets show that our method accurately predicts compression value, reduces unnecessary overhead, and offers reliable decision-making support for users. Zhenrui He, Zhixiong Xie, Dewen Zeng, Jianfeng Lu 0002, Zhiyong Xu 0003, Weijun Xiao, Yaping Wan |
IEEE Trans. Cloud Comput. | 4 |
| 2025 | IBNR-RD: Intra-Block Neighborhood Relationship-Based Resemblance Detection for High-Performance Multi-Node Post-DeduplicationabstractPost-deduplication in traditional cloud environments primarily focuses on single-node, where delta compression is performed on the same deduplication node located on server side. However, with data explosion, the multi-node post-deduplication, also called global deduplication, has become a hot issue in research communities, which aims to simultaneously execute delta compression on data distributed across all nodes. Simply setting up single-node deduplication systems on multi-node environments would significantly affect storage utilization and incur secondary overhead from file migration. Nevertheless, existing global deduplication solutions suffer from lower data compression ratios and high computational overhead due to their resemblance detection's inherent limitations and overly coarse granularities. Similar blocks typically have high correlations between sub-blocks; inspired by this observation, we propose IBNR (Intra-Block Neighborhood Relationship-Based Resemblance Detection for High-Performance Multi-Node Post-Deduplication), which introduces a novel resemblance detection based on relationships between sub-blocks and determines the ownership of blocks in entry stage to achieve efficient global deduplication. Furthermore, the by-products of IBNR have shown powerful scalability by replacing internal resemblance detection scheme with existing solutions on practical workloads. Experimental results indicate that IBNR outperforms state-of-the-art solutions, achieving an average 1.99× data reduction ratio and varying degrees of improvement across other key metrics. Dewen Zeng, Ruixuan Li 0001, Xuming Ye, Zhiyong Xu 0003 |
IEEE Trans. Cloud Comput. | 1 |
| 2024 | Enhancing 3D Transformer Segmentation Model for Medical Image with Token-level Representation LearningabstractIn the field of medical images, although various works find Swin Transformer has promising effectiveness on pixelwise dense prediction, whether pre-training these models without using extra dataset can further boost the performance for the downstream semantic segmentation remains unexplored. Applications of previous representation learning methods are hindered by the limited number of 3D volumes and high computational cost. In addition, most of pretext tasks designed specifically for Transformer are not applicable to hierarchical structure of Swin Transformer. Thus, this work proposes a token-level representation learning loss that maximizes agreement between token embeddings from different augmented views individually instead of volume-level global features. Moreover, we identify a potential representation collapse exclusively caused by this new loss. To prevent collapse, we invent a simple "rotate-and-restore" mechanism, which rotates and flips one augmented view of input volume, and later restores the order of tokens in the feature maps. We also modify the contrastive loss to address the discrimination between tokens at the same position but from different volumes. Results on a public experiment more improvement of our methods than other state-of-the-art pre-trainig methods. Xinrong Hu, Dewen Zeng, Yawen Wu, Yiyu Shi 0001 |
BIBM | 2 |
| 2024 | Contrastive Learning with Synthetic Positives
Dewen Zeng, Yawen Wu, Xinrong Hu, Xiaowei Xu 0004, Yiyu Shi 0001 |
ECCV (37) | 1 |
| 2024 | Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory ArchitecturesabstractLarge Language Models (LLMs) deployed on edge devices learn through fine-tuning and updating a certain portion of their parameters. Although such learning methods can be optimized to reduce resource utilization, the overall required resources remain a heavy burden on edge devices. Instead, Retrieval-Augmented Generation (RAG), a resource-efficient LLM learning method, can improve the quality of the LLM-generated content without updating model parameters. However, the RAG-based LLM may involve repetitive searches on the profile data in every user-LLM interaction. This search can lead to significant latency along with the accumulation of user data. Conventional efforts to decrease latency result in restricting the size of saved user data, thus reducing the scalability of RAG as user data continuously grows. It remains an open question: how to free RAG from the constraints of latency and scalability on edge devices? In this paper, we propose a novel framework to accelerate RAG via Computing-in-Memory (CiM) architectures. It accelerates matrix multiplications by performing in-situ computation inside the memory while avoiding the expensive data transfer between the computing unit and memory. Our framework, Robust CiM-backed RAG (RoCR), utilizing a novel contrastive learning-based training method and noise-aware training, can enable RAG to efficiently search profile data with CiM. To the best of our knowledge, this is the first work utilizing CiM to accelerate RAG. Ruiyang Qin, Zheyu Yan, Dewen Zeng, Zhenge Jia, Dancheng Liu, Ahmed Abbasi, Zhi Zheng 0002, Ningyuan Cao, Kai Ni 0004, Jinjun Xiong, Yiyu Shi 0001 |
ICCAD | 3 |
| 2024 | Achieving Fairness Through Channel Pruning for Dermatological Disease Diagnosis
Qingpeng Kong, Ching-Hao Chiu, Dewen Zeng, Tsung-Yi Ho, Jingtong Hu, Yiyu Shi 0001 |
MICCAI (10) | 3 |
| 2023 | Synthetic Data Can Also Teach: Synthesizing Effective Data for Unsupervised Visual Representation LearningabstractContrastive learning (CL), a self-supervised learning approach, can effectively learn visual representations from unlabeled data. Given the CL training data, generative models can be trained to generate synthetic data to supplement the real data. Using both synthetic and real data for CL training has the potential to improve the quality of learned representations. However, synthetic data usually has lower quality than real data, and using synthetic data may not improve CL compared with using real data. To tackle this problem, we propose a data generation framework with two methods to improve CL training by joint sample generation and contrastive learning. The first approach generates hard samples for the main model. The generator is jointly learned with the main model to dynamically customize hard samples based on the training state of the main model. Besides, a pair of data generators are proposed to generate similar but distinct samples as positive pairs. In joint learning, the hardness of a positive pair is progressively increased by decreasing their similarity. Experimental results on multiple datasets show superior accuracy and data efficiency of the proposed data generation methods applied to CL. For example, about 4.0%, 3.5%, and 2.6% accuracy improvements for linear classification are observed on ImageNet-100, CIFAR-100, and CIFAR-10, respectively. Besides, up to 2× data efficiency for linear classification and up to 5× data efficiency for transfer learning are achieved. Yawen Wu, Zhepeng Wang 0001, Dewen Zeng, Yiyu Shi 0001, Jingtong Hu |
AAAI | 3 |
| 2023 | Additional Positive Enables Better Representation Learning for Medical Images
Dewen Zeng, Yawen Wu, Xinrong Hu, Xiaowei Xu 0004, Jingtong Hu, Yiyu Shi 0001 |
MICCAI (1) | 1 |
| 2023 | Self-Supervised On-Device Federated Learning From Unlabeled StreamsabstractThe ubiquity of edge devices has led to a growing amount of unlabeled data produced at the edge. Deep learning models deployed on edge devices are required to learn from these unlabeled data to continuously improve accuracy. Self-supervised representation learning has achieved promising performances using centralized unlabeled data. However, the increasing awareness of privacy protection limits centralizing the distributed unlabeled image data on edge devices. While federated learning has been widely adopted to enable distributed machine learning with privacy preservation, without a data selection method to efficiently select streaming data, the traditional federated learning framework fails to handle these huge amounts of decentralized unlabeled data with limited storage resources on edge. To address these challenges, we propose a self-supervised on-device federated learning framework with coreset selection, which we call SOFed, to automatically select a coreset that consists of the most representative samples into the replay buffer on each device. It preserves data privacy as each client does not share raw data while learning good visual representations. Experiments demonstrate the effectiveness and significance of the proposed method in visual representation learning. Jiahe Shi, Yawen Wu, Dewen Zeng, Jun Tao 0001, Jingtong Hu, Yiyu Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Decentralized Unsupervised Learning of Visual RepresentationsabstractCollaborative learning enables distributed clients to learn a shared model for prediction while keeping the training data local on each client. However, existing collaborative learning methods require fully-labeled data for training, which is inconvenient or sometimes infeasible to obtain due to the high labeling cost and the requirement of expertise. The lack of labels makes collaborative learning impractical in many realistic settings. Self-supervised learning can address this challenge by learning from unlabeled data. Contrastive learning (CL), a self-supervised learning approach, can effectively learn visual representations from unlabeled image data. However, the distributed data collected on clients are usually not independent and identically distributed (non-IID) among clients, and each client may only have few classes of data, which degrades the performance of CL and learned representations. To tackle this problem, we propose a collaborative contrastive learning framework consisting of two approaches: feature fusion and neighborhood matching, by which a unified feature space among clients is learned for better data representations. Feature fusion provides remote features as accurate contrastive information to each client for better local learning. Neighborhood matching further aligns each client’s local features to the remote features such that well-clustered features among clients can be learned. Extensive experiments show the effectiveness of the proposed framework. It outperforms other methods by 11% on IID data and matches the performance of centralized learning. Yawen Wu, Zhepeng Wang 0001, Dewen Zeng, Meng Li 0004, Yiyu Shi 0001, Jingtong Hu |
IJCAI | 3 |
| 2022 | FairPrune: Achieving Fairness Through Pruning for Dermatological Disease Diagnosis
Yawen Wu, Dewen Zeng, Xiaowei Xu 0004, Yiyu Shi 0001, Jingtong Hu |
MICCAI (1) | 2 |
| 2022 | Distributed contrastive learning for medical image segmentation
Yawen Wu, Dewen Zeng, Zhepeng Wang 0001, Yiyu Shi 0001, Jingtong Hu |
Medical Image Anal. | 2 |
| 2021 | Enabling On-Device Self-Supervised Contrastive Learning with Selective Data ContrastabstractAfter a model is deployed on edge devices, it is desirable for these devices to learn from unlabeled data to continuously improve accuracy. Contrastive learning has demonstrated its great potential in learning from unlabeled data. However, the online input data are usually none independent and identically distributed (non-iid) and edge devices’ storages are usually too limited to store enough representative data from different data classes. We propose a framework to automatically select the most representative data from the unlabeled input stream, which only requires a small data buffer for dynamic learning. Experiments show that accuracy and learning speed are greatly improved. Yawen Wu, Zhepeng Wang 0001, Dewen Zeng, Yiyu Shi 0001, Jingtong Hu |
DAC | 3 |
| 2021 | Invited: Hardware-aware Real-time Myocardial Segmentation Quality Control in Contrast EchocardiographyabstractAutomatic myocardial segmentation of contrast echocardio-graphy has shown great potential in the quantification of myocardial perfusion parameters. Segmentation quality control is an important step to ensure the accuracy of segmentation results for quality research as well as its clinical application. Usually, the segmentation quality control happens after the data acquisition. At the data acquisition time, the operator could not know the quality of the segmentation results. On-the-fly segmentation quality control could help the operator to adjust the ultrasound probe or retake data if the quality is unsatisfied, which can greatly reduce the effort of time-consuming manual correction. However, it is infeasible to deploy state-of-the-art DNN-based models because the segmentation module and quality control module must fit in the limited hardware resource on the ultrasound machine while satisfying strict latency constraints. In this paper, we propose a hardware-aware neural architecture search framework for automatic myocardial segmentation and quality control of contrast echocardiography. We explicitly incorporate the hardware latency as a regularization term into the loss function during training. The proposed method searches the best neural network architecture for the segmentation module and quality prediction module with strict latency. Dewen Zeng, Yukun Ding, Haiyun Yuan, Meiping Huang, Xiaowei Xu 0004, Jian Zhuang, Jingtong Hu, Yiyu Shi 0001 |
DAC | 1 |
| 2021 | Federated Contrastive Learning for Dermatological Disease Diagnosis via On-device Learning (Invited Paper)abstractDeep learning models have been deployed in an increasing number of edge and mobile devices to provide healthcare. These models rely on training with a tremendous amount of labeled data to achieve high accuracy. However, for medical applications such as dermatological disease diagnosis, the private data collected by mobile dermatology assistants exist on distributed mobile devices of patients, and each device only has a limited amount of data. Directly learning from limited data greatly deteriorates the performance of learned models. Federated learning (FL) can train models by using data distributed on devices while keeping the data local for privacy. Existing works on FL assume all the data have ground-truth labels. However, medical data often comes without any accompanying labels since labeling requires expertise and results in prohibitively high labor costs. The recently developed self-supervised learning approach, contrastive learning (CL), can leverage the unlabeled data to pre-train a model for learning data representations, after which the learned model can be fine-tuned on limited labeled data to perform dermatological disease diagnosis. However, simply combining CL with FL as federated contrastive learning (FCL) will result in ineffective learning since CL requires diverse data for accurate learning but each device in FL only has limited data diversity. In this work, we propose an on-device FCL framework for dermatological disease diagnosis with limited labels. Features are shared among devices in the FCL pre-training process to provide diverse and accurate contrastive information without sharing raw data for privacy. After that, the pre-trained model is fine-tuned with local labeled data independently on each device or collaboratively with supervised federated learning on all devices. Experiments on dermatological disease datasets show that the proposed framework effectively improves the recall and precision of dermatological disease diagnosis compared with state-of-the-art methods. Yawen Wu, Dewen Zeng, Zhepeng Wang 0001, Yi Sheng 0001, Lei Yang 0018, Alaina J. James, Yiyu Shi 0001, Jingtong Hu |
ICCAD | 2 |
| 2021 | Contrastive Learning with Temporal Correlated Medical Images: A Case Study using Lung Segmentation in Chest X-Rays (Invited Paper)abstractContrastive learning has been proved to be a promising technique for image-level representation learning from unlabeled data. Many existing works have demonstrated improved results by applying contrastive learning in classification and object detection tasks for either natural images or medical images. However, its application to medical image segmentation tasks has been limited. In this work, we use lung segmentation in chest X-rays as a case study and propose a contrastive learning framework with temporal correlated medical images, named CL-TCI, to learn superior encoders for initializing the segmentation network. We adapt CL-TCI from two state-of-the-art contrastive learning methods-MoCo and SimCLR. Experiment results on three chest X-ray datasets show that under two different segmentation backbones, U-Net and Deeplab-V3, CL-TCI can outperform all baselines that do not incorporate any temporal correlation in both semi-supervised learning setting and transfer learning setting with limited annotation. This suggests that information among temporal correlated medical images can indeed improve contrastive learning performance. Between the two variations of CL-TCI, CL-TCI adapted from MoCo outperforms CL-TCI adapted from SimCLR in most settings, indicating that more contrastive samples can benefit the learning process and help the network learn high-quality representations. Code is available here https://github.com/dewenzeng/CL-TCI. Dewen Zeng, John N. Kheir, Yiyu Shi 0001 |
ICCAD | 1 |
| 2021 | Towards Efficient Human-Machine Collaboration: Real-Time Correction Effort Prediction for Ultrasound Data Acquisition
Yukun Ding, Dewen Zeng, Hongwen Fei, Haiyun Yuan, Meiping Huang, Jian Zhuang, Yiyu Shi 0001 |
MICCAI (1) | 2 |
| 2021 | Semi-supervised Contrastive Learning for Label-Efficient Medical Image Segmentation
Xinrong Hu, Dewen Zeng, Xiaowei Xu 0004, Yiyu Shi 0001 |
MICCAI (2) | 2 |
| 2021 | Federated Contrastive Learning for Volumetric Medical Image Segmentation
Yawen Wu, Dewen Zeng, Zhepeng Wang 0001, Yiyu Shi 0001, Jingtong Hu |
MICCAI (3) | 2 |
| 2021 | Positional Contrastive Learning for Volumetric Medical Image Segmentation
Dewen Zeng, Yawen Wu, Xinrong Hu, Xiaowei Xu 0004, Haiyun Yuan, Meiping Huang, Jian Zhuang, Jingtong Hu, Yiyu Shi 0001 |
MICCAI (2) | 1 |
| 2020 | Towards Cardiac Intervention Assistance: Hardware-aware Neural Architecture Exploration for Real-Time 3D Cardiac Cine MRI SegmentationabstractReal-time cardiac magnetic resonance imaging (MRI) plays an increasingly important role in guiding various cardiac interventions. In order to provide better visual assistance, the cine MRI frames need to be segmented on-the-fly to avoid noticeable visual lag. In addition, considering reliability and patient data privacy, the computation is preferably done on local hardware. State-of-the-art MRI segmentation methods mostly focus on accuracy only, and can hardly be adopted for real-time application or on local hardware. In this work, we present the first hardware-aware multi-scale neural architecture search (NAS) framework for real-time 3D cardiac cine MRI segmentation. The proposed framework incorporates a latency regularization term into the loss function to handle realtime constraints, with the consideration of underlying hardware. In addition, the formulation is fully differentiable with respect to the architecture parameters, so that stochastic gradient descent (SGD) can be used for optimization to reduce the computation cost while maintaining optimization quality. Experimental results on ACDC MICCAI 2017 dataset demonstrate that our hardware-aware multi-scale NAS framework can reduce the latency by up to 3.5× and satisfy the real-time constraints, while still achieving competitive segmentation accuracy, compared with the state-of-the-art NAS segmentation framework. Dewen Zeng, Weiwen Jiang, Xiaowei Xu 0004, Haiyun Yuan, Meiping Huang, Jian Zhuang, Jingtong Hu, Yiyu Shi 0001 |
ICCAD | 1 |
| 2019 | MDA: A Reconfigurable Memristor-Based Distance Accelerator for Time Series Mining on Data CentersabstractThe rapid development of Internet-of-Things is yielding a huge volume of time series data, the real-time mining of which becomes a major load for data centers. The computation bottleneck in time series data mining is distance function, which is the fundamental element of many high data mining tasks. Recently various software optimization and hardware acceleration techniques have been proposed to tackle the challenge. However, each of these techniques is only designed or optimized for a specific distance function. To address this problem, in this paper we propose MDA, a high-throughput reconfigurable memristor-based distance accelerator for real-time and energy-efficient data mining with time series in data centers. Common circuit structure is extracted for efficiency, and the circuit can be configured to any specific distance functions. Particularly, we adopt the emerging device memristor for the design of MDA. Comprehensive experiments are presented with public available datasets to evaluate the performance of the proposed MDA. Experimental results show that compared with existing works, MDA has achieved a speedup of 3.5×-376× on performance and an improvement of 1-3 orders of magnitude on energy efficiency with little accuracy loss. Xiaowei Xu 0004, Feng Lin 0004, Wenyao Xu, Xin-Wei Yao 0001, Yiyu Shi 0001, Dewen Zeng, Yu Hu 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2017 | An Efficient Memristor-based Distance Accelerator for Time Series Data Mining on Data CentersabstractThe rapid development of Internet-of-Things (IoT) is yielding a huge volume of time series data, the real-time mining of which becomes a major load for data centers. The computation bottleneck in time series data mining is the distance function, which has been tackled by various software optimization and hardware acceleration techniques recently. However, each of these techniques is only designed or optimized for a specific distance function. To address this problem, in this paper we propose an efficient and reconfigurable memristor-based distance accelerator for real-time and energy-efficient data mining with time series on data centers. Common circuit structure is extracted to save chip areas, and the circuit can be configured to any specific distance functions. Experimental results show that compared with existing works, our work has achieved a speedup of 3.5x-376x on performance and an improvement of 1-3 orders of magnitude on energy efficiency. Xiaowei Xu 0004, Dewen Zeng, Wenyao Xu, Yiyu Shi 0001, Yu Hu 0002 |
DAC | 2 |