EDBT 2026 Demo / reviewers in the wild / expert
Wei-Ming Chen
dblp:63/5986
· DBLP profile ↗
40ranked-venue papers
15as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021Computer networks · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A 16-Gb/s Baud-Rate CDR Circuit With One-Tap Speculative DFE and Wide Frequency Capture RangeabstractA 16-Gb/s baud-rate clock and data recovery (CDR) circuit with a one-tap decision-feedback equalizer (DFE) and a wide frequency capture range (FCR) is presented. The proposed asymmetrical pattern-based phase detectors are used to achieve a wide FCR. This quarter-rate CDR circuit is fabricated in 40-nm CMOS technology and the active area is 0.1094 mm2. For a 16 Gb/s PRBS of 27–1, the power of the CDR circuit is 38.4 mW and its calculated energy efficiency is 2.4 pJ/b. The measured FCR is 40.6%. Po-Yuan Chou, Wei-Ming Chen, Shen-Iuan Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | PockEngine: Sparse and Efficient Fine-tuning in a PocketabstractOn-device learning and efficient fine-tuning enable continuous and privacy-preserving customization (e.g., locally fine-tuning large language models on personalized data). However, existing training frameworks are designed for cloud servers with powerful accelerators (e.g., GPUs, TPUs) and lack the optimizations for learning on the edge, which faces challenges of resource limitations and edge hardware diversity. We introduce PockEngine: a tiny, sparse and efficient engine to enable fine-tuning on various edge devices. PockEngine supports sparse backpropagation: it prunes the backward graph and sparsely updates the model with measured memory saving and latency reduction while maintaining the model quality. Secondly, PockEngine is compilation first: the entire training graph (including forward, backward and optimization steps) is derived at compile-time, which reduces the runtime overhead and brings opportunities for graph transformations. PockEngine also integrates a rich set of training graph optimizations, thus can further accelerate the training cost, including operator reordering and backend switching. PockEngine supports diverse applications, frontends and hardware backends: it flexibly compiles and tunes models defined in PyTorch/TensorFlow/Jax and deploys binaries to mobile CPU/GPU/DSPs. We evaluated PockEngine on both vision models and large language models. PockEngine achieves up to 15 × speedup over off-the-shelf TensorFlow (Raspberry Pi), 5.6 × memory saving back-propagation (Jetson AGX Orin). Remarkably, PockEngine enables fine-tuning LLaMav2-7B on NVIDIA Jetson AGX Orin at 550 tokens/s, 7.9 × faster than the PyTorch. Ligeng Zhu, Lanxiang Hu, Ji Lin 0002, Wei-Ming Chen, Wei-Chen Wang 0002, Chuang Gan 0001, Song Han 0003 |
MICRO | 4 |
| 2022 | Lite Pose: Efficient Architecture Design for 2D Human Pose EstimationabstractPose estimation plays a critical role in human-centered vision applications. However, it is difficult to deploy state-of-the-art HRNet-based pose estimation models on resource-constrained edge devices due to the high computational cost (more than 150 GMACs per frame). In this paper, we study efficient architecture design for real-time multi-person pose estimation on edge. We reveal that HRNet's high-resolution branches are redundant for models at the low-computation region via our gradual shrinking experiments. Removing them improves both efficiency and performance. Inspired by this finding, we design LitePose, an efficient single-branch architecture for pose estimation, and introduce two simple approaches to enhance the capacity of LitePose, including fusion deconv head and large kernel conv. On mobile platforms, LitePose reduces the latency by up to$5.0\times$without sacrificing performance, compared with prior state-of-the-art efficient pose estimation models, pushing the frontier of real-time multi-person pose estimation on edge. Our code and pretrained models are released at https://github.com/mit-han-lab/litepose. Han Cai, Wei-Ming Chen, Song Han 0003 |
CVPR | 4 |
| 2022 | On-Device Training Under 256KB MemoryabstractOn-device training enables the model to adapt to new data collected from the sensors by fine-tuning a pre-trained model. Users can benefit from customized AI models without having to transfer the data to the cloud, protecting the privacy. However, the training memory consumption is prohibitive for IoT devices that have tiny memory resources. We propose an algorithm-system co-design framework to make on-device training possible with only 256KB of memory. On-device training faces two unique challenges: (1) the quantized graphs of neural networks are hard to optimize due to low bit-precision and the lack of normalization; (2) the limited hardware resource (memory and computation) does not allow full backpropagation. To cope with the optimization difficulty, we propose Quantization- Aware Scaling to calibrate the gradient scales and stabilize 8-bit quantized training. To reduce the memory footprint, we propose Sparse Update to skip the gradient computation of less important layers and sub-tensors. The algorithm innovation is implemented by a lightweight training system, Tiny Training Engine, which prunes the backward computation graph to support sparse updates and offload the runtime auto-differentiation to compile time. Our framework is the first practical solution for on-device transfer learning of visual recognition on tiny IoT devices (e.g., a microcontroller with only 256KB SRAM), using less than 1/1000 of the memory of PyTorch and TensorFlow while matching the accuracy. Our study enables IoT devices not only to perform inference but also to continuously adapt to new data for on-device lifelong learning. A video demo can be found here: https://youtu.be/XaDCO8YtmBw. Ji Lin 0002, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang 0002, Chuang Gan 0001, Song Han 0003 |
NeurIPS | 3 |
| 2022 | Intermittent-Aware Distributed Concurrency ControlabstractInternet of Things (IoT) devices are gradually adopting battery-less, energy harvesting solutions, thereby driving the development of an intermittent computing paradigm to accumulate computation progress across multiple power cycles. While many attempts have been made to enable standalone intermittent systems, little attention has focused on IoT networks formed by intermittent devices. We observe that the computation progress improved by distributed task concurrency in an intermittent network can be significantly offset by data unavailability due to frequent system failures. This article presents an intermittent-aware distributed concurrency control protocol which leverages existing data copies inherently created in the network to improve the computation progress of concurrently executed tasks. In particular, we propose a borrowing-based data management method to increase data availability and an intermittent two-phase commit procedure incorporated with distributed backward validation to ensure data consistency in the network. The proposed protocol was integrated into a FreeRTOS-extended intermittent operating system running on Texas Instruments devices. Experimental results show that the computation progress can be significantly improved, and this improvement is more apparent under weaker power, where more devices will remain offline for longer duration. Wei-Che Tsai, Wei-Ming Chen, Tei-Wei Kuo, Pi-Cheng Hsiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Memory-efficient Patch-based Inference for Tiny Deep LearningabstractTiny deep learning on microcontroller units (MCUs) is challenging due to the limited memory size. We find that the memory bottleneck is due to the imbalanced memory distribution in convolutional neural network (CNN) designs: the first several blocks have an order of magnitude larger memory usage than the rest of the network. To alleviate this issue, we propose a generic patch-by-patch inference scheduling, which operates only on a small spatial region of the feature map and significantly cuts down the peak memory. However, naive implementation brings overlapping patches and computation overhead. We further propose receptive field redistribution to shift the receptive field and FLOPs to the later stage and reduce the computation overhead. Manually redistributing the receptive field is difficult. We automate the process with neural architecture search to jointly optimize the neural architecture and inference scheduling, leading to MCUNetV2. Patch-based inference effectively reduces the peak memory usage of existing networks by4-8×. Co-designed with neural networks, MCUNetV2 sets a record ImageNetaccuracy on MCU (71.8%) and achieves >90% accuracy on the visual wake words dataset under only 32kB SRAM. MCUNetV2 also unblocks object detection on tiny devices, achieving 16.9% higher mAP on Pascal VOC compared to the state-of-the-art result. Our study largely addressed the memory bottleneck in tinyML and paved the way for various vision applications beyond image classification. Ji Lin 0002, Wei-Ming Chen, Han Cai, Chuang Gan 0001, Song Han 0003 |
NeurIPS | 2 |
| 2021 | A 10.4-16-Gb/s Reference-Less Baud-Rate Digital CDR With One-Tap DFE Using a Wide-Range FDabstractA 10.4-16-Gb/s reference-less and baud-rate clock and data recovery (CDR) circuit with a one-tap speculative decision feedback equalizer (DFE) is presented. The quarter-rate CDR circuit uses a pattern-based phase detector (PD) and the proposed FD. This wide-range FD is composed of a coarse FD and a fine FD (FFD) which share the same front-end comparators with the PD and the DFE. Thus, no extra comparators are required. In addition, by monitoring the drift direction of five samples on the five-bit data patterns, the FFD performs an error-free operation within a frequency error of 7.7%. Therefore, the CDR has a robust FD-to-PD transition. By using the proposed FD, this CDR circuit not only achieves a wide frequency capture range of 43%, but also has a short frequency settling time of$680~\mu \text{s}$. This CDR circuit is fabricated in 40-nm CMOS technology and occupies an active area of 0.1004 mm2. The total power of the receiver is 39.9mW at 16 Gb/s, and the calculated energy efficiency is 2.49pJ/b. Wei-Ming Chen, Yun-Sheng Yao, Shen-Iuan Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Heterogeneity-aware Multicore Synchronization for Intermittent SystemsabstractIntermittent systems enable batteryless devices to operate through energy harvesting by leveraging the complementary characteristics of volatile (VM) and non-volatile memory (NVM). Unfortunately, alternate and frequent accesses to heterogeneous memories for accumulative execution across power cycles can significantly hinder computation progress. The progress impediment is mainly due to more CPU time being wasted for slow NVM accesses than for fast VM accesses. This paper explores how to leverage heterogeneous cores to mitigate the progress impediment caused by heterogeneous memories. In particular, a delegable and adaptive synchronization protocol is proposed to allow memory accesses to be delegated between cores and to dynamically adapt to diverse memory access latency. Moreover, our design guarantees task serializability across multiple cores and maintains data consistency despite frequent power failures. We integrated our design into FreeRTOS running on a Cypress device featuring heterogeneous dual cores and hybrid memories. Experimental results show that, compared to recent approaches that assume single-core intermittent systems, our design can improve computation progress at least 1.8x and even up to 33.9x by leveraging core heterogeneity. Wei-Ming Chen, Tei-Wei Kuo, Pi-Cheng Hsiu |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2020 | MCUNet: Tiny Deep Learning on IoT DevicesabstractMachine learning on tiny IoT devices based on microcontroller units (MCU) is appealing but challenging: the memory of microcontrollers is 2-3 orders of magnitude smaller even than mobile phones. We propose MCUNet, a framework that jointly designs the efficient neural architecture (TinyNAS) and the lightweight inference engine (TinyEngine), enabling ImageNet-scale inference on microcontrollers. TinyNAS adopts a two-stage neural architecture search approach that first optimizes the search space to fit the resource constraints, then specializes the network architecture in the optimized search space. TinyNAS can automatically handle diverse constraints (i.e. device, latency, energy, memory) under low search costs. TinyNAS is co-designed with TinyEngine, a memory-efficient inference library to expand the search space and fit a larger model. TinyEngine adapts the memory scheduling according to the overall network topology rather than layer-wise optimization, reducing the memory usage by 3.4×, and accelerating the inference by 1.7-3.3× compared to TF-Lite Micro [3] and CMSIS-NN [28]. MCUNet is the first to achieves >70% ImageNet top1 accuracy on an off-the-shelf commercial microcontroller, using 3.5× less SRAM and 5.7× less Flash compared to quantized MobileNetV2 and ResNet-18. On visual&audio wake words tasks, MCUNet achieves state-of-the-art accuracy and runs 2.4-3.4× faster than Mo- bileNetV2 and ProxylessNAS-based solutions with 3.7-4.1× smaller peak SRAM. Our study suggests that the era of always-on tiny machine learning on IoT devices has arrived. Ji Lin 0002, Wei-Ming Chen, Yujun Lin 0001, John Cohn, Chuang Gan 0001, Song Han 0003 |
NeurIPS | 2 |
| 2020 | Enabling Failure-Resilient Intermittent Systems Without Runtime CheckpointingabstractSelf-powered intermittent systems typically adopt runtime checkpointing as a means to accumulate computation progress across power cycles and recover system status from power failures. However, existing approaches based on the checkpointing paradigm normally require system suspension and/or logging at runtime. This article presents a design which overcomes the drawbacks of checkpointing-based approaches, to enable failure-resilient intermittent systems. Our design allows accumulative execution and instant system recovery under frequent power failures while enforcing the serializability of concurrent task execution to improve computation progress and ensuring data consistency without system suspension during runtime, by leveraging the characteristics of data accessed in hybrid memory. We integrated the design into FreeRTOS running on a Texas Instruments device. The experimental results show that our design can still accumulate progress when the power source is too weak for checkpointing-based approaches to advance, and significantly improves the computation progress while reducing the recovery time. Wei-Ming Chen, Tei-Wei Kuo, Pi-Cheng Hsiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | LSIM: Ultra Lightweight Similarity Measurement for Mobile Graphics ApplicationsabstractPerceptual similarity measurement allows mobile applications to eliminate unnecessary computations without compromising visual experience. Existing pixel-wise measures incur significant overhead with increasing display resolutions and frame rates. This paper presents an ultra lightweight similarity measure called LSIM, which assesses the similarity between frames based on the transformation matrices of graphics objects. To evaluate its efficacy, we integrate LSIM into the Open Graphics Library and conduct experiments on an Android smartphone with various mobile 3D games. The results show that LSIM is highly correlated with the most widely used pixel-wise measure SSIM, yet three to five orders of magnitude faster. We also apply LSIM to a CPU-GPU governor to suppress the rendering of similar frames, thereby further reducing computation energy consumption by up to 27.3% while maintaining satisfactory visual quality. Yu-Chuan Chang, Wei-Ming Chen, Pi-Cheng Hsiu, Yen-Yu Lin, Tei-Wei Kuo |
DAC | 2 |
| 2019 | Enabling Failure-resilient Intermittently-powered Systems Without Runtime CheckpointingabstractSelf-powered intermittent systems enable accumulative execution in unstable power environments, where checkpointing is often adopted as a means to achieve data consistency and system recovery under power failures. However, existing approaches based on the checkpointing paradigm normally require system suspension and/or logging at runtime. This paper presents a design which enables failure-resilient intermittently-powered systems without runtime checkpointing. Our design enforces the consistency and serializability of concurrent task execution while maximizing computation progress, as well as allows instant system recovery after power resumption, by leveraging the characteristics of data accessed in hybrid memory. We integrated the design into FreeRTOS running on a Texas Instruments device. Experimental results show that our design achieves up to 11.8 times the computation progress achieved by checkpointing-based approaches, while reducing the recovery time by nearly 90%. Wei-Ming Chen, Pi-Cheng Hsiu, Tei-Wei Kuo |
DAC | 1 |
| 2019 | Multiversion Concurrency Control on Intermittent SystemsabstractConcurrency control allows multiple tasks that share data objects to be concurrently executed in a serializable order, thus significantly improving computation progress. However, to accumulate forward progress on energy-harvesting intermittent systems while achieving data consistency across power cycles, existing approaches based on the checkpointing paradigm typically require system suspension at runtime. The runtime overheads incurred by suspension will be more manifest when more tasks are suspended and resumed during checkpointing, offsetting the computation progress improved by concurrent task execution. This paper presents a multiversion concurrency control design, which enables concurrent task execution without system suspension during checkpointing, while maintaining the serializability of task execution and ensuring data consistency after system recovery. We integrated our design into FreeRTOS running on a Texas Instruments device. Experimental results show that, at the very best, our design can double computation progress by reducing the runtime overheads incurred by system checkpointing, especially when tasks are executed with high concurrency. Wei-Ming Chen, Pi-Cheng Hsiu, Tei-Wei Kuo |
ICCAD | 1 |
| 2019 | A virtual tutor movement learning system in eLearning
Hsin-Hung Chiang, Wei-Ming Chen, Han-Chieh Chao, De-Li Tsai |
Multim. Tools Appl. | 2 |
| 2019 | A User-Centric CPU-GPU Governing Framework for 3-D Mobile GamesabstractGraphics-intensive mobile games place different and varying levels of demand on the associated central processing units (CPUs) and graphics processing units (GPUs). In contrast to the workload variability that characterizes games, the current design of the energy governor employed by mobile systems appears to be outdated. In this paper, we review the energy-saving mechanism implemented in an Android system coupled with graphics-intensive gaming workloads from three perspectives: 1) user perception; 2) application status; and 3) the interplay between the CPU and GPU. We observe that there are information gaps in the current system, which may result in unnecessary energy wastage. To resolve the problem, we propose an online user-centric CPU-GPU governing framework. To bridge the identified information gaps, we classify rendered game frames into redundant/changing frames to satisfy user demand, categorize an application into GPU sensitive/insensitive phases to understand the application's demand, and determine the frequency scaling intents of the CPU and GPU to capture processor demand. In response to the measured demand, we employ a required workload estimator, a unified policy selector, and a frequency-scaling intent communicator in the framework to save energy. The proposed framework was implemented on an LG Nexus 5X smartphone, and extensive experiments with real-world 3-D gaming applications were conducted. According to the experiment results, for an application which is low interactive and infrequent phase changing, the proposed framework can, respectively, reduce energy consumption by 25.3% and 39% compared with our previous work and Android governors while maintaining user experience. Wei-Ming Chen, Sheng-Wei Cheng, Pi-Cheng Hsiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Differentiated handling of physical scenes and virtual objects for mobile augmented realityabstractMobile devices running augmented reality applications consume considerable energy for graphics-intensive workloads. This paper presents a scheme for the differentiated handling of camera-captured physical scenes and computer-generated virtual objects according to different perceptual quality metrics. We propose online algorithms and their real-time implementations to reduce energy consumption through dynamic frame rate adaptation while maintaining the visual quality required for augmented reality applications. To evaluate system efficacy, we integrate our scheme into Android and conduct extensive experiments on a commercial smartphone with various application scenarios. The results show that the proposed scheme can achieve energy savings of up to 39.1% in comparison to the native graphics system in Android while maintaining satisfactory visual quality. Chih-Hsuan Yen, Wei-Ming Chen, Pi-Cheng Hsiu, Tei-Wei Kuo |
ICCAD | 2 |
| 2017 | Design considerations and clinical applications of closed-loop neural disorder control SoCsabstractThis paper presents the closed-loop neural disorder control concept and some design considerations. Two architectures of closed-loop neuromodulation for Parkinson's disease and epileptic seizure are proposed. One is a closed-loop deep brain stimulator, which meets the IEC 60601-1 standard. The other one is an implantable SoC for epileptic seizure control, which is verified by animal experiment. Chung-Yu Wu, Cheng-Hsiang Cheng, Yi-Huan Ou-Yang, Chiung-Ghu Chen, Wei-Ming Chen, Ming-Dou Ker, Chen-Yi Lee, Sheng-Fu Liang, Fu-Zen Shaw |
ASP-DAC | 5 |
| 2017 | Image watermark protection based on self-recovery images and sparse approximation
Hao-Chun Wang, Ing-Yi Chen, Wei-Ming Chen |
Multim. Tools Appl. | 3 |
| 2016 | The design of 8-channel CMOS area-efficient low-power current-mode analog front-end amplifier for EEG signal recordingabstractIn this paper, an 8-channel area-efficient low-power current-mode analog front-end amplifier (AFEA) is designed for EEG signal recording. The AFEA is composed of eight capacitive coupled transconductors (CCGMs), current-mode band-pass filters (CMBPFs), and programmable current-gain amplifiers (PCGAs) with a multiplexer (MUX), a transimpedance amplifier (TIA), and an offset current cancellation loop (OCCL). The AFEA employs CCGM with only 2pF input capacitance to eliminate the electrode dc offset (EDO). The current-mode topology is adopted in the design of CCGMs, CMBPF s, and PCGAs to reduce the power consumption. The shared OCCL is designed to eliminate the output offset of CCGM, CMBPF and PCGA. The AFEA is designed and fabricated in 180-nm CMOS technology and the core area occupies only 1mm2. The measured maximum gain is 82 dB. The measured input-referred noise is 3.34μVrms within the bandwidth of 0.5-100 Hz. The measured maximum power consumption is 7.85 μW per channel under power supply of 1.2 V. The fabricated AFEA is applied to record the human EEG signal successfully. Ya-Syuan Sung, Wei-Ming Chen, Chung-Yu Wu |
ISCAS | 2 |
| 2016 | Value-Based Task Scheduling for Nonvolatile Processor-Based Embedded DevicesabstractEnergy harvesting wearable sensor nodes offer low maintenance and high mobility, but suffer from unstable power supply and insufficient energy. Featuring low standby power and instant backup and restore operations, nonvolatile processors have emerged as one of the most promising technologies for the improvement of energy harvesting wearable devices. However, device quality of service (QoS) is still limited by lack of sufficient energy. To address this problem, this paper attempts to maximize QoS by optimizing task scheduling for DVFS-enabled nonvolatile processor-based embedded devices. First, we model the task scheduling problem as an optimization problem to maximize the total value (which represents QoS) given available harvested energy and time. We then prove the problem to be NP-hard. In addition, we propose a pseudopolynomial-time optimal algorithm based on dynamic programming, as well as an approximation algorithm that allows for a trade-off between running time and total value. We evaluate algorithm performance on an ultra-lowpower platform produced by Texas Instruments, and conduct extensive simulations with different system models and task sets. The results demonstrate that the platform can execute optimal sequences derived by the dynamic-programming algorithm, with differences between simulation results and real traces of less than 1%. Also, our approximate algorithm executes approximately 10 5 times faster than the dynamic-programming algorithm at a value degradation cost of less than 3%. Wei-Ming Chen, Taisheng Cheng, Pi-Cheng Hsiu, Tei-Wei Kuo |
RTSS | 1 |
| 2016 | User-Centric Scheduling and Governing on Mobile Devices with big.LITTLE ProcessorsabstractMobile applications will become progressively more complicated and diverse. Heterogeneous computing architectures like big.LITTLE are a hardware solution that allows mobile devices to combine computing performance and energy efficiency. However, software solutions that conform to the paradigm of conventional fair scheduling and governing are not applicable to mobile systems, thereby degrading user experience or reducing energy efficiency. In this article, we exploit the concept of application sensitivity, which reflects the user’s attention on each application, and devise a user-centric scheduler and governor that allocate computing resources to applications according to their sensitivity. Furthermore, we integrate our design into the Android operating system. The results of experiments conducted on a commercial big.LITTLE smartphone with real-world mobile apps demonstrate that the proposed design can achieve significant gains in energy efficiency while improving the quality of user experience. Pi-Cheng Hsiu, Po-Hsien Tseng, Wei-Ming Chen, Chin-Chiang Pan, Tei-Wei Kuo |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | Identifying smallest unique subgraphs in a heterogeneous social networkabstractThis paper proposes to study a novel problem, discovering a Smallest Unique Subgraph (SUS) for any node of interest specified by user in a heterogeneous social network. The rationale of the SUS problem lies in how a person is different from any others in a social network, and how to represent the identity of a person using her surrounding relational structure in a social network. To deal with the proposed SUS problem, we develop an Ego-Graph Heuristic (EGH) method to efficiently solve the SUS problem in an approximated manner. EGH intelligently examine whether one graph is not isomorphic to the other, instead of using the conventional subgraph isomorphism test. We also prove SUS is a NP-complete problem through doing a reduction from Minimum Vertex Cover (MVC) in a homogeneous tree structure. Experimental results conducted on a real-world movie heterogeneous social network data show both the promising efficiency and compactness of our method. Yen-Kai Wang, Wei-Ming Chen, Cheng-Te Li, Shou-De Lin |
IEEE BigData | 2 |
| 2015 | A User-Centric CPU-GPU Governing Framework for 3D Games on Mobile DevicesabstractGraphics-intensive mobile games are becoming increasingly popular, but such applications place high demand on device CPUs and GPUs. The design of current mobile systems results in unnecessary energy waste due to lack of consideration of application phases and user attention (a “demand-level” gap) and because each processor administers power management autonomously (a “processor-level” gap). This paper proposes a user-centric CPU-GPU governing framework which aims to reduce energy consumption without significantly impacting the user experience. To bridge the gap at the demand level, we identify the user demand at runtime and accordingly determine appropriate governing policies for the respective processors. On the other hand, to bridge the gap at the processor level, the proposed framework interprets the frequency scaling intents of processors based on the observation of the CPU-GPU interaction and the processor status. We implemented our framework on a Samsung Galaxy S4, and conducted extensive experiments with real-world 3D gaming apps. Experimental results showed that, for an application being highly interactive and frequent phase changing, our proposed framework can reduce energy consumption by 45.1% compared with state-of-the-art policy without significantly impacting the user experience. Wei-Ming Chen, Sheng-Wei Cheng, Pi-Cheng Hsiu, Tei-Wei Kuo |
ICCAD | 1 |
| 2015 | An 8-channel power-efficient time-constant-enhanced analog front-end amplifier for neural signal acquisitionabstractIn this paper, an 8-channel low-power analog front-end amplifier (AFEA) with time-constant-enhanced topology is proposed for neural signal acquisition. The AFEA is composed of eight time-constant-enhanced amplifiers (TCEAs) and high-pass filters (HPFs), a multiplexed transconductor (MGM) and a transimpedance amplifier (TIA). The AFEA is designed and simulated in 65nm CMOS technology. The bandwidth of the AFEA is 7 kHz, and the maximum gain is 61.2 dB with power consumption of 1.77 μW per channel. The TCEA is designed to reduce the high-pass cut-off frequency to 0.16 Hz and achieve a high gain of 50.7 dB with only 4 pF input capacitance. The power consumption of TCEA is 1.08 μW under 1-V power supply. The noise efficiency factor (NEF) of the TCEA is 2.13 with 4.46 μVrmsinput-referred noise. Jung-Chen Chung, Wei-Ming Chen, Chung-Yu Wu |
ISCAS | 2 |
| 2014 | Backup routing firewall mechanism in P2P environment
Wei-Ming Chen, Hsin-Hung Chiang, Kai-Di Chang, Han-Chieh Chao, Jiann-Liang Chen |
Peer-to-Peer Netw. Appl. | 1 |
| 2013 | Effects of discrete hill climbing on model building forestimation of distribution algorithmsabstractHybridization of global and local searches is a well-known technique for optimization algorithms. Hill climbing is one of the local search methods. On estimation of distribution algorithms (EDAs), hill climbing strengthens the signals of dependencies on correlated variables and improves the quality of model building, which reduces the required population size and convergence time. However, hill climbing also consumes extra computational time. In this paper, analytical models are developed to investigate the effects of combining two different hill climbers with the extended compact genetic algorithm and the dependency structure matrix genetic algorithm. By using the one-max problem and the 5-bit non-overlapping trap problem as the test problems, the performances of different hill climbers are compared. Both analytical models and experiments reveal that the greedy hill climber reduces the number of function evaluations for EDAs to find the global optimum. Wei-Ming Chen, Chu-Yu Hsu, Tian-Li Yu 0001, Wei-Che Chien |
GECCO | 1 |
| 2013 | Design of test problems for discrete estimation of distribution algorithmsabstractTwo types of problem structures, overlapping and conflict structures, are challenging for the estimation of distribution algorithms (EDAs) to solve. To test the capabilities of different EDAs of dealing with overlapping and conflict structures, some test problems have been proposed. However, the upper-bound of the degree of overlap and the effect of conflict have not been fully investigated. This paper investigates how to properly define the degree of overlap and the degree of conflict to reflect the difficulties of problems for the EDAs. A new test problem is proposed with the new definitions of the degree of overlap and the degree of conflict. A framework for building the proposed problem is presented, and some model-building genetic algorithms are tested by the problem. This test problem can be applied to further researches on overlapping and conflict structures. Shih-Ming Wang, Jie-Wei Wu, Wei-Ming Chen, Tian-Li Yu 0001 |
GECCO | 3 |
| 2013 | A Low power 10bit 500kS/s delta-modulated SAR ADC (DMSAR ADC) for implantable medical devicesabstractAn architecture of SAR ADC called the delta-modulated SAR ADC (DMSAR ADC) is proposed and designed for medical device applications. In the proposed DMSAR ADC, only the voltage difference between two successive samples is resolved to reduce the conversion steps and decrease the power consumption per channel up to 66%. The experimental chip is implemented in 0.18μm CMOS technology. At the digital supply voltage of 1.35V and Vref=1.00V, the measured power consumption per channel is 1.38μW (2.71μW) and the measured SNDR is 56.68dB (53.89 dB) for Fin=10Hz (7kHz). The ENOB is 9.12b and FoM is 39.54fJ/step. Yuan-Fu Lyu, Chung-Yu Wu, Li-Chen Liu, Wei-Ming Chen |
ISCAS | 4 |
| 2013 | PhosphoChain: a novel algorithm to predict kinase and phosphatase networks from high-throughput expression dataabstractMOTIVATION: Protein phosphorylation is critical for regulating cellular activities by controlling protein activities, localization and turnover, and by transmitting information within cells through signaling networks. However, predictions of protein phosphorylation and signaling networks remain a significant challenge, lagging behind predictions of transcriptional regulatory networks into which they often feed. RESULTS: We developed PhosphoChain to predict kinases, phosphatases and chains of phosphorylation events in signaling networks by combining mRNA expression levels of regulators and targets with a motif detection algorithm and optional prior information. PhosphoChain correctly reconstructed ∼78% of the yeast mitogen-activated protein kinase pathway from publicly available data. When tested on yeast phosphoproteomic data from large-scale mass spectrometry experiments, PhosphoChain correctly identified ∼27% more phosphorylation sites than existing motif detection tools (NetPhosYeast and GPS2.0), and predictions of kinase-phosphatase interactions overlapped with ∼59% of known interactions present in yeast databases. PhosphoChain provides a valuable framework for predicting condition-specific phosphorylation events from high-throughput data. AVAILABILITY: PhosphoChain is implemented in Java and available at http://virgo.csie.ncku.edu.tw/PhosphoChain/ or http://aitchisonlab.com/PhosphoChain Wei-Ming Chen, Samuel A. Danziger, Jung-Hsien Chiang, John D. Aitchison |
Bioinform. | 1 |
| 2013 | Improving Graph Cuts algorithm to transform sequence of stereo image to depth map
Wei-Ming Chen, Sheng-Hao Jhang |
J. Syst. Softw. | 1 |
| 2012 | A test problem with adjustable degrees of overlap and conflict among subproblemsabstractIn the field of genetic algorithms (GAs), some researches on overlapping building blocks (BBs) have been proposed. To further study on overlapping BBs, we need to measure the performance of an algorithm to solve problems with over-lap among subproblems. Several test problems have been proposed, but the controllability over the degree of overlapping is not yet fully satisfactory. Our new test problem is designed with full controllability of overlapping as well as conflict, a specific type of overlap, among BBs. We present a framework for building the structure of this problem in this paper. Some model-building GAs are tested by the proposed problem. This test problem can be applied to further researches on overlapping and conflicting BBs. Wei-Ming Chen, Chung-Yu Shao, Po-Chun Hsu, Tian-Li Yu 0001 |
GECCO | 1 |
| 2012 | A low-power current-mode front-end acquisition system for biopotential signal recordingabstractIn this paper, a new current-mode front-end amplifier for biopotential signal recording system is proposed. A current-mode preamplifier incorporates with an active feedback loop to bypass any dc current generated by tissues is designed as the first stage. A programmable gain stage and a current-mode filter are designed to adjust both gain and the low-pass cutoff frequency, respectively. The current-mode front-end amplifier is designed and fabricated in 0.18-μm CMOS technology. The measured maximum current gain is 55.9 dB. The high-pass cutoff frequency can achieve as low as 0.3Hz and the low-pass cutoff frequency can be adjusted from 1 KHz to 10 KHz. The measured input-referred noise is 153fA/√Hz, and the power consumption is 13 μW at 1-V power supply. Wei-Ming Chen, Liang-Ting Kuo, Chung-Yu Wu |
ISCAS | 1 |
| 2011 | Design and implementation of a high performance closed-loop MIMO communications with ultra low complexity handsetabstractAn efficient and practicable MIMO transceiver in which transmitter antenna selection is applied to geometric mean decomposition (GMD) which is combined with Tomlinson-Harashima Precoder (THP) in TDD system is implemented. The proposed work can save more than 60% computational complexity at the handset compared with that of the GMD scheme is comparable to the conventional linear transceiver schemes. From the simulation results, the proposed transceiver can achieve about 7 dB SNR improvement over the open-loop VBLAST counterparts under i.i.d. channel. Finally, a MIMO joint transceiver is implemented on a SoC platform which is realized to do the hardware/software (HW/SW) co-verification strategy to debug the proposed architecture. In this paper, which introduces the figure file to be the transmission media. Designer could verify the decoded results in various environment by liquid crystal display (LCD) panel. Yu-Han Yuan, Wei-Ming Chen, Hsi-Pin Ma |
ASP-DAC | 2 |
| 2008 | A novel secure communication scheme in vehicular ad hoc networks
Neng-Wen Wang, Yueh-Min Huang, Wei-Ming Chen |
Comput. Commun. | 3 |
| 2007 | Jumping ant routing algorithm for sensor networks
Wei-Ming Chen, Chung-Sheng Li, Fu-Yu Chiang, Han-Chieh Chao |
Comput. Commun. | 1 |
| 2006 | An Energy-Aware Quality of Services Routing Protocol in Mobile Ad Hoc Networks
Yun-Sheng Yen, Chih-Shan Liao, Ruay-Shiung Chang, Han-Chieh Chao, Wei-Ming Chen |
WASA | 5 |
| 2005 | Design and analysis of GPRS-WLAN mobility gateway (GWMG)abstractThis paper presents the design and analysis of GPRS-WLAN mobility gateway (GWMG) for the integration of GPRS and wireless LANs (WLANs). The proposed architecture leverages mobile IP as the mobility management protocol over WLANs. The interworking between GPRS and WLANs is achieved by the GWMG which resides on the border of GPRS and WLAN systems. The design goal is to minimize the modifications in GPRS and WLANs as that both systems are widely available in the markets already. By deploying the GWMG, users can seamlessly roam among two systems. Both mathematical analysis and simulation are developed to analyze the performance. The proposed GWMG has also been implemented in a commercial GPRS system. The results show that the GWMG could achieve the design goal to effectively integrate GPRS and WLANs. Jyh-Cheng Chen, Wei-Ming Chen, Hong-Wei Lin |
ICC | 2 |
| 2005 | FPGA Authentication Header (AH) Implementation for Internet AppliancesabstractData integrity assurance and data origin authentication are essential security services in financial transactions, electronic commerce, electronic mail, software distribution, data storage and so on. Nowadays, consumer electronics has been shifted toward Internet or intelligent appliances (IA) with network capability to exchange information through Internet. Therefore, a hardware based security mechanism is essential to be combined into the IA so that security and performance can be both preserved. In the Internet protocol security (IPSec) mechanism, the authentication header (AH) is an important portion. The two authentication algorithms specified for AH are MD5 and SHA-1 which have been implemented and evaluated in FPGA. With the proposed enhanced (register usage and concurrent statement) operation core design, a 6% improvement for slice utilization plus 24% more throughput for MD5 are obtained comparing to the previous one. Chang-Chun Cheng, Wei-Ming Chen, Han-Chieh Chao, Yao-Po Wang |
PRDC | 2 |
| 1996 | Interframe difference quadtree edge-based side-match finite-state classified vector quantization for image sequence codingabstractVery low bit rate image sequence coding is very important for video transmission and storage applications. The fundamental goal of image sequence coding is to remove spatial redundancy and only update the moving parts in an image sequence so that the total information bits required for storage or transmission can be greatly reduced. In our proposed approach, each frame within an image sequence will be separated into moving and stationary blocks. Only the moving blocks need to be transmitted to the decoder, so that the total number of bits and computing time are greatly reduced. The moving blocks will be encoded by edge-based side-match finite-state classified vector quantization (EBSMCVQ). Moreover, a quadtree can be also used to represent the moving and edge information for each frame. In order to reduce the number of bits for the quadtree, we propose a new difference quadtree technique that only transmits the different parts between the previous frame's quadtree and the current frame's quadtree. In the proposed interframe difference quadtree EBSMCVQ image sequence coding scheme, the average bit rate of each frame is reduced to 0.0393 b/pixel and the PSNR is still up to 35.40 dB for image sequence Claire. Ruey-Feng Chang, Wei-Ming Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1996 | Adaptive edge-based side-match finite-state classified vector quantization with quadtree mapabstractVector quantization (VQ) is an effective image coding technique at low bit rate. The side-match finite-state vector quantizer (SMVQ) exploits the correlations between neighboring blocks (vectors) to avoid large gray level transition across block boundaries. A new adaptive edge-based side-match finite-state classified vector quantizer (classified FSVQ) with a quadtree map has been proposed. In classified FSVQ, blocks are arranged into two main classes, edge blocks and nonedge blocks, to avoid selecting a wrong state codebook for an input block. In order to improve the image quality, edge vectors are reclassified into 16 classes. Each class uses a master codebook that is different from the codebooks of other classes. In our experiments, results are given and comparisons are made between the new scheme and ordinary SMVQ and VQ coding techniques. As is shown, the improvement over ordinary SMVQ is up to 1.16 dB at nearly the same bit rate, moreover, the improvement over ordinary VQ can be up to 2.08 dB at the same bit rate for the image, Lena. Further, block boundaries and edge degradation are less visible because of the edge-vector classification. Hence, the perceptual image quality of classified FSVQ is better than that of ordinary SMVQ. Ruey-Feng Chang, Wei-Ming Chen |
IEEE Trans. Image Process. | 2 |