VLDB 2026 Research / reviewers in the wild / expert
Wei-Yu Chen
dblp:01/448
· DBLP profile ↗
58ranked-venue papers
17as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 5 since 2021Computer networks · 12 · 4 first-author · 9 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A High-Performance and Scalable sFlow Scheme with Zero Control-Plane Overhead
Shie-Yuan Wang, Wei-Yu Chen, Yen Wang, Xiang-Ling Lin |
ICC | 2 |
| 2026 | Multi-source domain adaptive object detection under different privacy levels
Peggy Joy Lu, Wei-Yu Chen, Chia-Yung Jui, Vincent S. Tseng, Jen-Hui Chuang |
Multim. Syst. | 2 |
| 2025 | A generative-adversarial-network-based temporal raw trace data augmentation framework for fault detection in semiconductor manufacturing
Shu-Kai S. Fan, Wei-Yu Chen |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Coherence as Texture - Passive Textureless 3D Reconstruction by Self-InterferenceabstractPassive depth estimation based on stereo or defocus relies on the presence of the texture on an object to resolve its depth. Hence, recovering the depth of a textureless object-for example, a large white wall-is not just hard but perhaps even impossible. Or is it? We show that spatial coherence, a property of natural light sources, can be used to resolve the depth of a scene point even when it is textureless. Our approach relies on the idea that natural light scattered off a scene point is locally coherent with itself, while incoherent with the light scattered from other surface points; we use this insight to design an optical setup that uses self-interference as a texture feature for estimating depth. Our lab prototype is capable of resolving depths of textureless objects in sunlight as well as indoor lights. Wei-Yu Chen, Aswin C. Sankaranarayanan, Anat Levin, Matthew O'Toole |
CVPR | 1 |
| 2024 | Enhancing ECAPA-TDNN with Feature Processing Module and Attention Mechanism for Speaker Verification
Shiu-Hsiang Liou, Po-Cheng Chan, Chia-Ping Chen, Tzu-Chieh Lin, Chung-Li Lu, Yu-Han Cheng, Hsiang-Feng Chuang, Wei-Yu Chen |
INTERSPEECH | 8 |
| 2024 | Federated Contrastive Domain Adaptation for Category-inconsistent Object DetectionabstractTo obtain diverse scenarios for collaboratively training a more generalized object detector, the multi-source domain adaptive object detection has been proposed. However, such scenario faces challenges related to data privacy, domain discrepancy and category inconsistency, thus we propose a framework called FedCoin: Federated Contrastive domain-adaptation for category-inconsistent object detection. On the client sides, a novel dynamic model contrastive strategy is proposed to reduce excessively domain-specific features from local models. On the server side, we design a two-stage teacher-student architecture to tackle the challenge of backbone aggregation and inconsistent categories integration. Our method outperforms SOTA methods across different domain adaptation tasks, with an average precision increase of 6% on various datasets, demonstrating its superiority over existing methods for category-inconsistent and privacy-preserving scenarios. The source code is available online: https://github.com/ccuvislab/FedCoin Wei-Yu Chen, Peggy Joy Lu, Vincent S. Tseng |
VCIP | 1 |
| 2024 | Joint Multi-User Grouping and AP Switch On/Off for Energy-Efficient Cell-Free Massive MIMOabstractThis paper investigates multi-user grouping (MUG) to improve energy efficiency (EE) in cell-free massive multiple-input multiple-output systems with access point (AP) switch on/off (ASO). ASO puts those APs into sleep mode which have little contribution with any spatially-multiplexed user equipments (UEs). However, in general, ASO can be less effective especially when spatially-multiplexed UEs are widely distributed in the area. Therefore, it is crucial to systematically integrate ASO and MUG, where UEs are divided into groups for spatial multiplexing. As a concrete procedure, we propose a practical MUG-ASO algorithm utilizing the criteria based on correlations of large-scale fading among UEs for effective EE improvement. The proposed algorithm can effectively select closely locating UEs into the same group for spatial multiplexing, and then it can put comparatively large number of APs into sleep mode at each time slot. Numerical example simulations show the effectiveness of the integration for EE improvement; in a typical case, the proposed algorithm can improve EE by 29.4% at 50-percentile compared to the conventional ASO algorithms with random MUG without severe degradation of spectral efficiency. Masaaki Ito, Issei Kanno, Yoji Kishi, Wei-Yu Chen, Andreas F. Molisch |
VTC Spring | 4 |
| 2024 | Impact of Hardware Impairment on the Joint Reconfigurable Intelligent Surface and Robust Transceiver Design in MU-MIMO SystemabstractReconfigurable intelligent surface (RIS) is a revolutionary passive radio technique to facilitate capacity enhancement beyond the current massive multiple-input multiple-output (MIMO) transmission. However, the potential hardware impairment (HWI) of the RIS usually causes inevitable performance degradation and the amplification of imperfect CSI. These impacts still lack full investigation in the RIS-assisted wireless network. This paper developed a robust joint RIS and transceiver design algorithm to minimize the worst-case mean square error (MSE) of the received signal under the HWI effect and imperfect channel state information (CSI) in the RIS-assisted multi-user MIMO (MU-MIMO) wireless network. Specifically, since the proposed robust joint RIS and transceiver design problem yields non-convex characteristics under severe HWI, an iterative three-step convex algorithm is developed to approach the optimality by relaxation and convex transformation. Compared with the state-of-the-art baselines that ignore the HWI, the proposed robust algorithm inhibits the destruction of HWI while raising the worst-case MSE effectively in several numerical simulations. Moreover, due to the properties of the HWI, the performance loss is notable under the magnification of the number of reflected elements in the RIS-assisted MU-MIMO wireless network. Wei-Yu Chen, Chih-Yu Wang 0001, Ren-Hung Hwang, Wen-Tsuen Chen, Sin-Yu Huang |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Efficient RRH Activation Management for 5G V2XabstractVehicle-to-everything (V2X) communication is one of the key technologies of 5G New Radio to support emerging applications such as autonomous driving. Due to the high density of vehicles, Remote Radio Heads (RRHs) will be deployed as Road Side Units to support V2X. Nevertheless, activation of all RRHs during low-traffic off-peak hours may cause energy wasting. The proper activation of RRH and association between vehicles and RRHs while maintaining the required service quality are the keys to reducing energy consumption. In this work, we first formulate the problem as an Integer Linear Programming optimization problem and prove that the problem is NP-hard. Then, we propose two novel algorithms, referred to as “Least Delete (LD)” and “Largest-First Rounding with Capacity Constraints (LFRCC).” The simulation results show that the proposed algorithms can achieve significantly better performance compared with existing solutions and are competitive with the optimal solution. Specifically, the LD and LFRCC algorithms can reduce the number of activated RRHs by 86$\%$and 89$\%$in low-density scenarios. In high-density scenarios, the LD algorithm can reduce the number of activated RRHs by 90$\%$. In addition, the solution of LFRCC is larger than that of the optimal solution within 7$\%$on average. Jing-Wen Ke, Ren-Hung Hwang, Chih-Yu Wang 0001, Jian-Jhih Kuo, Wei-Yu Chen |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Joint IoT Device Selection and Health-Aware Beamforming Design for MIMO-WPTabstractWireless power transfer (WPT) has emerged to enhance the robustness of the energy harvesting Internet of Things (EH-IoT), whereas beamforming has been leveraged to significantly boost the efficiency of far-field WPT. Nevertheless, potential negative impacts due to high electromagnetic fields (EMF) exposure for radiation-susceptible users have not been thoroughly considered in the design of WPT for EH-IoT with IoT application-level requirements (e.g., coverage). In this article, we explore the health-aware beamforming and IoT selection problem under the EH and human safety constraints. First, we formulate a new optimization problem Health-Aware Beamforming and IoT Selection (HABIS) and prove the NP-hardness. Second, we design an approximation algorithm, named Minimum Radiation Exposure and Maximum IoT Coverage (MREMIC), to exploit the EH-health dependency (EHHD) graph for properly addressing the trade-off between EH efficiency and potential EMF radiation exposure to human bodies. We also discover the optimal health-aware beamforming to minimize the total radiation energy absorption of humans. Simulation results show that MREMIC can effectively charge IoT devices and significantly outperforms existing EH approaches by more than 200% regarding human safety. Chih-Hang Wang, Yishuo Shi, De-Nian Yang, Wei-Yu Chen, Wen-Tsuen Chen |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Pointersect: Neural Rendering with Cloud-Ray IntersectionabstractWe propose a novel method that renders point clouds as if they are surfaces. The proposed method is differentiable and requires no scene-specific optimization. This unique capability enables, out-of-the-box, surface normal estimation, rendering room-scale point clouds, inverse rendering, and ray tracing with global illumination. Unlike existing work that focuses on converting point clouds to other representations—e.g., surfaces or implicit functions—our key idea is to directly infer the intersection of a light ray with the underlying surface represented by the given point cloud. Specifically, we train a set transformer that, given a small number of local neighbor points along a light ray, provides the intersection point, the surface normal, and the material blending weights, which are used to render the outcome of this light ray. Localizing the problem into small neighborhoods enables us to train a model with only 48 meshes and apply it to unseen point clouds. Our model achieves higher estimation accuracy than state-of-the-art surface reconstruction and point-cloud rendering methods on three test sets. When applied to room-scale point clouds, without any scene-specific optimization, the model achieves competitive quality with the state-of-the-art novel-view rendering methods. Moreover, we demonstrate ability to render and manipulate Lidar-scanned point clouds such as lighting control and object insertion. Jen-Hao Rick Chang, Wei-Yu Chen, Anurag Ranjan, Kwang Moo Yi, Oncel Tuzel |
CVPR | 2 |
| 2023 | Stochastic Geometry-Based Performance Analysis with Correlated Shadowing in Distributed Antenna SystemsabstractDistributed antenna systems (DAS) have attracted significant attention for next-generation wireless systems. This largely motivated by the inherent macrodiversity, i.e., the fact that shadowing of the links to the different access points (APs) is different, which plays a major role in the reliability and capacity of such systems. However, shadowing correlation might reduce the benefits. In this paper, we provide the first stochastic geometry-based analysis of the impact of correlated shadowing on the uplink performance of DAS. Since the statistics of the SNR are determined by the second moment, we use Fenton-Wilkinson (F-W) moment matching and the second moment measure, to formulate it as triple integral. From this, we further construct two closed-form approximations with different degrees of accuracy and simplicity. Extensive Monte Carlo simulations validate the theoretical inferences and approximation accuracy. The results show that a large decorrelation distance will increase the variance of uplink SNR, and a small path-loss exponent leads to a stronger dependence of the second moment on the decorrelation distance. The results can also serve as the basis for future investigations of cell-free massive MIMO systems. Wei-Yu Chen, Masaaki Ito, Issei Kanno, Thomas Choi 0001, Andreas F. Molisch |
GLOBECOM | 1 |
| 2023 | Adaptive Bit Allocation for SVD based Hybrid Processing of Uplink Cell-Free Massive MIMO under Limited Fronthaul CapacityabstractThis paper suggests and analyzes adaptive bit allocation for the quantization of uplink signals of a cell-free massive MIMO (CF-mMIMO) system under limited fronthaul capacity. Specifically, we consider a CF-mMIMO system with hybrid processing, where at each access point (AP) a singular-value decomposition (SVD) reduces the number of streams that need to hauled, each stream is quantized with an adaptive number of bits, and a central processing unit (CPU) decodes the uplink signals. The hybrid processing, which the authors previously proposed, had been shown its potential to reduce fronthaul load without severe degradation of the spectral efficiency. However, as the bandwidths of the wireless system increases, the fronthaul capacity becomes comparatively tight, and the quantization noise would degrade the spectral efficiency severely. In order to improve the performance under such a scenario, this paper proposes algorithms for adaptive bit allocation of the output streams, based on the optimization of the average SNR, or the sum capacity. In addition, appropriate selection of the number of streams of the hybrid processing in each AP is also discussed. Computer simulations verify the effectiveness of these proposed methods. Issei Kanno, Masaaki Ito, Yoshiaki Amano, Yoji Kishi, Thomas Choi 0001, Wei-Yu Chen, Andreas F. Molisch |
VTC2023-Spring | 6 |
| 2023 | Dual Pricing Optimization for Live Video Streaming in Mobile Edge Computing With Joint User Association and Resource ManagementabstractMobile live video streaming is expected to become mainstream in the fifth generation (5G) mobile networks. To boost the Quality of Experience (QoE) of streaming services, the integration of Scalable Video Coding (SVC) with Mobile Edge Computing (MEC) becomes a natural candidate due to its scalability and the reliable transmission supports for real-time interactions. However, it still takes efforts to integrate MEC into video streaming services to exploit its full potentials. We find that the efficiency of the MEC-enabled cellular system can be significantly improved when the requests of users can be redirected to proper MEC servers through optimal user associations. In light of this observation, we jointly address the caching placement, video quality decision, and user association problem in the live video streaming service. Since the proposed nonlinear integer optimization problem is NP-hard, we first develop a two-step approach from a Lagrangian optimization under the dual pricing specification. Further, to have a computation-efficient solution and less performance loss, we provide a one-step Lagrangian dual pricing algorithm by the convex transformation of non-convex constraints. The simulations show that the service quality of live video streaming can be remarkably enhanced by the proposed algorithms in the MEC-enabled cellular system. Wei-Yu Chen, Po-Yu Chou, Chih-Yu Wang 0001, Ren-Hung Hwang, Wen-Tsuen Chen |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Split-Lohmann Multifocal DisplaysabstractThis work provides the design of a multifocal display that can create a dense stack of focal planes in a single shot. We achieve this using a novel computational lens that provides spatial selectivity in its focal length, i.e, the lens appears to have different focal lengths across points on a display behind it. This enables a multifocal display via an appropriate selection of the spatially-varying focal length, thereby avoiding time multiplexing techniques that are associated with traditional focus tunable lenses. The idea central to this design is a modification of a Lohmann lens, a focus tunable lens created with two cubic phase plates that translate relative to each other. Using optical relays and a phase spatial light modulator, we replace the physical translation of the cubic plates with an optical one, while simultaneously allowing for different pixels on the display to undergo different amounts of translations and, consequently, different focal lengths. We refer to this design as a Split-Lohmann multifocal display. Split-Lohmann displays provide a large étendue as well as high spatial and depth resolutions; the absence of time multiplexing and the extremely light computational footprint for content processing makes it suitable for video and interactive experiences. Using a lab prototype, we show results over a wide range of static, dynamic, and interactive 3D scenes, showcasing high visual quality over a large working range. Yingsi Qin, Wei-Yu Chen, Matthew O'Toole, Aswin C. Sankaranarayanan |
ACM Trans. Graph. | 2 |
| 2022 | A Realistic Path Loss Model for Cell-Free Massive MIMO in Urban EnvironmentsabstractCell-free massive multi-input multi-output (CF-mMIMO) systems are one of the key technologies for 6G. Currently, performance assessment of such systems is hampered by the fact that there are no specific path loss (PL) models for CF-mMIMO. Conventional PL models based on Euclidean distance and log-normal shadowing assuming spatial stationarity across coverage area are usually employed for simplicity but show significant deviations from reality particularly in urban environments, which are the main deployment scenario for CF-mMIMO. In this work, we provide the first realistic channel model for CF-mMIMO systems in urban environments, introducing non-isotropic, non-stationary behavior in different parts of street canyon locations and incorporating both correlations between access points, and between user equipments. Simulation results demonstrate the superior reproduction of typical PL values in urban street canyons. Thomas Choi 0001, Issei Kanno, Masaaki Ito, Wei-Yu Chen, Andreas F. Molisch |
GLOBECOM | 4 |
| 2022 | On the Robustness of Cross-lingual Speaker Recognition using Transformer-based ApproachesabstractMost speaker recognition systems presume that the language for enrollment and testing is the same. Cross-lingual speaker recognition is rarely investigated. This study collected trilingual (including Mandarin, English, and Taiwanese) cross-language recordings named MET-40. A total of 40 participants (20 male, 20 female) contribute to the dataset which contains 740 minutes of audio. Spoken texts are mainly taken from elementary school textbooks, and some English texts use TIMIT.We employ ResNet, vision transformer (ViT), and convo-lutional vision transformer (CvT) in combination with three acoustic features, namely, spectrogram, Mel spectrogram, and Mel frequency cepstral coefficient for single, mixed and cross-language speaker recognition tasks. In the mixed-language setting, the language to be tested is included in the training set, while in the cross-language scenario the language to be tested is not used for training. Experimental results show that the highest accuracy is 97.16% for single language models. Mixture of two languages improves the performance to 99.17%. In cross-language situations, the accuracy drops significantly to 79.64%, as the spoken language is not present in the training data. When two languages are employed for training, the accuracy rose to 90.92%. In general, CvT-based models demonstrate the best stability in all cases.The robustness of the model is critical to security in practical applications. Therefore, we analyze how adversarial attacks impact different speaker identification models. The results show that although CvT-based model exhibits excellent performance, it is easily affected by the perturbation caused by the adversarial attack. The effect is less pronounced when more languages are used for training, with an average increase of 5.11% in accuracy. Finally, extra caution needs to be taken when MFCC is chosen to be the acoustic feature, as attacks can still take place without training data, and the recognition rate is reduced by 31.57% using FGSM cross-language attack. Wen-Hung Liao, Wei-Yu Chen, Yi-Chieh Wu |
ICPR | 2 |
| 2022 | Joint AP On/Off and User-Centric Clustering for Energy-Efficient Cell-Free Massive MIMO SystemsabstractCell-free massive multiple-input multiple-output systems are expected to provide faster and more robust connections to user equipments (UEs) by cooperation of a massive number of distributed access points (APs). Energy efficiency (EE) is becoming an important indicator to design and operate networks; to improve EE, use of sleep-mode of APs (SMA), also called AP switch on/off, for selected APs has been investigated. Although previous works analyze the performance of SMA in the presence of user-centric clustering (UCC), these two techniques are assumed to not affect each other. In this paper, we propose a new greedy combining algorithm (GCA), where SMA and UCC work alternately to obtain better performance, and show its superiority over a conventional algorithm. Example simulations show that GCA can achieve 44% higher total EE for 8 UEs and 59% for 16 UEs with 64 APs. Additionally, GCA also provides higher minimum spectral efficiency thanks to its structure of the algorithm. Masaaki Ito, Issei Kanno, Yoshiaki Amano, Yoji Kishi, Wei-Yu Chen, Thomas Choi 0001, Andreas F. Molisch |
VTC Fall | 5 |
| 2022 | Fronthaul Load-Reduced Scalable Cell-Free massive MIMO by Uplink Hybrid Signal ProcessingabstractThis paper proposes hybrid signal processing schemes for the uplink cell-free massive MIMO; these schemes serve to reduce fronthaul loads to obtain a scalable centralized processing architecture. In this architecture, received signals of multiple receive antennas at the access points (APs) are compressed into fewer streams by local spatial signal processing and then the streams are forwarded to a central processing unit (CPU) via fronthaul, and the CPU performs scalable processing for channel estimation and signal detection based on partial minimum mean squared error (PMMSE). We propose two kinds of concrete local signal processing methods for this hybrid processing architecture: one is based on MMSE, and the other is based on principal component analysis (PCA) with eigenvalue decomposition (EVD). For the EVD, a local vector selection based EVD (LVS-EVD) that selects uniform number of eigenvectors for each AP in a standalone way, and a global vector selection based EVD (GVS-EVD) that determines the dimensions of the weight vector of each AP in the CPU, are further considered. Computer simulations verify the approaches and compare their effectiveness. In addition, we show that the GVS-EVD scheme can be operated with significantly reduced fronthaul loads without severe performance degradation. Issei Kanno, Masaaki Ito, Takeo Ohseki, Kosuke Yamazaki, Yoji Kishi, Thomas Choi 0001, Wei-Yu Chen, Andreas F. Molisch |
VTC Spring | 7 |
| 2022 | Pricing-Based Deep Reinforcement Learning for Live Video Streaming With Joint User Association and Resource Management in Mobile Edge ComputingabstractMobile Edge Computing (MEC) is a promising technique in the 5G Era to improve the Quality of Experience (QoE) for online video streaming due to its ability to reduce the backhaul transmission by caching certain content. However, it still takes effort to address the user association and video quality selection problem under the limited resource of MEC to fully support the low-latency demand for live video streaming. We found the optimization problem to be a non-linear integer programming, which is impossible to obtain a globally optimal solution under polynomial time. In this paper, we formulate the problem and derive the closed-form solution in the form of Lagrangian multipliers; the searching of the optimal variables is formulated as a Multi-Arm Bandit (MAB) and we propose a Deep Deterministic Policy Gradient (DDPG) based algorithm exploiting the supply-demand interpretation of the Lagrange dual problem. Simulation results show that our proposed approach achieves significant QoE improvement, especially in the low wireless resource and high user number scenario compared to other baselines. Po-Yu Chou, Wei-Yu Chen, Chih-Yu Wang 0001, Ren-Hung Hwang, Wen-Tsuen Chen |
IEEE Trans. Wirel. Commun. | 2 |
| 2021 | On-Demand Service Function Chain Based on IPv6 Segment RoutingabstractSegment routing network is the evolution and innovation of IP routing technology. With the development of mobile services and cloud computing, the segment routing technology based on IPv6 has attracted more attentions. Segment Routing over IPv6 (SRv6) is a source routing solution for IPv6 networks. SRv6 has the ability of service chain and network programmability, which can strengthen the advantages of traffic engineering and cloud network to realize various innovative applications. This paper mainly utilizes SRv6 technology to implement a Service Function Chain (SFC) combined with network security applications. Experimental results show that using SRv6 to guide the SFC path for malicious attack detection and traffic restriction can effectively improve IPv6 network security and service quality. Chia-Wei Wu, Chia-Wei Tseng, Wei-Yu Chen, Li-Fan Wu, Shih-Chun Hsu, Sheng-Wang Yu |
APNOMS | 3 |
| 2021 | C-for-Metal: High Performance Simd Programming on Intel GPUsabstractThe SIMT execution model is commonly used for general GPU development. CUDA and OpenCL developers write scalar code that is implicitly parallelized by compiler and hardware. On Intel GPUs, however, this abstraction has profound performance implications as the underlying ISA is SIMD and important hardware capabilities cannot be fully utilized. To close this performance gap we introduce C- For- Metal (CM), an explicit SIMD programming framework designed to deliver close-to-the-metal performance on Intel GPUs. The CM programming language and its vector/matrix types provide an intuitive interface to exploit the underlying hardware features, allowing fine-grained register management, SIMD size control and cross-lane data sharing. Experimental results show that CM applications from different domains outperform the best-known SIMT-based OpenCL implementations, achieving up to 2.7x speedup on the latest Intel GPU. Guei-Yuan Lueh, Kaiyu Chen, Joel Fuentes, Wei-Yu Chen, Fangwen Fu, Hongzheng Li, Daniel Rhee |
CGO | 5 |
| 2021 | Reference Wave Design for Wavefront SensingabstractOne of the classical results in wavefront sensing is phase-shifting point diffraction interferometry (PS-PDI), where the phase of a wavefront is measured by interfering it with a planar reference created from the incident wave itself. The limiting drawback of this approach is that the planar reference, often created by passing light through a narrow pinhole, is dim and noise sensitive. We address this limitation with a novel approach called ReWave that uses a non-planar reference that is designed to be brighter. The reference wave is designed in a specific way that would still allow for analytic phase recovery, exploiting ideas of sparse phase retrieval algorithms. ReWave requires only four image intensity measurements and is significantly more robust to noise compared to PS-PDI. We validate the robustness and applicability of our approach using a suite of simulated and real results. Wei-Yu Chen, Anat Levin, Matthew O'Toole, Aswin C. Sankaranarayanan |
ICCP | 1 |
| 2021 | Lightweight fog computing-based authentication protocols using physically unclonable functions for internet of medical things
Tian-Fu Lee, Wei-Yu Chen |
J. Inf. Secur. Appl. | 2 |
| 2020 | Deep Reinforcement Learning for MEC Streaming with Joint User Association and Resource ManagementabstractMobile Edge Computing (MEC) is a promising technique in the 5G Era to improve the Quality of Experience (QoE) for online video streaming due to its ability to reduce the backhaul transmission by caching certain content. However, it still takes effort to address the user association and video quality selection problem under the limited resource of MEC to fully support the low-latency demand for live video streaming. We found the optimization problem to be a non-linear integer programming, which is impossible to obtain a globally optimal solution under polynomial time. In this paper, we first reformulate this problem as a Markov Decision Process (MDP) and develop a Deep Deterministic Policy Gradient (DDPG) based algorithm exploiting the supply-demand interpretation of the Lagrange dual problem. Simulation results show that our proposed approach achieves significant QoE improvement especially in the low wireless resource and high user number scenario compared to other baselines. Po-Yu Chou, Wei-Yu Chen, Chih-Yu Wang 0001, Ren-Hung Hwang, Wen-Tsuen Chen |
ICC | 2 |
| 2019 | IGC: The Open Source Intel Graphics CompilerabstractWith increasing general purpose programming capability, GPUs have become the mainstay for a wide variety of compute intensive tasks from cloud to edge computing. Because of its availability on nearly every desktop and mobile processor that Intel ships, Intel integrated GPU offers a plethora of opportunities for researchers and application developers to make significant real-world impact. In this paper we present the Intel Graphics Compiler (IGC), the LLVM-based production compiler for Intel HD and Iris graphics. IGC supports all major graphics and compute APIs, and its OpenCL compute stack including compute runtime, compiler frontend and backend, and architecture specification is fully open-source, giving a unique opportunity for developers to optimize the entire stack. We highlight several custom optimizations that address the challenges for GPU compilation. Examples include SIMD size selection, divergence analysis, instruction scheduling, addressing mode selection, and redundant copy elimination. These optimizations take advantage of features in the Intel GPU architecture such as a larger register file, indirect register addressing, and multiple memory addressing modes. Experimental results show that our optimizations deliver significant speedup on a number of OpenCL benchmarks; compared to the baseline, we see a geometric mean of 12% speed up across benchmarks with a peak gain of 45%. Anupama Chandrasekhar, Wei-Yu Chen, Junjie Gu, Shruthi Hebbur Prasanna Kumar, Guei-Yuan Lueh, Pankaj Mistry, Thomas Raoux, Konrad Trifunovic |
CGO | 4 |
| 2019 | A Closer Look at Few-shot Classification
Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, Jia-Bin Huang 0001 |
ICLR (Poster) | 1 |
| 2019 | Transfer Neural Trees: Semi-Supervised Heterogeneous Domain Adaptation and BeyondabstractHeterogeneous domain adaptation (HDA) addresses the task of associating data not only across dissimilar domains but also described by different types of features. Inspired by the recent advances of neural networks and deep learning, we propose a deep leaning model of Transfer Neural Trees (TNT), which jointly solves cross-domain feature mapping, adaptation, and classification in a unified architecture. As the prediction layer in TNT, we introduce Transfer Neural Decision Forest (Transfer- NDF), which is able to learn the neurons in TNT for adaptation by stochastic pruning. In order to handle semi-supervised HDA, a unique embedding loss term is introduced to TNT for preserving prediction and structural consistency between labeled and unlabeled target-domain data. We further show that our TNT can be extended to zero shot learning for associating image and attribute data with promising performance. Finally, experiments on different classification tasks across features, datasets, and modalities would verify the effectiveness of our TNT. Wei-Yu Chen, Tzu-Ming Harry Hsu, Yao-Hung Tsai, Ming-Syan Chen, Yu-Chiang Frank Wang |
IEEE Trans. Image Process. | 1 |
| 2018 | Register allocation for Intel processor graphicsabstractRegister allocation is a well-studied problem, but surprisingly little work has been published on assigning registers for GPU architectures. In this paper we present the register allocator in the production compiler for Intel HD and Iris Graphics. Intel GPUs feature a large byte-addressable register file organized into banks, an expressive instruction set that supports variable SIMD-sizes and divergent control flow, and high spill overhead due to relatively long memory latencies. These distinctive characteristics impose challenges for register allocation, as input programs may have arbitrarily-sized variables, partial updates, and complex control flow. Not only should the allocator make a program spill-free, but it must also reduce the number of register bank conflicts and anti-dependencies. Since compilation occurs in a JIT environment, the allocator also needs to incur little overhead. Wei-Yu Chen, Guei-Yuan Lueh, Pratik Ashar, Kaiyu Chen, Buqi Cheng |
CGO | 1 |
| 2018 | Robust Minimax MSE Transceiver and Power Splitting Design for Multiuser MIMO SWIPT: An LMI ApproachabstractThis paper proposes a robust joint transceiver and power splitting (PS) design according to a minimax mean-squar-error (MSE) scheme for multiuser multi- input multi-output (MU-MIMO)simultaneous wireless information and power transfer (SWIPT) system. The proposed scheme considers channel uncertainty as a bound on the spectral matrix norm of each estimated error. Since the proposed robust minimax MSE design problem jointly optimizes the precoder, the equalizers and the PS factors, the proposed design problem is non-convex and non-deterministic polynomial-time hard (NP-hard). Thus, we transform the original optimization problem into an iteratively linear matrix inequalities (LMIs)- constrained optimization problem, which can be solved efficiently by the convex optimization toolbox. According to the simulation results of the MU-MIMO PS SWIPT wireless system with the benchmark of robust precoder, the proposed method can validate the theoretical analysis to achieve robust transceiver design effectively. Wei-Yu Chen, Bor-Sen Chen, Wen-Tsuen Chen |
GLOBECOM | 1 |
| 2018 | Noncooperative Game Strategy in Cyber-Financial Systems With Wiener and Poisson Random Fluctuations: LMIs-Constrained MOEA ApproachabstractThe financial market is a nonlinear stochastic system with continuous Wiener and discontinuous Poisson random fluctuations. Most managers or investors hope their investment policies to be with the not only high profit but also low risk. Managers and investors involved pursue their own interests which are partly conflicting with others. Stochastic game theory has been widely applied to multiperson noncooperative decision making problem of financial market. However, for the nonlinear stochastic financial system with random fluctuations, it still lacks an analytical or computational scheme to effectively solve the complex noncooperative game strategy design problem. In this paper, the stochastic multiperson noncooperative game strategy in cyber-financial systems is transformed to a multituple Hamilton-Jacobi-Isacc inequalities (HJIIs)-constrained multiobjective optimization problem (MOP). This HJIIs-constrained MOP solution is also found to be the Nash equilibrium solution of multiperson noncooperative game strategy in nonlinear stochastic financial systems. In order to simplify design procedure by the global linearization theory, a set of local linear systems are interpolated to approximate the nonlinear stochastic financial system so that the m-tuple HJIIs-constrained MOP for noncooperative game strategy of cyber-financial system could be converted to a linear matrix inequalities (LMIs)-constrained MOP. Finally, an LMIs-constrained multiobjective evolution algorithm is explored for effectively solving the multiperson noncooperative game strategy in cyber-financial systems. Two design examples are also given for the illustration of the design procedure and the performance validation of the proposed stochastic noncooperative investment strategy in the nonlinear stochastic financial systems. Bor-Sen Chen, Wei-Yu Chen, Chun-Tao Young, Zhiguo Yan |
IEEE Trans. Cybern. | 2 |
| 2017 | Enhanced canonical correlation analysis with local density for cross-domain visual classificationabstractReal-world visual classification tasks typically need to deal with data observed from different domains. Inspired by canonical correlation analysis (CCA), we propose an enhanced CCA with local density for associating and recognizing cross-domain data. In addition to maximizing the correlation of the projected cross-domain data, our CCA model further exploits the local density information observed from each domain. As a result, our CCA not only exhibits excellent abilities in identifying representative data, noisy data like outliers can be further suppressed during the derivation of our CCA subspace. In our experiments, we successfully apply the proposed methods for solving two cross-domain classification tasks: person re-identification and cross-view action recognition. Wei-Jen Ko, Jheng-Ying Yu, Wei-Yu Chen, Yu-Chiang Frank Wang |
ICASSP | 3 |
| 2017 | No More Discrimination: Cross City Adaptation of Road Scene SegmentersabstractDespite the recent success of deep-learning based semantic segmentation, deploying a pre-trained road scene segmenter to a city whose images are not presented in the training set would not achieve satisfactory performance due to dataset biases. Instead of collecting a large number of annotated images of each city of interest to train or refine the segmenter, we propose an unsupervised learning approach to adapt road scene segmenters across different cities. By utilizing Google Street View and its time-machine feature, we can collect unannotated images for each road scene at different times, so that the associated static-object priors can be extracted accordingly. By advancing a joint global and class-specific domain adversarial learning framework, adaptation of pre-trained segmenters to that city can be achieved without the need of any user annotation or interaction. We show that our method improves the performance of semantic segmentation in multiple cities across continents, while it performs favorably against state-of-the-art approaches requiring annotated training data. Yi-Hsin Chen, Wei-Yu Chen, Bo-Cheng Tsai, Yu-Chiang Frank Wang |
ICCV | 2 |
| 2017 | An Approximation Algorithm for the Maximum-Lifetime Data Aggregation Tree Problem in Wireless Sensor NetworksabstractThis paper studies the problem of constructing maximum-lifetime data aggregation trees in wireless sensor networks for collecting sensor readings. This problem is known to be NP-hard. Wireless sensor networks in which transmission power levels of sensors are adjustable and heterogeneous are considered. An approximation algorithm is developed to construct a data aggregation tree whose inverse lifetime is guaranteed to be within a bound from the optimal one. Adjustable transmission power levels of the sensors introduce an additional term in the bound compared with the bound for networks in which transmission power levels of all sensors are fixed. The additional term is proportional to the difference between the maximum and minimum amounts of energy for a sensor to transmit a message using respectively its maximum and minimum transmission power levels. The proposed algorithm is further enhanced to obtain an improved version. Simulation results show that properly adjusting transmission power levels of the sensors yields higher lifetime of the network than keeping their transmission power levels at the maximum level. Hwa-Chun Lin, Wei-Yu Chen |
IEEE Trans. Wirel. Commun. | 2 |
| 2016 | Domain-Constraint Transfer Coding for Imbalanced Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) deals with the task that labeled training and unlabeled test data collected from source and target domains, respectively. In this paper, we particularly address the practical and challenging scenario of imbalanced cross-domain data. That is, we do not assume the label numbers across domains to be the same, and we also allow the data in each domain to be collected from multiple datasets/sub-domains. To solve the above task of imbalanced domain adaptation, we propose a novel algorithm of Domain-constraint Transfer Coding (DcTC). Our DcTC is able to exploit latent subdomains within and across data domains, and learns a common feature space for joint adaptation and classification purposes. Without assuming balanced cross-domain data as most existing UDA approaches do, we show that our method performs favorably against state-of-the-art methods on multiple cross-domain visual classification tasks. Yao-Hung Tsai, Cheng-An Hou, Wei-Yu Chen, Yi-Ren Yeh, Yu-Chiang Frank Wang |
AAAI | 3 |
| 2016 | Transfer Neural Trees for Heterogeneous Domain Adaptation
Wei-Yu Chen, Tzu-Ming Harry Hsu, Yao-Hung Tsai, Yu-Chiang Frank Wang, Ming-Syan Chen |
ECCV (5) | 1 |
| 2015 | Unsupervised Domain Adaptation with Imbalanced Cross-Domain DataabstractWe address a challenging unsupervised domain adaptation problem with imbalanced cross-domain data. For standard unsupervised domain adaptation, one typically obtains labeled data in the source domain and only observes unlabeled data in the target domain. However, most existing works do not consider the scenarios in which either the label numbers across domains are different, or the data in the source and/or target domains might be collected from multiple datasets. To address the aforementioned settings of imbalanced cross-domain data, we propose Closest Common Space Learning (CCSL) for associating such data with the capability of preserving label and structural information within and across domains. Experiments on multiple cross-domain visual classification tasks confirm that our method performs favorably against state-of-the-art approaches, especially when imbalanced cross-domain data are presented. Tzu-Ming Harry Hsu, Wei-Yu Chen, Cheng-An Hou, Yao-Hung Tsai, Yi-Ren Yeh, Yu-Chiang Frank Wang |
ICCV | 2 |
| 2015 | Connecting the dots without clues: Unsupervised domain adaptation for cross-domain visual classificationabstractMany real-world visual classification tasks require one to recognize test data in a particular domain of interest, while the training data can only be collected from a different domain. This can be viewed as the problem of unsupervised domain adaptation, in which the domain difference and the lack of cross-domain label/correspondence information make the recognition task very difficult. In this paper, we propose to exploit the cross-domain data correspondence using both observed data similarity and labels transferred from the source domain. This allows us to perform distribution matching for cross-domain data with recognition guarantees. Our experiments on three different cross-domain visual classification tasks would confirm the effectiveness of our method, which is shown to perform favorably against state-of-the-art unsupervised domain adaptation approaches. Wei-Yu Chen, Tzu-Ming Harry Hsu, Cheng-An Hou, Yi-Ren Yeh, Yu-Chiang Frank Wang |
ICIP | 1 |
| 2014 | A Novel Layout Decomposition Algorithm for Triple Patterning LithographyabstractWhile double patterning lithography (DPL) has been widely recognized as one of the most promising solutions for the sub-22 nm technology node to enhance pattern printability, triple patterning lithography (TPL) will be required for gate, contact, and metal-1 layers which are too complex and dense to be split into only two masks, for the 15 nm technology node and beyond. Nevertheless, there is very little research focusing on the layout decomposition for TPL. Recent work proposed the first systematic study on the layout decomposition for TPL. However, the proposed algorithm extending a stitch-finding method used in DPL may miss legal stitch locations and generate conflicts that can be resolved by inserting stitches for TPL. In this paper, we point out two main differences between DPL and TPL layout decompositions. Based on the two differences, we propose a novel TPL layout decomposition algorithm. We first present two new graph reduction techniques to reduce the problem size without degrading overall solution quality. We then propose a stitch-aware mask assignment algorithm, based on a heuristic that finds a mask assignment such that the conflicts among the features in the same mask are more likely to be resolved by inserting stitches. Finally, stitches are inserted to resolve as many conflicts as possible. Experimental results show that the proposed layout decomposition algorithm can achieve around 56% reduction of conflicts and more than 40X speed-up, as compared to the previous work. Shao-Yun Fang, Yao-Wen Chang, Wei-Yu Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2013 | Graph-Based Subfield Scheduling for Electron-Beam Photomask FabricationabstractElectron beam lithography has shown great promise in photomask fabrication; however, its successive heating process centralizing in a small region may cause a severe problem of critical dimension (CD) distortion. Consequently, subfield scheduling that reorders the sequence of the writing process is needed to avoid successive writing of neighboring subfields. In addition, the writing process of a subfield raises the temperature of neighboring regions and may block other subfields for writing. This paper presents the first work to solve the subfield scheduling problem while considering blocked regions by formulating the problem into a constrained maximum scatter traveling salesman problem (constrained MSTSP). To tackle the constrained MSTSP that can be shown to be NP-complete in general, we identify a special case thereof with points on two parallel lines and solve it optimally in linear time. We then decompose the constrained MSTSP into subproblems conforming to the special case, solve each subproblem optimally and efficiently by a graph-based algorithm, and then merge the subsolutions into a complete scheduling solution. We also extend our algorithm to handle the cases when the moving time of an e-beam writing head is comparable with the writing time of a subfield. Experimental results show that our algorithms are effective and efficient in finding good subfield scheduling solutions that can alleviate the successive heating problem (and thus reduce CD distortion) for e-beam photomask fabrication. Shao-Yun Fang, Wei-Yu Chen, Yao-Wen Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | A novel layout decomposition algorithm for triple patterning lithographyabstractWhile double patterning lithography (DPL) has been widely recognized as one of the most promising solutions for the sub-22nm technology node to enhance pattern printability, triple patterning lithography (TPL) will be required for gate, contact, and metal-1 layers which are too complex and dense to be split into only two masks, for the 15nm technology node and beyond. Nevertheless, there is very little research focusing on the layout decomposition for TPL. The recent work [16] proposed the first systematic study on the layout decomposition for TPL. However, the proposed algorithm extending a stitch-finding method used in DPL may miss legal stitch locations and generate conflicts that can be resolved by inserting stitches for TPL. In this paper, we point out two main differences between DPL and TPL layout decompositions. Based on the two differences, we propose a novel TPL layout decomposition algorithm. We first present two new graph reduction techniques to reduce the problem size without degrading overall solution quality. We then propose a stitch-aware mask assignment algorithm, based on a heuristic that finds a mask assignment such that the conflicts among the features in the same mask are more likely to be resolved by inserting stitches. Finally, stitches are inserted to resolve as many conflicts as possible. Experimental results show that the proposed layout decomposition algorithm can achieve around 56% reduction of conflicts and more than 40X speed-up compared to the previous work. Shao-Yun Fang, Yao-Wen Chang, Wei-Yu Chen |
DAC | 3 |
| 2012 | Graph-based subfield scheduling for electron-beam photomask fabricationabstractElectron beam lithography (EBL) has shown great promise for photomask fabrication; however, its successive heating process centralizing in a small region may cause a severe problem of critical dimension (CD) distortion. Consequently, subfield scheduling which reorders the sequence of the writing process is needed to avoid successive writing of neighboring subfields. In addition, the writing process of a subfield raises the temperature of neighboring regions and may block other subfields for writing. This paper presents the first work to solve the subfield scheduling problem while taking into account blocked regions by formulating the problem into a constrained maximum scatter travelling salesman problem (constrained MSTSP). To tackle the constrained MSTSP which can be shown to be NP-complete in general, we identify a special case thereof with points on two parallel lines and solve it optimally in linear time. We then decompose the constrained MSTSP into subproblems conforming to the special case, solve each subproblem optimally and efficiently by a graph-based algorithm, and then merge the sub-solutions into a complete scheduling solution. Experimental results show that our algorithm is effective and efficient in finding good subfield scheduling solutions that can alleviate the successive heating problem (and thus reduce CD distortion) for e-beam photomask fabrication. Shao-Yun Fang, Wei-Yu Chen, Yao-Wen Chang |
ISPD | 2 |
| 2010 | The scan-DFT features of AMD's next-generation microprocessor coreabstractThere is an ever-increasing demand for higher performance microprocessors within a given power budget. This demand forces design choices - that were once seen only in high-speed custom blocks - to spread throughout the microprocessor core. These unique design structures, combined with the nanometer technology test challenges such as crosstalk, process variations, power-supply noise, and resistive short and open defects, lead to unique test challenges for today's high-performance microprocessor core. In this paper, we present the scan architecture-related design-for-test (DFT) features and corresponding verification strategies of the nextgeneration Advanced Micro Devices (AMD) high-performance microprocessor core. Mahmut Yilmaz, Jayalakshmi Rajaraman, Tom Olsen, Kanwaldeep Sobti, Dwight Elvey, Jeff Fitzgerald, Grady Giles, Wei-Yu Chen |
ITC | 9 |
| 2009 | Robust adaptive fuzzy control for nonlinear uncertain systems with unknown dead-zone and unknown upper bound of uncertaintiesabstractIn this paper, a robust adaptive fuzzy control for a class of nonlinear uncertain systems preceded by an unknown dead-zone and with unknown upper bound of uncertainties is developed. The dead-zones are quite commonly encountered in many systems (e.g., DC servosystem, robot, and machine tools), are usually poorly known, and may severely limit the performance of control. In addition, the system uncertainties (e.g., parameter variations, or external load, unmodeled dynamics) often exist. Therefore, the controllers are required to deal with the robust stability and performance of the systems with unknown dead-zone and in the presence of uncertainties, whose upper bound is generally unknown. In the beginning, an adaptive dead-zone compensation is employed to improve system performance. Then the unknown system functions and the unknown upper bound of system uncertainties are respectively approximated by fuzzy logic systems with unknown weights. The unknown bounds caused by the learning error of the slope of dead-zone and the system functions are also tackled by an extra learning law. The above weights are all on-line learned to provide for the controller design. Moreover, the projection terms in these learning laws are designed such that the boundedness of the learning weight can be assured. Chih-Lyang Hwang, Chiang-Cheng Chiang, Wei-Yu Chen |
FUZZ-IEEE | 3 |
| 2009 | A High Performance Linear Current Mode Image SensorabstractA new linear current mode image sensor is proposed in this paper. The proposed circuit features high linearity, low power consumption, programmable multiple gain stages, wide input swing and correlated double sampling (CDS) technology. The signal swing of the linear current mode sensor is enhanced by the proposed multiple gain readout structure. A simple and accurate front-end programmable gain structure is proposed to improve the signal-to-noise ratio (SNR) with low power consumption. The function and performance has been verified by HSPICE simulation of 0.18 mum 3T sensor process. Chih-Cheng Hsieh, Wei-Yu Chen, Chung-Yu Wu |
ISCAS | 2 |
| 2007 | Code Generation and Optimization for Transactional Memory Constructs in an Unmanaged LanguageabstractTransactional memory offers significant advantages for concurrency control compared to locks. This paper presents the design and implementation of transactional memory constructs in an unmanaged language. Unmanaged languages pose a unique set of challenges to transactional memory constructs - for example, lack of type and memory safety, use of function pointers, aliasing of local variables, and others. This paper describes novel compiler and runtime mechanisms that address these challenges and optimize the performance of transactions in an unmanaged environment. We have implemented these mechanisms in a production-quality C compiler and a high-performance software transactional memory runtime. We measure the effectiveness of these optimizations and compare the performance of lock-based versus transaction-based programming on a set of concurrent data structures and the SPLASH-2 benchmark suite. On a 16 processor SMP system, the transaction-based version of the SPLASH-2 benchmarks scales much better than the coarse-grain locking version and performs comparably to the fine-grain locking version. Compiler optimizations significantly reduce the overheads of transactional memory so that, on a single thread, the transaction-based version incurs only about 6.4% overhead compared to the lock-based version for the SPLASH-2 benchmark suite. Thus, our system is the first to demonstrate that transactions integrate well with an unmanaged language, and can perform as well as fine-grain locking while providing the programming ease of coarse-grain locking even on an unmanaged environment Cheng Wang 0013, Wei-Yu Chen, Youfeng Wu, Bratin Saha, Ali-Reza Adl-Tabatabai |
CGO | 2 |
| 2007 | A technique for selecting CMOS transistor ordersabstractTransistor reordering has been known to be effective in reducing delays of a circuit with nearly zero penalties. However, techniques to determine good transistor orders have not been proposed in literature. Previous work on this has to resort to running SPICE for all meaningful transistor orders and selecting a best one, which is extremely time-consuming. This paper proposes an efficient and accurate technique for determining best transistor orders without running SPICE simulations. Experimental results from SPICE3 show that the predictions are very accurate. Ting Wei Chiang, C. Y. Roger Chen, Wei-Yu Chen |
ICCD | 3 |
| 2007 | An efficient gate delay model for VLSI designabstractAccurate estimation of gate delays is essential for timing-related CAD tools. CAD researchers tend to use Elmore delay model for estimating gate delays. Since Elmore delay model was primarily developed for estimating interconnection delay, when applied to gate delay estimation, there will be significant inaccuracy. In this paper, by embedding concepts of electronic theories into switch-level analysis, a simple and efficient delay model for gates of general types (such as NAND, NOR, and complex gates) is proposed. Experimental data show that the proposed gate delay model consistently achieves high accuracy (typically within around 2% of SPICE simulations). Ting Wei Chiang, C. Y. Roger Chen, Wei-Yu Chen |
ICCD | 3 |
| 2007 | Automatic nonblocking communication for partitioned global address space programsabstractOverlapping communication with computation is an important optimization on current cluster architectures; its importance is likely to increase as the doubling of processing power far outpaces any improvements in communication latency. PGAS languages offer unique opportunities for communication overlap, because their one-sided communication model enables low overhead data transfer. Recent results have shown the value of hiding latency by manually applying language-level nonblocking data transfer routines, but this process can be both tedious and error-prone. In this paper, we present a runtime framework that automatically schedules the data transfers to achieve overlap. The optimization framework is entirely transparent to the user, and aggressively reorders and aggregates both remote puts and gets. We preserve correctness via runtime conflict checks and temporary buffers, using several techniques to lower the overhead. Experimental results on application benchmarks suggest that our framework can be very effective at hiding communication latency on clusters, improving performance over the blocking code by an average of 16% for some of the NAS Parallel Benchmarks, 48% for GUPS, and over 25% for a multi-block fluid dynamics solver. While the system is not yet as effective as aggressive manual optimization, it increases programmers' productivity by freeing them from the details of communication management. Wei-Yu Chen, Dan Bonachea, Costin Iancu, Katherine A. Yelick |
ICS | 1 |
| 2007 | Ubiquitous e-Helpers: An UPnP-based home automation platformabstractThis paper describes a composeable service platform that supports opportunistic collaboration among smart appliances deployed sporadically and incrementally in living and working spaces. This ubiquitous e-Helper's platform is UPnP based and is consisted of a three-level device/service abstraction and a protocol translation proxy. As the first experiment, the platform is used to perform indoor luminance feedback control using wireless sensors and light dimmers. It outperforms similar systems in its responsiveness to dynamic device behaviors, and the modularity in its hardware/software implementation. John Kar-Kin Zao, Yu-Chih Liu, Ming-Hsiao Yang, Sheng-Kun Li, Wei-Yu Chen, Ching-Wei Chen, Kuo-Chin Huang, Jwu-Sheng Hu, Lun-Chia Kuo |
SMC | 5 |
| 2007 | Resource Allocation and Quality of Service Evaluation for Wireless Communication Systems Using Fluid ModelsabstractWireless systems offer a unique mixture of connectivity, flexibility, and freedom. It is therefore not surprising that wireless technology is being embraced with increasing vigor. For real-time applications, user satisfaction is closely linked to quantities such as queue length, packet loss probability, and delay. System performance is therefore related to, not only Shannon capacity, but also quality of service (QoS) requirements. This work studies the problem of resource allocation in the context of stringent QoS constraints. The joint impact of spectral bandwidth, power, and code rate is considered. Analytical expressions for the probability of buffer overflow, its associated exponential decay rate, and the effective capacity are obtained. Fundamental performance limits for Markov wireless channel models are identified. It is found that, even with an unlimited power and spectral bandwidth budget, only a finite arrival rate can be supported for a QoS constraint defined in terms of exponential decay rate Lingjia Liu 0001, Parimal Parag, Wei-Yu Chen, Jean-François Chamberland |
IEEE Trans. Inf. Theory | 4 |
| 2006 | Hold time validation on silicon and the relevance of hazards in timing analysisabstractIn this paper we motivate the explicit validation of hold-time violations in silicon and propose a method for doing so. New hold-time failure model and test pattern generation methodologies are defined.We outline conditions under which these tests can be applied reliably. We present results of applying these test patterns on a microprocessor and discuss the implications of intermittent failures on the relevance of hazards during timing analysis. Amitava Majumdar 0002, Wei-Yu Chen |
DAC | 2 |
| 2004 | Evaluating support for global address space languages on the Cray X1abstractThe Cray X1 was recently introduced as the first in a new line of parallel systems to combine high-bandwidth vector processing with an MPP system architecture. Alongside capabilities such as automatic fine-grained data parallelism through the use of vector instructions, the X1 offers hardware support for a transparent global-address space (GAS), which makes it an interesting target for GAS languages. In this paper, we describe our experience with developing a portable, open-source and high performance compiler for Unified Parallel C (UPC), a SPMD global-address space language extension of ISO C. As part of our implementation effort, we evaluate the X1's hardware support for GAS languages and provide empirical performance characterizations in the context of leveraging features such as vectorization and global pointers for the Berkeley UPC compiler. We discuss several difficulties encountered in the Cray C compiler which are likely to present challenges for many users, especially implementors of libraries and source-to-source translators. Finally, we analyze the performance of our compiler on some benchmark programs and show that, while there are some limitations of the current compilation approach, the Berkeley UPC compiler uses the X1 network more effectively than MPI or SHMEM, and generates serial code whose vectorizability is comparable to the original C code. Christian Bell, Wei-Yu Chen, Dan Bonachea, Katherine A. Yelick |
ICS | 2 |
| 2004 | Split vector-radix-2/8 2-D fast Fourier transformabstractThis letter presents an efficient split vector-radix-2/8 fast Fourier transform (FFT) algorithm. The split vector-radix-2/8 FFT algorithm saves 14% real multiplications and has much lower arithmetic complexity than the split vector-radix-2/4 FFT algorithm. Moreover, this algorithm reduces 25% data loads and stores compared with the split vector-radix-2/4 FFT algorithm. Soo-Chang Pei, Wei-Yu Chen |
IEEE Signal Process. Lett. | 2 |
| 2002 | Test Generation for Crosstalk-Induced Faults: Framework and Computational Results
Wei-Yu Chen, Sandeep Gupta 0001, Melvin A. Breuer |
J. Electron. Test. | 1 |
| 2002 | Analytical models for crosstalk excitation and propagation in VLSI circuitsabstractThe authors develop a general methodology to analyze crosstalk effects that are likely to cause errors in deep submicron high-speed circuits. They focus on crosstalk due to capacitive coupling between a pair of lines. Closed form equations are derived that quantify the severity of these effects and describe qualitatively the dependence of these effects on the values of circuit parameters, the rise/fall times of the input transitions, and the skew between the transitions. For noise propagation, they present a new way for predicting the output waveform produced by an inverter due to a nonsquare wave pulse at its input. To expedite the computation of the response of a logic gate to an input pulse, the authors have developed a novel way of modeling such gates by an equivalent inverter. The results of their analysis provide conditions that must be satisfied by a sequence of vectors used for validation of designs as well as post-manufacturing testing of devices in the presence of significant crosstalk. They present data to demonstrate accuracy of their results, including example runs of a test generator that uses these results. Wei-Yu Chen, Sandeep Gupta 0001, Melvin A. Breuer |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2000 | Test generation for crosstalk-induced faults: framework and computational resultabstractDue to technology scaling and increasing clock frequency, problems due to noise effects are leading to an increase in design/debugging efforts and a decrease in circuit performance. This paper addresses the problem of efficiently and accurately generating two-vector tests for crosstalk-induced effects, such as pulses, signal speedup and slowdown, in digital combinational circuits. We have developed a mixed-signal test generator, called XGEN, that incorporates classical static values as well as dynamic signals such as transitions and pulses, and timing information such as signal arrival times, rise/fall times and gate delay. In this paper, we first discuss the general framework of the test generation algorithm followed by computational results. A comparison of our results with SPICE simulations confirms the accuracy of this approach. Wei-Yu Chen, Sandeep Gupta 0001, Melvin A. Breuer |
Asian Test Symposium | 1 |
| 1999 | Test generation for crosstalk-induced delay in integrated circuitsabstractDue to technology scaling and increasing clock frequency, problems due to noise effects lead to an increase in design/debugging efforts and a decrease in circuit performance. This paper shows how crosstalk coupling between lines can affect the propagation delay of signals in integrated circuits. A model is presented to evaluate the effect of parasitic coupling crosstalk. Conditions for the creation of the worst-case coupling and propagation of a delayed signal are presented. A test pattern generation algorithm utilizing the above conditions is presented and applied to several example circuits. Wei-Yu Chen, Sandeep Gupta 0001, Melvin A. Breuer |
ITC | 1 |