Yi-Chang Lu

dblp:79/5056 · DBLP profile ↗
← Back
36ranked-venue papers
1as first author
9since 2021 · last 2024
0000-0002-7638-0367ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Deep Plug-and-play Nighttime Non-blind Deblurring with Saturated Pixel Handling Schemes
abstract
Due to the setting of shutter speeds, over-exposed blurry images can often be seen in nighttime photography. Although image deblurring is a classic problem in image restoration, state-of-the-art methods often fail in nighttime cases with saturated pixels. The primary reason is that those pixels are out of the sensor range and thus violate the assumption of the linear blur model. To address this issue, we propose a new nighttime non-blind deblurring algorithm with saturated pixel handling schemes, including a pixel stretching mask, an image segment mask, and a saturation awareness mechanism (SAM). Our algorithm achieves superior results by strategically adjusting mask configurations, making our method robust to various saturation levels. We formulate our task into two new optimization problems and introduce a unified framework based on the plug-and-play alternating direction method of multipliers (PnP-ADMM). We also evaluate our approach qualitatively and quantitatively to demonstrate its effectiveness. The results show that the proposed algorithm recovers sharp latent images with finer details and fewer artifacts than other state-of-the-art deblurring methods.
Hung-Yu Shu, Yi-Hsien Lin, Yi-Chang Lu
WACV3
2022 CF-Net: Complementary Fusion Network for Rotation Invariant Point Cloud Completion
abstract
Real-world point clouds usually have inconsistent orientations and often suffer from data missing issues. To solve this problem, we design a neural network, CF-Net, to address challenges in rotation invariant completion. In our network, we modify and integrate complementary operators to extract features that are robust against rotation and incompleteness. Our CF-Net can achieve competitive results both geometrically and semantically as demonstrated in this paper.
Bo-Fan Chen, Yang-Ming Yeh, Yi-Chang Lu
ICASSP3
2022 Shadow Removal Through Learning-Based Region Matching and Mapping Function Optimization
abstract
In this paper, we develop a novel shadow removal technique where the inputs are a single natural image to be restored and its corresponding shadow mask. We first decompose the image by super-pixels and cluster them into several sim-ilar regions. Then we train a random forest model to pre-dict matched pairs between shadow and non-shadow regions. By applying a distribution-based mapping function on the matched pairs, we can relight pixels in those shadow regions. An optimization framework based on half-quadratic splitting (HQS) method is also introduced to further improve the qual-ity of the mapping process. We also design a post-processing stage with a boundary inpainting function to generate bet-ter visual results. Our experiments show that the proposed method can remove shadows effectively and produce high-quality shadow-free images.
Shih-Wei Hsieh, Chih-Hsiang Yang, Yi-Chang Lu
ICME3
2022 An Alignment-Based Hardware Accelerator for Rapid Prediction of RNA Secondary Structures
abstract
The Zuker algorithm is a widely-adopted method for rapid prediction of RNA secondary structures. In this work, we present a new hardware architecture based on the algorithm and implement its accelerator using TSMC 40-nm technology. The proposed architecture is capable of handling more types of dangling bases when compared to the previous hardware work. Besides, with better data re-use, the idling time and memory usage of our hardware modules can be effectively reduced, and thus a higher processing speed can be achieved. Compared to the well-known software approach, RNAstructure, our 16-processing-element design has an $11\times$ speedup for RNA sequences around 3000 bases.
Shih-Shiuan Weng, Yang-Ming Yeh, Yi-Chang Lu
ISCAS4
2022 MSRCall: a multi-scale deep neural network to basecall Oxford Nanopore sequences
abstract
MOTIVATION: MinION, a third-generation sequencer from Oxford Nanopore Technologies, is a portable device that can provide long-nucleotide read data in real-time. It primarily aims to deduce the makeup of nucleotide sequences from the ionic current signals generated when passing DNA/RNA fragments through nanopores charged with a voltage difference. To determine nucleotides from measured signals, a translation process known as basecalling is required. However, compared to NGS basecallers, the calling accuracy of MinION still needs to be improved. RESULTS: In this work, a simple but powerful neural network architecture called multi-scale recurrent caller (MSRCall) is proposed. MSRCall comprises a multi-scale structure, recurrent layers, a fusion block and a connectionist temporal classification decoder. To better identify both short-and long-range dependencies, the recurrent layer is redesigned to capture various time-scale features with a multi-scale structure. The results show that MSRCall outperforms other basecallers in terms of both read and consensus accuracies. AVAILABILITY AND IMPLEMENTATION: MSRCall is available at: https://github.com/d05943006/MSRCall. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yang-Ming Yeh, Yi-Chang Lu
Bioinform.2
2022 Low-Light Enhancement Using a Plug-and-Play Retinex Model With Shrinkage Mapping for Illumination Estimation
abstract
Low-light photography conditions degrade image quality. This study proposes a novel Retinex-based low-light enhancement method to correctly decompose an input image into reflectance and illumination. Subsequently, we can improve the viewing experience by adjusting the illumination using intensity and contrast enhancement. Because image decomposition is a highly ill-posed problem, constraints must be properly imposed on the optimization framework. To meet the criteria of ideal Retinex decomposition, we design a nonconvexLpnorm and apply shrinkage mapping to the illumination layer. In addition, edge-preserving filters are introduced using the plug-and-play technique to improve illumination. Pixel-wise weights based on variance and image gradients are adopted to suppress noise and preserve details in the reflectance layer. We choose the alternating direction method of multipliers (ADMM) to solve the problem efficiently. Experimental results on several challenging low-light datasets show that our proposed method can more effectively enhance image brightness as compared with state-of-the-art methods. In addition to subjective observations, the proposed method also achieved competitive performance in objective image quality assessments.
Yi-Hsien Lin, Yi-Chang Lu
IEEE Trans. Image Process.2
2021 Power Reduction of a Set-Associative Instruction Cache Using a Dynamic Early Tag Lookup
abstract
An energy-efficient instruction cache lookup technique with low area overheads is proposed. The key concept of this Dynamic Early Tag Lookup (DETL) method is to exploit the presence of instruction fetch-bubble cycles. In a fetch-bubble cycle, the index of the matching cache set can be determined earlier. Hence, the dynamic energy for parallel memory accesses to irrelevant cache banks can be saved. We implemented the proposed DETL algorithm on a 4-way set-associative instruction cache in a RISC-V micro-architecture, and tested its performance using the SPEC CPU2006 benchmark suite. The experiment results showed a 19.38% dynamic power reduction with an area overhead smaller than 0.1 %.
Chun-Chang Yu, Yu Hen Hu, Yi-Chang Lu, Charlie Chung-Ping Chen
DATE3
2021 A Memory-Efficient Accelerator for DNA Sequence Alignment with Two-Piece Affine Gap Tracebacks
abstract
Previously, dynamic-programming-based DNA sequence aligners were mostly implemented with a penalty function of the one-piece affine gap model. When aligning sequences with longer gaps, the two-piece affine gap model provides better results at the cost of memory usage, which becomes an issue especially for aligners with memory-hungry traceback capabilities. In this paper, we design a memory-efficient scheme for traceback recording with the two-piece penalty scoring, so that the aligner can be realized on an ASIC. Our design is implemented with TSMC 40nm technology, and the proposed aligner can speed up pairwise alignment by 71× compared to the CPU approach.
Jing-Ping Wu, Yi-Chien Lin, Ying-Wei Wu, Shih-Wei Hsieh, Ching-Hsuan Tai, Yi-Chang Lu
ISCAS6
2021 Using Regularity Unit As Guidance For Summarization-Based Image Resizing
abstract
In this paper, we propose a novel algorithm for summarization-based image resizing. In the past, a process of detecting precise locations of repeating patterns is required before the pattern removal step in resizing. However, it is difficult to find repeating patterns which are illuminated under different lighting conditions and viewed from different perspectives. To solve the problem, we first identify the regularity unit of repeating patterns by statistics. Then we can use the regularity unit for shift-map optimization to obtain a better resized image. The experimental results show that our method is competitive with other well-known methods.
Fang-Tsung Hsiao, Yi-Hsien Lin, Yi-Chang Lu
VCIP3
2020 Comprehensive Study of Keywords for Sequence-Based Automatic Annotation of Protein Functions
abstract
Homology-based transfer is frequently used to predict protein functions of unannotated sequences through similarity analysis between the target and previously annotated sequences. The most direct and accessible homology-based transfer approach is sequence alignment. To assess the reliability of alignment-based prediction, we applied a 10-fold cross-validation test in SWISS-Prot database. We compared Matthews correlation coefficient, sensitivity, as well as precision, and examined different parameter settings used in the alignment-based methods, with BLASTp and PSI-BLAST. As the results shown in this paper, in the categories of domain, ligand, molecular function, biological process, cellular component, and PTM, the keywords can be confidently used for protein function predictions, whereas the others are less reliable.
Mao-Jan Lin, Xiao-Xuan Huang, Chien-Yu Chen 0001, Yi-Chang Lu
BIBE5
2020 An Image Deblurring Processor for Chromatic Aberration Based on the Primal-Dual Algorithm with Cross-Channel Prior
abstract
Modern camera technology adopts complicated lens systems to compensate for chromatic aberration. To reduce the cost and complexity of camera lenses, an option is to use a simple lens solution and recover blurry images computationally. However, deblurring algorithms often have low efficiency in time and consume a lot of power. Therefore, using the advantages of hardware computing, we have designed an image deblurring processor for the FlexISP algorithm. In the proposed hardware, two channels can be deblurred concurrently. We implemented the chip with TSMC 40 nm technology. The speed is 4.87 times faster than that of the software version, and it consumes only 118 mW when operating at 200 MHz. In addition, we propose a splitting method that allows the processor to deblur large-sized images. This method is hardware-friendly since it requires less memory when implemented with ASIC.
Chia-Han Huang, Yi-Chang Lu
ISCAS2
2020 HDR Deghosting Using Motion-Registration-Free Fusion in the Luminance Gradient Domain
abstract
For most of the existing high dynamic range (HDR) deghosting flows, they require a time-consuming motion registration step to generate ghost-free HDR results. Since the motion registration step usually becomes the bottleneck of the entire flow, in this paper, we propose a novel H DR deghosting flow which does not require any motion registration process. By taking channel properties into account, the luminance and chrominance channels are fused differently in the proposed flow. Our motion-registration-free fusion could generate high-quality HDR results swiftly even if the original Low Dynamic Range (LDR) images contain objects with large foreground motions.
Cheng-Yeh Liou, Cheng-Yen Chuang, Chia-Han Huang, Yi-Chang Lu
VCIP4
2019 Banded Pair-HMM Algorithm for DNA Variant Calling and Its Hardware Accelerator Design
abstract
In this paper, we propose a new pair hidden Markov model (Pair-HMM) algorithm, namely Banded Pair-HMM, which is a heuristic approach for variant calling applications. When compared to the conventional Pair-HMM, our Banded Pair-HMM can reduce the execution time at a minor cost in accuracy. In addition, a hardware accelerator is implemented using TSMC 40nm technology based on the proposed algorithm. As demonstrated later in the paper, the proposed hardware accelerator runs 4× faster than the conventional Pair-HMM hardware, and over 17,000× faster than the original Pair-HMM software.
Ming-Hung Chen, Mao-Jan Lin, Yi-Chang Lu
BIBE4
2019 A Special-Purpose Processor for FFT-Based Digital Refocusing using 4-D Light Field Data
abstract
In this paper, we use a pinhole-array-mask light field camera to generate duplicated spectra in the frequency domain. Based on Fourier-Slicing Theorem, we can collect light field information on different image planes by taking different data-slices in the frequency domain. When a series of sliced data is converted back to the spatial domain, digital-refocusing can be performed. Since the digital refocusing process is time-consuming, we design a special-purpose processor for acceleration. The chip is synthesized with TSMC 90nm technology. A 15X speedup can be achieved when compared to the software version.
Man-Rong Chen, Hao-Wei Liu, Yi-Hsien Lin, Yi-Chang Lu
ISCAS4
2018 Adaptively Banded Smith-Waterman Algorithm for Long Reads and Its Hardware Accelerator
abstract
In this paper, we propose hardware-compatible Adaptively Banded Smith-Waterman algorithm (ABSW) to align long genomic sequences. By utilizing banded Smith-Waterman algorithm to align subsequences of fixed lengths, ABSW finds alignment of a pair of arbitrarily long sequences with constant memory. In addition, a heuristic algorithm, dynamic overlapping, is proposed to make overlaps of bands of subsequences to improve accuracy. To enable hardware acceleration of ABSW, we further propose the hardware architecture of banded Smith-Waterman with traceback. Experiments show that ABSW produces near optimal alignment scores for sequences with up to 40% error rates. Our hardware implementation of ABSW demonstrates more than 200× speedup over software implementation.
Yi-Lun Liao, Nae-Chyun Chen, Yi-Chang Lu
ASAP4
2018 Subpixel-Level-Accurate Algorithm for Removing Double-Layered Reflections from a Single Image
abstract
When photographers take pictures through glasses, image quality is sometimes degraded by window reflections. In this paper, we propose an algorithm to enhance the effectiveness of removing double-layered reflections from a single image. First, we apply subpixel-level autocorrelation to calculate the shift kernel of reflections. Secondly, a Gaussian-mixture-model (GMM) based layer separation technique is applied to estimate the initial transmission and reflection layers. In the refinement stage, we propose to add another GMM optimization flow to obtain the accurate transmission layer. Finally, a color and brightness adjustment step in Lab domain is used to provide better image quality. The qualitative and quantitative results show that the proposed method is very robust and effective.
Shih-Wei Hsieh, Yao-Cheng Yang, Chi-Ming Yeh, Sheng-Jui Huang, Yi-Chang Lu
ICIP5
2018 A Memory-Efficient FM-Index Constructor for Next-Generation Sequencing Applications on FPGAs
abstract
FM-index is an efficient data structure for string search and is widely used in next-generation sequencing (NGS) applications such as sequence alignment and de novo assembly. Recently, FM-indexing is even performed down to the read level, raising a demand of an efficient algorithm for FM-index construction. In this work, we propose a hardware-compatible Self-Aided Incremental Indexing (SAII) algorithm and its hard-ware architecture. This novel algorithm builds FM-index with no memory overhead, and the hardware system for realizing the algorithm can be very compact. Parallel architecture and a special prefetch controller is designed to enhance computational efficiency. An SAII-based FM-index constructor is implemented on an Altera Stratix V FPGA board. The presented constructor can support DNA sequences of sizes up to 131,072-bp, which is enough for small-scale references and reads obtained from current major platforms. Because the proposed constructor needs very few hardware resource, it can be easily integrated into different hardware accelerators designed for FM-index-based applications.
Nae-Chyun Chen, Yi-Chang Lu
ISCAS3
2018 A High Dynamic Range Light Field Camera and Its Built-In Data Processor Design
abstract
Traditional High Dynamic Range (HDR) techniques require inputs of multiple Low Dynamic Range (LDR) images with different exposure time taken in serial shots or from a camera array. These methods are either suffered from ghosting effects or very expensive when setting up the system. In this paper, we design a hand-held pinhole-mask-based light field camera and use it to capture all LDR data in one shot. We also propose a light-field HDR algorithm to process LDR data collected. Finally, a hardware accelerator for the algorithm is implemented to complete calculations within 0.3 seconds.
Po-Hsiang Hsu, Yang-Ming Yeh, Chi-Ming Yeh, Yi-Chang Lu
ISCAS4
2018 A Hybrid Flow for Multiple Sequence Alignment with a BLASTn Based Pairwise Alignment Processor
abstract
In this paper, we design a special-purpose processor for pairwise alignment and propose to integrate this design into a Multiple Sequence Alignment (MSA) flow. The processor is based on a modified algorithm of Banded Two-Hit Basic Local Alignment Search Tool for Nucleotides (MB2-BLASTn). In the new hybrid MSA flow, our MB2-BLASTn processor is used to find pairwise alignment first. Then the results are sent to CPUs for tree-building and progressive alignment. Since the pairwise alignment takes about 90% of total computing time in the original software flow, 100X speedup achieved by MB2-BLASTn can greatly improve the speed of MSA.
Mao-Jan Lin, Chih-Yu Chang, Nae-Chyun Chen, Yi-Chang Lu
ISCAS5
2017 Deep Co-occurrence Feature Learning for Visual Object Recognition
abstract
This paper addresses three issues in integrating part-based representations into convolutional neural networks (CNNs) for object recognition. First, most part-based models rely on a few pre-specified object parts. However, the optimal object parts for recognition often vary from category to category. Second, acquiring training data with part-level annotation is labor-intensive. Third, modeling spatial relationships between parts in CNNs often involves an exhaustive search of part templates over multiple network streams. We tackle the three issues by introducing a new network layer, called co-occurrence layer. It can extend a convolutional layer to encode the co-occurrence between the visual parts detected by the numerous neurons, instead of a few pre-specified parts. To this end, the feature maps serve as both filters and images, and mutual correlation filtering is conducted between them. The co-occurrence layer is end-to-end trainable. The resultant co-occurrence features are rotation-and translation-invariant, and are robust to object deformation. By applying this new layer to the VGG-16 and ResNet-152, we achieve the recognition rates of 83.6% and 85.8% on the Caltech-UCSD bird benchmark, respectively. The source code is available at https://github.com/yafangshih/Deep-COOC.
Ya-Fang Shih, Yang-Ming Yeh, Yen-Yu Lin, Ming-Fang Weng, Yi-Chang Lu, Yung-Yu Chuang
CVPR5
2014 A pixel-based depth estimation algorithm and its hardware implementation for 4-D light field data
abstract
In this paper, we propose a hardware-compatible depth estimation algorithm for 4-D light field data. Compared to previous sub-image-based approaches, the proposed pixel-based method can provide depth maps with higher resolution since multiple disparity numbers can be assigned to different objects sharing the same sub-image. We also implement a hardware accelerator based on the proposed algorithm using TSMC 90 nm technology. With the adoption of modified raster scan, new data input sequence, and pipelining technology, the ASIC can operate up to 100 MHz with the power consumption of 105.7 mW. It takes only 0.02954 sec to complete the depth estimation procedure, where 270X speed-up is achieved when compared to the C++ version.
Man-Rong Chen, Po-Hsiang Hsu, Yi-Chang Lu
ISCAS4
2014 Testing of TSV-Induced Small Delay Faults for 3-D Integrated Circuits
abstract
Through silicon via (TSV) is a widely used interconnect technology in 3-D integrated circuits. This paper shows that defective TSVs can induce small delay faults in surrounding logic gates. We present simulation results of TSV-induced small delay fault (TSDF) because of mechanical stress or pinhole leakage. A test technique is proposed to detect TSDF using a physical-aware fault extractor and timing-aware automatic test pattern generation. This technique requires no DfT area overhead and no direct TSV probing. Experimental results on benchmark circuits show that test coverage can be improved by 22% and 10% for stress-induced and leakage-induced TSDF, respectively. In our results, the test length overheads of both TSDFs are <; 5%.
Chun-Yi Kuo, Chi-Jih Shih, Yi-Chang Lu, Chien-Mo James Li, Krishnendu Chakrabarty
IEEE Trans. Very Large Scale Integr. Syst.3
2013 Power distribution network modeling for 3-D ICs with TSV arrays
abstract
A coupling node insertion method (CNIM) is proposed to handle electrical coupling between top metals of on-chip interconnects and silicon substrate surfaces in three-dimensional integrated circuits (3-D ICs). This coupling effect should not be neglected especially as metal area is intentionally increased in order to reduce resistance values. In this paper, we illustrate how to build the CNIM model and incorporate it into power distribution networks. The CNIM model is validated by comparing our results to the one obtained from a full-wave simulator. The differences between two approaches are within 5% but our computation time is shorter than that required by a full-wave simulator.
Chi-Kai Shen, Yi-Chang Lu, Yih-Peng Chiou, Tai-Yu Cheng, Tzong-Lin Wu
ASP-DAC2
2013 Test Generation of Path Delay Faults Induced by Defects in Power TSV
abstract
This paper presents a novel test generation technique for defective power TSV induced path delay faults in 3D IC. This paper provides a simple close-form analysis to show that, in a regular 3D power grid model, open defects in power TSV do not induce serious IR drop. However, leakage defects in power TSV should be tested, even though the number of power TSV is large. This paper proposes a test generation flow to detect path delay faults induced by defective power TSV. The proposed technique is demonstrated on an 18-tier, 7 x 7 multi-core 3D IC model. In the experiment of b18 and b19 benchmark circuits, all detectable path delay faults induced by power TSV can be tested by around hundred test patterns. This technique requires no extra DfT hardware overhead.
Chi-Jih Shih, Shih-An Hsieh, Yi-Chang Lu, Chien-Mo James Li, Tzong-Lin Wu, Krishnendu Chakrabarty
Asian Test Symposium3
2013 Parallel architecture and hardware implementation of pre-processor and post-processor for sequence assembly
abstract
In current DNA sequence assembly flows, a time-consuming pre-processing routine is usually executed to remove low quality data for smaller problem sizes and better final results. The step usually takes 25%~90% of the total computation time. To speed up the process, based on the characteristics of DNA data, we propose to use digital processing hardware, with parallel and pipeline capabilities, to accelerate the pre-processing step. In our FPGA implementation, the computing speed can be improved up to 2,700 folds. With similar design concepts, we also implement a post-processor to convert assembly results from color space to base space whenever the step is necessary. Both designs enable faster sequence assembly, which will greatly benefit the research in genomics and related fields.
Yuan-Hsiang Kuo, Chun-Shen Liu, Yi-Chang Lu
ICASSP4
2013 Exploring synergistic DVFS control of cores and DRAMs for thermal efficiency in CMPs with 3D-stacked DRAMs
abstract
The three-dimensional (3D) integration technology that utilizes low-latency and high-density Through-Silicon Vias (TSVs) to integrate DRAMs and Chip-Multiprocessors(CMPs) in the third dimension has been demonstrated as a promising way to mitigate the memory wall problem in CMPs. In addition to the improved interconnection performance and heterogeneous integration, the 3D IC technology also provides the advantages of high packaging density and small chip area. However, the power density of 3D ICs increases with the number of active devices. Therefore, alleviating the thermal stress issue of 3D ICs is one of the major design challenges.
Ping-Sheng Lin, Yi-Jung Chen, Chia-Lin Yang, Yi-Chang Lu
ISLPED4
2013 LineDiff Entropy: Lossless Layout Data Compression Scheme for Maskless Lithography Systems
abstract
This letter presents a new lossless electron-beam layout data compression and decompression algorithm named LineDiff Entropy. The algorithm is designed to facilitate high volume data transfer and massively-parallel decompression in electron-beam direct-write lithography systems. LineDiff Entropy first compares consecutive electron-beam data scanlines and encodes the data based on change/no-change of pixel values and length of corresponding sequences. Then LineDiff Entropy utilizes the entropy encoding technique to assign unique short codes to data of frequent occurrence. Because the code format is simple and effective, LineDiff Entropy decompression can be achieved with limited computing resources. The benchmark results show that LineDiff Entropy is capable of achieving excellent compression factors and very fast decompression speed.
Chin-Khai Tang, Ming-Shing Su, Yi-Chang Lu
IEEE Signal Process. Lett.3
2011 Thermal Modeling and Analysis for 3-D ICs With Integrated Microchannel Cooling
abstract
Integrated microchannel liquid-cooling technology is envisioned as a viable solution to alleviate an increasing thermal stress imposed by 3-D stacked ICs. Thermal modeling for microchannel cooling is challenging due to its complicated thermal-wake effect, a localized temperature wake phenomenon downstream of a heated source in the flow. This paper presents a fast and accurate thermal-wake aware thermal model for integrated microchannel 3-D ICs. A combination of the microchannel thermal-wake function and the channel merging technique achieves more than 3300× speedup with less than 5% error in comparison with a commercial numerical finite volume simulation tool. With the proposed model, we characterize thermal behaviors of microchannel-cooled 3-D ICs and compare them with the case of conventional air-cooled 3-D ICs. We also demonstrate thermal-aware placements using our thermal model. It shows that the proposed model can be used to reduce peak temperatures, which is considered important for 3-D IC designs.
Hitoshi Mizunuma, Yi-Chang Lu, Chia-Lin Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2010 A new method to improve accuracy of parasitics extraction considering sub-wavelength lithography effects
abstract
Modern nanometer integrated circuits are patterned by sub-wavelength lithography with significant shape deviation from drawn layouts. Full-chip parasitics extraction faces new challenges since shape distortions such as line end rounding and corner rounding cannot be accurately characterized by existing layout parameter extraction (LPE) techniques which assume perfect polygons. A new LPE method and efficient shape approximation algorithms are proposed to account for the shape distortions. Preliminary results verified by field solver simulations indicate that accuracy of parasitics extraction can be significantly improved.
Kuen-Yu Tsai, Wei-Jhih Hsieh, Yuan-Ching Lu, Bo-Sen Chang, Sheng-Wei Chien, Yi-Chang Lu
ASP-DAC6
2010 Light field based digital refocusing using a DSLR camera with a pinhole array mask
abstract
In this paper, a computational photography system is utilized to sample 4D light fields. The system is implemented using a normal DSLR camera with a mask printed using Kodak LVT technique. We reconfigure the camera by inserting a pinhole array mask in front of the sensors. The mask blocks part of light and samples the rest that passes through the pinholes. With these 4D light field data, the refocused images are obtained by rearranging the captured sub-images. The range of refocusing is also studied to avoid diffraction blurring effect in the refocused images.
Chih-Chieh Chen, Yi-Chang Lu, Ming-Shing Su
ICASSP2
2010 Depth estimation of light field data from pinhole-masked DSLR cameras
abstract
In this paper, depth estimation of light field data obtained from the pinhole-masked camera system is proposed. In the first approach, the sharpness of objects in the refocused images is evaluated using the FFT method, from which an empirical formula can be obtained to estimate the depth of objects. The second method uses the raw data in sub-images to calculate an index, Average Absolute Difference (AAD), for depth estimation purposes. The advantages and disadvantages of two proposed approaches are discussed and compared.
Chih-Chieh Chen, Shih-Chieh Fan Chiang, Xiao-Xuan Huang, Ming-Shing Su, Yi-Chang Lu
ICIP5
2009 Thermal modeling for 3D-ICs with integrated microchannel cooling
abstract
Integrated microchannel liquid-cooling technology is envisioned as a viable solution to alleviate an increasing thermal stress imposed by 3D stacked ICs. Thermal modeling for microchannel cooling is challenging due to its complicated thermal-wake effect, a localized temperature wake phenomenon downstream of a heated source in the flow. This paper presents a fast and accurate thermal-wake aware thermal model for integrated microchannel 3D ICs. Validation results show the proposed thermal model achieves more than 400x speed up and only 2.0% error in comparison with a commercial numerical simulation tool. We also demonstrate the use of the proposed thermal model for thermal optimization during the IC placement stage. We find that due to the thermal-wake effect, tiles are placed in the descending order of power magnitude along the flow direction. We also find that modeling thermal-wakes is critical for generating a thermal-aware placement for integrated microchannel-cooled 3D IC. It could result in up to 25°C peak temperature difference according to our experiments.
Hitoshi Mizunuma, Chia-Lin Yang, Yi-Chang Lu
ICCAD3
2008 A new method to improve accuracy of leakage current estimation for transistors with non-rectangular gates due to sub-wavelength lithography effects
abstract
Non-ideal pattern transfer from drawn circuit layout to manufactured nanometer transistors can severely affect electrical characteristics such as drive current, leakage current, and threshold voltage. Obtaining accurate electrical models of non-rectangular transistors due to sub-wavelength lithography effects is indispensable for DFM-aware nanometer IC design. In this paper, TCAD device simulations are utilized to quantify the accuracy of a standard equivalent gate length extraction approach for non-rectangular transistors. It is verified that threshold voltage and current density are non-uniform along the channel width due to narrow-width related edge effects, leading to significant inaccuracy in the sub-threshold region. A new EGL extraction method utilizing location-dependent weighting factors and convex parameter extraction techniques is proposed to account for the current density non-uniformity. Preliminary results verified by TCAD simulations indicate that the accuracy of leakage current estimation for non-rectangular transistors can be significantly improved. The method is readily applicable to calibration with real silicon data.
Kuen-Yu Tsai, Meng-Fu You, Yi-Chang Lu, Philip C. W. Ng
ICCAD3
2007 Performance Benefits of Monolithically Stacked 3-D FPGA
abstract
The performance benefits of a monolithically stacked three-dimensional (3-D) field-programmable gate array (FPGA), whereby the programming overhead of an FPGA is stacked on top of a standard CMOS layer containing logic blocks (LBs) and interconnects, are investigated. A Virtex-II-style two-dimensional (2-D) FPGA fabric is used as a baseline architecture to quantify the relative improvements in logic density, delay, and power consumption achieved by such a 3-D FPGA. It is assumed that only the switch transistor and configuration memory cells can be moved to the top layers and that the 3-D FPGA employs the same LB and programmable interconnect architecture as the baseline 2-D FPGA. Assuming they are les 0.7, the area of a static random-access memory cell and switch transistors having the same characteristics as n-channel metal-oxide-semiconductor devices in the CMOS layer are used. It is shown that a monolithically stacked 3-D FPGA can achieve 3.2 times higher logic density, 1.7 times lower critical path delay, and 1.7 times lower total dynamic power consumption than the baseline 2-D FPGA fabricated in the same 65-nm technology node
Mingjie Lin, Abbas El Gamal, Yi-Chang Lu, S. Simon Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2006 Performance benefits of monolithically stacked 3D-FPGA
abstract
The performance benefits of a monolithically stacked 3D-FPGA, whereby the programming overhead of an FPGA is stacked on top of a standard CMOS layer containing the logic blocks and interconnects, are investigated. A Virtex-II style 2D-FPGA fabric is used as a baseline for quantifying the relative improvements in logic density, delay, and power consumption achieved by such a 3D-FPGA. It is assumed that only the pass-transistor switches and configuration memory cells can be moved to the top layers and that the 3D-FPGA employs the same logic block and programmable interconnect architecture as the baseline 2D-FPGA. Assuming a configuration memory cell that is ≤ 0.7 the area of an SRAM cell and pass-transistor switches having the same characteristics as nMOS devices in the CMOS layer are used, it is shown that a monolithically stacked 3D-FPGA can achieve 3.2 times higher logic density, 1.7 times lower critical path delay, and 1.7 times lower total dynamic power consumption than the baseline 2D-FPGA fabricated in the same 65nm technology node.
Mingjie Lin, Abbas El Gamal, Yi-Chang Lu, S. Simon Wong
FPGA3
2001 Min/max On-Chip Inductance Models and Delay Metrics
abstract
This paper proposes an analytical inductance extraction model for characterizing min/max values of typical on-chip global intercon-nect structures, and a corresponding delay metric that can be used to provide RLC delay prediction from physical geometries. The model extraction and analysis is efficient enough to be used within optimization and physical design exploration loops. The analytical min/max inductance approximations also provide insight into the effects caused by inductances.
Yi-Chang Lu, Mustafa Celik, Tak Young, Lawrence T. Pileggi
DAC1