Oliver Chiu-sing Choy

dblp:c/OliverChiusingChoy · also Chiu-Sing Choy, Chiu-sing Choy, Oliver C. S. Choy · DBLP profile ↗
← Back
40ranked-venue papers
4as first author
0since 2021 · last 2018
0000-0002-8370-3144ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 35 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Integrated circuit design · 87% Electronic design automation · 8% Hardware reliability and fault tolerance · 4%
Computer networks
1 paper
Internet of things and sensor networks · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Integrated circuit design
digital circuit design
0.422018
A Subthreshold Baseband Processor Core Design With Custom Modules and Cells for Passive RFID Tags · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
A comparison of via-programmable gate array logic cell circuits · FPGA 2009
Integrated circuit design › low-power circuit design
subthreshold circuit design
0.312018
A Subthreshold Baseband Processor Core Design With Custom Modules and Cells for Passive RFID Tags · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Internet of things and sensor networks › RFID systems › passive RFID
passive RFID tags
0.112018
A Subthreshold Baseband Processor Core Design With Custom Modules and Cells for Passive RFID Tags · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Integrated circuit design
low-power circuit design
0.112009
A comparison of via-programmable gate array logic cell circuits · FPGA 2009
Integrated circuit design
asynchronous circuit design
0.122004
A high-efficiency strongly self-checking asynchronous datapath · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
A New Control Circuit for Asynchronous Micropipelines · IEEE Trans. Computers 2001
Hardware reliability and fault tolerance
self-checking circuits
0.012004
A high-efficiency strongly self-checking asynchronous datapath · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2004
Integrated circuit design › asynchronous circuit design
micropipeline
0.012001
A New Control Circuit for Asynchronous Micropipelines · IEEE Trans. Computers 2001
Electronic design automation › hardware verification and test
design for testability
0.011996
Test Generation with Dynamic Probe Points in High Observability Testing Environment · IEEE Trans. Computers 1996
Electronic design automation › physical design › placement
detailed placement
0.011996
Incremental layout placement modification algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996
Electronic design automation
hardware verification and test
0.011996
Test Generation with Dynamic Probe Points in High Observability Testing Environment · IEEE Trans. Computers 1996
Electronic design automation
physical design
0.011996
Incremental layout placement modification algorithms · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1996
Electronic design automation › hardware verification and test
test generation
0.011996
Test Generation with Dynamic Probe Points in High Observability Testing Environment · IEEE Trans. Computers 1996
Electronic design automation › hardware verification and test › design for testability
test point insertion
0.011996
Test Generation with Dynamic Probe Points in High Observability Testing Environment · IEEE Trans. Computers 1996

Methods — techniques the papers use, named apart from their topics

ratioed logic · 0.7double-edge-triggered flip-flop · 0.7custom logic cell design · 0.7process variation analysis · 0.1circuit comparison · 0.1self-checking design · 0.0differential cascode voltage switch logic · 0.0incremental algorithms · 0.0
YearPublicationVenuePosition
2018 A Subthreshold Baseband Processor Core Design With Custom Modules and Cells for Passive RFID Tags
abstract
Sophisticated subthreshold passive radio frequency identification tag's baseband processor (BBP) core design for ultralow-power Internet of Things end devices is presented in this paper. Custom logic cells and tailored logic architectures are applied to eliminate timing violations when the operating voltage is much lower than nominal level. For the consideration of limited availability of radio frequency power, power-aware scheme is applied to the key modules, including PIE decoding and command receiving. Furthermore, Galois linear feedback shift register and double-edge-triggered techniques help to improve clock efficiency and reduce the impact of frequency variation in data link portions. Importantly, a novel custom ratioed logic style is adopted in key modules to fundamentally speed up signals' propagation at ultralow-voltage. The proposed BBP was fabricated in 90-nm CMOS as well as the regular design with the same function. It was also implemented in the tag chip's fabrication. In measurement the proposed design indicates good robustness and is much more competent for subthreshold operation. It can operate below 0.3 V with power consumption below 130 nW.
Weiwei Shi 0001, An Pan, Oliver Chiu-sing Choy
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2016 Cascaded Network Body Channel Model for Intrabody Communication
abstract
Intrabody communication has been of great research interest in recent years. This paper proposes a novel, compact but accurate body transmission channel model based on RC distribution networks and transmission line theory. The comparison between simulation and measurement results indicates that the proposed approach accurately models the body channel characteristics. In addition, the impedance-matching networks at the transmitter output and the receiver input further maximize the power transferred to the receiver, relax the receiver complexity, and increase the transmission performance. Based on the simulation results, the power gain can be increased by up to 16 dB after matching. A binary phase-shift keying modulation scheme is also used to evaluate the bit-error-rate improvement.
Hao Wang 0030, Xian Tang, Oliver Chiu-sing Choy, Gerald E. Sobelman
IEEE J. Biomed. Health Informatics3
2015 A 5.4-mW 180-cm Transmission Distance 2.5-Mb/s Advanced Techniques-Based Novel Intrabody Communication Receiver Analog Front End
abstract
This paper presents a low power, long-transmission distance, high data rate intrabody communication (IBC) analog receiver front end (RFE). First, to optimize the transmission performance, conventional transmission line analysis scheme is creatively adopted to the IBC design to characterize the body channel. Second, switched-capacitor filters based on sampling rate boosting technique are adopted for higher accuracy and lower power consumption. Third, a novel RFE topology is proposed to further enhance the IBC performance. The new RFE is designed and fabricated in a standard 180-nm CMOS process. Measurement results show that the RFE can successfully transmit data spanning the whole human body, around 180 cm, which is one of the longest transmission distances reported in related literatures. Furthermore, it reaches a maximum data rate of 2.5 Mb/s with a bit error rate less than 1e-7 and consumes 5.4 mW from a 1.8 V supply. The proposed RFE compares favorably to similar reported works.
Hao Wang 0030, Xian Tang, Oliver Chiu-sing Choy, Ka Nang Leung, Kong-Pang Pun
IEEE Trans. Very Large Scale Integr. Syst.3
2014 Reducing pin count on cross-referencing Digital Microfluidic Biochip
abstract
Digital Microfluidic Biochip(DMFB) allows traditional laboratory procedures to be conducted autonomously on a small chip. On a DMFB, the number of control pins is a limiting factor on maximum chip size and a major factor affecting the manufacturing cost. We have developed a methodology that reduces pin count in cross-referencing DMFB. Our algorithm simultaneously optimizing routing and control pin assignment. Experiments show that the proposed scheme can reduce pin-count by 23% to 32% with minimal effect on routing time.
Ho Chuen Jackson Yeung, Evangeline F. Y. Young, Oliver Chiu-sing Choy
ISCAS3
2013 Architecture and Design Flow for a Highly Efficient Structured ASIC
abstract
As fabrication process technology continues to advance, mask set costs have become prohibitively expensive. Structured application specific integrated circuits (sASICs) offer a middle ground in price and performance between ASICs and field-programmable gate arrays (FPGAs) by sharing masks across different designs. In this paper, two sASIC architectures are proposed, the first being based on three-input lookup-tables, and the second on AOI22 gates. The sASICs are programmed using a standard-cell compatible design flow. They are customized using a minimum of three masks, i.e., two metals and one via. The area and delay of the sASIC are compared with ASICs and FPGAs. Results over a set of benchmark circuits show that our AOI22-based sASIC had an average of 1.76x/1.41x increase in area/delay compared to ASICs, a considerable improvement compared with the 26.56x/5.09x increase for FPGAs. This is, to the best of our knowledge, the best performance reported in the literature for a practical sASIC. A prototype using the sASIC was fabricated using a universal machine control 0.13-μm mixed-mode/RF process. It was fully verified using scan and functional tests, and used in a demonstration system.
S. Man Ho Ho, Yanqing Ai, Thomas C. P. Chau, Steve C. L. Yuen, Oliver Chiu-sing Choy, Philip H. W. Leong, Kong-Pang Pun
IEEE Trans. Very Large Scale Integr. Syst.5
2011 Robust and efficient baseband receiver design for MB-OFDM UWB system
abstract
Robust, efficient and low complexity design methodologies for high speed multi-band orthogonal frequency division multiplexing ultra-wideband (MB-OFDM UWB) is presented. The proposed design is implemented in 0.13μm CMOS technology with the core area of 2.66mm × 0.94mm. Operating at 132MHz clock frequency, the estimated power consumption is 170mW.
Oliver Chiu-sing Choy
ASP-DAC2
2010 Rapid prototyping on a structured ASIC fabric
abstract
We describe the architecture of a structured ASIC fabric in which the logic and routing can be customized using three masks. A standard Cadence based design flow is employed, and using an active dynamic backlight controller as an example, performance is compared to that of an ASIC implementation in the same technology.
Steve C. L. Yuen, Yanqing Ai, Brian P. W. Chan, Thomas C. P. Chau, Sam M. H. Ho, Oscar K. L. Lau, Kong-Pang Pun, Philip H. W. Leong, Oliver Chiu-sing Choy
ASP-DAC9
2010 Design of a single layer programmable Structured ASIC library
abstract
A Structured Application-specific Integrated Circuit (SASIC) is a programmable fabric in which a small set of masks are customized for a particular application, serving to reduce the associated non-recurring engineering cost (NRE). In this paper we describe the implementation of a SASIC logic cell which is programmable via a single metal layer. A SASIC fabric prototype is fabricated and all implemented functions are verified on silicon. Experimental measurement verifies correct operation of our SASIC with a clock frequency of over 250 MHz.
Thomas C. P. Chau, David W. L. Wu, Yanqing Ai, Brian P. W. Chan, Sam M. H. Ho, Oscar K. L. Lau, Steve C. L. Yuen, Kong-Pang Pun, Oliver Chiu-sing Choy, Philip H. W. Leong
DDECS9
2010 Structured ASIC: Methodology and comparison
abstract
As fabrication process technology continues to advance, mask set costs have become prohibitively expensive. Structured ASICs can offer price and performance between ASICs and FPGAs. They are attractive for mid-volume production and offer good intellectual property security. In this paper, a structured ASIC methodology, where 2 metal- and 1 via-mask are customised, is described. The CAD tools are fully compatible with conventional ASIC design flows and a comparison of area and delay performance with ASICs and FPGAs is given. A prototype structured ASIC implementing an LED-backlit LCD controller was fabricated in a 0.13 μm CMOS process. It was verified and power consumption compared with an ASIC design.
Sam M. H. Ho, Steve C. L. Yuen, Hiu Ching Poon, Thomas C. P. Chau, Yanqing Ai, Philip H. W. Leong, Oliver Chiu-sing Choy, Kong-Pang Pun
FPT7
2010 A low-latency NoC router with lookahead bypass
abstract
Packet-switched networks on chip are emerging communication fabric to resolve the scalability and bandwidth limitation inherent in shared buses and dedicated links. However current state-of-the-art on-chip network routers suffer from latency overhead. In this work, we propose a new router which makes use of dynamic lookahead bypass to reduce latency. Special lookahead controlling pipeline is applied to speed up allocation computations so that the input buffers' bypassing rate increases. Lookahead pipeline and bypasses not only can reduce network latency but can save the energy due to writing and reading buffers. Analysis and simulation results using different traffic patterns prove that the architecture can significantly improve packet latency by up to 32.1% over a state-of-art router design and costs only a small silicon area overhead.
Ling Xin, Oliver Chiu-sing Choy
ISCAS2
2009 A comparison of via-programmable gate array logic cell circuits
abstract
Via-programmable gate arrays (VPGAs) offer a middle ground between application specific integrated circuits and field programmable gate arrays in terms of flexibility, manufactuing cost, speed, power and area. In this paper, we present a novel VPGA logic cell, the complementary universal logic gate (CULG) which can be used to implement both sequential and combinatorial elements. Its performance is compared with a number of other designs including transmission gate, differential cascode voltage switch with pass gate, and standard cell. The CULG is found to have comparable power-delay product and process variation sensitivity to the other designs while offering the lowest power consumption.
Thomas C. P. Chau, Philip H. W. Leong, Sam M. H. Ho, Brian P. W. Chan, Steve C. L. Yuen, Kong-Pang Pun, Oliver Chiu-sing Choy, Xinan Wang
FPGA7
2009 A Low-power Signal Processing Front-end and Decoder for UHF Passive RFID Transponders
abstract
In this paper, a low-power signal processing front-end and PIE decoder for use in UHF RFID passive transponders is presented. By merging with the decoder, the clock generator does not need to drive a large loading capacitor. Therefore, its power consumption can be greatly reduced. In addition, the ring oscillator of the generator was designed for low sensitivity to supply voltage variation and low power consumption. Fabricated in a 0.13-mum CMOS process, the measured power consumption consumes only 850 nW.
Chi Fat Chan, Weiwei Shi 0001, Kong-Pang Pun, Lincoln Lai Kan Leung, Ka Nang Leung, Oliver Chiu-sing Choy
ISCAS6
2009 Robust and Low Complexity Packet Detector Design for MB-OFDM UWB
abstract
Multiband orthogonal frequency division multiplexing (MB-OFDM) ultra wideband (UWB) systems have drawn much attention for its high spectrum efficiency and multiple access capability. However, its large throughput requirement and low power spectral density result in high hardware complexity and high power consumption, which are challenges of designing the packet detector. In this paper, a novel detection method is proposed with very good detection performance in low SNR. Low cost and low power schemes are also introduced in circuit design to save 70% area and 71% power. The proposed packet detector is synthesized with UMC 0.13 mum library at 132 MHz clock frequency. The hardware cost is 56.9 K gates and the power consumption is only 11.7 mW.
Oliver Chiu-sing Choy, Ka Nang Leung
ISCAS2
2009 A Novel Mismatch Cancellation and I/Q Channel Multiplexing Scheme for Quadrature Bandpass DeltaSigma Modulators
abstract
This paper investigates and resolves in-channel/quadrature channel (I/Q) imbalances in quadrature band pass delta sigma modulators. These mismatches result in image interference and noise being aliased into the desired signal band, thus degrading the dynamic range of the modulators. A novel dynamic element match shaping scheme is proposed to cancel the aliasing of image interference, noise, and self-image. The simulation results for the proposed scheme show that the SNDR is only degraded by 1.6 dB from the ideal case (without I/Q mismatch). Meanwhile, using the I/Q channel multiplexing technique, operational amplifiers, quantizers, and digital-analog converters can be shared between I/Q channels. As a result, silicon area can be reduced with the same power consumption.
Bing Li 0011, Cheong-Fat Chan, Kong-Pang Pun, Oliver Chiu-sing Choy
ISCAS4
2008 Low-Cost VC Allocator Design for Virtual Channel Wormhole Routers in Networks-on-Chip
Min Zhang 0012, Oliver Chiu-sing Choy
NOCS2
2008 A Five-Stage Pipeline, 204 Cycles/MB, Single-Port SRAM-Based Deblocking Filter for H.264/AVC
abstract
This paper describes the design and VLSI implementation of a highly efficient, single-port SRAM-based deblocking filter. It can achieve 204 cycles/macroblock throughput for H.264/AVC real-time decoding. Several deblocking filter designs in the literature have been compared and the possibility of realizing them in a pipeline is studied. Eventually we came up with a completely new design which has a five-stage pipeline with gated clock to increase system throughput while reducing power. Data hazards and structure hazards, which are the two most critical issues for a pipelined filter, are analyzed and resolved. Efficient memory organization for both on-chip SRAM and transposition buffers is employed. By using innovative hybrid edge filtering sequence and out-of-order memory update scenario, we obtain zero stall cycle in normal pipeline flow, making the best out of a pipelined architecture. Compared with existing designs, our design achieves at least 18% clock cycle reduction, as well as 20% lower power consumption owing to its efficient pipeline and memory architecture. The total gate count is comparable to other designs in literature without using any expensive two-port or dual-port on-chip SRAMs.
Ke Xu 0014, Oliver Chiu-sing Choy
IEEE Trans. Circuits Syst. Video Technol.2
2008 A Power-Efficient and Self-Adaptive Prediction Engine for H.264/AVC Decoding
abstract
Prediction, including intra prediction and inter prediction, is the most critical issue in H.264/AVC decoding in terms of processing cycles and computation complexity. These two predictions demand a huge number of memory accesses and account for up to 80% of the total decoding cycles. In this paper, we present the design and VLSI implementation of a novel power-efficient and highly self-adaptive prediction engine that utilizes a 4 times 4 block level pipeline. Based on the different prediction requirements, the prediction pipeline stages, as well as the correlated memory accesses and datapaths, are fully adjustable, which helps to reduce unnecessary decoding operations and energy dissipation while retaining the fixed real-time throughput. Compared with conventional designs, this paper has the advantage of higher efficiency and lower power consumption due to the elimination of all redundant operations and the wide employment of the pipeline and parallel processing. Under different prediction modes, our design is able to decode each macroblock within 500 cycles. A prototype H.264/AVC baseline decoder chip that utilizes the proposed prediction engine is fabricated with UMC 0.18-mu CMOS 1P6 M technology. The prediction engine contains 79 K gates and 2.8 kb single-port on-chip SRAM, and occupies half of the whole chip area. When running at 1.5 MHz for QCIF 30 f/s real-time decoding, the prediction engine dissipates 268 muW at a 1.8-V power supply.
Ke Xu 0014, Oliver Chiu-sing Choy
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Low-power H.264/AVC baseline decoder for portable applications
abstract
In this paper, we propose a low-power H.264/AVC baseline decoder. A systematic methodology for power reduction at all design levels for video decoding is proposed and applied. Power consumption is optimized at algorithm, architecture, circuit, and physical levels. The VLSI implementation results show that with UMC 180nm technology, the proposed design is able to decode QCIF 30fps at 1.5MHz. It consumes 698μW operated under 1.8V power supply. The decoder contains 169k logic gates and 2.5KB on-chip SRAM. The total chip area is 4.4x4.4mm2 in a CQFP 208 package. The low-power and real-time features make our design ideal for portable applications where video quality is often traded off for energy.
Ke Xu 0014, Oliver Chiu-sing Choy
ISLPED2
2006 A 6-digit CMOS current-mode analog-to-quaternary converter with RSD error correction algorithm
abstract
This paper presents a current-mode analog-to-quaternary (A/Q) converter using a 0.35/spl mu/m CMOS process. Redundant signed digit (RSD) technique is used to improve the resolution to 6 digits, which is equivalent to 12 binary bits. Simulations results show that the converter dissipates 382mW at 2.5V supply and 20MHz sampling rate. The converter achieves SNDR of 66.8dB, SFDR of 76.33dB and THD of -75.73dB. The effective number of bit is equal to 5.4 digits or 10.8 bits (binary).
Chi-Hong Chan, Cheong-Fat Chan, Oliver Chiu-sing Choy, Kong-Pang Pun
ISCAS3
2006 An efficient MFCC extraction method in speech recognition
abstract
This paper introduces a new algorithm of extracting MFCC for speech recognition. The new algorithm reduces the computation power by 53% compared to the conventional algorithm. Simulation results indicate the new algorithm has a recognition accuracy of 92.93%. There is only a 1.5% reduction in recognition accuracy compared to the conventional MFCC extraction algorithm, which has an accuracy of 94.43%. However, the number of logic gates required to implement the new algorithm is about half of the MFCC algorithm, which makes the new algorithm very efficient for hardware implementation.
Cheong-Fat Chan, Oliver Chiu-sing Choy, Kong-Pang Pun
ISCAS3
2006 A 0.5V fully differential OTA with local common feedback
abstract
This paper presents a fully differential OTA with a supply of 0.5V in a 0.18/spl mu/m CMOS process that has threshold voltages of 0.5V under zero body-source biasing. Unlike a recent two-stage 0.5V OTA architecture, this OTA employs two common mode feedback loops for obtaining a stable operating condition and thus a robust performance against process variations. Body terminals of the input differential pairs of each stage are used for common mode control. Simulation results indicate that this OTA has an open-loop gain of greater than 61dB, a unity gain-bandwidth of 41MHz with a load of 10pF. The power consumption is 510/spl mu/W.
Xiao-Yong He, Kong-Pang Pun, Oliver Chiu-sing Choy, Cheong-Fat Chan
ISCAS3
2006 An optimal normal basis elliptic curve cryptoprocessor for inductive RFID application
abstract
In this paper a 173-bit type II ONB ECC processor for inductive RFID applications is described. Compared with the standard mathematical expressions provided by A. J. Menezes (1993), the formula expressions adopted in the design can save one field multiplication operation in the curve addition. Therefore, the encryption/decryption time is reduced by approximately 7%. Furthermore, the system level architecture of the ECC processor is specially structured and optimized, which makes it faster and less power consuming, and is more favorable for inductive RFID systems.
Pak-Keung Leung, Oliver Chiu-sing Choy, Cheong-Fat Chan, Kong-Pang Pun
ISCAS2
2006 A fully differential low noise amplifier with real-time channel hopping for ultra-wideband wireless applications
abstract
In this paper, a CMOS low noise amplifier (LNA) employing the switched-capacitors is proposed to perform multi-band tuning for an MB-OFDM ultra-wideband (UWB) hopping system. By using the fully differential topology, the switching noise generated during each frequency transition interval is greatly suppressed. The proposed UWB LNA is implemented in a standard 0.18-/spl mu/m CMOS process. The simulation results show that it is capable of performing the fast switching with a settling time of less than 3 ns. Furthermore, the switching noise voltage at the output is rapidly reduced to less than 1 /spl mu/V within 5 ns. The designed circuit is optimized in each center frequency, achieving the maximum power gain of 16.2 dB and the minimum noise figure of 2.63 dB.
Siu-Kei Tang, Kong-Pang Pun, Oliver Chiu-sing Choy, Cheong-Fat Chan
ISCAS3
2006 An ECG measurement IC using driven-right-leg circuit
abstract
In this paper, an electrocardiographic (ECG) signal processing IC, which is used for portable biomedical application, was designed using continuous-time technique. The circuit consists of an instrumentation amplifier (INA) with driven-right-leg circuit (DRL), a 5th order G/sub m/-C low pass filter (G/sub m/-C LPF) operating in sub-threshold mode, and amplifiers. DRL circuit is used to detect small amplitude signal in the presence of large common-mode voltage from the human body. The CMRR of the INA is 78 dB and the G/sub m/-C LPF has a cutoff frequency of 18 Hz. As a result of using the DRL, a small signal can be detected in the presence of large common-mode differential. The circuit consumes 1.23 mW when operating from with a supply voltage of /spl plusmn/1.5-V and occupies a core area of 0.94 mm/sup 2/. The circuit was designed in a 0.35/spl mu/m CMOS process and simulation results have successfully demonstrated the functionalities.
Alex K. Y. Wong, Kong-Pang Pun, Yuan-Ting Zhang, Oliver Chiu-sing Choy
ISCAS4
2006 Power-efficient VLSI implementation of bitstream parsing in H.264/AVC decoder
abstract
In this paper, we propose a power-efficient bitstream parsing for H.264/AVC baseline profile decoding. It parses the input bitstream syntaxes and controls the following decoding steps. Various power reduction techniques, such as data-driven based on statistic results, nonuniform partition, precomputation, guarded evaluation, hierarchical FSM decomposition, clock gating etc., have been adopted in our design. The VLSI implementation results show that under UMC130nm technology with 1.08V supply voltage, the core power consumption is only 1.98mW@20MHz for real-time decoding. Total hardware costs are 49k gates and 1.2 /spl times/ 1.2mm/sup 2/ chip area. The power-efficient and real-time features make our design ideal for low-power video transmission applications such as mobile phone and PDA where video quality is often traded off for energy.
Ke Xu 0014, Oliver Chiu-sing Choy, Cheong-Fat Chan, Kong-Pang Pun
ISCAS2
2004 A low power asynchronous Java processor for contactless smart card
Chun-Pong Yu, Oliver Chiu-sing Choy, Hao Min, Cheong-Fat Chan, Kong-Pang Pun
ASP-DAC2
2004 Card-Centric Framework - Providing I/O Resources for Smart Cards
Pak-Kee Chan, Oliver Chiu-sing Choy, Cheong-Fat Chan, Kong-Pang Pun
CARDIS2
2004 A high-efficiency strongly self-checking asynchronous datapath
abstract
This work examines the inherent self-checking (SC) property of latch-free dynamic asynchronous datapath (LFDAD) using differential cascode voltage switch logic. Consequently, a highly efficient SC dynamic asynchronous datapath architecture is presented. In this architecture, no hardware needs to be added to the datapath to achieve SC. The presented implementation is efficient in terms of speed and area and represents a new approach to fault-tolerant design.
Jing-Ling Yang, Oliver Chiu-sing Choy, Cheong-Fat Chan, Kong-Pang Pun
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2003 Design for Self-Checking and Self-Timed Datapath
abstract
This work examines the inherent self-checking property of a latch-free dynamic asynchronous datapath (LFDAD) using differential cascode voltage switch logic (DCVSL). Consequently, a highly efficient self-checking (SC) dynamic asynchronous datapath architecture is presented. In this architecture, no hardware needs to be added to the datapath to achieve self-checking. The presented implementation is efficient in terms of speed and area and represents a new approach to fault-tolerant design.
Jing-Ling Yang, Oliver Chiu-sing Choy, Cheong-Fat Chan, Kong-Pang Pun
VTS2
2002 A Totally Self-Checking Dynamic Asynchronous Datapath
abstract
This paper investigates the inherent totally self-checking (TSC) property of one type of dynamic asynchronous datapath based on Differential Cascode Voltage Logic (DCVSL). As a result, a totally self-checking dynamic asynchronous datapath architecture is proposed. It is simpler than other similar approaches and represents a new approach to fault tolerant design.
Jing-Ling Yang, Oliver Chiu-sing Choy, Cheong-Fat Chan, Kong-Pang Pun
Asian Test Symposium2
2001 A New Control Circuit for Asynchronous Micropipelines
abstract
In this paper, we present a new configuration of an asynchronous micropipeline, called locally distributed asynchronous (LDA) micropipeline, using a new control circuit. The control circuit generates local control signals according to the 4-phase signaling protocol. The control circuit is used to control the dynamic logic in the datapath. Comparisons based on simulations with other earlier published asynchronous micropipelines are also presented.
Oliver Chiu-sing Choy, Jan Butas, Juraj Povazanec, Cheong-Fat Chan
IEEE Trans. Computers1
2000 Design of self-timed asynchronous Booth's multiplier
abstract
No abstract available.
Tin-Y. Tang, Oliver Chiu-sing Choy, Pui-Lam Siu, Cheong-Fat Chan
ASP-DAC2
2000 An 8×8 adiabatic quasi-static CMOS multiplier
abstract
This paper presents a new type of adiabatic logic. The new adiabatic circuit is named Adiabatic Quasi-Static CMOS (AqsCMOS), because the output is quasi-static. The AqsCMOS is totally compatible with conventional CMOS. Designers can easily reduce the power budget by replacing all or part of an existing CMOS circuit with AqsCMOS circuits to achieve for low power operation. We have designed and fabricated an 8/spl times/8 AqsCMOS multiplier to demonstrate the operation of AqsCMOS. The simulation results have indicated that the new AqsCMOS 8/spl times/8 multiplier consume 90% less power compare with a conventional 8/spl times/8 multiplier of similar architecture.
Wing-Sum Mak, Cheong-Fat Chan, Ka-Wai Cheung, Oliver Chiu-sing Choy
ISCAS4
2000 An ALU design using a novel asynchronous pipeline architecture
abstract
This paper presents a design of a 16-bit pipeline ALU. The ALU is implemented by using a novel asynchronous pipeline architecture. The architecture has simple handshake cells and these cells are embedded in the pipeline stage as normal logic cells. As a result, the speed of the ALU can be very fast.
Tin-Yau Tang, Oliver Chiu-sing Choy, Jan Butas, Cheong-Fat Chan
ISCAS2
1999 A self-timed ICT chip for image coding
abstract
This paper describes an asynchronous one-dimensional order-8 integer cosine transform chip, which can calculate either the forward or inverse transforms. The chip's performance is maximized with a fast computation algorithm and the self-timed circuit technique. The basic self-timed block is a microcoded programmable processor. Eight of these processors are used and achieve a data rate of up to 50 MHz in 0.7-/spl mu/m CMOS technology. If the delay in the handshake circuit can be minimized, the design suggests that asynchronous techniques are a feasible alternative to synchronous designs.
Tin-Chak Johnson Pang, Oliver Chiu-sing Choy, Cheong-Fat Chan, Wai-kuen Cham
IEEE Trans. Circuits Syst. Video Technol.2
1997 Self-timed 1-D ICT processor
abstract
This paper describes a LSI implementation of 1-D order-8 integer cosine transform (ICT) which can calculate either forward or reverse transformation. It is a standard-cell based design using 0.7 /spl mu/m CMOS SLP DLM process. The chip's performance is maximized with the fast computation algorithm and self-timed circuit technique. It consists of eight parallel self-timed pipelines. Each self-timed block is designed based on 2-phase handshaking protocol and variable delay concept. The die size is 5.7/spl times/4.1 mm with about 76 K transistors. This chip supports 16-bit I/O data and its data rate is up to 60 MHz.
Tin-Chak Johnson Pang, Oliver Chiu-sing Choy, Cheong-Fat Chan, Wai-kuen Cham
ASP-DAC2
1996 Test Generation with Dynamic Probe Points in High Observability Testing Environment
abstract
High observability testing environment allows internal circuit nodes to be used as test points. However, such flexibility requires the development of new ATPG algorithm. Previous reported algorithm does not guarantee full fault-coverage and assumes all internal circuit nodes are test points. The new algorithm described in this paper will generate a full fault-coverage test set for a fanout free combinational circuit. The main characteristic of the algorithm is that it generates test vectors as well as probe points. As a result, the probe points are different for each test vector, and the number of probe points is the minimum for test set generated. Results obtained show that an average of 30% test vector reduction is achieved compared with the conventional test method which uses only output pins as test points.
Oliver Chiu-sing Choy, Lap-kong Chan, Ray Chan, Cheong-Fat Chan
IEEE Trans. Computers1
1996 Incremental layout placement modification algorithms
abstract
Many circuit modifications require only a slight adjustment to the IC layouts. General purpose placement algorithms cannot take advantage of these situations because they are designed to generate a complete placement from scratch. In this paper, we present two new algorithms to effect incremental changes on a gate array layout automatically. The algorithms will selectively relocate a number of logic elements to vacate an empty slot. The empty slot is then ready for an added logic element. Results obtained prove that the two algorithms are superior over simple-minded layout modification methods. The computation time is of O(n/sup 3/2/) where n is the number of elements in the neighborhood of change in a layout. For conventional placement algorithms, n will include all the elements in the layout. Therefore, the incremental algorithms will be several orders of magnitude faster.
Oliver Chiu-sing Choy, Tsz-Shing Cheung, Kam-Keung Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1995 A Feedback Control Circuit Design Technique to Suppress Power Noise in High Speed Output Driver
abstract
In today's sub-micron CMOS integrated circuit technology, high speed output switching signals interacting with external inductance and capacitance produce noise which contaminates output signals and power buses. A Feedback Control Slew Rate Output Driver (FCSROD) which reduces the noise spike down to approximately 64% of a conventional output buffer without incurring the penalty of the propagation delay and even the rise/fall time is described. This effective power noise suppression is achieved by using distributed and weighted switching driver segments in conjunction with feedback control to control the output driver's slew rate. Dynamic short circuit current which is generated while both pFET and nFET are conducting is also minimized to reduce di/dt noise. FCSROD was compared with a conventional and the controlled slew rate output buffer, showing 64% noise reduction comparing to the conventional driver, and 22% improvement in both propagation delay and rise/fall time comparing with the controlled slew rate output driver.
Oliver Chiu-sing Choy, Cheong-Fat Chan, M. H. Ku
ISCAS1
1994 Hardware emulation board based on FPGAs and programmable interconnections
abstract
Describes a hardware emulation board based on field programmable gate arrays (FPGAs) and programmable interconnect switches to overcome the limitations of traditional verifications for ASIC designs. Both hardwired buses and programmable buses, via switches, contribute to the interconnection between the FPGAs. With a microprocessor and two EPROMs, the board is designed so that the microprocessor itself can be a part of emulation, in addition to downloading configuration data and testing. Finally presented is the software tool tailored to this board. It automatically partitions the design among multiple FPGAs, programs the switches and facilitates the design and verification process.>
W. Y. Lo, Oliver Chiu-sing Choy, Cheong-Fat Chan
RSP2