VLDB 2026 Research / reviewers in the wild / expert
Ngai Wong 0001
dblp:88/3656
· DBLP profile ↗
158ranked-venue papers
13as first author
76since 2021 · last 2026
0000-0002-3026-0108ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 95 · 10 first-author · 28 since 2021Artificial intelligence and machine learning · 45 · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 2 first-author · 20 since 2021Software engineering, systems software and programming languages · 12 · 1 first-author · 6 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting Model Interpolation for Efficient ReasoningabstractModel merging, typically on Instruct and Thinking models, has shown remarkable performance for efficient reasoning.In this paper, we systematically revisit the simplest merging method that interpolates two weights directly.Particularly, we observe that model interpolation follows a three-stage evolutionary paradigm with distinct behaviors on the reasoning trajectory.These dynamics provide a principled guide for navigating the performancecost trade-off.Empirical results demonstrate that a strategically interpolated model surprisingly surpasses sophisticated model merging baselines on both efficiency and effectiveness.We further validate our findings with extensive ablation studies on model layers, modules, and decoding strategies.Ultimately, this work demystifies model interpolation and offers a practical framework for crafting models with precisely targeted reasoning capabilities.Code is available at Github. Taiqiang Wu, Runming Yang, Jiahao Wang 0005, Ngai Wong 0001 |
ACL (1) | 5 |
| 2026 | Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher DistillationabstractHengyuan Zhang, Shiping Yang, Xiao Liang, Chenming Shang, Yuxuan Jiang, Chaofan Tao, Jing Xiong, Hayden Kwok-Hay So, Ruobing Xie, Angel X. Chang, Ngai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chenming Shang, Chaofan Tao, Hayden Kwok-Hay So, Ruobing Xie, Angel X. Chang, Ngai Wong 0001 |
ACL (1) | 11 |
| 2026 | Activation-Free Implicit Neural Representation via Finite-State-Machine Based Stochastic ComputingabstractImplicit neural representations (INRs) have revolutionized signal encoding by using neural networks to map coordinates to signal attributes. Despite their success, INRs present significant hardware implementation challenges due to complex activation functions and floating-point operations. Unlike previous efforts, such as model pruning or quantization, we address these challenges by introducing AIRFSC, a novel activationfree stochastic computing (SC) architecture that leverages finitestate machines (FSMs). AIRFSC eliminates complex activation functions and processes data efficiently through stochastic bitstreams. Our approach decomposes the input signal into a series of Fourier basis functions, enabling the FSM-based architecture to learn smooth coordinate-to-attribute mappings for accurate signal reconstruction. Extensive experiments on diverse signal types demonstrate that AIRFSC achieves reconstruction quality comparable to state-of-the-art (SOTA) INRs implemented with multi-layer perceptrons (MLPs), while significantly improving hardware efficiency. Specifically, AIRFSC reduces power and area by $\mathbf{9 6. 8 \%}$ and $\mathbf{7 2. 8 \%}$ compared to Sinusoidal Representation Networks (SIREN), and by 97.9% and 81.3% compared to Wavelet Implicit Representation (WIRE). Xincheng Feng, Wenyong Zhou, Taiqiang Wu, Meng Li 0004, Zhengwu Liu, Ngai Wong 0001 |
ASP-DAC | 6 |
| 2026 | A Unified Compute-In-Memory Framework for Multisensory Emotion Recognition
Yue Zhou 0015, Dirui Xie, Zhengwu Liu, Ngai Wong 0001 |
ASP-DAC | 7 |
| 2026 | Noise-Aware Adaptive Sampling for Robust Diffusion Models on Analog Compute-in-MemoryabstractDiffusion models achieve state-of-the-art image generation but impose heavy computational burdens on digital computers. Compute-in-memory (CIM) architectures offer promising acceleration, but inherent noise causes severe performance degradation through weight perturbations. We find that reducing sampling steps improves robustness but limits generation versatility, and that noise at earlier steps causes more severe degradation due to error accumulation. Based on these insights, we propose EtaMix, a novel noise-aware sampling strategy that interpolates between stochastic and deterministic sampling without requiring training or hardware modifications. EtaMix applies more stochastic sampling initially to offset weight perturbations, then gradually transitions to deterministic sampling. Experimental results show EtaMix achieves up to 2.01× and 5.12× FID improvements under different noise conditions for DDPM and DDIM, respectively. Yuannuo Feng, Wenyong Zhou, Yuexi Lv, Guangyao Wang, Zhengwu Liu, Ngai Wong 0001, Wang Kang 0001 |
DATE | 7 |
| 2026 | CAT++: Enhancing 3D Annotations with Hierarchical-Interleaved Encoding and Attention-Conditioned Implicit Representation
Xiaoyan Qian 0001, Chang Liu 0094, Xiaojuan Qi 0001, Siew-Chong Tan, Edmund Y. Lam, Ngai Wong 0001 |
Int. J. Comput. Vis. | 6 |
| 2026 | To Fold or Not to Fold: Graph Regularized Tensor Train for Visual Data CompletionabstractTensor train (TT) representation has achieved tremendous success in visual data completion tasks, especially when it is combined with tensor folding. However, folding an image or video tensor breaks the original data structure, leading to local information loss as nearby pixels may be assigned into different dimensions and become far away from each other. In this paper, to fully preserve the local information of the original visual data, we explore not folding the data tensor, and at the same time adopt graph information to regularize local similarity between nearby entries. To overcome the high computational complexity introduced by the graph-based regularization in the TT completion problem, we propose to break the original problem into multiple sub-problems with respect to each TT core fiber, instead of each TT core as in traditional methods. Furthermore, to avoid heavy parameter tuning, a sparsity-promoting probabilistic model is built based on the generalized inverse Gaussian (GIG) prior, and an inference algorithm is derived under the mean-field approximation. Experiments on both synthetic data and real-world visual data show the superiority of the proposed methods. Lei Cheng 0003, Ngai Wong 0001, Yik-Chung Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | From SMURF to HI-SMURF: Scalable Multivariate Nonlinear Function Approximation via Compact Stochastic Architectures
Xincheng Feng, Wenyong Zhou, Taiqiang Wu, Zhengwu Liu, Meng Li 0004, Ngai Wong 0001 |
IEEE Trans. Computers | 6 |
| 2026 | ReNN-RV: Run-Time PE Reconfiguration for DNN Inference Acceleration With Custom RISC-V ISAabstractDeep neural network (DNN) accelerators integrated with RISC-V Instruction Set Architecture (ISA) extensions have enabled efficient computing on resource-constrained platforms. However, their specialization in regular compute patterns limits effectiveness on irregular workloads, making it challenging to achieve high throughput and energy efficiency. To tackle these challenges, we present ReNN-RV, which integrates a computation-aware RISC-V ISA extension with an instructiondriven processing pipeline to efficiently accelerate run-time reconfigurable processing elements (RePEs). The computation-aware ISA employs configurable opcodes and custom encodings to support fine-grained task scheduling, while an instructiondriven pipeline implements it with minimal control complexity. Moreover, theRePEaccelerator provides seamless switching between multiply-accumulate (MAC) and non-MAC operations by configuring a path multiplexer to realize multiple operators at run time. Experimental results demonstrate that ReNN-RV achieves average reductions of 14.6× in cycle count and 15.3× in execution time across representative DNN workloads compared with the baseline RISC-V design. On average, ReNN-RV outperforms state-of-the-art designs by 10.1× for energy efficiency and 10.3× for computational throughput. Yueting Li 0001, Terry Tao Ye, Ngai Wong 0001, Zhenhua Zhu 0002, Yongfu Li 0002, Weisheng Zhao 0001 |
IEEE Trans. Computers | 3 |
| 2026 | Binary Weight Multibit Activation Quantization for Compute-in-Memory CNN AcceleratorsabstractCompute-in-memory (CIM) accelerators have emerged as a promising way for enhancing the energy efficiency of convolutional neural networks (CNNs). Deploying CNNs on CIM platforms generally requires quantization of network weights and activations to meet hardware constraints. However, existing approaches either prioritize hardware efficiency with binary weight and activation quantization at the cost of accuracy, or utilize multi-bit weights and activations for greater accuracy but limited efficiency. In this paper, we introduce a novel binary weight multi-bit activation (BWMA) method for CNNs on CIM-based accelerators. Our contributions include: deriving closed-form solutions for weight quantization in each layer, significantly improving the representational capabilities of binarized weights; and developing a differentiable function for activation quantization, approximating the ideal multi-bit function while bypassing the extensive search for optimal settings. Through comprehensive experiments on CIFAR-10 and ImageNet datasets, we show that BWMA achieves notable accuracy improvements over existing methods, registering gains of 1.44%-5.46% and 0.35%-5.37% on respective datasets. Moreover, hardware simulation results indicate that 4-bit activation quantization strikes the optimal balance between hardware cost and model performance. Wenyong Zhou, Zhengwu Liu, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory ArchitectureabstractLow-rank adaptation (LoRA) is a predominant parameter-efficient finetuning method for adapting large language models (LLMs) to downstream tasks. Meanwhile, Compute-in-Memory (CIM) architectures demonstrate superior energy efficiency due to their array-level parallel in-memory computing designs. In this article, we propose deploying the LoRA-finetuned LLMs on the hybrid CIM architecture (i.e., pretrained weights onto energy-efficient Resistive Random-Access Memory (RRAM) and LoRA branches onto noise-free Static Random-Access Memory (SRAM)), reducing the energy cost to about 3% compared with the Nvidia A100 GPU. However, the inherent noise of RRAM on the saved weights leads to performance degradation, simultaneously. To address this issue, we design a novel Hardware-aware Low-rank Adaptation (HaLoRA) method. The key insight is to train a LoRA branch that is robust toward such noise and then deploy it on noise-free SRAM, while the extra cost is negligible since the parameters of LoRAs are much fewer than pretrained weights (e.g., 0.15% for LLaMA-3.2 1B model). To improve the robustness towards the noise, we theoretically analyze the gap between the optimization trajectories of the LoRA branch under both ideal and noisy conditions and further design an extra loss to minimize the upper bound of this gap. Therefore, we can enjoy both energy efficiency and accuracy during inference. Experiments finetuning the Qwen and LLaMA series demonstrate the effectiveness of HaLoRA across multiple reasoning tasks, achieving up to 22.7 improvement in average score while maintaining robustness at various noise types and noise levels. Taiqiang Wu, Chenchen Ding, Wenyong Zhou, Yuxin Cheng, Xincheng Feng, Wendong Xu, Chufan Shi, Zhengwu Liu, Ngai Wong 0001 |
ACM Trans. Design Autom. Electr. Syst. | 10 |
| 2026 | PPD: A Portable and Highly Parallel Dispatching System for Deep Learning
Wendong Xu, Yuhao Ji, Yueting Li 0001, Yuxuan Zhao 0001, Zhengwu Liu, Bei Yu 0001, Ngai Wong 0001 |
ACM Trans. Design Autom. Electr. Syst. | 8 |
| 2025 | Stochastic Multivariate Universal-Radix Finite-State Machine: a Theoretically and Practically Elegant Nonlinear Function ApproximatorabstractNonlinearities are crucial for capturing complex input-output relationships especially in deep neural networks. However, nonlinear functions often incur various hardware and compute overheads. Meanwhile, stochastic computing (SC) has emerged as a promising approach to tackle this challenge by trading output precision for hardware simplicity. To this end, this paper proposes a first-of-its-kind stochastic multivariate universal-radix finite-state machine (SMURF) that harnesses SC for hardware-simplistic multivariate nonlinear function generation at high accuracy. We present the finite-state machine (FSM) architecture for SMURF, as well as analytical derivations of sampling gate coefficients for accurately approximating generic nonlinear functions. Experiments demonstrate the superiority of SMURF, requiring only 16.07% area and 14.45% power consumption of Taylor-series approximation, and merely 2.22% area of look-up table (LUT) schemes. Xincheng Feng, Guodong Shen, Jianhao Hu, Meng Li 0016, Ngai Wong 0001 |
ASP-DAC | 5 |
| 2025 | Invited paper: SPICE-Compatible Modeling and Design for Electronic-Photonic Integrated CircuitsabstractElectronic-photonic integrated circuit (EPIC) technologies are revolutionizing computing systems by improving their performance and energy efficiency. However, simulating EPIC is challenging and time consuming. In this paper, we propose the physics-informed neural network (PINN) based SPICE-compatible modeling method for EPICs. Experimental results show our method can speed up EPIC simulation by more than 100 times on average compared to FDTD method. Yinyi Liu, Ngai Wong 0001, Jiang Xu 0001 |
ASP-DAC | 3 |
| 2025 | A Custom RISC-V ISA with Scalable Processing Units for Efficient Neural Network InferenceabstractA customized RISC-V ISA with integrated digital accelerators offers a promising solution to improve energy efficiency in neural network inference.However, it often requires multiple instructions per accelerator operation, which limits computational efficiency during deep neural network inference.To overcome the instruction overhead, this design introduces a dedicated instruction set that enables scalable and fine-grained accelerator control.By incorporating the pattern-driven instruction mode, this design exploits the neural layer regularity to support efficient instruction iteration.Furthermore, this digital accelerator leverages hardware reuse for logic operations, forming a fusion-style architecture that integrates reconfigurable components.Experimental results demonstrate that the custom RISC-V ISA achieves an average runtime speedup of 8.26× and reduces the instruction count by 14.71×.This design also yields an average 8.73× reduction in cycles per instruction across MobileNetV2, ResNet50, VGG19, EfficientNet, and DenseNet-BC, validating its effectiveness across representative benchmarks.Additionally, it improves average energy efficiency by 1.74×, outperforming state-of-the-art designs. Yueting Li 0001, Wanshuang Lin, Wendong Xu, Ngai Wong 0001, Weisheng Zhao 0001 |
CF | 4 |
| 2025 | Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPsabstractDistilling high-accuracy Graph Neural Networks (GNNs) to low-latency multilayer perceptrons (MLPs) on graph tasks has become a hot research topic. However, conventional MLP learning relies almost exclusively on graph nodes and fails to effectively capture the graph structural information. Previous methods address this issue by processing graph edges into extra inputs for MLPs, but such graph structures may be unavailable for various scenarios. To this end, we propose Prototype-Guided Knowledge Distillation (PGKD), which does not require graph edges (edge-free setting) yet learns structure-aware MLPs. Our insight is to distill graph structural information from GNNs. Specifically, we first employ the class prototypes to analyze the impact of graph structures on GNN teachers, and then design two losses to distill such information from GNNs to MLPs. Experimental results on popular graph benchmarks demonstrate the effectiveness and robustness of the proposed PGKD. Taiqiang Wu, Zhe Zhao 0006, Jiahao Wang 0005, Xingyu Bai, Ngai Wong 0001, Yujiu Yang 0001 |
COLING | 6 |
| 2025 | Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language ModelsabstractKullback-Leiber divergence has been widely used in Knowledge Distillation (KD) to compress Large Language Models (LLMs). Contrary to prior assertions that reverse Kullback-Leibler (RKL) divergence is mode-seeking and thus preferable over the mean-seeking forward Kullback-Leibler (FKL) divergence, this study empirically and theoretically demonstrates that neither mode-seeking nor mean-seeking properties manifest in KD for LLMs. Instead, RKL and FKL are found to share the same optimization objective and both converge after a sufficient number of epochs. However, due to practical constraints, LLMs are seldom trained for such an extensive number of epochs. Meanwhile, we further find that RKL focuses on the tail part of the distributions, while FKL focuses on the head part at the beginning epochs. Consequently, we propose a simple yet effective Adaptive Kullback-Leiber (AKL) divergence method, which adaptively allocates weights to combine FKL and RKL. Metric-based and GPT-4-based evaluations demonstrate that the proposed AKL outperforms the baselines across various tasks and improves the diversity and quality of generated responses. Taiqiang Wu, Chaofan Tao, Jiahao Wang 0005, Runming Yang, Zhe Zhao 0006, Ngai Wong 0001 |
COLING | 6 |
| 2025 | DnLUT: Ultra-Efficient Color Image Denoising via Channel-Aware Lookup TablesabstractWhile deep neural networks have revolutionized image de-noising capabilities, their deployment on edge devices remains challenging due to substantial computational and memory requirements. To this end, we present DnLUT, an ultra-efficient lookup table-based framework that achieves high-quality color image denoising with minimal resource consumption. Our key innovation lies in two complementary components: a Pairwise Channel Mixer (PCM) that effectively captures inter-channel correlations and spatial dependencies in parallel, and a novel L-shaped convolution design that maximizes receptive field coverage while minimizing storage overhead. By converting these components into optimized lookup tables post-training, DnLUT achieves remarkable efficiency - requiring only 500KB storage and 0.1% energy consumption compared to its CNN contestant DnCNN, while delivering 20× faster inference. Extensive experiments demonstrate that DnLUT outperforms all existing LUT-based methods by over 1dB in PSNR, establishing a new state-of-the-art in resource-efficient color image de-noising. The project is available at https://github.com/Stephen0808/DnLUT. Sidi Yang, Binxiao Huang, Yulun Zhang 0001, Dahai Yu 0001, Yujiu Yang 0001, Ngai Wong 0001 |
CVPR | 6 |
| 2025 | NoiseZO: RRAM Noise-Driven Zeroth-Order Optimization for Efficient Forward-Only TrainingabstractCompute-in-memory using emerging resistive random-access memory (RRAM) demonstrates significant potential for building energy-efficient deep neural networks. However, RRAM-based network training faces challenges from computational noise and gradient calculation overhead. In this study, we introduce NoiseZO, a forward-only training framework that leverages intrinsic RRAM noise to estimate gradients via zeroth-order (ZO) optimization. The framework maps neural networks onto dual RRAM arrays, utilizing their inherent write noise as $\mathbf{Z O}$ perturbations for training. This enables network updates through only two forward computations. A fine-grained perturbation control strategy is further developed to enhance training accuracy. Extensive experiments on vowel and image datasets, implemented with typical networks, showcase the effectiveness of our framework. Compared to conventional complementary metal-oxide-semiconductor (CMOS) implementations, our approach achieves a 21-fold reduction in energy consumption. Zhengwu Liu, Chenchen Ding, Taiqiang Wu, Jiajun Zhou 0004, Ngai Wong 0001 |
DAC | 7 |
| 2025 | Towards Robust RRAM-Based Vision Transformer Models with Noise-Aware Knowledge DistillationabstractResistive random-access memory (RRAM)-based compute-in-memory (CIM) systems show promise in accelerating Transformer-based vision models but face challenges from inherent device non-idealities. In this work, we systematically investigate the vulnerability of Transformer-based vision models to RRAM-induced perturbations. Our analysis reveals that earlier Transformer layers are more vulnerable than later ones, and feed-forward networks (FFNs) are more susceptible to noise than multi-head self-attention (MHSA). Based on these observations, we propose a noise-aware knowledge distillation framework that enhances model robustness by aligning both intermediate features and final outputs between weight-perturbed and noise-free models. Experimental results demonstrate that our method improves accuracy by up to 1.54% and 1.49% on ViT and DeiT models under various noise conditions compared to their vanilla counterparts. Wenyong Zhou, Zhengwu Liu, Taiqiang Wu, Chenchen Ding, Ngai Wong 0001 |
DATE | 6 |
| 2025 | TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer ReviewabstractYuan Chang, Ziyue Li, Hengyuan Zhang, Yuanbo Kong, Yanru Wu, Hayden Kwok-Hay So, Zhijiang Guo, Liya Zhu, Ngai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yuanbo Kong, Yanru Wu, Hayden Kwok-Hay So, Zhijiang Guo, Liya Zhu, Ngai Wong 0001 |
EMNLP | 9 |
| 2025 | KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented GenerationabstractZiyi Guan, Jason Chun Lok Li, Zhijian Hou, Pingping Zhang, Donglai Xu, Yuzhi Zhao, Mengyang Wu, Jinpeng Chen, Thanh-Toan Nguyen, Pengfei Xian, Wenao Ma, Shengchao Qin, Graziano Chesi, Ngai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jason Chun Lok Li, Zhijian Hou, Donglai Xu, Yuzhi Zhao, Mengyang Wu, Jinpeng Chen 0003, Thanh-Toan Nguyen, Pengfei Xian, Wenao Ma, Shengchao Qin, Graziano Chesi, Ngai Wong 0001 |
EMNLP | 14 |
| 2025 | UNComp: Can Matrix Entropy Uncover Sparsity? - A Compressor Design from an Uncertainty-Aware PerspectiveabstractJing Xiong, Jianghan Shen, Fanghua Ye, Chaofan Tao, Zhongwei Wan, Jianqiao Lu, Xun Wu, Chuanyang Zheng, Zhijiang Guo, Min Yang, Lingpeng Kong, Ngai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jianghan Shen, Fanghua Ye 0001, Chaofan Tao, Zhongwei Wan, Jianqiao Lu, Chuanyang Zheng, Zhijiang Guo, Min Yang 0007, Lingpeng Kong, Ngai Wong 0001 |
EMNLP | 12 |
| 2025 | QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language ModelsabstractJiajun Zhou, Yifan Yang, Kai Zhen, Ziyue Liu, Yequan Zhao, Ershad Banijamali, Athanasios Mouchtaris, Ngai Wong, Zheng Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jiajun Zhou 0004, Kai Zhen, Ziyue Liu 0003, Yequan Zhao, Seyed Ershad Banijamali, Athanasios Mouchtaris, Ngai Wong 0001, Zheng Zhang 0005 |
EMNLP | 8 |
| 2025 | Enhancing Robustness of Implicit Neural Representations Against Weight PerturbationsabstractImplicit Neural Representations (INRs) encode discrete signals in a continuous manner using neural networks, demonstrating significant value across various multimedia applications. However, the vulnerability of INRs presents a critical challenge for their real-world deployments, as the network weights might be subjected to unavoidable perturbations. In this work, we investigate the robustness of INRs for the first time and find that even minor perturbations can lead to substantial performance degradation in the quality of signal reconstruction. To mitigate this issue, we formulate the robustness problem in INRs by minimizing the difference between loss with and without weight perturbations. Furthermore, we derive a novel robust loss function to regulate the gradient of the reconstruction loss with respect to weights, thereby enhancing the robustness. Extensive experiments on reconstruction tasks across multiple modalities demonstrate that our method achieves up to a 7.5 dB improvement in peak signal-to-noise ratio (PSNR) values compared to original INRs under noisy conditions. Wenyong Zhou, Yuxin Cheng, Zhengwu Liu, Taiqiang Wu, Ngai Wong 0001 |
ICASSP | 6 |
| 2025 | MINR: Efficient Implicit Neural Representations for Multi-Image EncodingabstractImplicit Neural Representations (INRs) aim to parameterize discrete signals through implicit continuous functions. However, formulating each image with a separate neural network (typically, a Multi-Layer Perceptron (MLP)) leads to computational and storage inefficiencies when encoding multi-images. To address this issue, we propose MINR, sharing specific layers to encode multi-image efficiently. We first compare the layer-wise weight distributions for several trained INRs and find that corresponding intermediate layers follow highly similar distribution patterns. Motivated by this, we share these intermediate layers across multiple images while preserving the input and output layers as input-specific. In addition, we design an extra novel projection layer for each image to capture its unique features. Experimental results on image reconstruction and super-resolution tasks demonstrate that MINR can save up to 60% parameters while maintaining comparable performance. Particularly, MINR scales effectively to handle 100 images, maintaining an average peak signal-to-noise ratio (PSNR) of 34 dB. Further analysis of various backbones proves the robustness of the proposed MINR. Wenyong Zhou, Taiqiang Wu, Zhengwu Liu, Yuxin Cheng, Ngai Wong 0001 |
ICASSP | 6 |
| 2025 | LLM-Barber: Block-Aware Rebuilder for Sparsity Mask in One-Shot for Large Language ModelsabstractLarge language models (LLMs) have seen substantial growth, necessitating efficient model pruning techniques. Existing post-training pruning methods primarily measure weight importance in converged dense models, often overlooking changes in weight significance during the pruning process, leading to performance degradation. To address this issue, we present LLM-Barber (Block-Aware Rebuilder for Sparsity Mask in One-Shot), a novel one-shot pruning framework that rebuilds the sparsity mask of pruned models without any retraining or weight reconstruction. LLM-Barber incorporates block-aware error optimization across Self-Attention and MLP blocks, facilitating global performance optimization. We are the first to employ the product of weights and gradients as a pruning metric in the context of LLM post-training pruning. This enables accurate identification of weight importance in massive models and significantly reduces computational complexity compared to methods using second-order information. Our experiments show that LLM-Barber efficiently prunes models from LLaMA and OPT families (7B to 13B) on a single A100 GPU in just 30 minutes, achieving state-of-the-art results in both perplexity and zero-shot performance across various language benchmarks. Yupeng Su, Xiaoqun Liu, Tianlai Jin, Dongkuan Wu, Zhengfei Chen, Graziano Chesi, Ngai Wong 0001, Hao Yu 0001 |
ICCAD | 8 |
| 2025 | Perspective-Aware 3D Gaussian Inpainting with Multi-View Consistencyabstract3D Gaussian inpainting, a critical technique for numerous applications in virtual reality and multimedia, has made significant progress with pretrained diffusion models. However, ensuring multi-view consistency, an essential requirement for high-quality inpainting, remains a key challenge. In this work, we present PAInpainter, a novel approach designed to advance 3D Gaussian inpainting by leveraging perspective-aware content propagation and consistency verification across multi-view inpainted images. Our method iteratively refines inpainting and optimizes the 3D Gaussian representation with multiple views adaptively sampled from a perspective graph. By propagating inpainted images as prior information and verifying consistency across neighboring views, PAInpainter substantially enhances global consistency and texture fidelity in restored 3D scenes. Extensive experiments demonstrate the superiority of PAInpainter over existing methods. Our approach achieves superior 3D inpainting quality, with PSNR scores of 26.03 dB and 29.51 dB on the SPIn-NeRF and NeRFiller datasets, respectively, highlighting its effectiveness and generalization capability. Yuxin Cheng, Binxiao Huang, Taiqiang Wu, Wenyong Zhou, Chenchen Ding, Zhengwu Liu, Graziano Chesi, Ngai Wong 0001 |
ICCV | 8 |
| 2025 | Distribution-Aware Hadamard Quantization for Hardware-Efficient Implicit Neural RepresentationsabstractImplicit Neural Representations (INRs) encode discrete signals using Multi-Layer Perceptrons (MLPs) with complex activation functions. While INRs achieve superior performance, they depend on full-precision number representation for accurate computation, resulting in significant hardware overhead. Previous INR quantization approaches have primarily focused on weight quantization, offering only limited hardware savings due to the lack of activation quantization. To fully exploit the hardware benefits of quantization, we propose DHQ, a novel distribution-aware Hadamard quantization scheme that targets both weights and activations in INRs. Our analysis shows that the weights in the first and last layers have distributions distinct from those in the intermediate layers, while the activations in the last layer differ significantly from those in the preceding layers. Instead of customizing quantizers individually, we utilize the Hadamard transformation to standardize these diverse distributions into a unified bell-shaped form, supported by both empirical evidence and theoretical analysis, before applying a standard quantizer. To demonstrate the practical advantages of our approach, we present an FPGA implementation of DHQ that highlights its hardware efficiency. Experiments on diverse image reconstruction tasks show that DHQ outperforms previous quantization methods, reducing latency by 32.7%, energy consumption by 40.1%, and resource utilization by up to 98.3% compared to full-precision counterparts. Wenyong Zhou, Jiachen Ren, Taiqiang Wu, Yuxin Cheng, Zhengwu Liu, Ngai Wong 0001 |
ICME | 6 |
| 2025 | ParallelComp: Parallel Long-Context Compressor for Length ExtrapolationabstractExtrapolating ultra-long contexts (text length $>$128K) remains a major challenge for large language models (LLMs), as most training-free extrapolation methods are not only severely limited by memory bottlenecks, but also suffer from the attention sink, which restricts their scalability and effectiveness in practice. In this work, we propose ParallelComp, a parallel long-context compression method that effectively overcomes the memory bottleneck, enabling 8B-parameter LLMs to extrapolate from 8K to 128K tokens on a single A100 80GB GPU in a training-free setting. ParallelComp splits the input into chunks, dynamically evicting redundant chunks and irrelevant tokens, supported by a parallel KV cache eviction mechanism. Importantly, we present a systematic theoretical and empirical analysis of attention biases in parallel attention—including the attention sink, recency bias, and middle bias—and reveal that these biases exhibit distinctive patterns under ultra-long context settings. We further design a KV cache eviction technique to mitigate this phenomenon. Experimental results show that ParallelComp enables an 8B model (trained on 8K context) to achieve 91.17% of GPT-4’s performance under ultra-long contexts, outperforming closed-source models such as Claude-2 and Kimi-Chat. We achieve a 1.76x improvement in chunk throughput, thereby achieving a 23.50x acceleration in the prefill stage with negligible performance loss and pave the way for scalable and robust ultra-long contexts extrapolation in LLMs. We release the code at https://github.com/menik1126/ParallelComp. Jianghan Shen, Chuanyang Zheng, Zhongwei Wan, Chiwun Yang, Fanghua Ye 0001, Hongxia Yang, Lingpeng Kong, Ngai Wong 0001 |
ICML | 10 |
| 2025 | Nonparametric Teaching for Graph Property LearnersabstractInferring properties of graph-structured data, e.g., the solubility of molecules, essentially involves learning the implicit mapping from graphs to their properties. This learning process is often costly for graph property learners like Graph Convolutional Networks (GCNs). To address this, we propose a paradigm called Graph Nonparametric Teaching (GraNT) that reinterprets the learning process through a novel nonparametric teaching perspective. Specifically, the latter offers a theoretical framework for teaching implicitly defined (i.e., nonparametric) mappings via example selection. Such an implicit mapping is realized by a dense set of graph-property pairs, with the GraNT teacher selecting a subset of them to promote faster convergence in GCN training. By analytically examining the impact of graph structure on parameter-based gradient descent during training, and recasting the evolution of GCNs—shaped by parameter updates—through functional gradient descent in nonparametric teaching, we show for the first time that teaching graph property learners (i.e., GCNs) is consistent with teaching structure-aware nonparametric learners. These new findings readily commit GraNT to enhancing learning efficiency of the graph property learner, showing significant reductions in training time for graph-level regression (-36.62%), graph-level classification (-38.19%), node-level regression (-30.97%) and node-level classification (-47.30%), all while maintaining its generalization performance. Weixin Bu, Zeyi Ren, Zhengwu Liu, Yik-Chung Wu, Ngai Wong 0001 |
ICML | 6 |
| 2025 | Poisoning-based Backdoor Attacks for Arbitrary Target Label with Positive TriggersabstractPoisoning-based backdoor attacks expose vulnerabilities during the data preparation phase of deep neural network (DNN) training. The DNNs trained on the poisoned dataset will be embedded with a backdoor, making them behave well on clean data while outputting malicious predictions whenever a trigger is applied. To exploit the abundant information contained in the input-to-label mapping, our scheme utilizes the network trained from the clean dataset as a trigger generator to produce poisons that significantly raise the success rate of backdoor attacks versus conventional approaches. Specifically, we introduce a new categorization of triggers inspired by adversarial techniques and propose a multi-label and multi-payload Poisoning-based backdoor attack with Positive Triggers (PPT), which strategically manipulates inputs to align them closer to the target label in the feature space of benign classifiers. Once the classifier is trained on the poisoned dataset, we can generate an input-label-aware trigger to make the infected classifier predict any given input to any target label with a high possibility. Through extensive experiments under both dirty-label and clean-label settings, we demonstrate empirically that the proposed attack achieves a high attack success rate without sacrificing accuracy across various datasets, including SVHN, CIFAR10, GTSRB, and Tiny ImageNet. Additionally, the PPT attack can elude a variety of classical backdoor defenses, proving its effectiveness. Binxiao Huang, Ngai Wong 0001 |
IJCAI | 2 |
| 2025 | Hybrid Mesh-Gaussian Representation for Efficient Indoor Scene Reconstructionabstract3D Gaussian splatting (3DGS) has demonstrated exceptional performance in image-based 3D reconstruction and real-time rendering. However, regions with complex textures require numerous Gaussians to capture significant color variations accurately, leading to inefficiencies in rendering speed. To address this challenge, we introduce a hybrid representation for indoor scenes that combines 3DGS with textured meshes. Our approach uses textured meshes to handle texture-rich flat areas, while retaining Gaussians to model intricate geometries. The proposed method begins by pruning and refining the extracted mesh to eliminate geometrically complex regions. We then employ a joint optimization for 3DGS and mesh, incorporating a warm-up strategy and transmittance-aware supervision to balance their contributions seamlessly.Extensive experiments demonstrate that the hybrid representation maintains comparable rendering quality and achieves superior frames per second FPS with fewer Gaussian primitives. Binxiao Huang, Zhihao Li 0002, Shiyong Liu, Jiajun Tang 0001, Yuxin Cheng, Ngai Wong 0001 |
IJCAI | 10 |
| 2025 | A 20.98TOPS/W Energy-Efficient Binary BERT Model on Group Vector Systolic CIM AcceleratorabstractTransformer-based large language models (LLMs) impose significant bandwidth and compute challenges when deployed on edge devices. SRAM-based compute-in-memory (CIM) accelerators offer a promising solution to reduce data movement but are still limited by model size. This work develops a ternary weight splitting (TWS) binarization to obtain Brain-Floating-Point-16×INT1 (BF16×1-b) and INT8×INT1 (8-b×1-b) based transformers that exhibit competitive accuracy while significantly reducing model size compared to full precision counterparts. Then, a fully digital SRAM-based CIM accelerator is designed incorporating a bit-parallel SRAM macro within a highly efficient group vector systolic architecture, which can store one column of BERT-Tiny model with stationary systolic data reuse. The design in a 28nm technology only requires 2KB SRAM with an area of 2mm2. It achieves a throughput of 6.55TOPS and consumes a total power of 312.5mW and 221mW at 400MHz, resulting in a state-of-the-art area efficiency of 3.3TOPS/mm2and normalized energy efficiency of 20.98TOPS/W and 34.35TOPS/W for BF16×1-b and 8-b×1-b respectively on BERT-Tiny model, demonstrating a 10.25× improvement in area efficiency and a 2.23× improvement in energy efficiency compared to other state-of-the-art counterparts. Additionally, our proposed configuration compresses the model size by 32% with only a 0.5% accuracy loss on SST-2. Dingbang Liu, Qilong Chen, Jingyun Gu, Jiaqi Yang 0009, Kai Li 0024, Wei Mao 0002, Ngai Wong 0001, Chang Wen Chen, Hao Yu 0001 |
ISLPED | 8 |
| 2025 | Re-Activating Frozen Primitives for 3D Gaussian Splatting
Yuxin Cheng, Binxiao Huang, Wenyong Zhou, Taiqiang Wu, Zhengwu Liu, Graziano Chesi, Ngai Wong 0001 |
ACM Multimedia | 7 |
| 2025 | Novel Partitioning-Based Approach for Electromigration Assessment With Neural NetworksabstractDue to continuing technology scaling, electromigration (EM) remains a prominent reliability concern in integrated circuit design. Traditional empirical methods often result in over-design in very large scale integration (VLSI) due to model inaccuracy. Recently, researchers have focused on analyzing EM susceptibility by tracking hydrostatic stress evolution in metal lines, governed by computationally expensive partial differential equations (PDEs). In this paper, we propose a partitioning-based approach using neural networks to efficiently forecast the stress evolution along interconnect trees during the void nucleation and growth phases. This approach begins by decomposing the interconnect tree into subcomponents, providing computationally efficient analytical solutions for predicting stress evolution within each subtree. Subsequently, we employ a lightweight neural network to reassemble these components with their corresponding solutions to the original structure, ensuring accurate stress prediction. This divide-and-conquer strategy can accommodate various tree structures, with offshoots at arbitrary junctions, and holds substantial promise for using NN-based methods to solve mesh-free stress evolution on much larger interconnect trees than previously possible, with reduced computational overhead and heightened accuracy. The proposed approach eliminates the need for time discretization and grid meshing typically required in numerical methods. Numerical results confirm its advantages in accuracy and computational efficiency. Tianshu Hou, Farid N. Najm, Ngai Wong 0001, Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Valence-Arousal Disentangled Representation Learning for Emotion Recognition in SSVEP-Based BCIsabstractSteady state visually evoked potential (SSVEP)-based brain-computer interfaces (BCIs), which are widely used in rehabilitation and disability assistance, can benefit from real-time emotion recognition to enhance human-machine interaction. However, the learned discri-minative latent representations in SSVEP-BCIs may generalize in an unintended direction, which can lead to reduced accuracy in detecting emotional states. In this paper, we introduce a Valence-Arousal Disentangled Representation Learning (VADL) method, drawing inspir-ation from the classical two-dimensional emotional model, to enhance the performance and generalization of emotion recognition within SSVEP-BCIs. VADL distinctly disentangles the latent variables of valence and arousal information to improve accuracy. It utilizes the structured state space duality model to thoroughly extract global emotional features. Additionally, we propose a Multisubject Gradient Blending training strategy that individually tailors the learning pace of reconstruction and discrimination tasks within VADL on-the-fly. To verify the feasibility of our method, we have developed a comprehensive database comprising 23 subjects, in which both the emotional states and SSVEPs were effectively elicited. Experimental results indicate that VADL surpasses existing state-of-the-art benchmark algorithms. Zhengwu Liu, Ngai Wong 0001, Zhiwei Ding, Jian Liu 0026, Edith C. H. Ngai |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Lite It Fly: An All-Deformable-Butterfly NetworkabstractMost deep neural networks (DNNs) consist fundamentally of convolutional and/or fully connected layers, wherein the linear transform can be cast as the product between a filter matrix and a data matrix obtained by arranging feature tensors into columns. Recently proposed deformable butterfly (DeBut) decomposes the filter matrix into generalized, butterfly-like factors, thus achieving network compression orthogonal to the traditional ways of pruning or low-rank decomposition. This work reveals an intimate link between DeBut and a systematic hierarchy of depthwise and pointwise convolutions, which explains the empirically good performance of DeBut layers. By developing an automated DeBut chain generator, we show for the first time the viability of homogenizing a DNN into all DeBut layers, thus achieving extreme sparsity and compression. Various examples and hardware benchmarks verify the advantages of All-DeBut networks. In particular, we show it is possible to compress a PointNet to <5% parameters with <5% accuracy drop, a record not achievable by other compression schemes. Jason Chun Lok Li, Jiajun Zhou 0004, Binxiao Huang, Jie Ran, Ngai Wong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Learning Spatially Collaged Fourier Bases for Implicit Neural RepresentationabstractExisting approaches to Implicit Neural Representation (INR) can be interpreted as a global scene representation via a linear combination of Fourier bases of different frequencies. However, such universal basis functions can limit the representation capability in local regions where a specific component is unnecessary, resulting in unpleasant artifacts. To this end, we introduce a learnable spatial mask that effectively dispatches distinct Fourier bases into respective regions. This translates into collaging Fourier patches, thus enabling an accurate representation of complex signals. Comprehensive experiments demonstrate the superior reconstruction quality of the proposed approach over existing baselines across various INR tasks, including image fitting, video representation, and 3D shape representation. Our method outperforms all other baselines, improving the image fitting PSNR by over 3dB and 3D reconstruction to 98.81 IoU and 0.0011 Chamfer Distance. Jason Chun Lok Li, Chang Liu 0094, Binxiao Huang, Ngai Wong 0001 |
AAAI | 4 |
| 2024 | Physics-Informed Learning for Versatile RRAM Reset and Retention SimulationabstractResistive random-access memory (RRAM) constitutes an emerging and promising platform for compute-inmemory (CIM) edge AI. However, the switching mechanism and controllability of RRAM are still under debate owing to the influence of multiphysics. Although physics-informed neural networks (PINNs) are successful in achieving mesh-free multiphysics solutions in many applications, the resultant accuracy is not satisfactory in RRAM analyses. This work investigates the characteristics of RRAM devices - retention and reset transition which are described in terms of the dissolution of a conductive filament (CF) in 3-D axis-symmetric geometry. Specifically, we provide a novel neural network characterization of ion migration, Joule heating, and carrier transport, governed by the solutions of partial differential equations (PDEs). Motivated by physics-informed learning, the separation of variables (SOV) method and the neural tangent kernel (NTK) theory, we propose a customized 3-channel fully-connected network and a modified random Fourier feature (mRFF) embedding strategy to capture multiscale properties and appropriate frequency features of the self-consistent multiphysics solutions. The proposed model eliminates the need for grid meshing and temporal iterations widely used in RRAM analysis. Experiments then confirm its superior accuracy over competing physics-informed methods. Tianshu Hou, Wenyong Zhou, Can Li 0024, Haibao Chen, Ngai Wong 0001 |
ASPDAC | 7 |
| 2024 | APTQ: Attention-aware Post-Training Mixed-Precision Quantization for Large Language ModelsabstractLarge Language Models (LLMs) have greatly advanced the natural language processing paradigm. However, the high computational load and huge model sizes pose a grand challenge for deployment on edge devices. To this end, we propose APTQ (Attention-aware Post-Training Mixed-Precision Quantization) for LLMs, which considers not only the second-order information of each layer's weights, but also, for the first time, the nonlinear effect of attention outputs on the entire model. We leverage the Hessian trace as a sensitivity metric for mixed-precision quantization, ensuring an informed precision reduction that retains model performance. Experiments show APTQ surpasses previous quantization methods, achieving an average of 4 bit width a 5.22 perplexity nearly equivalent to full precision in the C4 dataset. In addition, APTQ attains state-of-the-art zero-shot accuracy of 68.24% and 70.48% at an average bitwidth of 3.8 in LLaMa-7B and LLaMa-13B, respectively, demonstrating its effectiveness to produce high-quality quantized LLMs. Hantao Huang, Yupeng Su, Ngai Wong 0001, Hao Yu 0001 |
DAC | 5 |
| 2024 | An Isotropic Shift-Pointwise Network for Crossbar-Efficient Neural Network DesignabstractResistive random-access memory (RRAM), with its programmable and nonvolatile conductance, permits compute-in-memory (CIM) at a much higher energy efficiency than the traditional von Neumann architecture, making it a promising candidate for edge AI. Nonetheless, the fixed-size crossbar tiles on RRAM are inherently unfit for conventional pyramid-shape convolutional neural networks (CNNs) that incur low crossbar utilization. To this end, we recognize the mixed-signal (digital-analog) nature in RRAM circuits and customize an isotropic shift-pointwise network that exploits digital shift operations for efficient spatial mixing and analog pointwise operations for channel mixing. To fast ablate various shift-pointwise topologies, a new recon-figurable energy-efficient shift module is designed and packaged into a seamless mixed-domain simulator. The optimized design achieves a near-100% crossbar utilization, providing a state-of-the-art INT8 accuracy of 94.88% (76.55%) on the CIFAR-10 (CIFAR-100) dataset with 1.6M parameters, which sets a new standard for RRAM-based AI accelerators. Muqun Niu, Hantao Huang, Graziano Chesi, Hao Yu 0001, Ngai Wong 0001 |
DATE | 8 |
| 2024 | Taming Lookup Tables for Efficient Image Retouching
Sidi Yang, Binxiao Huang, Mingdeng Cao, Yatai Ji, Hanzhong Guo, Ngai Wong 0001, Yujiu Yang 0001 |
ECCV (58) | 6 |
| 2024 | Mixture-of-Subspaces in Low-Rank AdaptationabstractIn this paper, we introduce a subspace-inspired Low-Rank Adaptation (LoRA) method, which is computationally efficient, easy to implement, and readily applicable to large language, multimodal, and diffusion models. Initially, we equivalently decompose the weights of LoRA into two subspaces, and find that simply mixing them can enhance performance. To study such a phenomenon, we revisit it through a fine-grained subspace lens, showing that such modification is equivalent to employing a fixed mixer to fuse the subspaces. To be more flexible, we jointly learn the mixer with the original LoRA weights, and term the method as Mixture-of-Subspaces LoRA (MoSLoRA). MoSLoRA consistently outperforms LoRA on tasks in different modalities, including commonsense reasoning, visual instruction tuning, and subject-driven text-to-image generation, demonstrating its effectiveness and robustness. Taiqiang Wu, Jiahao Wang 0005, Zhe Zhao 0006, Ngai Wong 0001 |
EMNLP | 4 |
| 2024 | Hybrid Module with Multiple Receptive Fields and Self-Attention Layers for Medical Image SegmentationabstractRecent advances in medical image segmentation models combine convolution with the attention mechanism which provides an effective approach to formulate long-term dependencies. However, many works either replaced the convolutional layers with attention layers or embedded attention layers into convolutional neural network (CNN)-based models. To explore the potential of hybrid architecture, we propose a simple cascade module that builds up multiple receptive fields using convolutional kernels with different sizes and learns global context via self-attention layers. Benefiting from the powerful representation ability of the proposed module, multilayer perceptrons (MLPs) with shift operation are adopted to bridge the encoder and decoder to reduce the model size without losing accuracy. Experiments show that our model consistently outperforms the latest 2D and 3D models by large margins on three public tasks and is more resilient to shape, size, and boundary variations. The code is available at https://github.com/cicailalala/AERFNet. Wenbo Qi, Wenyong Zhou, Ngai Wong 0001, S. C. Chan 0001 |
ICASSP | 3 |
| 2024 | MCUBERT: Memory-Efficient BERT Inference on Commodity MicrocontrollersabstractIn this paper, we propose MCUBERT to enable language models like BERT on tiny microcontroller units (MCUs) through network and scheduling co-optimization. We observe the embedding table contributes to the major storage bottleneck for tiny BERT models. Hence, at the network level, we propose an MCU-aware two-stage neural architecture search algorithm based on clustered low-rank approximation for embedding compression. To reduce the inference memory requirements, we further propose a novel fine-grained MCU-friendly scheduling strategy. Through careful computation tiling and re-ordering as well as kernel design, we drastically increase the input sequence lengths supported on MCUs without any latency or accuracy penalty. MCUBERT reduces the parameter size of BERT-tiny and BERT-mini by 5.7× and 3.0× and the execution memory by 3.5× and 4.3×, respectively. MCUBERT also achieves 1.5× latency reduction. For the first time, MCUBERT enables lightweight BERT models on commodity MCUs and processing more than 512 tokens with less than 256KB of memory. Renze Chen, Taiqiang Wu, Ngai Wong 0001, Yun Liang 0001, Runsheng Wang, Ru Huang 0001, Meng Li 0004 |
ICCAD | 4 |
| 2024 | ASMR: Activation-Sharing Multi-Resolution Coordinate Networks for Efficient InferenceabstractCoordinate network or implicit neural representation (INR) is a fast-emerging method for encoding natural signals (such as images and videos) with the benefits of a compact neural representation. While numerous methods have been proposed to increase the encoding capabilities of an INR, an often overlooked aspect is the inference efficiency, usually measured in multiply-accumulate (MAC) count. This is particularly critical in use cases where inference bandwidth is greatly limited by hardware constraints. To this end, we propose the Activation-Sharing Multi-Resolution (ASMR) coordinate network that combines multi-resolution coordinate decomposition with hierarchical modulations. Specifically, an ASMR model enables the sharing of activations across grids of the data. This largely decouples its inference cost from its depth which is directly correlated to its reconstruction capability, and renders a near $O(1)$ inference complexity irrespective of the number of layers. Experiments show that ASMR can reduce the MAC of a vanilla SIREN model by up to 500$\times$ while achieving an even higher reconstruction quality than its SIREN baseline. Jason Chun Lok Li, Steven Tin Sui Luo, Ngai Wong 0001 |
ICLR | 4 |
| 2024 | Nonparametric Teaching of Implicit Neural RepresentationsabstractWe investigate the learning of implicit neural representation (INR) using an overparameterized multilayer perceptron (MLP) via a novel nonparametric teaching perspective. The latter offers an efficient example selection framework for teaching nonparametrically defined (viz. non-closed-form) target functions, such as image functions defined by 2D grids of pixels. To address the costly training of INRs, we propose a paradigm called Implicit Neural Teaching (INT) that treats INR learning as a nonparametric teaching problem, where the given signal being fitted serves as the target function. The teacher then selects signal fragments for iterative training of the MLP to achieve fast convergence. By establishing a connection between MLP evolution through parameter-based gradient descent and that of function evolution through functional gradient descent in nonparametric teaching, we show *for the first time* that teaching an overparameterized MLP is consistent with teaching a nonparametric learner. This new discovery readily permits a convenient drop-in of nonparametric teaching algorithms to broadly enhance INR training efficiency, demonstrating 30%+ training time savings across various input modalities. Steven Tin Sui Luo, Jason Chun Lok Li, Yik-Chung Wu, Ngai Wong 0001 |
ICML | 5 |
| 2024 | Hundred-Kilobyte Lookup Tables for Efficient Single-Image Super-Resolution
Binxiao Huang, Jason Chun Lok Li, Jie Ran, Jiajun Zhou 0004, Dahai Yu 0001, Ngai Wong 0001 |
IJCAI | 7 |
| 2024 | LoRETTA: Low-Rank Economic Tensor-Train Adaptation for Ultra-Low-Parameter Fine-Tuning of Large Language ModelsabstractYifan Yang, Jiajun Zhou, Ngai Wong, Zheng Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jiajun Zhou 0004, Ngai Wong 0001, Zheng Zhang 0005 |
NAACL-HLT | 3 |
| 2024 | Scaling Laws with Vocabulary: Larger Models Deserve Larger VocabulariesabstractResearch on scaling large language models (LLMs) has primarily focused on model parameters and training data size, overlooking the role of vocabulary size. We investigate how vocabulary size impacts LLM scaling laws by training models ranging from 33M to 3B parameters on up to 500B characters with various vocabulary configurations. We propose three complementary approaches for predicting the compute-optimal vocabulary size: IsoFLOPs analysis, derivative estimation, and parametric fit of the loss function. Our approaches converge on the conclusion that the optimal vocabulary size depends on the compute budget, with larger models requiring larger vocabularies. Most LLMs, however, use insufficient vocabulary sizes. For example, we predict that the optimal vocabulary size of Llama2-70B should have been at least 216K, 7 times larger than its vocabulary of 32K. We validate our predictions empirically by training models with 3B parameters across different FLOPs budgets. Adopting our predicted optimal vocabulary size consistently improves downstream performance over commonly used vocabulary sizes. By increasing the vocabulary size from the conventional 32K to 43K, we improve performance on ARC-Challenge from 29.1 to 32.0 with the same 2.3e21 FLOPs. Our work highlights the importance of jointly considering tokenization and model scaling for efficient pre-training. The code and demo are available at https://github.com/sail-sg/scaling-with-vocab and https://hf.co/spaces/sail/scaling-with-vocab-demo. Chaofan Tao, Qian Liu 0033, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo 0002, Ngai Wong 0001 |
NeurIPS | 8 |
| 2024 | Hyperdimensional Computing With Multiscale Local Binary Patterns for Scalp EEG-Based Epileptic Seizure DetectionabstractEpilepsy is a common condition that causes frequent seizures, significantly impacting patients’ daily lives. Non-invasive EEG is an effective tool for detecting seizure onset. Wearable EEG devices enable real-time monitoring and timely intervention but pose new algorithmic challenges on small model weight sizes and limited training data. Brain-inspired hyperdimensional computing (HDC) presents a potential solution for its small weight size and quick learning ability. Combining local binary pattern (LBP) codes with HDC can capture dynamic features in EEG time series. However, traditional LBP features may not offer sufficient robustness for trend modeling due to their high localization on individual samples, particularly on low-amplitude and non-stationary scalp EEG signals. To address the above challenges, this paper proposes a multi-scale LBP-based HDC (MSLBP-HDC) approach for scalp EEG analysis. Unlike traditional LBP-based HDC focusing only on the local change trend, the designed MSLBP-HDC extracts dynamic features at different time resolutions to detect abnormal cortical oscillations. The lengths of multiple temporal scales in MSLBP-HDC are determined based on the duration of spikes. Our results demonstrate that MSLBP-HDC has the highest specificity for all test seizure types and achieves competitive macroaveraging accuracy with the smallest model weight size in detection, compared to advanced deep learning, support vector machine, and HDC methods. Regarding few-shot learning performance, MSLBP-HDC outperforms existing approaches and achieves high accuracy using only 1% of the training data. Moreover, feature interpretability analysis from space and time domains highlights that MSLBP-HDC successfully extracts seizure-relevant features rather than noise or artifacts, ensuring the algorithm’s reliability. Ngai Wong 0001, Edith C. H. Ngai |
IEEE Internet Things J. | 3 |
| 2024 | DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network InferenceabstractTo accelerate the inference of deep neural networks (DNNs), quantization with low-bitwidth numbers is actively researched. A prominent challenge is to quantize the DNN models into low-bitwidth numbers without significant accuracy degradation, especially at very low bitwidths (< 8 bits). This work targets an adaptive data representation with variablelength encoding called DyBit. DyBit can dynamically adjust the precision and range of separate bit-fields to be adapted to the DNN weights/activations distribution. We also propose a hardware-aware quantization framework with a mixed-precision accelerator to trade-off the inference accuracy and speedup. Experimental results demonstrate that the ImageNet inference accuracy via DyBit is 1.97% higher than the state-of-the-art at 4-bit quantization, and the proposed framework can achieve up to 8.1× speedup compared with the original ResNet-50 model. Jiajun Zhou 0004, Jiajun Wu 0006, Yizhao Gao 0002, Yuhao Ding, Chaofan Tao, Fengbin Tu, Kwang-Ting Cheng, Hayden Kwok-Hay So, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2024 | FAT: Frequency-Aware Transformation for Bridging Full-Precision and Low-Precision Deep RepresentationsabstractLearning low-bitwidth convolutional neural networks (CNNs) is challenging because performance may drop significantly after quantization. Prior arts often quantize the network weights by carefully tuning hyperparameters such as nonuniform stepsize and layerwise bitwidths, which are complicated since the full- and low-precision representations have large discrepancies. This work presents a novel quantization pipeline, named frequency-aware transformation (FAT), that features important benefits: 1) instead of designing complicated quantizers, FAT learns to transform network weights in the frequency domain to remove redundant information before quantization, making them amenable to training in low bitwidth with simple quantizers; 2) FAT readily embeds CNNs in low bitwidths using standard quantizers without tedious hyperparameter tuning and theoretical analyses show that FAT minimizes the quantization errors in both uniform and nonuniform quantizations; and 3) FAT can be easily plugged into various CNN architectures. Using FAT with a simple uniform/logarithmic quantizer can achieve the state-of-the-art performance in different bitwidths on various model architectures. Consequently, FAT serves to provide a novel frequency-based perspective for model quantization. Chaofan Tao, Quan Chen 0007, Zhaoyang Zhang 0004, Ping Luo 0002, Ngai Wong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Context-Aware Transformer for 3D Point Cloud Automatic Annotationabstract3D automatic annotation has received increased attention since manually annotating 3D point clouds is laborious. However, existing methods are usually complicated, e.g., pipelined training for 3D foreground/background segmentation, cylindrical object proposals, and point completion. Furthermore, they often overlook the inter-object feature correlation that is particularly informative to hard samples for 3D annotation. To this end, we propose a simple yet effective end-to-end Context-Aware Transformer (CAT) as an automated 3D-box labeler to generate precise 3D box annotations from 2D boxes, trained with a small number of human annotations. We adopt the general encoder-decoder architecture, where the CAT encoder consists of an intra-object encoder (local) and an inter-object encoder (global), performing self-attention along the sequence and batch dimensions, respectively. The former models intra-object interactions among points and the latter extracts feature relations among different objects, thus boosting scene-level understanding. Via local and global encoders, CAT can generate high-quality 3D box annotations with a streamlined workflow, allowing it to outperform existing state-of-the-arts by up to 1.79% 3D AP on the hard task of the KITTI test set. Xiaoyan Qian 0001, Chang Liu 0094, Xiaojuan Qi 0001, Siew-Chong Tan, Edmund Y. Lam, Ngai Wong 0001 |
AAAI | 6 |
| 2023 | RankSearch: An Automatic Rank Search Towards Optimal Tensor Compression for Video LSTM Networks on EdgeabstractVarious industrial and domestic applications call for optimized lightweight video LSTM network models on edge. The recent tensor-train method can transform space-time features into tensors, which can be further decomposed into low-rank network models for lightweight video analysis on edge. The rank selection of tensor is however manually performed with no optimization. This paper formulates a rank search algorithm to automatically decide tensor ranks with consideration of the trade-off between network accuracy and complexity. A fast rank search method, called RankSearch, is developed to find optimized low-rank video LSTM network models on edge. Results from experiments show that RankSearch achieves a$4.84 >$reduction in model complexity, and$1.96\times$speed-up in run time while delivering a 3.86% accuracy improvement compared with the manual-ranked models. Changhai Man, Chenchen Ding, Shaobo Luo, Rumin Zhang, Ngai Wong 0001, Hao Yu 0001 |
DATE | 10 |
| 2023 | PECAN: A Product-Quantized Content Addressable Memory NetworkabstractA novel deep neural network (DNN) architecture is proposed wherein the filtering and linear transform are realized solely with product quantization (PQ). This results in a natural implementation via content addressable memory (CAM), which transcends regular DNN layer operations and requires only simple table lookup. Two schemes are developed for the end-to-end PQ prototype training, namely, through angle- and distance-based similarities, which differ in their multiplicative and additive natures with different complexity-accuracy tradeoffs. Even more, the distance-based scheme constitutes a truly multiplier-free DNN solution. Experiments confirm the feasibility of such Product-QuantizEd Content Addressable Memory Network (PECAN), which has strong implication on hardware-efficient deployments especially for in-memory computing. Jie Ran, Jason Chun Lok Li, Jiajun Zhou 0004, Ngai Wong 0001 |
DATE | 5 |
| 2023 | MSD: Mixing Signed Digit Representations for Hardware-efficient DNN Acceleration on FPGA with Heterogeneous ResourcesabstractBy quantizing weights with different precision for different parts of a network, mixed-precision quantization promises to reduce the hardware cost and improve the speed of deep neural network (DNN) accelerators that typically operate with a fixed quantization scheme. However, the additional control needed, and the decreased hardware efficiency arising from multi-precision operations have made mixed-precision quantization schemes challenging to deploy in practice. In this paper, a practical mixed-precision quantization framework called MSD that leverages the heterogeneous computing resources on FPGA to perform bit-serial and bit-parallel operations simultaneously is presented. MSD combines the use of a custom restricted signed digit (RSD) representation, which utilizes a limited number of effectual bits, and the conventional 2's complement representation to quantize DNN weights. Depending on the availability of fine-grained and coarse-grained resources, MSD encodes a subset of weights with RSD to allow highly efficient bit-serial multiply-accumulate implementation using LUT resources. Furthermore, the number of effectual bits used in RSD is optimized to match the bit-serial hardware latency to the bit-parallel operation on the coarse-grained resources to ensure the highest run-time utilization of all on-chip resources. Experiments show that MSD achieved a 1.36× speedup on the ResNet-18 model over the state-of-the-art, and a remarkable 4.91% higher accuracy on MobileNet-V2. Jiajun Wu 0006, Jiajun Zhou 0004, Yizhao Gao 0002, Yuhao Ding, Ngai Wong 0001, Hayden Kwok-Hay So |
FCCM | 5 |
| 2023 | Tensor train factorization under noisy and incomplete data with automatic rank estimation
Lei Cheng 0003, Ngai Wong 0001, Yik-Chung Wu |
Pattern Recognit. | 3 |
| 2023 | Analytical Post-Voiding Modeling and Efficient Characterization of EM Failure Effects Under Time-Dependent Current StressingabstractElectromigration (EM) has become the major concern for integrated circuits (ICs) in advanced technology nodes. Traditional empirical EM models, such as Black’s equation, show inaccurate estimation for the time-to-failure of ICs, thus resulting in unnecessary over-design. To address this drawback, we propose a few analytical solutions for calculating the transient stress evolution and void volume in straight multisegment interconnect trees during the post-voiding phase. By employing the Laplace transform, the proposed method aims at solving coupled partial differential equations (PDEs) governed by physics-based EM modeling. The analytical solutions can be tailored to expressions with required accuracy and computational savings, leading to a compact end-to-end system providing results of EM failure effects at arbitrary time instances and locations of interconnect trees with varying geometry under time-dependent current and temperature stressing. The EM lifetime such as the incubation time, related to the void volume evolution, at any desired precision, can be calculated by the analytical solutions. The proposed method shows its accuracy, scalability, and computational savings through results compared with the finite element method (FEM) tool COMSOL and the competing methods and can achieve up to$593\times $speedup with < 10% error in EM failure time estimation. Tianshu Hou, Ngai Wong 0001, Quan Chen 0007, Zhigang Ji, Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Multilayer Perceptron-Based Stress Evolution Analysis Under DC Current Stressing for Multisegment WiresabstractElectromigration (EM) is one of the major concerns in the reliability analysis of very large-scale integration (VLSI) systems due to the continuous technology scaling. Accurately predicting the time-to-failure of integrated circuits (ICs) becomes increasingly important for modern IC design. However, traditional methods are often not sufficiently accurate, leading to undesirable over-design especially in advanced technology nodes. In this article, we propose an approach using multilayer perceptrons (MLPs) to compute stress evolution in the interconnect trees during the void nucleation phase. The availability of a customized trial function for neural network training holds the promise of finding dynamic mesh-free stress evolution on complex interconnect trees under time-varying temperatures. Specifically, we formulate a new objective function considering the EM-induced coupled partial differential equations (PDEs), boundary conditions (BCs), and initial conditions to enforce the physics-based constraints in the spatial–temporal domain. The proposed model avoids meshing and reduces temporal iterations compared with conventional numerical approaches like finite element method. Numerical results confirm its advantages on accuracy and computational performance. Tianshu Hou, Peining Zhen, Ngai Wong 0001, Quan Chen 0007, Guoyong Shi, Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Compression of Generative Pre-trained Language Models via QuantizationabstractChaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang, Xin Jiang, Qun Liu, Ping Luo, Ngai Wong. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Chaofan Tao, Lu Hou 0002, Wei Zhang 0196, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Ping Luo 0002, Ngai Wong 0001 |
ACL (1) | 8 |
| 2022 | Multimodal Transformer for Automatic 3D Annotation and Object Detection
Chang Liu 0094, Xiaoyan Qian 0001, Binxiao Huang, Xiaojuan Qi 0001, Edmund Y. Lam, Siew-Chong Tan, Ngai Wong 0001 |
ECCV (38) | 7 |
| 2022 | Coarse to Fine: Image Restoration Boosted by Multi-Scale Low-Rank Tensor CompletionabstractExisting low-rank tensor completion (LRTC) approaches aim at restoring a partially observed tensor by imposing a global low-rank constraint on the underlying completed tensor. However, such a global rank assumption suffers the trade-off between restoring the originally details-lacking parts and neglecting the potentially complex objects, making the completion performance unsatisfactory on both sides. To address this problem, we propose a novel and practical strategy for image restoration that restores the partially observed tensor in a coarse-to-fine (C2F) manner, which gets rid of such trade-off by searching proper local ranks for both low- and high-rank parts. Extensive experiments are conducted to demonstrate the superiority of the proposed C2F scheme. The codes are available at: https://github.com/RuiLin0212/C2FLRTC. Cong Chen 0003, Ngai Wong 0001 |
ICPR | 3 |
| 2022 | MAP-Gen: An Automated 3D-Box Annotation Flow with Multimodal Attention Point GeneratorabstractManually annotating 3D point clouds is laborious and costly, limiting the training data preparation for deep learning in real-world object detection. While a few previous studies tried to automatically generate 3D bounding boxes from weak labels such as 2D boxes, the quality is sub-optimal compared to human annotators. This work proposes a novel autolabeler, called multimodal attention point generator (MAP-Gen), that generates high-quality 3D labels from weak 2D boxes. It leverages dense image information to tackle the sparsity issue of 3D point clouds, thus improving label quality. For each 2D pixel, MAP-Gen predicts its corresponding 3D coordinates by referencing context points based on their 2D semantic or geometric relationships. The generated 3D points densify the original sparse point clouds, followed by an encoder to regress 3D bounding boxes. Using MAP-Gen, object detection networks that are weakly supervised by 2D boxes can achieve 94 ~ 99% performance of those fully supervised by 3D annotations. It is hopeful this newly proposed MAP-Gen autolabeling flow can shed new light on utilizing multimodal information for enriching sparse point clouds. Chang Liu 0094, Xiaoyan Qian 0001, Xiaojuan Qi 0001, Edmund Y. Lam, Siew-Chong Tan, Ngai Wong 0001 |
ICPR | 6 |
| 2022 | ODG-Q: Robust Quantization via Online Domain GeneralizationabstractQuantizing neural networks to low-bitwidth is important for model deployment on resource-limited edge hardware. Although a quantized network has a smaller model size and memory footprint, it is fragile to adversarial attacks. However, few methods study the robustness and training efficiency of quantized networks. To this end, we propose a new method by recasting robust quantization as an online domain generalization problem, termed ODG-Q, which generates diverse adversarial data at a low cost during training. ODG-Q consistently outperforms existing works against various adversarial attacks. For example, on CIFAR-10 dataset, ODG-Q achieves 49.2% average improvements under five common white-box attacks and 21.7% average improvements under five common black-box attacks, with a training cost similar to that of natural training (viz. without adversaries). To our best knowledge, this work is the first work that trains both quantized and binary neural networks on ImageNet that consistently improves robustness under different attacks. We also provide a theoretical insight of ODG-Q that accounts for the bound of model risk on attacked data. Chaofan Tao, Ngai Wong 0001 |
ICPR | 2 |
| 2022 | FASSST: Fast Attention Based Single-Stage Segmentation Net for Real-Time Instance SegmentationabstractReal-time instance segmentation is crucial in various AI applications. This work designs a network named Fast Attention based Single-Stage Segmentation NeT (FASSST) that performs instance segmentation with video-grade speed. Using an instance attention module (IAM), FASSST quickly locates target instances and segments with region of interest (ROI) feature fusion (RFF) aggregating ROI features from pyramid mask layers. The module employs an efficient single-stage feature regression, straight from features to instance coordinates and class probabilities. Experiments on COCO and CityScapes datasets show that FASSST achieves state-of-the-art performance under competitive accuracy: real-time inference of 47.5FPS on a GTX1080Ti GPU and 5.3FPS on a Jetson Xavier NX board with only 71.6 GFLOPs. Peining Zhen, Tianshu Hou, Chiu Wa Ng, Haibao Chen, Hao Yu 0001, Ngai Wong 0001 |
WACV | 8 |
| 2022 | EZCrop: Energy-Zoned Channels for Robust Output PruningabstractRecent results have revealed an interesting observation in a trained convolutional neural network (CNN), namely, the rank of a feature map channel matrix remains surprisingly constant despite the input images. This has led to an effective rank-based channel pruning algorithm [23], yet the constant rank phenomenon remains mysterious and unexplained. This work aims at demystifying and interpreting such rank behavior from a frequency-domain perspective, which as a bonus suggests an extremely efficient Fast Fourier Transform (FFT)-based metric for measuring channel importance without explicitly computing its rank. We achieve remarkable CNN channel pruning based on this analytically sound and computationally efficient metric, and adopt it for repetitive pruning to demonstrate robustness via our scheme named Energy-Zoned Channels for Robust Output Pruning (EZCrop), which shows consistently better results than other state-of-the-art channel pruning methods. The codes and Appendix are publicly available at: https://github.com/ruilin0212/EZCrop. Jie Ran, Dongpeng Wang, King Hung Chiu, Ngai Wong 0001 |
WACV | 5 |
| 2022 | Kernelized support tensor train machines
Cong Chen 0003, Kim Batselier, Wenjian Yu, Ngai Wong 0001 |
Pattern Recognit. | 4 |
| 2022 | A Space-Time Neural Network for Analysis of Stress Evolution Under DC Current StressingabstractThe electromigration (EM)-induced reliability issues in very large-scale integration (VLSI) circuits have attracted increased attention due to the continuous technology scaling. Traditional EM models often lead to overly pessimistic predictions incompatible with the shrinking design margin in future technology nodes. Motivated by the latest success of neural networks in solving differential equations in physical problems, we propose a novel mesh-free model to compute EM-induced stress evolution in VLSI circuits. The model utilizes a specifically crafted space–time physics-informed neural network (STPINN) as the solver for EM analysis. By coupling the physics-based EM analysis with dynamic temperature incorporating Joule heating and via effect, we can observe stress evolution along multisegment interconnect trees under constant, time-dependent, and space–time-dependent temperature during the void nucleation phase. The proposed STPINN method obviates the time discretization and meshing required in conventional numerical stress evolution analysis and offers significant computational savings. Numerical comparison with competing schemes demonstrates a$2\times $–$52\times $speedup with a satisfactory accuracy. Tianshu Hou, Ngai Wong 0001, Quan Chen 0007, Zhigang Ji, Haibao Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | LiteGT: Efficient and Lightweight Graph TransformersabstractTransformers have shown great potential for modeling long-term dependencies for natural language processing and computer vision. However, little study has applied transformers to graphs, which is challenging due to the poor scalability of the attention mechanism and the under-exploration of graph inductive bias. To bridge this gap, we propose a Lite Graph Transformer (LiteGT) that learns on arbitrary graphs efficiently. First, a node sampling strategy is proposed to sparsify the considered nodes in self-attention with only O (Nlog N) time. Second, we devise two kernelization approaches to form two-branch attention blocks, which not only leverage graph-specific topology information, but also reduce computation further to O (1 over 2 Nlog N). Third, the nodes are updated with different attention schemes during training, thus largely mitigating over-smoothing problems when the model layers deepen. Extensive experiments demonstrate that LiteGT achieves competitive performance on both node classification and link prediction on datasets with millions of nodes. Specifically, Jaccard + Sampling + Dim. reducing setting reduces more than 100x computation and halves the model size without performance degradation. Cong Chen 0003, Chaofan Tao, Ngai Wong 0001 |
CIKM | 3 |
| 2021 | A Video-based Fall Detection Network by Spatio-temporal Joint-point Model on Edge DevicesabstractTripping or falling is among the top threats in elderly healthcare, and the development of automatic fall detection systems are of considerable importance. With the fast development of the Internet of Things (IoT), camera vision-based solutions have drawn much attention in recent years. The traditional fall video analysis on the cloud has significant communication overhead. This work introduces a fast and lightweight video fall detection network based on a spatio-temporal joint-point model to overcome these hurdles. Instead of detecting falling motion by the traditional Convolutional Neural Networks (CNNs), we propose a Long Short-Term Memory (LSTM) model based on time-series joint-point features, extracted from a pose extractor and then filtered from a geometric joint-point filter. Experiments are conducted to verify the proposed framework, which shows a high sensitivity of 98.46% on Multiple Cameras Fall Dataset and 100% on UR Fall Dataset. Furthermore, our model can achieve pose estimation tasks simultaneously, attaining 73.3 mAP in the COCO keypoint challenge dataset, which outperforms the OpenPose work by 8%. Shuwei Li, Changhai Man, Wei Mao 0002, Ngai Wong 0001, Hao Yu 0001 |
DATE | 6 |
| 2021 | Overfitting Avoidance in Tensor Train Factorization and Completion: Prior Analysis and InferenceabstractTensor train (TT) decomposition, a powerful tool for analyzing multidimensional data, exhibits superior performance in many machine learning tasks. However, existing methods for TT decomposition either suffer from noise overfitting, or require extensive fine-tuning of the balance between model complexity and representation accuracy. In this paper, a fully Bayesian treatment of TT decomposition is employed to avoid noise overfitting without parameter tuning. In particular, theoretical evidence is established for adopting a Gaussian-product-Gamma prior to induce sparsity on the slices of the TT cores. Furthermore, based on the proposed probabilistic model, an efficient learning algorithm is derived under the variational inference framework. Experiments on real-world data demonstrate the proposed algorithm performs better in image completion and image classification, compared to other existing TT decomposition algorithms. Lei Cheng 0003, Ngai Wong 0001, Yik-Chung Wu |
ICDM | 3 |
| 2021 | Deformable Butterfly: A Highly Structured and Sparse Linear TransformabstractWe introduce a new kind of linear transform named Deformable Butterfly (DeBut) that generalizes the conventional butterfly matrices and can be adapted to various input-output dimensions. It inherits the fine-to-coarse-grained learnable hierarchy of traditional butterflies and when deployed to neural networks, the prominent structures and sparsity in a DeBut layer constitutes a new way for network compression. We apply DeBut as a drop-in replacement of standard fully connected and convolutional layers, and demonstrate its superiority in homogenizing a neural network and rendering it favorable properties such as light weight and low inference complexity, without compromising accuracy. The natural complexity-accuracy tradeoff arising from the myriad deformations of a DeBut layer also opens up new rooms for analytical and practical research. The codes and Appendix are publicly available at: https://github.com/ruilin0212/DeBut. Jie Ran, King Hung Chiu, Graziano Chesi, Ngai Wong 0001 |
NeurIPS | 5 |
| 2021 | S3-Net: A Fast and Lightweight Video Scene Understanding Network by Single-shot SegmentationabstractReal-time understanding in video is crucial in various AI applications such as autonomous driving. This work presents a fast single-shot segmentation strategy for video scene understanding. The proposed net, called S3-Net, quickly locates and segments target sub-scenes, meanwhile extracts structured time-series semantic features as inputs to an LSTM-based spatio-temporal model. Utilizing ten-sorization and quantization techniques, S3-Net is intended to be lightweight for edge computing. Experiments using CityScapes, UCF11, HMDB51 and MOMENTS datasets demonstrate that the proposed S3-Net achieves an accuracy improvement of 8.1% versus the 3D-CNN based approach on UCF11, a storage reduction of 6.9× and an inference speed of 22.8 FPS on CityScapes with a GTX1080Ti GPU. Haibao Chen, Ngai Wong 0001, Hao Yu 0001 |
WACV | 4 |
| 2021 | S3-Net: A Fast Scene Understanding Network by Single-Shot Segmentation for Autonomous DrivingabstractReal-time segmentation and understanding of driving scenes are crucial in autonomous driving. Traditional pixel-wise approaches extract scene information by segmenting all pixels in a frame, and hence are inefficient and slow. Proposal-wise approaches only learn from the proposed object candidates, but still require multiple steps on the expensive proposal methods. Instead, this work presents a fast single-shot segmentation strategy for video scene understanding. The proposed net, called S3-Net, quickly locates and segments target sub-scenes , and meanwhile extracts attention-aware time-series sub-scene features ( ats-features ) as inputs to an attention-aware spatio-temporal model (ASM) . Utilizing tensorization and quantization techniques, S3-Net is intended to be lightweight for edge computing. Experiments results on CityScapes, UCF11, HMDB51, and MOMENTS datasets demonstrate that the proposed S3-Net achieves an accuracy improvement of 8.1% versus the 3D-CNN based approach on UCF11, a storage reduction of 6.9× and an inference speed of 22.8 FPS on CityScapes with a GTX1080Ti GPU. Haibao Chen, Ngai Wong 0001, Hao Yu 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2020 | Fastened CROWN: Tightened Neural Network Robustness CertificatesabstractThe rapid growth of deep learning applications in real life is accompanied by severe safety concerns. To mitigate this uneasy phenomenon, much research has been done providing reliable evaluations of the fragility level in different deep neural networks. Apart from devising adversarial attacks, quantifiers that certify safeguarded regions have also been designed in the past five years. The summarizing work in (Salman et al. 2019) unifies a family of existing verifiers under a convex relaxation framework. We draw inspiration from such work and further demonstrate the optimality of deterministic CROWN (Zhang et al. 2018) solutions in a given linear programming problem under mild constraints. Given this theoretical result, the computationally expensive linear programming based method is shown to be unnecessary. We then propose an optimization-based approach FROWN (Fastened CROWN): a general algorithm to tighten robustness certificates for neural networks. Extensive experiments on various networks trained individually verify the effectiveness of FROWN in safeguarding larger robust regions. Zhaoyang Lyu, Ching Yun Ko, Zhifeng Kong, Ngai Wong 0001, Dahua Lin, Luca Daniel |
AAAI | 4 |
| 2020 | A Portable Hong Kong Sign Language Translation Platform with Deep Learning and Jetson NanoabstractAs hearing loss is arousing more and more public concern, different researches have been conducted on translating the sign language into spoken language. However, most of these researches remain in a theoretical level and few of them investigate how to realize a real system. In this paper, we introduce an effective and portable Hong Kong sign language recognition platform which can translate the Hong Kong sign language within a few seconds. In this platform, there are mainly two parts: a mobile application and a Jetson Nano. The mobile application accounts for preprocessing the sign video and transferring the videos to Jetson Nano. Then, Jetson Nano will translate sign videos into spoken language with the pretrained deep learning model and return the results to the mobile application. With this platform, non-disabled people can easily translate and understand the sign performed by deaf people through mobile phones quickly. We believe that this platform can significantly facilitate the daily communication between deaf people and the others in Hong Kong. Zhenxing Zhou, Yisiang Neo, King-Shan Lui, Vincent W. L. Tam, Edmund Y. Lam, Ngai Wong 0001 |
ASSETS | 6 |
| 2020 | An Anomaly Comprehension Neural Network for Surveillance Videos on Terminal DevicesabstractAnomaly comprehension in surveillance videos is more challenging than detection. This work introduces the design of a lightweight and fast anomaly comprehension neural network. For comprehension, a spatio-temporal LSTM model is developed based on the structured, tensorized time-series features extracted from surveillance videos. Deep compression of network size is achieved by tensorization and quantization for the implementation on terminal devices. Experiments on large-scale video anomaly dataset UCF-Crime demonstrate that the proposed network can achieve an impressive inference speed of 266 FPS on a GTX-1080Ti GPU, which is 4.29 faster than ConvLSTM-based method; a 3.34% AUC improvement with 5.55% accuracy niche versus the 3D-CNN based approach; and at least 15k× parameter reduction and 228× storage compression over the RNN-based approaches. Moreover, the proposed framework has been realized on an ARM-core based IOT board with only 2.4W power consumption. Guangtai Huang, Peining Zhen, Haibao Chen, Ngai Wong 0001, Hao Yu 0001 |
DATE | 6 |
| 2020 | Learn to Floorplan through Acquisition of Effective Local Search HeuristicsabstractAutomatic heuristic design through reinforcement learning opens a promising direction for solving computationally difficult problems. Unlike most previous works that aimed at solution construction, we explore the possibility of acquiring local search heuristics through massive search experiments. To illustrate the applicability, an agent is trained to perform a walk in the search space by selecting a candidate neighbor solution at each step. Specifically, we target the floorplanning problem, where a neighbor solution is generated through perturbing the sequence pair encoding of a floorplan. Experimental results demonstrate the efficacy of the acquired heuristics as well as the potential of automatic heuristic design. Zhuolun He, Yuzhe Ma, Peiyu Liao, Ngai Wong 0001, Bei Yu 0001, Martin D. F. Wong |
ICCD | 5 |
| 2020 | Exploiting Elasticity in Tensor Ranks for Compressing Neural NetworksabstractElasticities in depth, width, kernel size and resolution have been explored in compressing deep neural networks (DNNs). Recognizing that the kernels in a convolutional neural network (CNN) are 4-way tensors, we further exploit a new elasticity dimension along the input-output channels. Specifically, a novel nuclear-norm rank minimization factorization (NRMF) approach is proposed to dynamically and globally search for the reduced tensor ranks during training. Correlation between tensor ranks across multiple layers is revealed, and a graceful tradeoff between model size and accuracy is obtained. Experiments then show the superiority of NRMF over the previous non-elastic variational Bayesian matrix factorization (VBMF) scheme. Jie Ran, Hayden Kwok-Hay So, Graziano Chesi, Ngai Wong 0001 |
ICPR | 5 |
| 2020 | DEEPEYE: A Deeply Tensor-Compressed Neural Network for Video Comprehension on Terminal DevicesabstractVideo object detection and action recognition typically require deep neural networks (DNNs) with huge number of parameters. It is thereby challenging to develop a DNN video comprehension unit in resource-constrained terminal devices. In this article, we introduce a deeply tensor-compressed video comprehension neural network, called DEEPEYE, for inference on terminal devices. Instead of building a Long Short-Term Memory (LSTM) network directly from high-dimensional raw video data input, we construct an LSTM-based spatio-temporal model from structured, tensorized time-series features for object detection and action recognition. A deep compression is achieved by tensor decomposition and trained quantization of the time-series feature-based LSTM network. We have implemented DEEPEYE on an ARM-core-based IOT board with 31 FPS consuming only 2.4W power. Using the video datasets MOMENTS, UCF11 and HMDB51 as benchmarks, DEEPEYE achieves a 228.1× model compression with only 0.47% mAP reduction; as well as 15 k × parameter reduction with up to 8.01% accuracy improvement over other competing approaches. Guangya Li, Ngai Wong 0001, Haibao Chen, Hao Yu 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Fast and Accurate Tensor Completion With Total Variation Regularized Tensor TrainsabstractWe propose a new tensor completion method based on tensor trains. The to-be-completed tensor is modeled as a low-rank tensor train, where we use the known tensor entries and their coordinates to update the tensor train. A novel tensor train initialization procedure is proposed specifically for image and video completion, which is demonstrated to ensure fast convergence of the completion algorithm. The tensor train framework is also shown to easily accommodate Total Variation and Tikhonov regularization due to their low-rank tensor train representations. Image and video inpainting experiments verify the superiority of the proposed scheme in terms of both speed and scalability, where a speedup of up to 155× is observed compared to state-of-the-art tensor completion methods at a similar accuracy. Moreover, we demonstrate the proposed scheme is especially advantageous over existing algorithms when only tiny portions (say, 1%) of the to-be-completed images/videos are known. Ching Yun Ko, Kim Batselier, Luca Daniel, Wenjian Yu, Ngai Wong 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Deep Model Compression and Inference Speedup of Sum-Product Networks on Tensor TrainsabstractSum-product networks (SPNs) constitute an emerging class of neural networks with clear probabilistic semantics and superior inference speed over other graphical models. This brief reveals an important connection between SPNs and tensor trains (TTs), leading to a new canonical form which we call tensor SPNs (tSPNs). Specifically, we demonstrate the intimate relationship between a valid SPN and a TT. For the first time, through mapping an SPN onto a tSPN and employing specially customized optimization techniques, we demonstrate improvements up to a factor of 100 on both model compression and inference speedup for various data sets with negligible loss in accuracy. Ching Yun Ko, Cong Chen 0003, Zhuolun He, Kim Batselier, Ngai Wong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2019 | DEEPEYE: A Deeply Tensor-Compressed Neural Network Hardware Accelerator: Invited PaperabstractVideo detection and classification constantly involve high dimensional data that requires a deep neural network (DNN) with huge number of parameters. It is thereby quite challenging to develop a DNN video comprehension at terminal devices. In this paper, we introduce a deeply tensor compressed video comprehension neural network called DEEPEYE for inference at terminal devices. Instead of building a Long Short-Term Memory (LSTM) network directly from raw video data, we build a LSTM-based spatio-temporal model from tensorized time-series features for object detection and action recognition. Moreover, a deep compression is achieved by tensor decomposition and trained quantization of the time-series feature-based spatio-temporal model. We have implemented DEEPEYE on an ARM-core based IOT board with only 2.4W power consumption. Using the video datasets MOMENTS and UCF11 as benchmarks, DEEPEYE achieves a 228.1× model compression with only 0.47% mAP deduction; as well as 15k× parameter reduction yet 16.27% accuracy improvement. Guangya Li, Ngai Wong 0001, Haibao Chen, Hao Yu 0001 |
ICCAD | 3 |
| 2019 | POPQORN: Quantifying Robustness of Recurrent Neural NetworksabstractThe vulnerability to adversarial attacks has been a critical issue for deep neural networks. Addressing this issue requires a reliable way to evaluate the robustness of a network. Recently, several methods have been developed to compute robustness quantification for neural networks, namely, certified lower bounds of the minimum adversarial perturbation. Such methods, however, were devised for feed-forward networks, e.g. multi-layer perceptron or convolutional networks. It remains an open problem to quantify robustness for recurrent networks, especially LSTM and GRU. For such networks, there exist additional challenges in computing the robustness quantification, such as handling the inputs at multiple steps and the interaction between gates and states. In this work, we propose POPQORN (Propagated-output Quantified Robustness for RNNs), a general algorithm to quantify robustness of RNNs, including vanilla RNNs, LSTMs, and GRUs. We demonstrate its effectiveness on different network architectures and show that the robustness quantification on individual steps can lead to new insights. Ching Yun Ko, Zhaoyang Lyu, Lily Weng, Luca Daniel, Ngai Wong 0001, Dahua Lin |
ICML | 5 |
| 2019 | MiSC: Mixed Strategies CrowdsourcingabstractPopular crowdsourcing techniques mostly focus on evaluating workers' labeling quality before adjusting their weights during label aggregation. Recently, another cohort of models regard crowdsourced annotations as incomplete tensors and recover unfilled labels by tensor completion. However, mixed strategies of the two methodologies have never been comprehensively investigated, leaving them as rather independent approaches. In this work, we propose MiSC ( Mixed Strategies Crowdsourcing), a versatile framework integrating arbitrary conventional crowdsourcing and tensor completion techniques. In particular, we propose a novel iterative Tucker label aggregation algorithm that outperforms state-of-the-art methods in extensive experiments. Ching Yun Ko, Ngai Wong 0001 |
IJCAI | 4 |
| 2019 | Matrix Product Operator Restricted Boltzmann MachinesabstractA restricted Boltzmann machine (RBM) learns a probability distribution over its input samples and has numerous uses like dimensionality reduction, classification and generative modeling. Conventional RBMs accept vectorized data that dismiss potentially important structural information in the original tensor (multi-way) input. Matrix-variate and tensor-variate RBMs, named MvRBM and TvRBM, have been proposed but are all restrictive by model construction and have weak model expression power. This work presents the matrix product operator RBM (MPORBM) that utilizes a tensor network generalization of Mv/TvRBM, preserves input formats in both the visible and hidden layers, and results in higher expressive power. A novel training algorithm integrating contrastive divergence and an alternating optimization procedure is also developed. Numerical experiments compare the MPORBM with the traditional RBM and MvRBM for data classification and image completion and denoising tasks. The expressive power of the MPORBM as a function of the MPO-rank is also investigated. Cong Chen 0003, Kim Batselier, Ching Yun Ko, Ngai Wong 0001 |
IJCNN | 4 |
| 2019 | A Support Tensor Train MachineabstractThere has been growing interest in extending traditional vector-based machine learning techniques to their tensor forms. Support tensor machine (STM) and support Tucker machine (STuM) are two typical tensor generalization of the conventional support vector machine (SVM). However, the expressive power of STM is restrictive due to its rank-one tensor constraint, and STuM is not scalable because of the exponentially sized Tucker core tensor. To overcome these limitations, we introduce a novel and effective support tensor train machine (STTM) by employing a general and scalable tensor train as the parameter model. Experiments validate and confirm the superiority of the STTM over SVM, STM and STuM. Cong Chen 0003, Kim Batselier, Ching Yun Ko, Ngai Wong 0001 |
IJCNN | 4 |
| 2018 | Parallelized Tensor Train Learning of Polynomial ClassifiersabstractIn pattern classification, polynomial classifiers are well-studied methods as they are capable of generating complex decision surfaces. Unfortunately, the use of multivariate polynomials is limited to kernels as in support-vector machines, because polynomials quickly become impractical for high-dimensional problems. In this paper, we effectively overcome the curse of dimensionality by employing the tensor train (TT) format to represent a polynomial classifier. Based on the structure of TTs, two learning algorithms are proposed, which involve solving different optimization problems of low computational complexity. Furthermore, we show how both regularization to prevent overfitting and parallelization, which enables the use of large training sets, are incorporated into these methods. The efficiency and efficacy of our tensor-based polynomial classifier are then demonstrated on the two popular data sets U.S. Postal Service and Modified NIST. Zhongming Chen, Kim Batselier, Johan A. K. Suykens, Ngai Wong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | An efficient homotopy-based Poincaré-Lindstedt method for the periodic steady-state analysis of nonlinear autonomous oscillatorsabstractThe periodic steady-state analysis of nonlinear systems has always been an important topic in electronic design automation (EDA). For autonomous systems, the mainstream approaches, like shooting Newton and harmonic balance, are difficult to employ since the period itself becomes an unknown. This paper presents an innovative state-space homotopy-based Poincaré-Lindstedt method, with a novel Padé approximation of the stretched time axis, that effectively overcomes this hurdle. Examples demonstrate the excellent efficiency and scalability of the proposed approach. Zhongming Chen, Kim Batselier, Ngai Wong 0001 |
ASP-DAC | 4 |
| 2017 | Tensor Computation: A New Framework for High-Dimensional Problems in EDAabstractMany critical electronic design automation (EDA) problems suffer from the curse of dimensionality, i.e., the very fast-scaling computational burden produced by large number of parameters and/or unknown variables. This phenomenon may be caused by multiple spatial or temporal factors (e.g., 3-D field solvers discretizations and multirate circuit simulation), nonlinearity of devices and circuits, large number of design or optimization parameters (e.g., full-chip routing/placement and circuit sizing), or extensive process variations (e.g., variability /reliability analysis and design for manufacturability). The computational challenges generated by such high-dimensional problems are generally hard to handle efficiently with traditional EDA core algorithms that are based on matrix and vector computation. This paper presents “tensor computation” as an alternative general framework for the development of efficient EDA algorithms and tools. A tensor is a high-dimensional generalization of a matrix and a vector, and is a natural choice for both storing and solving efficiently high-dimensional EDA problems. This paper gives a basic tutorial on tensors, demonstrates some recent examples of EDA applications (e.g., nonlinear circuit modeling and high-dimensional uncertainty quantification), and suggests further open EDA problems where the use of tensor computation could be of advantage. Zheng Zhang 0005, Kim Batselier, Luca Daniel, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2016 | STORM: A nonlinear model order reduction method via symmetric tensor decompositionabstractNonlinear model order reduction has always been a challenging but important task in various science and engineering fields. In this paper, a novel symmetric tensor-based order-reduction method (STORM) is presented for simulating large-scale nonlinear systems. The multidimensional data structure of symmetric tensors, as the higher order generalization of symmetric matrices, is utilized for the effective capture of high-order nonlinearities and efficient generation of compact models. Compared to the recent tensor-based nonlinear model order reduction (TNMOR) algorithm [1], STORM shows advantages in two aspects. First, STORM avoids the assumption of the existence of a low-rank tensor approximation. Second, with the use of the symmetric tensor decomposition, STORM allows significantly faster computation and less storage complexity than TNMOR. Numerical experiments demonstrate the superior computational efficiency and accuracy of STORM against existing nonlinear model order reduction methods. Kim Batselier, Yu-Kwong Kwok, Ngai Wong 0001 |
ASP-DAC | 5 |
| 2016 | A tensor-based volterra series black-box nonlinear system identification and simulation frameworkabstractTensors are a multi-linear generalization of matrices to their d-way counterparts, and are receiving intense interest recently due to their natural representation of high-dimensional data and the availability of fast tensor decomposition algorithms. Given the input-output data of a nonlinear system/circuit, this paper presents a non-linear model identification and simulation framework built on top of Volterra series and its seamless integration with tensor arithmetic. By exploiting partially-symmetric polyadic decompositions of sparse Toeplitz tensors, the proposed framework permits a pleasantly scalable way to incorporate high-order Volterra kernels. Such an approach largely eludes the curse of dimensionality and allows computationally fast modeling and simulation beyond weakly non-linear systems. The black-box nature of the model also hides structural information of the system/circuit and encapsulates it in terms of compact tensors. Numerical examples are given to verify the efficacy, efficiency and generality of this tensor-based modeling and simulation framework. Kim Batselier, Zhongming Chen, Ngai Wong 0001 |
ICCAD | 4 |
| 2015 | STAVES: Speedy Tensor-Aided Volterra-Based Electronic SimulatorabstractVolterra series is a powerful tool for black-box macro-modeling of nonlinear devices. However, the exponential complexity growth in storing and evaluating higher order Volterra kernels has limited so far its employment on complex practical applications. On the other hand, tensors are a higher order generalization of matrices that can naturally and efficiently capture multi-dimensional data. Significant computational savings can often be achieved when the appropriate low-rank tensor decomposition is available. In this paper we exploit a strong link between tensors and frequency-domain Volterra kernels in modeling nonlinear systems. Based on such link we have developed a technique called speedy tensor-aided Volterra-based electronic simulator (STAVES) utilizing high-order Volterra transfer functions for highly accurate time-domain simulation of nonlinear systems. The main computational tools in our approach are the canonical tensor decomposition and the inverse discrete Fourier transform. Examples demonstrate the efficiency of the proposed method in simulating some practical nonlinear circuit structures. Xiaoyan Y. Z. Xiong, Kim Batselier, Lijun Jiang, Luca Daniel, Ngai Wong 0001 |
ICCAD | 6 |
| 2015 | Model Reduction and Simulation of Nonlinear Circuits via Tensor DecompositionabstractModel order reduction of nonlinear circuits (especially highly nonlinear circuits) has always been a theoretically and numerically challenging task. In this paper, we utilize tensors (namely, a higher order generalization of matrices) to develop a tensor-based nonlinear model order reduction algorithm we named TNMOR for the efficient simulation of nonlinear circuits. Unlike existing nonlinear model order reduction methods, in TNMOR high-order nonlinearities are captured using tensors, followed by decomposition and reduction to a compact tensor-based reduced-order model. Therefore, TNMOR completely avoids the dense reduced-order system matrices, which in turn allows faster simulation and a smaller memory requirement if relatively low-rank approximations of these tensors exist. Numerical experiments on transient and periodic steady-state analyses confirm the superior accuracy and efficiency of TNMOR, particularly in highly nonlinear scenarios. Luca Daniel, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | H-Matrix-Based Finite-Element-Based Thermal Analysis for 3D ICsabstractIn this article, we propose an efficient finite-element-based (FE-based) method for both steady and transient thermal analyses of high-performance integrated circuits based on the hierarchical matrix ( H -matrix) representation. H -matrix has been shown to provide a data-sparse way to approximate the matrices and their inverses with almost linear-space and time complexities. In this work, we apply the H -matrix concept for solving heating diffusion problems modeled by parabolic partial differential equations (PDEs) based on the finite element method. We show that the matrix from a FE-based steady and transient thermal analysis can be represented by H -matrix without any approximation, and its inverse and Cholesky factors can be evaluated by H -matrix with controlled accuracy. We then show and prove that the memory and time complexities of the solver are bounded by O ( k 1 N log N ) and O ( k 1 2 N log 2 N ), respectively, where k 1 is a small quantity determined by accuracy requirements and N is the number of unknowns in the system. The comparison with existing product-quality LU solvers, CSPARSE and UMFPACK, on a number of 3D IC thermal matrices, shows that the new method is much more memory efficient than these methods, which however prevents CPU time comparison with those methods on large examples. But the proposed method can solve all the given thermal circuits with decent scalabilities, which shows good agreement with the predicted theoretical results. Haibao Chen, Ying-Chi Li, Sheldon X.-D. Tan, Xin Huang 0003, Hai Wang 0002, Ngai Wong 0001 |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2014 | Efficient matrix exponential method based on extended Krylov subspace for transient simulation of large-scale linear circuitsabstractMatrix exponential (MEXP) method has been demonstrated to be a competitive candidate for transient simulation of very large-scale integrated circuits. Nevertheless, the performance of MEXP based on ordinary Krylov subspace is unsatisfactory for stiff circuits, wherein the underlying Arnoldi process tends to oversample the high magnitude part of the system spectrum while undersampling the low magnitude part that is important to the final accuracy. In this work we explore the use of extended Krylov subspace to generate more accurate and efficient approximation for MEXP. We also develop a formulation that allows unequal positive and negative dimensions in the generated Krylov subspace for better performance. Numerical results demonstrate the efficacy of the proposed method. Quan Chen 0007, Ngai Wong 0001 |
ASP-DAC | 3 |
| 2014 | An Efficient Two-level DC Operating Points Finder for Transistor CircuitsabstractDC analysis, as a foundation for the simulation of many electronic circuits, is concerned with locating DC operating points. In this paper, a new and efficient algorithm to find all DC operating points is proposed for transistor circuits. The novelty of this DC operating points finder is its two-level simple implementation based on the affine arithmetic preconditioning and interval contraction method. Compared to traditional methods such as homotopy, this finder offers a dramatically faster way of computing all roots, without sacrificing any accuracy. Explicit numerical examples and comparative analysis are given to demonstrate the feasibility and accuracy of the proposed approach. Kim Batselier, Ngai Wong 0001 |
DAC | 4 |
| 2014 | A novel linear algebra method for the determination of periodic steady states of nonlinear oscillatorsabstractPeriodic steady-state (PSS) analysis of nonlinear oscillators has always been a challenging task in circuit simulation. We present a new way that uses numerical linear algebra to identify the PSS(s) of nonlinear circuits. The method works for both autonomous and excited systems. Using the harmonic balancing method, the solution of a nonlinear circuit can be represented by a system of multivariate polynomials. Then, a Macaulay matrix based root-finder is used to compute the Fourier series coefficients. The method avoids the difficult initial guess problem of existing numerical approaches. Numerical examples show the accuracy and feasibility over existing methods. Kim Batselier, Ngai Wong 0001 |
ICCAD | 3 |
| 2013 | Piecewise-polynomial associated transform macromodeling algorithm for fast nonlinear circuit simulationabstractWe present a piecewise-polynomial based associated transform algorithm (PWPAT) for macromodeling nonlinear circuits in system-level circuit design. The generated reduced model can provide both global and local accuracies with the most compact dimension. Numerical examples compare it with existing algorithms and verify its superior accuracy in higher order harmonics simulation over traditional Trajectory Piecewise-Linear (TPWL) approach. Neric Fong, Ngai Wong 0001 |
ASP-DAC | 3 |
| 2013 | Integration of a wireless sensor network project for introductory circuits and systems teachingabstractThis paper presents an integration of a wireless sensor network design project in an introductory course about circuits and systems. In the project, students will design a wireless sensor network that constitutes of sensors, for a creative surveillance application. Through a versatile project vehicle, project-oriented learning modules, a comprehensive assessment strategy and public learning communities, students can learn contemporary concepts of circuits and systems from the system perspective, as well as develop ability to design a basic electronic system. Chi-Un Lei, Ngai Wong 0001, Ka Lok Man |
ISCAS | 2 |
| 2013 | A hybrid MPPT method for Photovoltaic systems via estimation and revision methodabstractMaximum Power Point Tracking (MPPT) methods can be classified into direct and indirect approaches. They are used to improve the efficiency of power conversion in Photovoltaic (PV) systems. However, a review of present literature implies that the indirect methods never produce accurate results. Meanwhile, the conventional direct Perturb and Observe (P&O) method has two problems: oscillations at steady state and slow dynamic response under changing environment conditions. Estimation and Revision (ER) method is proposed in this paper to overcome these limitations by the alternative use of MPP estimation and MPP revision process. The efficiency of the ER method is verified in an MPPT system implemented with a specific DC-DC converter and an adopted PV module. Jieming Ma, Ka Lok Man, T. O. Ting, Chi-Un Lei, Ngai Wong 0001 |
ISCAS | 6 |
| 2013 | Low-cost global MPPT scheme for Photovoltaic systems under partially shaded conditionsabstractMaximum Power Point Tracking (MPPT) is a technique applied to improve the efficiency of power conversion in Photovoltaic (PV) systems. Under partially shadowed conditions, the Power-Voltage (P-V) characteristic exhibits multiple peaks and the existing MPPT methods such as the Perturb and Observe (P&O) are incapable of searching for the Global Maximum Power Point (GMPP). This paper proposes a low-cost on-line MPPT scheme to overcome this drawback. By using hybrid numerical searching process, the operating point approaches Local Maximum Power Points (LMPPs) gradually and the GMPP is caught by comparing all the LMPPs. Simulation results prove the effectiveness and correctness of the proposed method. Jieming Ma, Ka Lok Man, T. O. Ting, Chi-Un Lei, Ngai Wong 0001 |
ISCAS | 6 |
| 2013 | A Numerically Efficient Formulation for Time-Domain Electromagnetic-Semiconductor Cosimulation for Fast-Transient SystemsabstractWe report recent progress in developing a numerically efficient formulation for electromagnetic-technology computer-aided design cosimulation for fast-transient computations. The difficulties underlying the currently existing transient formulation stemming from the vector potential-scalar potential (A-V) framework are analyzed. A time-domain electric field-scalar potential (E-V) framework is then developed via equation and variable transformations. This results in better-conditioned systems that are friendly to iterative solutions at fast switching times. Numerical examples show that the proposed E-V solver renders a useful tool for addressing multidomain simulation. Quan Chen 0007, Wim Schoenmaker, Lijun Jiang, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2013 | Autonomous Volterra Algorithm for Steady-State Analysis of Nonlinear CircuitsabstractWe present a novel algorithm, named autonomous Volterra (AV), that achieves efficient steady-state analysis of nonlinear circuits. With elegant analytic forms and availability of efficient solvers, AV constitutes a competitive steady-state algorithm besides the two mainstreams, namely, shooting Newton (SN) and harmonic balance (HB). Nonlinear systems are first captured in nonlinear differential algebraic equations, followed by expansion into linear Volterra subsystems. A key step of steady-state analysis lies in modeling each Volterra subsystem with autonomous nonlinear inputs. The steady-state solution of these subsystems then proceeds with a series of Sylvester equation solves, completely avoiding the guesses of initial condition and time stepping as in SN, as well as the uncertain length of Fourier series as in HB. Error control in AV is also straightforward by monitoring the norms of the Sylvester equation solutions. We further demonstrate that AV is readily parallelizable with superior scalability toward large-scale problems. Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Weakly nonlinear circuit analysis based on fast multidimensional inverse Laplace transformabstractThere have been continuing thrusts in developing efficient modeling techniques for circuit simulation. However, most circuit simulation methods are time-domain solvers. In this paper we propose a frequency-domain simulation method based on Laguerre function expansion. The proposed method handles both linear and nonlinear circuits. The Laguerre method can invert multidimensional Laplace transform efficiently with a high accuracy, which is a key step of the proposed method. Besides, an adaptive mesh refinement (AMR) technique is developed and its parallel implementation is introduced to speed up the computation. Numerical examples show that our proposed method can accurately simulate large circuits while enjoying low computation complexity. Yuanzhe Wang, Ngai Wong 0001 |
ASP-DAC | 4 |
| 2012 | Fast nonlinear model order reduction via associated transforms of high-order volterra transfer functionsabstractWe present a new and fast way of computing the projection matrices serving high-order Volterra transfer functions in the context of (weakly and strongly) nonlinear model order reduction. The novelty is to perform, for the first time, the association of multivariate (Laplace) variables in high-order multiple-input multiple-output (MIMO) transfer functions to generate the standard single-s transfer functions. The consequence is obvious: instead of finding projection subspaces about every si, only that about a single s is required. This translates into drastic saving in computation and memory, and much more compact reduced-order nonlinear models, without compromising any accuracy. Qing Wang 0051, Neric Fong, Ngai Wong 0001 |
DAC | 5 |
| 2012 | An operational matrix-based algorithm for simulating linear and fractional differential circuitsabstractWe present a new time-domain simulation algorithm (named OPM) based on operational matrices, which naturally handles system models cast in ordinary differential equations (ODEs), differential algebraic equations (DAEs), high-order differential equations and fractional differential equations (FDEs). When applied to simulating linear systems (represented by ODEs or DAEs), OPM has similar performance to advanced transient analysis methods such as trapezoidal or Gear's method in terms of complexity and accuracy. On the other hand, OPM naturally handles FDEs without much extra effort, which can not be efficiently solved using existing time-domain methods. High-order differential systems, being special cases of FDEs, can also be simulated using OPM. Moreover, adaptive time step can be utilized in OPM to provide a more flexible simulation with low CPU time. Numerical results then validate OPM's wide applicability and superiority. Yuanzhe Wang, Grantham Pang, Ngai Wong 0001 |
DATE | 4 |
| 2012 | Efficient variation-aware EM-semiconductor coupled solver for the TSV structures in 3D ICabstractIn this paper, we present a variational electromagnetic-semiconductor coupled solver to assess the impacts of process variations on the 3D integrated circuit (3D IC) on-chip structures. The solver employs the finite volume method (FVM) to handle a system of equation considering both the full-wave electromagnetic effects and semiconductor effects. With a smart geometrical variation model for the FVM discretization, the solver is able to handle both small-size or large-size variations. Moreover, a weighted principle factor analysis (wPFA) technique is presented to reduce the random variables in both electromagnetic and semiconductor regions, and the spectral stochastic collocation method (SSCM) is used to generate the quadratic statistical model. Numerical results validate the accuracy and efficiency of this solver in dealing with process variations in hybrid material through-silicon via (TSV) structures. Yuanzhe Xu, Wenjian Yu, Quan Chen 0007, Lijun Jiang, Ngai Wong 0001 |
DATE | 5 |
| 2012 | Operation scheduling and architecture co-synthesis for energy-efficient dataflow computations on FPGAs (abstract only)abstractCompiling high-level user applications for execution on FPGAs often involves synthesizing dataflow graphs beyond the size of the available on-chip computational resources. One way to address this is by folding the execution of the given dataflow graphs onto an array of directly connected simple configurable processing elements (CPEs). Under this scenario, the performance and energy-efficiency of the resulting system depends not only on the mapping schedule of the compute operations on the CPEs, but also on the topology of the interconnect array that connects the CPEs. This paper presents a framework in which the operation scheduler and the underlying CPE interconnect network topology are co-optimized on a per-application basis for energy-efficient FPGA computation. Given the same application, more than 2.5x difference in energy-efficiency was achievable by the use of different common regular array topologies to connect the CPEs. Moreover, by using irregular application-specific interconnect topologies derived from a genetic algorithm, up to 50% improvement in energy-delay-product was achievable when compared to the use of even the best regular topology. The use of such framework is anticipated to serve as part of a rapid high-level FPGA application compiler since minimum hardware place-and-route is needed to generate the optimal schedule and topology. Colin Yu Lin, Ngai Wong 0001, Hayden Kwok-Hay So |
FPGA | 2 |
| 2012 | A fast time-domain EM-TCAD coupled simulation framework via matrix exponentialabstractWe present a fast time-domain multiphysics simulation framework that combines full-wave electromagnetism (EM) and carrier transport in semiconductor devices (TCAD). The proposed framework features a division of linear and nonlinear components in the EM-TCAD coupled system. The former is extracted and handled independently with high efficiency by a matrix exponential approach assisted with Krylov subspace method. The latter is treated by ordinary Newton's method yet with a much sparser Jacobian matrix that leads to substantial speedup in solving the linear system of equations. More convenient error management and adaptive control are also available through the linear and nonlinear decoupling. Quan Chen 0007, Wim Schoenmaker, Shih-Hung Weng, Chung-Kuan Cheng, Lijun Jiang, Ngai Wong 0001 |
ICCAD | 7 |
| 2012 | Circuit simulation via matrix exponential method for stiffness handling and parallel processingabstractWe propose an advanced matrix exponential method (MEXP) to handle the transient simulation of stiff circuits and enable parallel simulation. We analyze the rapid decaying of fast transition elements in Krylov subspace approximation of matrix exponential and leverage such scaling effect to leap larger steps in the later stage of time marching. Moreover, matrix-vector multiplication and restarting scheme in our method provide better scalability and parallelizability than implicit methods. The performance of ordinary MEXP can be improved up to 4.8 times for stiff cases, and the parallel implementation leads to another 11 times speedup. Our approach is demonstrated to be a viable tool for ultra-large circuit simulations (with 1.6M ~ 12M nodes) that are not feasible with existing implicit methods. Shih-Hung Weng, Quan Chen 0007, Ngai Wong 0001, Chung-Kuan Cheng |
ICCAD | 3 |
| 2012 | Co-simulation of RFIC with bondwire antenna via retarded PEEC methodabstractWe present an antenna modeling method based on partial element equivalent circuit (PEEC) theory. The antenna is modeled as an equivalent circuit of lumped circuit elements, which enables circuit co-simulation between the antenna and circuits in both time and frequency domains. Antenna radiation is captured as an equivalent radiation resistor. For verification, a 2.4-GHz transmitter with bondwire antenna was implemented in a standard digital 0.35μm CMOS technology. Measurement shows good agreement with the proposed model. This model can be applied to any other on-chip electromagnetic structures. N. H. W. Fong, David C. W. Ng, Ngai Wong 0001 |
ISCAS | 4 |
| 2012 | A Realistic Early-Stage Power Grid Verification Algorithm Based on Hierarchical ConstraintsabstractPower grid verification has become an indispensable step to guarantee a functional and robust chip design. Vectorless power grid verification methods, by solving linear programming (LP) problems under current constraints, enable worst-case voltage drop predictions at an early stage of design when the specific waveforms of current drains are unknown. In this paper, a novel power grid verification algorithm based on hierarchical constraints is proposed. By introducing novel power constraints, the proposed algorithm generates more realistic current patterns and provides less pessimistic voltage drop predictions. The model order reduction-based coefficient computation algorithm reduces the complexity of formulating the LP problems from being proportional to steps to being independent of steps. Utilizing the special hierarchical constraint structure, the submodular polyhedron greedy algorithm dramatically reduces the complexity of solving the LP problems from overO(km3) to roughlyO(kmlogkm), wherekmis the number of variables. Numerical results have shown that the proposed algorithm provides less pessimistic voltage drop prediction while at the same time achieves dramatic speedup. Yuanzhe Wang, Chung-Kuan Cheng, Grantham Pang, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2012 | Corrigendum to "A Realistic Early-Stage Power Grid Verification Algorithm Based on Hierarchical Constraints"abstractThe authors for the above titled paper were incorrectly listed. The correct list of authors is presented here. Yuanzhe Wang, Chung-Kuan Cheng, Grantham Pang, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2012 | Passivity Enforcement for Descriptor Systems Via Matrix Pencil PerturbationabstractPassivity is an important property of circuits and systems to guarantee stable global simulation. Nonetheless, nonpassive models may result from passive underlying structures due to numerical or measurement error/inaccuracy. A postprocessing passivity enforcement algorithm is therefore desirable to perturb the model to be passive under a controlled error. However, previous literature only reports such passivity enforcement algorithms for pole-residue models and regular systems (RSs). In this paper, passivity enforcement algorithms for descriptor systems (DSs, a superset of RSs) with possibly singular direct term (specifically,D+DTorI-DDT) are proposed. The proposed algorithms cover all kinds of state-space models (RSs or DSs, with direct terms being singular or nonsingular, in the immittance or scattering representation) and thus have a much wider application scope than existing algorithms. The passivity enforcement is reduced to two standard optimization problems that can be solved efficiently. The objective functions in both optimization problems are the error functions, hence perturbed models with adequate accuracy can be obtained. Numerical examples then verify the efficiency and robustness of the proposed algorithms. Yuanzhe Wang, Zheng Zhang 0005, Cheng-Kok Koh, Guoyong Shi, Grantham Pang, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2011 | Balanced truncation for time-delay systems via approximate GramiansabstractIn circuit simulation, when a large RLC network is connected with delay elements, such as transmission lines, the resulting system is a time-delay system (TDS). This paper presents a new model order reduction (MOR) scheme for TDSs with state time delays. It is the first time to reduce a TDS using balanced truncation. The Lyapunov-type equations for TDSs are derived, and an analysis of their computational complexity is presented. To reduce the computational cost, we approximate the controllability and observability Gramians in the frequency domain. The reduced-order models (ROMs) are then obtained by balancing and truncating the approximate Gramians. Numerical examples are presented to verify the accuracy and efficiency of the proposed algorithm. Qing Wang 0051, Zheng Zhang 0005, Quan Chen 0007, Ngai Wong 0001 |
ASP-DAC | 5 |
| 2011 | A moment-matching scheme for the passivity-preserving model order reduction of indefinite descriptor systems with possible polynomial partsabstractPassivity-preserving model order reduction (MOR) of descriptor systems (DSs) is highly desired in the simulation of VLSI interconnects and on-chip passives. One popular method is PRIMA, a Krylov-subspace projection approach which preserves the passivity of positive semidefinite (PSD) structured DSs. However, system passivity is not guaranteed by PRIMA when the system is indefinite. Furthermore, the possible polynomial parts of singular systems are normally not captured. For indefinite DSs, positive-real balanced truncation (PRBT) can generate passive reduced-order models (ROMs), whose main bottleneck lies in solving the dual expensive generalized algebraic Riccati equations (GAREs). This paper presents a novel moment-matching MOR for indefinite DSs, which preserves both the system passivity and, if present, also the improper polynomial part. This method only requires solving one GARE, therefore it is cheaper than existing PRBT schemes. On the other hand, the proposed algorithm is capable of preserving the passivity of indefinite DSs, which is not guaranteed by traditional moment-matching MORs. Examples are finally presented showing that our method is superior to PRIMA in terms of accuracy. Zheng Zhang 0005, Qing Wang 0051, Ngai Wong 0001, Luca Daniel |
ASP-DAC | 3 |
| 2011 | A block-diagonal structured model reduction scheme for power grid networksabstractWe propose a block-diagonal structured model order reduction (BDSM) scheme for fast power grid analysis. Compared with existing power grid model order reduction (MOR) methods, BDSM has several advantages. First, unlike many power grid reductions that are based on terminal reduction and thus error-prone, BDSM utilizes an exact column-by-column moment matching to provide higher numerical accuracy. Second, with similar accuracy and macromodel size, BDSM generates very sparse block-diagonal reduced-order models (ROMs) for massive-port systems at a lower cost, whereas traditional algorithms such as PRIMA produce full dense models inefficient for the subsequent simulation. Third, different from those MOR schemes based on extended Krylov subspace (EKS) technique, BDSM is input-signal independent, so the resulting ROM is reusable under different excitations. Finally, due to its blockdiagonal structure, the obtained ROM can be simulated very fast. The accuracy and efficiency of BDSM are verified by industrial power grid benchmarks. Zheng Zhang 0005, Chung-Kuan Cheng, Ngai Wong 0001 |
DATE | 4 |
| 2011 | Process-variation-aware electromagnetic-semiconductor coupled simulationabstractWe develop a new method based on the high-frequency electromagnetic (EM)-semiconductor coupled simulation to analyze the impact of multi-type process variations happen around semiconductor-metal structure. It is competent to simultaneously handle geometrical variations like surface roughness and material variations like semi-conductor doping profile, which are difficult for traditional "stand alone" simulation methods. A sparse grid based stochastic spectral collocation method (SSCM) combined with principle factor analysis (PFA) is implemented to accelerate the stochastic simulation. Numerical results confirm the validity and significance of our variational coupled simulation framework. Yuanzhe Xu, Quan Chen 0007, Lijun Jiang, Ngai Wong 0001 |
ISCAS | 4 |
| 2011 | More realistic power grid verification based on hierarchical current and power constraintsabstractVectorless power grid verification algorithms, by solving linear programming (LP) problems under current constraints, enable worst-case voltage drop predictions at an early design stage. However, worst-case current patterns obtained by many existing vectorless algorithms are time-invariant (i.e., are constant throughout the simulation time), which may result in an overly pessimistic voltage drop prediction. In this paper, a more realistic power grid verification algorithm based on hierarchical current and power constraints is proposed. The proposed algorithm naturally handles general RCL power grid models. Currents at different time steps are treated as independent variables and additional power constraints are introduced; this results in more realistic time-varying worst-case current patterns and less pessimistic worst-case voltage drop predictions. Moreover, a sorting-deletion algorithm is proposed to speed up solving LP problems by utilizing the hierarchical constraint structure. Experimental results confirm that worst-case current patterns and voltage drops obtained by the proposed algorithm are more realistic, and that the sorting-deletion algorithm reduces runtime needed to solve LP problems by 85%. Chung-Kuan Cheng, Andrew B. Kahng, Grantham Pang, Yuanzhe Wang, Ngai Wong 0001 |
ISPD | 6 |
| 2011 | Frequency-domain transient analysis of multitime partial differential equation systemsabstractMultitime partial differential equations (MPDEs) provide an efficient method to simulate circuits with widely separated rates of inputs. This paper proposes a fast and accurate frequency-domain multitime transient analysis method for MPDE systems, which fills in the gap for the lack of general frequency-domain solver for MPDE systems. A block-pulse function-based multidimensional inverse Laplace transform strategy is adopted. The method can be applied to discrete input systems. Numerical examples then confirm its superior accuracy, under similar efficiency, over time-domain solvers. Fengrui Shi, Yuanzhe Wang, Ngai Wong 0001 |
VLSI-SoC | 4 |
| 2011 | An Effective Formulation of Coupled Electromagnetic-TCAD Simulation for Extremely High Frequency OnwardabstractThis paper presents an effective formulation tailored for electromagnetic-technology computer-aided design coupled simulations for extremely-high-frequency ranges and beyond (>;50 GHz). A transformation of variables is exploited from the starting A-V formulation to the E-V formulation, combined with adopting the gauge condition as the equation for scalar potential. The transformation significantly reduces the cross-coupling between electric and magnetic systems at high frequencies, providing therefore much better convergence for iterative solution. The validation of such transformations is ensured through a careful analysis of redundancy in the coupled system and material properties. Employment of the advanced matrix permutation technique further alleviates the extra computational cost introduced by the variable transformation. Numerical experiments confirm the accuracy and efficiency of the proposed E-V formulation. Quan Chen 0007, Wim Schoenmaker, Peter Meuris, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2011 | A Sub-1 V, 26 muW, Low-Output-Impedance CMOS Bandgap Reference With a Low Dropout or Source Follower ModeabstractWe present a low-power bandgap reference (BGR), functional from sub-1 V to 5 V supply voltage with either a low dropout (LDO) regulator or source follower (SF) output stage, denoted as the LDO or SF mode, in a 0.5-μm standard digital CMOS process withVtn≈ 0.6 V and |Vtp| ≈ 0.7 V at 27°C. Both modes operate at sub-1 V under zero load with a power consumption of around 26 μW. At 1 V (1.1 V) supply, the LDO (SF) mode provides an output current up to 1.1 mA (0.35 mA), a load regulation of ±8.5 mV/mA (±33 mV/mA) with approximately 10 μs transient, a line regulation of ±4.2 mV/V ( ±50 μV/V), and a temperature compensated reference voltage of 0.228 V (0.235 V) with a temperature coefficient around 34 ppm/°C from -20°C to 120 °C. At 1.5 V supply, the LDO (SF) mode can further drive up to 9.6 mA (3.2 mA) before the reference voltage falls to 90% of its nominal value. Such low-supply-voltage and high-current-driving BGR in standard digital CMOS processes is highly useful in portable and switching applications. David C. W. Ng, David K. K. Kwong, Ngai Wong 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | VISA: versatile impulse structure approximation for time-domain linear macromodelingabstractWe develop a rational function macromodeling algorithm named VISA (Versatile Impulse Structure Approximation) for macromodeling of system responses with (discrete) time-sampled data. The ideas of Walsh theorem and complementary signal are introduced to convert the macromodeling problem into a non-pole-based Steiglitz-McBride (SM) iteration (a class of first- and second-order interpolations) without initial guess and eigenvalue computation. We demonstrate the fast convergence and the versatile macromodeling requirement adoption through a P-norm approximation expansion, using examples from practical data. Chi-Un Lei, Ngai Wong 0001 |
ASP-DAC | 2 |
| 2010 | An extension of the generalized Hamiltonian method to S-parameter descriptor systemsabstractA generalized Hamiltonian method (GHM) was recently proposed for the passivity test of hybrid descriptor systems. This paper extends the GHM theory to its S-parameter counterpart. Based on the S-parameter GHM, a passivity test flow is proposed, which is capable of detecting nonpassive regions of descriptor-form physical models. The proposed method is applicable to S-parameter and hybrid systems either in the standard state-space or descriptor forms. Experimental results confirm the effectiveness and accuracy of the proposed method. Zheng Zhang 0005, Ngai Wong 0001 |
ASP-DAC | 2 |
| 2010 | MFTI: matrix-format tangential interpolation for modeling multi-port systemsabstractNumerous algorithms to macromodel a linear time-invariant (LTI) system from its frequency-domain sampling data have been proposed in recent years [1, 2, 3, 4, 5, 6, 7, 8], among which Loewner matrix-based tangential interpolation proves to be especially suitable for modeling massive-port systems [6, 7, 8]. However, the existing Loewner matrix-based method follows vector-format tangential interpolation (VFTI), which fails to explore all the information contained in the frequency samples. In this paper, a novel matrix-format tangential interpolation (MFTI) is proposed, which requires much fewer samples to recover the system and yields better accuracy when handling under-sampled, noisy and/or ill-conditioned data. A recursive version of MFTI is proposed to further reduce the computational complexity. Numerical examples then confirm the superiority of MFTI over VFTI. Yuanzhe Wang, Chi-Un Lei, Grantham Pang, Ngai Wong 0001 |
DAC | 4 |
| 2010 | Design space exploration for sparse matrix-matrix multiplication on FPGAsabstractThe design and implementation of a sparse matrix-matrix multiplication architecture on FPGAs is presented. Performance of the design, in terms of computational latency, as well as the associated power-delay and energy-delay tradeoff are studied. Taking advantage of the sparsity of the input matrices, the proposed design allows user-tunable power-delay and energy-delay tradeoffs by employing different number of processing elements (PEs) in the architecture design and different block size in the blocking decomposition. Such ability allows designers to employ different on-chip computational architecture for different system power-delay and energy-delay requirements. It is in contrast to conventional dense matrix-matrix multiplication architectures that always favor the maximum number of PEs and largest block size. In our implementation, the better energy consumption and power-delay product favors less PEs and smaller block size for the 90%-sparsity matrix-matrix multiplications. While in order to achieve better energy-delay product, more PEs and larger block size are preferred. Colin Yu Lin, Zheng Zhang 0005, Ngai Wong 0001, Hayden Kwok-Hay So |
FPT | 3 |
| 2010 | PEDS: Passivity enforcement for descriptor systems via Hamiltonian-symplectic matrix pencil perturbationabstractPassivity is a crucial property of macromodels to guarantee stable global (interconnected) simulation. However, weakly nonpassive models may be generated for passive circuits and systems in various contexts, such as data fitting, model order reduction (MOR) and electromagnetic (EM) macromodeling. Therefore, a post-processing passivity enforcement algorithm is desired. Most existing algorithms are designed to handle pole-residue models. The few algorithms for state space models only handle regular systems (RSs) with a nonsingular D+DTterm. To the authors' best knowledge, no algorithm has been proposed to enforce passivity for more general descriptor systems (DSs) and state space models with singular D+DTterms. In this paper, a new post-processing passivity enforcement algorithm based on perturbation of Hamiltonian-symplectic matrix pencil, PEDS, is proposed. PEDS, for the first time, can enforce passivity for DSs. It can also handle all kinds of state space models (both RSs and DSs) with singular D+DTterms. Moreover, a criterion to control the error of perturbation is devised, with which the optimal passive models with the best accuracy can be obtained. Numerical examples then verify that PEDS is efficient, robust and relatively cheap for passivity enforcement of DSs with mild passivity violations. Yuanzhe Wang, Zheng Zhang 0005, Cheng-Kok Koh, Grantham Pang, Ngai Wong 0001 |
ICCAD | 5 |
| 2010 | An Efficient Projector-Based Passivity Test for Descriptor SystemsabstractAn efficient passivity test based on canonical projector techniques is proposed for descriptor systems (DSs) widely encountered in circuit and system modeling. The test features a natural flow that first evaluates the index of a DS, followed by possible decoupling into its proper and improper subsystems. Explicit state-space formulations for respective subsystems are derived to facilitate further processing such as model order reduction and/or passivity enforcement. Efficient projector construction and a fast generalized Hamiltonian test for the proper-part passivity are also elaborated. Numerical examples then confirm the superiority of the proposed method over existing passivity tests for DSs based on linear matrix inequalities or skew-Hamiltonian/Hamiltonian matrix pencils. Zheng Zhang 0005, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | An efficient passivity test for descriptor systems via canonical projector techniquesabstractAn efficient passivity test based on matrix chain and canonical projector techniques is proposed for descriptor systems (DSs) commonly encountered in VLSI modeling. The test entails a natural flow that first evaluates the index of a DS, followed by possible decoupling into its proper and improper (if any) subsystems. Compared to the recent DS passivity tests using linear matrix inequality (LMI)-based extended positive real lemma or skew-Hamiltonian/Hamiltonian (SHH) transformations, numerical examples show that the proposed projector approach is significantly more efficient. Explicit state space formulations for the decoupled subsystems that facilitate further processing such as passivity enforcement and model order reduction are also derived. Ngai Wong 0001 |
DAC | 1 |
| 2009 | New simulation methodology of 3D surface roughness loss for interconnects modelingabstractAs clock frequencies exceed giga-Hertz, the extra power loss due to conductor surface roughness in interconnects and packagings is more evident and thus demands a proper accounting for accurate prediction of signal integrity and energy consumption. Existing techniques based on analytical approximation often suffer from a narrow valid range, i.e., small or large limit of roughness. In this paper, we propose a new simulation methodology for surface roughness loss that is applicable to general surface roughness and a wide frequency range. The method is based on 3D statistical modeling of surface roughness and the numerical solution of scalar wave modeling (SWM) with the method of moments (MOM). The spectral stochastic collocation method (SSCM) is applied in association of random surface modeling to avoid the time-consuming Monte-Carlo (MC) simulation. Comparisons with existing methods in their respective valid region then verify the effectiveness of our approach. Quan Chen 0007, Ngai Wong 0001 |
DATE | 2 |
| 2009 | Operation scheduling for FPGA-based reconfigurable computersabstractMany high-performance applications involve large data sets that are impossible to fit entirely within on-chip memories of even the largest FPGAs. As a result, they must be stored in off-chip SDRAMs and loaded onto the FPGAs as computations progress. Because of the high latency and energy consumption associated with off-chip memory accesses, it is important to develop efficient operation schedules that not only minimize latency of computations, but also the amount of data I/Os. We formulate this problem as a modified resource-constrained job scheduling problem. The problem is then solved using a list scheduling algorithm that takes advantage of the fast burst-mode access of SDRAMs. Results have shown that for large problem sizes, the performance of our algorithm is within 1% of a hand-optimized matrix-matrix multiplication implementation, with no memory overhead, and is within 0.03% of the theoretical minimum latency of an 8-by-8 cofactor matrix computation. Colin Yu Lin, Ngai Wong 0001, Hayden Kwok-Hay So |
FPL | 2 |
| 2009 | GHM: A generalized Hamiltonian method for passivity test of impedance/admittance descriptor systemsabstractA generalized Hamiltonian method (GHM) is proposed for passivity test of descriptor systems (DSs) which describe impedance or admittance input-output responses. GHM can test passivity of DSs with any system index without minimal realization. This frequency-independent method can avoid the time-consuming system decomposition as required in many existing DS passivity test approaches. Furthermore, GHM can test systems with singular D + DT where traditional Hamiltonian method fails, and enjoys a more accurate passivity violation identification compared to frequency sweeping techniques. Numerical results have verified the effectiveness of GHM. The proposed method constitutes a versatile tool to speed up passivity check and enforcement of DSs and subsequently ensures globally stable simulations of electrical circuits and components. Zheng Zhang 0005, Chi-Un Lei, Ngai Wong 0001 |
ICCAD | 3 |
| 2009 | IIR Approximation of FIR Filters Via Discrete-Time Hybrid-Domain Vector FittingabstractWe present a discrete-time hybrid-domain vector fitting algorithm, called HD-VFz, for the IIR approximation of FIR filters with an arbitrary combination of time-and frequency-sampled responses. The core routine involves a two-step pole refinement process based on a linear least-squares solve and an eigenvalue problem. Through hybrid-domain data approximation and digital partial fraction basis with relative stability consideration, HD-VFz exhibits fast computation and remarkable fitting accuracy in both time and frequency domains. Chi-Un Lei, Ngai Wong 0001 |
IEEE Signal Process. Lett. | 2 |
| 2009 | Robust Simulation Methodology for Surface-Roughness Loss in Interconnect and Package ModelingsabstractIn multigigahertz integrated-circuit design, the extra energy loss caused by conductor surface roughness in metallic interconnects and packagings is more evident than ever before and demands explicit consideration for accurate prediction of signal integrity and energy consumption. Existing techniques based on analytical approximation, despite simple formulations, suffer from restrictive valid ranges, namely, either small or large roughness/frequencies. In this paper, we propose a robust and efficient numerical-simulation methodology applicable to evaluating general surface roughness, described by parameterized stochastic processes, across a wide frequency band. Traditional computation-intensive electromagnetic simulation is avoided via a tailored scalar-wave modeling to capture the power loss due to surface roughness. The spectral stochastic collocation method is applied to construct the complete statistical model. Comparisons with full wave simulation as well as existing methods in their respective valid ranges then verify the effectiveness of the proposed approach. Quan Chen 0007, Hoi Wai Choi, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | Efficient numerical modeling of random rough surface effects for interconnect internal impedance extractionabstractThis paper proposes an efficient model for numerically evaluating the impact of random surface roughness on the internal impedance for large-scale interconnect structures. The effective resistivity (ER) and effective permeability (EP) are numerically formulated to avoid the computationally prohibitive global discretization, while maintaining the model accuracy and flexibility. A modified stochastic integral equation (SIE) method is proposed to significantly speed up the computation for the mean values of ER and EP under the assumption of random surface roughness. Numerical experiments then verify the efficacy of our approach. Quan Chen 0007, Ngai Wong 0001 |
ASP-DAC | 2 |
| 2008 | Global optimization of common subexpressions for multiplierless synthesis of multiple constant multiplicationsabstractIn the context of multiple constant multiplication (MCM) design, we propose a novel common subexpression elimination (CSE) algorithm that models the optimal synthesis of coefficients into a 0–1 mixed-integer linear programming (MILP) problem. A time delay constraint is included for synthesis. We also propose coefficient decompositions that combine all minimal signed digit (MSD) representations and the shifted sum (difference) of coefficients. In the examples we demonstrate, the proposed solution space further reduces the number of adders/subtractors in the MCM synthesis. Yuen-Hong Alvin Ho, Chi-Un Lei, Hing-Kit Kwan, Ngai Wong 0001 |
ASP-DAC | 4 |
| 2008 | Direct sigma-delta modulated signal processing in FPGAabstractThe effectiveness of implementing bit-stream signal processing (BSSP) multiplier circuits in FPGAs, in terms of hardware resources and clock frequency, is presented. In particular, the result of realizing BSSP multipliers on FPGA architectures that utilize 6-input lookup tables (LUTs) is compared against architectures that utilize 4-input LUTs. It is found that architectures featuring 6-input LUTs suit well in BSSP applications where wide combinatorial paths are common. Furthermore, the performance of a BSSP multiplier is compared against conventional parallel multipliers in terms of LUT resource requirements. For a given resource requirement, it is found that an over-sampling ratio of less than 32 is required for a BSSP multiplier to outperform its parallel counterpart. Chiu-Wah Ng, Ngai Wong 0001, Hayden Kwok-Hay So, Tung-Sang Ng |
FPL | 2 |
| 2008 | Quad-level bit-stream signal processing on FPGAsabstractQuad-level bit-stream signal processing (BSSP) circuits are implemented and their performances are compared with previously published tri-level and bi-level BSSP implementations on FPGAs. BSSP refers to the process of performing computation directly on over-sampled delta-sigma modulated signals to eliminate the need of resource consuming decimators and interpolators. Quad-level BSSP offers better performance than their bi-and tri-level counterparts at the expense of higher resource utilization. Using a digital phase locked loop (DPLL) and a quadrature phase-shift keying (QPSK) demodulator as application examples, the effectiveness of quad-level BSSP on FPGAs is studied. The BSSP approach will be contrasted with conventional multi-bit implementations using built-in digital signal processing blocks in modern FPGAs. Chiu-Wah Ng, Ngai Wong 0001, Hayden Kwok-Hay So, Tung-Sang Ng |
FPT | 2 |
| 2008 | Design of hybrid continuous-time discrete-time delta-sigma modulatorsabstractRecent attention has been drawn to the hybrid Delta-Sigma (ΔΣ) structure featuring the integration of continuous-time (CT) and discrete-time (DT) structures in the loop filter. It combines the accurate loop filter characteristic of a DT ΔΣ modulator and the inherent anti-aliasing of a CT modulator. We present a design methodology for building a CT-DT ΔΣ modulator via the transformation from a DT ΔΣ modulator prototype. We also demonstrate the tradeoff of applying this structure to cascaded Delta-Sigma modulators compared to pure CT or DT implementations. Hing-Kit Kwan, Siu-Hong Lui, Chi-Un Lei, Ngai Wong 0001, Ka-Leung Ho |
ISCAS | 5 |
| 2008 | Efficient linear macromodeling via least-squares response approximationabstractWe present a least-squares (LS) algorithm for rational function macromodeling of port-to-port responses with discrete-time sampled data. The core routine involves over-determined equations and filtering operation, and avoids numerical-sensitive calculation and initial pole assignment. We demonstrate the fast computation and excellent accuracy and robustness, even with noisy data, in stable response approximation. Chi-Un Lei, Hing-Kit Kwan, Ngai Wong 0001 |
ISCAS | 4 |
| 2008 | Efficient Positive-Real Balanced Truncation of Symmetric Systems Via Cross-Riccati EquationsabstractWe present a highly efficient approach for realizing a positive-real balanced truncation (PRBT) of symmetric systems. The solution of a pair of dual algebraic Riccati equations in conventional PRBT, whose cost constrains practical large-scale deployment, is reduced to the solution of one cross-Riccati equation (XRE). The cross-Riccatian nature of the solution then allows a simple construction of PRBT projection matrices, using a Schur decomposition, without actual balancing. An invariant subspace method and a modified quadratic alternating-direction-implicit iteration scheme are proposed to efficiently solve the XRE. A low-rank variant of the latter is shown to offer a remarkably fast PRBT speed over the conventional implementations. The XRE-based framework can be applied to a large class of linear passive networks, and its effectiveness is demonstrated through numerical examples. Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | Fast positive-real balanced truncation of symmetric systems using cross Riccati equations
Ngai Wong 0001 |
DATE | 1 |
| 2007 | FIR Filter Approximation by IIR Filters Based on Discrete-Time Vector FittingabstractWe present a novel way of approximating FIR filters by IIR structures through generalizing the vector fitting (VF) algorithm, popularly used in continuous-time frequency-domain rational approximation, to its discrete-time counterpart called VFz. VFz progressively refines the filter poles for improved approximation. Linear equation and eigenvalue solves are used in each iteration, with real-arithmetic formulation to accommodate complex poles. Comparison against conventional algorithms confirms that VFz exhibits fast convergence and produces highly accurate IIR approximants. Ngai Wong 0001, Chi-Un Lei |
ISCAS | 1 |
| 2007 | Fast Positive-Real Balanced Truncation Via Quadratic Alternating Direction Implicit IterationabstractBalanced truncation (BT), as applied to date in model order reduction (MOR), is known for its superior accuracy and computable error bounds. Positive-real BT (PRBT) is a particular BT procedure that preserves passivity and stability and imposes no structural constraints on the original state space. However, PRBT requires solving two algebraic Riccati equations (AREs), whose computational complexity limits its practical use in large-scale systems. This paper introduces a novel quadratic extension of the alternating direction implicit (ADI) iteration, which is called quadratic ADI (QADI), that efficiently solves an ARE. A Cholesky factor version of QADI, which is called CFQADI, exploits low-rank matrices and further accelerates PRBT. Ngai Wong 0001, Venkataramanan Balakrishnan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Multi-shift quadratic alternating direction implicit iteration for high-speed positive-real balanced truncationabstractThis paper presents a multi-shift generalization of the recently proposed quadratic alternating direction implicit (QADI) iteration. QADI and its Cholesky Factor (CF) variant, CFQADI, have been shown to be efficient ways of solving the large-scale algebraic Riccati equations (AREs) required in positive-real balanced truncation (PRBT). However, only their single-shift implementations have been considered so far. Using linear fractional transformation (LFT), we present elegant multi-shift extensions of both QADI and CFQADI, thereby enabling even faster and more accurate PRBT. Ngai Wong 0001, Venkataramanan Balakrishnan |
DAC | 1 |
| 2006 | A fast passivity test for descriptor systems via structure-preserving transformations of Skew-Hamiltonian/Hamiltonian matrix pencilsabstractPassivity in a VLSI model is an important property to guarantee stable global simulation. Most VLSI models are naturally described as descriptor systems (DSs) or singular state spaces. Passivity tests for DSs, however, are much less developed compared to their non-singular state space counterparts. For large-scale DSs, the existing test based on linear matrix inequality (LMI) is computationally prohibitive. Other system decoupling techniques involve complicated coding and sometimes ill-conditioned transformations. This paper proposes a simple DS passivity test based on the key insight that the sum of a passive system and its adjoint must be impulse-free. A sidetrack shows that the proper (non-impulsive) part of a passive DS can be easily decoupled along the test flow. Numerical examples confirm the effectiveness of the proposed DS passivity test over conventional approaches. Ngai Wong 0001, C. K. Chu |
DAC | 1 |
| 2006 | Signal Detection for MIMO-OFDM Systems with Time OffsetsabstractIn this paper, signal detection for MIMO-OFDM systems operating on quasi-static frequency selective fading channels with time offsets is considered. A channel equalizer is first proposed to cancel most inter-block and inter-symbol interferences, based on second order statistics of the received signals. The transmitted signals are then detected from the equalizer output with the aid of a few pilot symbols. In most scenarios, the number of pilot symbols required here is less than that required in the existing algorithms. In addition, perfect time synchronization, time offset and channel length estimations are not needed. The algorithm is applicable irrespective of whether the channel length is shorter than, equal to or longer than the cyclic prefix (CP) length. Simulations confirm the effectiveness of the proposed algorithm and verify its robustness against time offsets, namely, different time delays for different users. Shaodan Ma, Ngai Wong 0001, Tung-Sang Ng |
GLOBECOM | 2 |
| 2006 | Time domain equalization for OFDM systemsabstractIn this paper, a time domain equalization algorithm is proposed for single-TX multiple-RX OFDM systems over frequency selective fading channels. The algorithm, which cancels most of the intersymbol interference (ISI), utilizes the orthogonality of the IFFT matrix and the second order statistics of the received signals. The signals are then detected, with the aid of only four pilots, from the equalizer output. The number of pilots required in the proposed algorithm is less than that in existing algorithms and channel length estimation is not needed. In addition, the proposed algorithm is applicable to both the case where the channel length is shorter than or equal to the length of cyclic prefix (CP), and the case where the channel length is longer than the length of cyclic prefix which results in interblock interference (IBI). Simulation results confirm the effectiveness of the proposed algorithm in both cases and indicate that it is more practical as there is no restriction on the channel and CP lengths. Shaodan Ma, Ngai Wong 0001, Tung-Sang Ng |
ISCAS | 2 |
| 2006 | Power optimization in a repeater-inserted interconnect via geometric programmingabstractWe present an innovative geometric programming (GP) approach for minimizing the power dissipation of an interconnect with repeater insertion, subject to delay, bandwidth and area constraints. Repeater sizes and segment lengths are globally optimized in various technology nodes with respect to International Technology Roadmap for Semiconductors (ITRS). Relative power dissipation due to different power components is analyzed. We show that, on average, the power dissipation per unit length can be reduced by over 30% when the timing constraint is relaxed by 5%. The optimum number of repeaters is always given as an integer in our design flow. The relationships between power dissipation and respective design constraints are easily visualized in tradeoff curves. Additional design criteria, such as reliability of the interconnect delay against process variations, are easily incorporated into the optimization. W. T. Cheung, Ngai Wong 0001 |
ISLPED | 2 |
| 2006 | Two Algorithms for Fast and Accurate Passivity-Preserving Model Order ReductionabstractThis paper presents two recently developed algorithms for efficient model order reduction. Both algorithms enable the fast solution of continuous-time algebraic Riccati equations (CAREs) that constitute the bottleneck in the passivity-preserving balanced stochastic truncation (BST). The first algorithm is a Smith-method-based Newton algorithm, called Newton/Smith CARE, that exploits low-rank matrices commonly found in physical system modeling. The second algorithm is a project-and-balance scheme that utilizes dominant eigenspace projection, followed by a simultaneous solution of a pair of dual CAREs through completely separating the stable and unstable invariant subspaces of a Hamiltonian matrix. The algorithms can be applied individually or together. Numerical examples show the proposed algorithms offer significant computational savings and better accuracy in reduced-order models over those from conventional schemes. Ngai Wong 0001, Venkataramanan Balakrishnan, Cheng-Kok Koh, Tung-Sang Ng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Fast balanced stochastic truncation via a quadratic extension of the alternating direction implicit iterationabstractBalanced truncation (BT) model order reduction (MOR) is known for its superior accuracy and computable error bounds. Balanced stochastic truncation (BST) is a particular BT procedure that provides a general, structure-independent MOR framework to preserve both passivity and stability of original models. Its application toward large scale systems, however, has been limited by the complexity of solving large size continuous time algebraic Riccati equations (CAREs). This paper introduces a novel quadratic extension of the alternating direction implicit (ADI) iteration, called QADI, that efficiently solves a CARE. A Cholesky factor variant of QADI, called CFQADI, further exploits low rank matrices and and produces solution in factor form that greatly accelerates BST. Remarkable efficiency of the proposed BST/(CF)QADI integration is demonstrated with numerical examples. Ngai Wong 0001, Venkataramanan Balakrishnan |
ICCAD | 1 |
| 2004 | Passivity-preserving model reduction via a computationally efficient project-and-balance schemeabstractThis paper presents an efficient t o-stage project-and-balance scheme for passivity-preserving model order reduction. Orthogonal dominant eigenspace projection is implemented by integrating the Smith method and Krylov subspace iteration. It is followed by stochastic balanced truncation herein a novel method, based on the complete separation of stable and unstable invariant subspaces of a Hamiltonian matrix, is used for solving two dual algebraic Riccati equations at the cost of essentially one. A fast-converging quadruple-shift bulge-chasing SR algorithm is also introduced for this purpose. Numerical examples confirm the quality of the reduced-order models over those from conventional schemes. Ngai Wong 0001, Venkataramanan Balakrishnan, Cheng-Kok Koh |
DAC | 1 |
| 2004 | A fast Newton/Smith algorithm for solving algebraic Riccati equations and its application in model order reductionabstractA very fast Smith-method-based Newton algorithm is introduced for the solution of large-scale continuous-time algebraic Riccati equations (CAREs). When the CARE contains low-rank matrices, as is common in the modeling of physical systems, the proposed algorithm, called the Newton/Smith CARE or NSCARE algorithm, offers significant computational savings over conventional CARE solvers. The effectiveness of the algorithm is demonstrated in the context of VLSI model order reduction, wherein stochastic balanced truncation (SBT) is used to reduce large-scale passive circuits. It is shown that the NSCARE algorithm exhibits guaranteed quadratic convergence under mild assumptions. Moreover, two large-sized matrix factorizations and one large-scale singular value decomposition (SVD), necessary for SBT, can be omitted by utilizing the Smith method output in each Newton iteration, thereby significantly speeding up the model reduction process. Ngai Wong 0001, Venkataramanan Balakrishnan, Cheng-Kok Koh, Tung-Sang Ng |
ICASSP (5) | 1 |
| 2003 | A geometrical approach to robust minimum variance beamformingabstractThis paper presents a highly efficient geometrical approach for designing robust minimum variance (RMV) beamformers against uncertainties in the array steering vector. Instead of the conventional approach of modeling the uncertainty region by a convex closed space, the proposed algorithm exploits the optimization constraint and shows that optimization only needs to be done on the intersection of a hyperplane and a second-order cone (SOC). The problem can then be cast as a second-order cone programming (SOCP) problem so as to enjoy the high efficiency of a class of interior point algorithms. A general case of modeling the uncertainties of an array using complex-plane trapezoids is investigated. The efficiency and tightness of the proposed method over other schemes are demonstrated with numerical examples. Ngai Wong 0001, Tung-Sang Ng, Venkataramanan Balakrishnan |
ICASSP (5) | 1 |
| 2001 | An efficient algorithm for downconverting multiple bandpass signals using bandpass samplingabstractBandpass sampling is a useful alternative for direct digital downconversion in software radio. It significantly relaxes the analog-to-digital converter (ADC) sampling rate requirement and facilitates the design goal of moving the ADC as close as possible to the antenna. This paper presents a modified interpretation to the graph of allowable sampling frequencies in uniform bandpass sampling. It is shown that the position and guard bands of a downconverted bandpass signal band are highly related to the order of the valid sampling range, called the wedge order. An efficient algorithm is then proposed, which significantly reduces the computational load in determining the valid sampling frequencies to downconvert multiple distinct bandpass signals. Conditions for the placement of bandpass signals to utilize a given sampled bandwidth are also discussed. Ngai Wong 0001, Tung-Sang Ng |
ICC | 1 |