EDBT 2026 Demo / reviewers in the wild / expert
Xinming Huang 0001
dblp:18/6823-1
· DBLP profile ↗
70ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0003-0584-3448ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 16 · 8 since 2021Computer networks · 11Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ECG Statement Classification and Lead Reconstruction Using CNN-Based ModelsabstractECG is an essential diagnostic tool that offers important insight into a person's cardiac and general health. The rise of intelligent wearable devices has opened a new avenue for clinicians and individuals to capture long term ECG data—albeit with fewer leads than the 12 leads that are typically used clinically, which can be vital for identifying and addressing health concerns. In this work, a multi-task convolutional neural network (CNN) classifier was used to study the influence of various combinations of ECG leads in interpretation of 71 cardiac statements spanning cardiac diagnostics, form, and rhythm. Results of this analysis suggest that the subset of limb leads I and II and chest leads V1, V3, and V6 can be used to identify several cardiac statements without loss of performance (average macro AUC of 0.903) when compared to a model trained using all 12- leads (average macro AUC of 0.905; p = 1). A hybrid CNNLSTM (long short-term memory) model was developed to reconstruct the missing chest leads. The highest performing lead reconstructor achieved an average R2 score of 0.835 when reconstructing three chest leads. This architecture was proposed as the foundation for a wearable system that could record a limited number of ECG leads while also providing a 12-lead ECG for clinical applications. Kiriaki J. Rajotte, Bashima Islam, Xinming Huang 0001, David D. McManus, Edward A. Clancy |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Comparing Quantization Methods for On-Edge ECG Interpretation Using Multi-Task CNNabstractWearable devices have begun to incorporate machine learning models to assist with detection of various cardiac conditions. In this work, we developed a multi-task convolutional neural network to simultaneously predict$\mathbf{7 5}$diagnostic, form and rhythm statements from 10-s duration, 12-lead ECGs. The model, originally developed off-line in TensorFlow, was converted to the FlatBuffers format for on-edge AI using the LiteRT toolset. Posttraining quantization was used to compare different numerical precisions in terms of model size, model performance and inference time. Classifier performance for the 12-lead configuration was consistent between the 32-bit floating point model (“float32” baseline), the dynamic range quantized model (DR) and the float16 model$(p=0.92)$with an average macro AUC score of 0.893 with all output statements considered. A large degradation in classification performance was observed for 8-bit integer quantization (int8) which yielded an average macro AUC score of 0.513 for the 12-lead configuration across all statements. To address class imbalance, minority classes were removed. Reducing the number of statements to$\mathbf{4 1}$classes increased macro F1 score by an average of 72.6% (to a mean value about 0.358) for the float32, float16 and DR quantized models. Kiriaki J. Rajotte, Bashima Islam, David D. McManus, Xinming Huang 0001, Edward A. Clancy |
BSN | 4 |
| 2025 | Parallel Lidar Ground Segmentation: From Mechanical to Solid-State SensorsabstractReal-time and efficient LiDAR ground segmentation is crucial for autonomous perception and edge computing applications. This paper introduces an FPGA-based parallel segmentation framework that significantly reduces computational latency while maintaining high accuracy. Through a stream-pipelined design, the proposed method achieves significant speedup over CPU-based implementations while benefiting from the denser and more evenly distributed point clouds of Solid-State LiDAR (SSL). Extensive evaluations on the SemanticKITTI dataset and a self-collected SSL dataset demonstrate the effectiveness of our approach across both mechanical and SSL systems, highlighting its potential for real-world deployment in autonomous vehicles and edge-based perception platforms. Zhanhong Huang, Antony García, Witek Jachimczyk, Xinming Huang 0001 |
VTC2025-Spring | 5 |
| 2025 | Stream-Based LiDAR Point Cloud Ground Segmentation for Autonomous VehicleabstractThis paper introduces a hardware-friendly refined depth ground segmentation method and its FPGA deployment, designed for real-time applications with stringent power and energy constraints. Our approach outperforms traditional CPU solutions, achieving over 10x lower power consumption and orders of magnitude reduction in energy usage per frame. Leveraging a one-pass stream-based architecture, the method supports various LiDAR configurations with low latency and consistent performance. Experimental results demonstrate superior segmentation accuracy in both point-wise and area-wise evaluations, establishing our approach as a practical and efficient solution for autonomous vehicle applications. Zhanhong Huang, Antony García, Witek Jachimczyk, Xinming Huang 0001 |
VTC2025-Spring | 5 |
| 2025 | An Efficient L-Shape Based Object Detection Using Density Distribution in LiDAR Point Cloud for Autonomous DrivingabstractThis paper presents a novel and efficient L-shape detection framework designed for autonomous driving applications. We propose a fast geometry-based L-shape fitting algorithm that ensures high computational efficiency and consistent execution across diverse data sources, while preserving accurate shape alignment. Furthermore, a cluster-wise, octree-based Bird's Eye View (BEV) classifier is introduced to determine both oriented and axis-aligned bounding box alignments. Extensive experiments validate the proposed method's effectiveness and accuracy in pose estimation, object detection, and speed estimation. The system achieves an object velocity estimation error within 0.5 meters per second. Real-world evaluations confirm the robustness and applicability of the approach in practical autonomous driving scenarios. Zhanhong Huang, Xinming Huang 0001 |
VTC2025-Spring | 3 |
| 2024 | A Novel, Efficient and Accurate Method for Lidar Camera CalibrationabstractAs autonomous systems evolve, the precise calibration of lidar and camera sensors remains a pivotal concern. Among the myriad of available techniques, target-based calibration methods, which employ planar boards with distinct geometry and image patterns, have been a popular choice. These methods simplify the task of extracting corresponding features between the image and lidar point cloud. But many of these approaches also face a significant challenge, which is their sensitivity to lidar resolution and Field of View (FOV), which may degrade the reliability of the calibration results. Therefore, our research introduces a novel calibration method using a uniquely designed acrylic checkerboard which allows the lidar beam to pass through the white grids and reflect back from the black grids. This innovative technique sidesteps the common challenges associated with lidar feature extraction. Our method’s distinct advantage lies in its ability to perform accurate calibrations at close distances, owing to the efficient feature extraction from both lidar and camera sensors. This novel, efficient, and accurate method can provide state-of-the-art results for camera lidar calibration in the field. Please also check our Github repository: https://github.com/WPI-APA-Lab/Acrylic-Board-Lidar-Camera-Calibration Zhanhong Huang, Antony García, Xinming Huang 0001 |
ICRA | 4 |
| 2023 | Power Consumption and Maximum Number of Supported Nodes for BLE Biosensor ApplicationsabstractThere has been significant growth in wearable wireless patient physiological monitoring over the last decade. Many applications require a real-time, low-latency, small profile, and low-power-battery system. In this work, various Bluetooth low energy (BLE) configurations were tested in a multi-channel, wireless system to determine the lowest peripheral power configuration and the number of supported peripherals for each BLE configuration. Using nodes that continuously sampled at 1 kHz, connection intervals from 10-100 ms and event lengths of 2500, 5000 and 7500 μs were tested. The lowest current consumption, 2.39 mA, was measured for a connection interval of 100 ms, event length of 2500 μs, and maximum transmission unit (MTU) of 247 bytes. The maximum number of supported peripheral connections was observed to be 11 for a connection interval of 100 ms, event lengths of 5000 and 7500 μs, and MTU size of 247 bytes. We found that using longer connection intervals led to decreases in power consumption and shorter event lengths allowed for support of more peripheral sensors nodes for a given connection interval, assuming the event length is long enough to transmit the desired amount of data. Future work should investigate techniques to optimize power consumption further and to extend the number of supported peripheral nodes. Kiriaki J. Rajotte, Anson Wooding, Jianan Li 0004, Benjamin E. McDonald, Xinming Huang 0001, Todd R. Farrell, Edward A. Clancy |
BSN | 5 |
| 2023 | PRISE: Demystifying Deep Lucas-Kanade with Strongly Star-Convex Constraints for Multimodel Image AlignmentabstractThe Lucas-Kanade (LK) method is a classic iterative homography estimation algorithm for image alignment, but often suffers from poor local optimality especially when image pairs have large distortions. To address this challenge, in this paper we propose a novel Deep Star-Convexified Lucas-Kanade (PRISE) method for multimodel image alignment by introducing strongly star-convex constraints into the optimization problem. Our basic idea is to enforce the neural network to approximately learn a star-convex loss landscape around the ground truth give any data to facilitate the convergence of the LK method to the ground truth through the high dimensional space defined by the network. This leads to a minimax learning problem, with contrastive (hinge) losses due to the definition of strong star-convexity that are appended to the original loss for training. We also provide an efficient sampling based algorithm to leverage the training cost, as well as some analysis on the quality of the solutions from PRISE. We further evaluate our approach on benchmark datasets such as MSCOCO, GoogleEarth, and GoogleMap, and demonstrate state-of-the-art results, especially for small pixel errors. Code can be downloaded from https://github.com/Zhang-VISLab. Yiqing Zhang 0003, Xinming Huang 0001 |
CVPR | 2 |
| 2023 | Revisiting 2D Convolutional Neural Networks for Graph-Based ApplicationsabstractGraph convolutional networks (GCNs) are widely used in graph-based applications such as graph classification and segmentation. However, current GCNs have limitations on implementation such as network architectures due to their irregular inputs. In contrast, convolutional neural networks (CNNs) are capable of extracting rich features from large-scale input data, but they do not support general graph inputs. To bridge the gap between GCNs and CNNs, in this paper we study the problem of how to effectively and efficiently map general graphs to 2D grids that CNNs can be directly applied to, while preserving graph topology as much as possible. We therefore propose two novel graph-to-grid mapping schemes, namely, graph-preserving grid layout (GPGL) and its extension Hierarchical GPGL (H-GPGL) for computational efficiency. We formulate the GPGL problem as integer programming and further propose an approximate yet efficient solver based on a penalized Kamada-Kawai method, a well-known optimization algorithm in 2D graph drawing. We propose a novel vertex separation penalty that encourages graph vertices to lay on the grid without any overlap. Along with this image representation, even extra 2D maxpooling layers contribute to the PointNet, a widely applied point-based neural network. We demonstrate the empirical success of GPGL on general graph classification with small graphs and H-GPGL on 3D point cloud segmentation with large graphs, based on 2D CNNs including VGG16, ResNet50 and multi-scale maxout (MSM) CNN. Yecheng Lyu, Xinming Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Exploiting Multi-View Part-Wise Correlation via an Efficient Transformer for Vehicle Re-IdentificationabstractImage-based vehicle re-identification (ReID) has witnessed much progress in recent years. However, most of existing works struggled to extract robust but discriminative features from a single image to represent one vehicle instance. We argue that images taken from distinct viewpoints,e.g.,front and back, have significantly different appearances and patterns for recognition. In order to identify each vehicle, these models have to capture consistent “ID codes” from totally different views, causing learning difficulties. Additionally, we claim that part-level correspondences among views,i.e.,various vehicle parts observed from the identical image and the same part visible from different viewpoints, contribute to instance-level feature learning as well. Motivated by these, we propose to extract comprehensive vehicle instance representations from multiple views through modelling part-wise correlations. To this end, we present our efficient transformer-based framework to exploit both inner- and inter-view correlations for vehicle ReID. In specific, we first adopt a convnet encoder to condense a series of patch embeddings from each view. Then our efficient transformer, consisting of a distillation token and a noise token in addition to a regular classification token, is constructed for enforcing these patch embeddings to interact with each other regardless of whether they are taken from identical or different views. We conduct extensive experiments on widely used vehicle ReID benchmarks, and our approach achieves the state-of-the-art performance, showing the effectiveness of our method. Ming Li 0073, Jun Liu 0036, Xinming Huang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Distance Transform Pooling Neural Network for LiDAR Depth CompletionabstractRecovering dense depth maps from sparse depth sensors, such as LiDAR, is a recently proposed task with many computer vision and robotics applications. Previous works have identified input sparsity as the key challenge of this task. To solve the sparsity challenge, we propose a recurrent distance transform pooling (DTP) module that aggregates multi-level nearby information prior to the backbone neural network. The intuition of this module is originated from the observation that most pixels within the receptive field of the network are zero. This indicates a deep and heavy network structure has to be used to enlarge the receptive field aiming at capturing enough useful information as most processed signals are uninformative zeros. Our recurrent DTP module can fill in empty pixels with the nearest value in a local patch and recurrently transform distance to reach farther nearest points. The output of the proposed DTP module is a collection of multi-level semi-dense depth maps from original sparse to almost full. Processing this collection of semi-dense depth maps alleviates the network from the input sparsity, which helps a lightweight simplified ResNet-18 with 1M parameters achieve state-of-the-art performance on the Karlsruhe Institute of Technology and Toyota Technological Institute (KITTI) depth completion benchmark with LiDAR only. Besides the sparsity, the input LiDAR map also contains some incorrect values due to the sensor error. Thus, we further enhance the DTP with an error correction (EC) module to avoid the spreading of the incorrect input values. At last, we discuss the benefit of only using LiDAR for nighttime driving and the potential extension of the proposed method for sensor fusion and the indoor scenario. The code has been released online at https://github.com/placeforyiming/DistanceTransform-DepthCompletion. Mahdi Elhousni, Xinming Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | A Divide-and-Merge Point Cloud Clustering Algorithm for LiDAR Panoptic SegmentationabstractClustering objects from the LiDAR point cloud is an important research problem with many applications such as autonomous driving. To meet the real-time requirement, existing research proposed to apply the connected-component-labeling (CCL) technique on LiDAR spherical range image with a heuristic condition to check if two neighbor points are connected. However, LiDAR range image is different from a binary image which has a deterministic condition to tell if two pixels belong to the same component. The heuristic condition used on the LiDAR range image only works empirically, which suggests the LiDAR clustering algorithm should be robust to potential failures of the empirical heuristic condition. To overcome this challenge, this paper proposes a divide-and-merge LiDAR clustering algorithm. This algorithm firstly conducts clustering in each evenly divided local region, then merges the local clustered small components by voting on edge point pairs. Assuming there are$N$LiDAR points of objects in total with$m$divided local regions, the time complexity of the proposed algorithm is$O(N)+O(m^{2})$. A smaller$m$means the voting will involve more neighbor points, but the time complexity will become larger. So the$m$controls the trade-off between the time complexity and the clustering accuracy. A proper$m$helps the proposed algorithm work in real-time as well as maintain good performance. We evaluate the divide-and-merge clustering algorithm on the SemanticKITTI panoptic segmentation benchmark by cascading it with a state-of-the-art semantic segmentation model. The final performance evaluated through the leaderboard achieves the best among all published methods. The proposed algorithm is implemented with C++ and wrapped as a python function. It can be easily used with the modern deep learning framework in python. We released the code under the following link11https://github.com/placeforyiming/Divide-and-Merge-LiDAR-Panoptic-Cluster. Xinming Huang 0001 |
ICRA | 3 |
| 2022 | A Near Sensor Edge Computing System for Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation has attracted attentions due to its robustness to light condition. This makes it an ideal semantic solution for autonomous driving. However, considering the large computation burden and bandwidth demanding of neural networks, putting all the computing into vehicle Electronic Control Unit (ECU) is not efficient or practical. In this paper, we proposed a light weighted point cloud semantic segmentation network based on range view. Due to its simple preprocessing and standard convolution, it is efficient when running on deep learning accelerator like DPU. Furthermore, a near sensor computing system is built for autonomous vehicles. In this system, a FPGA-based deep learning accelerator core (DPU) is placed next to the LiDAR sensor, to perform point cloud preprocessing and segmentation neural network. By leaving only the post-processing step to ECU, this solution heavily alleviate the computation burden of ECU and consequently shortens the decision making and vehicles reaction latency. Our semantic segmentation network achieved 10 frame per second (fps) on Xilinx DPU with computation efficiency 42.5 GOP/W. Lin Bai 0002, Xinming Huang 0001 |
ISCAS | 3 |
| 2022 | EllipsoidNet: Ellipsoid Representation for Point Cloud Classification and SegmentationabstractPoint cloud patterns are hard to learn because of the implicit local geometry features among the orderless points. In recent years, point cloud representation in 2D space has attracted increasing research interest since it exposes the local geometry features in a 2D space. By projecting those points to a 2D feature map, the relationship between points is inherited in the context between pixels, which are further extracted by a 2D convolutional neural network. However, existing 2D representing methods are either accuracy limited or time-consuming. In this paper, we propose a novel 2D representation method that projects a point cloud onto an ellipsoid surface space, where local patterns are well exposed in ellipsoid-level and point-level. Additionally, a novel convolutional neural network named EllipsoidNet is proposed to utilize those features for point cloud classification and segmentation applications. The proposed methods are evaluated in ModelNet40 and ShapeNet benchmarks, where the advantages are clearly shown over existing 2D representation methods. Yecheng Lyu, Xinming Huang 0001 |
WACV | 2 |
| 2021 | Deep Lucas-Kanade Homography for Multimodal Image AlignmentabstractEstimating homography to align image pairs captured by different sensors or image pairs with large appearance changes is an important and general challenge for many computer vision applications. In contrast to others, we propose a generic solution to pixel-wise align multimodal image pairs by extending the traditional Lucas-Kanade algorithm with networks. The key contribution in our method is how we construct feature maps, named as deep Lucas-Kanade feature map (DLKFM). The learned DLKFM can spontaneously recognize invariant features under various appearance-changing conditions. It also has two nice properties for the Lucas-Kanade algorithm: (1) The template feature map keeps brightness consistency with the input feature map, thus the color difference is very small while they are well-aligned. (2) The Lucas-Kanade objective function built on DLKFM has a smooth landscape around ground truth homography parameters, so the iterative solution of the Lucas-Kanade can easily converge to the ground truth. With those properties, directly updating the Lucas-Kanade algorithm on our feature maps will precisely align image pairs with large appearance changes. We share the datasets, code, and demo video online1. Xinming Huang 0001 |
CVPR | 2 |
| 2021 | Self-supervised Geometric Features Discovery via Interpretable Attention for Vehicle Re-Identification and BeyondabstractTo learn distinguishable patterns, most of recent works in vehicle re-identification (ReID) struggled to redevelop official benchmarks to provide various supervisions, which requires prohibitive human labors. In this paper, we seek to achieve the similar goal but do not involve more human efforts. To this end, we introduce a novel framework, which successfully encodes both geometric local features and global representations to distinguish vehicle instances, optimized only by the supervision from official ID labels. Specifically, given our insight that objects in ReID share similar geometric characteristics, we propose to borrow self-supervised representation learning to facilitate geometric features discovery. To condense these features, we introduce an interpretable attention module, with the core of local maxima aggregation instead of fully automatic learning, whose mechanism is completely understandable and whose response map is physically reasonable. To the best of our knowledge, we are the first that perform self-supervised learning to discover geometric features. We conduct comprehensive experiments on three most popular datasets for vehicle ReID, i.e., VeRi-776, CityFlow-ReID, and VehicleID. We report our state-of-the-art (SOTA) performances and promising visualization results. We also show the excel-lent scalability of our approach on other ReID related tasks, i.e., person ReID and multi-target multi-camera (MTMC) vehicle tracking. Ming Li 0073, Xinming Huang 0001 |
ICCV | 2 |
| 2021 | FIDNet: LiDAR Point Cloud Semantic Segmentation with Fully Interpolation DecodingabstractProjecting the point cloud on the 2D spherical range image transforms the LiDAR semantic segmentation to a 2D segmentation task on the range image. However, the LiDAR range image is still naturally different from the regular 2D RGB image; for example, each position on the range image encodes the unique geometry information. In this paper, we propose a new projection-based LiDAR semantic segmentation pipeline that consists of a novel network structure and an efficient post-processing step. In our network structure, we design a FID (fully interpolation decoding) module that directly upsamples the multi-resolution feature maps using bilinear interpolation. Inspired by the 3D distance interpolation used in PointNet++, we argue this FID module is a 2D version distance interpolation on (θ, ϕ) space. As a parameter-free decoding module, the FID largely reduces the model complexity by maintaining good performance. Besides the network structure, we empirically find that our model predictions have clear boundaries between different semantic classes. This makes us rethink whether the widely used K-nearest-neighbor post-processing is still necessary for our pipeline. Then, we realize the many-to-one mapping causes the blurring effect that some points are mapped into the same pixel and share the same label. Therefore, we propose to process those occluded points by assigning the nearest predicted label to them. This NLA (nearest label assignment) post-processing step shows a better performance than KNN with faster inference speed in the ablation study. On SemanticKITTI dataset, our pipeline achieves the best performance among all projection-based methods with 64×2048 resolution and all point-wise solutions. With a ResNet-34 as the backbone, both the training and testing of our model can be finished on a single RTX 2080 Ti with 11G memory. The code is released here.1 Lin Bai 0002, Xinming Huang 0001 |
IROS | 3 |
| 2021 | RoadNet-RT: High Throughput CNN Architecture and SoC Design for Real-Time Road SegmentationabstractIn recent years, convolutional neural network (CNN) has gained popularity in many engineering applications especially for computer vision. In order to achieve better performance, more complex structures and advanced operations are incorporated into neural networks, which results in very long inference time. For time-critical tasks such as autonomous driving and virtual reality, real-time processing is fundamental. In order to reach real-time processing speed, a lightweight, high-throughput CNN architecture namely RoadNet-RT is proposed for road segmentation in this article. It achieves 92.55% MaxF score on KITTI road segmentation dataset. The inference time is about 9 ms per frame when running on GTX 1080 GPU. Comparing to the state-of-the-art network, RoadNet-RT speeds up the inference time by a factor of 17.8 at the cost of only 3.75% loss in accuracy. What is more, on CamVid dataset its accuracy is 92.98%. Several techniques such as depthwise separable convolution and non-uniformed kernel size convolution are optimized in the hardware accelerator design. The proposed CNN architecture has been successfully implemented on a ZCU102 MPSoC FPGA that achieves the computation capability of 331 GOPS using INT8 quantization. The system throughput reaches 196.7 frames per second with input image size of 280 × 960 . The source code is published at https://github.com/linbaiwpi/RoadNet-RT. Lin Bai 0002, Yecheng Lyu, Xinming Huang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2020 | Automatic Building and Labeling of HD Maps with Deep LearningabstractIn a world where autonomous driving cars are becoming increasingly more common, creating an adequate infrastructure for this new technology is essential. This includes building and labeling high-definition (HD) maps accurately and efficiently. Today, the process of creating HD maps requires a lot of human input, which takes time and is prone to errors. In this paper, we propose a novel method capable of generating labelled HD maps from raw sensor data. We implemented and tested our methods on several urban scenarios using data collected from our test vehicle. The results show that the proposed deep learning based method can produce highly accurate HD maps. This approach speeds up the process of building and labeling HD maps, which can make meaningful contribution to the deployment of autonomous vehicles. Mahdi Elhousni, Yecheng Lyu, Xinming Huang 0001 |
AAAI | 4 |
| 2020 | Learning to Segment 3D Point Clouds in 2D Image SpaceabstractIn contrast to the literature where local patterns in 3D point clouds are captured by customized convolutional operators, in this paper we study the problem of how to effectively and efficiently project such point clouds into a 2D image space so that traditional 2D convolutional neural networks (CNNs) such as U-Net can be applied for segmentation. To this end, we are motivated by graph drawing and reformulate it as an integer programming problem to learn the topology-preserving graph-to-grid mapping for each individual point cloud. To accelerate the computation in practice, we further propose a novel hierarchical approximate algorithm. With the help of the Delaunay triangulation for graph construction from point clouds and a multi-scale U-Net for segmentation, we manage to demonstrate the state-of-the-art performance on ShapeNet and PartNet, respectively, with significant improvement over the literature. Code is available at https://github.com/Zhang-VISLab. Yecheng Lyu, Xinming Huang 0001 |
CVPR | 2 |
| 2020 | TreeRNN: Topology-Preserving Deep Graph Embedding and LearningabstractGeneral graphs are difficult for learning due to their irregular structures. Existing works employ message passing along graph edges to extract local patterns using customized graph kernels, but few of them are effective for the integration of such local patterns into global features. In contrast, in this paper we study the methods to transfer the graphs into trees so that explicit orders are learned to direct the feature integration from local to global. To this end, we apply the breadth first search (BFS) to construct trees from the graphs, which adds direction to the graph edges from the center node to the peripheral nodes. In addition, we proposed a novel projection scheme that transfer the trees to image representations, which is suitable for conventional convolution neural networks (CNNs) and recurrent neural networks (RNNs). To best learn the patterns from the graph-tree-images, we propose TreeRNN, a 2D RNN architecture that recurrently integrates the image pixels by rows and columns to help classify the graph categories. We evaluate the proposed method on several graph classification datasets, and manage to demonstrate comparable accuracy with the state-of-the-art on MUTAG, PTC-MR and NCI1 datasets. Yecheng Lyu, Ming Li 0073, Xinming Huang 0001, Ulkuhan Guler 0001, Patrick Schaumont |
ICPR | 3 |
| 2020 | PointNet on FPGA for Real-Time LiDAR Point Cloud ProcessingabstractLiDAR sensors have been widely used in many autonomous vehicle modalities, such as perception, mapping, and localization. This paper presents an FPGA-based deep learning platform for real-time point cloud processing targeted on autonomous vehicles. The software driver for the Velodyne LiDAR sensor is modified and moved into the on-chip processor system, while the programmable logic is designed as a customized hardware accelerator. As the state-of-art deep learning algorithm for point cloud processing, PointNet is successfully implemented on the proposed FPGA platform. Targeted on a Xilinx Zynq UltraScale+ MPSoC ZCU104 development board, the FPGA implementations of PointNet achieve the computing performance of 182.1 GOPS and 280.0 GOPS for classification and segmentation respectively. The proposed design can support an input up to 4096 points per frame. The processing time is 19.8 ms for classification and 34.6 ms for segmentation, which meets the real-time requirement for most of the existing LiDAR sensors. Lin Bai 0002, Yecheng Lyu, Xinming Huang 0001 |
ISCAS | 4 |
| 2020 | A Unified Hardware Architecture for Convolutions and Deconvolutions in CNNabstractDeconvolution plays an important role in the state-of-the-art convolutional neural networks (CNNs) for the tasks like semantic segmentation, image super resolution, etc. In this paper, a scalable neural network hardware architecture for image segmentation is proposed. By sharing the same computing resources, both convolution and deconvolution operations are handled by the same process element array. In addition, access to on-chip and off-chip memories is optimized to alleviate the burden introduced by partial sum. As an example, SegNet-Basic has been implemented using the proposed unified architecture by targeting on Xilinx ZC706 FPGA, which achieves the performance of 151.5 GOPS and 94.3 GOPS for convolution and deconvolution respectively. This unified convolution/deconvolution design is applicable to other CNNs with deconvolution. Lin Bai 0002, Yecheng Lyu, Xinming Huang 0001 |
ISCAS | 3 |
| 2020 | Pedestrian Tracking with Gated Recurrent Units and Attention MechanismsabstractPedestrian tracking has long been considered an important problem, especially in security applications. Previously, many approaches have been proposed with various types of sensors. One popular method is Pedestrian Dead Reckoning (PDR) [1] which is based on the inertial measurement unit (IMU) sensor. However PDR is an integration and threshold based method, which suffers from accumulation errors and low accuracy. In this paper, we propose a novel method in which the sensor data is fed into a deep learning model to predict the displacements and orientations of the pedestrian. We also devise a new apparatus to collect and construct databases containing synchronized IMU sensor data and precise locations measured by a LIDAR. The preliminary results are promising, and we plan to push this forward by collecting more data and adapting the deep learning model for all general pedestrian motions. Mahdi Elhousni, Xinming Huang 0001 |
ISCAS | 2 |
| 2020 | A Survey on 3D LiDAR Localization for Autonomous VehiclesabstractLiDAR sensors are becoming one of the most essential sensors in achieving full autonomy for self driving cars. LiDARs are able to produce rich, dense and precise spatial data, which can tremendously help in localizing and tracking a moving vehicle. In this paper, we review the latest finding in 3D LiDAR localization for autonomous driving cars, and analyse the results obtained by each method, in an effort to guide the research community towards the path that seems to be the most promising. Mahdi Elhousni, Xinming Huang 0001 |
IV | 2 |
| 2019 | MAP Joint Frequency and Channel Estimation for MIMO Systems with Spatial CorrelationabstractCarrier frequency offset (CFO) and channel estimation is a classic topic with a large body of prior work using the maximum likelihood (ML) approach together with the CramérRao lower bound (CRLB) analysis. We give the maximum a posteriori probability (MAP) estimation solution which is particularly useful for tracking. Unlike the ML cases, the corresponding Bayesian CRLB (BCRLB) shows a clear relation with parameters and a low complexity algorithm achieves the BCRLB in almost all SNR range. We allow the time invariant MIMO channel within a packet to have arbitrary spatial correlation and mean. The estimation is based on pilot signals. An unexpected result is that the joint MAP estimation is equivalent to an individual MAP estimation of the frequency offset first, again different from the ML results. We provide insight on the pilot/training signal design based on the BCRLB. Unlike past algorithms that trade performance and/or complexity for the accommodation of time varying channels, the MAP solution provides a different route for dealing with time variation. Mingda Zhou, Zhe Feng 0005, Xinming Huang 0001, Youjian Liu |
GLOBECOM | 3 |
| 2019 | A 20 TOp/s/W Binary Neural Network AcceleratorabstractThis paper presents the hardware architecture and VLSI implementation of a binarized neural network (BNN). As a modification of convolutional neural network (CNN), BNN constrains all activations and weights to be +1 or -1, making it very appealing to low-power ASIC design. In this paper, BNN is proven to be highly power efficient and accurate for computer vision tasks. We use pedestrian and car detections as examples to showcase the capability of the BNN chip design. The total memory use of all weights in the BNN is only 22K bytes, which is significantly less than a typical convolutional neural network. Evaluated using INRIA and CIFAR-10 datasets, our BNN chip can achieve an average accuracy of 96.5%, which is much higher than traditional computer vision approaches such as histogram of oriented gradients with support vector machine. Our design achieves a power efficiency of 20 TOp/s/w, far exceeding most of the mainstream CNN chips. Therefore, the proposed low power hardware architecture of BNN enables deep learning on mobile embedded platforms. Xinming Huang 0001, Yuteng Zhou |
ISCAS | 1 |
| 2019 | Road Segmentation using CNN and Distributed LSTMabstractIn automated driving systems (ADS) and advanced driver-assistance systems (ADAS), an efficient road segmentation is necessary to perceive the drivable region and build an occupancy map for path planning. The existing algorithms implement gigantic convolutional neural networks (CNNs) that are computationally expensive and time consuming. In this paper, we introduced distributed LSTM, a neural network widely used in audio and video processing, to process rows and columns in images and feature maps. We then propose a new network combining the convolutional and distributed LSTM layers to solve the road segmentation problem. In the end, the network is trained and tested in KITTI road benchmark. The result shows that the combined structure enhances the feature extraction and processing but takes less processing time than pure CNN structure. Yecheng Lyu, Lin Bai 0002, Xinming Huang 0001 |
ISCAS | 3 |
| 2018 | Real-Time Road Segmentation Using LiDAR Data Processing on an FPGAabstractThis paper presents the FPGA design of a convolutional neural network (CNN) based road segmentation algorithm for real-time processing of LiDAR data. For autonomous vehicles, it is important to perform road segmentation and obstacle detection such that the drivable region can be identified for path planning. Traditional road segmentation algorithms are mainly based on image data from cameras, which is subjected to the light condition as well as the quality of lane markings. LiDAR sensor can obtain the precise 3D geometry information of the vehicle surroundings. However, it is a computational challenge to process a large amount of LiDAR data at real-time. In this work, a convolutional neural network model is proposed and trained to perform semantic segmentation using the LiDAR sensor data. Furthermore, an efficient hardware design is implemented on the FPGA that can process each LiDAR scan in 16.9 ms, which is much faster than the previous works. Evaluated using KITTI road benchmarks, the proposed solution achieves high accuracy of road segmentation. Yecheng Lyu, Lin Bai 0002, Xinming Huang 0001 |
ISCAS | 3 |
| 2018 | Low-Power SDR Design on an FPGA for Intersatellite Communications
Mingda Zhou, Tian Xia 0005, Wai H. Fong, Wing-Tsz Lee, Xinming Huang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2018 | Genetic Algorithm Based QoS Perception Routing Protocol for VANETsabstractA genetic algorithm (GA) based QoS perception routing protocol (GABR) is proposed to guarantee the quality of service (QoS) influenced by broken links between vehicles and the failure of packets transmission in a vehicular ad hoc network (VANET). With the observation that all improvable paths are probed by the intersection based routing protocol, the genetic GA is utilized to optimize the global available paths which satisfies the QoS requirement. Moreover, by means of the numerical results, it is shown that the proposed scheme is significantly improved compared with protocols of the intersection based routing (IBR) and connectivity aware routing (CAR) in terms of transmission delay and packet loss rate. Guoan Zhang, Wei Duan 0001, Xinming Huang 0001 |
Wirel. Commun. Mob. Comput. | 4 |
| 2017 | An FPGA prototype of dual link algorithm for MIMO interference networkabstractThis paper presents an FPGA-based prototype of the dual link algorithm that maximize the achievable weighted sum rate for MIMO interference network. The iterative algorithm is fast monotone convergent but it must be completed quickly through pilot signaling. Therefore we propose an FPGA-based implementation, targeting software defined radio (SDR) platforms, that is designed to rapidly process the received pilot signals, estimate local channel information, and compute the transmit signal covariance matrices. Compared to its implementation on a CPU, the FPGA implementation is 2 to 4 times faster using fixed-point or floating-point designs. Mingda Zhou, Xinming Huang 0001, Yuteng Zhou, Youjian Liu |
ICASSP | 2 |
| 2017 | End-to-end learning for lane keeping of self-driving carsabstractLane keeping is an important feature for self-driving cars. This paper presents an end-to-end learning approach to obtain the proper steering angle to maintain the car in the lane. The convolutional neural network (CNN) model takes raw image frames as input and outputs the steering angles accordingly. The model is trained and evaluated using the comma.ai dataset, which contains the front view image frames and the steering angle data captured when driving on the road. Unlike the traditional approach that manually decomposes the autonomous driving problem into technical components such as lane detection, path planning and steering control, the end-to-end model can directly steer the vehicle from the front view camera data after training. It learns how to keep in lane from human driving data. Further discussion of this end-to-end approach and its limitation are also provided. Zhilu Chen, Xinming Huang 0001 |
Intelligent Vehicles Symposium | 2 |
| 2016 | A system-on-chip FPGA design for real-time traffic signal recognition systemabstractTraffic signal detection has long been an important function in an advanced driver assistance system (ADAS). This paper presents a complete system design based on the techniques of blob detection, histogram of oriented gradients (HOG) and support vector machine (SVM). Blob detection is applied to detect potential candidates, and then HOG and SVM is for feature classification. A novel hardware/software co-design architecture is developed for traffic light recognition at real-time. With well-balanced workload on FPGA fabric and the on-chip ARM processor, the entire system-on-chip can achieve a processing rate of 60 fps for XGA 1024-by-768 video. The system can achieve an accuracy rate of over 90% on both red lights and green lights. The proposed system can be improved by replacing HOG with more advanced feature algorithm to obtain higher accuracy. Yuteng Zhou, Zhilu Chen, Xinming Huang 0001 |
ISCAS | 3 |
| 2015 | FPGA Design for PCANet Deep Learning NetworkabstractIn recent years, deep learning has attracted lots of research interests for pattern recognition and artificial intelligence. PCA Network (PCANet) is a simple deep learning network with highly competitive performance for texture classification and object recognition. When compared to other deep neural networks such as convolutional neural network (CNN), PCANet has much simpler structure, which makes it attractive for hardware design on an FPGA. In this paper, an efficient, high-throughput, pipeline architecture is proposed for the PCANet classifier. The implementation on an FPGA is more than 1,000 times faster than software execution on a general purpose processor. When evaluated using the MNIST handwritten digits dataset, the PCANet design results an accuracy of about 99.46%. Yuteng Zhou, Wei Wang 0053, Xinming Huang 0001 |
FCCM | 3 |
| 2015 | A pipeline architecture for traffic sign classification on an FPGAabstractThis paper presents an efficient FPGA design that can classify 48 different traffic signs at real-time. The method is based on histogram of oriented gradients (HOG) feature extraction and support vector machine (SVM) for classification. A full-pipeline, resource efficient architecture is presented with detail design of each block. The FPGA implementation has a system clock of 241.7 MHz with an absolute response time of 6.5 s. Taking streaming pixel input, the system throughput is about 106 times faster than the same algorithm executed on a general purpose processor. Yuteng Zhou, Zhilu Chen, Xinming Huang 0001 |
ISCAS | 3 |
| 2015 | Road marking detection and classification using machine learning algorithmsabstractThis paper presents a novel approach for road marking detection and classification based on machine learning algorithms. Road marking recognition is an important feature of an intelligent transportation system (ITS). Previous works are mostly developed using image processing and decisions are often made using empirical functions, which makes it difficult to be generalized. Hereby, we propose a general framework for object detection and classification, aimed at video-based intelligent transportation applications. It is a two-step approach. The detection is carried out using binarized normed gradient (BING) method. PCA network (PCANet) is employed for object classification. Both BING and PCANet are among the latest algorithms in the field of machine learning. Practically the proposed method is applied to a road marking dataset with 1,443 road images. We randomly choose 60% images for training and use the remaining 40% images for testing. Upon training, the system can detect 9 classes of road markings with an accuracy better than 96.8%. The proposed approach is readily applicable to other ITS applications. Tairui Chen, Zhilu Chen, Xinming Huang 0001 |
Intelligent Vehicles Symposium | 4 |
| 2015 | Automatic detection of traffic lights using support vector machineabstractMany traffic accidents occurred at intersections are caused by drivers who miss or ignore the traffic signals. In this paper, we present a new method for automatic detection of traffic lights that integrates both image processing and support vector machine techniques. An experimental dataset with 21299 samples is built from the captured original videos while driving on the streets. When compared to the traditional object detection and existing methods, the proposed system provides significantly better performance with 96.97% precision and 99.43% recall. The system framework is extensible that users can introduce additional parameters to further improve the detection performance. Zhilu Chen, Xinming Huang 0001 |
Intelligent Vehicles Symposium | 3 |
| 2015 | TDD Channel Calibration for MIMO Interference NetworksabstractThis paper presents a simple method of internal channel calibration for MIMO interference networks to maintain TDD channel reciprocity. The method works for algorithms requiring channel reciprocity. In particular, it is shown how the method works for the distributed implementation of the Dual Link algorithm, which is a very fast and provably convergent algorithm for optimal interference management. Youjian Liu, Xinming Huang 0001 |
VTC Fall | 3 |
| 2015 | A Novel TDMA-MAC Protocol for VANET Using Cooperative and Opportunistic TransmissionsabstractThis paper presents a novel time division multiple access-medium access control (TDMA-MAC) protocol for vehicular ad hoc networks (VANETs), in which both cooperative and opportunistic transmissions are employed for enhanced communications. When vehicle density is low, the idle time slots of a licensed VANET channel are used for cooperative transmission through a relay. When the vehicle density is high, cognitive radio technique is applied to seek additional time slots available on the cognitive channels for opportunistic transmission. Simulation results show that the proposed TDMA-MAC protocol can reduce latency and packet loss rate significantly when compared with the existing protocols. Guoan Zhang, Xinming Huang 0001, Xiang Ye |
VTC Fall | 3 |
| 2015 | Security-quality aware routing for wireless multimedia sensor networks using secret sharingabstractAbstract Security and video quality are progressively significant attributes for wireless multimedia sensor networks. Most of existing research considers security and video quality separately. However, it is crucial to integrate security and video quality together for video transmission because delivering video data across a secure path does not often meet video quality requirements in many traditional approaches. Applying the general concept of secret sharing algorithm on a data packet and delivering it through disjoint multipaths can be considered to deliver the data securely. However, using the general concept of secret sharing is not efficient when large‐size video data are routed. To tackle these issues, we propose a novel security and quality aware routing (SQAR) protocol to address these two issues concurrently. We jointly consider security and video quality in wireless multimedia networks by proposing a video distortion model based on a new secret image sharing scheme. In SQAR, a secret image sharing is only applied on the intra‐frames of the video codec H.264 and can significantly reduce the transmission overheads. Simulation results show that SQAR scheme can achieve better trade‐off between the security and quality over the traditional routing protocols. Copyright © 2015 John Wiley & Sons, Ltd. Abdelnaser Rashwan, Honggang Wang 0001, Dalei Wu, Xinming Huang 0001 |
Secur. Commun. Networks | 4 |
| 2015 | Exploring the Feasibility of Fully Homomorphic EncryptionabstractIn 2010, Gentry and Halevi presented the first FHE implementation. FHE allows the evaluation of arbitrary functions directly on encrypted data on untrusted servers. However, even for the small setting with 2048 dimensions, the authors reported a performance of 1.8 s for a single bit encryption and 32 s for recryption on a high-end server. Much of the latency is due to computationally intensive multi-million-bit modular multiplications. In this paper, we introduce two optimizations coupled with a novel precomputation technique. In the first optimization called partial FFT, we adopt Strassen’s FFT-based multiplication algorithm along with Barret reduction to speedup modular multiplications. For the encrypt primitive, we employ a window-based evaluation technique along with a modest degree of precomputation. In the full FFT optimization, we delay modular reductions and change the window algorithm, which allows us to carry out the bulk of computations in the frequency domain. We manage to eliminate all FFT conversion except the final inverse transformation drastically reducing the computation latency for all FHE primitives. We implemented the GH FHE scheme on two GPUs to further speedup the operations. Our experimental results with small parameter setting show speedups of 174, 7.6, and 13.5 times for encryption, decryption, and recryption, respectively, when compared to the Gentry–Halevi implementation. The speedup is enhanced in the medium setting. However, in the large setting, memory becomes the bottleneck and the speedup is somewhat diminished. Wei Wang 0053, Lianmu Chen, Xinming Huang 0001, Berk Sunar |
IEEE Trans. Computers | 4 |
| 2015 | Multicast communications in cognitive radio networks using directional antennasabstractThis paper presents a study on multicast communications in cognitive radio networks CRNsusing directional antennas. The objective is to maximize the throughput of the CRN. The spectrum is divided into multiple channels and licensed to the primary network. While the CRN is accessing the spectrum, the interference power is carefully controlled to avoid impacting the operation of the primary network. The mathematical model is presented and subsequently formulated as a mixed integer non-linear programming MINLP problem, which is non-deterministic polynomial-time hard. Therefore, a greedy algorithm is designed to approximate the optimal performance. The MINLP problem is then relaxed and an upper bound is developed. Simulation results are presented to compare the performance of the greedy algorithm and the upper bound, which demonstrates the efficacy of the greedy algorithm as well as the tightness of the upper bound. Copyright © 2012 John Wiley & Sons, Ltd. Xinming Huang 0001 |
Wirel. Commun. Mob. Comput. | 2 |
| 2014 | Experimental studies on indoor sign recognition and classificationabstractPrevious works on outdoor traffic sign recognition and classification have been demonstrated useful to the driver assistant system and the possibility to the autonomous vehicles. This motivates our research on the assistance for visual impairment or visual disabled pedestrians in the indoor environment. In this paper, we build an indoor sign database and investigate the recognition and classification for the indoor sign problem. We adopt the classical techniques on extracting the features, including the principle component analysis (PCA), dense scale invariant feature transform (DSIFT), histogram of oriented gradients (HOG), and conduct the state-of-art classification techniques, such as the neural network (NN), support vector machine (SVM) and k-nearest neighbors (KNN). We provide the experimental results on this newly built database and also discuss the insight for the possibility of indoor navigation for the blind or visual-disabled people. Zhen Ni, Si-Yao Fu, Bo Tang 0011, Haibo He, Xinming Huang 0001 |
CIDM | 5 |
| 2014 | A fast deep learning system using GPUabstractThe invention of deep belief network (DBN) provides a powerful tool for data modeling. The key advantage of DBN is that it is driven by training data only, which can alleviate researchers from the routine of devising explicit models or features for data with complicated distributions. However, as the dimensionality and quantity of data increase, the computing load of training a DBN increases rapidly. Prospectively, the remarkable computing power provided by modern GPU devices can reduce the training time of DBN significantly. As highly efficient computational libraries become available, it provides additional support for GPU based parallel computing. Moreover, GPU server is more affordable and accessible compared with computer cluster or supercomputer. In this paper, we implement a variant of the DBNs, called folded-DBN, on NVIDA's Tesla K20 GPU. In our simulations, two sets of database are used to train the folded-DBNs on both CPU and GPU platforms. Comparing execution time of the fine-tuning process, the GPU implementation results 7 to 11 times speedup over the CPU platform. Zhilu Chen, Haibo He, Xinming Huang 0001 |
ISCAS | 4 |
| 2014 | Multilevel error correction scheme for MLC flash memoryabstractStoring multiple bits in a flash memory cell is a primary technique to linearly increase flash memory capacity, but memory endurance is tremendously sacrificed. This paper presents a multilevel fault tolerance technique for MLC flash memories. The main idea is to explore multi-level forward error correction (FEC) for multiple bits in a flash cell. The associated practical implementation issues are well addressed in this paper. Compared to the conventional error protection methods for flash memory, the proposed multi-level FEC approach can obtain much larger system coding gain using the same amount of redundant bits. As a result, the proposed technique reduces power consumption considerably compared to the conventional methods since the required throughout of LDPC codec is drastically reduced. It can also increase flash memory endurance as it can allocate more redundancy to LDPC code while maintaining overall redundancy ratio. Zhiqiang Cui, Zhongfeng Wang 0001, Xinming Huang 0001 |
ISCAS | 3 |
| 2014 | Performance comparison of hybrid partial response detectors over frequency-selective fading channelsabstractFrequency-selective fading channels are encountered in many modern wireless communication systems. In order to combat the intersymbol interference (ISI) introduced by such fading, equalization is required for reliable symbol detection. The maximum-likelihood sequence detector is the optimal equalization scheme; however its implementation complexity increases exponentially with the channel length and thus can be prohibitively high. In this paper, we compare the performance of two practically implementable suboptimal symbol detectors, including the partial response maximum-likelihood (PRML) detector and the partial response belief propagation (PRBP) detector, under frequency-selective fading channels. Both detectors employ a hybrid two-stage scheme, and allow a tradeoff between performance and complexity. The first stage is a partial response equalizer implemented as a linear filter which transforms the original channel impulse response to a target impulse response with reduced ISI. The residual ISI is then cancelled in the second stage using a more sophisticated nonlinear detector. In simulations, we consider a slow fading environment and use the ITU-R 3G channel models. From the numerical results, it is shown that in frequency-selective fading wireless channels, the PRBP detector provides superior performance over both the traditional minimum mean squared error linear equalizer and the PRML detector. Due to the effect of colored noise, the PRML detector in fading wireless channels is not as effective as it is in magnetic recording applications. Yanjie Peng, Xinming Huang 0001 |
ISCAS | 2 |
| 2014 | Hybrid DFSF-BP equalization for ATSC DTV receiversabstractSevere intersymbol interference (ISI) is one of the main obstacles for reliable signal reception in ATSC DTV systems. Decision feedback equalizers (DFEs) are commonly used to suppress the ISI. However, DFEs may suffer from error propagation due to incorrect symbol decisions from the symbol slicer. This phenomenon deteriorates the performance even more when the post-cursor ISI is strong. In order to reduce error propagation, we present a novel hybrid equalization scheme for ATSC channels. The proposed scheme consists of an adaptive decision feedback sparsening filter (DFSF), and an iterative maximum a posteriori (MAP) equalizer based on the belief propagation (BP) algorithm. In the first stage, instead of removing all the ISI from post cursors, the DFSF employs a modified feedback filter which leaves the strongest post-cursor ISI taps uncorrected. As a result, a long ISI channel is equalized to a sparse channel having only a small number of nonzero taps. In the second stage, a belief propagation algorithm is applied to mitigate the residual ISI. Since the channel is typically time-varying and suffers from Doppler fading, the DFSF is adapted using the least mean square (LMS) algorithm, such that the amplitude and the locations of the nonzero taps of the equalized sparse channel appear to be fixed. As such, the channel appears to be static during the second stage of equalization which consists of the BP detector. Simulation results demonstrate that the proposed scheme outperforms the traditional DFE in symbol error rate, under both static channels and dynamic ATSC channels. Yanjie Peng, Andrew G. Klein, Xinming Huang 0001 |
ISCAS | 3 |
| 2014 | Accelerating leveled fully homomorphic encryption using GPUabstractGentry introduced the first plausible fully homomorphic encryption (FHE) scheme, which was considered a major breakthrough in cryptography. Several FHE schemes have been proposed to make FHE more efficient for practical applications since then. The leveled fully homomorphic scheme is among the most well-known schemes. In leveled FHE scheme, large-number matrix-vector multiplication is a crucial part of the encryption algorithm. In this paper, Chinese Remainder Theorem (CRT) is employed to reduce the computational complexity of the large-number element-by-element modular multiplication. The first step is called decomposition, in which each large-number element in the matrix and vector is decomposed into many small words. The next step is vector operation that performs the modular multiplications and additions of the decomposed small words. Finally the matrix-vector multiplication results can be obtained through reconstruction. We compare the CRTbased method with Number Theory Library (NTL), showing the proposed method is about 7.8 times faster when executing on CPU. In addition, it is observed that vector operation takes up to 99.6% of the total computation time and the reconstruction only takes 0.4%. Therefore GPU acceleration is employed to speed up the vector operations. Experiment results show that the GPU implementation of the CRT-based method is 35.2 times faster than the same method implemented on CPU and is 273.6 times faster than the NTL library on CPU. Wei Wang 0053, Zhilu Chen, Xinming Huang 0001 |
ISCAS | 3 |
| 2014 | VLSI Design of a Large-Number Multiplier for Fully Homomorphic EncryptionabstractThis paper presents the design of a power- and area-efficient high-speed 768000-bit multiplier, based on fast Fourier transform multiplication for fully homomorphic encryption operations. A memory-based in-place architecture is presented for the FFT processor that performs 64000-point finite-field FFT operations using a radix-16 computing unit and 16 dual-port SRAMs. By adopting a special prime as the base of the finite field, the radix-16 calculations are simplified to requiring only additions and shift operations. A two-stage carry-look-ahead scheme is employed to resolve carries and obtain the multiplication result. The multiplier design is validated by comparing its results with the GNU Multiple Precision (GMP) arithmetic library. The proposed design has been synthesized using 90-nm process technology with an estimated die area of 45.3 mm2. At 200 MHz, the large-number multiplier offers roughly twice the performance of a previous implementation on an NVIDIA C2050 graphics processor unit and is 29 times faster than the Xeon X5650 CPU, while at the same time consuming a modest 0.97 W. Wei Wang 0053, Xinming Huang 0001, Niall Emmart, Charles C. Weems |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | An FPGA co-processor for adaptive lane departure warning systemabstractThis paper presents an FPGA co-processor design for adaptive lane departure warning system, which requires intensive computation for real-time video image processing. The main functions of the co-processor include color scheme conversion, 2D filtering, transferring intensity into binary pattern using Otsu's threshold method, and detecting lanes using Hough transform. The system design is implemented on a Xilinx Kintex FPGA with FMC DVI module connected to a camera. Our experimental tests prove the design is fully functional and the FPGA-based implementation meets the real-time requirement of reporting lane departure for practical applications. Wei Wang 0053, Xinming Huang 0001 |
ISCAS | 2 |
| 2013 | FPGA implementation of a large-number multiplier for fully homomorphic encryptionabstractThe first plausible scheme of fully homomorphic encryption (FHE), introduced by Gentry in 2009, was considered a major breakthrough in the field of information security. FHE allows the evaluation of arbitrary functions directly on encrypted data on untrusted servers. However, previous implementations of FHE on general-purpose processors had very long latency, which makes it impractical for cloud computing. The most computationally intensive components in the Gentry-Halevi FHE primitives are the large-number modular multiplications and additions. In this paper, we attempt to use customized circuits to speedup the large number multiplication. Strassen's algorithm is employed in the design of an efficient, high-speed large-number multiplier. In particular, we propose an architecture design of an 768K-bit multiplier. As a key compoment, an 64K-point finite-field fast Fourier transform (FFT) processor is designed and prototyped on the Stratix-V FPGA. At 100 MHz, the FPGA implementation is about twice as fast as the same FFT algorithm executed on the NVIDA C2050 GPU which has 448 cores running at 1.15 GHz but at much lower power consumption. Wei Wang 0053, Xinming Huang 0001 |
ISCAS | 2 |
| 2013 | Design and Implementation of a Low-Complexity Symbol Detector for Sparse ChannelsabstractIn this paper, we present a low-complexity symbol detector for communication channels which have long spanning durations but a sparse multipath structure. Traditional maximum-likelihood sequence estimation using the Viterbi algorithm can provide optimal error performance for eliminating the multipath effect, but the hardware complexity grows exponentially with channel length and it is not practical for long sparse channels. We implement a near-optimal algorithm and its architecture by cascading an adaptive partial response equalizer (PRE) with an iterative belief propagation (BP) detector. A sparse channel is first equalized by a PRE to a target impulse response (TIR) with only a few nonzero coefficients remaining. The residual intersymbol interference is then canceled by a BP detector whose complexity is solely dependent on the number of nonzero coefficients in the TIR. Moreover, we present a pipeline high-throughput implementation of the detector for channel length 30 with quadrature phase-shift keying modulation. The detector can achieve a maximum throughput of 206 Mb/s with an estimated core area of 3.162 mm2using 90-nm technology node. At a target frequency of 515 MHz, the dynamic power is about 1.096 W. Yanjie Peng, Xinming Huang 0001, Andrew G. Klein, Kai Zhang 0025 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Distributed cross-layer optimization for wireless regional area network-based cognitive radio networksabstractABSTRACT This paper presents a study of a cross‐layer design through joint optimization of spectrum allocation and power control for cognitive radio networks (CRNs). The spectrum of interest is divided into independent channels licensed to a set of primary users (PUs). The secondary users are activated only if the transmissions do not cause excessive interference to PUs. In particular, this paper studies the downlink channel assignment and power control in a CRN with the coexistence of PUs and secondary users. The objective was to maximize the total throughput of a CRN. A mathematical model is presented and subsequently formulated as a binary integer programming problem, which belongs to the class of non‐deterministic polynomial‐time hard problems. Subsequently, we develop a distributed algorithm to obtain sub‐optimal results with lower computational complexity. The distributed algorithm iteratively improves the network throughput, which consists of several modules including maximum power calculation, excluded channel sets recording, base station throughput estimation, base station sorting, and channel usage implementation. Through investigating the impacts of the different parameters, simulation results demonstrates that the distributed algorithm can achieve a better performance than two other schemes. Copyright © 2011 John Wiley & Sons, Ltd. Xinming Huang 0001 |
Wirel. Commun. Mob. Comput. | 2 |
| 2012 | High-Speed Low-Power Viterbi Decoder Design for TCM DecodersabstractHigh-speed, low-power design of Viterbi decoders for trellis coded modulation (TCM) systems is presented in this paper. It is well known that the Viterbi decoder (VD) is the dominant module determining the overall power consumption of TCM decoders. We propose a pre-computation architecture incorporated with T-algorithm for VD, which can effectively reduce the power consumption without degrading the decoding speed much. A general solution to derive the optimal pre-computation steps is also given in the paper. Implementation result of a VD for a rate-3/4 convolutional code used in a TCM system shows that compared with the full trellis VD, the precomputation architecture reduces the power consumption by as much as 70% without performance loss, while the degradation in clock speed is negligible. Jinjin He, Huaping Liu 0002, Zhongfeng Wang 0001, Xinming Huang 0001, Kai Zhang 0025 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | Design and implementation of a belief propagation detector for sparse channelsabstractIn this paper, we address the design and implementation of the symbol detector for sparse channels which are described as having long spanning durations but sparse multipath structure. The traditional maximum-likelihood (ML) algorithm provides an optimal performance to eliminate the multipath effect, however its complexity scales exponentially with the channel length. As a more efficient symbol detection algorithm through sparse channels, the iterative belief propagation (BP) algorithm has a complexity merely dependent on the number of nonzero channel coefficients, while achieving a near-optimal error performance. We present the architecture design for a reconfigurable low-complexity high-throughput BP detector. As an example, we implement a BP detector for quadrature phase-shift keying (QPSK) modulation on Xilinx Virtex 5 FPGA with a maximum frequency of 252 MHz and equivalently a throughput of 100.8 Mb/s at 5 iterations. Yanjie Peng, Kai Zhang 0025, Andrew G. Klein, Xinming Huang 0001 |
ASAP | 4 |
| 2011 | Achieving capacity fairness for wireless mesh networksabstractAbstract This paper addresses a joint problem of power control and channel assignment within a wireless mesh network. A wireless mesh network is made up of two kinds of nodes: mesh routers (MRs) and user nodes (UNs). The MRs form a backbone network, while UNs receive data from the backbone network by connecting to the MRs via one hop. This paper aims to find the optimal joint solution of power control and channel assignment of the wireless mesh networks such that the minimum capacity of all links is maximized. We develop an upper bound for the objective by relaxing the integer variables and linearization. Subsequently, we put forward a heuristic approach to approximate the optimal solution, which tries to increase the minimal capacity of all links via setting tighter constraint and solving a binary integer programming problem. Simulation results show that solutions obtained by this algorithm are very close to the upper bounds obtained via relaxation, thus suggesting that the solution produced by the algorithm is near‐optimal. Copyright © 2009 John Wiley & Sons, Ltd. Xinming Huang 0001 |
Wirel. Commun. Mob. Comput. | 2 |
| 2010 | Dynamic relay deployment for disaster area wireless networksabstractAbstract This paper investigates the disaster area communication system using relay‐assisted wireless network for first responders as mobile nodes (MNs). Firstly, a novel mobility model is proposed to describe the movement pattern of MNs within a large disaster area. Secondly, we study the relay management problem of finding a minimum number of relay nodes (RNs) and their dynamic locations to cover all the MNs within the disaster area. A square disk cover (SDC) problem is formulated and three different algorithms, including the two‐vertex square covering (TVSC) algorithm, the circle covering algorithm and the binary integer programming (BIP) algorithm, are proposed to solve the SDC problem. Simulation results are presented to validate the mobility model and compare the algorithms with respect to computational complexity and worst case performance ratio. Copyright © 2008 John Wiley & Sons, Ltd. Xinming Huang 0001, Youjian Liu |
Wirel. Commun. Mob. Comput. | 2 |
| 2009 | Mapping Parallel FFT Algorithm onto SmartCell Coarse-Grained Reconfigurable ArchitectureabstractThis paper presents the implementation of a novel parallel FFT algorithm on SmartCell, a coarse-grained reconfigurable architecture, which is targeted on data streaming applications. The proposed FFT algorithm achieves balanced workload and memory requirement among the computational units, while maintaining optimized data flow at low configuration and communication cost. The proposed parallel FFT algorithm is then mapped onto the SmartCell prototype device with 64 processing elements. Results show that the parallel FFT implementation on SmartCell is about 14.9 and 2.7 times faster than network-on-chip (NoC) and Morphosys, respectively. The implementation also shows about 3.6 times better energy efficiency when comparing with the pipelined FFT implementations on FPGA. Cao Liang, Xinming Huang 0001 |
ASAP | 2 |
| 2009 | An Area-Efficient LDPC Decoder Architecture and Implementation for CMMB SystemsabstractThis paper presents an area-efficient LDPC decoder architecture for the China multimedia mobile broadcasting (CMMB) standard. Several techniques are adopted to reduce memory size, including the min-sum algorithm (MSA), optimal bit-width quantization of the iterative messages and reduced complexity for the interconnect network. The decoder for the rate-1/2 9216-bit code is implemented using the 90 nm 1.0 V CMOS technology. It achieves the decoding throughput of 48 Mbps at 5 iterations when operating at 60 MHz and the power dissipation is only 34 mW. Kai Zhang 0025, Xinming Huang 0001, Zhongfeng Wang 0001 |
ASAP | 2 |
| 2009 | Design of a maximum-likelihood detector for cooperative communications in intersymbol interference channelsabstractRecently, cooperative communication has attracted a lot of attention for its potential to increase spatial diversity. However, limited attention has been paid to the physical layer and implementation issues. In this paper, we investigate the feasibility of building optimal detectors for cooperative communications in intersymbol interference channels. A novel system model with the amplify-and-forward half-duplex relay is first introduced. Subsequently, we propose an optimal detector for receiver design which can be realized with a whitening filter and maximum likelihood sequence estimator (MLSE) based on the Viterbi algorithm. The implementation of the proposed detector is discussed, and its complexity is analyzed based on an FPGA implementation result. Numerical simulations demonstrating the performance of the detector are provided. Yanjie Peng, Andrew G. Klein, Xinming Huang 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2009 | High-throughput layered decoder implementation for quasi-cyclic LDPC codesabstractThis paper presents a high-throughput decoder design for the Quasi-Cyclic (QC) Low-Density Parity-Check (LDPC) codes. Two new techniques are proposed, including parallel layered decoding architecture (PLDA) and critical path splitting. PLDA enables parallel processing for all layers by establishing dedicated message passing paths among them. The decoder avoids crossbar-based large interconnect network. Critical path splitting technique is based on articulate adjustment of the starting point of each layer to maximize the time intervals between adjacent layers, such that the critical path delay can be split into pipeline stages. Furthermore, min-sum and loosely coupled algorithms are employed for area efficiency. As a case study, a rate-1/2 2304-bit irregular LDPC decoder is implemented using ASIC design in 90 nm CMOS process. The decoder can achieve the maximum decoding throughput of 2.2 Gbps at 10 iterations. The operating frequency is 950 MHz after synthesis and the chip area is 2.9 mm2. Kai Zhang 0025, Xinming Huang 0001, Zhongfeng Wang 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2008 | Power Scheduling for MIMO Relay Channels Employing Rateless CodesabstractWe propose a simple power scheduling algorithm for relay channels with multiple antennas. The algorithm combines seamlessly with a low complexity communication protocol that employs rateless coding for the relay channel. The goal of the algorithm is to achieve a target average total transmission power. It requires little signaling and can be used whenever the destination node can predict future channel states. The scheduling works by a power on/off control of the source and relay nodes. It has about 1 dB gain in low SNR. It has little gain in moderate to high SNR due to the very small amount of signaling, but can be used to achieve a range of target average total transmission power without adjusting the on-power. Youjian Liu, Mahesh K. Varanasi, Xinming Huang 0001 |
ICC | 3 |
| 2008 | A high SFDR direct digital synthesizer with frequency error free outputabstractIn this paper, an error-compensation method is proposed to achieve high spurious free dynamic range (SFDR) in direct digital synthesizers (DDS) design. This method allows very small ROM storage, while leaving the resultant errors corrected by the error-compensation circuits. An extended phase accumulator (EPA) is also adopted to provide arbitrary output frequency, thus can achieve frequency error free output. Experimental results show that the DDS using only 256 bits of ROM can achieve 104 dBc SFDR for output signal at 20 MHz with a clock frequency of 100 MHz. The effect of EPA is also demonstrated in the design. Kai Zhang 0025, Xinming Huang 0001 |
ISCAS | 2 |
| 2008 | Mobility Model and Relay Management for Disaster Area Wireless Networks
Xinming Huang 0001 |
WASA | 2 |
| 2008 | On Relay Node Placement and Assignment for Two-tiered Wireless Networks
Xinming Huang 0001, Wenjing Lou, Cao Liang |
Mob. Networks Appl. | 2 |
| 2008 | System Architecture and Implementation of MIMO Sphere Decoders on FPGAabstractMultiple-input-multiple-output (MIMO) systems use multiple antennas in both transmitter and receiver ends for higher spectrum efficiency. The hardware implementation of MIMO detection becomes a challenging task as the computational complexity increases. This paper presents the architectures and implementations of two typical sphere decoding algorithms, including the Viterbo-Boutros (VB) algorithm and the Schnorr-Euchner (SE) algorithm. Hardware/software codesign technique is applied to partition the decoding algorithm on a single field-programmable gate array (FPGA) device. Three levels of parallelism are explored to improve the decoding rate: the concurrent execution of the channel matrix preprocessing on an embedded processor and the decoding functions on customized hardware modules, the parallel decoding of real/imaginary parts for complex constellation, and the concurrent execution of multiple steps during the closest lattice point search. The decoders for a 4times4 MIMO system with 16-QAM modulation are prototyped on a Xilinx XC2VP30 FPGA device with a MicroBlaze soft core processor. The hardware prototypes of the SE and VB algorithms show that they support up to 81.5 and 36.1 Mb/s data rates at 20 dB signal-to-noise ratio, which are about 22 and 97 times faster than their respective implementations in a digital signal processor. Xinming Huang 0001, Cao Liang, Jing Ma 0006 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2007 | Optimal Relay Node Association for Two-Tiered Wireless NetworksabstractWireless networks that operate on batteries are imposed with energy constraints and long distance communications between nodes are not desirable. Implementing relay nodes can improve network capacity and save communication energy. A two-hop relay routing scheme is considered, in which the relay nodes are temporarily placed and have energy constraints. This paper investigates a joint optimization problem on relay node placement and route assignment for two-tiered wireless networks. A recursive weighted clustering binary integer programming (WCBIP) algorithm is proposed to maximize the total number of information packets received at the BS during the network lifetime. We first present an optimization algorithm based on binary integer programming (BIP) for relay node assignment with the current node locations. Subsequently, a weighted clustering algorithm is applied to move the relay nodes to the best locations to best serve their respectively associated edge nodes. The algorithm has the complexity of O(2n). The simulation results show that the proposed algorithm has significantly better performance than the other two relay placement schemes. Both theoretical analysis and practical design procedures are also presented with details. Xinming Huang 0001 |
GLOBECOM | 2 |
| 2006 | Hardware/Software Co-Design Architecture for Lattice Decoding AlgorithmsabstractThis paper presents hardware/software co-design architecture targeted on a single FPGA for two typical lattice decoding algorithms in MIMO system. Two levels of parallelisms are analyzed for an efficient implementation with the preprocessing part on embedded MicroBlaze soft processor and the decoder part on customized hardware. The system prototypes of the AV and VB decoders show that they support up to 34.2 Mbps and 3.15 Mbps data rate respectively on XUP Virtex-II pro developing board, which are 19 and 16 times faster than their respective implementations on a DSP Cao Liang, Jing Ma 0006, Xinming Huang 0001 |
FCCM | 3 |
| 2005 | A System-on-Programmable Chip Approach for MIMO Sphere DecoderabstractThis paper presents a system-on-programmable chip approach for a Schnorr-Euchner strategy based sphere decoder. The decoding algorithm is partitioned with the division intensive matrix computation part on embedded microprocessor and the iterative lattice decoding function on programmable logics. Efficient hardware architectures are developed with three levels of parallelism. The system prototype of the decoder shows that it supports 35.75 Mbit/s data rate on an Altera Stratix EP1S10 FPGA located on Nios development board for a four antenna system with 16-QAM signal constellation, and is about 20 times faster than its implementation on a DSP. Jing Ma 0006, Xinming Huang 0001 |
FCCM | 2 |