EDBT 2026 Demo / reviewers in the wild / expert
Hongyi Pan
dblp:02/7832
· DBLP profile ↗
24ranked-venue papers
8as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VHU-Net: Variational hadamard U-Net for body MRI bias field correction
Xin Zhu 0005, A. Enis Çetin, Gorkem Durak, Batuhan Gündogdu, Ziliang Hong, Hongyi Pan, Halil Ertugrul Aktas, Elif Keles, Hatice Savas, Aytekin Oto, Hiten D. Patel, Adam B. Murphy, Ashley Ross, Frank H. Miller, Baris Turkbey, Ulas Bagci |
Medical Image Anal. | 6 |
| 2025 | Classroom Simulacra: Building Contextual Student Generative Agents in Online Education for Learning Behavioral SimulationabstractStudent simulation supports educators to improve teaching by interacting with virtual students.However, most existing approaches ignore the modulation effects of course materials because of two challenges: the lack of datasets with granularly annotated course materials, and the limitation of existing simulation models in processing extremely long textual data.To solve the challenges, we first run a 6-week education workshop from N = 60 students to collect fine-grained data using a custom built online education system, which logs students' learning behaviors as they interact with lecture materials over time.Second, we propose a transferable iterative reflection (TIR) module that augments both prompting-based and finetuning-based large language models (LLMs) for simulating learning behaviors.Our comprehensive experiments show that TIR enables the LLMs to perform more accurate student simulation than classical deep learning models, even with limited demonstration data.Our TIR approach better captures the granular dynamism of learning performance and inter-student correlations in classrooms, paving the way towards a "digital twin" for online education. Songlin Xu, Hao-Ning Wen, Hongyi Pan, Dallas Dominguez, Dongyin Hu, Xinyu Zhang 0003 |
CHI | 3 |
| 2025 | MDNet: Multi-Decoder Network for Abdominal CT Organs SegmentationabstractAccurate segmentation of organs from abdominal CT scans is essential for clinical applications such as diagnosis, treatment planning, and patient monitoring. To handle challenges of heterogeneity in organ shapes, sizes, and complex anatomical relationships, we propose a Multi decoder network (MDNet), an encoder-decoder network that uses the pre-trained MiT-B2 as the encoder and multiple different decoder networks. Each decoder network is connected to a different part of the encoder via a multi-scale feature enhancement dilated block. With each decoder, we increase the depth of the network iteratively and refine segmentation masks, enriching feature maps by integrating previous decoders’ feature maps. To refine the feature map further, we also utilize the predicted masks from the previous decoder to the current decoder to provide spatial attention across foreground and background regions. MDNet effectively refines the segmentation mask with a high dice similarity coefficient (DSC) of 0.9013 and 0.9169 on the Liver Tumor segmentation (LiTS) and MSD Spleen datasets. Additionally, it reduces Hausdorff distance (HD) to 3.79 for the LiTS dataset and 2.26 for the spleen segmentation dataset, underscoring the precision of MDNet in capturing the complex contours. Moreover, MDNet is more interpretable and robust compared to the other baseline models. The code for our architecture is available at https://github.com/DebeshJha/MDNet. Debesh Jha, Nikhil Kumar Tomar, Koushik Biswas, Gorkem Durak, Matthew Antalek, Zheyuan Zhang 0001, Bin Wang 0068, Md Mostafijur Rahman, Hongyi Pan, Alpay Medetalibeyoglu, Vandan Gorade, Yury Velichko, Daniela P. Ladner, Amir Borhani, Ulas Bagci |
ICASSP | 9 |
| 2025 | Frequency-Based Federated Domain Generalization for Polyp SegmentationabstractFederated Learning (FL) offers a powerful strategy for training machine learning models across decentralized datasets while maintaining data privacy, yet domain shifts among clients can degrade performance, particularly in medical imaging tasks like polyp segmentation. This paper introduces a novel Frequency-Based Domain Generalization (FDG) framework, utilizing soft-thresholding and hard-thresholding in the Fourier domain to address these challenges. By applying soft-thresholding and hard-thresholding to Fourier coefficients, our method generates new images with reduced background noise and enhances the model’s ability to generalize across diverse medical imaging domains. Extensive experiments demonstrate substantial improvements in segmentation accuracy and domain robustness over baseline methods. This innovation integrates frequency domain techniques into FL, presenting a resilient approach to overcoming domain variability in decentralized medical image analysis. Hongyi Pan, Debesh Jha, Koushik Biswas, Ulas Bagci |
ICASSP | 1 |
| 2025 | VideoAds for Fast-Paced Video Understanding
Zheyuan Zhang 0001, Wanying Dou, Linkai Peng, Hongyi Pan, Ulas Bagci, Boqing Gong |
ICCV | 4 |
| 2025 | Optimizing Neural Network Effectiveness via Non-monotonicity RefinementabstractActivation functions play a crucial role in artificial neural networks by introducing non-linearities that enable networks to learn complex patterns in data. An appropriate choice of an activation function plays a crucial role in the training dynamics of a neural network, which can boost network performance significantly. Rectified Linear Unit (ReLU) and its variants, like leaky ReLU and parametric ReLU, have emerged as the most popular activations due to their ability to enable faster training and generalization in deep neural networks despite having some significant issues like vanishing gradient problems. In this paper, we have proposed smooth functions, which we call the AMSU family, which are smooth approximations of the maximum function. We derive three activations from the AMSU family, namely AMSU-1, AMSU-2, & AMSU-3, and show their effectiveness in different deep learning problems. By simply replacing the ReLU function, Top-1 accuracy improves by 5.88%, 5.96%, and 5.32% on the CIFAR100 dataset on the ShuffleNet V2 model. Also, replacing ReLU with AMSU-1, AMSU-2, and AMSU-3, Top-1 accuracy improves by 8.50%, 8.29%, and 7.70% on the CIFAR100 dataset on the ShuffleNet V2 model with FGSM attack. Also, Replacing ReLU with AMSU-1, AMSU-2, and AMSU-3 on ImageNet-1K data, we got 3%-5% improvement on ShuffleNet and MobileNet models. The source code is publicly available at https://github.com/koushik313/AMSU. Koushik Biswas, Amit Reza, Meghana Karri, Debesh Jha, Hongyi Pan, Nikhil Kumar Tomar, Aliza Subedi, Smriti Regmi, Ulas Bagci |
WACV | 5 |
| 2025 | Edge-Fog Computing-Enabled EEG Data Compression via Asymmetrical Variational Discrete Cosine Transform NetworkabstractThe large volume of electroencephalograph (EEG) data produced by brain-computer interface (BCI) systems presents challenges for rapid transmission over bandwidth-limited channels in Internet of Things (IoT) networks. To address the issue, we propose a novel multichannel asymmetrical variational discrete cosine transform (DCT) network for EEG data compression within an edge-fog computing framework. At the edge level, low-complexity DCT compression units are designed using parallel trainable hard-thresholding and scaling operators to remove redundant data and extract the effective latent space representation. At the fog level, an adaptive filter bank is applied to merge important features from adjacent channels into each individual channel by leveraging interchannel correlations. Then, the inverse DCT reconstructed multihead attention is developed to capture both local and global dependencies and reconstruct the original signals. Furthermore, by applying the principles of variational inference, a new evidence lower bound is formulated as the loss function, driving the model to balance compression efficiency and reconstruction accuracy. Experimental results on two public datasets demonstrate that the proposed method achieves superior compression performance without sacrificing any useful information for BCI detection compared with state-of-the-art techniques, indicating a feasible solution for EEG data compression. Xin Zhu 0005, Hongyi Pan, A. Enis Çetin |
IEEE Internet Things J. | 2 |
| 2025 | Large-scale multi-center CT and MRI segmentation of pancreas with deep learningabstractAutomated volumetric segmentation of the pancreas on cross-sectional imaging is needed for diagnosis and follow-up of pancreatic diseases. While CT-based pancreatic segmentation is more established, MRI-based segmentation methods are understudied, largely due to a lack of publicly available datasets, benchmarking research efforts, and domain-specific deep learning methods. In this retrospective study, we collected a large dataset (767 scans from 499 participants) of T1-weighted (T1 W) and T2-weighted (T2 W) abdominal MRI series from five centers between March 2004 and November 2022. We also collected CT scans of 1,350 patients from publicly available sources for benchmarking purposes. We introduced a new pancreas segmentation method, called PanSegNet , combining the strengths of nnUNet and a Transformer network with a new linear attention module enabling volumetric computation. We tested PanSegNet ’s accuracy in cross-modality (a total of 2,117 scans) and cross-center settings with Dice and Hausdorff distance (HD95) evaluation metrics. We used Cohen’s kappa statistics for intra and inter-rater agreement evaluation and paired t-tests for volume and Dice comparisons, respectively. For segmentation accuracy, we achieved Dice coefficients of 88.3% (±7.2%, at case level) with CT, 85.0% (±7.9%) with T1 W MRI, and 86.3% (±6.4%) with T2 W MRI. There was a high correlation for pancreas volume prediction with R 2 of 0.91, 0.84, and 0.85 for CT, T1 W, and T2 W, respectively. We found moderate inter-observer (0.624 and 0.638 for T1 W and T2 W MRI, respectively) and high intra-observer agreement scores. All MRI data is made available at https://osf.io/kysnj/ . Our source code is available at https://github.com/NUBagciLab/PaNSegNet . • We develop a first-ever cross-platform compatible (T1 W, T2 W, and CT) pancreas segmentation tool, named PanSegNet . • PaNSegNet has innovative “linear self-attention” blocks to reduce computational cost significantly while operating on 3D. • We shared our both source code and multi-center multi-contrast MRI datasets with ground truths. • PaNSegNet underwent rigorous validation, including cross-domain and multi-center comparisons between CT and MRI scans. Zheyuan Zhang 0001, Elif Keles, Gorkem Durak, Yavuz Taktak, Onkar Susladkar, Vandan Gorade, Debesh Jha, Asli C. Ormeci, Alpay Medetalibeyoglu, Lanhong Yao, Bin Wang 0068, Ilkin Isler, Linkai Peng, Hongyi Pan, Camila Lopes Vendrami, Amir Bourhani, Yury Velichko, Boqing Gong, Concetto Spampinato, Ayis Pyrros, Pallavi Tiwari, Derk C. F. Klatte, Megan Engels, Sanne Hoogenboom, Candice W. Bolan, Emil Agarunov, Nassier Harfouch, Chenchan Huang, Marco J. Bruno, Ivo Schoots, Rajesh Keswani, Frank H. Miller, Tamas Gonda, Cemal Yazici, Temel Tirkes, Baris Turkbey, Michael B. Wallace, Ulas Bagci |
Medical Image Anal. | 14 |
| 2025 | Multichannel Orthogonal Transform-Based Perceptron Layers for Efficient ResNetsabstractIn this article, we propose a set of transform-based neural network layers as an alternative to the Conv2D layers in convolutional neural networks (CNNs). The proposed layers can be implemented based on orthogonal transforms, such as the discrete cosine transform (DCT), Hadamard transform (HT), and biorthogonal block wavelet transform (BWT). Furthermore, by taking advantage of the convolution theorems, convolutional filtering operations are performed in the transform domain using elementwise multiplications. Trainable soft-thresholding layers, that remove noise in the transform domain, bring nonlinearity to the transform domain layers. Compared with the Conv2D layer, which is spatial-agnostic and channel-specific, the proposed layers are location-specific and channel-specific. Moreover, these proposed layers reduce the number of parameters and multiplications significantly while improving the accuracy results of regular ResNets on the ImageNet-1K classification task. Furthermore, they can be inserted with a batch normalization (BN) layer before the global average pooling layer in the conventional ResNets as an additional layer to improve classification accuracy. Hongyi Pan, Emadeldeen Hamdan, Xin Zhu 0005, Salih Atici, A. Enis Çetin |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Domain Generalization with fourier Transform and soft thresholdingabstractDomain generalization aims to train models on multiple source domains so that they can generalize well to unseen target domains. Among many domain generalization methods, Fourier-transformbased domain generalization methods have gained popularity primarily because they exploit the power of Fourier transformation to capture essential patterns and regularities in the data, making the model more robust to domain shifts. The mainstream Fouriertransform-based domain generalization swaps the Fourier amplitude spectrum while preserving the phase spectrum between the source and the target images. However, it neglects background interference in the amplitude spectrum. To overcome this limitation, we introduce a soft-thresholding function in the Fourier domain. We apply this newly designed algorithm to retinal fundus image segmentation, which is important for diagnosing ocular diseases but the neural network’s performance can degrade across different sources due to domain shifts. The proposed technique basically enhances fundus image augmentation by eliminating small values in the Fourier domain and providing better generalization. The innovative nature of the soft thresholding fused with Fourier-transform-based domain generalization improves neural network models’ performance by reducing the target images’ background interference significantly. Experiments on public data validate our approach’s effectiveness over conventional and state-of-the-art methods with superior segmentation metrics. Hongyi Pan, Bin Wang 0068, Zheyuan Zhang 0001, Xin Zhu 0005, Debesh Jha, A. Enis Çetin, Concetto Spampinato, Ulas Bagci |
ICASSP | 1 |
| 2024 | Stein Variational Gradient Descent-Based Detection for Random Access with Preambles in MTCabstractTraditional preamble detection algorithms have low accuracy in the grant-based random access scheme in massive machine-type communication (mMTC). We present a novel preamble detection algorithm based on Stein variational gradient descent (SVGD) at the second step of the random access procedure. It efficiently leverages deterministic updates of particles for continuous inference. To further enhance the performance of the SVGD detector, especially in a dense user scenario, we propose a normalized SVGD detector with momentum. It utilizes the momentum and a bias correction term to reduce the preamble estimation errors during the gradient descent process. Simulation results show that the proposed algorithm performs better than Markov Chain Monte Carlo-based approaches in terms of detection accuracy. Xin Zhu 0005, Hongyi Pan, Salih Atici, A. Enis Çetin |
ICASSP | 2 |
| 2024 | Electroencephalogram Sensor Data Compression Using an Asymmetrical Sparse Autoencoder with a Discrete Cosine Transform LayerabstractElectroencephalogram (EEG) data compression is necessary for wireless recording applications to reduce the amount of data that needs to be transmitted. In this paper, an asymmetrical sparse autoencoder with a discrete cosine transform (DCT) layer is proposed to compress EEG signals. The encoder module of the autoencoder has a combination of a fully connected linear layer and the DCT layer to reduce redundant data using hard-thresholding nonlinearity. Furthermore, the DCT layer includes trainable hard-thresholding parameters and scaling layers to give emphasis or de-emphasis on individual DCT coefficients. Finally, the one-by-one convolutional layer generates the latent space. The sparsity penalty-based cost function is employed to keep the feature map as sparse as possible in the latent space. The latent space data is transmitted to the receiver. The decoder module of the autoencoder is designed using the inverse DCT and two fully connected linear layers to improve the accuracy of data reconstruction. In comparison to other state-of-the-art methods, the proposed method significantly improves the average quality score in various data compression experiments. Xin Zhu 0005, Hongyi Pan, Shuaiang Rong, A. Enis Çetin |
ICASSP | 2 |
| 2024 | AlN Sputtering Parameter Estimation Using A Multichannel Parallel DCT Neural NetworkabstractIn this paper, we present a method for estimating the deposition parameters of the thin film material Aluminum Nitride (AIN) using a deep neural network. The neural network predicts the AIN orientations, which are critical for micromachining Micro-Electro-Mechanical Systems (MEMS) transducers such as accelerometers and acoustic emission sensors. The network features three parallel channels, each equipped with a Discrete Cosine Transform (DCT) based layer that encodes the input parameters into a latent space. This DCT layer applies a hard-thresholding nonlinearity to eliminate noise from the input parameters, resulting in a sparse representation of the latent space. Trained with a dataset comprising AlN orientations parameters and their optimal values, our model is adept at simultaneously extracting and integrating various essential frequency components. Experimental results underscore the effectiveness of our proposed approach in achieving accurate and comprehensive estimation of AlN orientations and MEMS design parameters, thereby providing a promising path for advanced optimization. Yingyi Luo, Talha M. Khan, Emadeldeen Hamdan, Xin Zhu 0005, Hongyi Pan, Didem Ozevin, A. Enis Çetin |
VTS | 5 |
| 2024 | GazeGNN: A Gaze-Guided Graph Neural Network for Chest X-ray ClassificationabstractEye tracking research is important in computer vision because it can help us understand how humans interact with the visual world. Specifically for high-risk applications, such as in medical imaging, eye tracking can help us to comprehend how radiologists and other medical professionals search, analyze, and interpret images for diagnostic and clinical purposes. Hence, the application of eye tracking techniques in disease classification has become increasingly popular in recent years. Contemporary works usually transform gaze information collected by eye tracking devices into visual attention maps (VAMs) to supervise the learning process. However, this is a time-consuming preprocessing step, which stops us from applying eye tracking to radiologists’ daily work. To solve this problem, we propose a novel gaze-guided graph neural network (GNN), GazeGNN, to leverage raw eye-gaze data without being converted into VAMs. In GazeGNN, to directly integrate eye gaze into image classification, we create a unified representation graph that models both images and gaze pattern information. With this benefit, we develop a real-time, real-world, end-to-end disease classification algorithm for the first time in the literature. This achievement demonstrates the practicality and feasibility of integrating real-time eye tracking techniques into the daily work of radiologists. To our best knowledge, GazeGNN is the first work that adopts GNN to integrate image and eye-gaze data. Our experiments on the public chest X-ray dataset show that our proposed method exhibits the best classification performance compared to existing methods. The code is available at https://github.com/ukaukaaaa/GazeGNN. Bin Wang 0068, Hongyi Pan, Armstrong Aboah, Zheyuan Zhang 0001, Elif Keles, Drew A. Torigian, Baris Turkbey, Elizabeth A. Krupinski, Jayaram K. Udupa, Ulas Bagci |
WACV | 2 |
| 2024 | A novel asymmetrical autoencoder with a sparsifying discrete cosine Stockwell transform layer for gearbox sensor data compression
Xin Zhu 0005, Daoguang Yang, Hongyi Pan, Hamid Reza Karimi, Didem Ozevin, A. Enis Çetin |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | ADC/DAC-Free Analog Acceleration of Deep Neural Networks With Frequency TransformationabstractThe edge processing of deep neural networks (DNNs) is becoming increasingly important due to its ability to extract valuable information directly at the data source to minimize latency and energy consumption. Although pruning techniques are commonly used to reduce model size for edge computing, they have certain limitations. Frequency-domain model compression, such as with the Walsh–Hadamard transform (WHT), has been identified as an efficient alternative. However, the benefits of frequency-domain processing are often offset by the increased multiply-accumulate (MAC) operations required. This article proposes a novel approach to an energy-efficient acceleration of frequency-domain neural networks by utilizing analog-domain frequency-based tensor transformations. Our approach offers unique opportunities to enhance computational efficiency, resulting in several high-level advantages, including array microarchitecture with parallelism, analog-to-digital converter (ADC)/digital-to-analog converter (DAC)-free analog computations, and increased output sparsity. Our approach achieves more compact cells by eliminating the need for trainable parameters in the transformation matrix. Moreover, our novel array microarchitecture enablesadaptive stitchingof cells column-wise and row-wise, thereby facilitating perfect parallelism in computations. Additionally, our scheme enables ADC/DAC-free computations by training against highly quantized matrix-vector products, leveraging the parameter-free nature of matrix multiplications. Another crucial aspect of our design is its ability to handle signed-bit processing for frequency-based transformations. This leads to increased output sparsity and reduced digitization workload. On a$16 \ttimes 16$crossbars, for 8-bit input processing, the proposed approach achieves the energy efficiency of 801 tera operations per second per Watt (TOPS/W) without early termination strategy and 2655 TOPS/W with early termination strategy at VDD$=$0.85 V for 16-nm predictive technology models (PTM). Nastaran Darabi, Maeesha Binte Hashem, Hongyi Pan, A. Enis Çetin, Wilfred Gomes, Amit Ranjan Trivedi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | Classification of the Cervical Vertebrae Maturation (CVM) Stages Using the Tripod NetworkabstractWe present a novel deep learning method for fully automated detection and classification of the Cervical Vertebrae Maturation (CVM) stages. The deep convolutional neural network consists of three parallel networks (TriPodNet) independently trained with different initialization parameters. They also have a built-in set of novel directional filters that highlight the Cervical Vertebrae edges in X-ray images. Outputs of the three parallel networks are combined using a fully connected layer. 1018 cephalometric radiographs were labeled, divided by gender, and classified according to the CVM stages. Resulting images, using different training techniques and patches, were used to train TripodNet together with a set of tunable directional edge enhancers. Data augmentation is implemented to avoid overfitting. TripodNet achieves the state-of-the-art accuracy of 81.18% in female patients and 75.32% in male patients. The proposed TripodNet achieves a higher accuracy in our dataset than the Swin Transformers and the previous network models that we investigated for CVM stage estimation. Salih Atici, Hongyi Pan, Mohammed H. Elnagar, Veerasathpurush Allareddy, Omar Suhaym, Rashid Ansari, A. Enis Çetin |
ICASSP | 2 |
| 2023 | Real-Time Wireless ECG-Derived Respiration Rate Estimation using an Autoencoder with a DCT LayerabstractIn this paper, we present a wireless ECG-derived Respiration Rate (RR) estimation using an autoencoder with a DCT Layer. The wireless wearable system records the ECG data of the subject and the respiration rate is determined from the variations in the baseline level of the ECG data. A straightforward Fourier analysis of the ECG data obtained using the wireless wearable system may lead to incorrect results due to uneven breathing. To improve the estimation precision, we propose a neural network that uses a novel Discrete Cosine Transform (DCT) layer to denoise and decorrelates the data. The DCT layer has trainable weights and soft-thresholds in the transform domain. In our dataset, we improve the Mean Squared Error (MSE) and Mean Absolute Error (MAE) of the Fourier analysis-based approach using our novel neural network with the DCT layer. Hongyi Pan, Xin Zhu 0005, Zhilu Ye, Pai-Yen Chen, A. Enis Çetin |
ICASSP | 1 |
| 2023 | A Hybrid Quantum-Classical Approach based on the Hadamard Transform for the Convolutional LayerabstractIn this paper, we propose a novel Hadamard Transform (HT)-based neural network layer for hybrid quantum-classical computing. It implements the regular convolutional layers in the Hadamard transform domain. The idea is based on the HT convolution theorem which states that the dyadic convolution between two vectors is equivalent to the element-wise multiplication of their HT representation. Computing the HT is simply the application of a Hadamard gate to each qubit individually, so the HT computations of our proposed layer can be implemented on a quantum computer. Compared to the regular Conv2D layer, the proposed HT-perceptron layer is computationally more efficient. Compared to a CNN with the same number of trainable parameters and 99.26% test accuracy, our HT network reaches 99.31% test accuracy with 57.1% MACs reduced in the MNIST dataset; and in our ImageNet-1K experiments, our HT-based ResNet-50 exceeds the accuracy of the baseline ResNet-50 by 0.59% center-crop top-1 accuracy using 11.5% fewer parameters with 12.6% fewer MACs. Hongyi Pan, Xin Zhu 0005, Salih Atici, A. Enis Çetin |
ICML | 1 |
| 2023 | Hybrid Binary Neural Networks: A Tutorial ReviewabstractIn this article, we review neural networks which have neurons with binary operations or networks that use binary transforms such as the Walsh-Hadamard transform (WHT). Neural networks with binary neurons or binary layers can be used in edge applications and/or applications requiring energy-efficient decision-making. WHT-based network is as accurate as the regular neural networks in the CIFAR-10 and the Tiny ImageNet image databases. A. Enis Çetin, Hongyi Pan |
VTS | 2 |
| 2022 | Multiplication-Avoiding Variant of Power Iteration with ApplicationsabstractPower iteration is a fundamental algorithm in data analysis. It extracts the eigenvector corresponding to the largest eigenvalue of a given matrix. Applications include ranking algorithms, principal component analysis (PCA), among many others. Certain use cases may benefit from alternate, non-linear power methods with low complexity. In this paper, we introduce multiplication-avoiding power iteration (MAPI). MAPI replaces the standard ℓ2inner products that appear at the regular power iteration (RPI) with multiplication-free vector products, which are Mercer-type kernels that induce the ℓ1norm. For an n × n matrix, MAPI requires n multiplications, while RPI needs n2multiplications per iteration. Therefore, MAPI provides a significant reduction of the number of multiplication operations, which are known to be costly in terms of energy consumption. We provide applications of MAPI to PCA-based image reconstruction as well as to graph-based ranking algorithms. When compared to RPI, MAPI not only typically converges much faster, but also provides superior performance. Hongyi Pan, Diaa Badawi, Runxuan Miao, Erdem Koyuncu, A. Enis Çetin |
ICASSP | 1 |
| 2022 | Block Walsh-Hadamard Transform-based Binary Layers in Deep Neural NetworksabstractConvolution has been the core operation of modern deep neural networks. It is well known that convolutions can be implemented in the Fourier Transform domain. In this article, we propose to use binary block Walsh–Hadamard transform (WHT) instead of the Fourier transform. We use WHT-based binary layers to replace some of the regular convolution layers in deep neural networks. We utilize both one-dimensional (1D) and 2D binary WHTs in this article. In both 1D and 2D layers, we compute the binary WHT of the input feature map and denoise the WHT domain coefficients using a nonlinearity that is obtained by combining soft-thresholding with the tanh function. After denoising, we compute the inverse WHT. We use 1D-WHT to replace the 1 × 1 convolutional layers, and 2D-WHT layers can replace the 3 × 3 convolution layers and Squeeze-and-Excite layers. 2D-WHT layers with trainable weights can be also inserted before the Global Average Pooling layers to assist the dense layers. In this way, we can reduce the number of trainable parameters significantly with a slight decrease in trainable parameters. In this article, we implement the WHT layers into MobileNet-V2, MobileNet-V3-Large, and ResNet to reduce the number of parameters significantly with negligible accuracy loss. Moreover, according to our speed test, the 2D-FWHT layer runs about 24 times as fast as the regular 3 × 3 convolution with 19.51% less RAM usage in an NVIDIA Jetson Nano experiment. Hongyi Pan, Diaa Badawi, A. Enis Çetin |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2022 | Critique of "MemXCT: Memory-Centric X-Ray CT Reconstruction With Massive Parallelization" by SCC Team From University of California San DiegoabstractIn this article, we describe our efforts to reproduce results reported in the SC19 article by Hidayetoğluet al., titled“MemXCT: Memory-Centric X-ray CT Reconstruction with Massive Parallelization”.MemXCT's single-device performance, parallelized via OpenMP and MPI, was characterized using AMD Zen2 CPU cores and NVIDIA V100 GPU devices running on the Microsoft Azure cloud. We were able to reproduce most of the results, and exceed the performance of larger inputs, on an AMD EPYC HBv2 cluster. We were also able to reproduce the strong scaling trends for optimized CPU and GPU versions. Slight variations in performance of the CPU version were observed due to differences in the underlying hardware, input size, and number of available nodes. Digital artifacts from these experiments are available at: 10.5281/zenodo.5598108 Maximilian Apodaca, Arunav Gupta, Zihao Kong, Hongyi Pan, Mary P. Thomas, Martin Kandes, Mahidhar Tatineni, Lewis Carroll |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | Fourier Domain Pruning of MobileNet-V2 with Application to Video Based Wildfire DetectionabstractIn this paper, we propose a deep convolutional neural network for camera based wildfire detection. We train the neural network via transfer learning and use window based analysis strategy to increase the fire detection rate. To achieve computational efficiency, we calculate frequency response of the kernels in convolutional and dense layers and eliminate those filters with low energy impulse response. Moreover, to reduce the storage for edge devices, we compare the convolutional kernels in Fourier domain and discard similar filters using the cosine similarity measure in the frequency domain. We test the performance of the neural network with a variety of wildfire video clips and prune system performs as good as the regular network in daytime wild fire detection, and it also works well on some night wild fire video clips. Hongyi Pan, Diaa Badawi, A. Enis Çetin |
ICPR | 1 |