Siqi Deng

dblp:189/4604 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TernaryGNNs: A High-Throughput, Area-Efficient Ternary Weight GNNs Inference Framework on CPU-FPGA Heterogeneous Platform
abstract
Graph Neural Networks (GNNs) have achieved remarkable success in recent years due to their powerful ability to model non-Euclidean data structures and complex relationships. However, as graph sizes and model complexities continue to grow, the efficient deployment of GNNs across various hardware platforms presents significant challenges. The two main computation modes in GNN inference exhibit distinct characteristics in sparsity and computational density, which results in imbalanced workload distribution and inefficient utilization of compute and memory bandwidth. Furthermore, subgraph partitioning strategies, commonly adopted to address on-chip resource constraints, frequently introduce write-back overhead and port conflicts, thereby limiting overall system throughput. To address these challenges, we propose TernaryGNNs, a high-throughput, area-efficient ternary weight GNN inference framework targeting CPU-FPGA heterogeneous platforms. First, we introduce a precision-preserving ternary quantization method that maintains model accuracy with 1.016% degradation, while achieving an average weight sparsity of 81.1% and a parameter reduction rate of 94.71%. Next, by exploiting sparsity, we reformulate GNN inference as Sparse Matrix–Dense Matrix Multiplication (SpMM) or Sparse General Matrix–Matrix Multiplication (SpGEMM) computation and propose a unified sparse-optimized processor architecture. Finally, we present a comprehensive software–hardware co-design framework that ensures adaptability to the evolving landscape of diverse GNN model architectures. Our framework supports nine mainstream GNN models. Compared to the State-of-the-Art (SOTA) general-purpose GNN processor GraphOPU, TernaryGNNs achieves an average \(2.79\times\) hardware performance improvement, \(1.70\times\) end-to-end performance improvement, and \(2.83\times\) area-efficiency improvement. Compared to the leading overlay accelerator FP-GNN, it delivers \(6.73\times\) hardware performance improvement on average.
Zhuorong Liang, Siqi Deng
ACM Trans. Reconfigurable Technol. Syst.2
2024 FairRAG: Fair Human Generation via Fair Retrieval Augmentation
abstract
Existing text-to-image generative models reflect or even amplify societal biases ingrained in their training data. This is especially concerning for human image generation where models are biased against certain demographic groups. Existing attempts to rectify this issue are hindered by the inherent limitations of the pre-trained models and fail to substantially improve demographic diversity. In this work, we introduce Fair Retrieval Augmented Generation (FairRAG), a novel framework that conditions pre-trained generative models on reference images retrieved from an external image database to improve fairness in human generation. FairRAG enables conditioning through a lightweight linear module that projects reference images into the textual space. To enhance fairness, FairRAG applies simple-yet-effective debiasing strategies, providing images from diverse demographic groups during the generative process. Extensive experiments demonstrate that FairRAG outper-forms existing methods in terms of demographic diversity, image-text alignment and image fidelity while incurring minimal computational overhead during inference.
Robik Shrestha, Qiuyu Chen, Yusheng Xie, Siqi Deng
CVPR6
2024 Enhancing Consistent Federated Learning Objectives Through Uniform Feature Distributions
abstract
Federated Learning is a distributed paradigm that facilitates collaborative training of deep models among multiple parties without exchanging raw data. However, the common non-independent and identically distributed (Non-IID) data distribution among clients introduces discrepancies between local training objectives and the global goal. This misalignment results in a slow convergence of global model and a decrease in generalization performance. We propose a method to enhance consistent federated learning objectives through uniform and consistent feature distributions (FedUF). FedUF effectively captures the feature variation space rich in semantic information and integrates implicit semantic data augmentation and logits adjustment to establish a uniform and consistent global feature distribution. Through lightweight yet innovative adjustments to the client-side objective functions, we formulate a globally consistent objective function. The mutual promotion between globally consistent feature distribution and the objective function significantly alleviates the impact of Non-IID data, greatly enhancing the overall performance of Federated Learning. Moreover, extensive experiments conducted on Cifar10 and Cifar100 datasets convincingly validate the effectiveness of the proposed FedUF.
Siqi Deng, Liu Yang 0010
ICME1
2023 Prototype Contrastive Learning for Personalized Federated Learning
Siqi Deng
ICANN (3)1
2023 Harnessing Unrecognizable Faces for Improving Face Recognition
abstract
The common implementation of face recognition systems as a cascade of a detection stage and a recognition or verification stage can cause problems beyond failures of the detector. When the detector succeeds, it can detect faces that cannot be recognized, no matter how capable the recognition system is. Recognizability, a latent variable, should therefore be factored into the design and implementation of face recognition systems. We propose a measure of recognizability of a face image that leverages a key empirical observation: An embedding of face images, implemented by a deep neural network trained using mostly recognizable identities, induces a partition of the hypersphere whereby unrecognizable identities cluster together. This occurs regardless of the phenomenon that causes a face to be unrecognizable, be it optical or motion blur, partial occlusion, spatial quantization, or poor illumination. Therefore, we use the distance from such an "unrecognizable identity" as a measure of recognizability, and incorporate it into the design of the overall system. We show that accounting for recognizability reduces the error rate of single-image face recognition by 58% at FAR=1e-5 on the IJB-C Covariate Verification benchmark, and reduces the verification error rate by 24% at FAR=1e-5 in set-based recognition on the IJB-C benchmark.
Siqi Deng, Yuanjun Xiong, Wei Xia 0009, Stefano Soatto
WACV1
2022 Shasta Log Aggregation, Monitoring and Alerting in HPC Environments with Grafana Loki and ServiceNow
abstract
The ongoing deployments of post-petascale computing systems has led to the proliferation in the hybrid computing models and orchestration of many complex services leading to the growth in operational challenges. It becomes increasingly important to deploy new integrated and comprehensive event management and monitoring solutions to collect computing infrastructures health logs and metrics data, to correlate and analyze the gathered data for reducing response time and downtime in face of computational center critical issues caused due to the physical and the digital threats. To address the above mentioned challenges, in this paper we present an automated log aggregation, monitoring and alerting framework that leverages the Operations Monitoring and Networking Infrastructure (OMNI) when used with Hewlett Packard Enterprise (HPE) Shasta, Grafana Loki and ServiceNow platforms for enabling a comprehensive proactive event response management and monitoring. Moreover, herein we also present two case studies involving automated remediation workflows employing the proposed framework at the National Energy Research Scientific Computing Center (NERSC) at Lawrence Berkeley National Laboratory (LBNL) using the Perlmutter computational system for real-time collection, aggregation, correlation, analysis, management and visualization of the system health metrics and logs in a single pane of glass for enhancing proactive monitoring and operational efficiency.
Elizabeth Bautista, Nitin Sukhija, Siqi Deng
CLUSTER3
2022 Multi-Dimensional, Nuanced and Subjective - Measuring the Perception of Facial Expressions
abstract
Humans can perceive multiple expressions, each one with varying intensity, in the picture of a face. We propose a methodology for collecting and modeling multidimensional modulated expression annotations from human annotators. Our data reveals that the perception of some expressions can be quite different across observers; thus, our model is designed to represent ambiguity alongside intensity. An empirical exploration of how many dimensions are necessary to capture the perception of facial expression suggests six principal expression dimensions are sufficient. Using our method, we collected multidimensional modulated expression annotations for 1,000 images culled from the popular ExpW in-the-wild dataset. As a proof of principle of our improved measurement technique, we used these annotations to benchmark four public domain algorithms for automated facial expression prediction.
De'Aira Bryant, Siqi Deng, Nashlie Sephus, Wei Xia 0009, Pietro Perona
CVPR2
2022 A fast geometric prediction merge mode decision algorithm based on CU gradient for VVC
abstract
To address the problem of large computational redundancy in the geometric prediction merge mode with motion vector refinement (GPM with MMVD), a new decision algorithm based on CU gradient is proposed. By comparing the mean value of the gradient in four directions to determine whether GPM can be terminated early. The advance decision of GPM partition mode can be determined by the calculated gradient direction of CU. Meanwhile, the proposed algorithm saves coding time by reducing the number of candidate modes during the mode selection process.
Siqi Deng, Zhi Liu 0008
DCC2
2022 Unsupervised and Semi-supervised Bias Benchmarking in Face Recognition
Alexandra Chouldechova, Siqi Deng, Wei Xia 0009, Pietro Perona
ECCV (13)2
2022 The Caltech Fish Counting Dataset: A Benchmark for Multiple-Object Tracking and Counting
Justin Kay, Peter Kulits, Suzanne Stathatos, Siqi Deng, Erik Young, Sara Beery, Grant Van Horn, Pietro Perona
ECCV (8)4
2022 Two-Stream Communication-Efficient Federated Pruning Network
Shiqiao Gu, Siqi Deng, Zhengyi Xu
PRICAI (3)3
2021 Positive-Congruent Training: Towards Regression-Free Model Updates
abstract
Reducing inconsistencies in the behavior of different versions of an AI system can be as important in practice as reducing its overall error. In image classification, sample-wise inconsistencies appear as "negative flips": A new model incorrectly predicts the output for a test sample that was correctly classified by the old (reference) model. Positive-congruent (PC) training aims at reducing error rate while at the same time reducing negative flips, thus maximizing congruency with the reference model only on positive predictions, unlike model distillation. We propose a simple approach for PC training, Focal Distillation, which enforces congruence with the reference model by giving more weights to samples that were correctly classified. We also found that, if the reference model itself can be chosen as an ensemble of multiple deep neural networks, negative flips can be further reduced without affecting the new model’s accuracy.
Sijie Yan, Yuanjun Xiong, Kaustav Kundu, Shuo Yang 0003, Siqi Deng, Wei Xia 0009, Stefano Soatto
CVPR5
2020 Event Management and Monitoring Framework for HPC Environments using ServiceNow and Prometheus
abstract
The challenge of monitoring and event response management of a high performance computing facility grows significantly as the facilities employs and orchestrates more complex and heterogeneous systems and infrastructure. As the computational components encompassing the HPC facility system increases, the computational staff experiences rise in alert fatigue due to the false alarms and noise related to the similar events generated by monitoring tools. The National Energy Research Scientific Computing Center (NERSC) at the Lawrence Berkeley National Laboratory (LBNL) has begun to address the issues of duplication of alerts and alert remediation. However, more automation and integration is needed for collecting, aggregating, correlating, analyzing, managing and visualizing the scale of events that will be generated by the emergent hybrid computing infrastructures. In this paper, we present an event management and monitoring framework that addresses the operational needs of the future pre-exascale systems at the Lawrence Berkeley National Laboratory's National Energy Research Scientific Computing Center (NERSC). The framework integrates the Operations Monitoring and Notification Infrastructure (OMNI) at NERSC with the Prometheus, Grafana and ServiceNow platforms to help identify, diagnose, and resolve incidents in real-time, as well as conduct more thorough post-incident reviews enabled by the intuitive dashboards that provides a single pane of glass console for an efficient operations management and real-time proactive monitoring.
Nitin Sukhija, Elizabeth Bautista, Owen James, Daniel Gens, Siqi Deng, Yulok Lam, Tony Quan, Basil Lalli
MEDES5
2016 Online variational Bayesian Support Vector Regression
abstract
Traditional Support Vector Regression (SVR) solvers require user pre-specified penalty (regularization) parameter as input and typically model the training data with maximum a posterior (MAP) principle. The resultant point estimates can be affected seriously by inappropriate regularization, outliers and noise, especially when training online. In this paper, we address the aforementioned problems by developing a Bayesian SVR model with the pseudo-likelihood and data augmentation idea. Then we perform variational posterior inference in an augmented variable space and the approximate posterior of model weights, rather than point estimates as in traditional SVR, are used to make robust predictions. Besides, once the approximate posterior is obtained from a given set of data, we can regard it as model prior when dealing with new arrival data, which leads to a natural way to extend our batch model to the online scenario. Experiments on several benchmark regression problems as well as a real vehicle accident rate prediction task show that our models have superior performance while inferring penalty parameter automatically.
Siqi Deng, Kan Gao, Changying Du, Wenjing Ma, Guoping Long, Yucheng Li 0002
IJCNN1