Ghazanfar Ali

dblp:194/7300 · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 RIDGE: Rule-Infused Deep Learning for Realistic Co-Speech Gesture Generation
abstract
ABSTRACT Co‐speech gestures are essential for natural human communication, yet existing synthesis methods fall short in delivering semantically aligned and contextually appropriate motions. In this paper, we present RIDGE, a hybrid system that combines rule‐based and deep learning approaches to generate realistic gestures for virtual avatars and human‐computer interaction. RIDGE employs a high‐fidelity rule base, generated from motion capture data with the assistance of large language models, to select reliable gesture mappings. When a high‐confidence match is not available, a contrastively trained deep learning model steps in to produce semantically appropriate gestures. Evaluated using a novel Gesture Cluster Affinity (GCA) metric, our system outperforms existing baselines, achieving a GCA score of 0.73 compared to a rule‐based baseline of 0.6 and an end‐to‐end: 0.52, while the ground truth score was 0.90. Detailed analyses of system architecture, data preprocessing, and evaluation methodologies demonstrate RIDGE's potential to enhance gesture synthesis. Project Url: https://www.mrlab.co.kr/research/ridge .
Ghazanfar Ali, Hwang Youn Kim, Jae-In Hwang
Comput. Animat. Virtual Worlds1
2025 A Retrieval-Augmented Generation System for Accurate and Contextual Historical Analysis: AI-Agent for the Annals of the Joseon Dynasty
abstract
ABSTRACT In this article, we propose an AI‐agent that integrates a large language model (LLM) with a retrieval‐augmented generation (RAG) system to deliver reliable historical information from the Annals of the Joseon Dynasty through both objective facts and contextual analysis, achieving significant performance improvements over existing models. For an AI‐agent using the Annals of the Joseon Dynasty to deliver reliable historical information, clear source citations and systematic analysis are essential. The Annals, an official record spanning 472 years (1392–1897), offer a dense, chronological account of daily events and state administration that shaped Korea's cultural, political, and social foundations. We propose integrating a LLM with a RAG system to generate highly accurate responses based on this extensive dataset. This approach provides both objective information about historical figures and events from specific periods and subjective contextual analysis of the era, helping users gain a broader understanding. Our experiments demonstrate improvements of approximately 23 to 50 points on a 100‐point scale compared with the GPT‐4o and OpenAI AI‐Assistant v2 models.
Jeong Ha Lee, Ghazanfar Ali, Jae-In Hwang
Comput. Animat. Virtual Worlds2
2025 ASAP for multi-outputs: auto-generating storyboard and pre-visualization with virtual actors based on screenplay
Hanseob Kim, Ghazanfar Ali, Hwang Youn Kim, Hyemin Shin, Gerard Jounghyun Kim, Jae-In Hwang
Multim. Tools Appl.2
2024 Enhancing doctor-patient communication in surgical explanations: Designing effective facial expressions and gestures for animated physician characters
abstract
Abstract Paying close attention to facial expressions, gestures, and communication techniques is essential when creating animated physician characters that are realistic and captivating when describing surgical procedures. This paper emphasizes the integration of appropriate emotions, co‐speech gestures when medical experts explain the medical procedure, and designing animated characters. We can achieve healthy doctor‐patient relationships and improvement of patients' understanding by depicting these components truthfully. We suggest two critical approaches to developing virtual medical experts by incorporating these elements. First, doctors can generate the contents of the surgical procedure with a virtual doctor. Second, patients can listen to the surgical procedure described by the virtual doctor and ask if they have any questions. Our system helps patients by considering their psychology and adding medical professionals' opinions. These improvements ensure the animated virtual agent is comforting, reassuring, and emotionally supportive. Through a user study, we evaluated our hypothesis and gained insight into improvements.
Hwang Youn Kim, Ghazanfar Ali, Jae-In Hwang
Comput. Animat. Virtual Worlds2
2023 Workload Failure Prediction for Data Centers
abstract
Failed workloads that consumed significant computational resources in time and space affect the efficiency of HPC data centers significantly and thus limit the amount of scientific work that can be achieved. While the computational power has increased significantly over the years, detection and prediction of workload failures have lagged far behind and will become increasingly critical as the system scale and complexity further increase. In this study, we analyze workload traces collected from a production cluster and train machine learning models on a large amount of data sets to predict workload failures. Our prediction models consist of a queue-time model that estimates the probability of workload failures before execution and a runtime model that predicts failures at runtime. Evaluation results show that the queue-time model and runtime model can predict workload failures with a maximum precision score of 90.61% and 97.75%, respectively. By integrating the runtime model with the job scheduler, it helps reduce CPU time, and memory usage by up to 16.7% and 14.53%, respectively.
Jie Li 0057, Ghazanfar Ali, Tommy Dang, Alan Sill, Yong Chen 0001
CLOUD3
2023 Performance-Aware Energy-Efficient GPU Frequency Selection using DNN-based Models
abstract
Energy efficiency will be important in future accelerator-based HPC systems for sustainability and to improve overall performance. This study proposes a deep neural network (DNN)-based learning model for execution time and power consumption of workloads across GPUs DVFS design space. Micro-architectural data obtained by running SPEC-ACCEL, DGEMM, and STREAM benchmarks are used for model training. These features are consistent for a workload unaffected by frequency and input size reducing the data required significantly. For real-world applications - LAMMPS, NAMD, GROMACS, LSTM, BERT, and ResNet50 power and time models show 89% – 98% accuracy on NVIDIA Ampere. Multi-objective functions help select optimal frequencies that lower power and minimize performance impact showing maximum energy savings of 27% at a performance loss of 1.8%. The same models trained on Ampere showed an accuracy of greater than 93% on an NVIDIA Volta, thereby demonstrating model portability across architectures.
Ghazanfar Ali, Mert Side, Sridutt Bhalachandra, Nicholas J. Wright, Yong Chen 0001
ICPP1
2023 An automated and portable method for selecting an optimal GPU frequency
Ghazanfar Ali, Mert Side, Sridutt Bhalachandra, Nicholas J. Wright, Yong Chen 0001
Future Gener. Comput. Syst.1
2022 Automating CPU Dynamic Thermal Control for High Performance Computing
abstract
In a production high-performance computing (HPC) data center, numerous factors, including workload compute in-tensity, cooling infrastructure failure, and the use of economized cooling can substantially increase the CPU temperature. CPU thermal design-related studies have shown that slight variances in the operational temperature can significantly impact the lifetime, durability, and performance of a CPU. Therefore, it is critical to monitor and control the operating temperature of the CPU. In this study, we design an automated and continuous CPU thermal monitoring and control methodology to maintain and control a healthy CPU thermal state. This research utilizes the Redfish protocol to monitor the CPU temperature and dynamic voltage frequency scaling to control the temperature. We developed a reference implementation and evaluated our methodology using a cluster of 150 Raspberry Pi3 nodes. We performed extensive CPU thermal analyses in different scenarios. We analyzed how quickly a CPU can attain the maximum temperature under 100% load at room temperature. Based on our experiments, the temperature of a CPU with 100% load can increase to ~72°C (161.6°F) and ~86°C (186.8°F) with the lowest and highest CPU frequency configurations, respectively. We analyzed the impact of applying thermal control at eight temperature configurations on the thermal and frequency scaling behavior of a CPU. We observed that applying thermal control at lower temperature configurations (e.g., 70°C (158°F)) is a better configuration for healing an overheated CPU. As a result of the proposed model, the CPU operating at normal temperature will consume comparatively less energy, deliver higher performance, and augment its durability.
Ghazanfar Ali, Lowell Wofford, Yong Chen 0001
CCGRID1
2020 On-chip EOL Prognostics Using Data-Fusion of Embedded Instruments for Dependable MP-SoCs
abstract
The usage of embedded instruments (EIs) in a processor core to address dependability challenges of modern-day Multi-Processor System-on-Chip (MP-SoC) has been studied in literature extensively. Data from these EIs can be used in applications like end-of-lifetime (EOL) predictions. However, inaccuracies present in the data from these EIs, due to their self-aging and resolution limitations during digitization, can lead to an inaccurate EOL assessment. In this paper, it is presented that in the presence of such inaccuracies from EIs as well as correlation between EIs, principal component analysis (PCA) based data-fusion approach for determining the EOL of selected critical paths provided overall better EOL predictions as compared to EOL predictions based on standalone EIs. Verification was performed with a commercial software-based EOL predictor tool ARULE running on a personal computer. Moreover, the presented results on the computational requirements for the presented data-fusion approach showed little overhead in terms of memory, execution time and energy requirements.
Ghazanfar Ali, Leila Bagheriye, Hans A. R. Manhaeve, Hans G. Kerkhoff
ATS1
2020 MonSTer: An Out-of-the-Box Monitoring Tool for High Performance Computing Systems
abstract
Understanding the status of high-performance computing platforms and correlating applications to resource usage provide insight into the interactions among platform components. A lot of efforts have been devoted into developing monitoring solutions; however, a large-scale HPC system usually requires a combination of methods/tools to successfully monitor all metrics, which will lead to a huge effort in configuration and monitoring. Besides, monitoring tools are often left behind in the procurement of large-scale HPC systems. These challenges have motivated the development of a next-generation out-of-the-box monitoring tool that can be easily deployed without losing informative metrics. In this work, we introduce MonSTer, an “out-of-the-box” monitoring tool for high-performance computing platforms. MonSTer uses the evolving specification Redfish to retrieve sensor data from Baseboard Management Controller (BMC), and resource management tools such as Univa Grid Engine (UGE) or Slurm to obtain application information and resource usage data. Additionally, it also uses a time-series database (e.g. InfluxDB) for data storage. MonSTer correlates applications to resource usage and reveals insightful knowledge without having additional overhead on the application and computing nodes. This paper presents the design and implementation of MonSTer, as well as experiences gained through real-world deployment on the 467-node Quanah cluster at Texas Tech University's High Performance Computing Center (HPCC) over the past year. In this work, we introduce MonSTer, an “out-of-the-box” monitoring tool for high-performance computing platforms. MonSTer uses the evolving specification Redfish to retrieve sensor data from Baseboard Management Controller (BMC), and resource management tools such as Univa Grid Engine (UGE) or Slurm to obtain application information and resource usage data. Additionally, it also uses a time-series database (e.g. InfluxDB) for data storage. MonSTer correlates applications to resource usage and reveals insightful knowledge without having additional overhead on the application and computing nodes. This paper presents the design and implementation of MonSTer, as well as experiences gained through real-world deployment on the 467-node Quanah cluster at Texas Tech University's High Performance Computing Center (HPCC) over the past year.
Jie Li 0057, Ghazanfar Ali, Ngan V. T. Nguyen, Jon R. Hass, Alan Sill, Tommy Dang
CLUSTER2
2020 Life-Time Prognostics of Dependable VLSI-SoCs using Machine-learning
abstract
Recently, the usage of on-chip embedded instruments (EIs) to ensure dependable safety-critical systems is becoming inevitable. These EIs can help to provide self-awareness, and their feedback can be used in different applications, e.g. end-of-lifetime (EOL) predictions. However, inaccuracies present in data from these EIs, due to their resolution limitations, self-aging and quantization errors during digitization, can lead to an inaccurate EOL assessment. To address this challenge, a machine learning-based system-level approach for determining the EOL of a many-processor system-on-chip (MPSoC) is discussed. It is based on the synchronous data capture of different IJTAG compatible EIs. To this end, two different data fusion techniques have been used for enhancing the accuracy of lifetime prognostics of multiple EIs; use is made of Independent Component Analysis (ICA) and the auto-encoder (AE). Different combinations of fused EIs (based on ICA and AE) along with standalone EIs for four different critical paths (CPs) have been investigated. For lifetime prediction based on different EIs/fused EIs, a data-driven degradation model was derived, and nonlinear regression has been employed for parameter estimation. Results show that data fusion of different EIs helps in obtaining better estimation of the EOL as compared to using a standalone EI.
Leila Bagheriye, Ghazanfar Ali, Hans G. Kerkhoff
IOLTS2
2020 On-Chip Embedded Instruments Data Fusion and Life-Time Prognostics of Dependable VLSI-SoCs using Machine-Learning
abstract
Nowadays, a rapid introduction of very complex nanometer Many-Processor Systems-on-Chip in safety-critical applications is taking place. Unfortunately, it pairs with an unacceptable decrease in dependability of these complex nanosystems if no additional countermeasures are taken. To address this challenge, a promising approach is presented in this paper that uses a set of IJTAG compatible embedded instruments (EIs), in and around a processor cores to monitor their present health status. Data from these EIs is collected and fused for lifetime prognostics and hence dependability. For the EIs data fusion, use is made of principal component analysis (PCA) technique. For lifetime prediction based on different EIs, power-law degradation model was used.
Ghazanfar Ali, Leila Bagheriye, Hans G. Kerkhoff
ISCAS1
2020 Automatic text-to-gesture rule generation for embodied conversational agents
abstract
Abstract Interactions with embodied conversational agents can be enhanced using human‐like co‐speech gestures. Traditionally, rule‐based co‐speech gesture mapping has been utilized for this purpose. However, the creation of this mapping is laborious and often requires human experts. Moreover, human‐created mapping tends to be limited, therefore prone to generate repeated gestures. In this article, we present an approach to automate the generation of rule‐based co‐speech gesture mapping from publicly available large video data set without the intervention of human experts. At run‐time, word embedding is utilized for rule searching to get the semantic‐aware, meaningful, and accurate rule. The evaluation indicated that our method achieved comparable performance with the manual map generated by human experts, with a more variety of gestures activated. Moreover, synergy effects were observed in users' perception of generated co‐speech gestures when combined with the manual map.
Ghazanfar Ali, Myungho Lee, Jae-In Hwang
Comput. Animat. Virtual Worlds1
2019 Design of Seamless Multi-modal Interaction Framework for Intelligent Virtual Agents in Wearable Mixed Reality Environment
abstract
In this paper, we present the design of a multimodal interaction framework for intelligent virtual agents in wearable mixed reality environments, especially for interactive applications at museums, botanical gardens, and similar places. These places need engaging and no-repetitive digital content delivery to maximize user involvement. An intelligent virtual agent is a promising mode for both purposes. Premises of framework is wearable mixed reality provided by MR devices supporting spatial mapping. We envisioned a seamless interaction framework by integrating potential features of spatial mapping, virtual character animations, speech recognition, gazing, domain-specific chatbot and object recognition to enhance virtual experiences and communication between users and virtual agents. By applying a modular approach and deploying computationally intensive modules on cloud-platform, we achieved a seamless virtual experience in a device with limited resources. Human-like gaze and speech interaction with a virtual agent made it more interactive. Automated mapping of body animations with the content of a speech made it more engaging. In our tests, the virtual agents responded within 2-4 seconds after the user query. The strength of the framework is flexibility and adaptability. It can be adapted to any wearable MR device supporting spatial mapping.
Ghazanfar Ali, Hong-Quan Le, Junho Kim 0001, Seung-won Hwang, Jae-In Hwang
CASA1
2019 IJTAG Compatible Timing Monitor with Robust Self-Calibration for Environmental and Aging Variation
abstract
The deployment of embedded instruments (EI) for online monitoring of a cyber-physical system-on-chip (CPSoC) for safety-critical applications has started getting more attention in recent years. Among these different types of EIs, the timing embedded instruments (to observe timing violations) are widely adopted to ensure a dependable operation during its operational lifetime. However, infield temperature variations together with self-aging of the timing EI affects its output, making it less reliable. This paper presents an IJTAG compatible, in-situ slack-delay timing embedded instrument to monitor critical paths in a CPSoC. The main contribution of the proposed design is the infield, online calibration mechanism. It enables the EI to self-calibrate for relatively fast-changing environmental variations like temperature, as well as slow changing variations like aging. The design has been implemented using a TSMC 40nm LP standard cell library. Simulation results show an average resolution of 13ps with a monitoring window of 13ps-416ps, which is sufficient to monitor the CPSoC for small timing margins. To further demonstrate the IJTAG compatibility, an FPGA implementation of the proposed design is also presented.
Ghazanfar Ali, Jerrin Pathrose, Hans G. Kerkhoff
ETS1
2017 An automotive MP-SoC featuring an advanced embedded instrument infrastructure for high dependability
abstract
In safety-critical systems, many-processor Systems-on-Chip are being increasingly employed. An example is an imminent collision detection System-on-Chip for cars. Such a system requires zero downtime and a very high reliability despite aging issues under harsh environmental conditions. By monitoring the health status of processor cores and other IPs, and taking appropriate counteractions if required, we accomplished this goal via IJTAG compatible embedded instruments. This paper shows the design of the required IJTAG network, and a number of new IJTAG-compatible embedded instruments like slack-delay, power-supply current IDDT and Intermittent Resistive Fault monitors. In addition, we discuss their numbers and optimal locations in a processor core and provide a PDL description for one of our embedded instruments. In the case of for instance a four-processor implementation, requiring only two for actual data processing, the lifetime can increase by a factor of roughly three.
Hans G. Kerkhoff, Ghazanfar Ali, Hassan Ebrahimi, Ahmed Ibrahim 0001
ITC-Asia2
2017 Applying IJTAG-compatible embedded instruments for lifetime enhancement of analog front-ends of cyber-physical systems
abstract
In safety-critical cyber-physical systems, analog front-ends combined with many-processors are being increasingly employed. An example is an imminent collision detection chip for cars. Such a complex system requires zero downtime and a very high dependability despite aging issues under harsh environmental conditions. By on-line monitoring the health status of the processor cores and taking appropriate counteractions if required, we have accomplished this goal in the past via IJTAG compatible embedded instruments and appropriate embedded software. This paper extends this approach to the analog / mixed-signal frontends of these systems, thereby creating a new uniform approach in design & test methodology, as well as a streamlined fault management. An IJTAG-compatible voltage monitor is introduced, for measuring aging-generated offset in OpAmps and SAR ADCs, as well as a delay-monitoring embedded instrument for detecting timing issues in ADCs. In addition, two-stage counter measures, like digitized recalibration and subsequent replacement, are presented to increase the lifetime by factors of the analog front-end of Cyber-Physical Systems-on-Chips.
Hans G. Kerkhoff, Ghazanfar Ali, Jinbo Wan, Ahmed Ibrahim 0001, Jerrin Pathrose
VLSI-SoC2