EDBT 2026 Demo / reviewers in the wild / expert
Chen Liu 0001
dblp:10/2639-1
· DBLP profile ↗
31ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0003-1558-6836ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Security and privacy · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Federated Learning Approach Towards Cross-Architecture Ransomware Detection Using Low-Level Hardware Information
Chutitep Woralert, Chen Liu 0001 |
COMPSAC | 2 |
| 2026 | Extending HARD-Lite: Toward Scalable Hardware-Assisted Malware Detection
Chutitep Woralert, Chen Liu 0001 |
COMPSAC | 2 |
| 2025 | Scalable Malware Detection Framework Using Performance Counters and Gradient BoostingabstractFacing the challenge of increasingly sophisticated malware, it is imperative to develop an effective and adaptable malware detection framework. Hardware-level information has been shown to be very effective in detecting malware in the system through dynamic behavioral analysis at runtime. However, previous approaches suffer from the overhead of complex neural network models, as well as difficulty in scaling toward new attacks. In this work, we introduce a novel approach that leverages the Light Gradient-Boosting Machine (LightGBM) model, known for its efficiency and support for fast transfer learning method, to create a scalable and highly accurate malware detection system. Our framework achieves an exceptional detection accuracy of more than 99.90% for multiclass classification. Through transfer learning, the model can quickly adapt to new malware data and environments, significantly reducing the time and resources needed for retraining. Our results highlight the potential of using the LightGBM model to improve the malware detection framework. Chutitep Woralert, Chen Liu 0001, Zander Blasingame |
ASAP | 2 |
| 2025 | PaFi-GS: Gaussian Splatting via Propagation-Aware Filtering for Urban Street View Rendering
Ying Long, Zhiliu Yang, Chen Liu 0001 |
ICANN (2) | 5 |
| 2025 | Greed is Good: A Unifying Perspective on Guided GenerationabstractTraining-free guided generation is a widely used and powerful technique that allows the end user to exert further control over the generative process of flow/diffusion models.
Generally speaking, two families of techniques have emerged for solving this problem for *gradient-based guidance*: namely, *posterior guidance* (*i.e.*, guidance via projecting the current sample to the target distribution via the target prediction model) and *end-to-end guidance* (*i.e.*, guidance by performing backpropagation throughout the entire ODE solve).
In this work, we show that these two seemingly separate families can actually be *unified* by looking at posterior guidance as a *greedy strategy* of *end-to-end guidance*.
We explore the theoretical connections between these two families and provide an in-depth theoretical of these two techniques relative to the *continuous ideal gradients*.
Motivated by this analysis we then show a method for *interpolating* between these two families enabling a trade-off between compute and accuracy of the guidance gradients.
We then validate this work on several inverse image problems and property-guided molecular generation. Zander Blasingame, Chen Liu 0001 |
NeurIPS | 2 |
| 2024 | Greedy-DiM: Greedy Algorithms for Unreasonably Effective Face MorphsabstractMorphing attacks are an emerging threat to state-of-the-art Face Recognition (FR) systems, which aim to create a single image that contains the biometric information of multiple identities. Diffusion Morphs (DiM) are a recently proposed morphing attack that has achieved state-of-the-art performance for representation-based morphing attacks. However, none of the existing research on DiMs have leveraged the iterative nature of DiMs and left the DiM model as a black box, treating it no differently than one would a Generative Adversarial Network (GAN) or Varational AutoEncoder (VAE). We propose a greedy strategy on the iterative sampling process of DiM models which searches for an optimal step guided by an identity-based heuristic function. We compare our proposed algorithm against ten other state-of-the-art morphing algorithms using the open-source SYN-MAD 2022 competition dataset. We find that our proposed algorithm is unreasonably effective, fooling all of the tested FR systems with an Mated Morph Presentation Match Rate (MMPMR) of 100%, outperforming all other morphing algorithms compared. Zander Blasingame, Chen Liu 0001 |
IJCB | 2 |
| 2024 | The Impact of Print-Scanning in Heterogeneous Morph Evaluation ScenariosabstractFace morphing attacks pose an increasing threat to face recognition (FR) systems. A morphed photo contains biometric information from two different subjects to take advantage of vulnerabilities in FRs. These systems are particularly susceptible to attacks when the morphs are subjected to print-scanning to mask the artifacts generated during the morphing process. We investigate the impact of print-scanning on morphing attack detection through a series of evaluations on heterogeneous morphing attack scenarios. Our experiments show that we can increase the Mated Morph Presentation Match Rate (MMPMR) by up to 8.48%. Furthermore, when a Single-image Morphing Attack Detection (S-MAD) algorithm is not trained to detect print-scanned morphs the Morphing Attack Classification Error Rate (MACER) can increase by up to 96.12%, indicating significant vulnerability. Richard E. Neddo, Zander Blasingame, Chen Liu 0001 |
IJCB | 3 |
| 2024 | AdjointDEIS: Efficient Gradients for Diffusion ModelsabstractThe optimization of the latents and parameters of diffusion models with respect to some differentiable metric defined on the output of the model is a challenging and complex problem. The sampling for diffusion models is done by solving either the *probability flow* ODE or diffusion SDE wherein a neural network approximates the score function allowing a numerical ODE/SDE solver to be used. However, naive backpropagation techniques are memory intensive, requiring the storage of all intermediate states, and face additional complexity in handling the injected noise from the diffusion term of the diffusion SDE. We propose a novel family of bespoke ODE solvers to the continuous adjoint equations for diffusion models, which we call *AdjointDEIS*. We exploit the unique construction of diffusion SDEs to further simplify the formulation of the continuous adjoint equations using *exponential integrators*. Moreover, we provide convergence order guarantees for our bespoke solvers. Significantly, we show that continuous adjoint equations for diffusion SDEs actually simplify to a simple ODE. Lastly, we demonstrate the effectiveness of AdjointDEIS for guided generation with an adversarial attack in the form of the face morphing problem. Our code will be released on our project page [https://zblasingame.github.io/AdjointDEIS/](https://zblasingame.github.io/AdjointDEIS/) Zander Blasingame, Chen Liu 0001 |
NeurIPS | 2 |
| 2023 | HARD-Lite: A Lightweight Hardware Anomaly Realtime Detection Framework Targeting RansomwareabstractRecent years have witnessed a surge in ransomware attacks. Especially, many new variants of ransomware have continued to emerge, employing more advanced techniques to distribute the payload while avoiding detection. This renders the traditional static ransomware detection mechanism ineffective. In this paper, we present our Hardware Anomaly Realtime Detection-Lightweight (HARD-Lite) framework that employs a semi-supervised machine learning method to detect ransomware using low-level hardware information. By using an LSTM network with a weighted majority voting ensemble and exponential moving average, we are able to take into consideration the temporal aspect of hardware-level information formed as time series in order to detect deviation in system behavior, thereby increasing the detection accuracy whilst reducing the number of false positives. Testing against various ransomware families across multiple hardware platforms, HARD-Lite has demonstrated remarkable effectiveness, detecting all cases tested successfully. What’s more, by having a separate machine for the classifier while the user machine is under monitoring, it allows the classifier machine to enforce strict protection and offload the heavy-weight classification work, without impeding the functionality of the user machine. This hierarchical design enables good scalability for the proposed framework. Chutitep Woralert, Chen Liu 0001, Zander Blasingame |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Where's Waldo?: identifying anomalous behavior of data-only attacks using hardware featuresabstractIn recent years, multiple techniques have been proposed to defend computing systems against control-oriented attacks that hijack the control-flow of the victim program. Data-only attacks, on the other hand, are a less common and more subtle type of exploit which are more difficult to detect using traditional mitigation techniques that target control-oriented attacks. In this paper we introduce a novel methodology for the detection of data-only attacks through modeling the execution behavior of an application using low-level hardware information collected as a data series during execution. One unique aspect of the proposed methodology is that it uses a compilation flag based approach to collect hardware counts, eliminating the need for manual code instrumentation. Another unique aspect is the introduction of a data compression algorithm as the classifier. Using several representative real-world data-only exploits, our experiments show that data-only attacks can be detected with high accuracy using the proposed methodology. We also performed analysis on how to select the most relevant hardware events for the detection of the studied data-only attack, as well as a quantitative study of hardware events' sensitivity to interference. Gildo Torres, Chen Liu 0001 |
CF | 2 |
| 2021 | Feature Creation Towards the Detection of Non-control-Flow Hijacking Attacks
Zander Blasingame, Chen Liu 0001, Xin Yao 0001 |
ICANN (1) | 2 |
| 2021 | Leveraging Adversarial Learning for the Detection of Morphing AttacksabstractAn emerging threat towards face recognition systems (FRS) is face morphing attack, which involves the combination of two faces from two different identities into a singular image that would trigger an acceptance for either identity within the FRS. Many of the existing morphing attack detection (MAD) approaches have been trained and evaluated on datasets with limited variation of image characteristics, which can make the approach prone to overfitting. Additionally, there has been difficulty in developing MAD algorithms which can generalize beyond the morphing attack they were trained on, as shown by the most recent NIST FRVT MORPH report. Furthermore, the Single image based MAD (S-MAD) problem has had poor performance, especially when compared to its counterpart, Differential based MAD (D-MAD). In this work, we propose a novel architecture for training deep learning based S-MAD algorithms that leverages adversarial learning to train a more robust detector. The performance of the proposed S-MAD method is benchmarked against the state-of-the-art VGG19 based S-MAD algorithm over 36 experiments using the ISO-IEC 30107-3 evaluation metrics. The proposed method has demonstrated superior and robust detection performance of less than 5% D-EER when evaluated against different morphing attacks. Zander Blasingame, Chen Liu 0001 |
IJCB | 2 |
| 2021 | TUPPer-Map: Temporal and Unified Panoptic Perception for 3D Metric-Semantic MappingabstractIn this paper, we propose TUPPer-Map, a metric-semantic mapping framework based on the unified panoptic segmentation and temporal data association. In contrast to the previous mapping method, our framework integrates the data association stage into the holistic pixel-level segmentation stage in an end-to-end fashion, taking advantage of both intra-frame and inter-frame spatial and temporal knowledge. Firstly, we unify two-branch instance segmentation network and semantic segmentation network into a single network by sharing the backbone net, maximizing the 2D panoptic segmentation performance. Next, we leverage geometric segmentation to refine the segments predicted via deep learning. Then, we design a novel deep learning based data association module to track the object instances across different frames. Optical flow of consecutive frames and alignment of ROI (Region of Interest) candidates are learned to predict the frame-consistent instance label. At last, 2D semantics are integrated into 3D volume by TSDF raycasting to build the final map. We evaluated the performance of our framework extensively over the SceneNN, ScanNet v2 and Cityscapes-VPS datasets. Our experimental results demonstrate the superiority of TUPPer-Map over existing semantic mapping methods. Overall, our work illustrates that using learning based data association strategy can enable a more unified perception network for 3D mapping. Zhiliu Yang, Chen Liu 0001 |
IROS | 2 |
| 2021 | Heterogeneous computation in specific domain accelerations
Chen Liu 0001, Songwen Pei |
Future Gener. Comput. Syst. | 1 |
| 2020 | π-Map: A Decision-Based Sensor Fusion with Global Optimization for Indoor MappingabstractIn this paper, we propose π-map, a tightly coupled fusion mechanism that dynamically consumes LiDAR and sonar data to generate reliable and scalable indoor maps for autonomous robot navigation. The key novelty of π-map over previous attempts is the utilization of a fusion mechanism that works in three stages: the first LiDAR scan matching stage efficiently generates initial key localization poses; the second optimization stage is used to eliminate errors accumulated from the previous stage and guarantees that accurate large-scale maps can be generated; then the final revisit scan fusion stage effectively fuses the LiDAR map and the sonar map to generate a highly accurate representation of the indoor environment. We evaluate π-map on both large and small environments and verify its superiority over existing fusion methods. Zhiliu Yang, Bo Yu 0014, Jie Tang 0003, Shaoshan Liu, Chen Liu 0001 |
IROS | 6 |
| 2020 | On the code modernization of shared sampling alpha matting with OpenMP
Tien-Hsiung Weng, Kuanching Li, Zhiliu Yang, Chen Liu 0001 |
Future Gener. Comput. Syst. | 4 |
| 2018 | IoT Edge Device Based Key Frame Extraction for Face in Video RecognitionabstractFollowing the development of computing and communication technologies, the idea of Internet of Things (IoT) has been realized not only at research level but also at application level. Among various IoT-related application fields, biometrics applications, especially face recognition, are widely applied in video-based surveillance, access control, law enforcement and many other scenarios. In this paper, we introduce a Face in Video Recognition (FivR) framework which performs real-time key-frame extraction on IoT edge devices, then conduct face recognition using the extracted key-frames on the Cloud back-end. With our key-frame extraction engine, we are able to reduce the data volume hence dramatically relief the processing pressure of the cloud back-end. Our experimental results show with IoT edge device acceleration, it is possible to implement face in video recognition application without introducing the middle-ware or cloud-let layer, while still achieving real-time processing speed. Xuan Qi, Chen Liu 0001, Stephanie Schuckers |
CCGrid | 2 |
| 2018 | Teaching Autonomous Driving Using a Modular and Integrated ApproachabstractIntroduction: Teaching autonomous driving is a challenging task. Indeed, most existing autonomous driving teaching activities focus on a few of the technologies involved. This not only fails to provide a comprehensive coverage, but also sets a high entry barrier for students with different backgrounds. Objective: The primary objective of this study is to present a modular, integrated approach towards teaching autonomous driving. Methods: We organize the technologies used in autonomous driving into modules. This is described in the textbook we have developed as well as a series of multimedia online lectures designed to provide technical overview for each module. Once the students have understood these modules, the experimental platforms for integration we have developed allow the students to fully understand how the modules interact with each other. Results: To verify this teaching approach, we present three case studies: an introductory class on autonomous driving for students with only a basic technology background; a new session in an existing embedded systems class to demonstrate how embedded system technologies can be applied towards autonomous driving; and an industry professional training session to quickly bring up experienced engineers to work in autonomous driving. The results show that students can maintain a high interest level and make great progress by starting with familiar concepts before moving onto other modules. Conclusions: Autonomous driving is not one single technology, but rather a complex system integrating many technologies. Our modular and integrated approach is an effective method in teaching autonomous driving. Jie Tang 0003, Shaoshan Liu, Songwen Pei, Stéphane Zuckerman, Chen Liu 0001, Weisong Shi, Jean-Luc Gaudiot |
COMPSAC (1) | 5 |
| 2015 | How can Garbage Collection be energy efficient by dynamic offloading?abstractGarbage Collection (GC) is still a major issue in JVM for both mobile and cluster computing. GC offloading is proposed to improve the performance of GC by delivering part or all of the operations into another dedicated GC hardware. However, the traditional offloading just offloads directly not considering the phase change of GC behavior, which can be classified into two different groups: minor GC and major GC. The minor GC is fast and frequently invoked, while major GC is expensive in terms of time but seldom takes place. The direct offloading made GC workload frequently hopping between main processor and GC hardware, introduced a noticeable overhead and offset any possible benefits of workload loading. To solve this issue, we propose to offload GC dynamically by a careful selection of profitable and harmful GC operations. We also made a case study on Apache Spark, a lightning-fast cluster computing platform. It shows dynamic offloading can yield nearly 42.6% performance improvement with a concurrent 32.1% in energy cost reduction. Jie Tang 0003, Chen Liu 0001, Jean-Luc Gaudiot |
ASAP | 2 |
| 2015 | GPU-Accelerated Key Frame Analysis for Face Detection in VideoabstractSince video monitoring cameras now are implemented widely, video analytics has drawn much attention as a new research area. Correspondingly, there is an emerging need to detect faces through a period of video or a great number of videos to make comparisons with the stored identities in the database for personal identification or other purposes. Thus, face detection in video has gained great attention. However, when there are lots of videos or when the video is very long, the workload for face detection becomes very huge. As a result, the detecting time is prolonged. Thus, there is a need to detect faces quickly with reduced computation time. In this paper, we propose to add key frame analysis into the face detection process and employ the state-of-the-art graphic processing unit (GPU) platform to improve the overall performance. Our experimental results show that speedup as high as 3 folds can be achieved in terms of frames per second processed when GPU platform is introduced, compared with high-performance general-purpose CPUs. This shows great potential in applying GPU towards this kind of applications. Xuan Qi, Chen Liu 0001 |
CloudCom | 2 |
| 2014 | How many cores do we need to run a parallel workload: A test drive of the Intel SCC platform?
Chen Liu 0001, Pollawat Thanarungroj, Jean-Luc Gaudiot |
J. Parallel Distributed Comput. | 1 |
| 2013 | OCP: Offload Co-Processor for energy efficiency in embedded mobile systemsabstractIn current embedded mobile systems design, the application processor (AP) is often woken up to service interrupts and user requests. However, this kind of wakeups from sleep is very expensive in terms of battery usage. In the observation that the operating system/driver workloads are very light-weight, in this paper we propose the Offload Co-Processor (OCP) SoC architecture. In the OCP SoC design, when the device is idle, we offload the operating system workloads (mainly interrupt handling workloads) to an ultra-low-power coprocessor. This way, the co-processor would be able to handle most wake-up requests without awakening the heavy-weight AP, thus avoiding the overhead of AP spin-up/down. Using GPS continuous sampling workload as a case study, we show that the proposed OCP SoC design would extend battery life by 3.5 folds. Jie Tang 0003, Chen Liu 0001, Yu-Liang Chou, Shaoshan Liu |
ASAP | 2 |
| 2013 | A Power-Aware Study of Iris Matching Algorithms on Intel's SCCabstractBiometric applications become paramount across private sectors, industry, as well as government agencies. As large amount of data being collected from many different sources, managing such volumes of data and developing efficient and effective large-scale operational solutions are becoming a concern. For example, real-time identification of individuals with the purpose of allowing or denying their access to specific system or resource is challenging from the performance point of view. In addition, processing large amount of data would definitely consume a significant amount of energy. The Single-chip Cloud Computer (SCC) is an experimental processor created by Intel Labs. In this paper we employ SCC, which supports dynamic frequency and voltage scaling (DVFS), to investigate the power-aware computing and performance enhancement of an iris matching algorithm on such many-core architecture. This biometric application contains a large degree of parallelism that we can exploit by porting it onto the SCC. Results in terms of performance, power, energy, energy delay product (EDP), and power per speedup (PPS) metrics of executing the iris matching application under different number of cores, frequency, and voltage settings of the SCC platform are presented. We also analyze how the results for these metrics vary as we change these parameters. Gildo Torres, Jed Kao-Tung Chang, Fang Hua, Chen Liu 0001, Stephanie Schuckers |
ICPP | 4 |
| 2013 | Pinned OS/Services: A Case Study of XML Parsing on Intel SCC
Jie Tang 0003, Pollawat Thanarungroj, Chen Liu 0001, Shaoshan Liu, Zhimin Gu, Jean-Luc Gaudiot |
J. Comput. Sci. Technol. | 3 |
| 2013 | Acceleration of XML Parsing through PrefetchingabstractExtensible Markup Language (XML) has become a widely adopted standard for data representation and exchange. However, its features also introduce significant overhead threatening the performance of modern applications. In this paper, we present a study of XML parsing and determine that memory-side data loading in the parsing stage incurs a significant performance overhead, as much as the computation does. Hence, we propose memory-side acceleration which incorporates of data prefetching techniques, and can be applied on top of computation-side acceleration to speed up the XML data parsing. To this end, we study here the impact of our proposed scheme on the performance and energy consumption and demonstrated how it is capable of improving performance by up to 20 percent as well as produce up to 12.77 percent of energy saving when implemented in 32-nm technology. In addition, we implement a prefetcher on an platform in an effort to evaluate its implementation feasibility in terms of area and energy overhead. Jie Tang 0003, Shaoshan Liu, Chen Liu 0001, Zhimin Gu, Jean-Luc Gaudiot |
IEEE Trans. Computers | 3 |
| 2011 | Power and energy consumption analysis on intel SCC many-core systemabstractAs semiconductor manufacturing technology continues to scale down, it is possible to integrate more transistors into a processor, which gives birth to many-core processor design. This paper introduces an approach to analyze the power and energy consumption of a many-core research platform. The investigation has been done by using the Intel SCC system as an experimental platform. The approach is to collect the time and power profiling of an executing parallel application on the Intel SCC system. And then, the total energy consumed by the entire execution is calculated as a consequence. We studied the effects of power and energy consumption in many-core systems by varying different hardware configuration parameters such as number of cores, clock frequency and voltage level. Thus, the many-core system can be explored for its scalability and fitness in operational cost and performance. Pollawat Thanarungroj, Chen Liu 0001 |
IPCCC | 2 |
| 2011 | Memory-Side Acceleration for XML Parsing
Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Chen Liu 0001, Jean-Luc Gaudiot |
NPC | 4 |
| 2011 | Workload Characterization of Cryptography Algorithms for Hardware AccelerationabstractData encryption/decryption has become an essential component for modern information exchange. However, executing these cryptographic algorithms is often associated with huge overhead and the need to reduce this overhead arises correspondingly. In this paper, we select nine widely adopted cryptography algorithms and study their workload characteristics. Different from many previous works, we consider the overhead not only from the perspective of computation but also focusing on the memory access pattern. We break down the function execution time to identify the software bottleneck suitable for hardware acceleration. Then we categorize the operations needed by these algorithms. In particular, we introduce a concept called 'Load-Store Block' (LSB) and perform LSB identification of various algorithms. Our results illustrate that for cryptographic algorithms, the execution rate of most hotspot functions is more than 60%; memory access instruction ratio is mostly more than 60%; and LSB instructions account for more than 30% for selected benchmarks. Based on our findings, we suggest future directions in designing either the hardware accelerator associated with microprocessor or specific microprocessor for cryptography applications. Jed Kao-Tung Chang, Chen Liu 0001, Shaoshan Liu, Jean-Luc Gaudiot |
ICPE | 2 |
| 2010 | Hardware-assisted security mechanism: The acceleration of cryptographic operations with low hardware costabstractThis paper presents generic cryptographic accelerator. Certain "hotspot function" are found in the cryptographic algorithm which consume a substantial amount of execution time of the specific algorithm. INTEL performance analyzer VTune was used which analyzes the software performance on IA-32 and Intel64-based machines to examne the hotsport function. By moving the operations to hardware, we can reduce the overheads introduced by the crypto-computation so that the computing resource can focus on the useful work. Jed Kao-Tung Chang, Shaoshan Liu, Jean-Luc Gaudiot, Chen Liu 0001 |
IPCCC | 4 |
| 2010 | The Performance Analysis and Hardware Acceleration of Crypto-computations for Enhanced SecurityabstractSecurity is very important in modern life due to most information is now stored in digital format. A good security mechanism will keep information secrecy and integrity, hence, plays an important role in modern information exchange. However, cryptography algorithms are extremely expensive in terms of execution time. To make data not easily being cracked, many arithmetic and logical operations will be executed in the encryption/decryption process with many data movement. This means the cryptographic applications are both computation and memory intensive. Using a general-purpose processor for this scenario would not be very cost-effective. This study addresses this problem. Compared to the previous designs, we used a performance analyzer to identify “hotspot” functions across a set of benchmarks. The hotspot function consumes a substantial amount of the execution time of the specific algorithm. Then we translate these hotspot functions into hardware accelerators to improve the performance. Overall we achieve 34 - 83 folds of speedup. Jed Kao-Tung Chang, Shaoshan Liu, Jean-Luc Gaudiot, Chen Liu 0001 |
PRDC | 4 |
| 2009 | The Impact of Resource Sharing Control on the Design of Multicore Processors
Chen Liu 0001, Jean-Luc Gaudiot |
ICA3PP | 1 |