EDBT 2026 Demo / reviewers in the wild / expert
Shukai Chen
dblp:188/7248
· DBLP profile ↗
17ranked-venue papers
2as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Security and privacy · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BCLCM-Palm: Bézier curve-guided latent consistency model for fine-grained palmprint generation
Yuanpan Zhu, Kevin Chu, Shukai Chen, Weide Li |
Image Vis. Comput. | 3 |
| 2025 | DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion TransformerabstractSpeech-driven 3D facial animation is important for many multimedia applications. Recent work has shown promise in using either Diffusion models or Transformer architectures for this task. However, their mere aggregation does not lead to improved performance. We suspect this is due to a shortage of paired audio-4D data, which is crucial for the Transformer to effectively perform as a denoiser within the Diffusion framework. To tackle this issue, we present DiffSpeaker, a Transformer-based network equipped with novel biased conditional attention modules. These modules serve as substitutes for the traditional self/cross-attention in standard Transformers, incorporating thoughtfully designed biases that steer the attention mechanisms to concentrate on both the relevant task-specific and diffusion-related conditions. We also explore the trade-off between accurate lip synchronization and non-verbal facial expressions within the Diffusion paradigm. Experiments show our model achieves state-of-the-art performance on existing benchmarks, and fast inference speed owing to its ability to generate facial motions in parallel. Our code is avalable at https://github.com/theEricMa/DiffSpeaker. Zhiyuan Ma 0002, Xiangyu Zhu 0001, Chen Qian 0006, Shukai Chen, Guo-Jun Qi, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
IJCB | 4 |
| 2025 | StreamWMR: A Streaming Framework for Real-time 3D Whole-body Mesh Recoveryabstract3D whole-body mesh recovery aims to extract parameters for the human body, hands, and head from a single human image. Most applications related to human mesh recovery, such as physical fitness motion capture and operating room motion capture, necessitate real-time video stream processing. However, existing methods ignore the video processing and often require significant computational resources, making real-time performance unattainable and greatly limiting their practicality. Moreover, noticeable misalignments are often observed when concatenating them back to the body and reprojecting them onto the image. In this paper, we propose a streaming framework for whole-body mesh recovery in the video. First, we simplify pose regression by leveraging the root nodes of the hands and head to locate each component. Second, for temporal optimization, we incorporate attention mechanisms related to keypoint velocity to incorporate information from previous frames and achieve more stable and smooth motions. Finally, we propose a multi-view projection loss to eliminate the ambiguity caused by inaccurate 3D regression and pose estimation in computing reprojection errors. The combination of these enables our method to achieve real-time inference speed while maintaining accuracy and stability. Xiangyu Zhu 0001, Jinlin Wu, Zidu Wang, Shukai Chen, Dong Yi, Zhen Lei 0001 |
IJCB | 7 |
| 2025 | Exploiting Facial Discomfort Clues with Vision-Language Model for Generalizable Face Forgery DetectionabstractFace forgery detection is a challenging problem due to the diversity and rapid iteration of face manipulation methods, especially in detecting unknown forgery types. To address this challenge, we explore the common features shared among various forgery types. We find that even though multiple manipulation methods leave different invisible forgery traces, the fake faces often exhibit a similar overall pattern of discomfort. Such discomfort can serve as a universal clue across multiple forgery types, thereby possessing the potential to achieve strong generalization in face forgery detection. To this end, we utilize Vision-Language Models (VLMs) to simulate the cognitive process from perceiving the image to generating the sense of discomfort and propose a multi-task Forgery-Discomfort Joint Learning (FDJL) framework to leverage VLMs to perceive and identify fake faces by integrating discomfort cues. Specifically, we collect a Facial Discomfort dataset guided by the uncanny valley theory, enabling the model to extract and learn discomfort features. Extensive experiments demonstrate that our method achieves state-of-the-art performance and exhibits the best generalization for unknown forgery types. Tianshuo Zhang, Xiangyu Zhu 0001, Kai Pang, Shukai Chen, Zhen Lei 0001 |
IJCB | 5 |
| 2025 | ET-Talk: Effective Training Strategy to Enhance Synchrony and Fidelity for Talking Face GenerationabstractRecently, significant advancements have been made in audio-driven talking face generation. While GAN-based methods are widely used in this task, they struggle to achieve simultaneous lip accuracy and high-fidelity. Generated lip shapes tend to be overly influenced by the lip of reference images that provide identity information, leading to unstable and unsynchronized results. Moreover, the synthesized face frequently suffers from blurred teeth, skin textures, and compromised facial identity. To address these challenges, we propose an effective and innovative training strategy that simultaneously ensures lip synchrony and facial fidelity. First, we adaptively select the reference image using a hard-mining based strategy to prevent the network from simply copying the reference lip, enhancing the stability and synchronicity of lip movements. Second, we incorporate high-resolution facial images in training a quality discriminator within the GAN loss, improving the generated faces’ fidelity. Third, a global-to-detail training strategy is employed, starting with strengthening synchrony and then image quality to preserve identity and visual details. Experiments on the HDTF dataset demonstrate that our method achieves state-of-the-art performance in both lip accuracy and image quality. Baiqin Wang, Xiangyu Zhu 0001, Shukai Chen, Zhen Lei 0001 |
ICME | 5 |
| 2025 | CLDM-Palm: A controllable latent diffusion model for high-fidelity palmprint generation based on Bézier curves
Yuanpan Zhu, Donghuai Jia, Kevin Chu, Wenshuang Zhi, Weide Li, Shukai Chen |
Appl. Intell. | 6 |
| 2024 | A Secure, Flexible, and PPG-Based Biometric Scheme for Healthy IoT Using Homomorphic Random ForestabstractAdvances in the Internet of Things, such as biosensing and camera-capturing technologies, have also resulted in biometric-based authentication approaches becoming more viable. Photoplethysmography (PPG), for example, can be leveraged to provide better biometric features for continuous authentication, in comparison to other biometrics such as fingerprint. However, the fewer morphological features in PPG signals can complicate the accurate authentication of PPG signals. Furthermore, one has to also consider transmission security and PPG template storage security. These two considerations are typically not taken into consideration in most existing PPG-based authentication schemes. Therefore, we design a secure PPG-based biometric system to achieve accurate authentication with biometric privacy protection. In our design, Homomorphic Random Forest is adopted to classify the homomorphically encrypted biometric features, thus protecting the user’s PPG biometrics from being compromised in authentication. Furthermore, beat qualification screening is set up to avoid the interference of unqualified signals, and 19 features with the least redundancy from the 541 features extracted are selected as the biometric features of the user. Doing so allows us to ensure the accuracy of authentication, as demonstrated in our evaluations using five PPG databases collected by three different collection methods (contact, remote, and monitor). In addition, our experiments adopt PPG signals acquired from remote cameras, which are not considered in other PPG-based biometric systems. The experimental results show that the average accuracy of our biometric system is 96.4%, the F1 score is 96.1%, the equal error rate is 2.14%, and the authentication time is about 0.5 s. Liping Zhang 0003, Anzi Li, Shukai Chen, Wei Ren 0002, Kim-Kwang Raymond Choo |
IEEE Internet Things J. | 3 |
| 2024 | Network-level short-term traffic state prediction incorporating critical nodes: A knowledge-based deep fusion approach
Haipeng Cui, Shukai Chen, Qiang Meng 0001 |
Inf. Sci. | 2 |
| 2024 | Accurate authentication based on ECG using deep learningabstractBiometric-based authentication methods have been widely used, for example on portable devices (e.g., Android and iOS devices). However, there are several known limitations in existing authentication methods based on biometrics (e.g., those using facial, iris, and fingerprint). For example, in a healthcare context, a user may be physically incapable of completing the authentication due to his/her medical conditions. Hence, as a complementary authentication mechanism, there have been attempts to also utilize electrocardiogram (ECG). In this work, we propose an ECG authentication system that leverages deep learning. Specifically, to achieve generalization ability, complementary ensemble empirical decomposition (CEEMD) is introduced in our design. Moreover, a 1-D Multi-scale Convolutional Neural Network (1-D MCNN) is implemented to achieve accurate authentication. To evaluate the usability of our proposed approach, we have performed extensive experiments on eight databases, and the findings show that our proposed approach achieves good performance even on abnormal databases and can be adapted for different application environments. In addition, our adopted data from eight public databases requires theoretical statistical treatment for practical applications in real authentication scenarios. Liping Zhang 0003, Shukai Chen, Wei Ren 0002, Geyong Min, Kim-Kwang Raymond Choo |
J. Comput. Secur. | 2 |
| 2024 | 1DIEN: Cross-session Electrocardiogram Authentication Using 1D Integrated EfficientNetabstractThe potential of using electrocardiogram (ECG), an important physiological signal for humans, as a new biometric trait has been demonstrated, and ongoing efforts have focused on utilizing deep learning (e.g., 2D neural networks) to improve authentication accuracy (with some efficiency tradeoffs). In most of the existing ECG-based authentication approaches, the ECG recordings for enrollment and testing are collected within short intervals (e.g., within an hour). However, since ECG biometrics change over time, this design may decrease authentication accuracy when ECG recordings are collected weeks or even months prior. In this article, we propose 1D Integrated EfficientNet (1DIEN) to achieve cross-session ECG authentication. We adopt 1D neural networks as a lightweight alternative to 2D neural networks, and a voting scheme is designed to reduce variance and improve general authentication performance. We use three public ECG databases (i.e., an inter-session database, a mixed-session database, and an intra-session database) to evaluate our proposed 1DIEN under different authentication scenarios. The experimental results show that our approach achieves satisfactory performance for ECG authentication at a 3-month interval and is suitable for practical applications. Liping Zhang 0003, Shukai Chen, Fei Lin 0002, Wei Ren 0002, Kim-Kwang Raymond Choo, Geyong Min |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | An Efficient and Secure Health Data Propagation Scheme Using Steganography-Based Approach for Electronic Health NetworksabstractElectronic health (e-health) networks enable users to enjoy convenient, flexible, and low-cost medical services at home, so they attract great attention and spread into the market quickly. In e-health networks, large amounts of various health data including personal privacy information and physiological signals are transmitted, which raises security risks. To protect the health data transmitted in e-health networks, steganography-based solutions have been widely researched. Although existing steganography-based solutions successfully hide health data in physiological signals such as electrocardiograms (ECG), forward secrecy is not fully considered. This means that adversaries are able to extract users’ health data hidden in previous stego signals by using compromised long-term secrets. Moreover, to reduce communication overhead, compression techniques are introduced in some steganography-based methods. However, the imperceptibility and embedding capacity of these solutions are sacrificed. To solve the above issues, in this study, we adopt Singular Value Decomposition (SVD) and the Bose-Chaudhuri-Hocquenghem (BCH) codes to design an efficient and secure health data propagation scheme based on steganography and compression. In our design, the BCH codes are used to update the encryption key and change the embedding locations in each steganography process, thus achieving forward secrecy and further enhancing the security of steganography. Moreover, a two-stage compression method is proposed in our scheme to compress the signals during signal processing and compression phases, which effectively reduces the communication overhead. Security analysis and the experimental results show that our proposed scheme enhances security while achieving an elaborate balance between imperceptibility, embedding capacity, and compression. Liping Zhang 0003, Wenshuo Han, Shukai Chen, Kim-Kwang Raymond Choo |
IEEE/ACM Trans. Netw. | 3 |
| 2023 | Spatially Correlated Placement Policies for Wireless Content Caching NetworksabstractWe propose a geographic content placement policy for wireless caching networks. This policy, named Joint Caching Policy (JCP), jointly determines the caching strategy across the set of base stations (BSs) to improve the hit probability of an arbitrarily located user in the network versus caching policies where placement is independent and identically distributed across the BSs. To that end, JCP divides the BSs into groups, and executes a joint caching policy for each group, while content placement is independent across the groups. Existing joint caching policies require knowledge of user location to optimize content placement. On the other hand, JCP does not require any information on user location, and it provides a content placement policy that outperforms the one given by Independent Caching Policy (ICP), proposed by Błaszczyszyn and Giovanidis in 2015, under any user location distribution. We prove that the hit probability under JCP is lower bounded by that of ICP. We further propose an extension of JCP, named JCP-OPT, which improves the hit probability over JCP by solving a concave maximization problem, provided that there is side information about the user locations. We validate the performance of JCP and JCP-OPT via numerical evaluations and demonstrate that they can provide up to a 30% gain in hit probability over ICP. Shukai Chen, Derya Malak, Alhussein A. Abouzeid |
ICC | 1 |
| 2023 | An Adaptive Model Ensemble Adversarial Attack for Boosting Adversarial TransferabilityabstractWhile the transferability property of adversarial examples allows the adversary to perform black-box attacks (i.e., the attacker has no knowledge about the target model), the transfer-based adversarial attacks have gained great attention. Previous works mostly study gradient variation or image transformations to amplify the distortion on critical parts of inputs. These methods can work on transferring across models with limited differences, i.e., from CNNs to CNNs, but always fail in transferring across models with wide differences, such as from CNNs to ViTs. Alternatively, model ensemble adversarial attacks are proposed to fuse outputs from surrogate models with diverse architectures to get an ensemble loss, making the generated adversarial example more likely to transfer to other models as it can fool multiple models concurrently. However, existing ensemble attacks simply fuse the outputs of the surrogate models evenly, thus are not efficacious to capture and amplify the intrinsic transfer information of adversarial examples. In this paper, we propose an adaptive ensemble attack, dubbed AdaEA, to adaptively control the fusion of the outputs from each model, via monitoring the discrepancy ratio of their contributions towards the adversarial objective. Furthermore, an extra disparity-reduced filter is introduced to further synchronize the update direction. As a result, we achieve considerable improvement over the existing ensemble attacks on various datasets, and the proposed AdaEA can also boost existing transfer-based attacks, which further demonstrates its efficacy and versatility. The source code: https://github.com/CHENBIN99/AdaEA Bin Chen 0020, Jia-Li Yin, Shukai Chen, Ximeng Liu |
ICCV | 3 |
| 2023 | Spoof-Guided Image Decomposition for Face Anti-spoofing
Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Shukai Chen, Peng Li 0035, Zhen Lei 0001 |
PRCV (5) | 4 |
| 2023 | Camera-aware representation learning for person re-identification
Jinlin Wu, Zhen Lei 0001, Yang Yang 0062, Shukai Chen, Stan Z. Li |
Neurocomputing | 5 |
| 2022 | A Privacy-Preserving Proximity Testing Using Private Set Intersection for Vehicular Ad-Hoc NetworksabstractProximity testing technologies have been of increasing importance in vehicularad-hocnetworks (VANETs), especially in location-based services. However, there exist several known challenges in most existing proximity testing methods. For instance, during proximity testing, how to protect the location privacy of users, guarantee the fairness trait of both communication parties, and reduce computational costs is challenging. In this article, we present an efficient privacy-preserving proximity testing scheme using private set intersection (PSI) and differential privacy. In our design, a Chebyshev-based PSI is constructed to achieve location privacy with low energy consumption during the proximity testing process. Furthermore, geo-indistinguishability is employed in our scheme to generate virtual points as inputs set of PSI, which further protects the location privacy from exposure and provides resistance to collusion attacks. Fairness requirement is alsosatisfied in our scheme. The performance evaluation shows that the proposed scheme achieves good efficiency and is suitable for VANETs. Liping Zhang 0003, Wenhao Gao 0002, Shukai Chen, Wei Ren 0002, Kim-Kwang Raymond Choo, Naixue Xiong |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | An Optimal Dynamic Lane Reversal and Traffic Control Strategy for Autonomous VehiclesabstractThis paper studies an optimal dynamic lane reversal and traffic control (DLRTC) strategy in the presence of autonomous vehicles (AVs). A centralized controller is set to change lane directions dynamically and regulate traffic flow on a motorway network. Through vehicle to infrastructure (V2I) communication, the roadside sensors can send lane reversal information and flow control actions to the AVs which can perform lane-changing behaviors and adjust travel speed. To model the traffic dynamics under DLRTC, we propose a novel multi-lane cell transmission model (CTM). A logit model is used to characterize the lane-changing behaviors under uncontrolled cases. A mixed integer linear programming model (MILP) is formulated for DLRTC, and optimal control actions are implemented in a framework of model predictive control (MPC). The numerical experiments based on the Ayer Rajah Expressway (AYE) in Singapore are conducted to demonstrate the effectiveness of the proposed methods. The results show that the DLRTC strategy can effectively reduce road congestion and achieve better system performance compared to the benchmark method. Shukai Chen, Qiang Meng 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |