Cheng-Yuan Sun

dblp:339/8239 · DBLP profile ↗
← Back
2ranked-venue papers in the field
0as first author
2since 2021 · last 2023
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2023 LoRA-like Calibration for Multimodal Deception Detection using ATSFace Data
abstract
Recently, deception detection on human videos is an eye-catching techniques and can serve lots applications. AI model in this domain demonstrates the high accuracy, but AI tends to be a non-interpretable black box. We introduce an attention-aware neural network addressing challenges inherent in video data and deception dynamics. This model, through its continuous assessment of visual, audio, and text features, pinpoints deceptive cues. We employ a multimodal fusion strategy that enhances accuracy; our approach yields a 92% accuracy rate on a real-life trial dataset. Most important of all, the model indicates the attention focus in the videos, providing valuable insights on deception cues. Hence, our method adeptly detects deceit and elucidates the underlying process. We further enriched our study with an experiment involving students answering questions either truthfully or deceitfully, resulting in a new dataset of 309 video clips, named ATSFace. Using this, we also introduced a calibration method, which is inspired by Low-Rank Adaptation (LoRA), to refine individual-based deception detection accuracy.
Shun-Wen Hsiao, Cheng-Yuan Sun
IEEE Big Data2
2022 Attention-Aware Multi-modal RNN for Deception Detection
abstract
Nowadays, various video data are springing up. In the field of human-centric video analysis, deception detection becomes a crucial issue to us. Thanks to AI development, automated deception detection has been researched for a while. Nonetheless, the mainstream techniques work as a black box, which is unexplainable. This paper presents an attention mechanism on visual and audio features, which makes the detection interpretable. Besides, we embrace the approach of multi-modal by combining visual, audio, and transcription features as an ensemble model, which can achieve 96% of accuracy.
Shun-Wen Hsiao, Cheng-Yuan Sun
IEEE Big Data2