Jeng-Lin Li

dblp:213/8624 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0002-9261-1524ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 6 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 How Bias Binds: Measuring Hidden Associations for Bias Control in Text-to-Image Compositions
abstract
Text-to-image generative models often exhibit bias related to sensitive attributes. However, current research tends to focus narrowly on single-object prompts with limited contextual diversity. In reality, each object or attribute within a prompt can contribute to bias. For example, the prompt ``an assistant wearing a pink hat'' may reflect female-inclined biases associated with a pink hat. Neglecting joint semantic bindings in prompts leads to significant failures of current debiasing methods. We present a preliminary investigation into how bias manifests under semantic binding, where contextual associations between objects and attributes affect generative outcomes. We demonstrate that the underlying bias distribution can be amplified based on these associations. Therefore, we introduce a bias adherence score that quantifies how specific object-attribute bindings activate bias. To delve deeper, we develop a training-free context-bias control framework to explore how token decoupling can facilitate the debiasing of semantic bindings. This framework achieves over 10% debiasing improvement in compositional generation tasks. Our analysis of bias scores across various attribute-object bindings and token decorrelation highlights a fundamental challenge: reducing bias without disrupting essential semantic relationships. These findings expose critical limitations in current debiasing approaches when applied to semantically bound contexts, underscoring the need to reassess prevailing bias mitigation strategies.
Jeng-Lin Li, Ming-Ching Chang, Wei-Chao Chen
AAAI1
2026 PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection
Jeng-Lin Li, Po-Hsuan Huang, Ming-Ching Chang, Wei-Chao Chen
WACV2
2025 Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
abstract
Recent advancements in large vision-language models (LVLM) have significantly enhanced their ability to comprehend visual inputs alongside natural language. How-ever, a major challenge in their real-world application is hallucination, where LVLMs generate non-existent visual elements, eroding user trust. The underlying mechanism driving this multimodal hallucination is poorly understood. Minimal research has illuminated whether contexts such as sky, tree, or grass field involve the LVLM in hallucinating a frisbee. We hypothesize that hiddenfactors, such as objects, contexts, and semantic foreground-background structures, induce hallucination. This study proposes a novel causal approach: a hallucination probing system to identify these hidden factors. By analyzing the causality between images, text prompts, and network saliency, we systematically ex-plore interventions to block these factors. Our experimen-tal findings show that a straightforward technique based on our analysis can significantly reduce hallucinations. Additionally, our analyses indicate the potential to edit network internals to minimize hallucinated outputs.
Po-Hsuan Huang, Jeng-Lin Li, Chin-Po Chen, Ming-Ching Chang, Wei-Chao Chen
WACV2
2024 Improving Limited Supervised Foot Ulcer Segmentation Using Cross-Domain Augmentation Strategies
abstract
Diabetic foot ulcers pose health risks, including higher morbidity, mortality, and amputation rates. Monitoring wound areas is crucial for proper care, but manual segmentation is subjective due to complex wound features and background variation. Expert annotations are costly and time-intensive, thus hampering large dataset creation. Existing segmentation models relying on extensive annotations are impractical in real-world scenarios with limited annotated data. In this paper, we propose a cross-domain augmentation method named TransMix that combines Augmented Global Pre-training (AGP) and Localized CutMix Fine-tuning (LCF) to enrich wound segmentation data for model learning. TransMix can effectively improve the foot ulcer segmentation model training by leveraging other dermatology datasets not on ulcer skins or wounds. AGP effectively increases the overall image variability, while LCF increases the diversity of wound regions. Experimental results show that TransMix increases the variability of wound regions and substantially improves the Dice score for models trained with only 40 annotated images under various proportions.
Shang-Jui Kuo, Chia-Ching Lin, Jeng-Lin Li, Ming-Ching Chang
ICASSP4
2024 In-The-Wild Physiological-Based Stress Detection Using Federated Strategy
abstract
Continuously identifying day-to-day mental stress can be realized by accessing wearable devices to measure physiological indicators. However, the nature of bodily signals raises issues of privacy and data heterogeneity. Recent federated learning scheme provides a promising direction to alleviate the privacy concern, but the large inter-client differences can lead to a sub-optimal model performance. In this work, we propose a client-aware aggregation strategy to customize the global model forked by each client to conduct mutual learning in federated setting. Our proposed mixture Federated Mutual Learning (mixFML) weighs the distances of local models to generate a unique mixture of global model per client. We evaluated our method on the public TILES-2018 and an in-house Firefighters dataset for stress detection using HRV. Our proposed mixFML achieved 8.0% and 1.8% MCC improvement on two datasets compared to federated mutual learning.
Po-Chen Lin, Jeng-Lin Li, Woan-Shiuan Chien, Chi-Chun Lee
ICASSP2
2023 An Enroll-to-Verify Approach for Cross-Task Unseen Emotion Class Recognition
abstract
Most speech emotion recognition studies often focus on recognizing pre-set emotion classes. However, the task definition may change due to a shift in focus to a previously unseen class in real-world applications. This cross-task modeling has not been addressed previously. Lengthy data re-collection, model retraining, and the traditional adaptation and transfer learning approaches are not applicable to this cross-task setting. This study proposes an enroll-to-verify framework to avoid model retraining and rapidly perform a new task prediction using only a handful of enrolled samples. Specifically, we use negative angular margin prototypical loss in a pretrained multiclass network as an emotion encoder. Then, we enroll a few samples corresponding to emotion classes in the new task definition and simply compare the encoded embedding distance to perform recognition. In the experiments on the IEMOCAP dataset, given a four-class pretrained emotion encoder, we achieved a 71.9% unweighted average recall in the frustration (unseen) recognition task. The MELD dataset was used where the unseen class was surprise, fear, or disgust. The results revealed that enrolling only 20 samples without retraining was comparable to supervised training using the complete dataset. Further analyses were conducted to demonstrate the working mechanism of our proposed enroll-to-verify approach.
Jeng-Lin Li, Chi-Chun Lee
IEEE Trans. Affect. Comput.1
2022 Romantic and Family Movie Database: Towards Understanding Human Emotion and Relationship via Genre-Dependent Movies
abstract
Movie gains popularity by successfully immersing viewers into affective contents under an emotion perception and elicitation process. Film genre sets the tone of a movie to shape key emotional context during the storytelling procedure. In romantic and family movies, character relationship is a key contextual attribute guiding the whole storyline, which is expected to evoke a sense of being moved in the audience. In this work, we propose Romantic and Family Movie Database, which consists of 1029 movie clips from 10 romantic and family movies. We provide annotations including character relationship type, character relationship status, perceived emotions, induced emotions, and degree of being moved. Our analysis elaborates the inconsistency between perceived and induced valence with the support of annotations including character relationship status and degree of being moved. We also find that being moved and induced arousal are positively correlated in romantic and family movies. We present comprehensive baseline results in recognizing relationship status and emotions using different models and modalities. The database is publicly available at https://rfmd.ee.nthu.edu.tw/.
Po-Chien Hsu, Jeng-Lin Li, Chi-Chun Lee
ACII2
2022 An Audio-Saliency Masking Transformer for Audio Emotion Classification in Movies
abstract
The process of perception to affective response of humans is gated by a bottom-up saliency mechanism at the sensory level. In specifics, auditory saliency emphasizes audio segments that need to be attended to cognitively appraise and experience emotion. In this work, inspired by this mechanism, we propose an end-to-end feature masking network for audio emotion recognition in movies. Our proposed Audio-Saliency Masking Transformer (ASTM) adjusts feature embedding using two learnable masks; one of them cross-refers to an auditory saliency map, and the other one is through self-reference. By joint training for front-end mask gating and the transformer as the back-end emotion classifier, we achieve three-class UARs improvement of 1.74%, 1.27%, 0.95%, 0.82% when comparing to the best of the other models on experienced arousal, experienced valence, intended arousal, and intended valence, respectively. We further analyze which acoustic feature categories that our saliency mask attends to the most.
Ya-Tse Wu, Jeng-Lin Li, Chi-Chun Lee
ICASSP2
2022 A Chunking-for-Pooling Strategy for Cytometric Representation Learning for Automatic Hematologic Malignancy Classification
abstract
Differentiating types of hematologic malignancies is vital to determine therapeutic strategies for the newly diagnosed patients. Flow cytometry (FC) can be used as diagnostic indicator by measuring the multi-parameter fluorescent markers on thousands of antibody-bound cells, but the manual interpretation of large scale flow cytometry data has long been a time-consuming and complicated task for hematologists and laboratory professionals. Past studies have led to the development of representation learning algorithms to perform sample-level automatic classification. In this work, we propose a chunking-for-pooling strategy to include large-scale FC data into a supervised deep representation learning procedure for automatic hematologic malignancy classification. The use of discriminatively-trained representation learning strategy and the fixed-size chunking and pooling design are key components of this framework. It improves the discriminative power of the FC sample-level embedding and simultaneously addresses the robustness issue due to an inevitable use of down-sampling in conventional distribution based approaches for deriving FC representation. We evaluated our framework on two datasets. Our framework outperformed other baseline methods and achieved 92.3% unweighted average recall (UAR) for four-class recognition on the UPMC dataset and 85.0% UAR for five-class recognition on the hema.to dataset. We further compared the robustness of our proposed framework with that of the traditional downsampling approach. Analysis of the effects of the chunk size and the error cases revealed further insights about different hematologic malignancy characteristics in the FC data.
Jeng-Lin Li, Yun-Chun Lin, Yu-Fen Wang, Sara A. Monaghan, Bor-Sheng Ko, Chi-Chun Lee
IEEE J. Biomed. Health Informatics1
2020 Using Speaker-Aligned Graph Memory Block in Multimodally Attentive Emotion Recognition Network
Jeng-Lin Li, Chi-Chun Lee
INTERSPEECH1
2019 A Dual-Complementary Acoustic Embedding Network Learned from Raw Waveform for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) technology has recently become a trend in a broader field and has achieved remarkable recognition performances using deep learning technique. However, the recognition performances obtained using end-to-end learning directly from raw audio waveform still hardly exceed those based on hand-crafted acoustic descriptors. Instead of solely rely on raw waveform or acoustic descriptors for SER, we propose an acoustic space augmentation network, termed as Dual-Complementary Acoustic Embedding Network (DCaEN), that combines knowledge-based features with raw waveform embedding learned with a novel complementary constraint. DCaEN includes representations from eGeMAPS acoustic feature and raw waveform by specifying a negative cosine distance loss to explicitly constrain the raw waveform embedding to be different from eGeMAPS. Our experimental results demonstrate an improved emotion discriminative power on the IEMOCAP database, which achieves 59.31% in a four class emotion recognition. Our analysis also demonstrates that the learned raw waveform embedding of DCaEN converges close to reverse mirroring of the original eGeMAPS space.
Tzu-Yun Huang, Jeng-Lin Li, Chun-Min Chang, Chi-Chun Lee
ACII2
2019 Attention Learning with Retrievable Acoustic Embedding of Personality for Emotion Recognition
abstract
Modeling multimodal behavior streams to automatically identify emotion states of an individual has progressed extensively especially with the advancement of deep learning algorithms. Emotion, being an abstract internal state, creates substantial differences in an individual's behavior expressivity, the development of personalized recognition framework is a critical next step to improve algorithm's modeling capacity. In this work, we propose to integrate the target speaker's personality embedding into the learning of multimodal (speech and language) attention based network architecture to improve recognition performances. Specifically, we propose a Personal Attribute-Aware Attention Network (PAaAN) that learns its multimodal attention weights jointly with the target speaker's retrievable acoustic embedding of personality. Our acoustic domain adapted personality retrieval strategy mitigates the common issue on the lack of personality scores in the current available emotion databases, and our proposed PAaAN then learns its attention weight by jointly considering an individual target speaker's personality profile with his or her multimodal acoustic and lexical modalties. In this work, we achieve a 70% unweighted accuracy in the IEMOCAP 4-class multimodal emotion recognition task. Further analysis shows the effect of integrating personality on the variation of our attention weights of each acoustic and lexical behavior modality for each speaker in the IEMOCAP database.
Jeng-Lin Li, Chi-Chun Lee
ACII1
2019 Learning Semantic-preserving Space Using User Profile and Multimodal Media Content from Political Social Network
abstract
The use of social media in politics has dramatically changed the way campaigns are run and how elected officials interact with their constituents. An advanced algorithm is required to analyze and understand this large amount of heterogeneous social media data to investigate several key issues, such as stance and strategy, in political science. Most of previous works concentrate their studies using text-as-data approach, where the rich yet heterogeneous information in the user profile, social relationship, and multimodal media content is largely ignored. In this work, we propose a two-branch network that jointly maps the post contents and politician profile into the same latent space, which is trained using a large-margin objective that combines a cross-instance distance constraint with a within-instance semantic-preserving constraint. Our proposed political embedding space can be utilized not only in reliably identifying political spectrum and message type but also in providing a political representation space for interpretable ease-of-visualization.
Wei-Hao Chang, Jeng-Lin Li, Chi-Chun Lee
ICASSP2
2019 Learning Minimal Intra-Genre Multimodal Embedding from Trailer Content and Reactor Expressions for Box Office Prediction
abstract
Movie watching is one of the most popular leisure activities in our daily life. The box office revenue, especially in the first week, is critical for financial planning in the movie industry. Most existing movie box office prediction relies on meta data, viewer's comments, and trailer content. However, when viewers are immersed in a movie experience, they would naturally manifest expressions invoked by the media content. In this work, we propose a novel movie box office prediction framework by joint modeling meta attributes, trailer content, and viewer's natural expressions gathered from YouTube reactor videos. The proposed network learns a discriminabilityenhanced content and expression embeddings using a minimal intra-genre distance loss function. The proposed architecture achieves 79.07%, 73.79% and 76.82% for low/high movie box office tier classification (top 30%, top 10% and top 5%) on a large scale trailer-reactor database. Furthermore, we provide an analysis on the effectiveness of viewer's reaction and our intra-genre projection. Most existing movie box office prediction relies on meta data, viewer's comments, and trailer content. However, when viewers are immersed in a movie experience, they would naturally manifest expressions invoked by the media content. In this work, we propose a novel movie box office prediction framework by joint modeling meta attributes, trailer content, and viewer's natural expressions gathered from YouTube reactor videos. The proposed network learns a discriminability-enhanced content and expression embeddings using a minimal intra-genre distance loss function. The proposed architecture achieves 79.07%, 73.79% and 76.82% for low/high movie box office tier classification (top 30%, top 10% and top 5%) on a large scale trailer-reactor database. Furthermore, we provide an analysis on the effectiveness of viewer's reaction and our intra-genre projection.
Ming-Ya Ko, Jeng-Lin Li, Chi-Chun Lee
ICME2
2019 Investigating the Variability of Voice Quality and Pain Levels as a Function of Multiple Clinical Parameters
Hui-Ting Hong, Jeng-Lin Li, Yi-Ming Weng, Chip-Jin Ng, Chi-Chun Lee
INTERSPEECH2
2019 Attentive to Individual: A Multimodal Emotion Recognition Network with Personalized Attention Profile
Jeng-Lin Li, Chi-Chun Lee
INTERSPEECH1
2018 A Genre-Affect Relationship Network with Task-Specific Uncertainty Weighting foR Recognizing Induced Emotion in Music
abstract
Emotion is a core fundamental attribute of humans. Using music to induce emotional responses from subjects to better facilitate human behavior shaping have been effective across domains of health, education, and retail. Computationally model the musically-induced emotion provides necessary content-based analytics for large-scale and wide-applicability of such human-centered applications. In this work, we propose a relationship neural network architecture to learn to regress the induced emotion attributes with an auxiliary task of genre classification. Our proposed Genre-Affect Relationship Network with homoscedastic uncertainty weighting embeds the relationship between affect and genre as tensor normal prior within task-specific layers; the architecture is optimized further by incorporating task-specific uncertainty. The proposed architecture achieves a state-of-art 0.564 average Pearson correlation computed over nine induced emotion ratings in the Emotify database. Furthermore, we provide an analysis to understand the relationship between the induced emotions of these musical pieces and their associated genres.
Wei-Hao Chang, Jeng-Lin Li, Yun-Shao Lin, Chi-Chun Lee
ICME2
2018 Generating fMRI-Enriched Acoustic Vectors using a Cross-Modality Adversarial Network for Emotion Recognition
abstract
Automatic emotion recognition has long been developed by concentrating on modeling human expressive behavior. At the same time, neuro-scientific evidences have shown that the varied neuro-responses (i.e., blood oxygen level-dependent (BOLD) signals measured from the functional magnetic resonance imaging (fMRI)) is also a function on the types of emotion perceived. While past research has indicated that fusing acoustic features and fMRI improves the overall speech emotion recognition performance, obtaining fMRI data is not feasible in real world applications. In this work, we propose a cross modality adversarial network that jointly models the bi-directional generative relationship between acoustic features of speech samples and fMRI signals of human percetual responses by leveraging a parallel dataset. We encode the acoustic descriptors of a speech sample using the learned cross modality adversarial network to generate the fMRI-enriched acoustic vectors to be used in the emotion classifier. The generated fMRI-enriched acoustic vector is evaluated not only in the parallel dataset but also in an additional dataset without fMRI scanning. Our proposed framework significantly outperform using acoustic features only in a four-class emotion recognition task for both datasets, and the use of cyclic loss in learning the bi-directional mapping is also demonstrated to be crucial in achieving improved recognition rates.
Gao-Yi Chao, Chun-Min Chang, Jeng-Lin Li, Ya-Tse Wu, Chi-Chun Lee
ICMI3
2018 Encoding Individual Acoustic Features Using Dyad-Augmented Deep Variational Representations for Dialog-level Emotion Recognition
Jeng-Lin Li, Chi-Chun Lee
INTERSPEECH1
2018 Learning Conditional Acoustic Latent Representation with Gender and Age Attributes for Automatic Pain Level Recognition
Jeng-Lin Li, Yi-Ming Weng, Chip-Jin Ng, Chi-Chun Lee
INTERSPEECH1
2018 Self-Assessed Affect Recognition Using Fusion of Attentional BLSTM and Static Acoustic Features
Bo-Hao Su, Sung-Lin Yeh, Ming-Ya Ko, Huan-Yu Chen, Shun-Chang Zhong, Jeng-Lin Li, Chi-Chun Lee
INTERSPEECH6
2017 A bootstrapped multi-view weighted Kernel fusion framework for cross-corpus integration of multimodal emotion recognition
abstract
Recently the development of robust emotion recognition has been increasingly emphasized in order to handle situations of different cultures and languages. This has become critical due to the potential applicability of emotion recognizers across a wide range of application scenarios. Instead of conventional approach in deriving a single universal emotion recognition module across all languages, we have previously demonstrated a method based on integrating other database's useful information to improve the emotion recognition of the current data with fusion of multiple emotion perspectives. In this paper, we present an improved framework, i.e., a bootstrapped multi-view weighted kernel fusion, to further advance the recognition accuracies. We have also extended the modeling of speech-only modality to include video information. In specifics, we utilize two emotional corpora of different languages. Our proposed framework obtains improved recognition in regressing activation and valence attributes using audio and video modalities across both of the databases. We not only demonstrate that the weighted kernel fusion can provide additional modeling power but also present analyses on the complementary emotionally-relevant acoustic and visual behaviors computed from the multiple emotion perspectives.
Chun-Min Chang, Bo-Hao Su, Shih-Chen Lin, Jeng-Lin Li, Chi-Chun Lee
ACII4