Ling Lo

dblp:254/7931 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-9471-8528ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SeCo: Semantic-Guided Multimodal Color Splash Effects
abstract
Color splash is a widely used image editing effect that highlights selected regions by retaining color while rendering the rest of the image in grayscale. However, existing tools often struggle with achieving high precision, efficiency, and user flexibility in controlling the effect. In this article, we propose Semantic-Guided Multimodal Color Splash Effects (SeCo), a novel framework for generating stylized and customizable color splash effects from natural language instructions and color palettes. SeCo decomposes the task into two key components: Semantic-Guided Object Isolation (SGOI) and Palette-Driven Color Adjustment (PDCA). SGOI accurately identifies and isolates user-referred objects with fine-grained transparency, while the PDCA module recolors the isolated regions under user-specified palette guidance. Our approach supports arbitrary object selection, handles transparency, and enables diverse stylization patterns. Experimental results on both synthetic and real-world datasets demonstrate that SeCo outperforms existing methods in precision and controllability, offering a practical and expressive solution for visual editing and content creation.
Jing-Xuan Chen, Ling Lo, Si-Yu Lu, Wen-Huang Cheng, Jungwoo Huh, Sanghoon Lee 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Future Sight and Tough Fights: Revolutionizing Sequential Recommendation with FENRec
abstract
Sequential recommendation (SR) systems predict user preferences by analyzing time-ordered interaction sequences. A common challenge for SR is data sparsity, as users typically interact with only a limited number of items. While contrastive learning has been employed in previous approaches to address the challenges, these methods often adopt binary labels, missing finer patterns and overlooking detailed information in subsequent behaviors of users. Additionally, they rely on random sampling to select negatives in contrastive learning, which may not yield sufficiently hard negatives during later training stages. In this paper, we propose Future data utilization with Enduring Negatives for contrastive learning in sequential Recommendation (FENRec). Our approach aims to leverage future data with time-dependent soft labels and generate enduring hard negatives from existing data, thereby enhancing the effectiveness in tackling data sparsity. Experiment results demonstrate our state-of-the-art performance across four benchmark datasets, with an average improvement of 6.16% across all metrics.
Yu-Hsuan Huang 0002, Ling Lo, Hong-Han Shuai, Wen-Huang Cheng
AAAI2
2025 From Prompt to Progression: Taming Video Diffusion Models for Seamless Attribute Transition
Ling Lo, Kelvin C. K. Chan, Wen-Huang Cheng, Ming-Hsuan Yang 0001
ICCV1
2024 Distraction is All You Need: Memory-Efficient Image Immunization against Diffusion-Based Image Editing
abstract
Recent text-to-image (T2I) diffusion models have revolutionized image editing by empowering users to control out-comes using natural language. However, the ease of image manipulation has raised ethical concerns, with the poten-tial for malicious use in generating deceptive or harmful content. To address the concerns, we propose an image im-munization approach named semantic attack to protect our images from being manipulated by malicious agents using diffusion models. Our approach focuses on disrupting the semantic understanding of T2I diffusion models regarding specific content. By attacking the cross-attention mecha-nism that encodes image features with text messages during editing, we distract the model's attention regarding the con-tent of our concern. Our semantic attack renders the model uncertain about the areas to edit, resulting in poorly edited images and contradicting the malicious editing attempts. In addition, by shifting the attack target towards intermediate attention maps from the final generated image, our approach substantially diminishes computational burden and alleviates GPU memory constraints in comparison to pre-vious methods. Moreover, we introduce timestep universal gradient updating to create timestep-agnostic perturbations effective across different input noise levels. By treating the full diffusion process as discrete denoising timesteps during the attack, we achieve equivalent or even superior immu-nization efficacy with nearly half the memory consumption of the previous method. Our contributions include a prac-tical and effective approach to safeguard images against malicious editing, and the proposed method offers robust immunization against various image inpainting and editing approaches, showcasing its potential for real-world appli-cations.
Ling Lo, Cheng Yu Yeo, Hong-Han Shuai, Wen-Huang Cheng
CVPR1
2024 ReCorD: Reasoning and Correcting Diffusion for HOI Generation
abstract
Diffusion models revolutionize image generation by leveraging natural language to guide the creation of multimedia content. Despite significant advancements in such generative models, challenges persist in depicting detailed human-object interactions, especially regarding pose and object placement accuracy. We introduce a training-free method named Reasoning and Correcting Diffusion (ReCorD) to address these challenges. Our model couples Latent Diffusion Models with Visual Language Models to refine the generation process, ensuring precise depictions of HOIs. We propose an interaction-aware reasoning module to improve the interpretation of the interaction, along with an interaction correcting module to refine the output image for more precise HOI generation delicately. Through a meticulous process of pose selection and object positioning, ReCorD achieves superior fidelity in generated images while efficiently reducing computational requirements. We conduct comprehensive experiments on three benchmarks to demonstrate the significant progress in solving text-to-image generation tasks, showcasing ReCorD's ability to render complex interactions accurately by outperforming existing methods in HOI classification score, as well as FID and Verb CLIP-Score. Project website is available at https://alberthkyhky.github.io/ReCorD/ .
Jian-Yu Jiang-Lin, Kang-Yang Huang, Ling Lo, Yi-Ning Huang, Terence Lin, Jhih-Ciang Wu, Hong-Han Shuai, Wen-Huang Cheng
ACM Multimedia3
2024 Modeling Uncertainty for Low-Resolution Facial Expression Recognition
abstract
Recently, facial expression recognition techniques have made significant progress on high-resolution web images. However, in real-world applications, the obtained images are often with low resolution since they are mostly captured in a wide range of public spaces. As a result, the ambiguity of the expression labels hinders recognition performance due to not only subjective emotion annotations but also ambiguous images. Existing approaches tend to perform poorly when the resolution of face images decreases. In this work, we aim to model the aleatoric uncertainty induced by low-image-resolution and label ambiguity for robust facial expression recognition. We propose probabilistic data uncertainty learning to capture the ambiguity induced by poor image resolution. Additionally, we introduce the emotion wheel to learn the label-uncertainty-aware embedding. Moreover, we exploit the ambiguous nature of neutrality and propose a neutral expression constraint to learn more robust features for facial expression recognition. To the best of our knowledge, this is the first work utilizing the intrinsic nature of neutrality as a regularization to benefit model training. Extensive experimental results show the effectiveness and robustness of our approach. Under low-resolution conditions, our proposed method outperforms the state-of-the-art approaches by 3.02% and 3.16% in terms of accuracy on RAF-DB and FERPlus, respectively.
Ling Lo, Bo-Kai Ruan, Hong-Han Shuai, Hao-Wen Cheng
IEEE Trans. Affect. Comput.1
2023 An Overview of Facial Micro-Expression Analysis: Data, Methodology and Challenge
abstract
Facial micro-expressions indicate brief and subtle facial movements that appear during emotional communication. In comparison to macro-expressions, micro-expressions are more challenging to be analyzed due to the short span of time and the fine-grained changes. In recent years, micro-expression recognition (MER) has drawn much attention because it can benefit a wide range of applications, e.g. police interrogation, clinical diagnosis, depression analysis, and business negotiation. In this survey, we offer a fresh overview to discuss new research directions and challenges these days for MER tasks. For example, we review MER approaches from three novel aspects: macro-to-micro adaptation, recognition based on key apex frames, and recognition based on facial action units. Moreover, to mitigate the problem of limited and biased ME data, synthetic data generation is surveyed for the diversity enrichment of micro-expression data. Since micro-expression spotting can boost micro-expression analysis, the state-of-the-art spotting works are also introduced in this paper. At last, we discuss the challenges in MER research and provide potential solutions as well as possible directions for further investigation.
Ling Lo, Hong-Han Shuai, Wen-Huang Cheng
IEEE Trans. Affect. Comput.2
2022 Mimicking the Annotation Process for Recognizing the Micro Expressions
abstract
Micro-expression recognition (MER) has recently become a popular research topic due to its wide applications, e.g., movie rating and recognizing the neurological disorder. By virtue of deep learning techniques, the performance of MER has been significantly improved and reached unprecedented results. This paper proposes a novel architecture to mimic how the expressions are annotated. Specifically, during the annotation process in several datasets, the AU labels are first obtained with FACS, and the expression labels are then decided based on the combinations of the AU labels. Meanwhile, these AU labels describe either the eyes or mouth movements (mutually-exclusive). Following this idea, we design a dual-branch structure with a new augmentation method to separately capture the eyes and mouth features and teach the model what the general expressions should be. Moreover, to adaptively fuse the area features for different expressions, we propose Area Weighted Module to assign different weights to each region. Additionally, we set up an auxiliary task to align the AU similarity scores to help our model capture facial patterns further with AU labels. The proposed approach outperforms other state-of-the-art methods in terms of accuracy on the CASME II and SAMM datasets. Moreover, we provide a new visualization approach to show the relationship between the facial regions and AU features.
Bo-Kai Ruan, Ling Lo, Hong-Han Shuai, Wen-Huang Cheng
ACM Multimedia2
2022 Facial Chirality: From Visual Self-Reflection to Robust Facial Feature Learning
abstract
As a fundamental vision task, facial expression recognition has made substantial progress recently. However, the recognition performance often degrades significantly in real-world scenarios due to the lack of robust facial features. In this paper, we propose an effective facial feature learning method that takes the advantage of facial chirality to discover the discriminative features for facial expression recognition. Most previous studies implicitly assume that human faces are symmetric. However, our work reveals that the facial asymmetric effect can be a crucial clue. Given a face image and its reflection without additional labels, we decouple the emotion-invariant facial features from the input image pair to better capture the emotion-related facial features. Moreover, as our model aligns emotion-related features of the image pair to enhance the recognition performance, the value of precise facial landmark alignment as a pre-processing step is reconsidered in this paper. Experiments demonstrate that the learned emotion-related features outperform the state of the art methods on several facial expression recognition benchmarks as well as real-world occlusion datasets, which manifests the effectiveness and robustness of the proposed model.
Ling Lo, Hong-Han Shuai, Wen-Huang Cheng
IEEE Trans. Multim.1
2021 FashionMirror: Co-attention Feature-remapping Virtual Try-on with Sequential Template Poses
abstract
Virtual try-on tasks have drawn increased attention. Prior arts focus on tackling this task via warping clothes and fusing the information at the pixel level with the help of semantic segmentation. However, conducting semantic segmentation is time-consuming and easily causes error accumulation over time. Besides, warping the information at the pixel level instead of the feature level limits the performance (e.g., unable to generate different views) and is unstable since it directly demonstrates the results even with a misalignment. In contrast, fusing information at the feature level can be further refined by the convolution to obtain the final results. Based on these assumptions, we propose a co-attention feature-remapping framework, namely FashionMirror, that generates the try-on results according to the driven-pose sequence in two stages. In the first stage, we consider the source human image and the target try-on clothes to predict the removed mask and the try-on clothing mask, which replaces the pre-processed semantic segmentation and reduces the inference time. In the second stage, we first remove the clothes on the source human via the removed mask and warp the clothing features conditioning on the try-on clothing mask to fit the next frame human. Meanwhile, we predict the optical flows from the consecutive 2D poses and warp the source human to the next frame at the feature level. Then, we enhance the clothing features and source human features in every frame to generate realistic try-on results with spatiotemporal smoothness. Both qualitative and quantitative results show that FashionMirror outperforms the state-of-the-art virtual try-on approaches.
Chieh-Yun Chen, Ling Lo, Pin-Jui Huang, Hong-Han Shuai, Wen-Huang Cheng
ICCV2
2021 Facial Chirality: Using Self-Face Reflection to Learn Discriminative Features for Facial Expression Recognition
abstract
As a fundamental vision task, facial expression recognition has made substantial progress recently. However, the recognition performance often degrades largely in real-world scenarios due to the lack of robust facial features. In this paper, we propose a simple but effective facial feature learning method that takes the advantage of facial chirality to discover the discriminative features for facial expression recognition. Most previous studies implicitly assume that human faces are symmetric. However, our work reveals that the facial asymmetric effect can be a crucial clue. Given a face image and its reflection without additional labels, we decouple the reflection-invariant facial features from the input image pair and then demonstrate that the new features with a standard and lightweight learning model (e.g. ResNet-18) are sufficiently robust to outperform the state-of-the-art methods (e.g. SCN in CVPR 2020 and ESRs in AAAI 2020). Our experiments also show the potential of the new features for other facial vision tasks such as expression image retrieval.
Ling Lo, Hong-Han Shuai, Wen-Huang Cheng
ICME1
2020 AU-assisted Graph Attention Convolutional Network for Micro-Expression Recognition
abstract
Micro-expressions (MEs) are important clues for reflecting the real feelings of humans, and micro-expression recognition (MER) can thus be applied in various real-world applications. However, it is difficult to perceive and interpret MEs correctly. With the advance of deep learning technologies, the accuracy of micro-expression recognition is improved but still limited by the lack of large-scale datasets. In this paper, we propose a novel micro-expression recognition approach by combining Action Units (AUs) and emotion category labels. Specifically, based on facial muscle movements, we model different AUs based on relational information and integrate the AUs recognition task with MER. Besides, to overcome the shortcomings of limited and imbalanced training samples, we propose a data augmentation method that can generate nearly indistinguishable image sequences with AU intensity of real-world micro-expression images, which effectively improve the performance and are compatible with other micro-expression recognition methods. Experimental results on three mainstream micro-expression datasets, i.e., CASME II, SAMM, and SMIC, manifest that our approach outperforms other state-of-the-art methods on both single database and cross-database micro-expression recognition.
Ling Lo, Hong-Han Shuai, Wen-Huang Cheng
ACM Multimedia2
2019 Dressing for Attention: Outfit Based Fashion Popularity Prediction
abstract
Analysis of fashion trends is crucial. However, existing predictive algorithms of fashion popularity are restricted to be feasible on the coarse style level but not a finer item level. That is, they are only predictive in the future popularity of a given type of fashion styles (e.g., Rocker), but cannot be precisely down to a particular outfit look chosen by individuals. This paper thus proposes the first solution directly aimed at predicting the fine-grained fashion popularity of an outfit look by taking social media as the learning source. Particularly, a deep temporal sequence learning framework is developed and the proposed framework is evaluated on a real dataset of 380,000 street fashion images collected from the fashion website lookbook.nu. The experimental results show that our proposed framework outperforms the state-of-the-art approaches, with a relative increase of 11.51% to 27.62% (MSE metric) and 7.02% to 32.61% (CSE metric) in the prediction accuracy.
Ling Lo, Chia-Lin Liu, Rong-An Lin, Bo Wu 0018, Hong-Han Shuai, Wen-Huang Cheng
ICIP1