EDBT 2026 Demo / reviewers in the wild / expert
Yingjian Li 0001
dblp:202/4968-1
· DBLP profile ↗
20ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-0653-4535ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Latent-Condensed Transformer for Efficient Long Context ModelingabstractZeng You, Yaofo Chen, Qiuwu Chen, Ying Sun, Shuhai Zhang, Yingjian Li, Yaowei Wang, Mingkui Tan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zeng You, Yaofo Chen, Qiuwu Chen, Shuhai Zhang, Yingjian Li 0001, Yaowei Wang 0001, Mingkui Tan |
ACL (1) | 6 |
| 2026 | Modality augmentation and task-aware dual-modal LoRAs for multi-task multimodal federated learning
Yushi Zeng, Haopeng Ren, Yi Cai 0001, Yingjian Li 0001, Harry Qin, Yaowei Wang 0001 |
Inf. Process. Manag. | 4 |
| 2026 | Hierarchical Multi-Criteria Representation Fusion for Robust Incomplete Multimodal Sentiment Analysis
Yijing Dai, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
IEEE Trans. Affect. Comput. | 2 |
| 2026 | Rethinking the Knowledge Gap Between Cloud and Device Models for Effective Co-AdaptationabstractBy collaboratively updating the cloud (large-scale) and device (small-scale) models, co-adaptation aims to enhance the generalization performance of device models in response to the distribution shifts in the incoming data. Existing methods often rely on low-entropy samples that are selected by thedevice modelfor co-adaptation, which ignores the differences between the predictions of the cloud and device models that are caused by the knowledge gap. As a result, some of the selected samples are redundant and contribute limited value to cloud model updating and knowledge distillation. To this end, we propose a test-time co-adaptation method by Rethinking the Knowledge Gap (RKG) between cloud and device models, which effectively updates the models by informative sample selection and targeted knowledge distillation for image-based classification tasks. Specifically, we design a sample selection module that integrates semantic prediction entropy with object structure cues to identify valuable samples, which effectively alleviates the redundancy problem. Based on these selected samples, we further construct a reweighting module that measures the prediction consistency between the two models and assigns greater emphasis to samples with larger prediction discrepancies, i.e., larger knowledge gaps, to improve knowledge distillation. Furthermore, by jointly leveraging these two modules, RKG enables efficient and effective co-adaptation, thereby achieving robust model generalization to continuously changing data in classification scenarios. Extensive experiments demonstrate that RKG outperforms state-of-the-art methods while requiring fewer uploaded samples. Yingjian Li 0001, Yushi Zeng, Dongmei Jiang, Yaowei Wang 0001, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Class-Incremental Cloud-Device Collaborative Adaptation With Contrastive Learning in Dynamic Changing EnvironmentsabstractLightweight models are often deployed on edge devices (e.g., smartphones and wearable devices) to enhance their scalability and practicability. To enhance their generalization ability in dynamic changing environments, cloud-device collaborative learning (CDCL) is proposed to transfer the generalization ability from large models on cloud servers to lightweight models deployed on devices. However, current methods mainly focus on solving the data distribution shifts for a limited number of seen classes but ignore the continual incoming new classes. Though existing class-incremental learning (CIL) methods achieve impressive performance, two major challenges arise when adapting them into the CDCL setting: 1) poor generalization of lightweight models during CIL and 2) overfitting during data-incremental learning. In this article, we explore a new problem named class-incremental CI-CDCL, and propose a contrastive prototypical network based CI-CDCL framework, aiming to improve the effectiveness of cloud-device collaboration in both class-incremental and data-incremental learning. Extensive experiments are conducted on two public datasets and the experimental results can evaluate the effectiveness of our proposed model. Yushi Zeng, Haopeng Ren, Yi Cai 0001, Yingjian Li 0001, Yaowei Wang 0001, Qing Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | A coarse-to-fine registration network based on affine transformation and multi-scale pyramid
Dongming Li 0002, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
Expert Syst. Appl. | 2 |
| 2024 | Multi-modal graph context extraction and consensus-aware learning for emotion recognition in conversationabstractMulti-modal emotion recognition in conversation is challenging because of the difficulty to jointly leverage the information from heterogeneous text, acoustic, and visual modalities . Recent context-aware methods usually design a graph structure to model dependencies of utterances and speakers, or integrate Multi-modal information. However, they typically lack a sufficient extraction of unimodal context, and rarely explore the emotion consensus prototypes among different samples with the same label. For solving these problems, in this paper, we propose a Graph Context extraction and Consensus-aware Learning (GCCL) framework to excavate context-sensitive fusion features and simulate the emotion evocation process during the emotion consensus learning. Specifically, GCCL contains a well-designed graph-based module to capture speaker, temporal and modality dependencies and integrate information from different modalities. Then, we design an emotion consensus learning unit to mine the most typical feature of each category in each modality. A speaker-guided contrastive learning loss is further proposed to guarantee the diversity between different individuals and the semantic consistency between distinct modalities. Moreover, we construct a consensus-aware unit with an attention-based memory mechanism to preserve semantic correlations among different samples on the category-level. Extensive experimental results on two conversational datasets demonstrate that the proposed GCCL outperforms the state-of-art methods. Code is available at https://github.com/gityider/GCCL . Yijing Dai, Jinxing Li 0003, Yingjian Li 0001, Guangming Lu 0002 |
Knowl. Based Syst. | 3 |
| 2024 | Multimodal Decoupled Distillation Graph Neural Network for Emotion Recognition in ConversationabstractGraph Neural Networks (GNNs) have attracted increasing attentions for multimodal Emotion Recognition in Conversation (ERC) due to their good performance in contextual understanding. However, most existing GNN-based methods suffer from two challenges: 1) How to explore and propagate appropriate information in a conversational graph. Typical GNNs in ERC neglect to mine the emotion commonality and discrepancy in the local neighborhood, leading to learn similar embbedings for connected nodes. However, the embeddings of these connected nodes are supposed to be distinguishable as they belong to different speakers with different emotions. 2) Most existing works apply simple concatenation or co-occurrence prior for modality combination, failing to fully capture the emotional information of multiple modalities in relationship modeling. In this paper, we propose a multimodal Decoupled Distillation Graph Neural Network (D2GNN) to address the above challenges. Specifically, D2GNN decouples the input features into emotion-aware and emotion-agnostic ones on the emotion category-level, aiming to capture emotion commonality and implicit emotion information, respectively. Moreover, we design a new message passing mechanism to separately propagate emotion-aware and -agnostic knowledge between nodes according to speaker dependency in two GNN-based modules, exploring the correlations of utterances and alleviating the similarities of embeddings. Furthermore, a multimodal distillation unit is performed to obtain the distinguishable embeddings by aggregating unimodal decoupled features. Experimental results on two ERC benchmarks demonstrate the superiority of the proposed model. Code is available at https://github.com/gityider/D2GNN. Yijing Dai, Yingjian Li 0001, Dongpeng Chen, Jinxing Li 0003, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Learning With Noisy Labels by Semantic and Feature Space CollaborationabstractLearning with noisy labels has become more and more popular because of the expensive costs of collecting high-quality labels. To avoid the decrease in model performance caused by incorrect annotations, some existing methods try to select reliable samples based on the local structure of nearest neighbors in the feature space. However, the information from local neighbors is unreliable when encountering extremely noisy cases, and selecting samples only using the feature space may result in clear noise accumulation. To this end, we propose a Dual-Space Collaborative Learning (DSCL) framework to boost classification accuracy by jointly using the complementarity information from both semantic and feature spaces. Specifically, a collaborative selection module is designed by constructing a set of global prototypes and high-confidence semantic predictions, which enhances the robustness of the sample selection process. Moreover, a collaborative regularization module is constructed by the bidirectional adjustment between the semantic and feature spaces, which effectively alleviates the noise accumulation issue caused by sample selection bias in a single space. By simultaneously utilizing the two modules, our method improves the accuracy of sample selection and mitigates the degradation caused by noisy labels. Extensive experimental results indicate the superior performance of DSCL compared with various baselines. Yingjian Li 0001, Zheng Zhang 0006, Lei Zhu 0002, Yong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Differential Enhanced Siamese Segmentation Network for Printed Label Defect DetectionabstractMany vision-based methods have been widely used to detect defects in industrial printed labels. However, most of them still face challenges of detecting unseen defects, low-contrast defects, and false detections caused by artifacts. To address these problems, we propose a differential enhanced Siamese segmentation network (DESS-Net) for defect detection. This method is based on Siamese similarity comparison which has a better generalization ability for unseen defects. Moreover, we introduce the differential feature enhancement (DFE) modules into the Siamese network to focus on multiple differential feature information which contributes to identifying defects and reducing false detections caused by artifacts. Additionally, a multi-scale feature fusion (MFF) module is further designed to fuse multiple low-level differential features, which is conducive to recovering fine boundaries of low-contrast defects. Experimental results show that our DESS-Net outperforms other compared methods. Dongming Li 0002, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
ICIP | 2 |
| 2023 | Facial Expression Recognition in the Wild Using Multi-Level Features and Attention MechanismsabstractLearning discriminative features is of vital importance for automatic facial expression recognition (FER) in the wild. In this article, we propose a novel Slide-Patch and Whole-Face Attention model with SE blocks (SPWFA-SE), which jointly perceives the discriminative locality characteristics and informative global features of the face for effective FER. Specifically, the well-designed slide patches are proposed to extract local features. Different from the existing methods, our slide patches not only can maintain the information at the edge area of patches, but also do not need to detect facial landmarks. Moreover, to make the model adaptively focus on the distinguishable regions, an attention module is proposed in the patch level to learn the weight of each patch. Furthermore, squeeze-and-excitation blocks are explored in the channel level to learn the weight of each channel. As such, the proposed multi-level feature extraction and attention mechanisms can enhance the representative ability of the learned features. Extensive experiments on five challenging datasets demonstrate that our method can achieve state-of-the-art performance. Cross database experiments on another three databases show the superior generalization performance of our model. Furthermore, complexity analysis results show that our model contains fewer parameters with fast training advantages than other competing models. Yingjian Li 0001, Guangming Lu 0002, Jinxing Li 0003, Zheng Zhang 0006, David Zhang 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | Cross-Domain Facial Expression Recognition via Contrastive Warm up and Complexity-Aware Self-TrainingabstractUnsupervised cross-domain Facial Expression Recognition (FER) aims to transfer the knowledge from a labeled source domain to an unlabeled target domain. Existing methods strive to reduce the discrepancy between source and target domain, but cannot effectively explore the abundant semantic information of the target domain due to the absence of target labels. To this end, we propose a novel framework via Contrastive Warm up and Complexity-aware Self-Training (namely CWCST), which facilitates source knowledge transfer and target semantic learning jointly. Specifically, we formulate a contrastive warm up strategy via features, momentum features, and learnable category centers to concurrently learn discriminative representations and narrow the domain gap, which benefits domain adaptation by generating more accurate target pseudo labels. Moreover, to deal with the inevitable noise in pseudo labels, we develop complexity-aware self-training with a label selection module based on prediction entropy, which iteratively generates pseudo labels and adaptively chooses the reliable ones for training, ultimately yielding effective target semantics exploration. Furthermore, by jointly using the two mentioned components, our framework enables to effectively utilize the source knowledge and target semantic information by source-target co- training. In addition, our framework can be easily incorporated into other baselines with consistent performance improvements. Extensive experimental results on seven databases show the superior performance of the proposed method against various baselines. Yingjian Li 0001, Jiaxing Huang 0001, Shijian Lu, Zheng Zhang 0006, Guangming Lu 0002 |
IEEE Trans. Image Process. | 1 |
| 2023 | Deep Margin-Sensitive Representation Learning for Cross-Domain Facial Expression RecognitionabstractCross-domain Facial Expression Recognition (FER) aims to safely transfer the learned knowledge from labeled source data to unlabeled target data, which is challenging due to the subtle difference between various expressions and the large discrepancy between domains. Existing methods mainly focus on reducing the domain shift for transferable features but fail to learn discriminative representations for recognizing facial expression, which may result in negative transfer under cross-domain settings. To this end, we propose a novel Deep Margin-Sensitive Representation Learning (DMSRL) framework, which can extract multi-level discriminative features during sematic-aware domain adaptation. Specifically, we design a semantic metric learning module based on the category prior of source data and generated pseudo labels of target data, which can facilitate discriminative intra-domain representation learning and transferable inter-domain knowledge discovery by enlarging the category margin. Moreover, we develop a mutual information minimization module by simultaneously distilling the domain-invariant components and eliminating the domain-sensitive ones, which benefits discriminative transferable feature learning by generating accurate pseudo target labels. Furthermore, instead of only utilizing the global features, we formulate a multi-level feature extracting module to concurrently get the local ones, which contain detailed information to distinguish the small changes among different expressions. These modules are jointly utilized in our DMSRL in an end-to-end manner to ensure the positive transfer of source knowledge. Extensive experimental results on seven databases demonstrate that our DMSRL can achieve superior performance against state-of-the-art baselines. Yingjian Li 0001, Zheng Zhang 0006, Bingzhi Chen, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | PPR-Net: Patch-Based Multi-scale Pyramid Registration Network for Defect Detection of Printed Label
Dongming Li 0002, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
ACCV (2) | 2 |
| 2022 | Multi-Label Chest X-Ray Image Classification via Semantic Similarity Graph EmbeddingabstractAutomated multi-label chest X-ray (CXR) image classification has recently made significant progress in clinical diagnosis based on the advanced deep learning techniques. However, most existing methods mainly focus on analyzing locality visual cues from a single image but fail to leverage the underlying explicit correlations among different images for precise disease diagnosis. By contrast, an experienced radiologist expertizes in transferring knowledge from previous tasks to diagnose the present radiograph. To enable the machine like a radiologist, this paper proposes a novel Semantic Similarity Graph Embedding (SSGE) framework, which explicitly explores the semantic similarities among images to optimize the visual feature embedding for improving the performance of multi-label CXR images classification. Specifically, the proposed SSGE framework contains three main components: the image feature embedding (IFE) module, similarity graph construction (SGC) module, and semantic similarity learning (SSL) module. To realize interactive teaching and learning between visual and semantic information, the proposed SSGE framework is built on the “Teacher-Student” (semantic-visual) learning mechanism. With the guidance and supervision of the cross-image similarity graph generated by the SGC module, the SSL module leverages Graph Convolutional Network (GCN) to adaptively recalibrate the multi-image feature representations extracted from the IFE module, which guarantees their semantic consistency. Furthermore, we propose a novel re-weighting strategy to learn a more optimal semantic-similarity graph for the information propagation of the GCN layers. Extensive experiments on two benchmark datasets demonstrate the effectiveness of the proposed method in comparison with some state-of-the-art baselines. Bingzhi Chen, Zheng Zhang 0006, Yingjian Li 0001, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Self-Supervised Exclusive-Inclusive Interactive Learning for Multi-Label Facial Expression Recognition in the WildabstractFacial Expression Recognition (FER) is a long-standing but challenging research problem in computer vision. Existing approaches mainly focus on single-label emotional prediction, which cannot handle the complex multi-label FER task because of the coupling behavior of multiple emotions on a single facial image. To this end, in this paper, we propose a novel Self-supervised Exclusive-Inclusive Interactive Learning (SEIIL) method to facilitate discriminative multi-label FER in the wild, which can effectively handle the coupled multiple sentiments with limited unconstrained training data. Specifically, we construct an emotion disentangling module to capture the inclusive and exclusive characteristics of facial expressions, which can decouple the compound numerous emotions on an image. Moreover, an adaptively-weighted ensemble technique is conceived to aggregate category-level latent exclusive embeddings, and then a conditional adversarial interactive learning module is designed to fully leverage the complementary between the inclusive and formulated latent representations. Furthermore, to tackle the insufficient data for training, we introduce a self-supervised learning strategy to augment the amount and diversity of facial images, which can endow the model with advanced generalization ability. Under this strategy, the proposed two modules can be concurrently utilized in our SEIIL to jointly handle the coupled emotions and alleviate the overfitting problem. Extensive experimental results on six databases illustrate the superb performance of our method against state-of-the-art baselines. Yingjian Li 0001, Yingnan Gao, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Learning Informative and Discriminative Features for Facial Expression Recognition in the WildabstractThe informativeness and discriminativeness of features collaboratively ensure high-accuracy Facial Expression Recognition (FER) in the wild. Most of existing methods use the single-path deep convolutional neural network with softmax loss for basic FER, while they cannot deal with the challenging situations of the compound FER in the wild, because they fail to learn informative and discriminative features in a targeted manner. To this end, we present an Informative and Discriminative Feature Learning (IDFL) framework that consists of two key components: the Multi-Path Attention Convolutional Neural Network (MPACNN) and Balanced Separate loss (BS loss), for both basic and compound high-accuracy FER in the wild. Specifically, MPACNN leverages different paths to learn diverse features. These features are then adaptively fused into informative ones via an attention module, such that the model can adequately capture detailed information for both basic and compound FER. The BS loss maximizes the inter-class distance of features and minimizes the intra-class one. In this way, the features are discriminative enough for high-accuracy FER in the wild. Particularly, the BS loss is invoked as the objective function of MPACNN, so the model can learn informative and discriminative features at the same time, yielding better performance. Seven databases are utilized to evaluate the proposed method, and the results demonstrate that our method achieves state-of-the-art performance on both basic and compound expressions with good generalization ability. Moreover, our model contains fewer parameters and can be trained faster than other related models. Yingjian Li 0001, Yao Lu 0008, Bingzhi Chen, Zheng Zhang 0006, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | JDMAN: Joint Discriminative and Mutual Adaptation Networks for Cross-Domain Facial Expression RecognitionabstractCross-domain Facial Expression Recognition (FER) is challenging due to the difficulty of concurrently handling the domain shift and semantic gap during domain adaptation. Existing methods mainly focus on reducing the domain discrepancy for transferable features but fail to decrease the semantic one, which may result in negative transfer. To this end, we propose Joint Discriminative and Mutual Adaptation Networks (JDMAN), which collaboratively bridge the domain shift and semantic gap by domain- and category-level co-adaptation based on mutual information and discriminative metric learning techniques. Specifically, we design a mutual information minimization module for domain-level adaptation, which narrows the domain shift by simultaneously distilling the domain-invariant components and eliminating the untransferable ones lying in different domains. Moreover, we propose a semantic metric learning module for category-level adaptation, which can close the semantic discrepancy during discriminative intra-domain representation learning and transferable inter-domain knowledge discovery. These two modules are jointly leveraged in our JDMAN to safely transfer the source knowledge to target data in an end-to-end manner. Extensive experimental results on six databases show that our method achieves state-of-the-art performance. The code of our JDMAN is available at https://github.com/YingjianLi/JDMAN. Yingjian Li 0001, Yingnan Gao, Bingzhi Chen, Zheng Zhang 0006, Lei Zhu 0002, Guangming Lu 0002 |
ACM Multimedia | 1 |
| 2021 | Deep Active Context Estimation for Automated COVID-19 DiagnosisabstractMany studies on automated COVID-19 diagnosis have advanced rapidly with the increasing availability of large-scale CT annotated datasets. Inevitably, there are still a large number of unlabeled CT slices in the existing data sources since it requires considerable consuming labor efforts. Notably, cinical experience indicates that the neighboring CT slices may present similar symptoms and signs. Inspired by such wisdom, we propose DACE, a novel CNN-based deep active context estimation framework, which leverages the unlabeled neighbors to progressively learn more robust feature representations and generate a well-performed classifier for COVID-19 diagnosis. Specifically, the backbone of the proposed DACE framework is constructed by a well-designed Long-Short Hierarchical Attention Network (LSHAN), which effectively incorporates two complementary attention mechanisms, i.e., short-range channel interactions (SCI) module and long-range spatial dependencies (LSD) module, to learn the most discriminative features from CT slices. To make full use of such available data, we design an efficient context estimation criterion to carefully assign the additional labels to these neighbors. Benefiting from two complementary types of informative annotations from -nearest neighbors, i.e., the majority of high-confidence samples with pseudo labels and the minority of low-confidence samples with hand-annotated labels, the proposed LSHAN can be fine-tuned and optimized in an incremental learning manner. Extensive experiments on the Clean-CC-CCII dataset demonstrate the superior performance of our method compared with the state-of-the-art baselines. Bingzhi Chen, Yishu Liu 0001, Zheng Zhang 0006, Yingjian Li 0001, Zhao Zhang 0001, Guangming Lu 0002, Hongbing Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | Separate Loss for Basic and Compound Facial Expression Recognition in the WildabstractIn the past few years, facial expression recognition has made great progress because of the development of convolutional neural networks. However, the features learned only using the softmax loss are not discriminative enough for highly accurate facial expression recognition in the wild, especially for the compound facial expression recognition. To enhance the discriminative power of the learned features, we propose the separate loss for both basic and compound facial expression recognition in the wild in this paper. Such loss maximizes intra-class similarity while minimizing the similarity between different classes. The qualitative and quantitative analysis shows that the features learned using such loss function are characterized by intra-class compactness and inter-class separation. Experiments are performed on two databases in the wild and the proposed method achieves state-of-the-art results on both basic and compound expressions. Furthermore, another two databases are used to perform cross database experiments to show the generalization ability of our method. Yingjian Li 0001, Yao Lu 0008, Jinxing Li 0003, Guangming Lu 0002 |
ACML | 1 |