VLDB 2026 Research / reviewers in the wild / expert
Mingcheng Li
dblp:156/8265
· DBLP profile ↗
33ranked-venue papers
5as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 27 since 2021Artificial intelligence and machine learning · 19 · 4 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image ComprehensionabstractSatire, a form of artistic expression combining humor with implicit critique, holds significant social value by illuminating societal issues. Despite its cultural and societal significance, satire comprehension, particularly in purely visual forms, remains a challenging task for current vision-language models. This task requires not only detecting satire but also deciphering its nuanced meaning and identifying the implicated entities. Existing models often fail to effectively integrate local entity relationships with global context, leading to misinterpretation, comprehension biases, and hallucinations. To address these limitations, we propose SatireDecoder, a training-free framework designed to enhance satirical image comprehension. Our approach proposes a multi-agent system performing visual cascaded decoupling to decompose images into fine-grained local and global semantic representations. In addition, we introduce a chain-of-thought reasoning strategy guided by uncertainty analysis, which breaks down the complex satire comprehension process into sequential subtasks with minimized uncertainty. Our method significantly improves interpretive accuracy while reducing hallucinations. Experimental results validate that SatireDecoder outperforms existing baselines in comprehending visual satire, offering a promising direction for vision-language reasoning in nuanced, high-level semantic tasks. Haiwei Xue, Minghao Han, Mingcheng Li, Xiaolu Hou, Dingkang Yang, Lihua Zhang 0002, Xu Zheng 0002 |
AAAI | 4 |
| 2025 | BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene GenerationabstractWith the widespread use of virtual reality applications, 3D scene generation has become a new challenging research frontier. 3D scenes have highly complex structures and need to ensure that the output is dense, coherent, and contains all necessary structures. Many current 3D scene generation methods rely on pre-trained text-to-image diffusion models and monocular depth estimators. However, the generated scenes occupy large amounts of storage space and often lack effective regularisation methods, leading to geometric distortions. To this end, we propose BloomScene, a lightweight structured 3D Gaussian splatting for crossmodal scene generation, which creates diverse and high-quality 3D scenes from text or image inputs. Specifically, a crossmodal progressive scene generation framework is proposed to generate coherent scenes utilizing incremental point cloud reconstruction and 3D Gaussian splatting. Additionally, we propose a hierarchical depth prior-based regularization mechanism that utilizes multi-level constraints on depth accuracy and smoothness to enhance the realism and continuity of the generated scenes. Ultimately, we propose a structured context-guided compression mechanism that exploits structured hash grids to model the context of unorganized anchor attributes, which significantly eliminates structural redundancy and reduces storage overhead. Comprehensive experiments across multiple scenes demonstrate the significant potential and advantages of our framework compared with several baselines. Xiaolu Hou, Mingcheng Li, Dingkang Yang, Jiawei Chen 0012, Ziyun Qian, Jinjie Wei, Qingyao Xu, Lihua Zhang 0002 |
AAAI | 2 |
| 2025 | Debiased Multimodal Understanding for Human Language SequencesabstractHuman multimodal language understanding (MLU) is an indispensable component of expression analysis (e.g., sentiment or humor) from heterogeneous modalities, including visual postures, linguistic contents, and acoustic behaviours. Existing works invariably focus on designing sophisticated structures or fusion strategies to achieve impressive improvements. Unfortunately, they all suffer from the subject variation problem due to data distribution discrepancies among subjects. Concretely, MLU models are easily misled by distinct subjects with different expression customs and characteristics in the training data to learn subject-specific spurious correlations, limiting performance and generalizability across new subjects. Motivated by this observation, we introduce a recapitulative causal graph to formulate the MLU procedure and analyze the confounding effect of subjects. Then, we propose SuCI, a simple yet effective causal intervention module to disentangle the impact of subjects acting as unobserved confounders and achieve model training via true causal effects. As a plug-and-play component, SuCI can be widely applied to most methods that seek unbiased predictions. Comprehensive experiments on several MLU benchmarks clearly show the effectiveness of the proposed module. Zhi Xu 0010, Dingkang Yang, Mingcheng Li, Zhaoyu Chen 0001, Jiawei Chen 0012, Jinjie Wei, Lihua Zhang 0002 |
AAAI | 3 |
| 2025 | Improving Factuality in Large Language Models via Decoding-Time Hallucinatory and Truthful ComparatorsabstractDespite their remarkable capabilities, Large Language Models (LLMs) are prone to generate responses that contradict verifiable facts, i.e., unfaithful hallucination content. Existing efforts generally focus on optimizing model parameters or editing semantic representations, which compromise the internal factual knowledge of target LLMs. In addition, hallucinations typically exhibit multifaceted patterns in downstream tasks, limiting the model's holistic performance across tasks. In this paper, we propose a Comparator-driven Decoding-Time (CDT) framework to alleviate the response hallucination. Firstly, we construct hallucinatory and truthful comparators with multi-task fine-tuning samples. In this case, we present an instruction prototype-guided mixture of experts strategy to enhance the ability of the corresponding comparators to capture different hallucination or truthfulness patterns in distinct task instructions. CDT constrains next-token predictions to factuality-robust distributions by contrasting the logit differences between the target LLMs and these comparators. Systematic experiments on multiple downstream tasks show that our framework can significantly improve the model performance and response factuality. Dingkang Yang, Dongling Xiao, Jinjie Wei, Mingcheng Li, Zhaoyu Chen 0001, Ke Li 0015, Lihua Zhang 0002 |
AAAI | 4 |
| 2025 | MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image GenerationabstractDiffusion models have shown excellent performance in text-to-image generation. Nevertheless, existing methods often suffer from performance bottlenecks when handling complex prompts that involve multiple objects, characteristics, and relations. Therefore, we propose a Multi-agent Collaboration-based Compositional Diffusion (MCCD) for text-to-image generation for complex scenes. Specifically, we design a multi-agent collaboration-based scene parsing module that generates an agent system comprising multiple agents with distinct tasks, utilizing MLLMs to extract various scene elements effectively. In addition, Hierarchical Compositional diffusion utilizes a Gaussian mask and filtering to refine bounding box regions and enhance objects through region enhancement, resulting in the accurate and high-fidelity generation of complex scenes. Comprehensive experiments demonstrate that our MCCD significantly improves the performance of the baseline models in a training-free manner, providing a substantial advantage in complex scene generation. Mingcheng Li, Xiaolu Hou, Dingkang Yang, Ziyun Qian, Jiawei Chen 0012, Jinjie Wei, Qingyao Xu, Lihua Zhang 0002 |
CVPR | 1 |
| 2025 | CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report GenerationabstractAutomatic medical report generation (MRG), which possesses significant research value as it can aid radiologists in clinical diagnosis and report composition, has garnered increasing attention. Despite recent progress, generating accurate reports remains arduous due to the requirement for precise clinical comprehension and disease diagnosis inference. Furthermore, owing to the limited accessibility of medical data and the imbalanced distribution of diseases, the underrepresentation of rare diseases in training data makes large-scale medical visual language models prone to hallucinations, such as omissions or fabrications, severely undermining diagnostic performance and further intensifying the challenges for MRG in practice. In this study, to effectively mitigate hallucinations in medical report generation, we propose a chain-of-medical-thought approach (CoMT), which intends to imitate the cognitive process of human doctors by decomposing diagnostic procedures. The radiological features with different importance are structured into fine-grained medical thought chains to enhance the inferential ability during diagnosis, thereby alleviating hallucination problems and enhancing the diagnostic accuracy of MRG. Jiawei Chen 0012, Dingkang Yang, Mingcheng Li, Shunli Wang 0001, Ke Li 0015, Lihua Zhang 0002 |
ICASSP | 4 |
| 2025 | MAFD: Fine-Grained Motion Style Transfer with Adaptive Signal FusionabstractMotion style transfer allows for the swift switching of different styles within the same motion for virtual avatars, offering significant efficiency gains and enhanced motion diversity compared to traditional motion capture methods. However, many existing methods struggle with controlling fine details in complex motions, leading to models that capture only coarse-grained style characteristics. To overcome this limitation, we introduce the Motion Adaptive Fusion Diffusion (MAFD) framework, which leverages adaptive signal fusion to highlight essential style-defining features while minimizing redundant information. Moreover, current diffusion-based denoisers often fail to effectively capture the temporal relationships in motion sequences, producing rigid and fragmented stylized motions. Drawing inspiration from the Mamba model, we propose the Style Mamba Denoiser (SMD), which adopts a selection mechanism to preserve long-range dependencies and maintain temporal coherence. Extensive experiments show that our approach outperforms state-of-the-art methods in both qualitative and quantitative evaluations, achieving more refined and coherent stylized motions. Ziyun Qian, Dingkang Yang, Mingcheng Li, Dongliang Kou, Lihua Zhang 0002 |
ICASSP | 3 |
| 2025 | FSRF: Factorization-guided Semantic Recovery for Incomplete Multimodal Sentiment AnalysisabstractIn recent years, Multimodal Sentiment Analysis (MSA) has become a research hotspot that aims to utilize multimodal data for human sentiment understanding. Previous MSA studies have mainly focused on performing interaction and fusion on complete multimodal data, ignoring the problem of missing modalities in real-world applications due to occlusion, personal privacy constraints, and device malfunctions, resulting in low generalizability. To this end, we propose a Factorization-guided Semantic Recovery Framework (FSRF) to mitigate the modality missing problem in the MSA task. Specifically, we propose a de-redundant homo-heterogeneous factorization module that factorizes modality into modality-homogeneous, modality-heterogeneous, and noisy representations and design elaborate constraint paradigms for representation learning. Furthermore, we design a distribution-aligned self-distillation module that fully recovers the missing semantics by utilizing bidirectional knowledge transfer. Comprehensive experiments on two datasets indicate that FSRF has a significant performance advantage over previous methods with uncertain missing modalities. Pengjunfei Chu, Shuming Dong, Mingcheng Li |
ICME | 5 |
| 2025 | Robust Signed Distance Fields for Articulated Human Body Reconstruction via Multiresolution Hash EncodingabstractThe Signed Distance Fields (SDF) of the human body has broad applications in shape representation, collision handling, and medical image analysis, etc. However, due to the inherently high complexity of human motion, computing the SDFs of dynamic human bodies both accurately and efficiently has long been a challenging problem in computer graphics. In this paper, we demonstrate that by decomposing a widely used explicit human body model (SMPL) and modeling each component in a targeted manner, we can simultaneously realize both efficiency and accuracy. From a high level, the pipeline of the explicit model can be divided into Linear Blend Skinning (LBS) and Pose Space Deformation (PSD). By partitioning the human body into multiple parts and using the transformation matrix of each part, we apply inverse transformations to map spatial points from the posed space back to the canonical space. This eliminates the need for learning transformations and significantly reduces the difficulty of learning PSD. We observe that PSD is essentially a weighted sum of a series of fixed corrective shapes, where the only variable is the coefficient. We propose using Multiresolution Hash Encoding (MHE) to accurately capture the influence of each corrective shape on the SDF and aggregate the features in a manner similar to the explicit model. Our experiments show that our method is robust, effective, and highly efficient. Minzhe Tang, Dongliang Kou, Mingcheng Li, Lihua Zhang 0002 |
IJCNN | 3 |
| 2025 | UMSD: High Realism Motion Style Transfer via Unified Mamba-based DiffusionabstractMotion style transfer is a significant research area in computer vision, enabling the rapid switching of stylistic variations for the same motion in virtual digital humans. This dramatically enhances the richness and realism of motions, making it widely applicable in multimedia contexts such as film, gaming, and the Metaverse. However, most existing methods employ a two-stream structure, which often overlooks the intrinsic relationships between content and style motions, resulting in information loss and misalignment. Additionally, these methods struggle to capture temporal dependencies in long-range motion sequences, resulting in less natural outputs. To address these limitations, we propose a Unified Motion Style Diffusion (UMSD) Framework that simultaneously extracts features from content and style motions, achieving comprehensive information interaction. We also introduce the Motion Style Mamba (MSM) denoiser, which, for the first time in motion style transfer, leverages Mamba's powerful sequence modelling capability to produce more temporally coherent stylized motion sequences. Furthermore, we design a diffusion-based content consistency loss and a style consistency loss to ensure that the framework preserves content motion while effectively learning style motion features. Extensive experiments demonstrate that our approach outperforms State-Of-The-Art (SOTA) methods qualitatively and quantitatively, achieving more realistic and coherent motion style transfer. Ziyun Qian, Zeyu Xiao 0001, Xingliang Jin, Dingkang Yang, Mingcheng Li, Zhenyi Wu, Dongliang Kou, Peng Zhai, Lihua Zhang 0002 |
ACM Multimedia | 5 |
| 2025 | MSTDF: Motion Style Transfer Towards High Visual Fidelity Based on Dynamic FusionabstractEmotion-guided motion style transfer is a novel research direction, enabling the efficient generation of motion in various emotional styles for use in films, games, and other domains. However, existing methods primarily rely on global feature statistics for motion style transfer, neglecting local semantic structure and resulting in the degradation of motion content structure. This letter proposes a novel Motion Style Transfer based on Dynamic Fusion (MSTDF) framework, which treats content and style motion as distinct signals and employs dynamic fusion for high-fidelity motion style transfer. Additionally, to address the challenge of traditional discriminators capturing subtle motion style features, we propose the Motion Dynamic Fusion (MDF) discriminator to capture the details and fine-grained style characteristics of motion sequences, assisting the generator in producing higher-fidelity stylized motion. Finally, extensive experiments on the Xia dataset demonstrate that our method surpasses state-of-the-art methods in qualitative and quantitative comparisons. Ziyun Qian, Dingkang Yang, Mingcheng Li, Zeyu Xiao 0001, Lihua Zhang 0002 |
IEEE Signal Process. Lett. | 3 |
| 2024 | A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing ModalitiesabstractMultimodal Sentiment Analysis (MSA) has attracted widespread research attention recently. Most MSA studies are based on the assumption of modality completeness. However, many inevitable factors in real-world scenarios lead to uncertain missing modalities, which invalidate the fixed multimodal fusion approaches. To this end, we propose a Unified multimodal Missing modality self-Distillation Framework (UMDF) to handle the problem of uncertain missing modalities in MSA. Specifically, a unified self-distillation mechanism in UMDF drives a single network to automatically learn robust inherent representations from the consistent distribution of multimodal data. Moreover, we present a multi-grained crossmodal interaction module to deeply mine the complementary semantics among modalities through coarse- and fine-grained crossmodal attention. Eventually, a dynamic feature integration module is introduced to enhance the beneficial semantics in incomplete modalities while filtering the redundant information therein to obtain a refined and robust multimodal representation. Comprehensive experiments on three datasets demonstrate that our framework significantly improves MSA performance under both uncertain missing-modality and complete-modality testing conditions. Mingcheng Li, Dingkang Yang, Yuxuan Lei, Shunli Wang 0001, Shuaibing Wang, Liuzhen Su, Kun Yang 0010, Lihua Zhang 0002 |
AAAI | 1 |
| 2024 | SceneWeaver: Text-Driven Scene Generation with Geometry-aware Gaussian Splatting
Xiaolu Hou, Mingcheng Li, Jiawei Chen 0012, Dingkang Yang, Ziyun Qian, Lihua Zhang 0002 |
ACML | 2 |
| 2024 | CPR-Coach: Recognizing Composite Error Actions Based on Single-Class TrainingabstractFine- grained medical action analysis plays a vital role in improving medical skill training efficiency, but it faces the problems of data and algorithm shortage. Cardiopul-monary Resuscitation (CPR) is an essential skill in emer-gency treatment. Currently, the assessment of CPR skills mainly depends on dummies and trainers, leading to high training costs and low efficiency. For the first time, this pa-per constructs a vision-based system to complete error action recognition and skill assessment in CPR. Specifically, we define 13 types of single-error actions and 74 types of composite error actions during external cardiac compres-sion and then develop a video dataset named CPR-Coach. By taking the CPR-Coach as a benchmark, this paper in-vestigates and compares the performance of existing action recognition models based on different data modalities. To solve the unavoidable “Single-class Training & Multi-class Testing” problem, we propose a human-cognition-inspired framework named ImagineNet to improve the model's multi-error recognition performance under restricted supervision. Extensive comparison and actual deployment experiments verify the effectiveness of the framework. We hope this work could bring new inspiration to the computer vision and medical skills training communities simultaneously. The dataset and the code are publicly available on https://github.com/Shunli-Wang/CPR-Coach. Shunli Wang 0001, Shuaibing Wang, Dingkang Yang, Mingcheng Li, Haopeng Kuang, Liuzhen Su, Peng Zhai, Lihua Zhang 0002 |
CVPR | 4 |
| 2024 | Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete ModalitiesabstractMultimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However, in real-world applications, some practical factors cause uncertain modality missingness, which drastically degrades the model's performance. To this end, we propose a Correlation-decoupled Knowledge Distillation (CorrKD) framework for the MSA task under uncertain missing modalities. Specifically, we present a sample-level contrastive distillation mechanism that transfers comprehensive knowledge containing cross-sample correlations to reconstruct missing semantics. Moreover, a category-guided prototype distillation mechanism is introduced to capture cross-category correlations using category prototypes to align feature distributions and generate favorable joint representations. Eventually, we design a response-disentangled consistency distillation strategy to optimize the sentiment decision boundaries of the student network through response disentanglement and mutual information maximization. Comprehensive experiments on three datasets indicate that our framework can achieve favorable improvements compared with several baselines. Mingcheng Li, Dingkang Yang, Shuaibing Wang, Yan Wang 0068, Kun Yang 0010, Dongliang Kou, Ziyun Qian, Lihua Zhang 0002 |
CVPR | 1 |
| 2024 | Robust Emotion Recognition in Context DebiasingabstractContext-aware emotion recognition (CAER) has recently boosted the practical applications of affective computing techniques in unconstrained environments. Mainstream CAER methods invariably extract ensemble representations from diverse contexts and subject-centred characteristics to perceive the target person's emotional state. Despite advancements, the biggest challenge remains due to context bias interference. The harmful bias forces the models to rely on spurious correlations between background contexts and emotion labels in likelihood estimation, causing severe performance bottlenecks and confounding valuable context priors. In this paper, we propose a counterfactual emotion inference (CLEF) framework to address the above issue. Specifically, we first formulate a generalized causal graph to decouple the causal relationships among the variables in CAER. Following the causal graph, CLEF introduces a non-invasive context branch to capture the adverse direct effect caused by the context bias. During the inference, we eliminate the direct context effect from the total causal effect by comparing factual and counterfactual outcomes, resulting in bias mitigation and robust prediction. As a model-agnostic framework, CLEF can be readily integrated into existing methods, bringing consistent performance gains. Dingkang Yang, Kun Yang 0010, Mingcheng Li, Shunli Wang 0001, Shuaibing Wang, Lihua Zhang 0002 |
CVPR | 3 |
| 2024 | Towards Multimodal Sentiment Analysis Debiasing via Bias Purification
Dingkang Yang, Mingcheng Li, Dongling Xiao, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Peng Zhai, Ke Li 0015, Lihua Zhang 0002 |
ECCV (58) | 2 |
| 2024 | Can LLMs' Tuning Methods Work in Medical Multimodal Domain?
Jiawei Chen 0012, Dingkang Yang, Mingcheng Li, Jinjie Wei, Ziyun Qian, Lihua Zhang 0002 |
MICCAI (5) | 4 |
| 2024 | Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Models
Jiawei Chen 0012, Dingkang Yang, Mingcheng Li, Jinjie Wei, Xiaolu Hou, Lihua Zhang 0002 |
ACM Multimedia | 4 |
| 2024 | IF-Garments: Reconstructing Your Intersection-Free Multi-Layered Garments from Monocular VideosabstractReconstructing garments from monocular videos has attracted considerable attention as it provides a convenient and low-cost solution for clothing digitization. In reality, people wear clothing with countless variations and multiple layers. Existing studies attempt to extract garments from a single video. They either behave poorly in generalization due to reliance on limited clothing templates or struggle to handle the intersections of multi-layered clothing leading to the lack of physical plausibility. Besides, there are inevitable and undetectable overlaps for a single video that hinder researchers from modeling complete and intersection-free multi-layered clothing. To address the above limitations, in this paper, we propose a novel method to reconstruct multi-layered clothing from multiple monocular videos sequentially, which surpasses existing work in generalization and robustness against penetration. For each video, neural fields are employed to implicitly represent the clothed body, from which the meshes with frame-consistent structures are explicitly extracted. Next, we implement a template-free method for extracting a single garment by back-projecting the image segmentation labels of different frames onto these meshes. In this way, multiple garments can be obtained from these monocular videos and then aligned to form the whole outfit. However, intersection always occurs due to overlapping deformation in the real world and perceptual errors in monocular videos. To this end, we innovatively introduce a physics-aware module that combines neural fields with a position-based simulation framework to fine-tune the penetrating vertices of garments, ensuring robustly intersection-free. Additionally, we collect a mini dataset with fashionable garments to evaluate the quality of clothing reconstruction comprehensively. We release our code and data at https://github.com/SMY19999/IF-Garments. Qipeng Yan, Zhuoer Liang, Dongliang Kou, Dingkang Yang, Ruisheng Yuan, Mingcheng Li, Lihua Zhang 0002 |
ACM Multimedia | 8 |
| 2024 | MaskBEV: Towards A Unified Framework for BEV Detection and Map SegmentationabstractAccurate and robust multimodal multi-task perception is crucial for modern autonomous driving systems. However, current multimodal perception research follows independent paradigms designed for specific perception tasks, leading to a lack of complementary learning among tasks and decreased performance in multi-task learning (MTL) due to joint training. In this paper, we propose MaskBEV, a masked attention-based MTL paradigm that unifies 3D object detection and bird's eye view (BEV) map segmentation. MaskBEV introduces a task-agnostic Transformer decoder to process these diverse tasks, enabling MTL to be completed in a unified decoder without requiring additional design of specific task heads. To fully exploit the complementary information between BEV map segmentation and 3D object detection tasks in BEV space, we propose spatial modulation and scene-level context aggregation strategies. These strategies consider the inherent dependencies between BEV segmentation and 3D detection, naturally boosting MTL performance. Extensive experiments on nuScenes dataset show that compared with previous state-of-the-art MTL methods, MaskBEV achieves 1.3 NDS improvement in 3D object detection and 2.7 mIoU improvement in BEV map segmentation, while also demonstrating slightly leading inference speed. Xukun Zhang, Dingkang Yang, Mingcheng Li, Shunli Wang 0001, Lihua Zhang 0002 |
ACM Multimedia | 5 |
| 2024 | Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation LearningabstractMultimodal Sentiment Analysis (MSA) is an important research area that aims to understand and recognize human sentiment through multiple modalities. The complementary information provided by multimodal fusion promotes better sentiment analysis compared to utilizing only a single modality. Nevertheless, in real-world applications, many unavoidable factors may lead to situations of uncertain modality missing, thus hindering the effectiveness of multimodal modeling and degrading the model’s performance. To this end, we propose a Hierarchical Representation Learning Framework (HRLF) for the MSA task under uncertain missing modalities. Specifically, we propose a fine-grained representation factorization module that sufficiently extracts valuable sentiment information by factorizing modality into sentiment-relevant and modality-specific representations through crossmodal translation and sentiment semantic reconstruction. Moreover, a hierarchical mutual information maximization mechanism is introduced to incrementally maximize the mutual information between multi-scale representations to align and reconstruct the high-level semantics in the representations. Ultimately, we propose a hierarchical adversarial learning mechanism that further aligns and adapts the latent distribution of sentiment-relevant representations to produce robust joint multimodal representations. Comprehensive experiments on three datasets demonstrate that HRLF significantly improves MSA performance under uncertain modality missing cases. Mingcheng Li, Dingkang Yang, Yang Liu 0246, Shunli Wang 0001, Jiawei Chen 0012, Shuaibing Wang, Jinjie Wei, Qingyao Xu, Xiaolu Hou, Ziyun Qian, Dongliang Kou, Lihua Zhang 0002 |
NeurIPS | 1 |
| 2024 | PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric ApplicationsabstractDeveloping intelligent pediatric consultation systems offers promising prospects for improving diagnostic efficiency, especially in China, where healthcare resources are scarce. Despite recent advances in Large Language Models (LLMs) for Chinese medicine, their performance is sub-optimal in pediatric applications due to inadequate instruction data and vulnerable training procedures.
To address the above issues, this paper builds PedCorpus, a high-quality dataset of over 300,000 multi-task instructions from pediatric textbooks, guidelines, and knowledge graph resources to fulfil diverse diagnostic demands. Upon well-designed PedCorpus, we propose PediatricsGPT, the first Chinese pediatric LLM assistant built on a systematic and robust training pipeline.
In the continuous pre-training phase, we introduce a hybrid instruction pre-training mechanism to mitigate the internal-injected knowledge inconsistency of LLMs for medical domain adaptation. Immediately, the full-parameter Supervised Fine-Tuning (SFT) is utilized to incorporate the general medical knowledge schema into the models. After that, we devise a direct following preference optimization to enhance the generation of pediatrician-like humanistic responses. In the parameter-efficient secondary SFT phase,
a mixture of universal-specific experts strategy is presented to resolve the competency conflict between medical generalist and pediatric expertise mastery. Extensive results based on the metrics, GPT-4, and doctor evaluations on distinct downstream tasks show that PediatricsGPT consistently outperforms previous Chinese medical LLMs. The project and data will be released at https://github.com/ydk122024/PediatricsGPT. Dingkang Yang, Jinjie Wei, Dongling Xiao, Shunli Wang 0001, Mingcheng Li, Shuaibing Wang, Jiawei Chen 0012, Qingyao Xu, Ke Li 0015, Peng Zhai, Lihua Zhang 0002 |
NeurIPS | 7 |
| 2024 | CASSTIMP: Cascaded Architecture for Symptom Status Tracking with Inquiry-Aware Attention and Multi-Perception PoolingabstractSymptom status tracking poses a significant challenge due to the intricate nature of symptom identification and inference from medical doctor-patient dialogues. Numerous prior studies in this domain have relied on approaches involving multi-label classification and multi-task learning. Multi-label classification methods typically consider symptoms and statuses within a unified label space. Nevertheless, this approach frequently results in sparse predictions, eroding semantic relationships among labels and causing instability in prediction outcomes. In contrast, multi-task learning segregates symptom prediction and status prediction into separate tasks, thereby improving performance relative to conventional multi-label classification methods. Nonetheless, despite these advancements, the imbalance in training task weights persists, leading to suboptimal performance. To tackle these challenges, we employ a cascaded model structure rooted in the Question-Answering (QA) paradigm in this study. Our approach utilizes dialogue content to create context-inquiry pairs and introduces two novel modules: inquiry-aware attention and multi-perception pooling. Inquiry-aware attention enhances the contextual relationship between inquiries and dialogues, while multi-perception pooling extracts diverse semantics from the dialogue. The experimental results unequivocally demonstrate our method's efficiency, surpassing state-of-the-art techniques in symptom status tracking and indicating its superior effectiveness. Haowen Yu, Mingcheng Li, Lihua Zhang 0002 |
SMC | 2 |
| 2024 | Towards Asynchronous Multimodal Signal Interaction and Fusion via Tailored TransformersabstractThe signals from human expressions are usually multimodal, including natural language, facial gestures, and acoustic behaviors. A key challenge is how to fuse multimodal time-series signals with temporal asynchrony. To this end, we present a Transformer-driven Signal Interaction and Fusion (TSIF) approach to effectively model asynchronous multimodal signal sequences. TSIF consists of linear and cross-modal transformer modules with different duties. The linear transformer module efficiently performs the global interaction for multimodal signals, and the vital philosophy is to replace the dot product similarity with the Exponential Kernel while achieving linear complexity by a low-rank matrix decomposition. By targeting the language modality, the cross-modal transformer module aims to capture reliable element correlations among distinct signals and mitigate noise interference in audio and visual modalities. Numerous experiments on two multimodal benchmarks show that our TSIF comparably outperforms previous state-of-the-art models with lower space-time complexities. The systematic analysis also proves the effectiveness of the proposed modules. Dingkang Yang, Haopeng Kuang, Kun Yang 0010, Mingcheng Li, Lihua Zhang 0002 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic RepresentationsabstractUnderstanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial expressions, and auditory clues. Despite the impressive advancements of previous works via attention-based paradigms, the inherent temporal asynchrony and modality heterogeneity challenges remain in multimodal sequence fusion, causing adverse performance bottlenecks. To tackle these issues, we propose a Multimodal fusion approach for learning modality-Exclusive and modality-Agnostic representations (MEA) to refine multimodal features and leverage the complementarity across distinct modalities. On the one hand, MEA introduces a predictive self-attention module to capture reliable context dynamics within modalities and reinforce unique features over the modality-exclusive spaces. On the other hand, a hierarchical cross-modal attention module is designed to explore valuable element correlations among modalities over the modality-agnostic space. Meanwhile, a double-discriminator strategy is presented to ensure the production of distinct representations in an adversarial manner. Eventually, we propose a decoupled graph fusion mechanism to enhance knowledge exchange across heterogeneous modalities and learn robust multimodal representations for downstream tasks. Numerous experiments are implemented on three multimodal datasets with asynchronous sequences. Systematic analyses show the necessity of our approach. Dingkang Yang, Mingcheng Li, Linhao Qu, Kun Yang 0010, Peng Zhai, Song Wang 0002, Lihua Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Context De-Confounded Emotion RecognitionabstractContext-Aware Emotion Recognition (CAER) is a crucial and challenging task that aims to perceive the emotional states of the target person with contextual information. Recent approaches invariably focus on designing sophisticated architectures or mechanisms to extract seemingly meaningful representations from subjects and contexts. However, a long-overlooked issue is that a context bias in existing datasets leads to a significantly unbalanced distribution of emotional states among different context scenarios. Concretely, the harmful bias is a confounder that misleads existing models to learn spurious correlations based on conventional likelihood estimation, significantly limiting the models' performance. To tackle the issue, this paper provides a causality-based perspective to disentangle the models from the impact of such bias, and formulate the causalities among variables in the CAER task via a tailored causal graph. Then, we propose a Contextual Causal Intervention Module (CCIM) based on the backdoor adjustment to de-confound the confounder and exploit the true causal effect for model training. CCIM is plug-in and model-agnostic, which improves diverse state-of-the-art approaches by considerable margins. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our CCIM and the significance of causal insight. Dingkang Yang, Zhaoyu Chen 0001, Shunli Wang 0001, Mingcheng Li, Siao Liu, Zhiyan Dong, Peng Zhai, Lihua Zhang 0002 |
CVPR | 5 |
| 2023 | AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving PerceptionabstractDriver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIstive Driving pErception dataset (AIDE) that considers context information both inside and outside the vehicle in naturalistic scenarios. AIDE facilitates holistic driver monitoring through three distinctive characteristics, including multi-view settings of driver and scene, multi-modal annotations of face, body, posture, and gesture, and four pragmatic task designs for driving understanding. To thoroughly explore AIDE, we provide experimental benchmarks on three kinds of baseline frameworks via extensive methods. Moreover, two fusion strategies are introduced to give new insights into learning effective multi-stream/modal representations. We also systematically investigate the importance and rationality of the key components in AIDE and benchmarks. The project link is https://github.com/ydk122024/AIDE. Dingkang Yang, Zhi Xu 0010, Shunli Wang 0001, Mingcheng Li, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Yan Wang 0068, Jing Liu 0050, Peixuan Zhang, Peng Zhai, Lihua Zhang 0002 |
ICCV | 6 |
| 2023 | Spatio-Temporal Domain Awareness for Multi-Agent Collaborative PerceptionabstractMulti-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However, several challenges remain in achieving pragmatic information sharing in this emerging research. In this paper, we propose SCOPE, a novel collaborative perception frame-work that aggregates the spatio-temporal awareness characteristics across on-road agents in an end-to-end manner. Specifically, SCOPE has three distinct strengths: i) it considers effective semantic cues of the temporal context to enhance current representations of the target agent; ii) it aggregates perceptually critical spatial information from heterogeneous agents and overcomes localization errors via multi-scale feature interactions; iii) it integrates multi-source representations of the target agent based on their complementary contributions by an adaptive fusion paradigm. To thoroughly evaluate SCOPE, we consider both real-world and simulated scenarios of collaborative 3D object detection tasks on three datasets. Extensive experiments show the superiority of our approach and the necessity of the proposed components. The project link is https://ydk122024.github.io/SCOPE/. Kun Yang 0010, Dingkang Yang, Mingcheng Li, Yang Liu 0246, Jing Liu 0050, Hanqi Wang, Peng Sun 0007 |
ICCV | 4 |
| 2023 | HandGCAT: Occlusion-Robust 3D Hand Mesh Reconstruction from Monocular ImagesabstractWe propose a robust and accurate method for reconstructing 3D hand mesh from monocular images. This is a very challenging problem, as hands are often severely occluded by objects. Previous works often have disregarded 2D hand pose information, which contains hand prior knowledge that is strongly correlated with occluded regions. Thus, in this work, we propose a novel 3D hand mesh reconstruction network HandGCAT, that can fully exploit hand prior as compensation information to enhance occluded region features. Specifically, we designed the Knowledge-Guided Graph Convolution (KGC) module and the Cross-Attention Transformer (CAT) module. KGC extracts hand prior information from 2D hand pose by graph convolution. CAT fuses hand prior into occluded regions by considering their high correlation. Extensive experiments on popular datasets with challenging hand-object occlusions, such as HO3D v2, HO3D v3, and DexYCB demonstrate that our HandGCAT reaches state-of-the-art performance. The code is available at https://github.com/heartStrive/HandGCAT. Shuaibing Wang, Shunli Wang 0001, Dingkang Yang, Mingcheng Li, Ziyun Qian, Liuzhen Su, Lihua Zhang 0002 |
ICME | 4 |
| 2023 | Target and source modality co-reinforcement for emotion understanding from asynchronous multimodal sequences
Dingkang Yang, Yang Liu 0246, Can Huang 0002, Mingcheng Li, Kun Yang 0010, Yan Wang 0068, Peng Zhai, Lihua Zhang 0002 |
Knowl. Based Syst. | 4 |
| 2023 | Towards Robust Multimodal Sentiment Analysis Under Uncertain Signal MissingabstractMultimodal Sentiment Analysis (MSA) has attracted widespread research attention recently. Most MSA studies are based on the assumption of signal completeness. However, many inevitable factors in real applications lead to uncertain signal missing, causing significant degradation of model performance. To this end, we propose a Robust multimodal Missing Signal Framework (RMSF) to handle the problem of uncertain signal missing for MSA tasks and can be generalized to other multimodal patterns. Specifically, a hierarchical cross modal interaction module in RMSF exploits potential complementary semantics among modalities via coarse- and fine-grained cross modal attention. Furthermore, we design an adaptive feature refinement module to enhance the beneficial semantics of modalities and filter redundant features. Finally, we propose a knowledge integrated self-distillation module that enables dynamic knowledge integration and bidirectional knowledge transfer within a single network to precisely reconstruct missing semantics. Comprehensive experiments are conducted on two datasets, indicating that RMSF significantly improves MSA performance under both uncertain missing-signal and complete-signal cases. Mingcheng Li, Dingkang Yang, Lihua Zhang 0002 |
IEEE Signal Process. Lett. | 1 |
| 2022 | Emotion Recognition for Multiple Context Awareness
Dingkang Yang, Shunli Wang 0001, Yang Liu 0246, Peng Zhai, Liuzhen Su, Mingcheng Li, Lihua Zhang 0002 |
ECCV (37) | 7 |