Kehua Su

dblp:78/8499 · DBLP profile ↗
← Back
26ranked-venue papers
7as first author
15since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 EMOE: Modality-Specific Enhanced Dynamic Emotion Experts
abstract
Multimodal Emotion Recognition (MER) aims to predict human emotions by leveraging multiple modalities, such as vision, acoustics, and language. However, due to the heterogeneity of these modalities, MER faces two key challenges: modality balance dilemma and modality specialization disappearance. Existing methods often overlook the varying importance of modalities across samples in tackling the modality balance dilemma. Moreover, mainstream decoupling methods, while preserving modality-specific information, often neglect the predictive capability of unimodal data. To address these, we propose a novel model, Modality-Specific Enhanced Dynamic Emotion Experts (EMOE), consisting of: (1) Mixture of Modality Experts for dynamically adjusting modality importance based on sample features, and (2) Unimodal Distillation to retain single-modality predictive ability within fused features. EMOE enables adaptive fusion by learning a unique modality weight distribution for each sample, enhancing multi-modal predictions with single-modality predictions to balance invariant and specific features in emotion recognition. Experimental results on benchmark datasets show that EMOE achieves superior or comparable performance to state-of-the-art methods. Additionally, we extend EMOE to Multimodal Intent Recognition (MIR), further demonstrating its effectiveness and versatility.
Yiyang Fang, Wenke Huang 0003, Guancheng Wan, Kehua Su, Mang Ye
CVPR4
2025 Catch Your Emotion: Sharpening Emotion Perception in Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs) have achieved impressive progress in tasks such as visual question answering and visual understanding, but they still face significant challenges in emotional reasoning. Current methods to enhance emotional understanding typically rely on fine-tuning or manual annotations, which are resource-intensive and limit scalability. In this work, we focus on improving the ability of MLLMs to capture emotions during the inference phase. Specifically, MLLMs encounter two main issues: they struggle to distinguish between semantically similar emotions, leading to misclassification, and they are overwhelmed by redundant or irrelevant visual information, which distracts from key emotional cues. To address these, we propose Sharpening Emotion Perception in MLLMs (SEPM), which incorporates a Confidence-Guided Coarse-to-Fine Inference framework to refine emotion classification by guiding the model through simpler tasks. Additionally, SEPM employs Focus-on-Emotion Visual Augmentation to reduce visual redundancy by directing the attention of models to relevant emotional cues in images. Experimental results demonstrate that SEPM significantly improves MLLM performance on emotion-related tasks, providing a resource-efficient and scalable solution for emotion recognition.
Yiyang Fang, Jian Liang 0001, Wenke Huang 0003, He Li 0054, Kehua Su, Mang Ye
ICML5
2025 Anatomy-Conserving Unpaired CBCT-to-CT Translation via Schrödinger Bridge
Song Ouyang, Yong Luo 0002, Kehua Su, Zhiwen Liang, Bo Du 0001
MICCAI (4)5
2025 M3Site: multiclass multimodal learning for protein active site identification and classification
abstract
Accurately identifying and classifying protein active sites is crucial for understanding protein mechanisms, drug design, and synthetic biology. Current methods often rely on binary classification and single-modal data, limiting their scope. To address these limitations, we propose M$^{3}$Site, a multimodal framework that integrates protein sequence embeddings, structural graph representations, and functional text annotations for residue-level, multiclass active site prediction. Built upon a curated dataset of 25 883 proteins sourced from UniProt and AlphaFold2, M$^{3}$Site leverages pretrained protein language models, equivariant graph neural networks, and biomedical language models for feature extraction. The function informed cross-attention module enables cross-modal feature fusion, while the adaptive weighted fusion mechanism balances modality contributions. A compound loss function tackles class imbalance, ensuring robust performance. Experimental results show M$^{3}$Site significantly outperforms existing models, and an interactive application has been developed to enhance its practical utility for predictions and visualizations. The dataset, source code for experiments, and interactive application are publicly available at https://github.com/Gift-OYS/M3Site.
Song Ouyang, Yong Luo 0002, Huiyu Cai, Kehua Su, Na Zhan, Huangxuan Zhao, Tailang Yin, Dongjing Shan
Briefings Bioinform.4
2024 A Distributional Reinforcement Learning-Based Strategy for Pod Scheduling in Satellite Clusters
abstract
An effective Pod scheduling strategy is crucial for maintaining optimal performance and resource utilization in distributed satellite clusters. The inherent complexity of the satellite cluster environment, combined with the sequential decision-making nature of Pod scheduling, makes it challenging to develop a strategy that meets multiple optimization objectives. Previous studies addressing multi-objective Pod scheduling often rely on DQN-based algorithms that model cumulative returns as expectation values$(V(s)\ \mathbf{or}\ Q(s,\ a.))$leading to significant loss of detailed distributional information. Furthermore, traditional DQN methods struggle with the dynamic variability of multi-objective reward functions, making weighted summation an unstable approach for achieving load balance in satellite clusters. To overcome these limitations, we propose a novel Pod scheduling strategy based on distributional reinforcement learning (DRL), which models the distribution$Z(s,\ a)$of cumulative returns, thereby preserving detailed distributional information. To enhance learning in satellite cluster environments, we employ multi-dimensional reward functions capturing the joint distributions of source-specific rewards. Our strategy effectively integrates optimization objectives including CPU, memory, inter-satellite links, and load balancing, ensuring robust and efficient service quality. Experimental results demonstrate that our proposed strategy significantly improves scheduling success rates, resource utilization efficiency, and load balancing in complex satellite cluster environments, outperforming traditional scheduling algorithms and DQN-based methods.
Kehua Su, Ying Tao
MSN2
2024 MMSite: A Multi-modal Framework for the Identification of Active Sites in Proteins
abstract
The accurate identification of active sites in proteins is essential for the advancement of life sciences and pharmaceutical development, as these sites are of critical importance for enzyme activity and drug design. Recent advancements in protein language models (PLMs), trained on extensive datasets of amino acid sequences, have significantly improved our understanding of proteins. However, compared to the abundant protein sequence data, functional annotations, especially precise per-residue annotations, are scarce, which limits the performance of PLMs. On the other hand, textual descriptions of proteins, which could be annotated by human experts or a pretrained protein sequence-to-text model, provide meaningful context that could assist in the functional annotations, such as the localization of active sites. This motivates us to construct a $\textbf{ProT}$ein-$\textbf{A}$ttribute text $\textbf{D}$ataset ($\textbf{ProTAD}$), comprising over 570,000 pairs of protein sequences and multi-attribute textual descriptions. Based on this dataset, we propose $\textbf{MMSite}$, a multi-modal framework that improves the performance of PLMs to identify active sites by leveraging biomedical language models (BLMs). In particular, we incorporate manual prompting and design a MACross module to deal with the multi-attribute characteristics of textual descriptions. MMSite is a two-stage ("First Align, Then Fuse") framework: first aligns the textual modality with the sequential modality through soft-label alignment, and then identifies active sites via multi-modal fusion. Experimental results demonstrate that MMSite achieves state-of-the-art performance compared to existing protein representation learning methods. The dataset and code implementation are available at https://github.com/Gift-OYS/MMSite.
Song Ouyang, Huiyu Cai, Yong Luo 0002, Kehua Su, Lefei Zhang, Bo Du 0001
NeurIPS4
2024 MICCF: A Mutual Information Constrained Clustering Framework for Learning Clustering-Oriented Feature Representations
abstract
Deep clustering is a crucial task in machine learning and data mining that focuses on acquiring feature representations conducive to clustering. Previous research relies on self-supervised representation learning for general feature representations, such features may not be optimally suited for downstream clustering tasks. In this article, we introduce MICCF, a framework designed to bridge this gap and enhance clustering performance. MICCF enhances feature representations by combining mutual information constraints at different levels and employs an auxiliary alignment mutual information module for learning clustering-oriented features. To be specific, we propose a dual mutual information constraints module, incorporating minimal mutual information constraints at the feature level and maximal mutual information constraints at the instance level. This reduction in feature redundancy encourages the neural network to extract more discriminative features, while maximization ensures more unbiased and robust representations. To obtain clustering-oriented representations, the auxiliary alignment mutual information module utilizes pseudo-labels to maximize mutual information through a multi-classifier network, aligning features with the clustering task. The main network and the auxiliary module work in synergy to jointly optimize feature representations that are well-suited for the clustering task. We validate the effectiveness of our method through extensive experiments on six benchmark datasets. The results indicate that our method performs well in most scenarios, particularly on fine-grained datasets, where our approach effectively distinguishes subtle differences between closely related categories. Notably, our approach achieved a remarkable accuracy of 96.4% on the ImageNet-10 dataset, surpassing other comparison methods. The code is available at https://github.com/Li-Hyn/MICCF.git .
Hongyu Li 0004, Lefei Zhang, Kehua Su, Wei Yu 0009
ACM Trans. Knowl. Discov. Data3
2024 Textual Enhanced Adaptive Meta-Fusion for Few-Shot Visual Recognition
abstract
Few-shot learning (FSL) is a challenging task that aims to train a classifier to recognize novel categories, where only a few annotated examples are available in each category. Recently, many FSL approaches have been proposed based on the meta-learning paradigm, which attempts to learn transferable knowledge from similar tasks by designing a meta-learner. However, most of these approaches only exploit the information from visual modality and do not utilize ones from additional modalities (e.g., textual description). Since the labeled examples in FSL are limited, increasing the information on the examples is a probable solution to improve the classification performance. This motivates us to propose a novel meta-learning method, termed textual enhanced adaptive meta-fusion FSL (TAMF-FSL), which leverages both the visual information from the visual image and semantic information from language supervision. Specifically, TAMF-FSL exploits the semantic information of textual description to improve the visual-based models. We first employ a text encoder to learn the semantic features of each visual category, and then design a modality alignment module and meta-fusion module to align and fuse the visual and semantic features for final prediction. Extensive experiments show that the proposed method outperforms many recent or competitive FSL counterparts on two popular datasets.
Mengya Han, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Kehua Su, Bo Du 0001
IEEE Trans. Multim.5
2023 Fine-Grained Position Helps Memorizing More, a Novel Music Compound Transformer Model with Feature Interaction Fusion
abstract
Due to the particularity of the simultaneous occurrence of multiple events in music sequences, compound Transformer is proposed to deal with the challenge of long sequences. However, there are two deficiencies in the compound Transformer. First, since the order of events is more important for music than natural language, the information provided by the original absolute position embedding is not precise enough. Second, there is an important correlation between the tokens in the compound word, which is ignored by the current compound Transformer. Therefore, in this work, we propose an improved compound Transformer model for music understanding. Specifically, we propose an attribute embedding fusion module and a novel position encoding scheme with absolute-relative consideration. In the attribute embedding fusion module, different attributes are fused through feature permutation by using a multi-head self-attention mechanism in order to capture rich interactions between attributes. In the novel position encoding scheme, we propose RoAR position encoding, which realizes rotational absolute position encoding, relative position encoding, and absolute-relative position interactive encoding, providing clear and rich orders for musical events. Empirical study on four typical music understanding tasks shows that our attribute fusion approach and RoAR position encoding brings large performance gains. In addition, we further investigate the impact of masked language modeling and casual language modeling pre-training on music understanding.
Zuchao Li, Ruhan Gong, Yineng Chen, Kehua Su
AAAI4
2023 Dual Mutual Information Constraints for Discriminative Clustering
abstract
Deep clustering is a fundamental task in machine learning and data mining that aims at learning clustering-oriented feature representations. In previous studies, most of deep clustering methods follow the idea of self-supervised representation learning by maximizing the consistency of all similar instance pairs while ignoring the effect of feature redundancy on clustering performance. In this paper, to address the above issue, we design a dual mutual information constrained clustering method named DMICC which is based on deep contrastive clustering architecture, in which the dual mutual information constraints are particularly employed with solid theoretical guarantees and experimental validations. Specifically, at the feature level, we reduce the redundancy among features by minimizing the mutual information across all the dimensionalities to encourage the neural network to extract more discriminative features. At the instance level, we maximize the mutual information of the similar instance pairs to obtain more unbiased and robust representations. The dual mutual information constraints happen simultaneously and thus complement each other to jointly optimize better features that are suitable for the clustering task. We also prove that our adopted mutual information constraints are superior in feature extraction, and the proposed dual mutual information constraints are clearly bounded and thus solvable. Extensive experiments on five benchmark datasets show that our proposed approach outperforms most other clustering algorithms. The code is available at https://github.com/Li-Hyn/DMICC.
Hongyu Li 0004, Lefei Zhang, Kehua Su
AAAI3
2023 FedABC: Targeting Fair Competition in Personalized Federated Learning
abstract
Federated learning aims to collaboratively train models without accessing their client's local private data. The data may be Non-IID for different clients and thus resulting in poor performance. Recently, personalized federated learning (PFL) has achieved great success in handling Non-IID data by enforcing regularization in local optimization or improving the model aggregation scheme on the server. However, most of the PFL approaches do not take into account the unfair competition issue caused by the imbalanced data distribution and lack of positive samples for some classes in each client. To address this issue, we propose a novel and generic PFL framework termed Federated Averaging via Binary Classification, dubbed FedABC. In particular, we adopt the ``one-vs-all'' training strategy in each client to alleviate the unfair competition between classes by constructing a personalized binary classification problem for each class. This may aggravate the class imbalance challenge and thus a novel personalized binary classification loss that incorporates both the under-sampling and hard sample mining strategies is designed. Extensive experiments are conducted on two popular datasets under different settings, and the results demonstrate that our FedABC can significantly outperform the existing counterparts.
Dui Wang, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Kehua Su, Yonggang Wen 0001, Dacheng Tao
AAAI5
2023 Improving Heterogeneous Model Reuse by Density Estimation
abstract
This paper studies multiparty learning, aiming to learn a model using the private data of different participants. Model reuse is a promising solution for multiparty learning, assuming that a local model has been trained for each party. Considering the potential sample selection bias among different parties, some heterogeneous model reuse approaches have been developed. However, although pre-trained local classifiers are utilized in these approaches, the characteristics of the local data are not well exploited. This motivates us to estimate the density of local data and design an auxiliary model together with the local classifiers for reuse. To address the scenarios where some local models are not well pre-trained, we further design a multiparty cross-entropy loss for calibration. Upon existing works, we address a challenging problem of heterogeneous model reuse from a decision theory perspective and take advantage of recent advances in density estimation. Experimental results on both synthetic and benchmark data demonstrate the superiority of the proposed method.
Anke Tang, Yong Luo 0002, Han Hu 0003, Fengxiang He, Kehua Su, Bo Du 0001, Yixin Chen 0001, Dacheng Tao
IJCAI5
2023 Composed Image Retrieval via Cross Relation Network With Hierarchical Aggregation Transformer
abstract
Composing Text and Image to Image Retrieval (CTI-IR) aims at finding the target image, which matches the query image visually along with the query text semantically. However, existing works ignore the fact that the reference text usually serves multiple functions, e.g., modification and auxiliary. To address this issue, we put forth a unified solution, namely Hierarchical Aggregation Transformer incorporated with Cross Relation Network (CRN). CRN unifies modification and relevance manner in a single framework. This configuration shows broader applicability, enabling us to model both modification and auxiliary text or their combination in triplet relationships simultaneously. Specifically, CRN includes: 1) Cross Relation Network comprehensively captures the relationships of various composed retrieval scenarios caused by two different query text types, allowing a unified retrieval model to designate adaptive combination strategies for flexible applicability; 2) Hierarchical Aggregation Transformer aggregates top-down features with Multi-layer Perceptron (MLP) to overcome the limitations of edge information loss in a window-based multi-stage Transformer. Extensive experiments demonstrate the superiority of the proposed CRN over all three fashion-domain datasets. Code is available at github.com/yan9qu/crn.
Qu Yang, Mang Ye, Zhaohui Cai, Kehua Su, Bo Du 0001
IEEE Trans. Image Process.4
2023 Cross-Modality Pyramid Alignment for Visual Intention Understanding
abstract
Visual intention understanding is the task of exploring the potential and underlying meaning expressed in images. Simply modeling the objects or backgrounds within the image content leads to unavoidable comprehension bias. To alleviate this problem, this paper proposes a Cross-modality Pyramid Alignment with Dynamic optimization (CPAD) to enhance the global understanding of visual intention with hierarchical modeling. The core idea is to exploit the hierarchical relationship between visual content and textual intention labels. For visual hierarchy, we formulate the visual intention understanding task as a hierarchical classification problem, capturing multiple granular features in different layers, which corresponds to hierarchical intention labels. For textual hierarchy, we directly extract the semantic representation from intention labels at different levels, which supplements the visual content modeling without extra manual annotations. Moreover, to further narrow the domain gap between different modalities, a cross-modality pyramid alignment module is designed to dynamically optimize the performance of visual intention understanding in a joint learning manner. Comprehensive experiments intuitively demonstrate the superiority of our proposed method, outperforming existing visual intention understanding methods.
Mang Ye, Qinghongya Shi, Kehua Su, Bo Du 0001
IEEE Trans. Image Process.3
2021 Learning behaviour recognition based on multi-object image in single viewpoint
Kehua Su, Chengcheng Zhou
Pers. Ubiquitous Comput.1
2020 Mesh Parametrization Driven by Unit Normal Flow
abstract
Abstract Based on mesh deformation, we present a unified mesh parametrization algorithm for both planar and spherical domains. Our approach can produce intermediate frames from the original meshes to the targets. We derive and define a novel geometric flow: ‘unit normal flow (UNF)’ and prove that if UNF converges, it will deform a surface to a constant mean curvature (CMC) surface, such as planes and spheres. Our method works by deforming meshes of disk topology to planes, and spherical meshes to spheres. Our algorithm is robust, efficient, simple to implement. To demonstrate the robustness and effectiveness of our method, we apply it to hundreds of models of varying complexities. Our experiments show that our algorithm can be a competing alternative approach to other state‐of‐the‐art mesh parametrization methods. The unit normal flow also suggests a potential direction for creating CMC surfaces.
Kehua Su, Na Lei, Steven J. Gortler, Xianfeng Gu
Comput. Graph. Forum2
2019 Curvature adaptive surface remeshing by sampling normal cycle
Kehua Su, Na Lei, Wei Chen 0130, Hang Si, Shikui Chen, Xianfeng Gu
Comput. Aided Des.1
2019 A geometric view of optimal transportation and generative model
Na Lei, Kehua Su, Shing-Tung Yau, Xianfeng Gu
Comput. Aided Geom. Des.2
2019 Discrete Lie flow: A measure controllable parameterization method
Kehua Su, Shifan Zhao, Na Lei, Xianfeng Gu
Comput. Aided Geom. Des.1
2019 Discrete Calabi Flow: A Unified Conformal Parameterization Method
abstract
Abstract Conformal parameterization for surfaces into various parameter domains is a fundamental task in computer graphics. Prior research on discrete Ricci flow provided us with promising inspirations from methods derived via Riemannian geometry, which is rigorous in theory and effective inpractice. In this paper, we propose a unified conformal parameterization approachfor turning triangle meshes into planar and spherical domains using discrete Calabi flow onpiecewise linear metric. We incorporate edge‐flipping surgery to guarantee convergence as well as other significant improvements including approximate Newton's method, optimal step‐lengths, priority embedding and boundary customizing, which achieve better performance and functionality with robustness and accuracy.
Kehua Su, Yuming Zhou, Xianfeng Gu
Comput. Graph. Forum1
2018 A 3D face registration algorithm based on conformal mapping
abstract
Summary Recently, 3D facial datasets are more easily available, and at the same time, the research of 3D face becomes more and more important. One of the most important research fields is 3D face registration, which plays an important role in face recognition, face shape analysis, and face animation. However, one of the most challenging issues in 3D face registration is to obtain a unique mapping for faces with different expression and landmark constraints. In this paper, we propose a novel conformal mapping algorithm to deal with the 3D face registration. Besides, the calculation is about harmonic energy, which makes our method applicable to low‐quality meshes. We begin with a harmonic mapping, then minimize the harmonic energy by a specific boundary condition on surfaces, and to obtain the conformal mapping, finally, we use a landmark‐constrained surface registration algorithm to register faces. Numerical experiments on various surfaces demonstrate the efficiency and robustness of our method.
Kun Qian 0013, Kehua Su, Jialing Zhang
Concurr. Comput. Pract. Exp.2
2018 A measure-driven method for normal mapping and normal map design of 3D models
Kun Qian 0013, Kehua Su, Jialing Zhang
Multim. Tools Appl.3
2017 Volume preserving mesh parameterization based on optimal mass transportation
Kehua Su, Wei Chen 0130, Na Lei, Junwei Zhang 0010, Kun Qian 0013, Xianfeng Gu
Comput. Aided Des.1
2017 Robust surface registration using optimal mass transport and Teichmüller mapping
Ming Ma 0003, Na Lei, Wei Chen 0130, Kehua Su, Xianfeng Gu
Graph. Model.4
2016 Measure controllable volumetric mesh parameterization
Kehua Su, Wei Chen 0130, Na Lei, Xianfeng Gu
Comput. Aided Des.1
2016 Area-preserving mesh parameterization for poly-annulus surfaces based on optimal mass transportation
Kehua Su, Kun Qian 0013, Na Lei, Junwei Zhang 0010, Min Zhang 0069, Xianfeng Gu
Comput. Aided Geom. Des.1