EDBT 2026 Demo / reviewers in the wild / expert
Jiaxin Ye
dblp:151/9288
· DBLP profile ↗
15ranked-venue papers
7as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DepMamba: Progressive Fusion Mamba for Multimodal Depression DetectionabstractDepression is a common mental disorder that affects millions of people worldwide. Although promising, current multimodal methods hinge on aligned or aggregated multi-modal fusion, suffering two significant limitations: (i) inefficient long-range temporal modeling, and (ii) sub-optimal multimodal fusion between intermodal fusion and intramodal processing. In this paper, we propose an audio-visual progressive fusion Mamba for multimodal depression detection, termed DepMamba. DepMamba features two core designs: hierarchical contextual modeling and progressive multimodal fusion. On the one hand, hierarchical modeling introduces convolution neural networks and Mamba to extract the local-to-global features within long-range sequences. On the other hand, the progressive fusion first presents a multimodal collaborative State Space Model (SSM) extracting intermodal and intramodal information for each modality, and then utilizes a multimodal enhanced SSM for modality cohesion. Extensive experimental results on two large-scale depression datasets demonstrate the superior performance of our DepMamba over existing state-of-the-art methods. Code is available at https://github.com/Jiaxin-Ye/DepMamba. Jiaxin Ye, Junping Zhang, Hongming Shan |
ICASSP | 1 |
| 2025 | Emotional Face-to-SpeechabstractHow much can we infer about an emotional voice solely from an expressive face? This intriguing question holds great potential for applications such as virtual character dubbing and aiding individuals with expressive language disorders. Existing face-to-speech methods offer great promise in capturing identity characteristics but struggle to generate diverse vocal styles with emotional expression. In this paper, we explore a new task, termed emotional face-to-speech, aiming to synthesize emotional speech directly from expressive facial cues. To that end, we introduce DEmoFace, a novel generative framework that leverages a discrete diffusion transformer (DiT) with curriculum learning, built upon a multi-level neural audio codec. Specifically, we propose multimodal DiT blocks to dynamically align text and speech while tailoring vocal styles based on facial emotion and identity. To enhance training efficiency and generation quality, we further introduce a coarse-to-fine curriculum learning algorithm for multi-level token processing. In addition, we develop an enhanced predictor-free guidance to handle diverse conditioning scenarios, enabling multi-conditional generation and disentangling complex attributes effectively. Extensive experimental results demonstrate that DEmoFace generates more natural and consistent speech compared to baselines, even surpassing speech-driven methods. Demos of DEmoFace are shown at our project https://demoface.github.io. Jiaxin Ye, Boyuan Cao, Hongming Shan |
ICML | 1 |
| 2025 | RepLDM: Reprogramming Pretrained Latent Diffusion Models for High-Quality, High-Efficiency, High-Resolution Image GenerationabstractWhile latent diffusion models (LDMs), such as Stable Diffusion, are designed for high-resolution image generation, they often struggle with significant structural distortions when generating images at resolutions higher than their training one.
Instead of relying on extensive retraining, a more resource-efficient approach is to reprogram the pretrained model for high-resolution (HR) image generation; however, existing methods often result in poor image quality and long inference time.
We introduce RepLDM, a novel reprogramming framework for pretrained LDMs that enables high-quality, high-efficiency, high-resolution image generation; see Fig. 1. RepLDM consists of two stages: (i) an attention guidance stage, which generates a latent representation of a higher-quality training-resolution image using a novel parameter-free self-attention mechanism to enhance the structural consistency; and (ii) a progressive upsampling stage, which progressively performs upsampling in pixel space to mitigate the severe artifacts caused by latent space upsampling. The effective initialization from the first stage allows for denoising at higher resolutions with significantly fewer steps, improving the efficiency.
Extensive experimental results demonstrate that RepLDM significantly outperforms state-of-the-art methods in both quality and efficiency for HR image generation, underscoring its advantages for real-world applications.
Codes: https://github.com/kmittle/RepLDM. Boyuan Cao, Jiaxin Ye, Yujie Wei 0001, Hongming Shan |
NeurIPS | 2 |
| 2024 | EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
Ziyang Ma 0001, Hezhao Zhang, Zhisheng Zheng, Xiquan Li, Jiaxin Ye, Xie Chen 0001, Thomas Hain |
INTERSPEECH | 7 |
| 2024 | TWACapsNet: a capsule network with two-way attention mechanism for speech emotion recognition
Xin-Cheng Wen, Kunhong Liu 0001, Jiaxin Ye, Li-Yan Chen |
Soft Comput. | 4 |
| 2024 | Positive-Sample-Free Object Tracking via a Soft ConstraintabstractMost of the existing bounding box-based trackers rely on a classification subnetwork and a regression subnetwork to predict the location and scale of the bounding box. They learn the classification subnetwork by processing each sample individually and applying the suggested classification confidence to produce the final prediction. They typically involve heuristic positive sample configurations, which inevitably introduce mislabelled training samples and therefore deteriorate their tracking performance. Moreover, the parallel prediction of the bounding box position and scale may lead to misalignment of classification and regression. To address these issues,we propose a simple yet effective soft constraint-based tracking framework without positive samples (named SoftCT). SoftCT adaptively senses the target’s pixel position through a soft constraint mechanism, which eliminates potential performance gaps caused by artificially marking the target’s pixel position. In addition, SoftCT computes the state of the bounding box by aggregating such positional information, thereby allowing the tracker to avoid misalignment in classification and regression due to uninformed communication. Specifically, SoftCT directly senses the position of the target pixel and fuses this information into the bounding box prediction, rather than requiring explicit annotation or regression of the target pixel. Extensive experiments on six tracking benchmarks including GOT-10k, TrackingNet, LaSOT, UAV123, LaSOText and TNL2K demonstrate that our tracker achieves state-of-the-art performance, confirming its effectiveness and efficiency. Jiaxin Ye, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Xianxian Li, Rongrong Ji |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Meta-Path Based Attentional Graph Learning Model for Vulnerability DetectionabstractIn recent years, deep learning (DL)-based methods have been widely used in code vulnerability detection. The DL-based methods typically extract structural information from source code, e.g., code structure graph, and adopt neural networks such as Graph Neural Networks (GNNs) to learn the graph representations. However, these methods fail to consider the heterogeneous relations in the code structure graph, i.e., the heterogeneous relations mean that the different types of edges connect different types of nodes in the graph, which may obstruct the graph representation learning. Besides, these methods are limited in capturing long-range dependencies due to the deep levels in the code structure graph. In this paper, we propose aMeta-path basedAttentionalGraph learning model for code vulNErability deTection, calledMAGNET. MAGNET constructs a multi-granularity meta-path graph for each code snippet, in which the heterogeneous relations are denoted as meta-paths to represent the structural information. A meta-path based hierarchical attentional graph neural network is also proposed to capture the relations between distant nodes in the graph. We evaluate MAGNET on three public datasets and the results show that MAGNET outperforms the best baseline method in terms of F1 score by 6.32%, 21.50%, and 25.40%, respectively. MAGNET also achieves the best performance among all the baseline methods in detecting Top-25 most dangerous Common Weakness Enumerations (CWEs), further demonstrating its effectiveness in vulnerability detection. Xin-Cheng Wen, Cuiyun Gao 0001, Jiaxin Ye, Yichen Li 0003, Zhihong Tian 0001, Yan Jia 0001, Xuan Wang 0002 |
IEEE Trans. Software Eng. | 3 |
| 2023 | Temporal Modeling Matters: A Novel Temporal Emotional Modeling Approach for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) plays a vital role in improving the interactions between humans and machines by inferring human emotion and affective states from speech signals. Whereas recent works primarily focus on mining spatiotemporal information from hand-crafted features, we explore how to model the temporal patterns of speech emotions from dynamic temporal scales. Towards that goal, we introduce a novel temporal emotional modeling approach for SER, termed Temporal-aware bI-direction Multi-scale Network (TIM-Net), which learns multi-scale contextual affective representations from various time scales. Specifically, TIM-Net first employs temporal-aware blocks to learn temporal affective representation, then integrates complementary information from the past and the future to enrich contextual representations, and finally fuses multiple time scale features for better adaptation to the emotional variation. Extensive experimental results on six benchmark SER datasets demonstrate the superior performance of TIM-Net, gaining 2.34% and 2.61% improvements of the average UAR and WAR over the second-best on each corpus. The source code is available at https://github.com/Jiaxin-Ye/TIM-Net_SER. Jiaxin Ye, Xin-Cheng Wen, Yujie Wei 0001, Yong Xu 0009, Kunhong Liu 0001, Hongming Shan |
ICASSP | 1 |
| 2023 | Online Prototype Learning for Online Continual LearningabstractOnline continual learning (CL) studies the problem of learning continuously from a single-pass data stream while adapting to new data and mitigating catastrophic forgetting. Recently, by storing a small subset of old data, replay-based methods have shown promising performance. Unlike previous methods that focus on sample storage or knowledge distillation against catastrophic forgetting, this paper aims to understand why the online learning models fail to generalize well from a new perspective of shortcut learning. We identify shortcut learning as the key limiting factor for online CL, where the learned features may be biased, not generalizable to new tasks, and may have an adverse impact on knowledge distillation. To tackle this issue, we present the online prototype learning (OnPro) framework for online CL. First, we propose online prototype equilibrium to learn representative features against shortcut learning and discriminative features to avoid class confusion, ultimately achieving an equilibrium status that separates all seen classes well while learning new classes. Second, with the feedback of online prototypes, we devise a novel adaptive prototypical feedback mechanism to sense the classes that are easily misclassified and then enhance their boundaries. Extensive experimental results on widely-used benchmark datasets demonstrate the superior performance of OnPro over the state-of-the-art baseline methods. Source code is available at https://github.com/weilllllls/OnPro. Yujie Wei 0001, Jiaxin Ye, Zhizhong Huang, Junping Zhang, Hongming Shan |
ICCV | 2 |
| 2023 | Deep Learning for Effective Gender Classification of Tasmania Giant CrabsabstractThe giant crab fishery in southeast Australia currently suffers from a lack of accurate size and sex data to establish population dynamics essential for the management of the industry. Determining these traits manually on boats or from observers using video would be time-consuming and prone to observer error. This research aims to find an efficient, accurate, and fast way to identify giant crabs' gender to eliminate the manual cost and operation time. This problem can be solved by artificial intelligence technologies, particularly Convolutional Neural Networks (CNN). However, CNNs can detect the crabs but find it challenging to identify their genders from the top view (carapace). Other issues include the lack of training data and the demand for compact system to be deployed on boat easily. With such constraints of effectiveness and efficiency, we address the problem of crab gender classification by proposing a cascading architecture. First, we simplify a light-weight object detection model (MobileNet) for carapace localisation. After that our model extract the carapace area on crab images for classification modelling. According to our experiments' results, with the use of light-weight CNNs for our cascading architecture, we achieved the highest accuracy of up to 96.34% with an efficiency of around 2 frames per second in a Raspberry Pi V4. Jianping Yao, Son N. Tran, Lianxue Zhang, Jiaxin Ye, Ananda Maiti, Scott Hadley |
IJCNN | 4 |
| 2023 | Emo-DNA: Emotion Decoupling and Alignment Learning for Cross-Corpus Speech Emotion RecognitionabstractCross-corpus speech emotion recognition (SER) seeks to generalize the ability of inferring speech emotion from a well-labeled corpus to an unlabeled one, which is a rather challenging task due to the significant discrepancy between two corpora. Existing methods, typically based on unsupervised domain adaptation (UDA), struggle to learn corpus-invariant features by global distribution alignment, but unfortunately, the resulting features are mixed with corpus-specific features or not class-discriminative. To tackle these challenges, we propose a novel Emotion Decoupling aNd Alignment learning framework (EMO-DNA) for cross-corpus SER, a novel UDA method to learn emotion-relevant corpus-invariant features. The novelties of EMO-DNA are two-fold: contrastive emotion decoupling and dual-level emotion alignment. On one hand, our contrastive emotion decoupling achieves decoupling learning via a contrastive decoupling loss to strengthen the separability of emotion-relevant features from corpus-specific ones. On the other hand, our dual-level emotion alignment introduces an adaptive threshold pseudo-labeling to select confident target samples for class-level alignment, and performs corpus-level alignment to jointly guide model for learning class-discriminative corpus-invariant features across corpora. Extensive experimental results demonstrate the superior performance of EMO-DNA over the state-of-the-art methods in several cross-corpus scenarios. Source code is available at https://github.com/Jiaxin-Ye/Emo-DNA. Jiaxin Ye, Yujie Wei 0001, Xin-Cheng Wen, Chenglong Ma 0002, Zhizhong Huang, Kunhong Liu 0001, Hongming Shan |
ACM Multimedia | 1 |
| 2022 | CTL-MTNet: A Novel CapsNet and Transfer Learning-Based Mixed Task Net for Single-Corpus and Cross-Corpus Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) has become a growing focus of research in human-computer interaction. An essential challenge in SER is to extract common attributes from different speakers or languages, especially when a specific source corpus has to be trained to recognize the unknown data coming from another speech corpus. To address this challenge, a Capsule Network (CapsNet) and Transfer Learning based Mixed Task Net (CTL-MTNet) are proposed to deal with both the single-corpus and cross-corpus SER tasks simultaneously in this paper. For the single-corpus task, the combination of Convolution-Pooling and Attention CapsNet module (CPAC) is designed by embedding the self-attention mechanism to the CapsNet, guiding the module to focus on the important features that can be fed into different capsules. The extracted high-level features by CPAC provide sufficient discriminative ability. Furthermore, to handle the cross-corpus task, CTL-MTNet employs a Corpus Adaptation Adversarial Module (CAAM) by combining CPAC with Margin Disparity Discrepancy (MDD), which can learn the domain-invariant emotion representations through extracting the strong emotion commonness. Experiments including ablation studies and visualizations on both single- and cross-corpus tasks using four well-known SER datasets in different languages are conducted for performance evaluation and comparison. The results indicate that in both tasks the CTL-MTNet showed better performance in all cases compared to a number of state-of-the-art methods. The source code and the supplementary materials are available at: https://github.com/MLDMXM2017/CTLMTNet. Xin-Cheng Wen, Jiaxin Ye, Yong Xu 0009, Xuan-Ze Wang, Chang-Li Wu, Kunhong Liu 0001 |
IJCAI | 2 |
| 2022 | GM-TCNet: Gated Multi-scale Temporal Convolutional Network using Emotion Causality for Speech Emotion Recognition
Jiaxin Ye, Xin-Cheng Wen, Xuan-Ze Wang, Yong Xu 0009, Chang-Li Wu, Li-Yan Chen, Kunhong Liu 0001 |
Speech Commun. | 1 |
| 2021 | AHOA: Adaptively Hybrid Optimization Algorithm for Flexible Job-shop Scheduling Problem
Jiaxin Ye, Dejun Xu, Haokai Hong, Yongxuan Lai, Min Jiang 0005 |
ICA3PP (1) | 1 |
| 2014 | On-board inertial-assisted visual odometer on an embedded systemabstractIn this paper, we propose a novel inertial-assisted visual odometry system intended for low-cost micro aerial vehicles (MAVs). The system sensor assembly consists of two downward-facing cameras and an inertial measurement unit (IMU) with three-axis accelerometers/gyroscopes. Real-time implementation of the system is enabled by a low-cost embedded system via two important features: firstly, simple pixel-level algorithms are integrated in a low-end FPGA and accelerated via pipeline and combinational logic techniques; secondly, a fast yaw-and-translation estimation algorithm works well with a novel outlier rejection scheme based on probabilistic predetermined operations rather than hypothesis testing iterations. We illustrate the performance of our system by hovering a MAV in a GPS-denied environment. Its feasibility and robustness is also illustrated in complex outdoor environments. Guyue Zhou, Jiaxin Ye, Zexiang Li 0001 |
ICRA | 2 |