Jincai Chen

dblp:21/7549 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0001-7368-1677ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Enhancing Speech Emotion Recognition with Speech Dynamic Modeling and Multi-Modal Knowledge Distillation
abstract
Complementary semantic information from the text modality, obtained through runtime transcription, plays a crucial role in Speech Emotion Recognition (SER). However, it introduces additional computational overhead and potential errors. To address these issues, we propose the SDMMKD framework, which directly leverages multimodal knowledge without runtime transcription. Specifically, SDMMKD distills emotion knowledge at both the feature and logit levels from a pre-trained multimodal teacher during training. During inference, SDMMKD relies solely on speech signals to perform unimodal SER. Additionally, we utilize a Mamba block to enhance dynamic temporal features. Experimental results on the widely used IEMOCAP dataset demonstrate that our proposed SDMMKD framework outperforms state-of-the-art methods, achieving a WAR of 75.89% and a UAR of 77.37%.
Chuanbo Zhu 0002, Yifan Liu 0015, Jincai Chen
ICASSP4
2025 You cannot handle the weather: Progressive amplified adverse-weather-gradient projection adversarial attack
Yifan Liu 0015, Min Chen 0003, Chuanbo Zhu 0002, Jincai Chen
Expert Syst. Appl.5
2025 Multi-level Multi-task representation learning with adaptive fusion for multimodal sentiment analysis
Chuanbo Zhu 0002, Min Chen 0003, Yifan Liu 0015, Jincai Chen
Neural Comput. Appl.8
2025 Listen With Seeing: Cross-Modal Contrastive Learning for Audio-Visual Event Localization
abstract
In real-world physiological and psychological scenarios, there often exists a robust complementary correlation between audio and visual signals. Audio-Visual Event Localization (AVEL) aims to identify segments with Audio-Visual Events (AVEs) that contain both audio and visual tracks in unconstrained videos. Prior studies have predominantly focused on audio-visual cross-modal fusion methods, overlooking the fine-grained exploration of the cross-modal information fusion mechanism. Moreover, due to the inherent heterogeneity of multi-modal data, inevitable new noise is introduced during the audio-visual fusion process. To address these challenges, we propose a novel Cross-modal Contrastive Learning Network (CCLN) for AVEL, comprising a backbone network and a branch network. In the backbone network, drawing inspiration from physiological theories of sensory integration, we elucidate the process of audio-visual information fusion, interaction, and integration from an information-flow perspective. Notably, the Self-constrained Bi-modal Interaction (SBI) module is a bi-modal attention structure integrated with audio-visual fusion information, and through gated processing of the audio-visual correlation matrix, it effectively captures inter-modal correlation. The Foreground Event Enhancement (FEE) module emphasizes the significance of event-level boundaries by elongating the distance between scene events during training through adaptive weights. Furthermore, we introduce weak video-level labels to constrain the cross-modal semantic alignment of audio-visual events and design a weakly supervised cross-modal contrastive learning loss (WCCL Loss) function, which enhances the quality of fusion representation in the dual-branch contrastive learning framework. Extensive experiments conducted on the AVE dataset for both fully supervised and weakly supervised event localization, as well as Cross-Modal Localization (CML) tasks, demonstrate the superior performance of our model compared to state-of-the-art approaches.
Min Chen 0003, Chuanbo Zhu 0002, Ping Lu 0006, Jincai Chen
IEEE Trans. Multim.6
2023 SCLAV: Supervised Cross-modal Contrastive Learning for Audio-Visual Coding
abstract
Audio and vision are important senses for high-level cognition, and their special strong correlation makes audio-visual coding a crucial factor in many multimodal tasks. However, there are two challenges in audio-visual coding. First, the heterogeneity of multimodal data often leads to misalignment of cross-modal features under the same sample, which reduces their representation quality. Second, most self-supervised learning frameworks are constructed based on instance semantics, and the generated pseudo labels introduce additional classification noise. To address these challenges, we propose a Supervised Cross-modal Contrastive Learning Framework for Audio-Visual Coding (SCLAV). Our framework includes an audio-visual coding network composed of an inter-modal attention interaction module and an intra-modal self-integration module, which leverage multimodal complementary and hidden information for better representation. Additionally, we introduce a supervised cross-modal contrastive loss to minimize the distance between audio and vision features of the same instance, and use weak labels of multimodal data to eliminate the feature-oriented classification noise. Extensive experiments on the AVE and XD-Violence datasets demonstrate that SCLAV outperforms the state-of-the-art results, even with limited computational resources.
Min Chen 0003, Jialiang Cheng, Chuanbo Zhu 0002, Jincai Chen
ACM Multimedia6
2023 Dynamic interactive learning network for audio-visual event localization
Jincai Chen, Ruili Wang 0001, Jiangfeng Zeng, Ping Lu 0006
Appl. Intell.1
2023 Snowed autoencoders are efficient snow removers
Yifan Liu 0015, Jincai Chen, Ping Lu 0006, Chuanbo Zhu 0002, Yugen Jian
Comput. Graph.2
2023 MOONLIT: momentum-contrast and large-kernel for multi-fine-grained deraining
Yifan Liu 0015, Jincai Chen, Ping Lu 0006, Chuanbo Zhu 0002, Yugen Jian
J. Supercomput.2
2023 Self-Supervised Learning With Data-Efficient Supervised Fine-Tuning for Crowd Counting
abstract
Due to the expensive and laborious annotations of labeled data required by fully-supervised learning in the crowd counting task, it is desirable to explore a method to reduce the labeling burden. There exists a large number of unlabeled images in the wild that can be easily obtained compared to labeled datasets. Based on the characteristics of consistent spatial transformation with the annotations of heads and image, this paper proposes a self-supervised learning framework with unlabeled and limited labeled data for pre-training and fine-tuning crowd counting model (SSL-FT). It includes an online network and a target network that receive the same images but are randomly processed by two defined augmentation transformations. We leverage unlabeled data to pre-train the online network based on a self-supervised loss and small-scale labeled data to transfer the model to a specific domain based on a fully-supervised loss. We demonstrate the effectiveness of the SSL-FT on four public datasets including ShanghaiTech PartA, PartB, UCF-QNRF and WorldExpo'10 utilizing a classical counting model. Experimental results show that our approach performs better than state-of-art semi-supervised methods.
Rui Wang 0077, Yixue Hao, Long Hu, Jincai Chen, Min Chen 0003, Di Wu 0001
IEEE Trans. Multim.4
2022 Multi-level, multi-modal interactions for visual question answering over text in images
Jincai Chen, Jiangfeng Zeng, Fuhao Zou, Yuan-Fang Li, Ping Lu 0006
World Wide Web1
2021 Considering anatomical prior information for low-dose CT image enhancement using attribute-augmented Wasserstein generative adversarial networks
Zhenxing Huang, Xinfeng Liu, Rongpin Wang, Jincai Chen, Ping Lu 0006, Qiyang Zhang 0002, Changhui Jiang, Yongfeng Yang, Xin Liu 0053, Hairong Zheng, Dong Liang 0001, Zhanli Hu
Neurocomputing4
2021 Combining cross-modal knowledge transfer and semi-supervised learning for speech emotion recognition
Min Chen 0003, Jincai Chen, Yuan-Fang Li, Yiling Wu, Minglei Li 0001, Chuanbo Zhu 0002
Knowl. Based Syst.3
2019 RobustiQ: A Robust ANN Search Method for Billion-scale Similarity Search on GPUs
abstract
GPU-based methods represent state-of-the-art in approximate nearest neighbor (ANN) search, as they are scalable (billion-scale), accurate (high recall) as well as efficient (sub-millisecond query speed). Faiss, the representative GPU-based ANN system, achieves considerably faster query speed than the representative CPU-based systems. The query accuracy of Faiss critically depends on the number of indexing regions, which in turn is dependent on the amount of available memory. At the same time, query speed deteriorates dramatically with the increase in the number of partition regions. Thus, it can be observed that Faiss suffers from a lack of robustness, that the fine-grained partitioning of datasets is achieved at the expense of search speed, and vice versa. In this paper, we introduce a new GPU-based ANN search method, Robust Quantization (RobustiQ), that addresses the robustness limitations of existing GPU-based methods in a holistic way. We design a novel hierarchical indexing structure using vector and bilayer line quantization. This indexing structure, together with our indexing and encoding methods, allows RobustiQ to avoid the need for maintaining a large lookup table, hence reduces not only memory consumption but also query complexity. Our extensive evaluation on two public billion-scale benchmark datasets, SIFT1B and DEEP1B, shows that RobustiQ consistently obtains 2-3 × speedup over Faiss while achieving better query accuracy for different codebook sizes. Compared to the best CPU-based ANN systems, RobustiQ achieves even more pronounced average speedups of 51.8 × and 11 × respectively.
Wei Chen 0154, Jincai Chen, Fuhao Zou, Yuan-Fang Li, Ping Lu 0006
ICMR2
2019 CCPNC: A Cooperative Caching Strategy Based on Content Popularity and Node Centrality
abstract
The in-network caching mechanism is one of the core technologies of the Content Centric Network (CCN) and has been increasingly concerned. In order to improve the cache hit ratio of the content centric network cache system and increase the content diversity of the cache system, this paper proposes a cooperative caching strategy based on content popularity and node centrality, called CCPNC. The CCPNC caching strategy comprehensively considers content popularity and node distribution rules. It can separately cache content objects based on different popularity and mobilize the core routing nodes in the network to work together with non-core routing nodes. The CCPNC caching strategy not only makes use of the core routing node cache resources to provide faster popular content services for a wide range of users, but also avoids unnecessary high-frequency cache replacement of the core routing nodes. Meanwhile, it utilizes the cache resources of non-core routing nodes to provide more convenient non-popular content services. Through simulation experiments, it is found that the CCPNC caching strategy can effectively balance the distribution of content objects in the cache system and improve the cache hit ratio of the content centric network, while reducing the average routing hop and average request latency of content backhaul.
Yunming Mo, Jinxing Bao, Shaobing Wang, Yaxiong Ma, Jiabao Huang, Ping Lu 0006, Jincai Chen
NAS8
2019 Vector and line quantization for billion-scale similarity search on GPUs
Wei Chen 0154, Jincai Chen, Fuhao Zou, Yuan-Fang Li, Ping Lu 0006
Future Gener. Comput. Syst.2
2011 Auxiliary Storage and Dynamic Configuration for Open Cloud Storage
Jincai Chen, Yangfeng Huang, Minghui Lai, Ping Lu 0006
CLOSER1
2010 Recording Performance Analysis of Pre-patterned-deposition Bit Patterned Media through Micromagnetic Simulation
abstract
Deposition of magnetic material onto pre-patterned disks is a promising choice for Bit Patterned Media (BPM) fabrication. Nevertheless, understanding the relationship between recording characteristics and the patterning parameters is still lacking. In this article, we present a theoretical study on the recording and noise properties of the BPM made from pre-patterned substrates through micro magnetic simulation, and propose a new Signal-to-Noise Ratio (SNR) simulating method, which can avoid the complicated calculation of recording signal sensed in GMR head. The dependence of SNR on the lattice geometry, such as dot size, spacing, pillar height and head-media spacing are carefully simulated and discussed. Finally, based on all the simulation, we conclude that 0.5~0.7 is preferred for island fill factor in practical preparation to keep high SNR, and decreasing the head-media spacing is another effective way to improve SNR, especially in the tall pillar situation.
Jincai Chen, Gongye Zhou
PDCAT3
2009 TS-A: A Hierarchy Extended Cellular Automata Model for Complex Networking Storage System
abstract
In order to analyze the dynamic behaviors of complex networking storage system, a hybrid analysis model called TS-A (Transit-Stub-Automata) is proposed. This analysis model is a hierarchy extended cellular automata; it extends cellular lattice to the hierarchy domains topology, so it can reflect more characteristics of the real storage network, and it is very suitable for simulating the evolution of the dynamic networking storage system.
Jincai Chen, Lanlan Yuan, Gongye Zhou
NAS1
2009 Stress Distributions on the Slider with Different Accommodation Coefficients
abstract
Recently the influence of accommodation coefficient, which determines the behaviors of the reflected molecules at the boundary walls, was studied a lot for exploring the flying characteristics of head-disk interface. However, all of the researches are based on the slider with unique accommodation coefficient. In this paper, the different correction models of Reynolds equation, the free molecular gas film lubrication (MGL) equation and direct simulation Monte Carlo (DSMC) method are compared. Stress distributions on the slider with different accommodation coefficients were calculated by using the free molecular MGL equation and DSMC method respectively. The results show that the shear stress is different with the slider with different accommodation coefficients, and hence the shear stress can be controlled by giving the slider proper accommodation coefficients, without influencing the pressure. This provides a theory basis for the design of slider of HDD with high stability.
Jincai Chen, Gongye Zhou, Libang Zhang
NAS1