Mohan Jing

dblp:344/5508 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
0009-0003-3891-9207ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Video understanding and tracking · 48% Vision and language · 36% Question answering and dialogue systems · 8%
Network and information security
1 paper
Digital forensics and information hiding · 100%
Human-computer interaction and pervasive computing
1 paper
Usability and user experience research · 50% Human-robot interaction · 50%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
multimodal reasoning
1.522025
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation · ICLR 2025
Answer-Based Entity Extraction and Alignment for Visual Text Question Answering · ACM Multimedia 2023
Computer vision › Video understanding and tracking
action recognition
0.912025
Hierarchical Multi-Feature Extraction and Aggregation for Micro-Action Recognition · ACM Multimedia 2025
Computer vision › Video understanding and tracking › action recognition
fine-grained action recognition
0.912025
Hierarchical Multi-Feature Extraction and Aggregation for Micro-Action Recognition · ACM Multimedia 2025
Computer vision › Video understanding and tracking › action recognition › fine-grained action recognition
micro-action recognition
0.912025
Hierarchical Multi-Feature Extraction and Aggregation for Micro-Action Recognition · ACM Multimedia 2025
Program synthesis and code generation › code generation with language models
chart-to-code generation
0.912025
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation · ICLR 2025
Computer vision › Video understanding and tracking
action detection
0.812024
End-to-end Spatio-Temporal Information Aggregation For Micro-Action Detection · ACM Multimedia 2024
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatiotemporal feature learning
0.812024
End-to-end Spatio-Temporal Information Aggregation For Micro-Action Detection · ACM Multimedia 2024
Multimedia analysis and retrieval
audio-visual analysis
0.812024
Building Robust Video-Level Deepfake Detection via Audio-Visual Local-Global Interactions · ACM Multimedia 2024
Digital forensics and information hiding
deepfake detection
0.812024
Building Robust Video-Level Deepfake Detection via Audio-Visual Local-Global Interactions · ACM Multimedia 2024
Digital forensics and information hiding › deepfake detection
video deepfake detection
0.812024
Building Robust Video-Level Deepfake Detection via Audio-Visual Local-Global Interactions · ACM Multimedia 2024
Natural language and speech › Question answering and dialogue systems
multi-hop reasoning
0.712023
Answer-Based Entity Extraction and Alignment for Visual Text Question Answering · ACM Multimedia 2023
Computer vision › Vision and language › visual question answering
text-based visual question answering
0.712023
Answer-Based Entity Extraction and Alignment for Visual Text Question Answering · ACM Multimedia 2023
Computer vision › Vision and language
visual question answering
0.712023
Answer-Based Entity Extraction and Alignment for Visual Text Question Answering · ACM Multimedia 2023
Usability and user experience research › user modeling
engagement estimation
0.712023
Sliding Window Seq2seq Modeling for Engagement Estimation · ACM Multimedia 2023
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model evaluation
0.312025
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation · ICLR 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.312025
Hierarchical Multi-Feature Extraction and Aggregation for Micro-Action Recognition · ACM Multimedia 2025
Machine learning › Deep learning architectures and training
multi-scale feature fusion
0.212024
End-to-end Spatio-Temporal Information Aggregation For Micro-Action Detection · ACM Multimedia 2024
Natural language and speech › Information extraction and text analysis
named entity recognition
0.212023
Answer-Based Entity Extraction and Alignment for Visual Text Question Answering · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

benchmark evaluation · 1.7local-global interaction · 1.5data augmentation · 1.5adaptive modality selection · 1.5two-stage training · 0.9feature decoupling · 0.93d-resnet adapter · 0.9feature pyramid · 0.8cross-attention · 0.83D-SENet adapter · 0.8transformer · 0.7sliding window seq2seq modeling · 0.7entity extraction · 0.7cross-modal alignment · 0.7BiLSTM · 0.7
YearPublicationVenuePosition
2025 ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
abstract
We introduce a new benchmark, ChartMimic, aimed at assessing the visually-grounded code generation capabilities of large multimodal models (LMMs). ChartMimic utilizes information-intensive visual charts and textual instructions as inputs, requiring LMMs to generate the corresponding code for chart rendering. ChartMimic includes $4,800$ human-curated (figure, instruction, code) triplets, which represent the authentic chart use cases found in scientific papers across various domains (e.g., Physics, Computer Science, Economics, etc). These charts span $18$ regular types and $4$ advanced types, diversifying into $201$ subcategories. Furthermore, we propose multi-level evaluation metrics to provide an automatic and thorough assessment of the output code and the rendered charts. Unlike existing code generation benchmarks, ChartMimic places emphasis on evaluating LMMs' capacity to harmonize a blend of cognitive capabilities, encompassing visual understanding, code generation, and cross-modal reasoning. The evaluation of $3$ proprietary models and $14$ open-weight models highlights the substantial challenges posed by ChartMimic. Even the advanced GPT-4o, InternVL2-Llama3-76B only achieved an average score across Direct Mimic and Customized Mimic tasks of $82.2$ and $61.6$, respectively, indicating significant room for improvement. We anticipate that ChartMimic will inspire the development of LMMs, advancing the pursuit of artificial general intelligence.
Cheng Yang 0002, Chufan Shi, Bo Shui, Junjie Wang 0011, Mohan Jing, Linran Xu, Siheng Li, Gongye Liu, Xiaomei Nie, Deng Cai 0002, Yujiu Yang 0001
ICLR6
2025 Hierarchical Multi-Feature Extraction and Aggregation for Micro-Action Recognition
abstract
Micro-action refers to subtle, low-intensity non-verbal behaviors that can provide insights into an individual's underlying emotions and intentions. Due to its brief duration and significant overlap, identifying these micro-actions poses a challenge for current models. In response to these challenges, this paper proposes a novel multi-feature fusion framework, which extracts coarse-grained body features and fine-grained action features separately. Specifically, we present Temporal Contextualization for fine-grained learning, a cross-frame injection mechanism designed to capture essential spatio-temporal information and introduce a 3D-ResNet Adapter for coarse-grained learning, which aggregates temporal data and facilitates parameter-efficient fine-tuning. In consideration of the task dataset distribution's long-tail nature, the implementation of Feature Decoupling is undertaken, adopting a two-stage training strategy. By conducting experiments, the aforementioned hierarchical multi-feature extraction and aggregation approach has been demonstrated to yield substantial enhancement in Micro-Action Recognition. Our method attains an F1-mean score of 77.75% on the MA-52 dataset, ranking 1st in the 2nd Micro-Action Analysis Grand Challenge in Conjunction with ACM MM'25.
Zhichao Xia, Yanjun Chi, Lingsi Zhu, Mohan Jing, Jun Yu 0001
ACM Multimedia5
2024 Building Robust Video-Level Deepfake Detection via Audio-Visual Local-Global Interactions
abstract
The continual advancements in Generative Artificial Intelligence have created substantial hurdles for accurate deepfake detection, leading to limitations of currently popular detection methods across content-driven video-level deepfake detection scenarios. In this paper, we present the solutions to the Video-Level Deepfake Detection task. Our empirical findings demonstrate that modeling correlations of audio-visual modalities is important for video-level deepfake detection. Therefore, we introduce the model denoted Audio-Visual Local-Global Neural Network (i.e., AV-LGNN) in which the core design is the proposed AV-LGI Module (Audio-Visual Local-Global Interaction Module). The AV-LGI Module is composed of three stages: Local Intra-Region Interaction, Global Inter-Region Interaction, and Local-Global Interaction, which can better capture detailed information at local-level and efficiently learn the fine-grained correlations of inter-modalities in video deepfake detection under lower computational overheads. We further propose an adaptive modality selection strategy to facilitate model learning. Besides, a variety of data augmentation techniques are incorporated for audio-visual branches to enhance the robustness of the AV-LGNN. The experimental results verify the effectiveness of our model.
Jia Zhang 0016, Mohan Jing, Keda Lu, Jun Yu 0001, Wen Su 0004, Fang Gao 0001, Jianqing Sun, Jiaen Liang
ACM Multimedia4
2024 End-to-end Spatio-Temporal Information Aggregation For Micro-Action Detection
abstract
Micro-actions convey the emotions of characters in daily communication and offer richer semantic information compared to conventional actions. Accurate detection of these micro-actions is essential for video understanding. Due to their short duration, low intensity, and high overlap, micro-actions require more detailed video features, presenting a significant challenge for accurate detection. To address these challenges, we propose the 3D-SENet Adapter, which aggregates spatio-temporal information and enables end-to-end online video feature learning. We also find that incorporating background information significantly enhances the detection of small-scale micro-actions. Thus we develop the Cross-Attention Aggregation Detection Head, which integrates multi-scale features within the feature pyramid, thereby improving the detection accuracy of micro-actions occupying small regions in video frames. Our approach achieves first place in the Multi-label Micro-Action Detection (MMAD) and second place in the Micro-Action Recognition (MAR) of Micro-Action Analysis Grand Challenge.
Jun Yu 0001, Mohan Jing, Guopeng Zhao, Keda Lu, Feng Zhao 0005, Jiaqing Sun, Jiaen Liang
ACM Multimedia2
2023 Answer-Based Entity Extraction and Alignment for Visual Text Question Answering
abstract
As a variant of visual question answering (VQA), visual text question answering (VTQA) provides a text-image pair for each question. Text utilizes named entities to describe corresponding image. Consequently, the ability to perform multi-hop reasoning using named entities between text and image becomes critically important. However, existing models pay relatively less attention to this aspect. Therefore, we propose Answer-Based Entity Extraction and Alignment Model (AEEA) to enable a comprehensive understanding and support multi-hop reasoning. The core of AEEA lies in two main components: AKECMR and answer aware predictor. The former emphasizes the alignment of modalities and effectively distinguishes between intra-modal and inter-modal information, and the latter prioritizes the full utilization of intrinsic semantic information contained in answers during training. Our model outperforms the baseline by 2.24% on test-dev set and 1.06% on test set, securing the third place in VTQA2023(English).
Jun Yu 0001, Mohan Jing, Weihao Liu 0004, Tongxu Luo, Keda Lu, Fangyu Lei, Jianqing Sun, Jiaen Liang
ACM Multimedia2
2023 Sliding Window Seq2seq Modeling for Engagement Estimation
abstract
Engagement estimation in human conversations has been one of the most important research issues for natural human-robot interaction. However, previous datasets and studies mainly focus on the video-wise level of engagement estimation, therefore, can hardly reflect human's constantly changing engagement. Fortunately, the MultiMediate '23 challenge provides the frame-wise level of engagement estimation task. In this paper, we propose Sliding Window Seq2seq Modeling by BiLSTM and Transformer with powerful sequence modeling capabilities. Our method fully utilizes the global and local multi-modal feature information in the participants' videos and accurately expresses the engagement of the participants at each moment. Our method achieves the state-of-the-art CCC result of 0.71 for engagement estimation on the corresponding test sets.
Jun Yu 0001, Keda Lu, Mohan Jing, Ziqi Liang, Jianqing Sun, Jiaen Liang
ACM Multimedia3
2023 Promoting Open-Domain Dialogue Generation Through Learning Pattern Information Between Contexts and Responses
Mengjuan Liu, Yunfan Yang, Mohan Jing
NLPCC (2)5