Ming Kong 0001

dblp:121/1621-1 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-6177-3707ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Placing Any Object at Any 3D Position
Ming Kong 0001, Zhanbin Hu, Zhijie Xu
AAAI2
2026 UrbanGeoEval: A City-Scale Benchmark for Evaluating Large Language Models in Geospatial Reasoning
abstract
Mutian Bao, Qiuyi Qi, Tian Liang, Jinjian Zhang, Wei Zhou, Ming Kong, Linjian Mo, Qiang Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mutian Bao, Qiuyi Qi, Jinjian Zhang, Ming Kong 0001, Linjian Mo
ACL (1)6
2026 STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training
abstract
Qiuyi Qi, Tian Liang, Mutian Bao, Jinjian Zhang, Dongnan Liu, Wei Zhou, Linjian Mo, Ming Kong, Jie Liu, Feng Zhang, Qiang Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qiuyi Qi, Mutian Bao, Jinjian Zhang, Dongnan Liu, Linjian Mo, Ming Kong 0001
ACL (1)8
2026 Demystifying 3D Spatial Awareness via LLM Router
He Tao, Luyuan Chen, Tesi Lin, Jinjian Zhang, Ming Kong 0001
ICPR (7)7
2025 MHBench: Demystifying Motion Hallucination in VideoLLMs
abstract
Similar to Language or Image LLMs, VideoLLMs are also plagued by hallucination issues. Hallucinations in videos not only manifest in the spatial dimension regarding the perception of the existence of visual objects (static) but also the temporal dimension influencing the perception of actions and events (dynamic). This paper introduces the concept of Motion Hallucination for the first time, exploring the hallucination phenomena caused by insufficient motion perception capabilities in VideoLMMs, as well as how to detect, evaluate, and mitigate the hallucination. To this end, we propose the first benchmark for assessing motion hallucination MHBench, which consists of 1,200 videos of 20 different action categories. By constructing a collection of adversarial triplet types of videos (original/antonym/incomplete), we achieve a comprehensive evaluation of motion hallucination. Furthermore, we present a Motion Contrastive Decoding (MotionCD) method, which employs bidirectional motion elimination between the original video and its reverse playback to construct an amateur model that removes the influence of motion while preserving visual information, thereby effectively suppressing motion hallucination. Extensive experiments on MHBench reveal that current state-of-the-art VideoLLMs significantly suffer from motion hallucination, while the introduction of MotionCD can effectively mitigate this issue, achieving up to a 15.1% performance improvement. We hope this work will guide future efforts in avoiding and mitigating hallucinations in VideoLLMs.
Ming Kong 0001, Xianzhou Zeng, Luyuan Chen
AAAI1
2025 MoLE: Decoding by Mixture of Layer Experts Alleviates Hallucination in Large Vision-Language Models
abstract
Recent advancements in Large Vision-Language Models (LVLMs) highlight their ability to integrate and process multi-modal information. However, hallucinations—where generated content is inconsistent with input vision and instructions—remain a challenge. In this paper, we analyze LVLMs' layer-wise decoding and identify that hallucinations can arise during the reasoning and factual information injection process. Additionally, as the number of generated tokens increases, the forgetting of the original prompt may also lead to hallucinations.To address this, we propose a training-free decoding method called Mixture of Layer Experts (MoLE). MoLE leverages a heuristic gating mechanism to dynamically select multiple layers of LVLMs as expert layers: the Final Expert, the Second Opinion expert, and the Prompt Retention Expert. By the cooperation of each expert, MoLE enhances the robustness and faithfulness of the generation process. Our extensive experiments demonstrate that MoLE significantly reduces hallucinations, outperforming the current state-of-the-art decoding techniques across three mainstream LVLMs and two established hallucination benchmarks. Moreover, our method reveals the potential of LVLMs to independently produce more reliable and accurate outputs.
Yuetian Du, Ming Kong 0001, Luyuan Chen, Siye Chen
AAAI4
2025 Confidence Calibration for Multimodal LLMs: An Empirical Study Through Medical VQA
Yuetian Du, Ming Kong 0001, Qiang Long, Bingdi Chen
MICCAI (6)3
2025 Explainable ADHD Diagnostic Framework Using Weakly-Supervised Action Recognition
Ninghan Fan, Ming Kong 0001, Bingdi Chen
MICCAI (8)2
2025 MoCo-ANA: MoCo-Like Adaptive Neighbor Aggregation for Face Clustering
Ming Kong 0001, Congquan Yan
PRICAI (5)1
2025 Progressive semantic learning for unsupervised skeleton-based action recognition
Luyuan Chen, Ming Kong 0001, Xianzhou Zeng, Mengxu Lu
Mach. Learn.3
2025 Semantic-aware contrastive learning via multi-prompt alignment
Ming Kong 0001, Luyuan Chen, Di Xie
Mach. Learn.3
2024 Querying as Prompt: Parameter-Efficient Learning for Multimodal Language Model
abstract
Recent advancements in language models pre-trained on large-scale corpora have significantly propelled developments in the NLP domain and advanced progress in multimodal tasks. In this paper, we propose a Parameter-Efficient multimodal language model learning strategy, named QaP (Querying as Prompt). Its core innovation is a novel modality-bridging method that allows a set of modality-specific queries to be input as soft prompts into a frozen pre-trained language model. Specifically, we introduce an efficient Text-Conditioned Resampler that is easy to incorporate into the language models, which enables adaptive injection of text-related multimodal information at different levels of the model through query learning. This approach effectively bridges multimodal information to the language models while fully leveraging its token fusion and representation potential. We validated our method across four datasets in three distinct multimodal tasks. The results demonstrate that our QaP multimodal language model achieves state-of-the-art performance in various tasks with training only 4.6% parameters. Code is available at https://github.com/RainltIQaP.
Ming Kong 0001, Luyuan Chen
CVPR3
2024 Probablistic Restoration with Adaptive Noise Sampling for 3D Human Pose Estimation
abstract
The accuracy and robustness of 3D human pose estimation (HPE) are limited by 2D pose detection errors and 2D to 3D ill-posed challenges, which have drawn great attention to Multi-Hypothesis HPE research. Most existing MH-HPE methods are based on generative models, which are computationally expensive and difficult to train. In this study, we propose a Probabilistic Restoration 3D Human Pose Estimation framework (PRPose) that can be integrated with any lightweight single-hypothesis model. Specifically, PRPose employs a weakly supervised approach to fit the hidden probability distribution of the 2D-to-3D lifting process in the Single-Hypothesis HPE model and then reverse-map the distribution to the 2D pose input through an adaptive noise sampling strategy to generate reasonable multi-hypothesis samples effectively. Extensive experiments on 3D HPE benchmarks (Human3.6M and MPI-INF-3DHP) highlight the effectiveness and efficiency of PRPose. Code is available at: https://github.com/xzhouzeng/PRPose.
Xianzhou Zeng, Ming Kong 0001, Luyuan Chen
ICME3
2024 MS-DETR: Exploiting Modality Synergy for Moment Retrieval and Highlight Detection
Luyuan Chen, Ming Kong 0001, Jianwu Wu
PRCV (10)3
2024 Simulating doctors' thinking logic for chest X-ray report generation via Transformer-based Semantic Query learning
Danyang Gao, Ming Kong 0001, Yongrui Zhao, Zhengxing Huang, Kun Kuang 0001, Fei Wu 0001
Medical Image Anal.2
2023 Temporal RPN Learning for Weakly-Supervised Temporal Action Localization
Ming Kong 0001, Luyuan Chen
ACML2
2023 CLAP: Contrastive Language-Audio Pre-training Model for Multi-modal Sentiment Analysis
abstract
Multi-modal Sentiment Analysis (MSA) is a hotspot of multi-modal fusion. To make full use of the correlation and complementarity between modalities in the process of fusing multi-modal data, we propose a two-stage framework of Contrastive Language-Audio Pre-training (CLAP) for the MSA task: 1) Making contrastive pre-training on an unlabeled large-scaled external data to yield better single-modal representations; 2) Adopting a Transformer-based multi-modal fusion module, to achieve further single-modal feature optimization and sentiment prediction via the task-driven training process. Our work fully demonstrates the importance and necessity of core elements such as pre-training, contrastive learning, and representation learning for the MSA task and significantly outperforms existing methods on two well-recognized MSA benchmarks.
Ming Kong 0001, Kun Kuang 0001, Fei Wu 0001
ICMR2
2022 TranSQ: Transformer-Based Semantic Query for Medical Report Generation
Ming Kong 0001, Zhengxing Huang, Kun Kuang 0001, Fei Wu 0001
MICCAI (8)1
2022 Attribute-aware interpretation learning for thyroid ultrasound diagnosis
Ming Kong 0001, Shuowen Zhou, Mengze Li 0001, Kun Kuang 0001, Zhengxing Huang, Fei Wu 0001
Artif. Intell. Medicine1
2021 Auxiliary diagnostic system for ADHD in children based on AI technology
abstract
Traditional diagnosis of attention deficit hyperactivity disorder (ADHD) in children is primarily through a questionnaire filled out by parents/teachers and clinical observations by doctors. It is inefficient and heavily depends on the doctor’s level of experience. In this paper, we integrate artificial intelligence (AI) technology into a software-hardware coordinated system to make ADHD diagnosis more efficient. Together with the intelligent analysis module, the camera group will collect the eye focus, facial expression, 3D body posture, and other children’s information during the completion of the functional test. Then, a multi-modal deep learning model is proposed to classify abnormal behavior fragments of children from the captured videos. In combination with other system modules, standardized diagnostic reports can be automatically generated, including test results, abnormal behavior analysis, diagnostic aid conclusions, and treatment recommendations. This system has participated in clinical diagnosis in Department of Psychology, The Children’s Hospital, Zhejiang University School of Medicine, and has been accepted and praised by doctors and patients.
Yanyi Zhang, Ming Kong 0001, Wenchen Hong, Di Xie, Chunmao Wang, Rongwang Yang
Frontiers Inf. Technol. Electron. Eng.2
2020 ADHD Intelligent Auxiliary Diagnosis System Based on Multimodal Information Fusion
abstract
The traditional medical diagnosis methods of ADHD mainly rely on scale evaluation and interview observation. The diagnosis conclusion is subjective and extremely dependent on the doctor's experience level. There is an urgent need to improve diagnosis efficiency and improve the diagnosis standard through other technical means in the clinical process. We have designed and developed the ADHD intelligent auxiliary diagnosis system with software and hardware cooperation. The system performs a set of functional test tasks, uses a camera module to capture multimodal information such as facial expressions, eye movements, limb movements, language expressions and reaction abilities of children during task completion, and uses computer vision technology to automatically extract measurable characteristics. Finally, deep learning technology is used to detect children's specific behaviors in the video, which is complementary to the existing doctor's diagnosis basis. This system was deployed in the Department of Psychology of Children's Hospital of Zhejiang University in July 2019 and has been used in actual clinical diagnosis to date. It has completed the testing and evaluation of hundreds of ADHD children.
Yanyi Zhang, Ming Kong 0001, Wenchen Hong, Fei Wu 0001
ACM Multimedia2
2017 Community-Based Question Answering via Contextual Ranking Metric Network Learning
abstract
The exponential growth of information on Community-based Question Answering (CQA) sites has raised the challenges for the accurate matching of high-quality answers to the given questions. Many existing approaches learn the matching model mainly based on the semantic similarity between questions and answers, which can not effectively handle the ambiguity problem of questions and the sparsity problem of CQA data. In this paper, we propose to solve these two problems by exploiting users' social contexts. Specifically, we propose a novel framework for CQA task by exploiting both the question-answer content in CQA site and users' social contexts. The experiment on real-world dataset shows the effectiveness of our method.
Hanqing Lu, Ming Kong 0001
AAAI2
2016 Social recommendation via multi-view user preference learning
Hanqing Lu, Chaochao Chen 0001, Ming Kong 0001, Hanyi Zhang, Zhou Zhao 0001
Neurocomputing3