Mingjie Han

dblp:72/11082 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 44% Representation and self-supervised learning · 22% 3D vision · 14%
Computer graphics and multimedia
1 paper
Virtual and augmented reality · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked prediction
1.012026
Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs · AAAI 2026
Computer vision › Vision and language
multimodal reasoning
1.012026
Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs · AAAI 2026
Computer vision › Vision and language
vision-language model
1.012026
Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs · AAAI 2026
Machine learning › Generative modeling
image generation
0.512021
A Generative Model-Based Predictive Display for Robotic Teleoperation · ICRA 2021
Computer vision › 3D vision
photorealistic rendering
0.512021
A Generative Model-Based Predictive Display for Robotic Teleoperation · ICRA 2021
Virtual and augmented reality › teleoperation
predictive display
0.512021
A Generative Model-Based Predictive Display for Robotic Teleoperation · ICRA 2021
Virtual and augmented reality
teleoperation
0.512021
A Generative Model-Based Predictive Display for Robotic Teleoperation · ICRA 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.312026
Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs · AAAI 2026
Robotics › Robot navigation and mapping › SLAM
3d mapping
0.112021
A Generative Model-Based Predictive Display for Robotic Teleoperation · ICRA 2021
Computer vision › 3D vision
3d reconstruction
0.112021
A Generative Model-Based Predictive Display for Robotic Teleoperation · ICRA 2021

Methods — techniques the papers use, named apart from their topics

reinforcement fine-tuning · 1.0prior sampling · 1.0generative model · 1.0RGB-D imaging · 1.0
YearPublicationVenuePosition
2026 Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs
abstract
Recent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet, a significant gap persists in their adaptation to real-world multimodal scenarios, most notably, vision-language tasks, due to a heavy focus on single-modal language settings. While efforts to transplant reinforcement learning techniques from NLP to Visual Language Models (VLMs) have emerged, these approaches often remain confined to perception-centric tasks or reduce images to textual summaries, failing to fully exploit visual context and commonsense knowledge, ultimately constraining the generalization of reasoning capabilities across diverse multimodal environments. To address this limitation, we introduce a novel fine-tuning task, Masked Prediction via Context and Commonsense (MPCC), which forces models to integrate visual context and commonsense reasoning by reconstructing semantically meaningful content from occluded images, thereby laying the foundation for generalized reasoning. To systematically evaluate the model’s performance in generalized reasoning, we developed a specialized evaluation benchmark, MPCC-Eval, and employed various fine-tuning strategies to guide reasoning. Among these, we introduced an innovative training method, Reinforcement Fine-Tuning with Prior Sampling, which not only enhances model performance but also improves its generalized reasoning capabilities in out-of-distribution (OOD) and cross-task scenarios.
Jiaao Yu 0001, Shenwei Li, Mingjie Han, Yifei Yin, Wenzheng Song, Chenghao Jia, Man Lan
AAAI3
2026 SecOutPIR: privacy preservation for data owner and access control for data user in outsourced private information retrieval
abstract
Abstract Private Information Retrieval (PIR) is a cryptographic technique that allows Data User (DU) to retrieve data from a SERVER without revealing which specific data item is being accessed. Traditional PIR protocols typically assume that the data is locally stored and directly controlled by Data Owner (DO), but in real-world scenarios, data is often hosted on untrusted third-party SERVERs, making it difficult for DO to effectively restrict the SERVER’s access to their data or control which DU is authorized to retrieve the data. Consequently, malicious SERVERs or unauthorized DU may infringe upon the privacy rights of DO. This paper presents SecOutPIR, a novel outsourced PIR system that addresses two key challenges: privacy preservation for DO and access control for DU. SecOutPIR integrates attribute-based encryption for fine-grained retrieval access control to ensure that only DU with valid retrieval can access the data, while also utilizing a decentralized identity management system based on decentralized identifiers and verifiable credentials to authenticate DU requests. The proposed system ensures that the DO’s data privacy is protected during data storage and retrieval, while also ensuring that only DU with authorized retrieval can make retrieval requests, thus preventing unauthorized access. We provide a detailed description of the system model, security requirements, and an in-depth security analysis. Furthermore, experimental results demonstrate that SecOutPIR significantly enhances the practicality and efficiency of PIR in outsourced settings by enabling fine-grained retrieval access control without degrading query performance. Our implementation demonstrates that the SERVER reply time increases with the dataset size, from 82.5 ms (1000 entries) to 113.8 ms (2000 entries) and 199.6 ms (5000 entries), while the query generation time remains approximately constant at around 2.0 ms.
Fei Tang 0001, Ruixue Li, Huihui Zhu 0001, Mingjie Han
Cybersecur.4
2023 Generating Questions via Unexploited OCR Texts: Prompt-Based Data Augmentation for TextVQA
abstract
Text-based Visual Question Answering (TextVQA) tasks rely on Optical Character Recognition (OCR) text to answer. There have been many models successfully exploring multi-modal features fusing and knowledge reasoning. However, current TextVQA datasets are few and the cost of using manual annotation is too high. So generating pseudo-labeled data is a better choice. In this paper, a prompt-based data augmentation method is proposed. The problems of current data augmentation are solved: 1) the distribution of the number of answer words in the pseudo-labeled data is not consistent with the real dataset. 2) the question forms in the pseudo-labeled data are not diverse. Specifically, prompt words are first matched to the constraints in the questions by finding the same words in the vocabulary. So, our generating model can generate different types of questions when the different prompt words are input. Experiments show that our method is significantly better than other state-of-the-art methods on TextVQA.
Mingjie Han, Wancong Lin, Liang Qiao 0001
IJCNN1
2021 A Generative Model-Based Predictive Display for Robotic Teleoperation
abstract
We propose a new generative model-based predictive display for robotic teleoperation over high-latency communication links. Our method is capable of rendering photo-realistic images of the scene to the human operator in real time from RGB-D images acquired by the remote robot. A preliminary exploration stage is used to build a coarse 3D map of the remote environment and to train a generative model, both of which are then used to generate photo-realistic images for the human operator based on the commanded pose of the robot. Data captured by the remote robot is used to dynamically update the 3D map, enabling teleoperation in the presence of new and relocated objects. Various experiments validate our proposed method’s performance and benefits over alternative methods.
Bowen Xie, Mingjie Han, Jun Jin 0001, Martin Barczyk, Martin Jägersand
ICRA2
2021 Image-Based Joint State Estimation Pipeline for Sensorless Manipulators
abstract
Motion planning is a largely solved problem for robot arms with joint state feedback, but remains an area of research for sensorless manipulators such as toy robot arms and heavy equipment such as excavators and cranes. A promising approach to this problem is deep learning, which employs a pre-trained convolutional neural network to identify manipulator links and estimate joint states from a monocular camera video feed. Whereas manual labeling of training image sets is tedious and non-transferable, a simulation environment can automatically generate labeled training image sets of any size. The issue is the gap between simulated and real-world images. This paper solves this problem by implementing a Generative Adversarial Network. The complete joint state estimation pipeline is implemented and tested in hardware experiments to validate our proposed approach.
Mingjie Han, Bowen Xie, Martin Barczyk, Alireza Bayat
IROS1