VLDB 2026 Research / reviewers in the wild / expert
Minghao Han
dblp:183/0607
· DBLP profile ↗
24ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0002-2027-328XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image ComprehensionabstractSatire, a form of artistic expression combining humor with implicit critique, holds significant social value by illuminating societal issues. Despite its cultural and societal significance, satire comprehension, particularly in purely visual forms, remains a challenging task for current vision-language models. This task requires not only detecting satire but also deciphering its nuanced meaning and identifying the implicated entities. Existing models often fail to effectively integrate local entity relationships with global context, leading to misinterpretation, comprehension biases, and hallucinations. To address these limitations, we propose SatireDecoder, a training-free framework designed to enhance satirical image comprehension. Our approach proposes a multi-agent system performing visual cascaded decoupling to decompose images into fine-grained local and global semantic representations. In addition, we introduce a chain-of-thought reasoning strategy guided by uncertainty analysis, which breaks down the complex satire comprehension process into sequential subtasks with minimized uncertainty. Our method significantly improves interpretive accuracy while reducing hallucinations. Experimental results validate that SatireDecoder outperforms existing baselines in comprehending visual satire, offering a promising direction for vision-language reasoning in nuanced, high-level semantic tasks. Haiwei Xue, Minghao Han, Mingcheng Li, Xiaolu Hou, Dingkang Yang, Lihua Zhang 0002, Xu Zheng 0002 |
AAAI | 3 |
| 2026 | Nested resolution mesh-graph CNN for automated extraction of liver surface anatomical landmarksabstractThe anatomical landmarks on the liver (mesh) surface, including the falciform ligament and liver ridge, are composed of triangular meshes of varying shapes, sizes, and positions, making them highly complex. Extracting and segmenting these landmarks is critical for augmented reality-based intraoperative navigation and monitoring. The key to this task lies in comprehensively understanding the overall geometric shape and local topological information of the liver mesh. However, due to the liver's variations in shape and appearance, coupled with limited data, deep learning methods often struggle with automatic liver landmark segmentation. To address this, we propose a two-stage automatic framework combining mesh-CNN and graph-CNN. In the first stage, dynamic graph convolution (DGCNN) is employed on low-resolution meshes to achieve rapid global understanding, generating initial landmark proposals at two levels, "dilation" and "erosion", and mapping them onto the original high-resolution surface. Subsequently, a refinement network based on mesh convolution fuses these landmark proposals from edge features along the local topology of the high-resolution mesh surface, producing refined segmentation results. Additionally, we incorporate an anatomy-aware Dice loss to address resolution imbalance and better handle sparse anatomical regions. Extensive experiments on two liver datasets, both in-distribution and out-of-distribution, demonstrate that our method accurately processes liver meshes of different resolutions, outperforming state-of-the-art methods. The reconstructed liver mesh dataset and the source code are available at https://github.com/xukun-zhang/MeshGraphCNN. Xukun Zhang, Jinghui Feng, Peng Liu 0074, Minghao Han, Yanlan Kang, Sharib Ali, Lihua Zhang 0002 |
Medical Image Anal. | 4 |
| 2026 | Towards unified molecule-enhanced pathology image representation learning via integrating spatial transcriptomics
Minghao Han, Dingkang Yang, Jiabei Cheng, Xukun Zhang, Zizhi Chen, Haopeng Kuang, Lihua Zhang 0002 |
Pattern Recognit. | 1 |
| 2025 | MamKO: Mamba-based Koopman operator for modeling and predictive controlabstractThe Koopman theory, which enables the transformation of nonlinear systems into linear representations, is a powerful and efficient tool to model and control nonlinear systems. However, the ability of the Koopman operator to model complex systems, particularly time-varying systems, is limited by the fixed linear state-space representation. To address the limitation, the large language model, Mamba, is considered a promising strategy for enhancing modeling capabilities while preserving the linear state-space structure.
In this paper, we propose a new framework, the Mamba-based Koopman operator (MamKO), which provides enhanced model prediction capability and adaptability, as compared to Koopman models with constant Koopman operators. Inspired by the Mamba structure, MamKO generates Koopman operators from online data; this enables the model to effectively capture the dynamic behaviors of the nonlinear system over time. A model predictive control system is then developed based on the proposed MamKO model. The modeling and control performance of the proposed method is evaluated through experiments on benchmark time-invariant and time-varying systems. The experimental results demonstrate the superiority of the proposed approach. Additionally, we perform ablation experiments to test the effectiveness of individual components of MamKO. This approach unlocks new possibilities for integrating large language models with control frameworks, and it achieves a good balance between advanced modeling capabilities and real-time control implementation efficiency. Minghao Han, Xunyuan Yin |
ICLR | 2 |
| 2025 | VGAT: A Cancer Survival Analysis Framework Transitioning from Generative Visual Question Answering to Genomic ReconstructionabstractMultimodal learning combining pathology images and genomic sequences enhances cancer survival analysis but faces clinical implementation barriers due to limited access to genomic sequencing in under-resourced regions. To enable survival prediction using only whole-slide images (WSI), we propose the Visual-Genomic Answering-Guided Transformer (VGAT), a framework integrating Visual Question Answering (VQA) techniques for genomic modality reconstruction. By adapting VQA’s text feature extraction approach, we derive stable genomic representations that circumvent dimensionality challenges in raw genomic data. Simultaneously, a cluster-based visual prompt module selectively enhances discriminative WSI patches, addressing noise from unfiltered image regions. Evaluated across five TCGA datasets, VGAT outperforms existing WSI-only methods, demonstrating the viability of genomic-informed inference without sequencing. This approach bridges multimodal research and clinical feasibility in resource-constrained settings. The code link is https://github.com/CZZZZZZZZZZZZZZZZZ/VGAT. Zizhi Chen, Minghao Han, Xukun Zhang, Shuwei Ma, Tao Liu 0050, Lihua Zhang 0002 |
ICME | 2 |
| 2025 | Economic Model Predictive Control of Time-Varying Nonlinear Systems Using Transformer-based Koopman OperatorabstractTime-varying systems commonly exist in modern industrial processes. This paper addresses the problem of learning-based modeling and economic control of time-varying nonlinear systems. By developing a deep time-varying Koopman operator model, the future information of the system related to economic costs and critical outputs is learned directly from data. Transformer architecture is employed to learn the observable functions and to generate time-varying Koopman operators. An efficient economic model predictive control (EMPC) problem is formulated based on the learned transformer-based Koopman model to achieve the economic operations of the system. The proposed method is applied to a membrane-based wastewater treatment process. The performance of the proposed method is compared to the baseline. Minghao Han, Minh Thu Hoang, Ryan Rui En Tan, Xunyuan Yin |
IECON | 1 |
| 2025 | VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual StainingabstractIn histopathology, tissue sections are typically stained using common H&E staining or special stains (MAS, PAS, PASM, etc. ) to clearly visualize specific tissue structures. The rapid advancement of deep learning offers an effective solution for generating virtually stained images, significantly reducing the time and labor costs associated with traditional histochemical staining. However, a new challenge arises in separating the fundamental visual characteristics of tissue sections from the visual differences induced by staining agents. Additionally, virtual staining often overlooks essential pathological knowledge and the physical properties of staining, resulting in only style-level transfer. To address these issues, we introduce, for the first time in virtual staining tasks, a pathological vision-language large model (VLM) as an auxiliary tool. We integrate contrastive learnable prompts, foundational concept anchors for tissue sections, and staining-specific concept anchors to leverage the extensive knowledge of the pathological VLM. This approach is designed to describe, frame, and enhance the direction of virtual staining. Furthermore, we have developed a data augmentation method based on the constraints of the VLM. This method utilizes the VLM's powerful image interpretation capabilities to further integrate image style and structural information, proving beneficial in high-precision pathological diagnostics. Extensive evaluations on publicly available multi-domain unpaired staining datasets demonstrate that our method can generate highly realistic images and enhance the accuracy of downstream tasks, such as glomerular detection and segmentation. Our code. https://github.com/CZZZZZZZZZZZZZZZZZ/VPGAN-HARBOR is available. Zizhi Chen, Minghao Han, Yizhou Liu 0002, Ziyun Qian, Xukun Zhang, Jingwei Wei, Lihua Zhang 0002 |
ACM Multimedia | 3 |
| 2025 | MSCPT: Few-Shot Whole Slide Image Classification With Multi-Scale and Context-Focused Prompt TuningabstractMultiple instance learning (MIL) has become a standard paradigm for the weakly supervised classification of whole slide images (WSIs). However, this paradigm relies on using a large number of labeled WSIs for training. The lack of training data and the presence of rare diseases pose significant challenges for these methods. Prompt tuning combined with pre-trained Vision-Language models (VLMs) is an effective solution to the Few-shot Weakly Supervised WSI Classification (FSWC) task. Nevertheless, applying prompt tuning methods designed for natural images to WSIs presents three significant challenges: 1) These methods fail to fully leverage the prior knowledge from the VLM's text modality; 2) They overlook the essential multi-scale and contextual information in WSIs, leading to suboptimal results; and 3) They lack exploration of instance aggregation methods. To address these problems, we propose a Multi-Scale and Context-focused Prompt Tuning (MSCPT) method for FSWC task. Specifically, MSCPT employs the frozen large language model to generate pathological visual language prior knowledge at multiple scales, guiding hierarchical prompt tuning. Additionally, we design a graph prompt tuning module to learn essential contextual information within WSI, and finally, a non-parametric cross-guided instance aggregation module has been introduced to derive the WSI-level features. Extensive experiments, visualizations, and interpretability analyses were conducted on five datasets and three downstream tasks using three VLMs, demonstrating the strong performance of our MSCPT. All codes have been made publicly accessible at https://github.com/Hanminghao/MSCPT. Minghao Han, Linhao Qu, Dingkang Yang, Xukun Zhang, Lihua Zhang 0002 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Robust Learning-Based Control for Uncertain Nonlinear Systems With Validation on a Soft RobotabstractExisting modeling and control methods for real-world systems typically deal with uncertainty and nonlinearity on a case-by-case basis. We present a universal and robust control framework for the general class of uncertain nonlinear systems. Our data-driven deep stochastic Koopman operator (DeSKO) model and robust learning control framework guarantee robust stability. DeSKO learns the uncertainty of dynamical systems by inferring a distribution of observables. The inferred distribution is used in our robust and stabilizing closed-loop controller for dynamical systems. We also develop a model predictive control framework with integral action to compensate for run-time parametric uncertainty, such as manipulating unknown objects. Modeling and control experiments in simulation show that our presented framework is more robust and scalable for robotic systems than state-of-the-art controllers using deep Koopman operators and reinforcement learning (RL) methods. We demonstrate that our method resists previously unseen uncertainties, such as external disturbances, at a magnitude of up to five times the maximum control input. Furthermore, we test our DeSKO-based control framework on a real-world soft robotic arm. It shows that our framework outperforms model-based controllers that have full knowledge of the model parameters, and the controller can conduct object pick-and-place tasks without further training. Our approach opens up new possibilities in robustly managing internal or external uncertainty while controlling high-dimensional nonlinear systems in a learning framework. This approach serves as a foundation to greatly simplify high-level control and decision-making for robots. Minghao Han, Kiwan Wong, Jacob Euler-Rolle, Lixian Zhang 0001, Robert K. Katzschmann |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Koopman-Constrained Hierarchical Deep State Space Model for Industrial Quality Prediction via Cloud-Edge Collaborative FrameworkabstractIn cloud manufacturing of industrial processes, the accurate online prediction of product quality is the basis for realizing decision-making and control of the manufacturing process. However, frequent fluctuations in working conditions and data noise restrict the application of data-driven methods in industrial sites. In addition, the constrained resources on edge devices limit their ability to automatically update or deploy complex models. To address these issues, this study proposes a Koopman-constrained hierarchical deep state-space model (KHSSM) and incorporates it into the innovative cloud-edge collaboration framework for industrial quality prediction. First, KHSSM integrates a state-space model, leveraging its advantage in modeling noisy dynamic data. Second, the Koopman operator is introduced to constrain the latent variables in the measurement space, enabling it to interpretably reflect the evolution dynamics of the system. In addition, novel strategies for model mismatch detection and model simplification are designed and deployed to improve the predictive accuracy and real-time efficiency of the cloud-edge collaboration framework. Finally, the effectiveness of the proposed method is verified by extensive experiments in a numerical simulation and a real-world industrial process. Qingkai Sui, Yalin Wang 0003, Chenliang Liu, Minghao Han, Chunhua Yang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2024 | Multi-Scale Heterogeneity-Aware Hypergraph Representation for Histopathology Whole Slide ImagesabstractSurvival prediction is a complex ordinal regression task that aims to predict the survival coefficient ranking among a cohort of patients, typically achieved by analyzing patients’ whole slide images. Existing deep learning approaches mainly adopt multiple instance learning or graph neural networks under weak supervision. Most of them are unable to uncover the diverse interactions between different types of biological entities(e.g., cell cluster and tissue block) across multiple scales, while such interactions are crucial for patient survival prediction. In light of this, we propose a novel multi-scale heterogeneity-aware hypergraph representation framework. Specifically, our framework first constructs a multi-scale heterogeneity-aware hypergraph and assigns each node with its biological entity type. It then mines diverse interactions between nodes on the graph structure to obtain a global representation. Experimental results demonstrate that our method outperforms state-of-the-art approaches on three benchmark datasets. Code is publicly available at https://github.com/Hanminghao/H2GT. Minghao Han, Xukun Zhang, Dingkang Yang, Tao Liu 0050, Haopeng Kuang, Jinghui Feng, Lihua Zhang 0002 |
ICME | 1 |
| 2024 | Causal Context Adjustment Loss for Learned Image CompressionabstractIn recent years, learned image compression (LIC) technologies have surpassed conventional methods notably in terms of rate-distortion (RD) performance. Most present learned techniques are VAE-based with an autoregressive entropy model, which obviously promotes the RD performance by utilizing the decoded causal context. However, extant methods are highly dependent on the fixed hand-crafted causal context. The question of how to guide the auto-encoder to generate a more effective causal context benefit for the autoregressive entropy models is worth exploring. In this paper, we make the first attempt in investigating the way to explicitly adjust the causal context with our proposed Causal Context Adjustment loss (CCA-loss). By imposing the CCA-loss, we enable the neural network to spontaneously adjust important information into the early stage of the autoregressive entropy model. Furthermore, as transformer technology develops remarkably, variants of which have been adopted by many state-of-the-art (SOTA) LIC techniques. The existing computing devices have not adapted the calculation of the attention mechanism well, which leads to a burden on computation quantity and inference latency. To overcome it, we establish a convolutional neural network (CNN) image compression model and adopt the unevenly channel-wise grouped strategy for high efficiency. Ultimately, the proposed CNN-based LIC network trained with our Causal Context Adjustment loss attains a great trade-off between inference latency and rate-distortion performance. Minghao Han, Shiyin Jiang, Shengxi Li, Xin Deng 0002, Mai Xu, Ce Zhu, Shuhang Gu |
NeurIPS | 1 |
| 2024 | Dual knowledge-guided two-stage model for precise small organ segmentation in abdominal CT imagesabstractAbstract Multi‐organ segmentation from abdominal CT scans is crucial for various medical examinations and diagnoses. Despite the remarkable achievements of existing deep‐learning‐based methods, accurately segmenting small organs remains challenging due to their small size and low contrast. This article introduces a novel knowledge‐guided cascaded framework that utilizes two types of knowledge—image intrinsic (anatomy) and clinical expertise (radiology)—to improve the segmentation accuracy of small abdominal organs. Specifically, based on the anatomical similarities in abdominal CT scans, the approach employs entropy‐based registration techniques to map high‐quality segmentation results onto inaccurate results from the first stage, thereby guiding precise localization of small organs. Additionally, inspired by the practice of annotating images from multiple perspectives by radiologists, novel Multi‐View Fusion Convolution (MVFC) operator is developed, which can extract and adaptively fuse features from various directions of CT images to refine segmentation of small organs effectively. Simultaneously, the MVFC operator offers a seamless alternative to conventional convolutions within diverse model architectures. Extensive experiments on the Abdominal Multi‐Organ Segmentation (AMOS) dataset demonstrate the superiority of the method, setting a new benchmark in the segmentation of small organs. Tao Liu 0050, Xukun Zhang, Zhongwei Yang, Minghao Han, Haopeng Kuang, Shuwei Ma, Lihua Zhang 0002 |
IET Image Process. | 4 |
| 2024 | Quantized output-feedback control of piecewise-affine systems with reachable regions of quantized measurements
Zepeng Ning, Wenqiang Ji, Yanzheng Zhu, Minghao Han |
Inf. Sci. | 5 |
| 2024 | Robust Learning and Control of Time-Delay Nonlinear Systems With Deep Recurrent Koopman OperatorsabstractIn this work, we consider the problem of Koopman modeling and data-driven predictive control for a class of uncertain nonlinear systems subject to time delays. A robust deep learning-based approach–deep recurrent Koopman operator is proposed. Without requiring the knowledge of system uncertainties or information on the time delays, the proposed deep recurrent Koopman operator method is able to learn the dynamics of the nonlinear systems autonomously. A robust predictive control framework is established based on the deep Koopman operator. Conditions on the stability of the closed-loop system are presented. The proposed approach is applied to a chemical process example. The results confirm the superiority of the proposed framework as compared to baselines. Minghao Han, Zhaojian Li 0001, Xiang Yin 0003, Xunyuan Yin |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Anatomical-Aware Point-Voxel Network for Couinaud Segmentation in Liver CT
Xukun Zhang, Yang Liu 0007, Sharib Ali, Minghao Han, Tao Liu 0050, Peng Zhai, Zhiming Cui 0001, Peixuan Zhang, Lihua Zhang 0002 |
MICCAI (3) | 6 |
| 2023 | Data-Driven Linear Predictive Control of Nonlinear Processes Based on Reduced-Order Koopman OperatorabstractIn this paper, we propose an efficient data-driven predictive control approach for general nonlinear processes based on a reduced-order Koopman operator. A Kalman-based sparse identification of nonlinear dynamics method is employed to select lifting functions for Koopman identification. The selected lifting functions are used to project the original nonlinear state space into a higher-dimensional linear function space, in which Koopman-based linear models may be constructed for the underlying nonlinear process. To address the potential issue of a significant increase in the dimensionality of the resulting full-order Koopman models caused by the use of lifting functions, we propose a reduced-order Koopman modeling approach based on proper orthogonal decomposition. A computationally efficient linear robust predictive control scheme is established based on the reduced-order Koopman model. A case study on a benchmark chemical process is conducted to illustrate the proposed framework. Xuewen Zhang, Minghao Han, Xunyuan Yin |
SMC | 2 |
| 2022 | DeSKO: Stability-Assured Robust Control with a Deep Stochastic Koopman Operator
Minghao Han, Jacob Euler-Rolle, Robert K. Katzschmann |
ICLR | 1 |
| 2021 | Poster: A Real-time Social Distance Measurement and Record System for COVID-19
Weijun Wang 0001, Tingting Yuan 0001, Minghao Han, Meng Li 0010, Sripriya Srikant Adhatarao, Xiaoming Fu 0001 |
EWSN | 3 |
| 2021 | Safe Reinforcement Learning With Stability Guarantee for Motion Planning of Autonomous VehiclesabstractReinforcement learning with safety constraints is promising for autonomous vehicles, of which various failures may result in disastrous losses. In general, a safe policy is trained by constrained optimization algorithms, in which the average constraint return as a function of states and actions should be lower than a predefined bound. However, most existing safe learning-based algorithms capture states via multiple high-precision sensors, which complicates the hardware systems and is power-consuming. This article is focused on safe motion planning with the stability guarantee for autonomous vehicles with limited size and power. To this end, the risk-identification method and the Lyapunov function are integrated with the well-known soft actor-critic (SAC) algorithm. By borrowing the concept of Lyapunov functions in the control theory, the learned policy can theoretically guarantee that the state trajectory always stays in a safe area. A novel risk-sensitive learning-based algorithm with the stability guarantee is proposed to train policies for the motion planning of autonomous vehicles. The learned policy is implemented on a differential drive vehicle in a simulation environment. The experimental results show that the proposed algorithm achieves a higher success rate than the SAC. Lixian Zhang 0001, Ruixian Zhang, Tong Wu 0013, Rui Weng, Minghao Han, Ye Zhao 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Self-paced and auto-weighted multi-view clustering
Yazhou Ren 0001, Shudong Huang, Minghao Han, Zenglin Xu |
Neurocomputing | 4 |
| 2018 | Distributed State Estimation of Sensor-Network Systems Subject to Markovian Channel Switching With Application to a Chemical ProcessabstractThis paper addresses a distributed estimator design problem for linear systems deployed over sensor networks within a multiple communication channels (MCCs) framework. A practical scenario is taken into account such that the channel used for communication can be switched and the switching is governed by a Markov chain. With the existence of communicational imperfections and external disturbances, an estimation algorithm is proposed such that the developed distributed estimators are able to give accurate state estimates against the channel switching phenomenon. The distributed estimation framework is applied to a chemical process to illustrate the effectiveness of the proposed methodology and the superiority of the MCCs framework featured by channel switching. Xunyuan Yin, Zhaojian Li 0001, Lixian Zhang 0001, Minghao Han |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2017 | Semi-time-dependent asynchronously switched control of continuous-time switched systems with persistent dwell timeabstractThis paper mainly concerns the issue of asynchronously switched control for a class of continuous-time switched linear systems with persistent dwell time (PDT). Asynchronous switching implies that when a subsystem switches the matched controller remains active for a finite period of time, such that the Lyapunov-like function increases. A semi-time-dependent (STD) multiple Lyapunov-like function with special form suitable to asynchronously switched systems is proposed, upon which a set of STD and mode-dependent stabilizing controllers is designed. The validity and advantage of the proposed results is demonstrated through a numerical example. Tianhe Liu, Changhong Wang 0003, Zeyang Fan, Minghao Han |
IECON | 5 |
| 2016 | Hierarchical Random Walk Inference in Knowledge GraphsabstractRelational inference is a crucial technique for knowledge base population. The central problem in the study of relational inference is to infer unknown relations between entities from the facts given in the knowledge bases. Two popular models have been put forth recently to solve this problem, which are the latent factor models and the random-walk models, respectively. However, each of them has their pros and cons, depending on their computational efficiency and inference accuracy. In this paper, we propose a hierarchical random-walk inference algorithm for relational learning in large scale graph-structured knowledge bases, which not only maintains the computational simplicity of the random-walk models, but also provides better inference accuracy than related works. The improvements come from two basic assumptions we proposed in this paper. Firstly, we assume that although a relation between two entities is syntactically directional, the information conveyed by this relation is equally shared between the connected entities, thus all of the relations are semantically bidirectional. Secondly, we assume that the topology structures of the relation-specific subgraphs in knowledge bases can be exploited to improve the performance of the random-walk based relational inference algorithms. The proposed algorithm and ideas are validated with numerical results on experimental data sampled from practical knowledge bases, and the results are compared to state-of-the-art approaches. Qiao Liu 0003, Liuyi Jiang, Minghao Han, Yao Liu 0019, Zhiguang Qin |
SIGIR | 3 |