VLDB 2026 Research / reviewers in the wild / expert
Qinghai Miao
dblp:33/1250
· DBLP profile ↗
19ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0003-1213-1123ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 48% Vision and language · 20% Representation and self-supervised learning · 10% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › prompt tuning
context optimization |
0.9 | 1 | 2025 | Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models · IEEE Trans. Multim. 2025 |
Machine learning › Reinforcement learning › offline reinforcement learning
model-based offline reinforcement learning |
0.9 | 1 | 2025 | Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.9 | 1 | 2025 | Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025 |
Computer vision › Vision and language › vision-language model
prompt learning |
0.9 | 1 | 2025 | Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models · IEEE Trans. Multim. 2025 |
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
0.9 | 1 | 2025 | Rank-Awareness and Angular Constraints: A New Perspective on Learning Sentence Embeddings from NLI Data · EMNLP 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models · IEEE Trans. Multim. 2025 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
0.9 | 1 | 2025 | Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
0.8 | 1 | 2024 | RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences · ICML 2024 |
Machine learning › Reinforcement learning
reward learning |
0.8 | 1 | 2024 | RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences · ICML 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.8 | 1 | 2024 | RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences · ICML 2024 |
Machine learning › Efficient and distributed learning › data-efficient learning
data-efficient fine-tuning |
0.3 | 1 | 2025 | Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025 |
Information retrieval › similarity measure
semantic textual similarity |
0.3 | 1 | 2025 | Rank-Awareness and Angular Constraints: A New Perspective on Learning Sentence Embeddings from NLI Data · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
natural language inference · 1.7contrastive learning · 1.7transformer · 0.9temporal difference learning · 0.9semantic knowledge mapper · 0.9prompt tuning · 0.9planning · 0.9warm start · 0.8sample selection · 0.8discriminator · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating the Perceptual Robustness of Vision-Language Models for Autonomous Driving in Corner Cases
Peizhe Gong, Enming Zhang, Ruixi Qiao, Xingyuan Dai, Xiaoyan Gong, Qinghai Miao |
IV | 6 |
| 2025 | Rank-Awareness and Angular Constraints: A New Perspective on Learning Sentence Embeddings from NLI DataabstractLearning high-quality sentence embeddings from Natural Language Inference (NLI) data is often challenged by a critical signal conflict between discrete labels and the continuous spectrum of semantic similarity, as well as information loss from discarded neutral sentence pairs during training.To address this, we introduce Rank-Awareness and Angular Optimization Embeddings (RAOE), a framework that leverages the full NLI dataset (Entailment, Neutral, Contradiction) augmented with precomputed continuous similarity scores (S).RAOE employs a novel composite objective which features: (1) a Rank Margin objective that enforces rank consistency against S using an explicit margin, and (2) a Gated Angular objective that conditionally refines embedding geometry based on NLI label (L) and S score agreement.Extensive evaluations on STS tasks and the MTEB benchmark demonstrate RAOE's effectiveness.Our generalpurpose RAOE-S1 model (BERT-base) significantly outperforms strong baselines, achieving an average Spearman's correlation of 85.11 (vs.SimCSE's 81.57and AnglE's 82.43), and shows consistent improvements on MTEB.Further STS-specialized fine-tuning (RAOE-S2) establishes new state-of-the-art performance on STS (88.17 with BERT-base).These results confirm RAOE's ability to efficiently learn robust and nuanced sentence representations through the synergy of rankawareness and conditional angular constraints. Zicheng Zhou, Min Huang 0009, Qinghai Miao |
EMNLP | 3 |
| 2025 | Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model PretrainingabstractA significant aspiration of offline reinforcement learning (RL) is to develop a generalist agent with high capabilities from large and heterogeneous datasets. However, prior approaches that scale offline RL either rely heavily on expert trajectories or struggle to generalize to diverse unseen tasks. Inspired by the excellent generalization of world model in conditional video generation, we explore the potential of image observation-based world model for scaling offline RL and enhancing generalization on novel tasks. In this paper, we introduce JOWA: Jointly-Optimized World-Action model, an offline model-based RL agent pretrained on multiple Atari games with 6 billion tokens data to learn general-purpose representation and decision-making ability. Our method jointly optimizes a world-action model through a shared transformer backbone, which stabilize temporal difference learning with large models during pretraining. Moreover, we propose a provably efficient and parallelizable planning algorithm to compensate for the Q-value estimation error and thus search out better policies. Experimental results indicate that our largest agent, with 150 million parameters, achieves 78.9% human-level performance on pretrained games using only 10% subsampled offline data, outperforming existing state-of-the-art large-scale offline RL baselines by 31.6% on averange. Furthermore, JOWA scales favorably with model capacity and can sample-efficiently transfer to novel games using only 5k offline fine-tuning data (approximately 4 trajectories) per game, demonstrating superior generalization. Jie Cheng 0009, Ruixi Qiao, Yingwei Ma, Binhua Li, Gang Xiong 0001, Qinghai Miao |
ICLR | 6 |
| 2025 | MiniDrive: More Efficient Vision-Language Models with Multi-level 2D Features as Text Tokens for Autonomous Driving
Enming Zhang, Xingyuan Dai, Min Huang 0009, Qinghai Miao |
PRCV (11) | 5 |
| 2025 | Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language ModelsabstractVision-language models (VLMs), such as CLIP, play a foundational role in various cross-modal applications. To fully leverage the potential of VLMs in adapting to downstream tasks, context optimization methods such as prompt tuning are essential. However, one key limitation is the lack of diversity in prompt templates, whether they are hand-crafted or learned through additional modules. This limitation restricts the capabilities of pretrained VLMs and can result in incorrect predictions in downstream tasks. To address this challenge, we propose context optimization with multi-knowledge representation (CoKnow), a framework that enhances prompt learning for VLMs with rich contextual knowledge. To facilitate CoKnow during inference, we train lightweight semantic knowledge mappers, which are capable of generating multi-knowledge representations for an input image without requiring additional priors. Experimentally, we conduct extensive experiments on 11 publicly available datasets, demonstrating that CoKnow outperforms a series of previous methods. Enming Zhang, Bingke Zhu, Yingying Chen 0003, Qinghai Miao, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Multim. | 4 |
| 2024 | DimA: A Parameter-efficient Fine-tuning Method with Knowledge Transfer Based on TransformerabstractFine-tuning is a widely used technique for leveraging pre-trained language models (PLMs) in downstream tasks, but it can be computationally expensive and storage-intensive. To address this challenge, researchers have developed parameter-efficient methods that balance performance and resource cost. However, these methods often come with trade-offs like increased inference latency, token length usage, or limited adaptability for multitasking scenarios. This paper introduces a novel parameter-efficient method called DimA(Dimensionality Augmentation), which enhances the Transformer architecture by increasing the dimensionality. DimA achieves state-of-the-art results in GLUE and XSUM tasks while utilizing less than 1% of the original model’s parameters. Moreover, DimA introduces a novel approach to knowledge transfer that enables the simultaneous utilization of knowledge learned from multiple tasks to handle new tasks. This method significantly enhances the performance of the model on new tasks. Its versatility in model structure also enables its application to various Transformer-based models. Min Huang 0009, Zhuoyang Song, Qinghai Miao |
LREC/COLING | 4 |
| 2024 | RIME: Robust Preference-based Reinforcement Learning with Noisy PreferencesabstractPreference-based Reinforcement Learning (PbRL) circumvents the need for reward engineering by harnessing human preferences as the reward signal. However, current PbRL methods excessively depend on high-quality feedback from domain experts, which results in a lack of robustness. In this paper, we present RIME, a robust PbRL algorithm for effective reward learning from noisy preferences. Our method utilizes a sample selection-based discriminator to dynamically filter out noise and ensure robust training. To counteract the cumulative error stemming from incorrect selection, we suggest a warm start for the reward model, which additionally bridges the performance gap during the transition from pre-training to online training in PbRL. Our experiments on robotic manipulation and locomotion tasks demonstrate that RIME significantly enhances the robustness of the state-of-the-art PbRL method. Code is available at https://github.com/CJReinforce/RIME_ICML2024. Jie Cheng 0009, Gang Xiong 0001, Xingyuan Dai, Qinghai Miao, Fei-Yue Wang 0001 |
ICML | 4 |
| 2024 | Mitigating Spurious Correlations in Named Entity Recognition Models Through Counterfactual Data AugmentationabstractNamed Entity Recognition (NER) models perform well on standard benchmarks, but they often lack robustness when dealing with out-of-domain data. Recent studies have highlighted their limitations in genuine sentence comprehension, often due to their reliance on memorized entities or context, leading to spurious correlations. In this paper, we present Structural Causal Model-based techniques for detecting and addressing two types of spurious correlations in NER models: contextual and entity spurious correlations. Additionally, we introduce four data augmentation methods to improve training data and enable NER models to learn contextual and entity information effectively. Experimental results reveal that several NER models suffer from spurious correlations, and our approach mitigates them effectively. Furthermore, in comparison to the baseline methods, our approach demonstrates competitive performance on the in-domain dataset while achieving state-of-the-art results on the out-of-domain dataset. Min Huang 0009, Qinghai Miao |
IJCNN | 3 |
| 2024 | Parallel Learning for Legal Intelligence: A HANOI Approach Based on Unified PromptingabstractPretrained language models (PLMs) have made significant progress on various NLP tasks recently. However, PLMs encounter challenges when it comes to domain-specific tasks such as legal AI. These tasks often involve intricate expertise, expensive data annotation, and limited training data availability. To tackle this problem, we propose a human-oriented artificial–natural parallel system for organized intelligence (HANOI)-Legal based on the parallel learning (PL) framework. First, by regarding the description in PL as the pretraining process based on a large-scale corpus, we setup an artificial system based on a PLM. Second, to adapt the PLM to legal tasks with limited resources, we propose UniPrompt as a prescription. UniPrompt serves as a unified prompt-based training framework, enabling the utilization of diverse open datasets for these tasks. Third, we labeled a few task-specific legal data through distributed autonomous operations (DAO-II) for further fine-tuning. By combining a scalable unified-task-format reformulation and a unified-prompt-based training pipeline, HANOI-Legal leverages PLMs’ linguistic capabilities acquired from a variety of open datasets to generate task-specific models. Our experiments in two legal domain tasks show that HANOI-Legal achieved an excellent performance in low-resource scenarios compared to the state-of-the-art prompt-based approach. Zhuoyang Song, Min Huang 0009, Qinghai Miao, Fei-Yue Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | TrajSGAN: A Semantic-Guiding Adversarial Network for Urban Trajectory GenerationabstractSimulating human mobility contributes to city behavior discovery and decision-making. Although the sequence-based and image-based approaches have made impressive achievements, they still suffer from respective deficiencies such as omitting the depiction of spatial properties or ordinal dependency in trajectory. In this article, we take advantage of the above two paradigms and propose a semantic-guiding adversarial network (TrajSGAN) for generating human trajectories. Specifically, we first devise an attention-based generator to yield trajectory locations in a sequence-to-sequence manner. The encoded historical visits are queried with semantic knowledge (e.g., travel modes and trip purposes) and their important features are enhanced by the multihead attention mechanism. Then, we designate a rollout module to complete the unfinished trajectory sequence and transform it into an image that can depict its spatial structure. Finally, a convolutional neural network (CNN)-based discriminator signifies how “real” the trajectory image looks, and its output is regarded as a reward signal to update the generator by the policy gradient. Experimental results show that the proposed TrajSGAN model significantly outperforms the benchmarks under the MTL-Trajet mobility dataset, with the divergence of spatial-related metrics such as radius of gyration and travel distance reduced by 10%–27%. Furthermore, we apply the real and synthetic trajectories, respectively, to simulate the COVID-19 epidemic spreading under three preventive actions. The coefficient of determination metric between real and synthetic results achieves 91%–98%, indicating that the synthesized data from TrajSGAN can be leveraged to study the epidemic diffusion with an acceptable difference. All of these results verify the superiority and utility of our proposed method. Gang Xiong 0001, Zhishuai Li, Meihua Zhao, Qinghai Miao, Fei-Yue Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2023 | Instance-Proxy Loss for Semi-supervised Learning with Coarse Labels
Qinghai Miao, Haiyun Guo, Min Huang 0009, Jinqiao Wang |
PRCV (12) | 2 |
| 2022 | Hybrid Attention-based Transformer for Long-range Document ClassificationabstractTransformer with the self-attention mechanism, which allows fully-connected contextual encoding over input tokens, has achieved outstanding performances in various NLP tasks, but it suffers from quadratic complexity with the input sequence length. Long-range contexts are often tackled by Transformer in chunks using a sliding window to avoid GPU memory overflow. However, how to achieve considerable performance on downstream tasks under the premise of modeling sequences as long as possible with limited GPU resources is still a problem to be explored. To address this issue, we propose a new framework using hybrid attention-based Transformer to capture long-range contextual features. More specifically, we make a combination of three types of attention, i.e. sliding window local attention, clustering-based long-range attention and specific global attention. Experiments comparing the performance of our model with mainstream efficient improved Transformers are conducted on document classification task in public datasets IMDB and CAIL-Long. Experimental results demonstrate that the proposed approach can outperform the state-of-the-art models on the adopted datasets, which verifies the effectiveness of our model on long-range document classification task. Ruyu Qin, Min Huang 0009, Qinghai Miao |
IJCNN | 4 |
| 2022 | Graph Neural Networks Based Multi-granularity Feature Representation Learning for Fine-Grained Visual Categorization
Haiyun Guo, Qinghai Miao, Min Huang 0009, Jinqiao Wang |
MMM (2) | 3 |
| 2022 | Multi-Granularity Mutual Learning Network for Object Re-IdentificationabstractObject re-identification (re-ID), which is key and fundamental technology for intelligent transportation systems, is a challenging task including person re-ID and vehicle re-ID. It aims to retrieve a given target object from the gallery images captured by different cameras. In this task, it is necessary to extract fine-grained and discriminative features to deal with complex inter-class and intra-class variations caused by the changes of camera viewpoints and object poses. Existing methods focus on learning discriminative local features to improve the re-ID performance. Some state-of-the-art methods use key point detection model to locate local features, which also increases the additional computational cost as side effect. Another type of method focuses on how to learn features of different granularity from rigid stripes of different scales. However, there is little attention paid to how to effectively coalesce multi-granularity features without additional calculation cost. To tackle this issue, this paper proposes the Multi-granularity Mutual Learning Network (MMNet) and makes two contributions. 1) We introduce the multi-granularity jigsaw puzzle module into object re-ID to impel the network to learn local discriminative features from multiple visual granularities by breaking spatial correlation in original images. 2) We propose a parameter-free multi-scale feature reconstruction module to facilitate mutual learning of features at multiple grain levels, thereby both global features and local features have strong representation capabilities. Extensive experiments demonstrate the effectiveness of our proposed modules and the superiority of our method over various state-of-the-art methods on both person and vehicle re-ID benchmarks. Mingfei Tu, Kuan Zhu, Haiyun Guo, Qinghai Miao, Chaoyang Zhao, Guibo Zhu, Honglin Qiao, Gaopan Huang, Ming Tang 0001, Jinqiao Wang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | Displacement-Correlated XFEM for Simulating Brittle FractureabstractAbstract We present a remeshing‐free brittle fracture simulation method under the assumption of quasi‐static linear elastic fracture mechanics (LEFM). To achieve this, we devise two algorithms. First, we develop an approximate volumetric simulation, based on the extended Finite Element Method (XFEM), to initialize and propagate Lagrangian crack‐fronts. We model the geometry of fracture explicitly as a surface mesh, which allows us to generate high‐resolution crack surfaces that are decoupled from the resolution of the deformation mesh. Our second contribution is a mesh cutting algorithm, which produces fragments of the input mesh using the fracture surface. We do this by directly operating on the half‐edge data structures of two surface meshes, which enables us to cut general surface meshes including those of concave polyhedra and meshes with abutting concave polygons. Since we avoid triangulation for cutting, the connectivity of the resulting fragments is identical to the (uncut) input mesh except at edges introduced by the cut. We evaluate our simulation and cutting algorithms and show that they outperform state‐of‐the‐art approaches both qualitatively and quantitatively. Floyd M. Chitalu, Qinghai Miao, Kartic Subr, Taku Komura |
Comput. Graph. Forum | 2 |
| 2019 | Pose-Weighted Gan for Photorealistic Face FrontalizationabstractFace recognition methods have achieved high accuracy when faces are captured in frontal pose and constrained scenes. However, severe drop in accuracy is observed when large pose variations exist. The main reason is that the large yaw angle leads to ID information loss. In this paper, we intend to solve the large pose variations in a generation manner. Specifically, we propose a Pose-Weighted Generative Adversarial Network (PW-GAN) for photorealistic frontal view synthesis. We find frontalizing the faces in large poses (yaw angle larger than 60°) is so difficult that the results are not photorealistic and the ID information is lost. To simplify the problem, we first frontalize the face image through 3D face model, which is then used to guide the network predicting. Second, we refine the pose code in the loss function to make the network pay more attention to large poses. Quantitative and qualitative experimental results on the Multi-PIE and LFW demonstrate our method achieves state of the art. Su-Fang Zhang, Qinghai Miao, Min Huang 0009, Xiangyu Zhu 0001, Yingying Chen 0003, Zhen Lei 0001, Jinqiao Wang |
ICIP | 2 |
| 2017 | Deep embedding network for robust age estimationabstractEstimating age through a single facial image is a classic and challenging topic in computer vision. Since facial images of the same age vary considerably, while those from different ages may look very similar. To address these problems, we propose an end-to-end deep embedding neural network for robust age estimation. Specifically, we jointly use classification loss and triplet-based ranking loss to train a deep embedding network, which maps the input facial images into an embedding metric space where features of the same age are compact and those from different ages are pushed away. Thus the deep embedding network can learn more discriminative features and improves the performance for age estimation. Additionally, to accelerate the convergence of the network, we adopt an online hard negative mining strategy during the triplet loss computation. Experimental results on public datasets MORPH II and FG-NET show the superiority of our approach compared to the state-of-the-art. Yating He, Min Huang 0009, Qinghai Miao, Haiyun Guo, Jinqiao Wang |
ICIP | 3 |
| 2014 | Parallel Public Transportation System and Its Application in Evaluating Evacuation Plans for Large-Scale ActivitiesabstractThis paper proposes a method based on the Artificial societies, computational experiments, and Parallel execution (ACP) approach to build parallel public transportation systems (PPTSs). The framework and components of a PPTS are analyzed, and some details for building the PPTS are discussed. One prototype based on intelligent traffic clouds is established. One specific PPTS is developed for the Guangzhou 2010 Asian Games in the case study, and its effectiveness is verified through the evaluation of two evacuation plans for the Asian Games. Fenghua Zhu, Songhang Chen, Zhi-Hong Mao, Qinghai Miao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2011 | A Game-Engine-Based Platform for Modeling and Computing Artificial Transportation SystemsabstractA game-engine-based modeling and computing platform for artificial transportation systems (ATSs) is introduced. As an important feature, the artificial-population module (APM) is described in both its macroscopic and microcosmic aspects. In this module, each person is designed similarly to the actors in games. The traffic-simulation module (TSM) is another important module, which takes advantage of Delta3D to construct a 3-D simulation environment. All mobile actors are also managed by this module with the help of the dynamic-actor-layer (DAL) mechanism that is offered by Delta3D. The platform is designed as agent-oriented, modularized, and distributed. Both modules, together with components that are responsible for message processing, rules, network, and interactions, are organized by the game manager (GM) in a flexible architecture. With the help of the network component, the platform can be constructed to implement a distributed simulation. Finally, four experiments are introduced to show functions and features of the platform. Qinghai Miao, Fenghua Zhu, Changjian Cheng, Xiaogang Qiu |
IEEE Trans. Intell. Transp. Syst. | 1 |