VLDB 2026 Research / reviewers in the wild / expert
Mingxiang Liao
dblp:332/5990
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Segmentation and scene understanding · 42% Vision and language · 34% Image recognition and object detection · 18% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 50% Multimedia systems and quality of experience · 50% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
instance segmentation |
1.5 | 2 | 2025 | Discriminatively Matched Part Tokens for Pointly Supervised Instance Segmentation · Int. J. Comput. Vis. 2025 AttentionShift: Iteratively Estimated Part-Based Attention Map for Pointly Supervised Instance Segmentation · CVPR 2023 |
Computer vision › Segmentation and scene understanding › instance segmentation
pointly supervised instance segmentation |
1.5 | 2 | 2025 | Discriminatively Matched Part Tokens for Pointly Supervised Instance Segmentation · Int. J. Comput. Vis. 2025 AttentionShift: Iteratively Estimated Part-Based Attention Map for Pointly Supervised Instance Segmentation · CVPR 2023 |
Computer vision › Vision and language
image captioning |
0.9 | 1 | 2025 | DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution · CVPR 2025 |
Computer vision › Vision and language › image captioning › grounded image captioning
region captioning |
0.9 | 1 | 2025 | DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution · CVPR 2025 |
Visual content generation and editing › video generation
text-to-video generation |
0.8 | 1 | 2024 | Evaluation of Text-to-Video Generation Models: A Dynamics Perspective · NeurIPS 2024 |
Multimedia systems and quality of experience
video quality assessment |
0.8 | 1 | 2024 | Evaluation of Text-to-Video Generation Models: A Dynamics Perspective · NeurIPS 2024 |
Computer vision › Image recognition and object detection › object detection
object proposal generation |
0.6 | 1 | 2022 | End-to-End Weakly Supervised Object Detection with Sparse Proposal Evolution · ECCV (9) 2022 |
Computer vision › Image recognition and object detection › object detection
weakly supervised object detection |
0.6 | 1 | 2022 | End-to-End Weakly Supervised Object Detection with Sparse Proposal Evolution · ECCV (9) 2022 |
Natural language and speech › Information extraction and text analysis › open vocabulary learning
open-vocabulary recognition |
0.3 | 1 | 2025 | DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution · CVPR 2025 |
Computer vision › Video understanding and tracking › temporal modeling
temporal dynamics |
0.2 | 1 | 2024 | Evaluation of Text-to-Video Generation Models: A Dynamics Perspective · NeurIPS 2024 |
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation |
0.2 | 1 | 2023 | AttentionShift: Iteratively Estimated Part-Based Attention Map for Pointly Supervised Instance Segmentation · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
human rating correlation · 1.5dynamics scoring · 1.5nested views · 0.9multimodal alignment · 0.9dynamic resolution · 0.9discriminatively matched part tokens · 0.9vision transformer · 0.7token querying · 0.7key-point shift · 0.7weakly supervised learning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DynRefer: Delving into Region-level Multimodal Tasks via Dynamic ResolutionabstractOne fundamental task of multimodal models is to translate referred image regions to human preferred language descriptions. Existing methods, however, ignore the resolution adaptability needs of different tasks, which hinders them to find out precise language descriptions. In this study, we propose a DynRefer approach, to pursue high-accuracy region-level referring through mimicking the resolution adaptability of human visual cognition. During training, DynRefer stochastically aligns language descriptions of multimodal tasks with images of multiple resolutions, which are constructed by nesting a set of random views around the referred region. During inference, DynRefer performs selectively multimodal referring by sampling proper region representations for tasks from the nested views based on image and task priors. This allows the visual information for referring to better match human preferences, thereby improving the representational adaptability of region-level multimodal models. Experiments show that DynRefer brings mutual improvement upon broad tasks including region-level captioning, open-vocabulary region recognition and attribute detection. Furthermore, DynRefer achieves state-of-the-art results on multiple region-level multimodal tasks using a single model. Code is available at https://github.com/callsys/DynRefer. Yuzhong Zhao, Feng Liu 0050, Mingxiang Liao, Chen Gong 0005, Qixiang Ye, Fang Wan 0001 |
CVPR | 4 |
| 2025 | Discriminatively Matched Part Tokens for Pointly Supervised Instance Segmentation
Zonghao Guo, Fang Wan 0001, Mingxiang Liao, Qixiang Ye |
Int. J. Comput. Vis. | 3 |
| 2025 | Hierarchical AttentionShift for Pointly Supervised Instance SegmentationabstractPointly supervised instance segmentation (PSIS) remains a challenging task when appearance variances across object parts cause semantic inconsistency. In this article, we propose a hierarchical AttentionShift approach, to solve the semantic inconsistency issue through exploiting the hierarchical nature of semantics and the flexibility of key-point representation. The estimation of hierarchical attention is defined upon key-point sets. The representative key points are iteratively estimated spatially and in the feature space to capture the fine-grained semantics and cover the full object extent. Hierarchical AttentionShift is performed at instance, part, and fine-grained levels, optimizing object semantics while promoting the conventional self-attention activation to hierarchical activation with local refinement. Experiments on PASCAL VOC 2012 Aug and MS-COCO 2017 benchmarks show that hierarchical AttentionShift improves the state-of-the-art (SOTA) method by 10.4% and 7.0% upon mean average precision (mAP)50, respectively. When applying hierarchical AttentionShift to the segment anything model (SAM), 9.4% AP improvement on the COCO test-dev is achieved. Hierarchical AttentionShift provides a fresh insight to regularize the self-attention mechanism for fine-grained vision tasks. The code is available at github.com/MingXiangL/AttentionShift. Mingxiang Liao, Fang Wan 0001, Zonghao Guo, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Evaluation of Text-to-Video Generation Models: A Dynamics PerspectiveabstractComprehensive and constructive evaluation protocols play an important role when developing sophisticated text-to-video (T2V) generation models. Existing evaluation protocols primarily focus on temporal consistency and content continuity, yet largely ignore dynamics of video content. Such dynamics is an essential dimension measuring the visual vividness and the honesty of video content to text prompts. In this study, we propose an effective evaluation protocol, termed DEVIL, which centers on the dynamics dimension to evaluate T2V generation models, as well as improving existing evaluation metrics. In practice, we define a set of dynamics scores corresponding to multiple temporal granularities, and a new benchmark of text prompts under multiple dynamics grades. Upon the text prompt benchmark, we assess the generation capacity of T2V models, characterized by metrics of dynamics ranges and T2V alignment. Moreover, we analyze the relevance of existing metrics to dynamics metrics, improving them from the perspective of dynamics. Experiments show that DEVIL evaluation metrics enjoy up to about 90\% consistency with human ratings, demonstrating the potential to advance T2V generation models. Mingxiang Liao, Hannan Lu, Qixiang Ye, Wangmeng Zuo, Fang Wan 0001, Tianyu Wang 0028, Yuzhong Zhao, Jingdong Wang 0001, Xinyu Zhang 0017 |
NeurIPS | 1 |
| 2023 | AttentionShift: Iteratively Estimated Part-Based Attention Map for Pointly Supervised Instance SegmentationabstractPointly supervised instance segmentation (PSIS) learns to segment objects using a single point within the object extent as supervision. Challenged by the non-negligible semantic variance between object parts, however, the single supervision point causes semantic bias and false segmentation. In this study, we propose an AttentionShift method, to solve the semantic bias issue by iteratively decomposing the instance attention map to parts and estimating fine-grained semantics of each part. AttentionShift consists of two modules plugged on the vision transformer backbone: (i) token querying for pointly supervised attention map generation, and (ii) key-point shift, which re-estimates part-based attention maps by key-point filtering in the feature space. These two steps are iteratively performed so that the part-based attention maps are optimized spatially as well as in the feature space to cover full object extent. Experiments on PASCAL VOC and MS COCO 2017 datasets show that AttentionShift respectively improves the state-of-the-art of by 7.7% and 4.8% under [email protected], setting a solid PSIS baseline using vision transformer. Mingxiang Liao, Zonghao Guo, Yuze Wang 0004, Bailan Feng, Fang Wan 0001 |
CVPR | 1 |
| 2022 | End-to-End Weakly Supervised Object Detection with Sparse Proposal Evolution
Mingxiang Liao, Fang Wan 0001, Zhenjun Han, Jialing Zou, Yuze Wang 0004, Bailan Feng, Qixiang Ye |
ECCV (9) | 1 |