VLDB 2026 Research / reviewers in the wild / expert
Mang Tik Chiu
dblp:239/8558
· DBLP profile ↗
5ranked-venue papers
3as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Segmentation and scene understanding · 61% Generative modeling · 15% Vision and language · 14% | |
| Computer graphics and multimedia
3 papers |
Image and video processing · 75% Visual content generation and editing · 25% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
image segmentation |
1.0 | 1 | 2026 | From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual Segmentation · ACL (1) 2026 |
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation |
1.0 | 1 | 2026 | From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual Segmentation · ACL (1) 2026 |
Computer vision › Vision and language › cross-modal alignment
visual-semantic embedding |
0.8 | 1 | 2024 | Brush2Prompt: Contextual Prompt Generator for Object Inpainting · CVPR 2024 |
Image and video processing › image restoration › image inpainting
object inpainting |
0.8 | 1 | 2024 | Brush2Prompt: Contextual Prompt Generator for Object Inpainting · CVPR 2024 |
Image and video processing › image restoration
image inpainting |
0.7 | 1 | 2023 | Automatic High Resolution Wire Segmentation and Removal · CVPR 2023 |
Image and video processing
image segmentation |
0.7 | 1 | 2023 | Automatic High Resolution Wire Segmentation and Removal · CVPR 2023 |
Computer vision › Segmentation and scene understanding › semantic segmentation › remote sensing image segmentation
aerial image segmentation |
0.4 | 1 | 2020 | Agriculture-Vision: A Large Aerial Image Database for Agricultural Pattern Analysis · CVPR 2020 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.4 | 1 | 2020 | Agriculture-Vision: A Large Aerial Image Database for Agricultural Pattern Analysis · CVPR 2020 |
Machine learning › Generative modeling › generative adversarial network
image-to-image translation |
0.4 | 1 | 2019 | LADN: Local Adversarial Disentangling Network for Facial Makeup and De-Makeup · ICCV 2019 |
Natural language and speech › Language models and text generation › controllable text generation › text style transfer
unsupervised text style transfer |
0.4 | 1 | 2019 | LADN: Local Adversarial Disentangling Network for Facial Makeup and De-Makeup · ICCV 2019 |
Visual content generation and editing
face editing |
0.4 | 1 | 2019 | LADN: Local Adversarial Disentangling Network for Facial Makeup and De-Makeup · ICCV 2019 |
Visual content generation and editing › style transfer
makeup transfer |
0.4 | 1 | 2019 | LADN: Local Adversarial Disentangling Network for Facial Makeup and De-Makeup · ICCV 2019 |
Computer vision › Segmentation and scene understanding
video segmentation |
0.3 | 1 | 2026 | From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual Segmentation · ACL (1) 2026 |
Machine learning › Generative modeling
diffusion model |
0.2 | 1 | 2024 | Brush2Prompt: Contextual Prompt Generator for Object Inpainting · CVPR 2024 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.2 | 1 | 2024 | Brush2Prompt: Contextual Prompt Generator for Object Inpainting · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
knowledge distillation · 1.5concept diffusion · 1.5CLIP embedding · 1.5large multimodal model · 1.0large language model · 1.0generative adversarial network · 0.8asymmetric loss · 0.8adversarial disentangling · 0.8two-stage segmentation · 0.7tile-based inpainting · 0.7semantic segmentation model · 0.4deep learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual SegmentationabstractVisual segmentation, the task of segmenting an image into semantically meaningful regions, is a cornerstone in machine learning and has widespread applications in industry. Nevertheless, visual segmentation with instruction has been a challenging task for many years. This largely stems from the cross-modal discrepancy between language and image domains, resulting in difficulty in relating the instruction semantics and the pixel-level predictions. In recent years, the remarkable reasoning capabilities of Large Language Models (LLMs) and Large Multimodal Models (LMMs) have spurred a new wave of research aiming to bridge the disparity between natural language instructions and pixel-level understanding. This survey offers the first comprehensive overview of the rapidly evolving field of LLM-driven visual segmentation. We categorize existing approaches based on their core objectives and methodologies, including reasoning-based segmentation, open-vocabulary segmentation, grounding techniques connecting language to pixels, and extensions to video domains. We review recent seminal works in LLM-based visual segmentation, analyzing their architectural innovations, training strategies, and benchmark performance. Furthermore, we discuss the common datasets, evaluation metrics, and identify key challenges and promising future directions at the intersection of language and visual segmentation. We hope this survey serves as a valuable resource for researchers and practitioners seeking to understand the current landscape and future directions of leveraging LLMs for sophisticated visual segmentation tasks and applications. The resource summary is available at https://github.com/wyzjack/Awesome-LLM-Visual-Segmentation. Yizhou Wang 0006, Mang Tik Chiu, Lingzhi Zhang, Xuan Shen, Sohrab Amirghodsi, Yun Fu 0001 |
ACL (1) | 2 |
| 2024 | Brush2Prompt: Contextual Prompt Generator for Object InpaintingabstractObject inpainting is a task that involves adding objects to real images and seamlessly compositing them. With the recent commercialization of products like Stable Diffusion and Generative Fill, inserting objects into images by using prompts has achieved impressive visual results. In this paper, we propose a prompt suggestion model to simplify the process of prompt input. When the user provides an image and a mask, our model predicts suitable prompts based on the partial contextual information in the masked image, and the shape and location of the mask. Specifically, we introduce a concept-diffusion in the CLIP space that predicts CLIP-text embeddings from a masked image. These diffused embeddings can be directly injected into open-source in-painting models like Stable Diffusion and its variants. Alternatively, they can be decoded into natural language for use in other publicly available applications such as Generative Fill. Our prompt suggestion model demonstrates a balanced accuracy and diversity, showing its capability to be both contextually aware and creatively adaptive. Mang Tik Chiu, Yuqian Zhou, Lingzhi Zhang, Zhe Lin 0001, Connelly Barnes, Sohrab Amirghodsi, Eli Shechtman, Humphrey Shi |
CVPR | 1 |
| 2023 | Automatic High Resolution Wire Segmentation and RemovalabstractWires and powerlines are common visual distractions that often undermine the aesthetics of photographs. The manual process of precisely segmenting and removing them is extremely tedious and may take up hours, especially on high-resolution photos where wires may span the entire space. In this paper, we present an automatic wire clean-up system that eases the process of wire segmentation and removal/inpainting to within a few seconds. We observe several unique challenges: wires are thin, lengthy, and sparse. These are rare properties of subjects that common segmentation tasks cannot handle, especially in high-resolution images. We thus propose a two-stage method that leverages both global and local contexts to accurately segment wires in high-resolution images efficiently, and a tile-based inpainting strategy to remove the wires given our predicted segmentation masks. We also introduce the first wire segmentation benchmark dataset, WireSegHR. Finally, we demonstrate quantitatively and qualitatively that our wire clean-up system enables fully automated wire removal with great generalization to various wire appearances. Mang Tik Chiu, Xuaner Cecilia Zhang, Zijun Wei, Yuqian Zhou, Eli Shechtman, Connelly Barnes, Zhe Lin 0001, Florian Kainz, Sohrab Amirghodsi, Humphrey Shi |
CVPR | 1 |
| 2020 | Agriculture-Vision: A Large Aerial Image Database for Agricultural Pattern AnalysisabstractThe success of deep learning in visual recognition tasks has driven advancements in multiple fields of research. Particularly, increasing attention has been drawn towards its application in agriculture. Nevertheless, while visual pattern recognition on farmlands carries enormous economic values, little progress has been made to merge computer vision and crop sciences due to the lack of suitable agricultural image datasets. Meanwhile, problems in agriculture also pose new challenges in computer vision. For example, semantic segmentation of aerial farmland images requires inference over extremely large-size images with extreme annotation sparsity. These challenges are not present in most of the common object datasets, and we show that they are more challenging than many other aerial image datasets. To encourage research in computer vision for agriculture, we present Agriculture-Vision: a large-scale aerial farmland image dataset for semantic segmentation of agricultural patterns. We collected 94,986 high-quality aerial images from 3,432 farmlands across the US, where each image consists of RGB and Near-infrared (NIR) channels with resolution as high as 10 cm per pixel. We annotate nine types of field anomaly patterns that are most important to farmers. As a pilot study of aerial agricultural semantic segmentation, we perform comprehensive experiments using popular semantic segmentation models; we also propose an effective model designed for aerial agricultural pattern recognition. Our experiments demonstrate several challenges Agriculture-Vision poses to both the computer vision and agriculture communities. Future versions of this dataset will include even more aerial images, anomaly patterns and image channels. Mang Tik Chiu, Xingqian Xu, Yunchao Wei, Alexander G. Schwing, Robert Brunner, Hrant Khachatrian, Hovnatan Karapetyan, Ivan Dozier, Greg Rose, Adrian Tudor, Naira Hovakimyan, Thomas S. Huang, Humphrey Shi |
CVPR | 1 |
| 2019 | LADN: Local Adversarial Disentangling Network for Facial Makeup and De-MakeupabstractWe propose a local adversarial disentangling network (LADN) for facial makeup and de-makeup. Central to our method are multiple and overlapping local adversarial discriminators in a content-style disentangling network for achieving local detail transfer between facial images, with the use of asymmetric loss functions for dramatic makeup styles with high-frequency details. Existing techniques do not demonstrate or fail to transfer high-frequency details in a global adversarial setting, or train a single local discriminator only to ensure image structure consistency and thus work only for relatively simple styles. Unlike others, our proposed local adversarial discriminators can distinguish whether the generated local image details are consistent with the corresponding regions in the given reference image in cross-image style transfer in an unsupervised setting. Incorporating these technical contributions, we achieve not only state-of-the-art results on conventional styles but also novel results involving complex and dramatic styles with high-frequency details covering large areas across multiple facial features. A carefully designed dataset of unpaired before and after makeup images is released at https://georgegu1997.github.io/LADN-project-page. Qiao Gu, Guanzhi Wang, Mang Tik Chiu, Yu-Wing Tai, Chi-Keung Tang |
ICCV | 3 |