VLDB 2026 Research / reviewers in the wild / expert
Weilin Zhuang
dblp:344/1427
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2024
0009-0001-6940-8384ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 67% Geometric modeling and processing · 33% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Artificial intelligence
2 papers |
Vision and language · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
cross-modal alignment |
0.7 | 1 | 2023 | Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval · ACM Multimedia 2023 |
Information retrieval
cross-modal retrieval |
0.7 | 1 | 2023 | Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval · ACM Multimedia 2023 |
Information retrieval › cross-modal retrieval
text-based person retrieval |
0.7 | 1 | 2023 | Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval · ACM Multimedia 2023 |
Visual content generation and editing
3d content creation |
0.7 | 1 | 2023 | X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual Guidance · ICCV 2023 |
Geometric modeling and processing
mesh processing |
0.7 | 1 | 2023 | X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual Guidance · ICCV 2023 |
Visual content generation and editing › style transfer › 3d content stylization
text-driven 3d stylization |
0.7 | 1 | 2023 | X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual Guidance · ICCV 2023 |
Computer vision › Vision and language › vision-language generation
text-guided generation |
0.2 | 1 | 2023 | X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual Guidance · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
bi-directional embedding alignment · 1.3attention module · 1.3CLIP loss · 1.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Omni-supervised Referring Expression SegmentationabstractReferring Expression Segmentation (RES) is a challenging task in computer vision that involves segmenting image instances using textual descriptions. Conventional approaches suffer from the high cost of acquiring segmentation labels. To overcome this, we propose a novel learning task called Omni-supervised Referring Expression Segmentation (Omni-RES) which leverages unlabeled, fully labeled, and weakly labeled data, such as referring points or bounding boxes, for efficient RES training. Our approach is based on a teacher-student learning framework, where weak labels guide the selection and refinement of high-quality pseudo-masks for training, rather than serving as direct supervision signals. We tested Omni-RES on various state-of-the-art RES models and datasets, demonstrating its superiority over fully-supervised and semi-supervised methods. Remarkably, with just 10% fully labeled data, Omni-RES can match the performance of 100% supervised training. Additionally, it enables using large-scale vision-language datasets like Visual Genome for cost-effective RES training, setting a new state-of-the-art performance in RES, such as 80.66 on RefCOCO. Our code is released at: https://github.com/nineblu/omni-res Minglang Huang, Yiyi Zhou, Gen Luo, Guannan Jiang, Weilin Zhuang, Xiaoshuai Sun |
ICME | 5 |
| 2023 | X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual GuidanceabstractText-driven 3D stylization is a complex and crucial task in the fields of computer vision (CV) and computer graphics (CG), aimed at transforming a bare mesh to fit a tar-get text. Prior methods adopt text-independent multilayer perceptrons (MLPs) to predict the attributes of the target mesh with the supervision of CLIP loss. However, such text-independent architecture lacks textual guidance during predicting attributes, thus leading to unsatisfactory stylization and slow convergence. To address these limitations, we present X-Mesh, an innovative text-driven 3D stylization framework that incorporates a novel Text-guided Dynamic Attention Module (TDAM). The TDAM dynamically integrates the guidance of the target text by utilizing text-relevant spatial and channel-wise attentions during vertex feature extraction, resulting in more accurate attribute prediction and faster convergence speed. Furthermore, existing works lack standard benchmarks and automated metrics for evaluation, often relying on subjective and non-reproducible user studies to assess the quality of stylized 3D assets. To overcome this limitation, we introduce a new standard text-mesh benchmark, namely MIT-30, and two automated metrics, which will enable future research to achieve fair and objective comparisons. Our extensive qualitative and quantitative experiments demonstrate that X-Mesh outperforms previous state-of-the-art methods. Our codes and results are available at our project webpage: https://xmu-xiaoma666.github.io/Projects/X-Mesh/ Haowei Wang 0001, Guannan Jiang, Xiaoshuai Sun, Weilin Zhuang, Jiayi Ji, Rongrong Ji |
ICCV | 6 |
| 2023 | Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person RetrievalabstractText-based person retrieval (TPR) is a challenging task that involves retrieving a specific individual based on a textual description. Despite considerable efforts to bridge the gap between vision and language, the significant differences between these modalities continue to pose a challenge. Previous methods have attempted to align text and image samples in a modal-shared space, but they face uncertainties in optimization directions due to the movable features of both modalities and the failure to account for one-to-many relationships of image-text pairs in TPR datasets. To address this issue, we propose an effective bi-directional one-to-many embedding paradigm that offers a clear optimization direction for each sample, thus mitigating the optimization problem. Additionally, this embedding scheme generates multiple features for each sample without introducing trainable parameters, making it easier to align with several positive samples. Based on this paradigm, we propose a novel Bi-directional one-to-many Embedding Alignment (Beat) model to address the TPR task. Our experimental results demonstrate that the proposed Beat model achieves state-of-the-art performance on three popular TPR datasets, including CUHK-PEDES (65.61 R@1), ICFG-PEDES (58.25 R@1), and RSTPReID (48.10 R@1). Furthermore, additional experiments on MS-COCO, CUB, and Flowers datasets further demonstrate the potential of Beat to be applied to other image-text retrieval tasks. Xiaoshuai Sun, Jiayi Ji, Guannan Jiang, Weilin Zhuang, Rongrong Ji |
ACM Multimedia | 5 |