Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Weilin Zhuang

dblp:344/1427 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2024
0009-0001-6940-8384ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Visual content generation and editing · 67% Geometric modeling and processing · 33%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Artificial intelligence
2 papers
Vision and language · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
cross-modal alignment
0.712023
Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval · ACM Multimedia 2023
Information retrieval
cross-modal retrieval
0.712023
Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval · ACM Multimedia 2023
Information retrieval › cross-modal retrieval
text-based person retrieval
0.712023
Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval · ACM Multimedia 2023
Visual content generation and editing
3d content creation
0.712023
X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual Guidance · ICCV 2023
Geometric modeling and processing
mesh processing
0.712023
X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual Guidance · ICCV 2023
Visual content generation and editing › style transfer › 3d content stylization
text-driven 3d stylization
0.712023
X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual Guidance · ICCV 2023
Computer vision › Vision and language › vision-language generation
text-guided generation
0.212023
X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual Guidance · ICCV 2023

Methods — techniques the papers use, named apart from their topics

bi-directional embedding alignment · 1.3attention module · 1.3CLIP loss · 1.3
YearPublicationVenuePosition
2024 Towards Omni-supervised Referring Expression Segmentation
abstract
Referring Expression Segmentation (RES) is a challenging task in computer vision that involves segmenting image instances using textual descriptions. Conventional approaches suffer from the high cost of acquiring segmentation labels. To overcome this, we propose a novel learning task called Omni-supervised Referring Expression Segmentation (Omni-RES) which leverages unlabeled, fully labeled, and weakly labeled data, such as referring points or bounding boxes, for efficient RES training. Our approach is based on a teacher-student learning framework, where weak labels guide the selection and refinement of high-quality pseudo-masks for training, rather than serving as direct supervision signals. We tested Omni-RES on various state-of-the-art RES models and datasets, demonstrating its superiority over fully-supervised and semi-supervised methods. Remarkably, with just 10% fully labeled data, Omni-RES can match the performance of 100% supervised training. Additionally, it enables using large-scale vision-language datasets like Visual Genome for cost-effective RES training, setting a new state-of-the-art performance in RES, such as 80.66 on RefCOCO. Our code is released at: https://github.com/nineblu/omni-res
Minglang Huang, Yiyi Zhou, Gen Luo, Guannan Jiang, Weilin Zhuang, Xiaoshuai Sun
ICME5
2023 X-Mesh: Towards Fast and Accurate Text-driven 3D Stylization via Dynamic Textual Guidance
abstract
Text-driven 3D stylization is a complex and crucial task in the fields of computer vision (CV) and computer graphics (CG), aimed at transforming a bare mesh to fit a tar-get text. Prior methods adopt text-independent multilayer perceptrons (MLPs) to predict the attributes of the target mesh with the supervision of CLIP loss. However, such text-independent architecture lacks textual guidance during predicting attributes, thus leading to unsatisfactory stylization and slow convergence. To address these limitations, we present X-Mesh, an innovative text-driven 3D stylization framework that incorporates a novel Text-guided Dynamic Attention Module (TDAM). The TDAM dynamically integrates the guidance of the target text by utilizing text-relevant spatial and channel-wise attentions during vertex feature extraction, resulting in more accurate attribute prediction and faster convergence speed. Furthermore, existing works lack standard benchmarks and automated metrics for evaluation, often relying on subjective and non-reproducible user studies to assess the quality of stylized 3D assets. To overcome this limitation, we introduce a new standard text-mesh benchmark, namely MIT-30, and two automated metrics, which will enable future research to achieve fair and objective comparisons. Our extensive qualitative and quantitative experiments demonstrate that X-Mesh outperforms previous state-of-the-art methods. Our codes and results are available at our project webpage: https://xmu-xiaoma666.github.io/Projects/X-Mesh/
Haowei Wang 0001, Guannan Jiang, Xiaoshuai Sun, Weilin Zhuang, Jiayi Ji, Rongrong Ji
ICCV6
2023 Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval
abstract
Text-based person retrieval (TPR) is a challenging task that involves retrieving a specific individual based on a textual description. Despite considerable efforts to bridge the gap between vision and language, the significant differences between these modalities continue to pose a challenge. Previous methods have attempted to align text and image samples in a modal-shared space, but they face uncertainties in optimization directions due to the movable features of both modalities and the failure to account for one-to-many relationships of image-text pairs in TPR datasets. To address this issue, we propose an effective bi-directional one-to-many embedding paradigm that offers a clear optimization direction for each sample, thus mitigating the optimization problem. Additionally, this embedding scheme generates multiple features for each sample without introducing trainable parameters, making it easier to align with several positive samples. Based on this paradigm, we propose a novel Bi-directional one-to-many Embedding Alignment (Beat) model to address the TPR task. Our experimental results demonstrate that the proposed Beat model achieves state-of-the-art performance on three popular TPR datasets, including CUHK-PEDES (65.61 R@1), ICFG-PEDES (58.25 R@1), and RSTPReID (48.10 R@1). Furthermore, additional experiments on MS-COCO, CUB, and Flowers datasets further demonstrate the potential of Beat to be applied to other image-text retrieval tasks.
Xiaoshuai Sun, Jiayi Ji, Guannan Jiang, Weilin Zhuang, Rongrong Ji
ACM Multimedia5