Weicong Liang

dblp:330/4850 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2027
0000-0003-0496-4774ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 24% Image recognition and object detection · 22% Segmentation and scene understanding · 17%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 88% Image and video coding · 12%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
dense prediction
1.322024
Expediting Large-Scale Vision Transformer for Dense Prediction Without Fine-Tuning · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning · NeurIPS 2022
Computer vision › Image recognition and object detection
object detection
1.132024
Rank-DETR for High Quality Object Detection · NeurIPS 2023
Expediting Large-Scale Vision Transformer for Dense Prediction Without Fine-Tuning · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning · NeurIPS 2022
Machine learning › Efficient and distributed learning › inference acceleration
vision transformer acceleration
0.812024
Expediting Large-Scale Vision Transformer for Dense Prediction Without Fine-Tuning · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Visual content generation and editing › image generation
text-to-image generation
0.812024
Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering · ECCV (75) 2024
Computer vision › Image recognition and object detection › object detection
detection transformer
0.712023
Rank-DETR for High Quality Object Detection · NeurIPS 2023
Machine learning › Generative modeling
diffusion model
0.712023
GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023
Machine learning › Learning theory
ranking
0.712023
Rank-DETR for High Quality Object Detection · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.712023
GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023
Visual content generation and editing
visual text generation
0.712023
GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023
Machine learning › Efficient and distributed learning
model acceleration
0.612022
Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning · NeurIPS 2022
Machine learning › Efficient and distributed learning
token reduction
0.612022
Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning · NeurIPS 2022
Computer vision › 3D vision
depth estimation
0.212024
Expediting Large-Scale Vision Transformer for Dense Prediction Without Fine-Tuning · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Image and video coding
quality assessment
0.212023
GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

text encoder fine-tuning · 1.5contrastive learning · 1.5token reconstruction · 1.3token clustering · 1.3glyph conditioning · 1.3diffusion model · 1.3vision transformer · 0.8ranking loss · 0.7matching cost · 0.7detection transformer · 0.7
YearPublicationVenuePosition
2027 BTCMOEA: A bidirectional knowledge transfer driven algorithm for constrained multiobjective optimization
Lun Zhou, Weicong Liang, Xinglin Chen, Huayan Pu, Jun Luo 0006
Expert Syst. Appl.2
2025 Improving Nested Named Entity Recognition with Collaborative Models
abstract
Named Entity Recognition involves identifying specific spans with distinct meanings within text. Despite achieving state-of-the-art performance in NER, traditional small fine-tuned models (FTMs) based on non-generative pre-trained language models still suffer from insensitivity to long-tail entities and poor classification capabilities. In contrast, generative large pretrained language models (LLMs) can naturally address these issues. However, LLMs still suffer from significant shortcomings in task adaptation. Recognizing the complementarities between these models, we propose a novel collaborative framework, comprising the ICL-assisted recognition strategy and the task-specific output fusion strategy, to identify entities with diverse structures. The ICL-assisted recognition strategy focuses on designing high quality prompt words to generate the desired entities. while the task-specific output fusion strategy aims to filter out incorrect spans or categories from the fusion of predictions from LLMs and FTMs. We systematically conduct extensive experiments on publicly available datasets for nested NER tasks. The experimental results show that our framework outperforms LLMs and FTMs, respectively, and significantly surpasses supervised state-of-the-art methods.
Zhiguo Fan, Weicong Liang
BIBM4
2024 Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering
Weicong Liang, Chong Luo 0001, Ji Li 0006, Gao Huang 0001, Yuhui Yuan
ECCV (75)2
2024 Expediting Large-Scale Vision Transformer for Dense Prediction Without Fine-Tuning
abstract
In a wide range of dense prediction tasks, large-scale Vision Transformers have achieved state-of-the-art performance while requiring expensive computation. In contrast to most existing approaches accelerating Vision Transformers for image classification, we focus on accelerating Vision Transformers for dense prediction without any fine-tuning. We present two non-parametric operators specialized for dense prediction tasks, a token clustering layer to decrease the number of tokens for expediting and a token reconstruction layer to increase the number of tokens for recovering high-resolution. To accomplish this, the following steps are taken: i) token clustering layer is employed to cluster the neighboring tokens and yield low-resolution representations with spatial structures; ii) the following transformer layers are performed only to these clustered low-resolution tokens; and iii) reconstruction of high-resolution representations from refined low-resolution representations is accomplished using token reconstruction layer. The proposed approach shows promising results consistently on 6 dense prediction tasks, including object detection, semantic segmentation, panoptic segmentation, instance segmentation, depth estimation, and video instance segmentation. Additionally, we validate the effectiveness of the proposed approach on the very recent state-of-the-art open-vocabulary recognition methods. Furthermore, a number of recent representative approaches are benchmarked and compared on dense prediction tasks.
Yuhui Yuan, Weicong Liang, Henghui Ding, Chao Zhang 0001, Han Hu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Rank-DETR for High Quality Object Detection
abstract
Modern detection transformers (DETRs) use a set of object queries to predict a list of bounding boxes, sort them by their classification confidence scores, and select the top-ranked predictions as the final detection results for the given input image. A highly performant object detector requires accurate ranking for the bounding box predictions. For DETR-based detectors, the top-ranked bounding boxes suffer from less accurate localization quality due to the misalignment between classification scores and localization accuracy, thus impeding the construction of high-quality detectors. In this work, we introduce a simple and highly performant DETR-based object detector by proposing a series of rank-oriented designs, combinedly called Rank-DETR. Our key contributions include: (i) a rank-oriented architecture design that can prompt positive predictions and suppress the negative ones to ensure lower false positive rates, as well as (ii) a rank-oriented loss function and matching cost design that prioritizes predictions of more accurate localization accuracy during ranking to boost the AP under high IoU thresholds. We apply our method to improve the recent SOTA methods (e.g., H-DETR and DINO-DETR) and report strong COCO object detection results when using different backbones such as ResNet-$50$, Swin-T, and Swin-L, demonstrating the effectiveness of our approach. Code is available at \url{https://github.com/LeapLabTHU/Rank-DETR}.
Yifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan, Yukang Yang, Chao Zhang 0001, Han Hu 0001, Gao Huang 0001
NeurIPS2
2023 GlyphControl: Glyph Conditional Control for Visual Text Generation
abstract
Recently, there has been an increasing interest in developing diffusion-based text-to-image generative models capable of generating coherent and well-formed visual text. In this paper, we propose a novel and efficient approach called GlyphControl to address this task. Unlike existing methods that rely on character-aware text encoders like ByT5 and require retraining of text-to-image models, our approach leverages additional glyph conditional information to enhance the performance of the off-the-shelf Stable-Diffusion model in generating accurate visual text. By incorporating glyph instructions, users can customize the content, location, and size of the generated text according to their specific requirements. To facilitate further research in visual text generation, we construct a training benchmark dataset called LAION-Glyph. We evaluate the effectiveness of our approach by measuring OCR-based metrics, CLIP score, and FID of the generated visual text. Our empirical evaluations demonstrate that GlyphControl outperforms the recent DeepFloyd IF approach in terms of OCR accuracy, CLIP score, and FID, highlighting the efficacy of our method.
Yukang Yang, Dongnan Gui, Yuhui Yuan, Weicong Liang, Haisong Ding, Han Hu 0001, Kai Chen 0001
NeurIPS4
2022 Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning
abstract
Vision transformers have recently achieved competitive results across various vision tasks but still suffer from heavy computation costs when processing a large number of tokens. Many advanced approaches have been developed to reduce the total number of tokens in the large-scale vision transformers, especially for image classification tasks. Typically, they select a small group of essential tokens according to their relevance with the [\texttt{class}] token, then fine-tune the weights of the vision transformer. Such fine-tuning is less practical for dense prediction due to the much heavier computation and GPU memory cost than image classification.In this paper, we focus on a more challenging problem, \ie, accelerating large-scale vision transformers for dense prediction without any additional re-training or fine-tuning. In response to the fact that high-resolution representations are necessary for dense prediction, we present two non-parametric operators, a \emph{token clustering layer} to decrease the number of tokens and a \emph{token reconstruction layer} to increase the number of tokens. The following steps are performed to achieve this: (i) we use the token clustering layer to cluster the neighboring tokens together, resulting in low-resolution representations that maintain the spatial structures; (ii) we apply the following transformer layers only to these low-resolution representations or clustered tokens; and (iii) we use the token reconstruction layer to re-create the high-resolution representations from the refined low-resolution representations. The results obtained by our method are promising on five dense prediction tasks including object detection, semantic segmentation, panoptic segmentation, instance segmentation, and depth estimation. Accordingly, our method accelerates $40\%\uparrow$ FPS and saves $30\%\downarrow$ GFLOPs of ``Segmenter+ViT-L/$16$'' while maintaining $99.5\%$ of the performance on ADE$20$K without fine-tuning the official weights.
Weicong Liang, Yuhui Yuan, Henghui Ding, Xiao Luo 0001, Weihong Lin, Ding Jia, Zheng Zhang 0022, Chao Zhang 0001, Han Hu 0001
NeurIPS1