VLDB 2026 Research / reviewers in the wild / expert
Weicong Liang
dblp:330/4850
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2027
0000-0003-0496-4774ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Efficient and distributed learning · 24% Image recognition and object detection · 22% Segmentation and scene understanding · 17% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 88% Image and video coding · 12% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
dense prediction |
1.3 | 2 | 2024 | Expediting Large-Scale Vision Transformer for Dense Prediction Without Fine-Tuning · IEEE Trans. Pattern Anal. Mach. Intell. 2024 Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning · NeurIPS 2022 |
Computer vision › Image recognition and object detection
object detection |
1.1 | 3 | 2024 | Rank-DETR for High Quality Object Detection · NeurIPS 2023 Expediting Large-Scale Vision Transformer for Dense Prediction Without Fine-Tuning · IEEE Trans. Pattern Anal. Mach. Intell. 2024 Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning · NeurIPS 2022 |
Machine learning › Efficient and distributed learning › inference acceleration
vision transformer acceleration |
0.8 | 1 | 2024 | Expediting Large-Scale Vision Transformer for Dense Prediction Without Fine-Tuning · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Visual content generation and editing › image generation
text-to-image generation |
0.8 | 1 | 2024 | Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering · ECCV (75) 2024 |
Computer vision › Image recognition and object detection › object detection
detection transformer |
0.7 | 1 | 2023 | Rank-DETR for High Quality Object Detection · NeurIPS 2023 |
Machine learning › Generative modeling
diffusion model |
0.7 | 1 | 2023 | GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023 |
Machine learning › Learning theory
ranking |
0.7 | 1 | 2023 | Rank-DETR for High Quality Object Detection · NeurIPS 2023 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.7 | 1 | 2023 | GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023 |
Visual content generation and editing
visual text generation |
0.7 | 1 | 2023 | GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023 |
Machine learning › Efficient and distributed learning
model acceleration |
0.6 | 1 | 2022 | Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning · NeurIPS 2022 |
Machine learning › Efficient and distributed learning
token reduction |
0.6 | 1 | 2022 | Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning · NeurIPS 2022 |
Computer vision › 3D vision
depth estimation |
0.2 | 1 | 2024 | Expediting Large-Scale Vision Transformer for Dense Prediction Without Fine-Tuning · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Image and video coding
quality assessment |
0.2 | 1 | 2023 | GlyphControl: Glyph Conditional Control for Visual Text Generation · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
text encoder fine-tuning · 1.5contrastive learning · 1.5token reconstruction · 1.3token clustering · 1.3glyph conditioning · 1.3diffusion model · 1.3vision transformer · 0.8ranking loss · 0.7matching cost · 0.7detection transformer · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | BTCMOEA: A bidirectional knowledge transfer driven algorithm for constrained multiobjective optimization
Lun Zhou, Weicong Liang, Xinglin Chen, Huayan Pu, Jun Luo 0006 |
Expert Syst. Appl. | 2 |
| 2025 | Improving Nested Named Entity Recognition with Collaborative ModelsabstractNamed Entity Recognition involves identifying specific spans with distinct meanings within text. Despite achieving state-of-the-art performance in NER, traditional small fine-tuned models (FTMs) based on non-generative pre-trained language models still suffer from insensitivity to long-tail entities and poor classification capabilities. In contrast, generative large pretrained language models (LLMs) can naturally address these issues. However, LLMs still suffer from significant shortcomings in task adaptation. Recognizing the complementarities between these models, we propose a novel collaborative framework, comprising the ICL-assisted recognition strategy and the task-specific output fusion strategy, to identify entities with diverse structures. The ICL-assisted recognition strategy focuses on designing high quality prompt words to generate the desired entities. while the task-specific output fusion strategy aims to filter out incorrect spans or categories from the fusion of predictions from LLMs and FTMs. We systematically conduct extensive experiments on publicly available datasets for nested NER tasks. The experimental results show that our framework outperforms LLMs and FTMs, respectively, and significantly surpasses supervised state-of-the-art methods. Zhiguo Fan, Weicong Liang |
BIBM | 4 |
| 2024 | Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering
Weicong Liang, Chong Luo 0001, Ji Li 0006, Gao Huang 0001, Yuhui Yuan |
ECCV (75) | 2 |
| 2024 | Expediting Large-Scale Vision Transformer for Dense Prediction Without Fine-TuningabstractIn a wide range of dense prediction tasks, large-scale Vision Transformers have achieved state-of-the-art performance while requiring expensive computation. In contrast to most existing approaches accelerating Vision Transformers for image classification, we focus on accelerating Vision Transformers for dense prediction without any fine-tuning. We present two non-parametric operators specialized for dense prediction tasks, a token clustering layer to decrease the number of tokens for expediting and a token reconstruction layer to increase the number of tokens for recovering high-resolution. To accomplish this, the following steps are taken: i) token clustering layer is employed to cluster the neighboring tokens and yield low-resolution representations with spatial structures; ii) the following transformer layers are performed only to these clustered low-resolution tokens; and iii) reconstruction of high-resolution representations from refined low-resolution representations is accomplished using token reconstruction layer. The proposed approach shows promising results consistently on 6 dense prediction tasks, including object detection, semantic segmentation, panoptic segmentation, instance segmentation, depth estimation, and video instance segmentation. Additionally, we validate the effectiveness of the proposed approach on the very recent state-of-the-art open-vocabulary recognition methods. Furthermore, a number of recent representative approaches are benchmarked and compared on dense prediction tasks. Yuhui Yuan, Weicong Liang, Henghui Ding, Chao Zhang 0001, Han Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Rank-DETR for High Quality Object DetectionabstractModern detection transformers (DETRs) use a set of object queries to predict a list of bounding boxes, sort them by their classification confidence scores, and select the top-ranked predictions as the final detection results for the given input image. A highly performant object detector requires accurate ranking for the bounding box predictions. For DETR-based detectors, the top-ranked bounding boxes suffer from less accurate localization quality due to the misalignment between classification scores and localization accuracy, thus impeding the construction of high-quality detectors. In this work, we introduce a simple and highly performant DETR-based object detector by proposing a series of rank-oriented designs, combinedly called Rank-DETR. Our key contributions include: (i) a rank-oriented architecture design that can prompt positive predictions and suppress the negative ones to ensure lower false positive rates, as well as (ii) a rank-oriented loss function and matching cost design that prioritizes predictions of more accurate localization accuracy during ranking to boost the AP under high IoU thresholds. We apply our method to improve the recent SOTA methods (e.g., H-DETR and DINO-DETR) and report strong COCO object detection results when using different backbones such as ResNet-$50$, Swin-T, and Swin-L, demonstrating the effectiveness of our approach. Code is available at \url{https://github.com/LeapLabTHU/Rank-DETR}. Yifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan, Yukang Yang, Chao Zhang 0001, Han Hu 0001, Gao Huang 0001 |
NeurIPS | 2 |
| 2023 | GlyphControl: Glyph Conditional Control for Visual Text GenerationabstractRecently, there has been an increasing interest in developing diffusion-based text-to-image generative models capable of generating coherent and well-formed visual text. In this paper, we propose a novel and efficient approach called GlyphControl to address this task. Unlike existing methods that rely on character-aware text encoders like ByT5 and require retraining of text-to-image models, our approach leverages additional glyph conditional information to enhance the performance of the off-the-shelf Stable-Diffusion model in generating accurate visual text. By incorporating glyph instructions, users can customize the content, location, and size of the generated text according to their specific requirements. To facilitate further research in visual text generation, we construct a training benchmark dataset called LAION-Glyph. We evaluate the effectiveness of our approach by measuring OCR-based metrics, CLIP score, and FID of the generated visual text. Our empirical evaluations demonstrate that GlyphControl outperforms the recent DeepFloyd IF approach in terms of OCR accuracy, CLIP score, and FID, highlighting the efficacy of our method. Yukang Yang, Dongnan Gui, Yuhui Yuan, Weicong Liang, Haisong Ding, Han Hu 0001, Kai Chen 0001 |
NeurIPS | 4 |
| 2022 | Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuningabstractVision transformers have recently achieved competitive results across various vision tasks but still suffer from heavy computation costs when processing a large number of tokens. Many advanced approaches have been developed to reduce the total number of tokens in the large-scale vision transformers, especially for image classification tasks. Typically, they select a small group of essential tokens according to their relevance with the [\texttt{class}] token, then fine-tune the weights of the vision transformer. Such fine-tuning is less practical for dense prediction due to the much heavier computation and GPU memory cost than image classification.In this paper, we focus on a more challenging problem, \ie, accelerating large-scale vision transformers for dense prediction without any additional re-training or fine-tuning. In response to the fact that high-resolution representations are necessary for dense prediction, we present two non-parametric operators, a \emph{token clustering layer} to decrease the number of tokens and a \emph{token reconstruction layer} to increase the number of tokens. The following steps are performed to achieve this: (i) we use the token clustering layer to cluster the neighboring tokens together, resulting in low-resolution representations that maintain the spatial structures; (ii) we apply the following transformer layers only to these low-resolution representations or clustered tokens; and (iii) we use the token reconstruction layer to re-create the high-resolution representations from the refined low-resolution representations. The results obtained by our method are promising on five dense prediction tasks including object detection, semantic segmentation, panoptic segmentation, instance segmentation, and depth estimation. Accordingly, our method accelerates $40\%\uparrow$ FPS and saves $30\%\downarrow$ GFLOPs of ``Segmenter+ViT-L/$16$'' while maintaining $99.5\%$ of the performance on ADE$20$K without fine-tuning the official weights. Weicong Liang, Yuhui Yuan, Henghui Ding, Xiao Luo 0001, Weihong Lin, Ding Jia, Zheng Zhang 0022, Chao Zhang 0001, Han Hu 0001 |
NeurIPS | 1 |